read-only demo

Videos fnLBmfsI_Fg

Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft

index_state ready data_status ok

AI Engineer· published 2026-07-20· 0:05:44· en-US· indexed 2026-08-10 19:47

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:13, 1 of 1 keyframes kept
  2. Shot 1, 0:13 to 0:30, 1 of 1 keyframes kept
  3. Shot 2, 0:30 to 0:42, 1 of 1 keyframes kept
  4. Shot 3, 0:42 to 1:02, 1 of 1 keyframes kept
  5. Shot 4, 1:02 to 1:28, 1 of 1 keyframes kept
  6. Shot 5, 1:28 to 1:54, 0 of 1 keyframes kept
  7. Shot 6, 1:54 to 2:19, 0 of 1 keyframes kept
  8. Shot 7, 2:19 to 2:45, 0 of 1 keyframes kept
  9. Shot 8, 2:45 to 3:11, 0 of 1 keyframes kept
  10. Shot 9, 3:11 to 3:43, 1 of 1 keyframes kept
  11. Shot 10, 3:43 to 4:14, 0 of 1 keyframes kept
  12. Shot 11, 4:14 to 4:43, 1 of 1 keyframes kept
  13. Shot 12, 4:43 to 5:08, 1 of 1 keyframes kept
  14. Shot 13, 5:08 to 5:33, 0 of 1 keyframes kept
  15. Shot 14, 5:33 to 5:44, 1 of 1 keyframes kept

15 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
60
whisperx 60
chunks
11
from 60 cues
keyframes
9
kept of 15 captured
frames with text
9
81 lines read
chapters
0
from the source metadata
keyframe bytes
958.7 kB
word timings on 60 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 04:27 1m 26s
stt done 2026-08-10 04:28 9s
chunk done 2026-08-10 04:28 0s
text_embed done 2026-08-10 19:47 0s
keyframe done 2026-08-10 04:28 15s
ocr done 2026-08-10 04:29 3s
frame_embed done 2026-08-10 19:47 1s

Frames, and what the machine read

  • 0:07 #0 done11 line(s)

    shot 0·sharpness 858.7

    1. Ace AI ENGINEER0.99
    2. LIGHTNING TALK1.00
    3. Ornella Bahidika0.96
    4. Joel Allou0.99
    5. Your voice agent1.00
    6. doesn't need a0.99
    7. frontier model.1.00
    8. Joel Allou1.00
    9. Ornella Bahidika1.00
    10. ENGINEERING1.00
    11. PRODUCT1.00
  • 0:23 #1 done5 line(s)

    shot 1·sharpness 1204.3

    1. Ace•0.87
    2. A few seconds pause feels1.00
    3. broken.1.00
    4. // dead air · 3.0s0.95
    5. Voice has no spinner.0.98
  • 0:41 #2 done5 line(s)

    shot 2·sharpness 440.7

    1. Ace•0.85
    2. Ornella Bahidka0.98
    3. THE INSTINCT0.96
    4. bigger modelbetter0.99
    5. In voice, that's exactly backwards.1.00
  • 1:00 #3 done6 line(s)

    shot 3·sharpness 510.7

    1. Ace •0.88
    2. Joel Allou0.97
    3. 9501.00
    4. ms1.00
    5. TIME TO FIRST TOKEN0.95
    6. Our budget was never IQ. It's milliseconds.1.00
  • 1:17 #4 done16 line(s)

    shot 4·sharpness 1247.7

    1. Ace•0.90
    2. Take the hard jobs out of the model.1.00
    3. Control flow0.97
    4. What they know0.99
    5. What's next1.00
    6. what happens next0.98
    7. mastery tracking1.00
    8. planning1.00
    9. state machine1.00
    10. BKT + FSRS1.00
    11. prereq scheduler0.99
    12. DETERMINISTIC CODE1.00
    13. structured brief, every turn1.00
    14. The model1.00
    15. just talks·small + fast LLM1.00
    16. voice out · ≈150ms to first token0.99
  • 1:31 #5 skipped

    shot 5·duplicate of #4

  • 2:16 #6 skipped

    shot 6·duplicate of #4

  • 2:22 #7 skipped

    shot 7·duplicate of #4

  • 3:00 #8 skipped

    shot 8·duplicate of #4

  • 3:39 #9 done19 line(s)

    shot 9·sharpness 844.2

    1. Ace •0.95
    2. English Foundations: Parts of Speech, Sentence Structure, and Verb Tense1.00
    3. CH. ENGLISH FOUNDA..0.94
    4. □ END0.88
    5. Joel Allou0.99
    6. Hey Jordan, great to see you again.0.98
    7. 4 lssues::2113 liddem0.82
    8. sse-livekit-lesson,ts:3730.84
    9. Defaul levels0.91
    10. Today we're building your grammar0.99
    11. toolkit for the A C0.99
    12. Hey. What are we leaming today?0.97
    13. The Eight Parts of Speech0.99
    14. Perfect timing. We're diving into0.97
    15. the1.00
    16. cad 1to turn on code0.88
    17. It feels instant.0.99
    18. The smart part already happened. Before1.00
    19. the model opened its mouth.0.99
  • 3:52 #10 skipped

    shot 10·duplicate of #9

  • 4:31 #11 done7 line(s)

    shot 11·sharpness 918.3

    1. Ace1.00
    2. Small model + no scaffolding =0.99
    3. FAILURE 011.00
    4. FAILURE 021.00
    5. Drifts on long structure.1.00
    6. Needs strict rules to stay organized.1.00
    7. // you pay the scaffolding once, in code. Not on every turn.1.00
  • 4:58 #12 done4 line(s)

    shot 12·sharpness 957.8

    1. Ace·0.91
    2. Pick the fastest model your latency allows.0.99
    3. Spend the rest on scaffolding.0.99
    4. voice· realtime, high-volume· the model is the smallest part0.98
  • 5:23 #13 skipped

    shot 13·duplicate of #12

  • 5:42 #14 done8 line(s)

    shot 14·sharpness 365.2

    1. Ace •0.89
    2. Ornella Bahidika1.00
    3. Joel Allou0.99
    4. LET THE MODEL TALK.1.00
    5. Thanks.1.00
    6. Joel & Ornella1.00
    7. aceactprep.com1.00
    8. Engineering - Product0.96

Transcript

60 cues· 882 words· 4,658 chars

  1. 0:00 Hi, I'm Ornella, and that's Joelle, and we built Ace, a live AI voice tutor.
  2. 0:07 It's run on a small model on purpose, and I want to tell you more why that's not a compromise.
  3. 0:14 Quick guide check.
  4. 0:19 That silence, on a voice call, that's the difference between a tutor and a broken up.
  5. 0:25 When a voice agent pause for even a second, your brain says it's dead.
  6. 0:31 So when the answer feels a little off, every instant stay waits for the smartest, biggest model.
  7. 0:38 In voice, that instant is actually a backward.
  8. 0:44 Because our budget was never IQ, it's millisecond.
  9. 0:49 The AI model needs to start talking in about 950 milliseconds.
  10. 0:54 A frontier model that thinks for a full second has already lost the room, no matter how good the answer is.
  11. 1:04 So we made the model small and took the hardest jobs away from it.
  12. 1:10 It doesn't decide what happened in the lesson.
  13. 1:14 It doesn't track what the student knows.
  14. 1:17 It doesn't plan what's next.
  15. 1:19 We have a system in place to do that.
  16. 1:21 And it hands the model a summary every turn.
  17. 1:26 What's left for the model is one thing it's really good at, talking.
  18. 1:30 And that's the theory.
  19. 1:32 So I'll go ahead and show them what it actually feel like.
  20. 1:35 Yeah, if maybe I can add some color to what Arnella was mentioning.
  21. 1:40 So if you think about the models of today, especially the frontier model, let's take Cloud 4.7, which is from Anthropic.
  22. 1:49 The model is really good at reasoning.
  23. 1:51 You can give it a problem, in this case a lesson, and it can reason through it.
  24. 1:55 It can reason through what the student is asking, and it can come up with the answer.
  25. 2:00 But that is actually precisely the problem because the reasoning can take a couple of seconds.
  26. 2:07 And those seconds are really valuable when you are building voice applications.
  27. 2:12 So what we are doing is saying, hey, let's extract all of the thinking away from the model so that the model focuses on only what matters, which is speaking in our case.
  28. 2:25 So all of the thinking is extracted into a state machine.
  29. 2:29 So for ACE, we have thought about all the scenarios that are needed for a lesson.
  30. 2:34 We have built a state machine that is able to coordinate each step to the next.
  31. 2:39 And we also added intelligent layer on top to derive some of the mastery that a student might need for the lesson to be complete.
  32. 2:49 So everything when it comes to what happens next, when it comes to what needs to be displayed, when it comes to how to actually answer a question, it's all done outside of the model.
  33. 3:01 And we simply feed that output to the model to speak out.
  34. 3:07 and so let's go ahead and look at an example and see how that works in real time so the first video here is without the implementation we've done so it's a simple opus 4.7 we ask a very simple question and as you can see the model is thinking it's reasoning and it takes a couple of seconds to return the answer back to the user hey in this video we've added everything we just talked about on haiku 4.5 which is a much smaller model
  35. 3:36 same question but now you see that the answer comes in about 900 milliseconds and so that's the beauty of building around the model so by removing all of the thinking out of the logic all of the reasoning from the model and actually putting it within the code we actually saves a lot of time and allows us to use smaller models which are cost effective and actually better at real-time voice applications
  36. 4:04 And as you can see, this feels almost instant.
  37. 4:07 And again, that's because all of the smart parts have already happened prior to the model actually speaking.
  38. 4:15 But I have to be honest because this isn't necessarily free.
  39. 4:20 It has a cost, right?
  40. 4:22 A small model like the Haiku 4.5, if it doesn't have any scaffolding, tend to drift on long structure and really needs straight rules in order to be able to stay organized.
  41. 4:35 So the scaffolding piece is the price.
  42. 4:38 But the good thing is you pay it once and in code, right?
  43. 4:41 Not on every single turn.
  44. 4:44 So here's the rule.
  45. 4:45 Pick the fastest model that your latency budget allows, and then spend the rest of your time actually building the scaffolding.
  46. 4:54 So in our case, maybe you build a state machine.
  47. 4:58 You build a reasoning process.
  48. 4:59 You think about scenarios.
  49. 5:00 What happens if this happens?
  50. 5:02 How should your model handle it?

Open at this second