read-only demo

Videos hMlLw1LeIK8

Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS

index_state ready data_status ok

AI Engineer· published 2026-07-20· 0:32:56· en-US· indexed 2026-08-10 19:44

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:27, 1 of 1 keyframes kept
  2. Shot 1, 0:27 to 0:54, 0 of 1 keyframes kept
  3. Shot 2, 0:54 to 1:28, 1 of 1 keyframes kept
  4. Shot 3, 1:28 to 2:03, 1 of 1 keyframes kept
  5. Shot 4, 2:03 to 2:37, 0 of 1 keyframes kept
  6. Shot 5, 2:37 to 2:41, 1 of 1 keyframes kept
  7. Shot 6, 2:41 to 3:11, 1 of 1 keyframes kept
  8. Shot 7, 3:11 to 3:40, 0 of 1 keyframes kept
  9. Shot 8, 3:40 to 3:53, 1 of 1 keyframes kept
  10. Shot 9, 3:53 to 4:18, 1 of 1 keyframes kept
  11. Shot 10, 4:18 to 4:44, 0 of 1 keyframes kept
  12. Shot 11, 4:44 to 5:09, 0 of 1 keyframes kept
  13. Shot 12, 5:09 to 5:29, 1 of 1 keyframes kept
  14. Shot 13, 5:29 to 5:58, 1 of 1 keyframes kept
  15. Shot 14, 5:58 to 6:27, 0 of 1 keyframes kept
  16. Shot 15, 6:27 to 6:56, 1 of 1 keyframes kept
  17. Shot 16, 6:56 to 7:27, 0 of 1 keyframes kept
  18. Shot 17, 7:27 to 7:57, 0 of 1 keyframes kept
  19. Shot 18, 7:57 to 8:27, 1 of 1 keyframes kept
  20. Shot 19, 8:27 to 8:38, 1 of 1 keyframes kept
  21. Shot 20, 8:38 to 9:04, 1 of 1 keyframes kept
  22. Shot 21, 9:04 to 9:29, 0 of 1 keyframes kept
  23. Shot 22, 9:29 to 9:54, 0 of 1 keyframes kept
  24. Shot 23, 9:54 to 10:11, 1 of 1 keyframes kept
  25. Shot 24, 10:11 to 10:39, 1 of 1 keyframes kept
  26. Shot 25, 10:39 to 11:06, 0 of 1 keyframes kept
  27. Shot 26, 11:06 to 11:33, 0 of 1 keyframes kept
  28. Shot 27, 11:33 to 11:59, 0 of 1 keyframes kept
  29. Shot 28, 11:59 to 12:25, 0 of 1 keyframes kept
  30. Shot 29, 12:25 to 12:50, 0 of 1 keyframes kept
  31. Shot 30, 12:50 to 13:16, 0 of 1 keyframes kept
  32. Shot 31, 13:16 to 13:51, 0 of 1 keyframes kept
  33. Shot 32, 13:51 to 14:26, 0 of 1 keyframes kept
  34. Shot 33, 14:26 to 15:07, 1 of 1 keyframes kept
  35. Shot 34, 15:07 to 15:36, 1 of 1 keyframes kept
  36. Shot 35, 15:36 to 16:05, 0 of 1 keyframes kept
  37. Shot 36, 16:05 to 16:34, 0 of 1 keyframes kept
  38. Shot 37, 16:34 to 17:03, 0 of 1 keyframes kept
  39. Shot 38, 17:03 to 17:37, 1 of 1 keyframes kept
  40. Shot 39, 17:37 to 18:12, 0 of 1 keyframes kept
  41. Shot 40, 18:12 to 18:37, 0 of 1 keyframes kept
  42. Shot 41, 18:37 to 19:03, 0 of 1 keyframes kept
  43. Shot 42, 19:03 to 19:28, 0 of 1 keyframes kept
  44. Shot 43, 19:28 to 19:53, 0 of 1 keyframes kept
  45. Shot 44, 19:53 to 20:27, 1 of 1 keyframes kept
  46. Shot 45, 20:27 to 21:00, 0 of 1 keyframes kept
  47. Shot 46, 21:00 to 21:33, 0 of 1 keyframes kept
  48. Shot 47, 21:33 to 21:41, 1 of 1 keyframes kept
  49. Shot 48, 21:41 to 22:24, 1 of 1 keyframes kept
  50. Shot 49, 22:24 to 22:58, 1 of 1 keyframes kept
  51. Shot 50, 22:58 to 23:01, 0 of 1 keyframes kept
  52. Shot 51, 23:01 to 23:04, 0 of 1 keyframes kept
  53. Shot 52, 23:04 to 23:14, 1 of 1 keyframes kept
  54. Shot 53, 23:14 to 23:34, 1 of 1 keyframes kept
  55. Shot 54, 23:34 to 23:42, 0 of 1 keyframes kept
  56. Shot 55, 23:42 to 24:21, 1 of 1 keyframes kept
  57. Shot 56, 24:21 to 24:23, 1 of 1 keyframes kept
  58. Shot 57, 24:23 to 24:56, 1 of 1 keyframes kept
  59. Shot 58, 24:56 to 25:25, 1 of 1 keyframes kept
  60. Shot 59, 25:25 to 25:55, 0 of 1 keyframes kept
  61. Shot 60, 25:55 to 26:24, 1 of 1 keyframes kept
  62. Shot 61, 26:24 to 26:26, 0 of 1 keyframes kept
  63. Shot 62, 26:26 to 26:43, 0 of 1 keyframes kept
  64. Shot 63, 26:43 to 26:52, 0 of 1 keyframes kept
  65. Shot 64, 26:52 to 27:18, 0 of 1 keyframes kept
  66. Shot 65, 27:18 to 27:43, 0 of 1 keyframes kept
  67. Shot 66, 27:43 to 28:09, 0 of 1 keyframes kept
  68. Shot 67, 28:09 to 28:35, 0 of 1 keyframes kept
  69. Shot 68, 28:35 to 29:01, 0 of 1 keyframes kept
  70. Shot 69, 29:01 to 29:27, 0 of 1 keyframes kept
  71. Shot 70, 29:27 to 29:53, 0 of 1 keyframes kept
  72. Shot 71, 29:53 to 29:55, 1 of 1 keyframes kept
  73. Shot 72, 29:55 to 30:03, 0 of 1 keyframes kept
  74. Shot 73, 30:03 to 30:13, 0 of 1 keyframes kept
  75. Shot 74, 30:13 to 30:41, 1 of 1 keyframes kept
  76. Shot 75, 30:41 to 31:08, 1 of 1 keyframes kept
  77. Shot 76, 31:08 to 31:35, 0 of 1 keyframes kept
  78. Shot 77, 31:35 to 32:02, 0 of 1 keyframes kept
  79. Shot 78, 32:02 to 32:29, 0 of 1 keyframes kept
  80. Shot 79, 32:29 to 32:56, 0 of 1 keyframes kept

80 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
266
whisperx 266
chunks
56
from 266 cues
keyframes
32
kept of 80 captured
frames with text
32
1,730 lines read
chapters
0
from the source metadata
keyframe bytes
10.8 MB
word timings on 266 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 00:28 1m 55s
stt done 2026-08-10 00:30 30s
chunk done 2026-08-10 00:31 0s
text_embed done 2026-08-10 19:43 1s
keyframe done 2026-08-10 00:31 1m 35s
ocr done 2026-08-10 00:33 34s
frame_embed done 2026-08-10 19:43 5s

Frames, and what the machine read

  • 0:08 #0 done8 line(s)

    shot 0·sharpness 1283.9

    1. Chintan Agrawal1.00
    2. AI ENGINEER WORLD'S FAIR 20260.99
    3. Voice agents0.97
    4. that handle1.00
    5. interrupts1.00
    6. The real-time audio engineering nobody talks about1.00
    7. Chintan Agrawal & Daniel Wirjo· AWS Solutions Architects0.99
    8. 2026-06-23 23:58:520.98
  • 0:50 #1 skipped

    shot 1·duplicate of #0

  • 0:58 #2 done24 line(s)

    shot 2·sharpness 2534.0

    1. amazon0.96
    2. aws1.00
    3. Chintan Agrawal1.00
    4. WHAT WE'LL COVER1.00
    5. Eight stops1.00
    6. 011.00
    7. What does a natural-sounding Al voice agent feel1.00
    8. 051.00
    9. Level 2 — built-in turn detection0.99
    10. like?1.00
    11. 021.00
    12. The200ms constraint1.00
    13. 061.00
    14. Level 3 — Smart Turn, own your stack0.98
    15. 031.00
    16. Architecture — the pipeline0.98
    17. 071.00
    18. Latency budget1.00
    19. 041.00
    20. Level 1 — Silero VAD0.99
    21. 081.00
    22. Production lessons & open problems0.99
    23. 02 / 230.97
    24. 2026-06-23 23:59:520.98
  • 1:42 #3 done21 line(s)

    shot 3·sharpness 2356.8

    1. aws1.00
    2. 70.94
    3. Chintan Agrawal1.00
    4. 011.00
    5. INTRODUCTION1.00
    6. Same user. Same sentence. Different result.0.98
    7. The difference between a voice agent that feels broken and one that feels natural comes down to how well it1.00
    8. handles interrupts and turn-taking.0.98
    9. BROKEN1.00
    10. WORKING1.00
    11. User: "I want to fly to—"1.00
    12. User: "I want to fly to—"0.97
    13. Agent keeps talking — ignores interruption for 1.8 s0.99
    14. STOP — interruption detected0.96
    15. User: "no wait, a HOTEL—" (ignored)0.98
    16. User: "actually, a hotel in Sydney"1.00
    17. 1.8 s delay0.97
    18. Agent: "Hotel in Sydney — got it!"0.97
    19. 180 ms1.00
    20. 03 / 230.93
    21. 2026-06-2400:00:451.00
  • 2:33 #4 skipped

    shot 4·duplicate of #3

  • 2:38 #5 done10 line(s)

    shot 5·sharpness 1208.5

    1. aws1.00
    2. Chintan Agrawal1.00
    3. 021.00
    4. . THE CONSTRAINT0.94
    5. The 200 ms0.99
    6. constraint1.00
    7. 04 / 230.93
    8. 4/230.91
    9. ResetR1.00
    10. 2026-06-24 00:01:521.00
  • 3:01 #6 done25 line(s)

    shot 6·sharpness 2145.1

    1. aws1.00
    2. Chintan Agrawal1.00
    3. THE 200 MS TARGET0.99
    4. Human turn-gap: measured across 10 languages1.00
    5. 200 ms is the time from when the user stops speaking to when the agent starts responding. Below that,0.99
    6. conversation feels natural. Above it, users pause, repeat themselves, or assume the call dropped. Chat agents0.99
    7. get seconds — voice agents don't get a second chance.0.98
    8. 0 ms0.97
    9. 200 ms1.00
    10. 400 ms1.00
    11. 800 ms0.99
    12. 1500 ms+0.98
    13. NATURAL1.00
    14. ACCEPTABLE1.00
    15. NOTICEABLE1.00
    16. BROKEN1.00
    17. CALL DROPPED?1.00
    18. 200 ms1.00
    19. 755 ms0.99
    20. Best measured human TTFA1.00
    21. Best measured TTFA — cascaded pipeline0.99
    22. Stivers et al., PNAS 2009 — 10 languages, 10K+ conversations0.99
    23. Qiu et al., Salesforce 20260.99
    24. 05 / 230.97
    25. 2026-06-24 00:02:200.99
  • 3:28 #7 skipped

    shot 7·duplicate of #6

  • 3:51 #8 done10 line(s)

    shot 8·sharpness 1229.5

    1. amazon1.00
    2. aws1.00
    3. Chintan Agrawal1.00
    4. 031.00
    5. ARCHITECTURE1.00
    6. The voice0.95
    7. agent pipeline1.00
    8. 031.00
    9. 06 / 230.88
    10. 2026-06-24 00:03:200.98
  • 3:58 #9 done36 line(s)

    shot 9·sharpness 2523.3

    1. amazon0.99
    2. aws1.00
    3. Chintan Agrawal0.99
    4. FRAME-BASED· ASYNC· BACKPRESSURE-AWARE0.98
    5. Voice agents need a pipeline0.99
    6. Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS) are all critical components of a voice agent. But0.99
    7. VAD — voice activity detection — is an often-overlooked layer. It decides when the user has finished speaking. Get it wrong0.98
    8. and everything downstream breaks.1.00
    9. WebRTC1.00
    10. VAD1.00
    11. STT0.84
    12. LLM1.00
    13. TTS1.00
    14. WebRTC1.00
    15. DAILY1.00
    16. SILERO V51.00
    17. CARTESIA INK0.96
    18. CLAUDE / GPT0.98
    19. 1.00
    20. CARTESIA1.00
    21. DAILY1.00
    22. 20 MS FRAMES0.95
    23. 32 MS FRAMES1.00
    24. STREAMING1.00
    25. TOKEN STREAM1.00
    26. PCM 24 KHZ1.00
    27. OUT1.00
    28. SMART TURN DETECTION1.00
    29. INTERRUPTION HANDLER1.00
    30. Listens to VAD + audio features. Decides: respond or wait.0.99
    31. Propagates InterruptionFrame — flushes TTS + LLM buffers.0.99
    32. WHAT IS PIPECAT?1.00
    33. 10k+0.97
    34. Open-source framework for real-time voice agents — wires STT, LLM, TTS, and VAD into a frame-based pipeline with backpressure, interruption handling, and streaming built in.1.00
    35. GITHUB STARS1.00
    36. 2026-06-24 00:03:290.99
  • 4:21 #10 skipped

    shot 10·duplicate of #9

  • 4:47 #11 skipped

    shot 11·duplicate of #9

  • 5:17 #12 done9 line(s)

    shot 12·sharpness 1384.0

    1. aws1.00
    2. Chintan Agrawal1.00
    3. 04. LEVEL 10.91
    4. Silero VAD:1.00
    5. take control1.00
    6. The simplest component you fully own. This is where most production1.00
    7. systems start.1.00
    8. 08/231.00
    9. 2026-06-24 00:05:030.99
  • 5:54 #13 done32 line(s)

    shot 13·sharpness 1662.7

    1. amazon1.00
    2. aws1.00
    3. 70.91
    4. Chintan Agrawal0.98
    5. 041.00
    6. LEVEL11.00
    7. Silero VAD — 309K parameters, <1 ms0.95
    8. STFT1.00
    9. 4×Conv1d1.00
    10. LSTM (128)0.97
    11. Linear1.00
    12. 0→11.00
    13. 16 KHZ PCM0.98
    14. SPECTRAL1.00
    15. FEATURES1.00
    16. + RELU0.98
    17. TEMPORAL0.99
    18. CONTEXT1.00
    19. P(SPEECH)1.00
    20. SIGMOID0.95
    21. SPEECH PROB0.99
    22. 309K1.00
    23. 2 MB0.96
    24. <1 ms0.98
    25. MCC 0.720.96
    26. PARAMETERS1.00
    27. MODEL SIZE0.99
    28. PER 32 MS FRAME1.00
    29. VS WEBRTC 0.411.00
    30. vad = SileroVADAnalyzer(params=VADParams(stop_secs=0.3)) # confirmed: Qiu et al. (2026) same config0.99
    31. 09 / 230.97
    32. 2026-06-24 00:05:480.99
  • 6:04 #14 skipped

    shot 14·duplicate of #13

  • 6:50 #15 done31 line(s)

    shot 15·sharpness 1514.2

    1. amazor0.99
    2. aws1.00
    3. Chintan Agrawal1.00
    4. 041.00
    5. LEVEL 11.00
    6. Tuning stop_secs — the critical parameter0.99
    7. SALES / 0.2 50.97
    8. GENERAL / 0.5 50.96
    9. HEALTH / 0.8 S0.95
    10. DATA ENTRY / 1.2 S0.97
    11. Snappy. May cut off1.00
    12. Balanced. Most users1.00
    13. Patient. Respects1.00
    14. Never interrupts. Dead-air1.00
    15. thinking.1.00
    16. happy.1.00
    17. pauses.1.00
    18. risk.1.00
    19. # The critical parameter1.00
    20. class VADParams:0.99
    21. threshold1.00
    22. = 0.50.92
    23. min_speech_ms1.00
    24. = 2500.93
    25. min_silence_ms1.00
    26. = 3000.94
    27. # THIS ONE0.96
    28. speech_pad_ms1.00
    29. = 1000.95
    30. 10 / 230.97
    31. 2026-06-2400:06:551.00
  • 7:12 #16 skipped

    shot 16·duplicate of #9

  • 7:36 #17 skipped

    shot 17·duplicate of #9

  • 8:12 #18 done24 line(s)

    shot 18·sharpness 2406.4

    1. amazon0.99
    2. aws1.00
    3. Chintan Agrawal1.00
    4. THREE LEVELS OF SOLVING THIS - INCREASING CONTROL, INCREASING COMPLEXITY0.99
    5. What VAD alone can't solve0.97
    6. 011.00
    7. SILENCE DETECTION0.99
    8. 021.00
    9. BARGE-IN1.00
    10. 031.00
    11. TURN DETECTION1.00
    12. When did they stop talking?0.99
    13. They started talking. Stop, fade,0.99
    14. Same silence means four things1.00
    15. or finish?1.00
    16. Is the user done, or just thinking? 300 ms and 1200 ms0.99
    17. Complete sentence. Incomplete thought. Thinking1.00
    18. silence look identical to raw VAD.0.99
    19. The decision depends on context — correction,0.99
    20. pause. Backchannel "yeah". VAD fires on all of them.0.99
    21. affirmation, or ambient noise. Most agents don't0.99
    22. distinguish.1.00
    23. 11 / 230.93
    24. 2026-06-24 00:08:330.98
  • 8:28 #19 done13 line(s)

    shot 19·sharpness 1399.6

    1. amagon0.99
    2. aws0.99
    3. Chintan Agrawal1.00
    4. 05LEVEL 21.00
    5. Built-in turn1.00
    6. detection1.00
    7. Smarter than raw VAD. State-of-the-art STT models are now starting to incorporate1.00
    8. turn detection.1.00
    9. 050.99
    10. 12/231.00
    11. 12/231.00
    12. ResetR1.00
    13. 2026-06-2400:08:521.00
  • 8:51 #20 done22 line(s)

    shot 20·sharpness 2507.9

    1. amazon1.00
    2. aws1.00
    3. 70.91
    4. Chintan Agrawal1.00
    5. 051.00
    6. LEVEL 20.97
    7. Cartesia Ink-2 and Deepgram Flux1.00
    8. CARTESIA INK-21.00
    9. DEEPGRAM FLUX1.00
    10. Turn detection built into the STT WebSocket protocol — emits turn. start0.99
    11. Endpointing built into the Deepgram streaming API — single-service turn0.99
    12. / turn.end events0.99
    13. detection1.00
    14. No local VAD or turn analyzer needed — the server drives boundaries1.00
    15. Configurable endpointing parameter controls silence threshold0.99
    16. P50 STT latency: 299 ms · P95: 328 ms0.97
    17. P50 STT latency: 247 ms · P95: 298 ms (nova-3)0.98
    18. Pipecat: CartesiaTurnsSTTService0.99
    19. Trade-off: no visibility into model decisions, vendor-specific behaviour0.98
    20. Both are Level 2 — smarter than raw VAD, no local model required. Best for prototypes and latency-sensitive deployments where control is less critical.0.98
    21. 13 / 230.89
    22. 2026-06-24 00:09:200.99
  • 9:16 #21 skipped

    shot 21·duplicate of #3

  • 9:41 #22 skipped

    shot 22·duplicate of #20

  • 10:06 #23 done10 line(s)

    shot 23·sharpness 1284.5

    1. amazon0.98
    2. aws1.00
    3. Chintan Agrawal1.00
    4. 06. LEVEL 30.93
    5. Silero VAD +1.00
    6. Smart Turn0.99
    7. Full control. Transparent. Tunable. Portable.1.00
    8. 061.00
    9. 14/231.00
    10. 2026-06-2400:10:501.00

Transcript

266 cues· 4,590 words· 24,790 chars

  1. 0:03 Hey, everyone.
  2. 0:04 I'm Chintan, and I have my colleague Daniel with me.
  3. 0:07 We are solutions architect on the AWS APG startup team.
  4. 0:11 So today, we are going to talk about something we've been working on for a while, which is turn-taking and voice agents.
  5. 0:18 What I mean by that is, how does the system actually know that you have finished speaking and it's time for the agent to respond?
  6. 0:25 These are all audio engineering problems.
  7. 0:27 They are not LLM problems, because you can have the perfect model, perfect track,
  8. 0:33 But the experience still might feel broken if the turn taking is off.
  9. 0:38 So we are going to walk you through the concepts and three different approaches to solving this.
  10. 0:44 And Daniel is going to make it more concrete by running all the three demo for you.
  11. 0:56 So to kick things off, we'll start with the 200 millisecond constraint, which is basically the physics of why voice is hard.
  12. 1:06 Then we look into the pipeline architecture, what are the components, where turn taking actually lives.
  13. 1:12 We'll then look at three levels of solving this problem, going from simple silence detection all the way up to running your own turn detection model.
  14. 1:21 and then we'll also talk about the latency the budget and production issues that we've seen so but before any of that i just want to show you the problem statement more clearly because you have the same user you have the same sentence on the both side but the user outcome is very different because
  15. 1:45 On the left, when the user says, I want to fly, and midway they try to correct themselves, the agent just does not notice because it keeps going on for almost two seconds while the person is sitting there trying to get a word in.
  16. 2:00 On the right side, the same exact interaction is happening, but the agent is able to capture the interruption in under 200 milliseconds and immediately backs out so the user cannot speak.
  17. 2:13 Both the scenarios, the LLM was identical.
  18. 2:15 It was the same model.
  19. 2:16 It was the same problem.
  20. 2:17 But the difference in the user experience is so different.
  21. 2:21 And the difference is purely because of the audio pipeline.
  22. 2:26 And how fast does it notice us?
  23. 2:28 Like someone else is talking and it knows when to shut up.
  24. 2:31 Like that's the problem we are trying to solve today.
  25. 2:37 So we cover this 200 millisecond constraint because 200 millisecond is how fast humans switch turns with each other in a conversation.
  26. 2:47 And the implications are pretty brutal because at 800 milliseconds, things start to feel off while at 1.5 second, your user would just hang up on you because you have like your HR chat agents might get five seconds to respond and nobody will care.
  27. 3:05 But with the voice agents, you don't get that luxury.
  28. 3:08 We've seen recently like Salesforce was able to pay for the pipeline and they published their results in March 26.
  29. 3:15 But even their best measure response time was 755 milliseconds.
  30. 3:19 So that's like almost 4x more than how humans naturally take turns while talking.
  31. 3:26 So the question here becomes, how do we close this gap?
  32. 3:29 And where we cannot close it, at least how do we get the turn taking right?
  33. 3:34 So the user experience does not feel as bad as the raw numbers might suggest here.
  34. 3:40 So if you're working with a budget of 755 millisecond and every millisecond count, you also need to understand what's actually in the pipeline.
  35. 3:49 Like where does the time go and where does this turn taking fit?
  36. 3:54 so the pipeline itself you probably might know the main components the std the llm the tts you have the audio in text out text in audio but the piece that's often missing like from people while people are thinking about this is this voice activity detection commonly called as bad it's a tiny component that sits right at the front and it is the one that controls term thinking
  37. 4:21 Its job is to basically detect whether the user has stopped talking or not.
  38. 4:26 And the two other things that you said alongside the main pipeline are also very critical for turn-taking.
  39. 4:33 So you have the smart turn detection, which watches the VRE signal, plus the audio features, and then it has to make that decision that whether should it respond now or should it wait because the user might still be talking.
  40. 4:47 And the interruption handler is triggered when someone barges in.
  41. 4:51 So what it'll do is it'll propagate a flush of downstream, like the TTS would stop and the LLM generation need to be canceled so that the pipeline is ready for the new input within about 15 millisecond.
  42. 5:05 So if you remember, I was talking about how there are different levels to building this turn taking in voice agent.
  43. 5:15 So we'll start with the level one.
  44. 5:17 Level one is your cellular VAD.
  45. 5:19 So this is a simple component which you fully own.
  46. 5:22 And like a lot of production system, even today, we see are running agents just by using this.
  47. 5:30 So coming to cellular VAD, it's a small 300,000 parameter model.
  48. 5:37 It takes in a short-term Fourier transform, and it takes raw audio, and it will convert it into spectral features.
  49. 5:46 It has four convolutional layers to pick up the pattern, an LSTM that gives it memory across the frame, so it's just not looking at one chunk in isolation.
  50. 5:57 And then it has a sigmoid that gives a probability of speech, and it's like a very small 2-megabyte model.

Open at this second