read-only demo

Videos Bc6Ojl2XS1w

From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind

index_state ready data_status ok

AI Engineer· published 2026-06-09· 0:19:33· en-US· indexed 2026-08-11 11:01

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:15, 1 of 1 keyframes kept
  4. Shot 3, 0:15 to 0:38, 1 of 1 keyframes kept
  5. Shot 4, 0:38 to 0:45, 1 of 1 keyframes kept
  6. Shot 5, 0:45 to 0:49, 0 of 1 keyframes kept
  7. Shot 6, 0:49 to 0:50, 0 of 1 keyframes kept
  8. Shot 7, 0:50 to 0:55, 1 of 1 keyframes kept
  9. Shot 8, 0:55 to 0:56, 0 of 1 keyframes kept
  10. Shot 9, 0:56 to 1:01, 0 of 1 keyframes kept
  11. Shot 10, 1:01 to 1:02, 0 of 1 keyframes kept
  12. Shot 11, 1:02 to 1:06, 0 of 1 keyframes kept
  13. Shot 12, 1:06 to 1:08, 0 of 1 keyframes kept
  14. Shot 13, 1:08 to 1:12, 1 of 1 keyframes kept
  15. Shot 14, 1:12 to 1:13, 0 of 1 keyframes kept
  16. Shot 15, 1:13 to 1:21, 1 of 1 keyframes kept
  17. Shot 16, 1:21 to 1:23, 0 of 1 keyframes kept
  18. Shot 17, 1:23 to 2:13, 1 of 1 keyframes kept
  19. Shot 18, 2:13 to 2:59, 1 of 1 keyframes kept
  20. Shot 19, 2:59 to 3:05, 1 of 1 keyframes kept
  21. Shot 20, 3:05 to 3:10, 0 of 1 keyframes kept
  22. Shot 21, 3:10 to 3:16, 0 of 1 keyframes kept
  23. Shot 22, 3:16 to 3:22, 0 of 1 keyframes kept
  24. Shot 23, 3:22 to 3:28, 0 of 1 keyframes kept
  25. Shot 24, 3:28 to 3:34, 0 of 1 keyframes kept
  26. Shot 25, 3:34 to 3:38, 1 of 1 keyframes kept
  27. Shot 26, 3:38 to 3:40, 1 of 1 keyframes kept
  28. Shot 27, 3:40 to 3:44, 0 of 1 keyframes kept
  29. Shot 28, 3:44 to 3:51, 0 of 1 keyframes kept
  30. Shot 29, 3:51 to 3:56, 0 of 1 keyframes kept
  31. Shot 30, 3:56 to 4:03, 0 of 1 keyframes kept
  32. Shot 31, 4:03 to 4:07, 0 of 1 keyframes kept
  33. Shot 32, 4:07 to 4:12, 0 of 1 keyframes kept
  34. Shot 33, 4:12 to 4:36, 1 of 1 keyframes kept
  35. Shot 34, 4:36 to 4:40, 1 of 1 keyframes kept
  36. Shot 35, 4:40 to 4:41, 0 of 1 keyframes kept
  37. Shot 36, 4:41 to 5:30, 1 of 1 keyframes kept
  38. Shot 37, 5:30 to 6:06, 0 of 1 keyframes kept
  39. Shot 38, 6:06 to 6:39, 1 of 1 keyframes kept
  40. Shot 39, 6:39 to 7:23, 1 of 1 keyframes kept
  41. Shot 40, 7:23 to 7:25, 1 of 1 keyframes kept
  42. Shot 41, 7:25 to 7:31, 0 of 1 keyframes kept
  43. Shot 42, 7:31 to 7:37, 0 of 1 keyframes kept
  44. Shot 43, 7:37 to 7:43, 1 of 1 keyframes kept
  45. Shot 44, 7:43 to 7:48, 0 of 1 keyframes kept
  46. Shot 45, 7:48 to 7:54, 0 of 1 keyframes kept
  47. Shot 46, 7:54 to 8:00, 0 of 1 keyframes kept
  48. Shot 47, 8:00 to 8:06, 0 of 1 keyframes kept
  49. Shot 48, 8:06 to 8:33, 1 of 1 keyframes kept
  50. Shot 49, 8:33 to 9:00, 1 of 1 keyframes kept
  51. Shot 50, 9:00 to 9:14, 1 of 1 keyframes kept
  52. Shot 51, 9:14 to 9:16, 0 of 1 keyframes kept
  53. Shot 52, 9:16 to 9:18, 1 of 1 keyframes kept
  54. Shot 53, 9:18 to 9:39, 1 of 1 keyframes kept
  55. Shot 54, 9:39 to 10:19, 1 of 1 keyframes kept
  56. Shot 55, 10:19 to 10:38, 1 of 1 keyframes kept
  57. Shot 56, 10:38 to 11:07, 1 of 1 keyframes kept
  58. Shot 57, 11:07 to 11:40, 0 of 1 keyframes kept
  59. Shot 58, 11:40 to 11:41, 1 of 1 keyframes kept
  60. Shot 59, 11:41 to 11:54, 0 of 1 keyframes kept
  61. Shot 60, 11:54 to 11:55, 1 of 1 keyframes kept
  62. Shot 61, 11:55 to 11:57, 1 of 1 keyframes kept
  63. Shot 62, 11:57 to 12:03, 0 of 1 keyframes kept
  64. Shot 63, 12:03 to 12:04, 0 of 1 keyframes kept
  65. Shot 64, 12:04 to 12:27, 1 of 1 keyframes kept
  66. Shot 65, 12:27 to 12:33, 1 of 1 keyframes kept
  67. Shot 66, 12:33 to 13:00, 1 of 1 keyframes kept
  68. Shot 67, 13:00 to 13:02, 1 of 1 keyframes kept
  69. Shot 68, 13:02 to 13:18, 1 of 1 keyframes kept
  70. Shot 69, 13:18 to 13:20, 1 of 1 keyframes kept
  71. Shot 70, 13:20 to 13:39, 0 of 1 keyframes kept
  72. Shot 71, 13:39 to 13:41, 1 of 1 keyframes kept
  73. Shot 72, 13:41 to 13:56, 1 of 1 keyframes kept
  74. Shot 73, 13:56 to 14:12, 1 of 1 keyframes kept
  75. Shot 74, 14:12 to 14:29, 1 of 1 keyframes kept
  76. Shot 75, 14:29 to 15:07, 0 of 1 keyframes kept
  77. Shot 76, 15:07 to 15:09, 0 of 1 keyframes kept
  78. Shot 77, 15:09 to 15:27, 1 of 1 keyframes kept
  79. Shot 78, 15:27 to 15:54, 1 of 1 keyframes kept
  80. Shot 79, 15:54 to 16:01, 1 of 1 keyframes kept
  81. Shot 80, 16:01 to 16:06, 0 of 1 keyframes kept
  82. Shot 81, 16:06 to 16:08, 1 of 1 keyframes kept
  83. Shot 82, 16:08 to 16:14, 1 of 1 keyframes kept
  84. Shot 83, 16:14 to 16:15, 1 of 1 keyframes kept
  85. Shot 84, 16:15 to 16:18, 1 of 1 keyframes kept
  86. Shot 85, 16:18 to 16:39, 1 of 1 keyframes kept
  87. Shot 86, 16:39 to 16:55, 0 of 1 keyframes kept
  88. Shot 87, 16:55 to 16:57, 1 of 1 keyframes kept
  89. Shot 88, 16:57 to 16:59, 1 of 1 keyframes kept
  90. Shot 89, 16:59 to 17:25, 1 of 1 keyframes kept
  91. Shot 90, 17:25 to 17:50, 1 of 1 keyframes kept
  92. Shot 91, 17:50 to 18:16, 0 of 1 keyframes kept
  93. Shot 92, 18:16 to 18:41, 0 of 1 keyframes kept
  94. Shot 93, 18:41 to 18:42, 1 of 1 keyframes kept
  95. Shot 94, 18:42 to 18:45, 1 of 1 keyframes kept
  96. Shot 95, 18:45 to 18:49, 0 of 1 keyframes kept
  97. Shot 96, 18:49 to 18:54, 0 of 1 keyframes kept
  98. Shot 97, 18:54 to 18:55, 0 of 1 keyframes kept
  99. Shot 98, 18:55 to 18:56, 1 of 1 keyframes kept
  100. Shot 99, 18:56 to 18:59, 0 of 1 keyframes kept
  101. Shot 100, 18:59 to 19:02, 1 of 1 keyframes kept
  102. Shot 101, 19:02 to 19:03, 1 of 1 keyframes kept
  103. Shot 102, 19:03 to 19:07, 0 of 1 keyframes kept
  104. Shot 103, 19:07 to 19:09, 0 of 1 keyframes kept
  105. Shot 104, 19:09 to 19:18, 1 of 1 keyframes kept
  106. Shot 105, 19:18 to 19:33, 1 of 1 keyframes kept

106 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
240
whisperx 240
chunks
34
from 240 cues
keyframes
60
kept of 106 captured
frames with text
60
1,571 lines read
chapters
0
from the source metadata
keyframe bytes
11.2 MB
word timings on 240 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 10:56 1m 20s
stt done 2026-08-11 10:58 20s
chunk done 2026-08-11 10:58 0s
text_embed done 2026-08-11 10:58 0s
keyframe done 2026-08-11 10:58 2m 14s
ocr done 2026-08-11 11:00 33s
frame_embed done 2026-08-11 11:01 10s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:27 #3 done5 line(s)

    shot 3·sharpness 675.6

    1. WHAT'S NEW IN AI AUDIO?1.00
    2. at1.00
    3. Google DeepMind1.00
    4. Thor Schaeff I @thorwebdev0.96
    5. ALFngineer0.93
  • 0:41 #4 done6 line(s)

    shot 4·sharpness 2225.6

    1. AIE1.00
    2. Thor Schaeff0.95
    3. Developer Relations Engineer,0.98
    4. @thorwebdev1.00
    5. Google DeepMind0.99
    6. AlEngineer0.99
  • 0:47 #5 skipped

    shot 5·duplicate of #4

  • 0:50 #6 skipped

    shot 6·duplicate of #4

  • 0:53 #7 done6 line(s)

    shot 7·sharpness 2225.3

    1. AIE1.00
    2. Thor Schaeff0.97
    3. Developer Relations Engineer,1.00
    4. @thorwebdev1.00
    5. Engineering the future of Al1.00
    6. AlEngineer1.00
  • 0:56 #8 skipped

    shot 8·duplicate of #4

  • 0:58 #9 skipped

    shot 9·duplicate of #7

  • 1:02 #10 skipped

    shot 10·duplicate of #7

  • 1:04 #11 skipped

    shot 11·duplicate of #7

  • 1:07 #12 skipped

    shot 12·duplicate of #7

  • 1:10 #13 done7 line(s)

    shot 13·sharpness 2289.6

    1. AIE1.00
    2. Thor.Schaeff:0.93
    3. Developer Relations Engineer,0.98
    4. @thorwebdev1.00
    5. Braintrust1.00
    6. WorkOS OpenAI0.95
    7. AIE1.00
  • 1:12 #14 skipped

    shot 14·duplicate of #13

  • 1:14 #15 done31 line(s)

    shot 15·sharpness 1706.3

    1. EchoScript | Google Al S0.92
    2. Gemini Voice Library0.98
    3. Ooogle AI Studilo0.87
    4. [Public] Live Jukebax | Ooog0.87
    5. +Ask Oemini0.89
    6. alstudio.google.com/apps/bundled/echoscript?showPreview=true&showAssistant=true&fullscreenApplet=true0.98
    7. 0.59
    8. Ca0.67
    9. EchoScript1.00
    10. Remix0.92
    11. L Device0.86
    12. 650.71
    13. EchoScript Al0.99
    14. Powered by Gemini 3 Flash Preview0.99
    15. Turn your audio into accurate text1.00
    16. Upload a file or record directly to get speaker-identified transcripts with1.00
    17. timestamps and language detection instantly.1.00
    18. 1.00
    19. AIE1.00
    20. Record Audio1.00
    21. Upload File0.99
    22. 1.00
    23. 1.00
    24. 1.00
    25. 1.00
    26. Recording...1.00
    27. 01:011.00
    28. Stop Recording0.99
    29. Braintrust1.00
    30. WorkOS OpenAI0.93
    31. Al0.90
  • 1:22 #16 skipped

    shot 16·duplicate of #7

  • 1:34 #17 done47 line(s)

    shot 17·sharpness 1833.6

    1. Shipping at relentless pace1.00
    2. Frontier Models0.97
    3. Gemini 2.5 Pro0.99
    4. Gemini 2.5 Pro1.00
    5. Gemini 2.5 Pro0.98
    6. Gemini 2.5 Flash-Lite0.99
    7. 1.00
    8. Geminl 2.5 Flash0.98
    9. AIE1.00
    10. 1.00
    11. Geminl1.5 Pro0.95
    12. Gemini 2.00.99
    13. Flash1.00
    14. Gemini 2.00.97
    15. Flash-Lite1.00
    16. Geminl 3.1 Flash-Lite0.97
    17. 1.00
    18. 1.00
    19. 1.00
    20. 1.00
    21. Gemini 1.5 Flash0.97
    22. Flash Thinking0.98
    23. Gemini 2.00.98
    24. Gemini 2.0 Pro0.99
    25. Geminl 3.0 Pro0.96
    26. Gemini 3.1 Pro0.99
    27. (Preview)1.00
    28. 2024May-Jun-Jul-Aug-Sep-Oct-Nov-Dec2025Jan-Feb-Mar-Apr-May-Jun-Jul-Aug-Sep-Oct-Nov-Dec1.00
    29. 2026Jan - Feb -Mar - Apr0.95
    30. Gemma 20.99
    31. Gemma 30.94
    32. Gemma 40.99
    33. 98.2780.93
    34. 18. 48. 128, 2780.89
    35. E28, E480.268 A48, 3180.82
    36. Gemma 21.00
    37. Gemma 3n0.96
    38. E28. E480.87
    39. Gemma 2 for Japan0.97
    40. Gemma 30.99
    41. 270M1.00
    42. Open Models0.96
    43. 20240.91
    44. AlEngineer0.97
    45. AlEngineer0.99
    46. EUROPE1.00
    47. 20260.96
  • 2:40 #18 done47 line(s)

    shot 18·sharpness 1964.6

    1. genMedia Models0.98
    2. Imagen 4 Standard/UItra0.98
    3. Imagen 4 Standard/UItra0.99
    4. (Preview)0.93
    5. Veo 3 & Veo 3 Fast0.98
    6. Veo 3.1 Lite1.00
    7. (Prevlew)1.00
    8. Lyria RealTime0.99
    9. Imagen 4 Fast1.00
    10. dard. Ultra GA)0.93
    11. Lyria 3 Clip & Pro0.97
    12. *★*0.52
    13. AIE1.00
    14. 1.00
    15. 1.00
    16. 1.00
    17. Imagen 31.00
    18. Veo1.00
    19. Genie 21.00
    20. (Preview)0.92
    21. Veo 20.92
    22. Imagen 3 0020.99
    23. Gemini 2.0 Flash Image0.98
    24. Veo 20.96
    25. (GA)1.00
    26. Nano Banana1.00
    27. (Preview)0.94
    28. Veo 3.11.00
    29. Veo 3.1 Fast1.00
    30. Nano Banana Pro0.99
    31. Nano-Banana 21.00
    32. 1.00
    33. 1.00
    34. 1.00
    35. 2024May-Jun-Jul-Aug-Sep-Oct-Nov-Dec2025Jan-Feb-Mar-Apr-May- Jun-Jul-Aug-Sep-Oct-Nov-Dec 20260.99
    36. Jan - Feb - Mar - Apr0.91
    37. Gemini 3.1 Flash Live0.98
    38. Chirp 30.99
    39. Gemini 2.5 Flash/Pro TTS0.98
    40. Gemini 2.51.00
    41. Flash/Pro TTS1.00
    42. Gemini 2.5 Flash Native Audio0.99
    43. Gemini 2.5 Flash Native Audio1.00
    44. Gemini 2.0 Live0.96
    45. Audio models1.00
    46. Engineering the future of Al0.99
    47. AlEng0.96
  • 3:02 #19 done5 line(s)

    shot 19·sharpness 1492.6

    1. Audio1.00
    2. AIE1.00
    3. Understanding1.00
    4. Engineering the future of Al0.99
    5. AlEng0.95
  • 3:07 #20 skipped

    shot 20·duplicate of #19

  • 3:13 #21 skipped

    shot 21·duplicate of #19

  • 3:19 #22 skipped

    shot 22·duplicate of #19

  • 3:25 #23 skipped

    shot 23·duplicate of #19

Transcript

240 cues· 2,734 words· 14,590 chars

  1. 0:15 What's new in AI audio?
  2. 0:17 I'm sorry, it's a little bit misleading because the title leaves out the Google DeepMind.
  3. 0:24 So we're just kind of looking at what we've been working on at DeepMind.
  4. 0:28 If we were to look at everything in AI audio, we'd be spending a lot of time here.
  5. 0:33 But I'd love to show you what we're working on at DeepMind.
  6. 0:40 Yeah, this is me.
  7. 0:41 Hi, everyone.
  8. 0:42 I'm Thor.
  9. 0:43 I work on the developer experience at Google DeepMind, working on the Gemini API and Google AI Studio.
  10. 0:51 Hello, everyone.
  11. 0:53 Welcome.
  12. 0:53 My name is Thorsten.
  13. 0:55 Bonjour.
  14. 0:56 Je m'appelle Thor.
  15. 0:58 Je suis très désolé.
  16. 0:59 Mon français, c'est très mauvais.
  17. 1:02 Konnichiwa.
  18. 1:02 Au revoir.
  19. 1:03 Writing da.
  20. 1:07 Okay, that was for the demo and now I just need to make sure last time I did this demo, I recorded over it and then it was all gone.
  21. 1:21 That was very sad, but we'll come back to that in a bit.
  22. 1:26 Yeah, what have we been up to at DeepMind?
  23. 1:29 There's been a couple releases.
  24. 1:32 I actually joined the team in November, literally the day before Gemini 3 was released.
  25. 1:38 So I joined, and they told me, tomorrow we're releasing Gemini 3.
  26. 1:42 And I was like, yay!
  27. 1:43 Didn't do anything, but it was great.
  28. 1:46 It was a great time.
  29. 1:48 Most recently, on the open model side, we released Gemma 4, I think literally last week.
  30. 1:54 And yeah, pretty incredible.
  31. 1:57 Some cool stuff you can do there.
  32. 1:58 Multi modality as well, baked into Gemma 4.
  33. 2:01 So there's audio understanding in the Gemma 4 models, and you can do that on device, on kind of edge devices as well.
  34. 2:10 So that is some very exciting stuff.
  35. 2:14 In terms of gen media and audio, you're probably very familiar with our image generation models, video generation models.
  36. 2:23 Obviously, Veo has audio generation in there as well.
  37. 2:28 So this is kind of where
  38. 2:30 The progression is there most recently with VO 3.1 Lite on the Gen Media model site.
  39. 2:36 And then on the audio models, we recently launched Gemini 3.1 Flash Life, which is our kind of full duplex, you know, sound-to-sound, real-time conversational model.
  40. 2:50 Also multimodal, so you can ingest real-time text, voice, vision, which we'll look at in a bit.
  41. 3:01 So on audio, very, very broad topic, but the baseline of everything we do are the frontier Gemini models.
  42. 3:10 And so Gemini 3 is incredibly good at understanding audio.
  43. 3:16 And that's not just transcribing it, but really understanding all the nuances that are in there.
  44. 3:23 So that might be obviously speech, but also the context of the speech, the emotion,
  45. 3:29 your pacing, your sort of, yeah, anything that sort of swings within the audio that's not just text.
  46. 3:40 So on the audio understanding, our goal is to build models that deeply comprehend, richly transcribe, and robustly reason through audio, seamlessly handling a large mix of different languages, dialect, accents, and modalities.
  47. 3:56 And sort of anywhere and always.
  48. 3:59 Gemini is really good at transcribing even people that are talking over each other, which is pretty incredible, seamlessly switching between different languages.
  49. 4:10 That was sort of the demo we're looking at now.
  50. 4:14 So EchoScript is kind of Gemini 3 Flash preview to sort of analyze audio recordings and extract sort of information out of it.

Open at this second