Videos Bc6Ojl2XS1w
From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
Scene timeline
106 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 240
- whisperx 240
- chunks
- 34
- from 240 cues
- keyframes
- 60
- kept of 106 captured
- frames with text
- 60
- 1,571 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 11.2 MB
- word timings on 240 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 10:56 | 1m 20s |
stt |
done | — | 2026-08-11 10:58 | 20s |
chunk |
done | — | 2026-08-11 10:58 | 0s |
text_embed |
done | — | 2026-08-11 10:58 | 0s |
keyframe |
done | — | 2026-08-11 10:58 | 2m 14s |
ocr |
done | — | 2026-08-11 11:00 | 33s |
frame_embed |
done | — | 2026-08-11 11:01 | 10s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- WHAT'S NEW IN AI AUDIO?1.00
- at1.00
- Google DeepMind1.00
- Thor Schaeff I @thorwebdev0.96
- ALFngineer0.93
-
- AIE1.00
- Thor Schaeff0.95
- Developer Relations Engineer,0.98
- @thorwebdev1.00
- Google DeepMind0.99
- AlEngineer0.99
-
- AIE1.00
- Thor Schaeff0.97
- Developer Relations Engineer,1.00
- @thorwebdev1.00
- Engineering the future of Al1.00
- AlEngineer1.00
-
- AIE1.00
- Thor.Schaeff:0.93
- Developer Relations Engineer,0.98
- @thorwebdev1.00
- Braintrust1.00
- WorkOS OpenAI0.95
- AIE1.00
-
- EchoScript | Google Al S0.92
- Gemini Voice Library0.98
- Ooogle AI Studilo0.87
- [Public] Live Jukebax | Ooog0.87
- +Ask Oemini0.89
- alstudio.google.com/apps/bundled/echoscript?showPreview=true&showAssistant=true&fullscreenApplet=true0.98
- 白0.59
- Ca0.67
- EchoScript1.00
- Remix0.92
- L Device0.86
- 650.71
- EchoScript Al0.99
- Powered by Gemini 3 Flash Preview0.99
- Turn your audio into accurate text1.00
- Upload a file or record directly to get speaker-identified transcripts with1.00
- timestamps and language detection instantly.1.00
- ★1.00
- AIE1.00
- Record Audio1.00
- Upload File0.99
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Recording...1.00
- 01:011.00
- Stop Recording0.99
- Braintrust1.00
- WorkOS OpenAI0.93
- Al0.90
-
- Shipping at relentless pace1.00
- Frontier Models0.97
- Gemini 2.5 Pro0.99
- Gemini 2.5 Pro1.00
- Gemini 2.5 Pro0.98
- Gemini 2.5 Flash-Lite0.99
- ★1.00
- Geminl 2.5 Flash0.98
- AIE1.00
- ★1.00
- Geminl1.5 Pro0.95
- Gemini 2.00.99
- Flash1.00
- Gemini 2.00.97
- Flash-Lite1.00
- Geminl 3.1 Flash-Lite0.97
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Gemini 1.5 Flash0.97
- Flash Thinking0.98
- Gemini 2.00.98
- Gemini 2.0 Pro0.99
- Geminl 3.0 Pro0.96
- Gemini 3.1 Pro0.99
- (Preview)1.00
- 2024May-Jun-Jul-Aug-Sep-Oct-Nov-Dec2025Jan-Feb-Mar-Apr-May-Jun-Jul-Aug-Sep-Oct-Nov-Dec1.00
- 2026Jan - Feb -Mar - Apr0.95
- Gemma 20.99
- Gemma 30.94
- Gemma 40.99
- 98.2780.93
- 18. 48. 128, 2780.89
- E28, E480.268 A48, 3180.82
- Gemma 21.00
- Gemma 3n0.96
- E28. E480.87
- Gemma 2 for Japan0.97
- Gemma 30.99
- 270M1.00
- Open Models0.96
- 20240.91
- AlEngineer0.97
- AlEngineer0.99
- EUROPE1.00
- 20260.96
-
- genMedia Models0.98
- Imagen 4 Standard/UItra0.98
- Imagen 4 Standard/UItra0.99
- (Preview)0.93
- Veo 3 & Veo 3 Fast0.98
- Veo 3.1 Lite1.00
- (Prevlew)1.00
- Lyria RealTime0.99
- Imagen 4 Fast1.00
- dard. Ultra GA)0.93
- Lyria 3 Clip & Pro0.97
- *★*0.52
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Imagen 31.00
- Veo1.00
- Genie 21.00
- (Preview)0.92
- Veo 20.92
- Imagen 3 0020.99
- Gemini 2.0 Flash Image0.98
- Veo 20.96
- (GA)1.00
- Nano Banana1.00
- (Preview)0.94
- Veo 3.11.00
- Veo 3.1 Fast1.00
- Nano Banana Pro0.99
- Nano-Banana 21.00
- ★1.00
- ★1.00
- ★1.00
- 2024May-Jun-Jul-Aug-Sep-Oct-Nov-Dec2025Jan-Feb-Mar-Apr-May- Jun-Jul-Aug-Sep-Oct-Nov-Dec 20260.99
- Jan - Feb - Mar - Apr0.91
- Gemini 3.1 Flash Live0.98
- Chirp 30.99
- Gemini 2.5 Flash/Pro TTS0.98
- Gemini 2.51.00
- Flash/Pro TTS1.00
- Gemini 2.5 Flash Native Audio0.99
- Gemini 2.5 Flash Native Audio1.00
- Gemini 2.0 Live0.96
- Audio models1.00
- Engineering the future of Al0.99
- AlEng0.96
-
- Audio1.00
- AIE1.00
- Understanding1.00
- Engineering the future of Al0.99
- AlEng0.95
Transcript
240 cues· 2,734 words· 14,590 chars
- 0:15 What's new in AI audio?
- 0:17 I'm sorry, it's a little bit misleading because the title leaves out the Google DeepMind.
- 0:24 So we're just kind of looking at what we've been working on at DeepMind.
- 0:28 If we were to look at everything in AI audio, we'd be spending a lot of time here.
- 0:33 But I'd love to show you what we're working on at DeepMind.
- 0:40 Yeah, this is me.
- 0:41 Hi, everyone.
- 0:42 I'm Thor.
- 0:43 I work on the developer experience at Google DeepMind, working on the Gemini API and Google AI Studio.
- 0:51 Hello, everyone.
- 0:53 Welcome.
- 0:53 My name is Thorsten.
- 0:55 Bonjour.
- 0:56 Je m'appelle Thor.
- 0:58 Je suis très désolé.
- 0:59 Mon français, c'est très mauvais.
- 1:02 Konnichiwa.
- 1:02 Au revoir.
- 1:03 Writing da.
- 1:07 Okay, that was for the demo and now I just need to make sure last time I did this demo, I recorded over it and then it was all gone.
- 1:21 That was very sad, but we'll come back to that in a bit.
- 1:26 Yeah, what have we been up to at DeepMind?
- 1:29 There's been a couple releases.
- 1:32 I actually joined the team in November, literally the day before Gemini 3 was released.
- 1:38 So I joined, and they told me, tomorrow we're releasing Gemini 3.
- 1:42 And I was like, yay!
- 1:43 Didn't do anything, but it was great.
- 1:46 It was a great time.
- 1:48 Most recently, on the open model side, we released Gemma 4, I think literally last week.
- 1:54 And yeah, pretty incredible.
- 1:57 Some cool stuff you can do there.
- 1:58 Multi modality as well, baked into Gemma 4.
- 2:01 So there's audio understanding in the Gemma 4 models, and you can do that on device, on kind of edge devices as well.
- 2:10 So that is some very exciting stuff.
- 2:14 In terms of gen media and audio, you're probably very familiar with our image generation models, video generation models.
- 2:23 Obviously, Veo has audio generation in there as well.
- 2:28 So this is kind of where
- 2:30 The progression is there most recently with VO 3.1 Lite on the Gen Media model site.
- 2:36 And then on the audio models, we recently launched Gemini 3.1 Flash Life, which is our kind of full duplex, you know, sound-to-sound, real-time conversational model.
- 2:50 Also multimodal, so you can ingest real-time text, voice, vision, which we'll look at in a bit.
- 3:01 So on audio, very, very broad topic, but the baseline of everything we do are the frontier Gemini models.
- 3:10 And so Gemini 3 is incredibly good at understanding audio.
- 3:16 And that's not just transcribing it, but really understanding all the nuances that are in there.
- 3:23 So that might be obviously speech, but also the context of the speech, the emotion,
- 3:29 your pacing, your sort of, yeah, anything that sort of swings within the audio that's not just text.
- 3:40 So on the audio understanding, our goal is to build models that deeply comprehend, richly transcribe, and robustly reason through audio, seamlessly handling a large mix of different languages, dialect, accents, and modalities.
- 3:56 And sort of anywhere and always.
- 3:59 Gemini is really good at transcribing even people that are talking over each other, which is pretty incredible, seamlessly switching between different languages.
- 4:10 That was sort of the demo we're looking at now.
- 4:14 So EchoScript is kind of Gemini 3 Flash preview to sort of analyze audio recordings and extract sort of information out of it.
loading