Videos r305-aQTaU0
Text Diffusion — Brendan O’Donoghue, Google DeepMind
Scene timeline
101 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 361
- whisperx 361
- chunks
- 50
- from 361 cues
- keyframes
- 44
- kept of 101 captured
- frames with text
- 44
- 581 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 10.4 MB
- word timings on 361 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 14:05 | 2m 16s |
stt |
done | — | 2026-08-10 14:08 | 30s |
chunk |
done | — | 2026-08-10 14:08 | 0s |
text_embed |
done | — | 2026-08-10 19:49 | 1s |
keyframe |
done | — | 2026-08-10 14:08 | 2m 36s |
ocr |
done | — | 2026-08-10 14:11 | 13s |
frame_embed |
done | — | 2026-08-10 19:49 | 8s |
Frames, and what the machine read
-
- Al Engineer0.93
- EUROPE1.00
-
- PRESENTING SPONSOR0.99
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.95
- WorkOS OpenAI0.96
-
- Text Diffusion1.00
- Speaker: Brendan O'Donoghue0.99
- irector of Research, Google DeepMind1.00
- AlEngineer0.99
- EUROPE1.00
-
- Image diffusion1.00
- Adding noise1.00
- ***0.73
- AIE1.00
- ★1.00
- Denoising1.00
- ★1.00
- ★1.00
- ★1.00
- Google DeepMind1.00
- Engineering the future of Al1.00
- AlEngineer0.99
- 20260.93
-
- Text diffusion1.00
- The cat was sitting on the table while looking through the window1.00
- *★*0.65
- ★1.00
- AIE1.00
- ★0.99
- Adding noise1.00
- ★1.00
- Denoing0.73
- The cat was plant on the table while aboveBxo through the window0.99
- ★1.00
- ★1.00
- ★1.00
- The cat was plant on the table hんに aboveBxo consider the window0.97
- not 上 John plant include a words んに aboveBxo consider 34 puede0.98
- Braintrust1.00
- WorkOS OpenAI0.98
- AlEngineer0.99
- 20260.93
-
- Gemini1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 1YEARAGO1.00
- 日0.57
- Google 1/O 25 Keynote0.95
- Google0.94
- 3ShareShop ± Dowmioad0.79
- X p0.54
- 口Save0.91
- Engineering the future of Al0.99
- AlEngineer0.98
- EUROPE1.00
- 20260.95
Transcript
361 cues· 4,637 words· 24,753 chars
- 0:15 So people are still filtering into the room.
- 0:18 It's mostly intro stuff for the first couple of slides, so they won't miss anything.
- 0:21 OK, welcome, everybody.
- 0:22 My name's Brendan.
- 0:23 I'm a research scientist at DeepMind.
- 0:25 I'm talking today about text diffusion, which is kind of a more forward-looking research area at DeepMind.
- 0:32 So you're probably familiar with image and video diffusion, which is kind of state of the art for these modalities right now, where you take ground truth, say image, you add noise to it in training, and then you train a neural network to remove that noise gradually.
- 0:48 And then at inference time, you just initialize
- 0:51 the picture with pure noise, and then you iteratively refine out the noise to recover back to whatever image or video or audio or whatever you're looking for.
- 1:00 And the principle is essentially the same for text, for text diffusion, where you start with a clean sequence of tokens, so a clean sentence or something like that.
- 1:10 And then you gradually add noise.
- 1:12 You corrupt it somehow.
- 1:13 There's lots of different ways to do that.
- 1:15 You can do it in a continuous or discrete way.
- 1:17 But let's just say discrete for now, which would, in this case, just mean adding random tokens or replacing tokens with other random tokens.
- 1:23 And you do that for a bunch of different noise levels.
- 1:25 And you train the neural network to try to fill in, to try to correct
- 1:30 the mistakes basically in the text.
- 1:32 And then at inference time, you initialize the sequence of tokens to just pure noise, like pure random discrete tokens from the vocabulary.
- 1:40 And then you iteratively refine through that to fill in the information in the order that the neural network wants to do to recover back to, say, a clean sentence.
- 1:49 And then in practice, I showed you some GIFs here of what it looks like for images.
- 1:54 You get very similar looking outputs for text, where it kind of starts off all noisy and then gradually fills in the text.
- 2:01 And then you get relatively clean outputs at the end.
- 2:06 OK, so the team I'm on, we had a research demo release one year ago now called Gemini Diffusion, which was a variant of a Gemini model with text diffusion instead of autoregressive next token generation.
- 2:19 And that was like a research preview that was open to about 100k people.
- 2:23 uh and you know we're still you know keep keep posted for new new developments in that direction uh upcoming soon um and we did have some good numbers at the time but again it's a year ago which is like you know prehistoric times in this field uh our kind of main comparator model was gemini 2.0 flashlight at the time because that was the architecture we were branching from the text fusion model and we basically had very similar quality across the board there
- 2:50 Mostly a little bit of advantage in code, a little bit of disadvantage in some other areas, but kind of relatively similar performance at much better latencies.
- 3:00 But again, this is a year ago, so I wouldn't fixate too much on these numbers.
- 3:04 OK, so what's the difference between autoregressive generation and diffusion?
- 3:06 So in the standard vanilla Gemini, Gemma, GPT, whatever, generation of text, you do this.
- 3:15 You have some context that comes in, and you want to generate some response to that.
- 3:18 And the model does it one token at a time.
- 3:20 So it generates the first token, and then conditional on that generates the next token, and so on.
- 3:24 Whereas in diffusion, a context will come in, whatever that is, and it'll initialize, like I mentioned, like a long sequence of tokens, could be hundreds, could be thousands, could be shorter, depends on the model, to be random noise.
- 3:37 And then it iteratively refines that canvas to remove the noise over the course of a few denoising steps.
- 3:43 So rather than one token at a time, it does the entire block together.
- 3:47 but over a couple of iterations.
- 3:49 So it's not just one pass, it has multiple passes, but it gets to attend to the future tokens and so on.
- 3:55 So it's kind of a different way of generating text.
- 4:00 So that obviously has some pros and cons.
- 4:02 So the main pro that people really like, and probably is the biggest advantage that text diffusion models have, is that it's just faster inference, it just generates faster tokens per second, because it makes much better use of the hardware, the TPU and the GPU.
- 4:15 And I have some slides on that to explain why.
- 4:17 But some other advantages are it can do bidirectional attention within this canvas of tokens.
- 4:22 So, you know, autoregressive models can only attend to the past.
- 4:26 They have causal attention within their, you know, their transformer.
- 4:30 Whereas a text fusion model is not restricted, though.
- 4:32 They can attend to the future.
- 4:34 And that has some interesting properties, like it can do self-corrected generation based on future tokens.
- 4:39 So we could do some reasoning, see that I got the answer incorrect, and then go back and fix the reasoning and do it again.
- 4:45 I have a demo of that.
loading
Chapters
- 0:00 Introduction to Text Diffusion
- 1:02 How Text Diffusion Works (Training and Inference)
- 2:06 Gemini Diffusion Research Preview
- 3:04 Difference Between Autoregressive and Diffusion Models
- 4:02 Pros and Cons of Text Diffusion
- 6:13 Hardware Efficiency: Why Text Diffusion is Faster
- 8:47 Bidirectional Reasoning and Self-Correction
- 12:00 Dynamic and Adaptive Computation
- 14:26 In-place Text Editing
- 16:09 Low Latency Applications and Demos
- 20:05 Q&A Session