read-only demo

Videos r305-aQTaU0

Text Diffusion — Brendan O’Donoghue, Google DeepMind

index_state ready data_status ok

AI Engineer· published 2026-06-04· 0:28:03· en-US· indexed 2026-08-10 19:50

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:31, 1 of 1 keyframes kept
  5. Shot 4, 0:31 to 1:00, 1 of 1 keyframes kept
  6. Shot 5, 1:00 to 1:49, 1 of 1 keyframes kept
  7. Shot 6, 1:49 to 1:51, 0 of 1 keyframes kept
  8. Shot 7, 1:51 to 1:54, 0 of 1 keyframes kept
  9. Shot 8, 1:54 to 1:58, 0 of 1 keyframes kept
  10. Shot 9, 1:58 to 2:01, 0 of 1 keyframes kept
  11. Shot 10, 2:01 to 2:05, 0 of 1 keyframes kept
  12. Shot 11, 2:05 to 2:31, 1 of 1 keyframes kept
  13. Shot 12, 2:31 to 3:02, 0 of 1 keyframes kept
  14. Shot 13, 3:02 to 3:30, 0 of 1 keyframes kept
  15. Shot 14, 3:30 to 3:59, 0 of 1 keyframes kept
  16. Shot 15, 3:59 to 4:25, 0 of 1 keyframes kept
  17. Shot 16, 4:25 to 4:52, 0 of 1 keyframes kept
  18. Shot 17, 4:52 to 5:19, 0 of 1 keyframes kept
  19. Shot 18, 5:19 to 5:45, 0 of 1 keyframes kept
  20. Shot 19, 5:45 to 6:12, 0 of 1 keyframes kept
  21. Shot 20, 6:12 to 6:41, 0 of 1 keyframes kept
  22. Shot 21, 6:41 to 7:10, 0 of 1 keyframes kept
  23. Shot 22, 7:10 to 7:39, 0 of 1 keyframes kept
  24. Shot 23, 7:39 to 8:08, 0 of 1 keyframes kept
  25. Shot 24, 8:08 to 8:13, 1 of 1 keyframes kept
  26. Shot 25, 8:13 to 8:28, 0 of 1 keyframes kept
  27. Shot 26, 8:28 to 8:37, 1 of 1 keyframes kept
  28. Shot 27, 8:37 to 8:45, 0 of 1 keyframes kept
  29. Shot 28, 8:45 to 8:54, 1 of 1 keyframes kept
  30. Shot 29, 8:54 to 9:03, 0 of 1 keyframes kept
  31. Shot 30, 9:03 to 9:12, 0 of 1 keyframes kept
  32. Shot 31, 9:12 to 9:14, 0 of 1 keyframes kept
  33. Shot 32, 9:14 to 9:40, 1 of 1 keyframes kept
  34. Shot 33, 9:40 to 10:05, 0 of 1 keyframes kept
  35. Shot 34, 10:05 to 10:31, 0 of 1 keyframes kept
  36. Shot 35, 10:31 to 10:57, 1 of 1 keyframes kept
  37. Shot 36, 10:57 to 11:27, 1 of 1 keyframes kept
  38. Shot 37, 11:27 to 11:56, 0 of 1 keyframes kept
  39. Shot 38, 11:56 to 12:00, 0 of 1 keyframes kept
  40. Shot 39, 12:00 to 12:38, 0 of 1 keyframes kept
  41. Shot 40, 12:38 to 13:11, 0 of 1 keyframes kept
  42. Shot 41, 13:11 to 13:44, 0 of 1 keyframes kept
  43. Shot 42, 13:44 to 13:51, 1 of 1 keyframes kept
  44. Shot 43, 13:51 to 14:20, 0 of 1 keyframes kept
  45. Shot 44, 14:20 to 14:49, 1 of 1 keyframes kept
  46. Shot 45, 14:49 to 15:18, 0 of 1 keyframes kept
  47. Shot 46, 15:18 to 15:19, 0 of 1 keyframes kept
  48. Shot 47, 15:19 to 15:21, 1 of 1 keyframes kept
  49. Shot 48, 15:21 to 15:46, 0 of 1 keyframes kept
  50. Shot 49, 15:46 to 16:04, 0 of 1 keyframes kept
  51. Shot 50, 16:04 to 16:09, 1 of 1 keyframes kept
  52. Shot 51, 16:09 to 16:43, 1 of 1 keyframes kept
  53. Shot 52, 16:43 to 16:52, 1 of 1 keyframes kept
  54. Shot 53, 16:52 to 17:00, 0 of 1 keyframes kept
  55. Shot 54, 17:00 to 17:11, 0 of 1 keyframes kept
  56. Shot 55, 17:11 to 17:13, 0 of 1 keyframes kept
  57. Shot 56, 17:13 to 17:15, 0 of 1 keyframes kept
  58. Shot 57, 17:15 to 17:17, 0 of 1 keyframes kept
  59. Shot 58, 17:17 to 17:23, 1 of 1 keyframes kept
  60. Shot 59, 17:23 to 17:25, 0 of 1 keyframes kept
  61. Shot 60, 17:25 to 17:50, 1 of 1 keyframes kept
  62. Shot 61, 17:50 to 17:52, 1 of 1 keyframes kept
  63. Shot 62, 17:52 to 17:59, 1 of 1 keyframes kept
  64. Shot 63, 17:59 to 18:01, 0 of 1 keyframes kept
  65. Shot 64, 18:01 to 18:07, 1 of 1 keyframes kept
  66. Shot 65, 18:07 to 18:09, 0 of 1 keyframes kept
  67. Shot 66, 18:09 to 18:12, 1 of 1 keyframes kept
  68. Shot 67, 18:12 to 18:38, 1 of 1 keyframes kept
  69. Shot 68, 18:38 to 18:59, 1 of 1 keyframes kept
  70. Shot 69, 18:59 to 19:03, 1 of 1 keyframes kept
  71. Shot 70, 19:03 to 19:10, 0 of 1 keyframes kept
  72. Shot 71, 19:10 to 19:13, 0 of 1 keyframes kept
  73. Shot 72, 19:13 to 19:20, 0 of 1 keyframes kept
  74. Shot 73, 19:20 to 19:22, 1 of 1 keyframes kept
  75. Shot 74, 19:22 to 19:35, 0 of 1 keyframes kept
  76. Shot 75, 19:35 to 19:36, 0 of 1 keyframes kept
  77. Shot 76, 19:36 to 19:39, 0 of 1 keyframes kept
  78. Shot 77, 19:39 to 19:56, 0 of 1 keyframes kept
  79. Shot 78, 19:56 to 20:14, 0 of 1 keyframes kept
  80. Shot 79, 20:14 to 20:41, 1 of 1 keyframes kept
  81. Shot 80, 20:41 to 21:08, 1 of 1 keyframes kept
  82. Shot 81, 21:08 to 21:35, 1 of 1 keyframes kept
  83. Shot 82, 21:35 to 22:03, 0 of 1 keyframes kept
  84. Shot 83, 22:03 to 22:30, 1 of 1 keyframes kept
  85. Shot 84, 22:30 to 22:57, 1 of 1 keyframes kept
  86. Shot 85, 22:57 to 23:24, 0 of 1 keyframes kept
  87. Shot 86, 23:24 to 23:51, 1 of 1 keyframes kept
  88. Shot 87, 23:51 to 24:18, 0 of 1 keyframes kept
  89. Shot 88, 24:18 to 24:21, 1 of 1 keyframes kept
  90. Shot 89, 24:21 to 24:23, 1 of 1 keyframes kept
  91. Shot 90, 24:23 to 24:25, 0 of 1 keyframes kept
  92. Shot 91, 24:25 to 24:51, 0 of 1 keyframes kept
  93. Shot 92, 24:51 to 25:16, 1 of 1 keyframes kept
  94. Shot 93, 25:16 to 25:41, 1 of 1 keyframes kept
  95. Shot 94, 25:41 to 26:07, 1 of 1 keyframes kept
  96. Shot 95, 26:07 to 26:32, 1 of 1 keyframes kept
  97. Shot 96, 26:32 to 26:57, 1 of 1 keyframes kept
  98. Shot 97, 26:57 to 27:23, 1 of 1 keyframes kept
  99. Shot 98, 27:23 to 27:48, 0 of 1 keyframes kept
  100. Shot 99, 27:48 to 28:02, 1 of 1 keyframes kept
  101. Shot 100, 28:02 to 28:02, 0 of 1 keyframes kept

101 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
361
whisperx 361
chunks
50
from 361 cues
keyframes
44
kept of 101 captured
frames with text
44
581 lines read
chapters
11
from the source metadata
keyframe bytes
10.4 MB
word timings on 361 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 14:05 2m 16s
stt done 2026-08-10 14:08 30s
chunk done 2026-08-10 14:08 0s
text_embed done 2026-08-10 19:49 1s
keyframe done 2026-08-10 14:08 2m 36s
ocr done 2026-08-10 14:11 13s
frame_embed done 2026-08-10 19:49 8s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 668.7

    1. Al Engineer0.93
    2. EUROPE1.00
  • 0:06 #1 done2 line(s)

    shot 1·sharpness 819.1

    1. PRESENTING SPONSOR0.99
    2. Google DeepMind1.00
  • 0:11 #2 done3 line(s)

    shot 2·sharpness 939.2

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.95
    3. WorkOS OpenAI0.96
  • 0:16 #3 done5 line(s)

    shot 3·sharpness 410.6

    1. Text Diffusion1.00
    2. Speaker: Brendan O'Donoghue0.99
    3. irector of Research, Google DeepMind1.00
    4. AlEngineer0.99
    5. EUROPE1.00
  • 0:54 #4 done13 line(s)

    shot 4·sharpness 1641.0

    1. Image diffusion1.00
    2. Adding noise1.00
    3. ***0.73
    4. AIE1.00
    5. 1.00
    6. Denoising1.00
    7. 1.00
    8. 1.00
    9. 1.00
    10. Google DeepMind1.00
    11. Engineering the future of Al1.00
    12. AlEngineer0.99
    13. 20260.93
  • 1:06 #5 done19 line(s)

    shot 5·sharpness 1927.6

    1. Text diffusion1.00
    2. The cat was sitting on the table while looking through the window1.00
    3. *★*0.65
    4. 1.00
    5. AIE1.00
    6. 0.99
    7. Adding noise1.00
    8. 1.00
    9. Denoing0.73
    10. The cat was plant on the table while aboveBxo through the window0.99
    11. 1.00
    12. 1.00
    13. 1.00
    14. The cat was plant on the table hんに aboveBxo consider the window0.97
    15. not 上 John plant include a words んに aboveBxo consider 34 puede0.98
    16. Braintrust1.00
    17. WorkOS OpenAI0.98
    18. AlEngineer0.99
    19. 20260.93
  • 1:51 #6 skipped

    shot 6·duplicate of #5

  • 1:52 #7 skipped

    shot 7·duplicate of #4

  • 1:57 #8 skipped

    shot 8·duplicate of #5

  • 2:01 #9 skipped

    shot 9·duplicate of #5

  • 2:04 #10 skipped

    shot 10·duplicate of #5

  • 2:28 #11 done17 line(s)

    shot 11·sharpness 1505.9

    1. Gemini1.00
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. 1.00
    7. 1YEARAGO1.00
    8. 0.57
    9. Google 1/O 25 Keynote0.95
    10. Google0.94
    11. 3ShareShop ± Dowmioad0.79
    12. X p0.54
    13. 口Save0.91
    14. Engineering the future of Al0.99
    15. AlEngineer0.98
    16. EUROPE1.00
    17. 20260.95
  • 2:44 #12 skipped

    shot 12·duplicate of #5

  • 3:19 #13 skipped

    shot 13·duplicate of #5

  • 3:37 #14 skipped

    shot 14·duplicate of #5

  • 4:04 #15 skipped

    shot 15·duplicate of #5

  • 4:44 #16 skipped

    shot 16·duplicate of #5

  • 4:55 #17 skipped

    shot 17·duplicate of #5

  • 5:24 #18 skipped

    shot 18·duplicate of #5

  • 5:54 #19 skipped

    shot 19·duplicate of #5

  • 6:29 #20 skipped

    shot 20·duplicate of #5

  • 7:04 #21 skipped

    shot 21·duplicate of #5

  • 7:14 #22 skipped

    shot 22·duplicate of #5

  • 8:05 #23 skipped

    shot 23·duplicate of #5

Transcript

361 cues· 4,637 words· 24,753 chars

  1. 0:15 So people are still filtering into the room.
  2. 0:18 It's mostly intro stuff for the first couple of slides, so they won't miss anything.
  3. 0:21 OK, welcome, everybody.
  4. 0:22 My name's Brendan.
  5. 0:23 I'm a research scientist at DeepMind.
  6. 0:25 I'm talking today about text diffusion, which is kind of a more forward-looking research area at DeepMind.
  7. 0:32 So you're probably familiar with image and video diffusion, which is kind of state of the art for these modalities right now, where you take ground truth, say image, you add noise to it in training, and then you train a neural network to remove that noise gradually.
  8. 0:48 And then at inference time, you just initialize
  9. 0:51 the picture with pure noise, and then you iteratively refine out the noise to recover back to whatever image or video or audio or whatever you're looking for.
  10. 1:00 And the principle is essentially the same for text, for text diffusion, where you start with a clean sequence of tokens, so a clean sentence or something like that.
  11. 1:10 And then you gradually add noise.
  12. 1:12 You corrupt it somehow.
  13. 1:13 There's lots of different ways to do that.
  14. 1:15 You can do it in a continuous or discrete way.
  15. 1:17 But let's just say discrete for now, which would, in this case, just mean adding random tokens or replacing tokens with other random tokens.
  16. 1:23 And you do that for a bunch of different noise levels.
  17. 1:25 And you train the neural network to try to fill in, to try to correct
  18. 1:30 the mistakes basically in the text.
  19. 1:32 And then at inference time, you initialize the sequence of tokens to just pure noise, like pure random discrete tokens from the vocabulary.
  20. 1:40 And then you iteratively refine through that to fill in the information in the order that the neural network wants to do to recover back to, say, a clean sentence.
  21. 1:49 And then in practice, I showed you some GIFs here of what it looks like for images.
  22. 1:54 You get very similar looking outputs for text, where it kind of starts off all noisy and then gradually fills in the text.
  23. 2:01 And then you get relatively clean outputs at the end.
  24. 2:06 OK, so the team I'm on, we had a research demo release one year ago now called Gemini Diffusion, which was a variant of a Gemini model with text diffusion instead of autoregressive next token generation.
  25. 2:19 And that was like a research preview that was open to about 100k people.
  26. 2:23 uh and you know we're still you know keep keep posted for new new developments in that direction uh upcoming soon um and we did have some good numbers at the time but again it's a year ago which is like you know prehistoric times in this field uh our kind of main comparator model was gemini 2.0 flashlight at the time because that was the architecture we were branching from the text fusion model and we basically had very similar quality across the board there
  27. 2:50 Mostly a little bit of advantage in code, a little bit of disadvantage in some other areas, but kind of relatively similar performance at much better latencies.
  28. 3:00 But again, this is a year ago, so I wouldn't fixate too much on these numbers.
  29. 3:04 OK, so what's the difference between autoregressive generation and diffusion?
  30. 3:06 So in the standard vanilla Gemini, Gemma, GPT, whatever, generation of text, you do this.
  31. 3:15 You have some context that comes in, and you want to generate some response to that.
  32. 3:18 And the model does it one token at a time.
  33. 3:20 So it generates the first token, and then conditional on that generates the next token, and so on.
  34. 3:24 Whereas in diffusion, a context will come in, whatever that is, and it'll initialize, like I mentioned, like a long sequence of tokens, could be hundreds, could be thousands, could be shorter, depends on the model, to be random noise.
  35. 3:37 And then it iteratively refines that canvas to remove the noise over the course of a few denoising steps.
  36. 3:43 So rather than one token at a time, it does the entire block together.
  37. 3:47 but over a couple of iterations.
  38. 3:49 So it's not just one pass, it has multiple passes, but it gets to attend to the future tokens and so on.
  39. 3:55 So it's kind of a different way of generating text.
  40. 4:00 So that obviously has some pros and cons.
  41. 4:02 So the main pro that people really like, and probably is the biggest advantage that text diffusion models have, is that it's just faster inference, it just generates faster tokens per second, because it makes much better use of the hardware, the TPU and the GPU.
  42. 4:15 And I have some slides on that to explain why.
  43. 4:17 But some other advantages are it can do bidirectional attention within this canvas of tokens.
  44. 4:22 So, you know, autoregressive models can only attend to the past.
  45. 4:26 They have causal attention within their, you know, their transformer.
  46. 4:30 Whereas a text fusion model is not restricted, though.
  47. 4:32 They can attend to the future.
  48. 4:34 And that has some interesting properties, like it can do self-corrected generation based on future tokens.
  49. 4:39 So we could do some reasoning, see that I got the answer incorrect, and then go back and fix the reasoning and do it again.
  50. 4:45 I have a demo of that.

Chapters

  1. 0:00 Introduction to Text Diffusion
  2. 1:02 How Text Diffusion Works (Training and Inference)
  3. 2:06 Gemini Diffusion Research Preview
  4. 3:04 Difference Between Autoregressive and Diffusion Models
  5. 4:02 Pros and Cons of Text Diffusion
  6. 6:13 Hardware Efficiency: Why Text Diffusion is Faster
  7. 8:47 Bidirectional Reasoning and Self-Correction
  8. 12:00 Dynamic and Adaptive Computation
  9. 14:26 In-place Text Editing
  10. 16:09 Low Latency Applications and Demos
  11. 20:05 Q&A Session

Open at this second