read-only demo

Videos gHs5ZiY80PM

You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia

index_state ready data_status ok

AI Engineer· published 2026-06-16· 0:18:46· en-US· indexed 2026-08-11 10:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:20, 1 of 1 keyframes kept
  5. Shot 4, 0:20 to 0:22, 1 of 1 keyframes kept
  6. Shot 5, 0:22 to 0:25, 1 of 1 keyframes kept
  7. Shot 6, 0:25 to 0:57, 1 of 1 keyframes kept
  8. Shot 7, 0:57 to 0:59, 1 of 1 keyframes kept
  9. Shot 8, 0:59 to 1:49, 1 of 1 keyframes kept
  10. Shot 9, 1:49 to 2:17, 1 of 1 keyframes kept
  11. Shot 10, 2:17 to 2:45, 0 of 1 keyframes kept
  12. Shot 11, 2:45 to 3:13, 0 of 1 keyframes kept
  13. Shot 12, 3:13 to 3:40, 1 of 1 keyframes kept
  14. Shot 13, 3:40 to 3:51, 1 of 1 keyframes kept
  15. Shot 14, 3:51 to 4:37, 1 of 1 keyframes kept
  16. Shot 15, 4:37 to 5:05, 1 of 1 keyframes kept
  17. Shot 16, 5:05 to 5:34, 0 of 1 keyframes kept
  18. Shot 17, 5:34 to 6:03, 0 of 1 keyframes kept
  19. Shot 18, 6:03 to 6:31, 0 of 1 keyframes kept
  20. Shot 19, 6:31 to 6:33, 1 of 1 keyframes kept
  21. Shot 20, 6:33 to 7:01, 1 of 1 keyframes kept
  22. Shot 21, 7:01 to 7:29, 1 of 1 keyframes kept
  23. Shot 22, 7:29 to 7:57, 0 of 1 keyframes kept
  24. Shot 23, 7:57 to 8:25, 0 of 1 keyframes kept
  25. Shot 24, 8:25 to 8:53, 1 of 1 keyframes kept
  26. Shot 25, 8:53 to 9:21, 1 of 1 keyframes kept
  27. Shot 26, 9:21 to 9:49, 1 of 1 keyframes kept
  28. Shot 27, 9:49 to 10:18, 1 of 1 keyframes kept
  29. Shot 28, 10:18 to 10:44, 1 of 1 keyframes kept
  30. Shot 29, 10:44 to 11:10, 0 of 1 keyframes kept
  31. Shot 30, 11:10 to 11:36, 1 of 1 keyframes kept
  32. Shot 31, 11:36 to 12:02, 0 of 1 keyframes kept
  33. Shot 32, 12:02 to 12:28, 0 of 1 keyframes kept
  34. Shot 33, 12:28 to 12:54, 1 of 1 keyframes kept
  35. Shot 34, 12:54 to 13:22, 1 of 1 keyframes kept
  36. Shot 35, 13:22 to 13:51, 1 of 1 keyframes kept
  37. Shot 36, 13:51 to 14:19, 1 of 1 keyframes kept
  38. Shot 37, 14:19 to 14:47, 1 of 1 keyframes kept
  39. Shot 38, 14:47 to 15:16, 0 of 1 keyframes kept
  40. Shot 39, 15:16 to 15:58, 1 of 1 keyframes kept
  41. Shot 40, 15:58 to 16:47, 1 of 1 keyframes kept
  42. Shot 41, 16:47 to 17:13, 1 of 1 keyframes kept
  43. Shot 42, 17:13 to 17:39, 1 of 1 keyframes kept
  44. Shot 43, 17:39 to 18:05, 1 of 1 keyframes kept
  45. Shot 44, 18:05 to 18:31, 1 of 1 keyframes kept
  46. Shot 45, 18:31 to 18:45, 1 of 1 keyframes kept

46 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
194
whisperx 194
chunks
32
from 194 cues
keyframes
35
kept of 46 captured
frames with text
35
713 lines read
chapters
0
from the source metadata
keyframe bytes
5.9 MB
word timings on 194 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 10:33 1m 38s
stt done 2026-08-11 10:35 20s
chunk done 2026-08-11 10:35 0s
text_embed done 2026-08-11 10:35 0s
keyframe done 2026-08-11 10:35 1m 35s
ocr done 2026-08-11 10:37 18s
frame_embed done 2026-08-11 10:37 6s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 667.1

    1. Al Engineer0.94
    2. EUROPE1.00
  • 0:07 #1 done2 line(s)

    shot 1·sharpness 824.4

    1. PRESENTING SPONSOR0.99
    2. Google DeepMind1.00
  • 0:11 #2 done3 line(s)

    shot 2·sharpness 938.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.95
    3. WorkOS OpenAI0.95
  • 0:17 #3 done6 line(s)

    shot 3·sharpness 616.4

    1. NVIDIA0.97
    2. You Might Not Need1.00
    3. 50 Diffusion Steps1.00
    4. Ziv Ilan | Al labs @NVIDIA0.96
    5. AlEngineer0.99
    6. EUROPE1.00
  • 0:21 #4 done6 line(s)

    shot 4·sharpness 575.4

    1. NVIDIA0.96
    2. You Might Not Need1.00
    3. 50 Diffusion Steps1.00
    4. Ziv Ilan | Al labs @NVIDIA0.97
    5. AlEngineer0.99
    6. EUROPE1.00
  • 0:23 #5 done7 line(s)

    shot 5·sharpness 600.1

    1. NVIDIA0.97
    2. You Might Not Need1.00
    3. 50 Diffusion Steps1.00
    4. Ziv Ilan | Al labs @NVIDIA0.98
    5. AlEngineer0.99
    6. EUROPE1.00
    7. 0261.00
  • 0:50 #6 done6 line(s)

    shot 6·sharpness 600.9

    1. NVIDIA0.97
    2. You Might Not Need1.00
    3. 50 Diffusion Steps1.00
    4. Ziv Ilan | AIl labs @NVIDIA0.95
    5. AlEngineer0.99
    6. EUROPE1.00
  • 0:58 #7 done8 line(s)

    shot 7·sharpness 2155.2

    1. NVIDIA0.98
    2. You Might Not Need1.00
    3. 50 Diffusion Steps1.00
    4. Ziv Ilan | AI labs @NVIDIA0.93
    5. AIE1.00
    6. Google DeepMind0.99
    7. AlEngineer1.00
    8. 20260.98
  • 1:38 #8 done24 line(s)

    shot 8·sharpness 3972.8

    1. Diffusion models: What & Why1.00
    2. The dominant architecture for image and video generation0.99
    3. Generate images/video by iteratively denoising from random noise0.98
    4. Each step: neural network predicts and removes noise → refines output0.98
    5. Quality comes from many refinement passes (typically 20-50 steps)0.99
    6. 1.00
    7. AIE0.84
    8. 0.97
    9. Powering today's leading models:0.98
    10. 1.00
    11. FLUX.2 (Black Forest Labs) — SOTA text-to-image, photorealistic0.98
    12. 1.00
    13. 0.99
    14. LTX-2.3 (Lightricks) — 22B params, 4K@50fps video + audio0.98
    15. Wan 2.7, HunyuanVideo, Seedance 2.00.98
    16. Used for: text-to-image, text-to-video, image editing, 3D generation, inpainting, super-resolution, world simulation,1.00
    17. scientific modeling1.00
    18. Pure Noise (t=T)0.99
    19. Partial Denoise (t=T/2)1.00
    20. Clean Output (t=0)1.00
    21. nVIDIA0.94
    22. Braintrust1.00
    23. WorkOS OpenAI0.96
    24. AlEngineer1.00
  • 2:14 #9 done46 line(s)

    shot 9·sharpness 3066.2

    1. The Problem: Why 50 Steps Hurts0.99
    2. Three barriers blocking diffusion from reaching its potential1.00
    3. Use Case1.00
    4. The Mature AR1.00
    5. The Latency Wall1.00
    6. Enablement1.00
    7. Ecosystem1.00
    8. 1.00
    9. 1.00
    10. AIE1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. 1.00
    15. - 50 steps = 30-600.97
    16. - Real-time image1.00
    17. seconds per video.0.97
    18. editing, live video1.00
    19. - LLMs: 1 forward1.00
    20. Real-time needs <10.99
    21. stylization, interactive1.00
    22. pass per token.0.98
    23. second1.00
    24. world models1.00
    25. Massive optimization1.00
    26. - That's a 50x gap —0.97
    27. - None of these work1.00
    28. ecosystem (vLLM,1.00
    29. blocking gaming, live1.00
    30. at 50 steps. They1.00
    31. TRT-LLM, SGLang)0.99
    32. broadcasting,0.99
    33. need less steps to be0.98
    34. - Diffusion: 50 passes0.98
    35. interactive apps1.00
    36. viable1.00
    37. per output1.00
    38. - We need to bring1.00
    39. diffusion inference to1.00
    40. the same maturity0.98
    41. level as LLMs1.00
    42. NVIDIA0.95
    43. AlEngineer0.96
    44. AlEngineer0.99
    45. EUROPE1.00
    46. EUROPE1.00
  • 2:42 #10 skipped

    shot 10·duplicate of #9

  • 2:59 #11 skipped

    shot 11·duplicate of #9

  • 3:27 #12 done7 line(s)

    shot 12·sharpness 1977.3

    1. Closing the Gap:1.00
    2. Quantization + Caching + Distillation1.00
    3. AIE1.00
    4. Engineering the future of Al0.99
    5. AlEngineer1.00
    6. EUROPE1.00
    7. 20260.99
  • 3:46 #13 done42 line(s)

    shot 13·sharpness 2715.2

    1. Model Compression Strategies1.00
    2. Range0.92
    3. Precision1.00
    4. mantissa1.00
    5. FP32 3mmmmmmm0.69
    6. m231.00
    7. m101.00
    8. BF161.00
    9. m71.00
    10. FP81.00
    11. m21.00
    12. AIE1.00
    13. 1.00
    14. FP81.00
    15. (E5M2)0.86
    16. (E4M3)0.94
    17. m31.00
    18. 1.00
    19. Quantization1.00
    20. Pruning1.00
    21. 1.00
    22. 1.00
    23. 1.00
    24. 0000000.93
    25. 000000.93
    26. 0000000.95
    27. 0000000.83
    28. 000000.74
    29. 080000.88
    30. Teacher1.00
    31. back prop.0.98
    32. Loss1.00
    33. dataset1.00
    34. Student1.00
    35. Dense Matrix1.00
    36. Sparse Matrix1.00
    37. Distillation1.00
    38. Sparsity1.00
    39. NVIDIA0.94
    40. Engineering the future of Al1.00
    41. AlEngineer0.99
    42. EUROPE1.00
  • 4:05 #14 done57 line(s)

    shot 14·sharpness 3725.7

    1. Quantization: Make Each Step Cheaper1.00
    2. Reduce precision → faster math, less memory, same quality0.99
    3. Post-training quantization(PTQ)1.00
    4. Quantization-aware training (QAT)1.00
    5. Pre-trained1.00
    6. 1.00
    7. Calibration data1.00
    8. model1.00
    9. AIE1.00
    10. 1.00
    11. Post-training quantization1.00
    12. Quantization-aware1.00
    13. (PTQ)1.00
    14. training (QAT)1.00
    15. 1.00
    16. 1.00
    17. 1.00
    18. 1.00
    19. Pre-trained1.00
    20. model1.00
    21. Start with a pre-trained model and1.00
    22. evaluate it on a calibration dataset.0.97
    23. introduce quantization ops at various0.99
    24. Start with a pre-trained model and0.98
    25. Add QAT ops0.99
    26. layers1.00
    27. Gather layer0.99
    28. statistics1.00
    29. Calibration data is used to calibrate the0.99
    30. data.1.00
    31. model. It can be a subset of training0.99
    32. Finetune it for a small number of1.00
    33. epochs.1.00
    34. Finetune with1.00
    35. QAT ops0.96
    36. Calculate dynamic ranges of weights1.00
    37. Simulates the quantization process0.99
    38. and activations in the network to1.00
    39. that occurs during inference.0.99
    40. Compute1.00
    41. compute quantization parameters (q-1.00
    42. params).1.00
    43. Store q-params1.00
    44. q-params1.00
    45. Quantize the network using q-params1.00
    46. The goal is to learn the q-params1.00
    47. and run inference1.00
    48. which can help to reduce the accuracy0.99
    49. drop between the quantized model1.00
    50. and pre-trained model.1.00
    51. Quantize model1.00
    52. Quantize model1.00
    53. for inference0.99
    54. NVIDIA0.93
    55. AlEngineer0.96
    56. AlEngineer0.99
    57. EUROPE1.00
  • 5:02 #15 done19 line(s)

    shot 15·sharpness 4077.1

    1. Quantization: Performance & Quality1.00
    2. FLUX.2[dev] on Blackwell | Quality preserved at every precision level0.99
    3. BF161.00
    4. NVFP41.00
    5. BF161.00
    6. NVFP41.00
    7. AIE1.00
    8. 1.00
    9. 1.00
    10. 1.00
    11. DIY with TensorRT-LLM - Visual Gen0.98
    12. Run a pre-quantized ckpt from HF0.99
    13. # With NVFP4 quantization0.99
    14. black-forest-labs/FLUX.2-dev-NVFP41.00
    15. python visual_gen_flux.py --model_path black-forest-0.99
    16. labs/FLUX.2-dev --prompt "A cat" --linear_type trtllm-nvfp40.97
    17. nVIDIA0.98
    18. Engineering the future of Al0.99
    19. AlEngineer0.96
  • 5:28 #16 skipped

    shot 16·duplicate of #15

  • 5:46 #17 skipped

    shot 17·duplicate of #15

  • 6:06 #18 skipped

    shot 18·duplicate of #15

  • 6:33 #19 done9 line(s)

    shot 19·sharpness 1957.3

    1. Architecture Optimization:1.00
    2. Skip What You Don't Need1.00
    3. AIE1.00
    4. 1.00
    5. 0.99
    6. AlEngineer0.97
    7. AlEngineer1.00
    8. EUROPE1.00
    9. EUROPE1.00
  • 6:55 #20 done21 line(s)

    shot 20·sharpness 4601.7

    1. Caching: Skip Redundant Computation1.00
    2. If the input barely changed, reuse the previous output0.99
    3. Common example - TeaCache (CVPR 2025) — Intra-request caching0.97
    4. Monitors timestep embedding changes between consecutive steps0.99
    5. When change is small → reuse cached transformer output, skip computation0.99
    6. Uses polynomial fitting to predict when caching is safe0.99
    7. AIE1.00
    8. 0.99
    9. Result: ~2x speedup, skips ~16 of 50 steps with <0.07% quality loss0.97
    10. 1.00
    11. Integrated in TRT-LLM VisualGen, ComfyUI, vLLM-Omni -> --enable_teacache + --0.99
    12. 1.00
    13. 1.00
    14. 1.00
    15. teacache_thresh1.00
    16. FLUX.2-dev Inference Speedup1.00
    17. 10.2×x0.87
    18. 0F16 (0200)0.89
    19. nVIDIA0.96
    20. Engineering the future of Al0.99
    21. AlEngineer1.00
  • 7:20 #21 done21 line(s)

    shot 21·sharpness 4591.1

    1. Caching: Skip Redundant Computation1.00
    2. If the input barely changed, reuse the previous output0.99
    3. Common example - TeaCache (CVPR 2025) — Intra-request caching0.98
    4. Monitors timestep embedding changes between consecutive steps0.99
    5. When change is small → reuse cached transformer output, skip computation0.99
    6. 1.00
    7. Uses polynomial fitting to predict when caching is safe0.99
    8. AIE1.00
    9. 0.99
    10. Result: ~2x speedup, skips ~16 of 50 steps with <0.07% quality loss0.98
    11. 1.00
    12. Integrated in TRT-LLM VisualGen, ComfyUI, vLLM-Omni -> --enable_teacache + --0.99
    13. 1.00
    14. 1.00
    15. 0.99
    16. teacache_thresh1.00
    17. FLUX.2-dev Inference Speedup1.00
    18. 0F16 (0200)0.91
    19. nVIDIA0.96
    20. Engineering the future of Al0.99
    21. AlEngineer0.98
  • 7:38 #22 skipped

    shot 22·duplicate of #20

  • 8:16 #23 skipped

    shot 23·duplicate of #20

Transcript

194 cues· 2,870 words· 15,715 chars

  1. 0:15 Okay.
  2. 0:17 Hope everyone are awake after lunch and nice to meet you all.
  3. 0:21 I'm Ziv.
  4. 0:22 I'm in the AI Labs team in NVIDIA based out of Paris and working with different frontier model builders across a lot of domains.
  5. 0:32 Of course, Diffusion is one of them and we'll hear about a couple of examples of work we do with them.
  6. 0:38 We have only 20 minutes, so obviously it's kind of a mix between going deep and going very high level.
  7. 0:44 Each of these topics will probably be a full day or full conference to cover, so I'll try to cover everything I can within this timeframe, but feel free to reach out afterwards, either through LinkedIn or I'll stay here a few minutes afterwards.
  8. 1:01 Without further ado, the fusion models, I assume everyone here knows about it.
  9. 1:05 Anyone doesn't know what video gen, image gen, how they work on a high level, denoising?
  10. 1:11 Perfect, OK.
  11. 1:12 The idea is that, of course, unlike autoregressive architectures, LLM, the idea is that you have a lot of iterations to denoise the image or the video, usually between 20 to 50 steps.
  12. 1:26 And we see, I think in the last year, an influx of very good high quality models, both for image generation, whether it's flux to video generation, LTX2, 1,
  13. 1:41 Google with the nano banana and the later generations.
  14. 1:45 And we do see a lot of more practical use cases for that.
  15. 1:48 And the main challenge, once we have some interesting use cases, is how to make it actually usable, right?
  16. 1:54 We know that
  17. 1:56 It's cool to generate videos or to generate images.
  18. 1:59 But now if we talk about a developer context or enterprise context, this should be fast.
  19. 2:06 We want it to be mature.
  20. 2:07 We want it to be scalable.
  21. 2:09 And these are usually the challenges that are hard to solve as this ecosystem is not as mature as the autoregressive LLM, VLM ecosystem.
  22. 2:19 So we try to borrow a lot of the concepts we see work very well for LLM.
  23. 2:24 and we gradually kind of distill them, if I'll steal this terminology, into the world of diffusion models, okay?
  24. 2:32 We'll cover a few of the topics here, but again, it's every day we see more and more research in this domain, and I expect this world to be even more mature in the next AI engineer.
  25. 2:44 Use case enablement, real-time image, real-time video is obviously the holy grail.
  26. 2:49 Imagine how many new use cases where it's world models for robotics, for computer games, for content generation.
  27. 2:58 It opens a lot of new avenues for companies and developers to use it.
  28. 3:03 And the big challenge together is, of course, the latency.
  29. 3:05 It takes a lot of time to get a first image and then to obviously get high quality, if we talk about 1080p or 720p content out there.
  30. 3:15 And to bridge this gap, I'll talk about three concepts.
  31. 3:18 Of course, it's not covering all the ways you can optimize your video gen, image gen models, but I'll touch on quantization, caching, and distillation.
  32. 3:29 It's not necessarily the order you'll deploy it yourself, okay?
  33. 3:32 Usually you'll start at distillation, then do some quantization, then some caching, but I started from the simple to the more complex, okay?
  34. 3:42 Simple is usually quantization.
  35. 3:44 For those of you who tried it in LLMs, concepts are quite similar.
  36. 3:49 Then we'll talk about caching and distillation.
  37. 3:52 When we talk about quantization, we have two approaches, post-training quantization and quantization-aware training.
  38. 4:00 I'd say in many cases, of course, we do want to use the more simple approach like PTQ, but we know that at least to maintain the image
  39. 4:09 quality the video quality it's a little bit more complex with the diffusion models okay and we also know that um these type of models are more attention heavy which means that the impact of doing quantization is not as impactful as the llms vlms but it is still quite a low hanging fruit when we talk about taking advantage of the more advanced
  40. 4:34 features of Blackwell, for example, and more modern compute.
  41. 4:38 In this example, just the work we did with Black Forest Labs on Flux2, you can see that using usually dynamic quantization, we can use static, which means that we compute all the range of all the different parameters upfront.
  42. 4:54 deploy it, and use this static range for the quantization.
  43. 4:57 In this case, we use dynamic approach, which means that some of the range will be computed on the fly.
  44. 5:04 Again, to make sure that the distribution is in line with the different data distribution that you'll probably want to use when running these models.
  45. 5:13 It's something that you can either do it yourself.
  46. 5:15 We recently released a good example in our tiered TLLM visual gen repository.
  47. 5:20 Open source, you can start using it and see how it goes.
  48. 5:24 What we also try to do to, again, help the community to adopt it is also to help our partners to do pre-quantized checkpoints.
  49. 5:34 So you can just go to Hugging Face, load the quantized checkpoint, and start using it.
  50. 5:39 If you don't need to fine tune or to do some lower adapters afterwards, it's something, again, it's quite handy.

Open at this second