Videos gHs5ZiY80PM
You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia
Scene timeline
46 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 194
- whisperx 194
- chunks
- 32
- from 194 cues
- keyframes
- 35
- kept of 46 captured
- frames with text
- 35
- 713 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.9 MB
- word timings on 194 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 10:33 | 1m 38s |
stt |
done | — | 2026-08-11 10:35 | 20s |
chunk |
done | — | 2026-08-11 10:35 | 0s |
text_embed |
done | — | 2026-08-11 10:35 | 0s |
keyframe |
done | — | 2026-08-11 10:35 | 1m 35s |
ocr |
done | — | 2026-08-11 10:37 | 18s |
frame_embed |
done | — | 2026-08-11 10:37 | 6s |
Frames, and what the machine read
-
- Al Engineer0.94
- EUROPE1.00
-
- PRESENTING SPONSOR0.99
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.95
- WorkOS OpenAI0.95
-
- NVIDIA0.97
- You Might Not Need1.00
- 50 Diffusion Steps1.00
- Ziv Ilan | Al labs @NVIDIA0.96
- AlEngineer0.99
- EUROPE1.00
-
- NVIDIA0.96
- You Might Not Need1.00
- 50 Diffusion Steps1.00
- Ziv Ilan | Al labs @NVIDIA0.97
- AlEngineer0.99
- EUROPE1.00
-
- NVIDIA0.97
- You Might Not Need1.00
- 50 Diffusion Steps1.00
- Ziv Ilan | Al labs @NVIDIA0.98
- AlEngineer0.99
- EUROPE1.00
- 0261.00
-
- NVIDIA0.97
- You Might Not Need1.00
- 50 Diffusion Steps1.00
- Ziv Ilan | AIl labs @NVIDIA0.95
- AlEngineer0.99
- EUROPE1.00
-
- NVIDIA0.98
- You Might Not Need1.00
- 50 Diffusion Steps1.00
- Ziv Ilan | AI labs @NVIDIA0.93
- AIE1.00
- Google DeepMind0.99
- AlEngineer1.00
- 20260.98
-
- Diffusion models: What & Why1.00
- The dominant architecture for image and video generation0.99
- Generate images/video by iteratively denoising from random noise0.98
- Each step: neural network predicts and removes noise → refines output0.98
- Quality comes from many refinement passes (typically 20-50 steps)0.99
- ★1.00
- AIE0.84
- ★0.97
- Powering today's leading models:0.98
- ★1.00
- FLUX.2 (Black Forest Labs) — SOTA text-to-image, photorealistic0.98
- ★1.00
- ★0.99
- LTX-2.3 (Lightricks) — 22B params, 4K@50fps video + audio0.98
- Wan 2.7, HunyuanVideo, Seedance 2.00.98
- Used for: text-to-image, text-to-video, image editing, 3D generation, inpainting, super-resolution, world simulation,1.00
- scientific modeling1.00
- Pure Noise (t=T)0.99
- Partial Denoise (t=T/2)1.00
- Clean Output (t=0)1.00
- nVIDIA0.94
- Braintrust1.00
- WorkOS OpenAI0.96
- AlEngineer1.00
-
- The Problem: Why 50 Steps Hurts0.99
- Three barriers blocking diffusion from reaching its potential1.00
- Use Case1.00
- The Mature AR1.00
- The Latency Wall1.00
- Enablement1.00
- Ecosystem1.00
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- - 50 steps = 30-600.97
- - Real-time image1.00
- seconds per video.0.97
- editing, live video1.00
- - LLMs: 1 forward1.00
- Real-time needs <10.99
- stylization, interactive1.00
- pass per token.0.98
- second1.00
- world models1.00
- Massive optimization1.00
- - That's a 50x gap —0.97
- - None of these work1.00
- ecosystem (vLLM,1.00
- blocking gaming, live1.00
- at 50 steps. They1.00
- TRT-LLM, SGLang)0.99
- broadcasting,0.99
- need less steps to be0.98
- - Diffusion: 50 passes0.98
- interactive apps1.00
- viable1.00
- per output1.00
- - We need to bring1.00
- diffusion inference to1.00
- the same maturity0.98
- level as LLMs1.00
- NVIDIA0.95
- AlEngineer0.96
- AlEngineer0.99
- EUROPE1.00
- EUROPE1.00
-
- Closing the Gap:1.00
- Quantization + Caching + Distillation1.00
- AIE1.00
- Engineering the future of Al0.99
- AlEngineer1.00
- EUROPE1.00
- 20260.99
-
- Model Compression Strategies1.00
- Range0.92
- Precision1.00
- mantissa1.00
- FP32 3mmmmmmm0.69
- m231.00
- m101.00
- BF161.00
- m71.00
- FP81.00
- m21.00
- AIE1.00
- ★1.00
- FP81.00
- (E5M2)0.86
- (E4M3)0.94
- m31.00
- ★1.00
- Quantization1.00
- Pruning1.00
- ★1.00
- ★1.00
- ★1.00
- 0000000.93
- 000000.93
- 0000000.95
- 0000000.83
- 000000.74
- 080000.88
- Teacher1.00
- back prop.0.98
- Loss1.00
- dataset1.00
- Student1.00
- Dense Matrix1.00
- Sparse Matrix1.00
- Distillation1.00
- Sparsity1.00
- NVIDIA0.94
- Engineering the future of Al1.00
- AlEngineer0.99
- EUROPE1.00
-
- Quantization: Make Each Step Cheaper1.00
- Reduce precision → faster math, less memory, same quality0.99
- Post-training quantization(PTQ)1.00
- Quantization-aware training (QAT)1.00
- Pre-trained1.00
- ★1.00
- Calibration data1.00
- model1.00
- AIE1.00
- ★1.00
- Post-training quantization1.00
- Quantization-aware1.00
- (PTQ)1.00
- training (QAT)1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Pre-trained1.00
- model1.00
- Start with a pre-trained model and1.00
- evaluate it on a calibration dataset.0.97
- introduce quantization ops at various0.99
- Start with a pre-trained model and0.98
- Add QAT ops0.99
- layers1.00
- Gather layer0.99
- statistics1.00
- Calibration data is used to calibrate the0.99
- data.1.00
- model. It can be a subset of training0.99
- Finetune it for a small number of1.00
- epochs.1.00
- Finetune with1.00
- QAT ops0.96
- Calculate dynamic ranges of weights1.00
- Simulates the quantization process0.99
- and activations in the network to1.00
- that occurs during inference.0.99
- Compute1.00
- compute quantization parameters (q-1.00
- params).1.00
- Store q-params1.00
- q-params1.00
- Quantize the network using q-params1.00
- The goal is to learn the q-params1.00
- and run inference1.00
- which can help to reduce the accuracy0.99
- drop between the quantized model1.00
- and pre-trained model.1.00
- Quantize model1.00
- Quantize model1.00
- for inference0.99
- NVIDIA0.93
- AlEngineer0.96
- AlEngineer0.99
- EUROPE1.00
-
- Quantization: Performance & Quality1.00
- FLUX.2[dev] on Blackwell | Quality preserved at every precision level0.99
- BF161.00
- NVFP41.00
- BF161.00
- NVFP41.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- DIY with TensorRT-LLM - Visual Gen0.98
- Run a pre-quantized ckpt from HF0.99
- # With NVFP4 quantization0.99
- black-forest-labs/FLUX.2-dev-NVFP41.00
- python visual_gen_flux.py --model_path black-forest-0.99
- labs/FLUX.2-dev --prompt "A cat" --linear_type trtllm-nvfp40.97
- nVIDIA0.98
- Engineering the future of Al0.99
- AlEngineer0.96
-
- Architecture Optimization:1.00
- Skip What You Don't Need1.00
- AIE1.00
- ★1.00
- ★0.99
- AlEngineer0.97
- AlEngineer1.00
- EUROPE1.00
- EUROPE1.00
-
- Caching: Skip Redundant Computation1.00
- If the input barely changed, reuse the previous output0.99
- Common example - TeaCache (CVPR 2025) — Intra-request caching0.97
- Monitors timestep embedding changes between consecutive steps0.99
- When change is small → reuse cached transformer output, skip computation0.99
- Uses polynomial fitting to predict when caching is safe0.99
- AIE1.00
- ★0.99
- Result: ~2x speedup, skips ~16 of 50 steps with <0.07% quality loss0.97
- ★1.00
- Integrated in TRT-LLM VisualGen, ComfyUI, vLLM-Omni -> --enable_teacache + --0.99
- ★1.00
- ★1.00
- ★1.00
- teacache_thresh1.00
- FLUX.2-dev Inference Speedup1.00
- 10.2×x0.87
- 0F16 (0200)0.89
- nVIDIA0.96
- Engineering the future of Al0.99
- AlEngineer1.00
-
- Caching: Skip Redundant Computation1.00
- If the input barely changed, reuse the previous output0.99
- Common example - TeaCache (CVPR 2025) — Intra-request caching0.98
- Monitors timestep embedding changes between consecutive steps0.99
- When change is small → reuse cached transformer output, skip computation0.99
- ★1.00
- Uses polynomial fitting to predict when caching is safe0.99
- AIE1.00
- ★0.99
- Result: ~2x speedup, skips ~16 of 50 steps with <0.07% quality loss0.98
- ★1.00
- Integrated in TRT-LLM VisualGen, ComfyUI, vLLM-Omni -> --enable_teacache + --0.99
- ★1.00
- ★1.00
- ★0.99
- teacache_thresh1.00
- FLUX.2-dev Inference Speedup1.00
- 0F16 (0200)0.91
- nVIDIA0.96
- Engineering the future of Al0.99
- AlEngineer0.98
Transcript
194 cues· 2,870 words· 15,715 chars
- 0:15 Okay.
- 0:17 Hope everyone are awake after lunch and nice to meet you all.
- 0:21 I'm Ziv.
- 0:22 I'm in the AI Labs team in NVIDIA based out of Paris and working with different frontier model builders across a lot of domains.
- 0:32 Of course, Diffusion is one of them and we'll hear about a couple of examples of work we do with them.
- 0:38 We have only 20 minutes, so obviously it's kind of a mix between going deep and going very high level.
- 0:44 Each of these topics will probably be a full day or full conference to cover, so I'll try to cover everything I can within this timeframe, but feel free to reach out afterwards, either through LinkedIn or I'll stay here a few minutes afterwards.
- 1:01 Without further ado, the fusion models, I assume everyone here knows about it.
- 1:05 Anyone doesn't know what video gen, image gen, how they work on a high level, denoising?
- 1:11 Perfect, OK.
- 1:12 The idea is that, of course, unlike autoregressive architectures, LLM, the idea is that you have a lot of iterations to denoise the image or the video, usually between 20 to 50 steps.
- 1:26 And we see, I think in the last year, an influx of very good high quality models, both for image generation, whether it's flux to video generation, LTX2, 1,
- 1:41 Google with the nano banana and the later generations.
- 1:45 And we do see a lot of more practical use cases for that.
- 1:48 And the main challenge, once we have some interesting use cases, is how to make it actually usable, right?
- 1:54 We know that
- 1:56 It's cool to generate videos or to generate images.
- 1:59 But now if we talk about a developer context or enterprise context, this should be fast.
- 2:06 We want it to be mature.
- 2:07 We want it to be scalable.
- 2:09 And these are usually the challenges that are hard to solve as this ecosystem is not as mature as the autoregressive LLM, VLM ecosystem.
- 2:19 So we try to borrow a lot of the concepts we see work very well for LLM.
- 2:24 and we gradually kind of distill them, if I'll steal this terminology, into the world of diffusion models, okay?
- 2:32 We'll cover a few of the topics here, but again, it's every day we see more and more research in this domain, and I expect this world to be even more mature in the next AI engineer.
- 2:44 Use case enablement, real-time image, real-time video is obviously the holy grail.
- 2:49 Imagine how many new use cases where it's world models for robotics, for computer games, for content generation.
- 2:58 It opens a lot of new avenues for companies and developers to use it.
- 3:03 And the big challenge together is, of course, the latency.
- 3:05 It takes a lot of time to get a first image and then to obviously get high quality, if we talk about 1080p or 720p content out there.
- 3:15 And to bridge this gap, I'll talk about three concepts.
- 3:18 Of course, it's not covering all the ways you can optimize your video gen, image gen models, but I'll touch on quantization, caching, and distillation.
- 3:29 It's not necessarily the order you'll deploy it yourself, okay?
- 3:32 Usually you'll start at distillation, then do some quantization, then some caching, but I started from the simple to the more complex, okay?
- 3:42 Simple is usually quantization.
- 3:44 For those of you who tried it in LLMs, concepts are quite similar.
- 3:49 Then we'll talk about caching and distillation.
- 3:52 When we talk about quantization, we have two approaches, post-training quantization and quantization-aware training.
- 4:00 I'd say in many cases, of course, we do want to use the more simple approach like PTQ, but we know that at least to maintain the image
- 4:09 quality the video quality it's a little bit more complex with the diffusion models okay and we also know that um these type of models are more attention heavy which means that the impact of doing quantization is not as impactful as the llms vlms but it is still quite a low hanging fruit when we talk about taking advantage of the more advanced
- 4:34 features of Blackwell, for example, and more modern compute.
- 4:38 In this example, just the work we did with Black Forest Labs on Flux2, you can see that using usually dynamic quantization, we can use static, which means that we compute all the range of all the different parameters upfront.
- 4:54 deploy it, and use this static range for the quantization.
- 4:57 In this case, we use dynamic approach, which means that some of the range will be computed on the fly.
- 5:04 Again, to make sure that the distribution is in line with the different data distribution that you'll probably want to use when running these models.
- 5:13 It's something that you can either do it yourself.
- 5:15 We recently released a good example in our tiered TLLM visual gen repository.
- 5:20 Open source, you can start using it and see how it goes.
- 5:24 What we also try to do to, again, help the community to adopt it is also to help our partners to do pre-quantized checkpoints.
- 5:34 So you can just go to Hugging Face, load the quantized checkpoint, and start using it.
- 5:39 If you don't need to fine tune or to do some lower adapters afterwards, it's something, again, it's quite handy.
loading