Videos J4_jCrTxMkk
Compression at the Edge — NVIDIA, Unsloth, HuggingFace, Ollama
Scene timeline
162 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 531
- whisperx 531
- chunks
- 79
- from 531 cues
- keyframes
- 137
- kept of 162 captured
- frames with text
- 137
- 239 lines read
- chapters
- 18
- from the source metadata
- keyframe bytes
- 23.0 MB
- word timings on 531 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:14 | 0s |
stt |
done | — | 2026-08-08 23:51 | 56s |
chunk |
done | — | 2026-08-08 23:52 | 0s |
text_embed |
done | — | 2026-08-10 19:34 | 1s |
keyframe |
done | — | 2026-08-08 23:52 | 8m 29s |
ocr |
done | — | 2026-08-09 00:05 | 21s |
frame_embed |
done | — | 2026-08-10 19:34 | 23s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.97
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- 2:25PM-3:10PM1.00
- Compression at the Edge1.00
- Chris Alexiuk0.98
- Daniel Han0.97
- Asma Beevi1.00
- Merve Noyan1.00
- Parth Sareen1.00
- Member of1.00
- Sr Product Research Manager0.99
- Co-founder1.00
- Deep Learning Algorithms1.00
- Builder1.00
- Technical Staff1.00
- Ollama0.96
- Nemotron1.00
- NVIDIA.0.97
- unsloth1.00
- NVIDIA.0.96
- Hugging Face1.00
-
- nsloth0.98
-
- HUGGING ACE0.82
- nsloth0.98
- Norld's Fair0.94
-
- HUGGING.ACE0.94
- nsloth0.99
-
- HUGGING.ACE0.95
-
- HUGGING ACE0.96
-
- HUSGING ACE0.85
- Norld'sFair0.95
-
- sloth0.94
-
- sloth0.97
-
- sloth0.99
-
- sloth1.00
-
- HUGGING ACE0.98
-
- HUGGING ACE0.98
-
- HUGGING TACE0.96
-
- HUGGING FACE0.99
- sloth0.99
-
- isloth0.96
-
- sloth0.99
-
- HUBGING ACE0.95
- isloth0.94
-
- isloth0.97
-
- isloth0.97
Transcript
531 cues· 7,453 words· 40,170 chars
- 0:12 Hello, everybody.
- 0:13 Welcome to Compression at the Edge, the panel that we'll be conducting for the next bit here.
- 0:22 Very nice to meet you all.
- 0:23 I'll be your trusty moderator today.
- 0:25 My name is Chris Oleksiak.
- 0:26 I'm a Product Research Engineer at NVIDIA.
- 0:28 I work on NemoTron.
- 0:29 Let's go.
- 0:30 Okay.
- 0:30 We are joined by Daniel.
- 0:33 Yes.
- 0:34 Hello, everyone.
- 0:35 I'm from Onslaught.
- 0:36 Yeah.
- 0:37 Thanks for coming, everyone.
- 0:38 Excellent and.
- 0:39 Let's go I'm I work as a machine learning engineer at hugging face.
- 0:43 I'm part I work at.
- 0:44 So.
- 0:54 Compression, a big topic.
- 0:57 We're going to set some context, hopefully, in order to kind of launch into this.
- 1:03 So maybe just in each of your own words, if you want to kind of define what you think about compression, let us know how you engage with technology that ultimately is designed to make models that are bigger be a little bit smaller.
- 1:19 That's the general idea.
- 1:22 Maybe we'll just go in reverse over your parts.
- 1:23 If you want to kick us off and what is compression to you?
- 1:26 Yeah, I think honestly with Ollama, and for those of you who are not familiar, Ollama, one of the easiest ways to run local models.
- 1:36 And for us,
- 1:37 honestly, like rose us to popularity was being able to run a larger model on a relatively small machine through quantization, which I'm sure we'll talk a lot about today.
- 1:46 And to me, compression is so, so important because it actually makes these giant models viable for most people.
- 1:55 I think for me, it's just shrinking something without losing information, but there's absolutely zero free lunch.
- 2:02 So at the end of the day, you still spend on something, whether it's latency or quality at the short.
- 2:11 And I kind of agree.
- 2:12 I feel like compression is much more than that definition because it democratizes the models for everyone at edge devices, at your computer.
- 2:21 I'm sure you are all running some Gemma 4.
- 2:25 Quant at the moment, or QN 3.6, those are the hot ones these days.
- 2:30 And like, it just works so well.
- 2:33 So yeah, this is my definition of this.
- 2:36 It democratizes things.
- 2:38 Cool.
- 2:39 So the way I think about is same cost, more intelligence.
- 2:43 So compression accelerates and enables.
- 2:46 To give like a quick example, originally we started with training in FP32, right?
- 2:51 And now we are talking about FP4.
- 2:53 So that is 8X more compression and almost same intelligence without not much degradation, yeah.
- 3:01 Same cost, more intelligence.
- 3:03 Yeah, how we see quantization is you take a big model like GLM 5.2.
- 3:10 It's 1.5 terabytes, which is definitely ginormous.
- 3:13 But then the trick is you can actually quantize it and shrink it to 250 GB.
- 3:18 so you can make it 86% smaller.
- 3:21 But with tricks of quantization, it will not become 86%.
loading
Chapters
- 0:00 Welcome and the panel
- 0:53 What compression means to each of them
- 3:05 GLM 5.2 from 1.5 terabytes to 250 GB
- 4:08 When each of them got the compression bug
- 8:19 QLoRA and finetuning on a T4
- 11:44 86% smaller without being 86% dumber
- 12:46 Why layer importance is so uneven
- 14:26 The super weight: one number, 20% dumber
- 14:51 Evaluating the quantized checkpoints
- 16:37 What NVFP4 actually is
- 17:55 Does compression matter beyond the toaster
- 21:49 Why compress a big model instead of using a small one
- 24:30 Where Ollama fits
- 28:54 How hard NVFP4 is to produce
- 32:46 The cursed era of model architectures
- 35:17 Why linear attention layers break quantization
- 37:22 Where compression goes next
- 43:22 How do you know a quant is any good