Videos fLUtUkqYHnQ
Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI
Scene timeline
52 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 217
- whisperx 217
- chunks
- 35
- from 217 cues
- keyframes
- 38
- kept of 52 captured
- frames with text
- 38
- 752 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 5.2 MB
- word timings on 217 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 01:24 | 1m 49s |
stt |
done | — | 2026-08-10 01:26 | 21s |
chunk |
done | — | 2026-08-10 01:26 | 0s |
text_embed |
done | — | 2026-08-10 19:45 | 0s |
keyframe |
done | — | 2026-08-10 01:26 | 1m 38s |
ocr |
done | — | 2026-08-10 01:28 | 14s |
frame_embed |
done | — | 2026-08-10 19:45 | 7s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- En0.55
- Liquid1.00
- AlEngin0.97
- Everything I Learned Training0.99
- Frontier Small Models1.00
- Maxime Labonne1.00
- Al Engineer Europe, London1.00
- 9 April, 20260.94
- AlEngine0.97
-
- 3B1.00
- LFM2-24B-A2B1.00
- Text1.00
- LFM2-2.6B1.00
- 2.5B1.00
- Text1.00
- *★*0.62
- ★1.00
- Actitaad paters0.63
- 2B1.00
- AIE1.00
- LFM2.5-VL-1.6B1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 1.5B1.00
- LFM2.5-Audio-1.5B1.00
- Audio1.00
- Vision0.99
- Yext0.90
- LFM2-8B-A1B1.00
- LFM2.5-1.2B1.00
- Text1.00
- 1B1.00
- LFM2-700M1.00
- LFM2.5-VL-450MNEW0.99
- 0.5B1.00
- Wision0.93
- LFM2.5-350M1.00
- Text1.00
- θB0.85
- 300M0.98
- 500M0.95
- 1B0.98
- 2B0.98
- 5B1.00
- 10B0.99
- 20B0.98
- 40B0.98
- Total parameters (log scale)1.00
- Braintrust1.00
- WorkOS OpenAI0.97
- AlEngineer0.94
-
- Characteristics of edge model deployment1.00
- 目0.62
- *★*0.63
- AIE1.00
- Memory-bound1.00
- Task-specific1.00
- Latency sensitive1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- <3B parameters0.99
- ≠general-purpose1.00
- Sub-100ms responses0.99
- chatbots1.00
- Edge models are not just scaled-down versions of bigger models0.99
- Edger0.96
- Braintrust1.00
- WorkOSOpenAI0.95
- AlEnai0.98
-
- Characteristics of edge model deployment1.00
- *★★0.62
- ★1.00
- AIE1.00
- ★1.00
- Memory bound0.96
- Task-specific1.00
- Latency sensitive0.97
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Low knowledge capacity1.00
- Easy to train and adapt to1.00
- Fast prefill is a0.99
- new data1.00
- requirement1.00
- Edge models are not just scaled-down versions of bigger models0.99
- Edger0.97
- AlEngineer0.97
- EUROPE1.00
- IEngineer0.91
-
- AIE1.00
- Architecture1.00
- ★1.00
- ★1.00
- Engineering the future of Al1.00
- AlEnginee0.99
-
- Gemma 3 270M (LLM)0.99
- Qwen3.5-0.8B (VLM)0.99
- RMSNorm1.00
- Feedforward1.00
- ★1.00
- AIE1.00
- 18rs0.80
- RMSNorm1.00
- Feedforward1.00
- 2rs0.97
- RMSNorm1.00
- RMSNorm1.00
- 5:1 SWA/GQA1.00
- 3:1 GDN/Gated Attention1.00
- RMSNorm1.00
- RMSNorm0.99
- Embedding1.00
- Embedding1.00
- Engineering the future of Al0.99
- AlEnginee0.96
-
- Effective size = 100M0.99
- Effective size = 600M0.98
- Gemma 3-270M (LLM)0.98
- Qwen3.5-0.8B (VLM)0.99
- 18yrs0.64
- RMSNorm1.00
- Feedforward1.00
- RMSNorm1.00
- RMSI0.99
- Norm1.00
- ★1.00
- ★1.00
- 5:1 SWA/GQA0.97
- ★0.82
- AIE1.00
- RMSNorm1.00
- Feedforward1.00
- ★1.00
- 2yrs0.80
- ★1.00
- ★1.00
- RMSNorm1.00
- 3:1 GDN/Gated Attention1.00
- 63% of1.00
- Embedding1.00
- RMSNorm1.00
- model params1.00
- ← Distillation0.99
- Embedding1.00
- 29% of0.93
- model params0.97
- AlEngineer0.97
- EUROPE1.00
- AlEnginee0.98
-
- LFM2.5-350M (LLM)1.00
- Output text1.00
- Effective size = 287M0.98
- Linear (tied)1.00
- RMSNorm1.00
- AIE1.00
- Feedforward1.00
- ★1.00
- ★0.99
- ★1.00
- 1yrs0.66
- RMSNorm1.00
- 3:1 ShortConv/GQA1.00
- RMSNorm1.00
- ✓ 19% of0.88
- Embedding1.00
- model params1.00
- Input text1.00
- AlEngineer0.96
- EUROPE1.00
- AlEngineer0.98
-
- LFM2.5-350M (LLM)0.97
- On-device profiling1.00
- Output text1.00
- Linear (tied)1.00
- RMSNorm1.00
- ★1.00
- AIE1.00
- ★1.00
- Feedforward1.00
- ★1.00
- ★1.00
- 1rs0.83
- RMSNorm1.00
- RYZEN AI0.96
- AMDA0.89
- Geallaaxy0.69
- S24itra0.63
- 3:1 ShortConv/GQA1.00
- Ryzen HX 3700.99
- Galaxy S24 Ultra1.00
- RMSNorm1.00
- Embedding1.00
- Inference metrics1.00
- Input text1.00
- Engineering the future of Al0.99
- AlEngin0.95
-
- LFM2.5-350M (LLM)1.00
- Cost ratio of operators1.00
- Output text1.00
- M4 Max CPU (decode)1.00
- Linear (tied)0.99
- 2.51.00
- RMSNorm1.00
- 21.00
- ★1.00
- AIE1.00
- Feedforward1.00
- 1.51.00
- ★1.00
- ★1.00
- 1rs0.86
- RMSNorm1.00
- 0.51.00
- 3:1 ShortConv/GQA1.00
- RMSNorm1.00
- ShortConvSWA(Gemma3)GDN(Qwen3.5)GLAGQA1.00
- Embedding1.00
- Input text1.00
- Engineering the future of Al1.00
- NIEnginee0.88
- EUROPE1.00
-
- CPU inference metrics1.00
- Llama.cpp | 4-bit quantization |Input: 2K tokens0.98
- AMD Ryzen0.97
- 2,9911.00
- 3131.00
- 8811.00
- AIE1.00
- ★1.00
- ★1.00
- LFM2.5-350M1.00
- Granite-4.0-350M1.00
- Al Max+ 3950.98
- Granite-4.0-H-350M1.00
- Gemma31BIT1.00
- Qwen3.5-0.8B1.00
- 2,0731.00
- IBM0.94
- 2,2621.00
- YBM0.82
- 1,4510.99
- 9051.00
- 1801.00
- LBM0.56
- 2111.00
- 1131.00
- 1091.00
- 4341.00
- à0.56
- 4491.00
- IBM0.99
- IBM0.99
- 3831.00
- 7261.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Prefill (tok/s)1.00
- Decode (tok/s)1.00
- Memory (MB)1.00
- Higher is better1.00
- Higher is better1.00
- Lower is better0.96
- Qualcomm Snapdragon1.00
- 1,2861.00
- 1881.00
- 1,4091.00
- Gen4 (Samsung0.98
- Galaxy S25 Utra)1.00
- LFM2.5-350M0.99
- IBM0.88
- 9951.00
- 1071.00
- TBM0.89
- 1521.00
- G1.00
- 9991.00
- Granite-4.0-350M1.00
- Granite-4.0-H-350M1.00
- Gemma 31BIT0.96
- 5211.00
- YBM0.97
- 5061.00
- 4711.00
- TBM0.68
- 671.00
- 交0.50
- 741.00
- 4251.00
- à0.82
- 5201.00
- TEM0.67
- 4321.00
- Qwen3.5-0.880.99
- Prefill (tok/s)1.00
- Decode (tok/s)1.00
- Memory (MB)1.00
- Higher is better1.00
- Higher is better0.99
- Lower is better0.99
- AlEngineer0.96
- EUROPE1.00
- Engin0.93
-
- AIE1.00
- Training1.00
- ★1.00
- ★1.00
- AlEngineer0.97
- EUROPE1.00
- AlEngin0.98
-
- Pre-training a 350M model on 28T tokens?0.99
- Optimal D/N1.00
- Optimal N1.00
- Optimal D1.00
- **0.77
- ★1.00
- 1090.87
- ☆女0.88
- Chinchilla (70B)1.00
- LFM2.5 (350M)1.00
- 10110.99
- ★1.00
- ★0.99
- Chinchilla (70B)1.00
- LFM2.5 (350M)1.00
- 10160.86
- ★0.98
- ★0.99
- Chinchilla (70B)0.98
- LFM2.5 (350M)1.00
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Tor ser0.72
- 1070.86
- 10^{50.78
- Hoffmann et al. (2022)0.99
- T2 Approach 2 (Acc)0.93
- T² Approach 1 (NLL.)0.90
- ★1.00
- Paraers0.73
- 10100.75
- T2 Approach 1 (NLL.)0.97
- Hoffmann et al. (2022)0.98
- T2 Approach 2 (Acc0.91
- Tokens0.79
- 10140.84
- 10130.99
- 10{150.69
- Hoffmann et al. (2022)0.99
- T2 Approach 1 (NLL.)0.93
- T2 Approach 2 (Acc)0.94
- 1090.98
- 10120.91
- 1030.91
- ★1.00
- 10110.95
- 1080.87
- 10100.81
- 10^10.86
- 1090.92
- 1070.93
- 1080.96
- 10^{170.76
- 10^{190.85
- 10210.95
- 10{230.89
- 10250.80
- 10170.96
- 10191.00
- 10^{210.85
- 10230.98
- 10^{250.67
- 10170.93
- 10190.99
- 10210.94
- 10^{230.84
- 10{250.91
- Training FLOPs1.00
- Training FLOPs1.00
- Training FLOPs1.00
- Roberts et al. "Test-Time Scaling Makes Overtraining Compute-Optimal." arXiv preprint arXiv:2604.01411, April 2026.0.99
- Engineering the future of Al0.99
- AlEn0.89
-
- More pre-training works, even at the smallest scale!0.99
- LFM2.5-350M1.00
- ●LFM2-350M0.98
- Granite-4.0-H-350M1.00
- Gemma 3 1B IT0.97
- Qwen3.5-0.8B (Instruct)1.00
- *★★0.65
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 401.00
- 201.00
- 30.641.00
- 27.580.99
- 00.64
- 22.321.00
- IBM0.80
- 23.891.00
- 27.411.00
- 251.00
- 501.00
- 00.97
- 40.691.00
- 18.201.00
- 17.221.00
- IBM0.93
- 20.331.00
- 22.870.93
- 201.00
- 401.00
- 32.451.00
- 00.65
- 11.670.99
- 12.441.00
- IBM0.99
- 2.281.00
- G1.00
- 13.831.00
- GPQA Diamond1.00
- IFBench1.00
- CaseReportBench1.00
- 301.00
- 21.861.00
- 201.00
- 18.861.00
- 201.00
- 17.841.00
- 151.00
- 12.291.00
- 13.281.00
- 18.701.00
- 101.00
- 10.821.00
- 13.741.00
- IEM0.80
- 9.361.00
- 12.571.00
- 101.00
- 00.81
- 00.59
- IBM0.68
- 7.170.99
- 5.560.98
- 6.141.00
- 6.431.00
- 6.141.00
- à0.66
- IBM0.73
- BFCLv40.99
- τ2-Bench Telecom0.96
- t²-Bench Retail0.93
- Braintrust1.00
- WorkOs OpenAI0.95
-
- Post-training small vs. big models1.00
- ***0.55
- AIE1.00
- ★1.00
- ★1.00
- Supervised1.00
- Preference1.00
- Reinforcement1.00
- Fine-Tuning1.00
- Alignment1.00
- Learning1.00
- More task-specific = better0.99
- AlEngineer0.97
- EUROPE1.00
-
- Post-training small vs. big models1.00
- ***0.55
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Supervised1.00
- Preference1.00
- Reinforcement1.00
- Fine-Tuning1.00
- Alignment1.00
- Learning1.00
- More task-specific = better1.00
- General improvements1.00
- beyond benchmarks1.00
- AlEngineer0.97
- EUROPE1.00
Transcript
217 cues· 3,289 words· 17,792 chars
- 0:14 Hi, everyone.
- 0:15 My name is Maxime Le Bon.
- 0:17 In this presentation, I want to talk about the lessons I've learned push training small models.
- 0:23 So for context, I work at Liquid AI as head of push training.
- 0:26 At Liquid, we mostly focus on edge models for on-device deployment.
- 0:31 And as you can see here, we have models from
- 0:34 350 million parameters to 24 billion parameters.
- 0:38 So this is very, very small.
- 0:40 And yesterday, we released our new VLM for 50M.
- 0:45 And the week before, we released the new version of the 350M model for text.
- 0:51 So this is what we do.
- 0:52 We work across text, vision, and audio.
- 0:56 And yeah, the models are available on Hugging Face if you want to try them out.
- 1:02 In this presentation, I want to talk about what separates small models and big models.
- 1:08 And there are three main characteristics I want to talk about.
- 1:10 So first of all, the small models, they are memory bound because the hardware is what it is, right?
- 1:17 On the phone, in a car, et cetera.
- 1:20 We can't really use super big models, which is why we try to keep the size quite small.
- 1:25 And because of that, we have low knowledge capacity compared to bigger models.
- 1:29 Then the models are task specific, which is great, because if you have small knowledge capacity, you can at least focus on one thing very well.
- 1:38 And so that means that they are usually not general purpose chatbots like ChatGPT, they are a lot more narrow in terms of focus, and they can do something like summarization to use very, very well.
- 1:49 So that's the second aspect.
- 1:51 And the final one is that it's very latency sensitive.
- 1:55 And that means that you need to have very, very fast throughput.
- 2:00 So all these characteristics are very important.
- 2:02 And we'll see in this presentation how they play with each other and how we can do better.
- 2:07 But the main lesson I want you to retain from this presentation is that small models are not just scaled on versions of bigger models.
- 2:14 They also have their unique challenges and we will see about how we do it in this presentation.
- 2:20 The first thing I want to talk about is the model architecture because there's a lot of interesting things that we can do here for edge models.
- 2:28 I want to first talk about GemR3-270M and Quen 3.5-0.8B.
- 2:34 So these models are the smallest version of their respective family.
- 2:38 And you can see that both of them, they adopt a hybrid architecture.
- 2:42 GemR3 has sliding window attention and GQA hybrid, Quen 3.5.
- 2:48 As an architecture with gated delta net and gated tension, this is great because this is a lot faster.
- 2:55 But what I'm interested in here is actually the embedding layer.
- 2:58 Because if you look at the size of the embedding layer compared to all the parameters of the model, you see that actually Gemma3-270M is mostly an embedding layer.
- 3:08 It's 63% of the total parameters.
- 3:11 And even Quen 3.5-0.80, it's still like 29% of the parameters.
- 3:17 So that's not super efficient, because the effective parameters, the parameters that are really used for reasoning, for knowledge capacity, and all that stuff, are not the embedding parameters.
- 3:29 It's the rest.
- 3:30 So the effective size is actually a lot smaller, and it means that you could squeeze more reasoning and more performance from the same memory footprint.
- 3:40 And the reason why they do that is because they use distillation to train the models, so they distill
- 3:46 these models, like those are the student models, and they have teacher models with a huge vocabulary sizes.
- 3:53 And this is why we have these super big embedding layers.
- 3:58 All right, let's talk about the LFM2 architecture now.
- 4:02 As you can see, the LFM2 architecture is actually not that different in terms of just layers.
- 4:09 We also have a hybrid architecture.
- 4:11 And this time, we have short convolutions and GQA.
- 4:16 And I want to talk a bit about, well, first, you can see that the embedding layer is actually a lot smaller compared to the others.
- 4:23 It's like 90% of the parameters.
loading
Chapters
- 0:00 Start
- 0:14 Introduction to frontier small models at Liquid AI
- 1:02 Characteristics: memory-bound, task-specific, latency-sensitive
- 2:20 Architecture: why large embedding layers are inefficient
- 4:01 LFM2 architecture: using gated short convolutions for speed
- 6:09 LFM 2.5 recipe: 28T tokens and post-training stages
- 8:34 Post-training: SFT, preference alignment, and RL best practices
- 10:43 Identifying "doom loops" in reasoning models
- 11:34 Solutions: mitigating loops via preference alignment and RL
- 15:29 Future focus: using agentic tools to overcome memory limits
- 17:58 Q&A: real-world applications for small vs. large models