read-only demo

Videos fLUtUkqYHnQ

Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI

index_state ready data_status ok

AI Engineer· published 2026-04-29· 0:20:13· en-US· indexed 2026-08-10 19:45

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:22, 1 of 1 keyframes kept
  5. Shot 4, 0:22 to 1:00, 1 of 1 keyframes kept
  6. Shot 5, 1:00 to 1:26, 1 of 1 keyframes kept
  7. Shot 6, 1:26 to 1:53, 0 of 1 keyframes kept
  8. Shot 7, 1:53 to 2:19, 1 of 1 keyframes kept
  9. Shot 8, 2:19 to 2:26, 1 of 1 keyframes kept
  10. Shot 9, 2:26 to 2:56, 1 of 1 keyframes kept
  11. Shot 10, 2:56 to 3:27, 1 of 1 keyframes kept
  12. Shot 11, 3:27 to 3:57, 0 of 1 keyframes kept
  13. Shot 12, 3:57 to 4:28, 1 of 1 keyframes kept
  14. Shot 13, 4:28 to 5:00, 1 of 1 keyframes kept
  15. Shot 14, 5:00 to 5:33, 1 of 1 keyframes kept
  16. Shot 15, 5:33 to 6:06, 1 of 1 keyframes kept
  17. Shot 16, 6:06 to 6:10, 1 of 1 keyframes kept
  18. Shot 17, 6:10 to 6:38, 1 of 1 keyframes kept
  19. Shot 18, 6:38 to 7:06, 0 of 1 keyframes kept
  20. Shot 19, 7:06 to 7:34, 0 of 1 keyframes kept
  21. Shot 20, 7:34 to 8:02, 1 of 1 keyframes kept
  22. Shot 21, 8:02 to 8:30, 0 of 1 keyframes kept
  23. Shot 22, 8:30 to 8:57, 1 of 1 keyframes kept
  24. Shot 23, 8:57 to 9:23, 1 of 1 keyframes kept
  25. Shot 24, 9:23 to 9:49, 1 of 1 keyframes kept
  26. Shot 25, 9:49 to 10:16, 0 of 1 keyframes kept
  27. Shot 26, 10:16 to 10:42, 1 of 1 keyframes kept
  28. Shot 27, 10:42 to 11:08, 1 of 1 keyframes kept
  29. Shot 28, 11:08 to 11:34, 1 of 1 keyframes kept
  30. Shot 29, 11:34 to 12:02, 1 of 1 keyframes kept
  31. Shot 30, 12:02 to 12:29, 0 of 1 keyframes kept
  32. Shot 31, 12:29 to 12:56, 1 of 1 keyframes kept
  33. Shot 32, 12:56 to 13:27, 1 of 1 keyframes kept
  34. Shot 33, 13:27 to 13:57, 0 of 1 keyframes kept
  35. Shot 34, 13:57 to 14:27, 1 of 1 keyframes kept
  36. Shot 35, 14:27 to 14:57, 0 of 1 keyframes kept
  37. Shot 36, 14:57 to 15:26, 0 of 1 keyframes kept
  38. Shot 37, 15:26 to 15:52, 1 of 1 keyframes kept
  39. Shot 38, 15:52 to 16:17, 0 of 1 keyframes kept
  40. Shot 39, 16:17 to 16:43, 0 of 1 keyframes kept
  41. Shot 40, 16:43 to 17:08, 0 of 1 keyframes kept
  42. Shot 41, 17:08 to 17:49, 1 of 1 keyframes kept
  43. Shot 42, 17:49 to 17:57, 1 of 1 keyframes kept
  44. Shot 43, 17:57 to 18:01, 1 of 1 keyframes kept
  45. Shot 44, 18:01 to 18:03, 1 of 1 keyframes kept
  46. Shot 45, 18:03 to 18:08, 1 of 1 keyframes kept
  47. Shot 46, 18:08 to 18:36, 1 of 1 keyframes kept
  48. Shot 47, 18:36 to 19:03, 1 of 1 keyframes kept
  49. Shot 48, 19:03 to 19:31, 1 of 1 keyframes kept
  50. Shot 49, 19:31 to 19:58, 1 of 1 keyframes kept
  51. Shot 50, 19:58 to 20:12, 1 of 1 keyframes kept
  52. Shot 51, 20:12 to 20:12, 0 of 1 keyframes kept

52 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
217
whisperx 217
chunks
35
from 217 cues
keyframes
38
kept of 52 captured
frames with text
38
752 lines read
chapters
11
from the source metadata
keyframe bytes
5.2 MB
word timings on 217 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 01:24 1m 49s
stt done 2026-08-10 01:26 21s
chunk done 2026-08-10 01:26 0s
text_embed done 2026-08-10 19:45 0s
keyframe done 2026-08-10 01:26 1m 38s
ocr done 2026-08-10 01:28 14s
frame_embed done 2026-08-10 19:45 7s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:17 #3 done9 line(s)

    shot 3·sharpness 445.6

    1. En0.55
    2. Liquid1.00
    3. AlEngin0.97
    4. Everything I Learned Training0.99
    5. Frontier Small Models1.00
    6. Maxime Labonne1.00
    7. Al Engineer Europe, London1.00
    8. 9 April, 20260.94
    9. AlEngine0.97
  • 0:55 #4 done44 line(s)

    shot 4·sharpness 2013.2

    1. 3B1.00
    2. LFM2-24B-A2B1.00
    3. Text1.00
    4. LFM2-2.6B1.00
    5. 2.5B1.00
    6. Text1.00
    7. *★*0.62
    8. 1.00
    9. Actitaad paters0.63
    10. 2B1.00
    11. AIE1.00
    12. LFM2.5-VL-1.6B1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. 1.00
    17. 1.5B1.00
    18. LFM2.5-Audio-1.5B1.00
    19. Audio1.00
    20. Vision0.99
    21. Yext0.90
    22. LFM2-8B-A1B1.00
    23. LFM2.5-1.2B1.00
    24. Text1.00
    25. 1B1.00
    26. LFM2-700M1.00
    27. LFM2.5-VL-450MNEW0.99
    28. 0.5B1.00
    29. Wision0.93
    30. LFM2.5-350M1.00
    31. Text1.00
    32. θB0.85
    33. 300M0.98
    34. 500M0.95
    35. 1B0.98
    36. 2B0.98
    37. 5B1.00
    38. 10B0.99
    39. 20B0.98
    40. 40B0.98
    41. Total parameters (log scale)1.00
    42. Braintrust1.00
    43. WorkOS OpenAI0.97
    44. AlEngineer0.94
  • 1:03 #5 done20 line(s)

    shot 5·sharpness 2775.6

    1. Characteristics of edge model deployment1.00
    2. 0.62
    3. *★*0.63
    4. AIE1.00
    5. Memory-bound1.00
    6. Task-specific1.00
    7. Latency sensitive1.00
    8. 1.00
    9. 1.00
    10. 1.00
    11. 1.00
    12. <3B parameters0.99
    13. ≠general-purpose1.00
    14. Sub-100ms responses0.99
    15. chatbots1.00
    16. Edge models are not just scaled-down versions of bigger models0.99
    17. Edger0.96
    18. Braintrust1.00
    19. WorkOSOpenAI0.95
    20. AlEnai0.98
  • 1:49 #6 skipped

    shot 6·duplicate of #5

  • 2:06 #7 done22 line(s)

    shot 7·sharpness 3026.3

    1. Characteristics of edge model deployment1.00
    2. *★★0.62
    3. 1.00
    4. AIE1.00
    5. 1.00
    6. Memory bound0.96
    7. Task-specific1.00
    8. Latency sensitive0.97
    9. 1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. Low knowledge capacity1.00
    14. Easy to train and adapt to1.00
    15. Fast prefill is a0.99
    16. new data1.00
    17. requirement1.00
    18. Edge models are not just scaled-down versions of bigger models0.99
    19. Edger0.97
    20. AlEngineer0.97
    21. EUROPE1.00
    22. IEngineer0.91
  • 2:20 #8 done6 line(s)

    shot 8·sharpness 912.2

    1. AIE1.00
    2. Architecture1.00
    3. 1.00
    4. 1.00
    5. Engineering the future of Al1.00
    6. AlEnginee0.99
  • 2:38 #9 done20 line(s)

    shot 9·sharpness 2848.9

    1. Gemma 3 270M (LLM)0.99
    2. Qwen3.5-0.8B (VLM)0.99
    3. RMSNorm1.00
    4. Feedforward1.00
    5. 1.00
    6. AIE1.00
    7. 18rs0.80
    8. RMSNorm1.00
    9. Feedforward1.00
    10. 2rs0.97
    11. RMSNorm1.00
    12. RMSNorm1.00
    13. 5:1 SWA/GQA1.00
    14. 3:1 GDN/Gated Attention1.00
    15. RMSNorm1.00
    16. RMSNorm0.99
    17. Embedding1.00
    18. Embedding1.00
    19. Engineering the future of Al0.99
    20. AlEnginee0.96
  • 3:23 #10 done34 line(s)

    shot 10·sharpness 3156.8

    1. Effective size = 100M0.99
    2. Effective size = 600M0.98
    3. Gemma 3-270M (LLM)0.98
    4. Qwen3.5-0.8B (VLM)0.99
    5. 18yrs0.64
    6. RMSNorm1.00
    7. Feedforward1.00
    8. RMSNorm1.00
    9. RMSI0.99
    10. Norm1.00
    11. 1.00
    12. 1.00
    13. 5:1 SWA/GQA0.97
    14. 0.82
    15. AIE1.00
    16. RMSNorm1.00
    17. Feedforward1.00
    18. 1.00
    19. 2yrs0.80
    20. 1.00
    21. 1.00
    22. RMSNorm1.00
    23. 3:1 GDN/Gated Attention1.00
    24. 63% of1.00
    25. Embedding1.00
    26. RMSNorm1.00
    27. model params1.00
    28. ← Distillation0.99
    29. Embedding1.00
    30. 29% of0.93
    31. model params0.97
    32. AlEngineer0.97
    33. EUROPE1.00
    34. AlEnginee0.98
  • 3:30 #11 skipped

    shot 11·duplicate of #10

  • 4:18 #12 done21 line(s)

    shot 12·sharpness 2422.6

    1. LFM2.5-350M (LLM)1.00
    2. Output text1.00
    3. Effective size = 287M0.98
    4. Linear (tied)1.00
    5. RMSNorm1.00
    6. AIE1.00
    7. Feedforward1.00
    8. 1.00
    9. 0.99
    10. 1.00
    11. 1yrs0.66
    12. RMSNorm1.00
    13. 3:1 ShortConv/GQA1.00
    14. RMSNorm1.00
    15. ✓ 19% of0.88
    16. Embedding1.00
    17. model params1.00
    18. Input text1.00
    19. AlEngineer0.96
    20. EUROPE1.00
    21. AlEngineer0.98
  • 4:32 #13 done26 line(s)

    shot 13·sharpness 3156.2

    1. LFM2.5-350M (LLM)0.97
    2. On-device profiling1.00
    3. Output text1.00
    4. Linear (tied)1.00
    5. RMSNorm1.00
    6. 1.00
    7. AIE1.00
    8. 1.00
    9. Feedforward1.00
    10. 1.00
    11. 1.00
    12. 1rs0.83
    13. RMSNorm1.00
    14. RYZEN AI0.96
    15. AMDA0.89
    16. Geallaaxy0.69
    17. S24itra0.63
    18. 3:1 ShortConv/GQA1.00
    19. Ryzen HX 3700.99
    20. Galaxy S24 Ultra1.00
    21. RMSNorm1.00
    22. Embedding1.00
    23. Inference metrics1.00
    24. Input text1.00
    25. Engineering the future of Al0.99
    26. AlEngin0.95
  • 5:10 #14 done25 line(s)

    shot 14·sharpness 2828.9

    1. LFM2.5-350M (LLM)1.00
    2. Cost ratio of operators1.00
    3. Output text1.00
    4. M4 Max CPU (decode)1.00
    5. Linear (tied)0.99
    6. 2.51.00
    7. RMSNorm1.00
    8. 21.00
    9. 1.00
    10. AIE1.00
    11. Feedforward1.00
    12. 1.51.00
    13. 1.00
    14. 1.00
    15. 1rs0.86
    16. RMSNorm1.00
    17. 0.51.00
    18. 3:1 ShortConv/GQA1.00
    19. RMSNorm1.00
    20. ShortConvSWA(Gemma3)GDN(Qwen3.5)GLAGQA1.00
    21. Embedding1.00
    22. Input text1.00
    23. Engineering the future of Al1.00
    24. NIEnginee0.88
    25. EUROPE1.00
  • 5:53 #15 done83 line(s)

    shot 15·sharpness 2588.3

    1. CPU inference metrics1.00
    2. Llama.cpp | 4-bit quantization |Input: 2K tokens0.98
    3. AMD Ryzen0.97
    4. 2,9911.00
    5. 3131.00
    6. 8811.00
    7. AIE1.00
    8. 1.00
    9. 1.00
    10. LFM2.5-350M1.00
    11. Granite-4.0-350M1.00
    12. Al Max+ 3950.98
    13. Granite-4.0-H-350M1.00
    14. Gemma31BIT1.00
    15. Qwen3.5-0.8B1.00
    16. 2,0731.00
    17. IBM0.94
    18. 2,2621.00
    19. YBM0.82
    20. 1,4510.99
    21. 9051.00
    22. 1801.00
    23. LBM0.56
    24. 2111.00
    25. 1131.00
    26. 1091.00
    27. 4341.00
    28. à0.56
    29. 4491.00
    30. IBM0.99
    31. IBM0.99
    32. 3831.00
    33. 7261.00
    34. 1.00
    35. 1.00
    36. 1.00
    37. 1.00
    38. Prefill (tok/s)1.00
    39. Decode (tok/s)1.00
    40. Memory (MB)1.00
    41. Higher is better1.00
    42. Higher is better1.00
    43. Lower is better0.96
    44. Qualcomm Snapdragon1.00
    45. 1,2861.00
    46. 1881.00
    47. 1,4091.00
    48. Gen4 (Samsung0.98
    49. Galaxy S25 Utra)1.00
    50. LFM2.5-350M0.99
    51. IBM0.88
    52. 9951.00
    53. 1071.00
    54. TBM0.89
    55. 1521.00
    56. G1.00
    57. 9991.00
    58. Granite-4.0-350M1.00
    59. Granite-4.0-H-350M1.00
    60. Gemma 31BIT0.96
    61. 5211.00
    62. YBM0.97
    63. 5061.00
    64. 4711.00
    65. TBM0.68
    66. 671.00
    67. 0.50
    68. 741.00
    69. 4251.00
    70. à0.82
    71. 5201.00
    72. TEM0.67
    73. 4321.00
    74. Qwen3.5-0.880.99
    75. Prefill (tok/s)1.00
    76. Decode (tok/s)1.00
    77. Memory (MB)1.00
    78. Higher is better1.00
    79. Higher is better0.99
    80. Lower is better0.99
    81. AlEngineer0.96
    82. EUROPE1.00
    83. Engin0.93
  • 6:07 #16 done7 line(s)

    shot 16·sharpness 881.2

    1. AIE1.00
    2. Training1.00
    3. 1.00
    4. 1.00
    5. AlEngineer0.97
    6. EUROPE1.00
    7. AlEngin0.98
  • 6:35 #17 done78 line(s)

    shot 17·sharpness 3506.8

    1. Pre-training a 350M model on 28T tokens?0.99
    2. Optimal D/N1.00
    3. Optimal N1.00
    4. Optimal D1.00
    5. **0.77
    6. 1.00
    7. 1090.87
    8. ☆女0.88
    9. Chinchilla (70B)1.00
    10. LFM2.5 (350M)1.00
    11. 10110.99
    12. 1.00
    13. 0.99
    14. Chinchilla (70B)1.00
    15. LFM2.5 (350M)1.00
    16. 10160.86
    17. 0.98
    18. 0.99
    19. Chinchilla (70B)0.98
    20. LFM2.5 (350M)1.00
    21. 1.00
    22. 1.00
    23. AIE1.00
    24. 1.00
    25. 1.00
    26. 1.00
    27. 1.00
    28. Tor ser0.72
    29. 1070.86
    30. 10^{50.78
    31. Hoffmann et al. (2022)0.99
    32. T2 Approach 2 (Acc)0.93
    33. T² Approach 1 (NLL.)0.90
    34. 1.00
    35. Paraers0.73
    36. 10100.75
    37. T2 Approach 1 (NLL.)0.97
    38. Hoffmann et al. (2022)0.98
    39. T2 Approach 2 (Acc0.91
    40. Tokens0.79
    41. 10140.84
    42. 10130.99
    43. 10{150.69
    44. Hoffmann et al. (2022)0.99
    45. T2 Approach 1 (NLL.)0.93
    46. T2 Approach 2 (Acc)0.94
    47. 1090.98
    48. 10120.91
    49. 1030.91
    50. 1.00
    51. 10110.95
    52. 1080.87
    53. 10100.81
    54. 10^10.86
    55. 1090.92
    56. 1070.93
    57. 1080.96
    58. 10^{170.76
    59. 10^{190.85
    60. 10210.95
    61. 10{230.89
    62. 10250.80
    63. 10170.96
    64. 10191.00
    65. 10^{210.85
    66. 10230.98
    67. 10^{250.67
    68. 10170.93
    69. 10190.99
    70. 10210.94
    71. 10^{230.84
    72. 10{250.91
    73. Training FLOPs1.00
    74. Training FLOPs1.00
    75. Training FLOPs1.00
    76. Roberts et al. "Test-Time Scaling Makes Overtraining Compute-Optimal." arXiv preprint arXiv:2604.01411, April 2026.0.99
    77. Engineering the future of Al0.99
    78. AlEn0.89
  • 6:41 #18 skipped

    shot 18·duplicate of #17

  • 7:09 #19 skipped

    shot 19·duplicate of #17

  • 7:59 #20 done79 line(s)

    shot 20·sharpness 2432.8

    1. More pre-training works, even at the smallest scale!0.99
    2. LFM2.5-350M1.00
    3. ●LFM2-350M0.98
    4. Granite-4.0-H-350M1.00
    5. Gemma 3 1B IT0.97
    6. Qwen3.5-0.8B (Instruct)1.00
    7. *★★0.65
    8. 1.00
    9. 1.00
    10. AIE1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. 1.00
    17. 401.00
    18. 201.00
    19. 30.641.00
    20. 27.580.99
    21. 00.64
    22. 22.321.00
    23. IBM0.80
    24. 23.891.00
    25. 27.411.00
    26. 251.00
    27. 501.00
    28. 00.97
    29. 40.691.00
    30. 18.201.00
    31. 17.221.00
    32. IBM0.93
    33. 20.331.00
    34. 22.870.93
    35. 201.00
    36. 401.00
    37. 32.451.00
    38. 00.65
    39. 11.670.99
    40. 12.441.00
    41. IBM0.99
    42. 2.281.00
    43. G1.00
    44. 13.831.00
    45. GPQA Diamond1.00
    46. IFBench1.00
    47. CaseReportBench1.00
    48. 301.00
    49. 21.861.00
    50. 201.00
    51. 18.861.00
    52. 201.00
    53. 17.841.00
    54. 151.00
    55. 12.291.00
    56. 13.281.00
    57. 18.701.00
    58. 101.00
    59. 10.821.00
    60. 13.741.00
    61. IEM0.80
    62. 9.361.00
    63. 12.571.00
    64. 101.00
    65. 00.81
    66. 00.59
    67. IBM0.68
    68. 7.170.99
    69. 5.560.98
    70. 6.141.00
    71. 6.431.00
    72. 6.141.00
    73. à0.66
    74. IBM0.73
    75. BFCLv40.99
    76. τ2-Bench Telecom0.96
    77. t²-Bench Retail0.93
    78. Braintrust1.00
    79. WorkOs OpenAI0.95
  • 8:24 #21 skipped

    shot 21·duplicate of #20

  • 8:49 #22 done14 line(s)

    shot 22·sharpness 2175.8

    1. Post-training small vs. big models1.00
    2. ***0.55
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. Supervised1.00
    7. Preference1.00
    8. Reinforcement1.00
    9. Fine-Tuning1.00
    10. Alignment1.00
    11. Learning1.00
    12. More task-specific = better0.99
    13. AlEngineer0.97
    14. EUROPE1.00
  • 9:15 #23 done17 line(s)

    shot 23·sharpness 2413.8

    1. Post-training small vs. big models1.00
    2. ***0.55
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. 1.00
    7. Supervised1.00
    8. Preference1.00
    9. Reinforcement1.00
    10. Fine-Tuning1.00
    11. Alignment1.00
    12. Learning1.00
    13. More task-specific = better1.00
    14. General improvements1.00
    15. beyond benchmarks1.00
    16. AlEngineer0.97
    17. EUROPE1.00

Transcript

217 cues· 3,289 words· 17,792 chars

  1. 0:14 Hi, everyone.
  2. 0:15 My name is Maxime Le Bon.
  3. 0:17 In this presentation, I want to talk about the lessons I've learned push training small models.
  4. 0:23 So for context, I work at Liquid AI as head of push training.
  5. 0:26 At Liquid, we mostly focus on edge models for on-device deployment.
  6. 0:31 And as you can see here, we have models from
  7. 0:34 350 million parameters to 24 billion parameters.
  8. 0:38 So this is very, very small.
  9. 0:40 And yesterday, we released our new VLM for 50M.
  10. 0:45 And the week before, we released the new version of the 350M model for text.
  11. 0:51 So this is what we do.
  12. 0:52 We work across text, vision, and audio.
  13. 0:56 And yeah, the models are available on Hugging Face if you want to try them out.
  14. 1:02 In this presentation, I want to talk about what separates small models and big models.
  15. 1:08 And there are three main characteristics I want to talk about.
  16. 1:10 So first of all, the small models, they are memory bound because the hardware is what it is, right?
  17. 1:17 On the phone, in a car, et cetera.
  18. 1:20 We can't really use super big models, which is why we try to keep the size quite small.
  19. 1:25 And because of that, we have low knowledge capacity compared to bigger models.
  20. 1:29 Then the models are task specific, which is great, because if you have small knowledge capacity, you can at least focus on one thing very well.
  21. 1:38 And so that means that they are usually not general purpose chatbots like ChatGPT, they are a lot more narrow in terms of focus, and they can do something like summarization to use very, very well.
  22. 1:49 So that's the second aspect.
  23. 1:51 And the final one is that it's very latency sensitive.
  24. 1:55 And that means that you need to have very, very fast throughput.
  25. 2:00 So all these characteristics are very important.
  26. 2:02 And we'll see in this presentation how they play with each other and how we can do better.
  27. 2:07 But the main lesson I want you to retain from this presentation is that small models are not just scaled on versions of bigger models.
  28. 2:14 They also have their unique challenges and we will see about how we do it in this presentation.
  29. 2:20 The first thing I want to talk about is the model architecture because there's a lot of interesting things that we can do here for edge models.
  30. 2:28 I want to first talk about GemR3-270M and Quen 3.5-0.8B.
  31. 2:34 So these models are the smallest version of their respective family.
  32. 2:38 And you can see that both of them, they adopt a hybrid architecture.
  33. 2:42 GemR3 has sliding window attention and GQA hybrid, Quen 3.5.
  34. 2:48 As an architecture with gated delta net and gated tension, this is great because this is a lot faster.
  35. 2:55 But what I'm interested in here is actually the embedding layer.
  36. 2:58 Because if you look at the size of the embedding layer compared to all the parameters of the model, you see that actually Gemma3-270M is mostly an embedding layer.
  37. 3:08 It's 63% of the total parameters.
  38. 3:11 And even Quen 3.5-0.80, it's still like 29% of the parameters.
  39. 3:17 So that's not super efficient, because the effective parameters, the parameters that are really used for reasoning, for knowledge capacity, and all that stuff, are not the embedding parameters.
  40. 3:29 It's the rest.
  41. 3:30 So the effective size is actually a lot smaller, and it means that you could squeeze more reasoning and more performance from the same memory footprint.
  42. 3:40 And the reason why they do that is because they use distillation to train the models, so they distill
  43. 3:46 these models, like those are the student models, and they have teacher models with a huge vocabulary sizes.
  44. 3:53 And this is why we have these super big embedding layers.
  45. 3:58 All right, let's talk about the LFM2 architecture now.
  46. 4:02 As you can see, the LFM2 architecture is actually not that different in terms of just layers.
  47. 4:09 We also have a hybrid architecture.
  48. 4:11 And this time, we have short convolutions and GQA.
  49. 4:16 And I want to talk a bit about, well, first, you can see that the embedding layer is actually a lot smaller compared to the others.
  50. 4:23 It's like 90% of the parameters.

Chapters

  1. 0:00 Start
  2. 0:14 Introduction to frontier small models at Liquid AI
  3. 1:02 Characteristics: memory-bound, task-specific, latency-sensitive
  4. 2:20 Architecture: why large embedding layers are inefficient
  5. 4:01 LFM2 architecture: using gated short convolutions for speed
  6. 6:09 LFM 2.5 recipe: 28T tokens and post-training stages
  7. 8:34 Post-training: SFT, preference alignment, and RL best practices
  8. 10:43 Identifying "doom loops" in reasoning models
  9. 11:34 Solutions: mitigating loops via preference alignment and RL
  10. 15:29 Future focus: using agentic tools to overcome memory limits
  11. 17:58 Q&A: real-world applications for small vs. large models

Open at this second