read-only demo

Videos _A367W_qvc8

Gemma 4 Deep Dive — Cassidy Hardin, Researcher, Google DeepMind

index_state ready data_status ok

AI Engineer· published 2026-04-27· 0:19:02· en-US· indexed 2026-08-10 19:50

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:49, 1 of 1 keyframes kept
  5. Shot 4, 0:49 to 0:51, 1 of 1 keyframes kept
  6. Shot 5, 0:51 to 1:38, 1 of 1 keyframes kept
  7. Shot 6, 1:38 to 1:53, 1 of 1 keyframes kept
  8. Shot 7, 1:53 to 2:24, 1 of 1 keyframes kept
  9. Shot 8, 2:24 to 2:53, 1 of 1 keyframes kept
  10. Shot 9, 2:53 to 3:22, 1 of 1 keyframes kept
  11. Shot 10, 3:22 to 3:51, 0 of 1 keyframes kept
  12. Shot 11, 3:51 to 4:12, 1 of 1 keyframes kept
  13. Shot 12, 4:12 to 4:21, 1 of 1 keyframes kept
  14. Shot 13, 4:21 to 4:23, 1 of 1 keyframes kept
  15. Shot 14, 4:23 to 5:02, 1 of 1 keyframes kept
  16. Shot 15, 5:02 to 5:39, 1 of 1 keyframes kept
  17. Shot 16, 5:39 to 5:51, 1 of 1 keyframes kept
  18. Shot 17, 5:51 to 6:23, 1 of 1 keyframes kept
  19. Shot 18, 6:23 to 6:47, 1 of 1 keyframes kept
  20. Shot 19, 6:47 to 7:18, 1 of 1 keyframes kept
  21. Shot 20, 7:18 to 7:40, 1 of 1 keyframes kept
  22. Shot 21, 7:40 to 8:09, 1 of 1 keyframes kept
  23. Shot 22, 8:09 to 8:38, 0 of 1 keyframes kept
  24. Shot 23, 8:38 to 9:12, 1 of 1 keyframes kept
  25. Shot 24, 9:12 to 9:39, 1 of 1 keyframes kept
  26. Shot 25, 9:39 to 10:05, 1 of 1 keyframes kept
  27. Shot 26, 10:05 to 10:31, 1 of 1 keyframes kept
  28. Shot 27, 10:31 to 11:01, 1 of 1 keyframes kept
  29. Shot 28, 11:01 to 11:35, 0 of 1 keyframes kept
  30. Shot 29, 11:35 to 12:19, 1 of 1 keyframes kept
  31. Shot 30, 12:19 to 12:44, 1 of 1 keyframes kept
  32. Shot 31, 12:44 to 12:54, 1 of 1 keyframes kept
  33. Shot 32, 12:54 to 12:58, 0 of 1 keyframes kept
  34. Shot 33, 12:58 to 13:27, 0 of 1 keyframes kept
  35. Shot 34, 13:27 to 13:53, 1 of 1 keyframes kept
  36. Shot 35, 13:53 to 14:19, 1 of 1 keyframes kept
  37. Shot 36, 14:19 to 14:45, 0 of 1 keyframes kept
  38. Shot 37, 14:45 to 15:11, 0 of 1 keyframes kept
  39. Shot 38, 15:11 to 15:37, 0 of 1 keyframes kept
  40. Shot 39, 15:37 to 16:09, 0 of 1 keyframes kept
  41. Shot 40, 16:09 to 16:29, 1 of 1 keyframes kept
  42. Shot 41, 16:29 to 16:57, 1 of 1 keyframes kept
  43. Shot 42, 16:57 to 17:23, 1 of 1 keyframes kept
  44. Shot 43, 17:23 to 17:38, 1 of 1 keyframes kept
  45. Shot 44, 17:38 to 18:03, 1 of 1 keyframes kept
  46. Shot 45, 18:03 to 18:36, 1 of 1 keyframes kept
  47. Shot 46, 18:36 to 18:47, 0 of 1 keyframes kept
  48. Shot 47, 18:47 to 19:01, 1 of 1 keyframes kept
  49. Shot 48, 19:01 to 19:02, 0 of 1 keyframes kept

49 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
152
whisperx 152
chunks
32
from 152 cues
keyframes
38
kept of 49 captured
frames with text
38
792 lines read
chapters
11
from the source metadata
keyframe bytes
4.5 MB
word timings on 152 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 14:11 1m 35s
stt done 2026-08-10 14:13 20s
chunk done 2026-08-10 14:13 0s
text_embed done 2026-08-10 19:50 1s
keyframe done 2026-08-10 14:13 1m 32s
ocr done 2026-08-10 14:15 11s
frame_embed done 2026-08-10 19:50 6s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:32 #3 done2 line(s)

    shot 3·sharpness 430.7

    1. Gemma 40.95
    2. AlEn0.87
  • 0:50 #4 done6 line(s)

    shot 4·sharpness 1023.9

    1. 1.00
    2. AIE1.00
    3. Gemma 40.99
    4. 1.00
    5. Google DeepMind1.00
    6. AIEn0.76
  • 1:33 #5 done37 line(s)

    shot 5·sharpness 1924.5

    1. Gemma 4 Models1.00
    2. Feature1.00
    3. E2B1.00
    4. E4B1.00
    5. 26B 4A1.00
    6. 31B1.00
    7. *★★0.57
    8. AIE1.00
    9. 1.00
    10. Architecture1.00
    11. Dense1.00
    12. Dense1.00
    13. MoE (4B Active)1.00
    14. Dense1.00
    15. 1.00
    16. 1.00
    17. 1.00
    18. 1.00
    19. Context Window1.00
    20. 128K1.00
    21. 128K1.00
    22. 256K1.00
    23. 256K1.00
    24. Modalities1.00
    25. Text, Vision, Audio1.00
    26. Text, Vision, Audio1.00
    27. Text, Vision0.98
    28. Text, Vision1.00
    29. Target Hardware1.00
    30. Pixel & Qualcomm1.00
    31. Pixel & Qualcomm1.00
    32. MacBook/ Cloud0.98
    33. V100 32GB1.00
    34. WorkOs OpenAI0.94
    35. Braintrust1.00
    36. AIEr0.88
    37. EUR0.96
  • 1:50 #6 done34 line(s)

    shot 6·sharpness 1200.7

    1. Model Performance VS Size1.00
    2. 14601.00
    3. kimi-k2.5-thinking0.94
    4. 14501.00
    5. 14401.00
    6. 14301.00
    7. AIE1.00
    8. 14201.00
    9. 1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. Eo Core0.66
    14. 14001.00
    15. 14101.00
    16. qwen3.5-27b1.00
    17. 13901.00
    18. 13801.00
    19. 13701.00
    20. 13601.00
    21. gpt-oss-120b0.99
    22. 13501.00
    23. 201.00
    24. 301.00
    25. 40 500.99
    26. 1001.00
    27. 2001.00
    28. 400 500 6000.98
    29. 10001.00
    30. Total Model Size (Billion Parameters)1.00
    31. AlEngineer0.96
    32. EUROPE1.00
    33. AlEn0.88
    34. EURC0.92
  • 2:12 #7 done8 line(s)

    shot 7·sharpness 1225.8

    1. AIE1.00
    2. 1.00
    3. 1.00
    4. Apache 2.0 License0.98
    5. AlEngineer0.97
    6. EUROPE1.00
    7. AIEn0.79
    8. EUR1.00
  • 2:33 #8 done29 line(s)

    shot 8·sharpness 2604.6

    1. Gemma 4 31B: Elite Dense Performance0.99
    2. A state-of-the-art dense model purpose-built for advanced reasoning.1.00
    3. *★★0.59
    4. 1.00
    5. Architecture & Scale0.99
    6. Elite Benchmarks1.00
    7. AIE1.00
    8. 1.00
    9. Wide dense model with 60 layers.0.98
    10. #3 spot on the global Arena Al0.99
    11. 1.00
    12. leaderboard—rivaling models 20x its0.98
    13. 1.00
    14. 1.00
    15. 1.00
    16. size.1.00
    17. Context Capacity1.00
    18. Agentic Foundation1.00
    19. A 256K context length with 1024 token0.99
    20. Purpose-built for autonomous0.99
    21. sliding window, and 5:1 local global layer0.99
    22. workflows with native support for1.00
    23. ratio.1.00
    24. thinking mode, function calling, and1.00
    25. structured JSON output.0.99
    26. AIEngineer0.95
    27. EUROPE1.00
    28. AlEng0.95
    29. EUROI0.92
  • 3:10 #9 done24 line(s)

    shot 9·sharpness 2640.0

    1. Gemma 4 26B4A: Specialized MoE Efficiency0.97
    2. Advanced 128-Expert architecture for hyper-granular domain intelligence.1.00
    3. *★*0.55
    4. The Expert Mix1.00
    5. Active Parameters1.00
    6. AIE1.00
    7. 1.00
    8. Utilizes 128 Total Experts plus 1 Constant1.00
    9. Activates only 3.8B parameters during0.99
    10. 1.00
    11. Shared Expert. 8 active experts are used0.99
    12. any forward pass1.00
    13. 1.00
    14. during inference.1.00
    15. Context Capacity1.00
    16. KV Cache Reuse0.98
    17. A 256K context length with 1024 token0.98
    18. Shared KV Cache across 8 layers to0.99
    19. sliding window, and 5:1 local global layer0.99
    20. significantly reduce memory overhead.1.00
    21. ratio.1.00
    22. Engineering the future of Al1.00
    23. AIEn0.93
    24. EUR0.99
  • 3:26 #10 skipped

    shot 10·duplicate of #9

  • 4:04 #11 done72 line(s)

    shot 11·sharpness 1792.8

    1. Benchmark1.00
    2. Gemma 41.00
    3. 31B IT0.92
    4. Thinking1.00
    5. 26B A4B IT0.98
    6. Gemma 40.99
    7. Thinking0.99
    8. Gemma 41.00
    9. E4B IT0.98
    10. Thinking1.00
    11. Gemma 40.99
    12. Thinking1.00
    13. E2B IT0.96
    14. Gemma 30.99
    15. 27B IT0.96
    16. As of 4/2/260.96
    17. Arena Al (text)0.98
    18. 14521.00
    19. 14411.00
    20. 13651.00
    21. 1.00
    22. ★★0.83
    23. MMMLU1.00
    24. Multilingual Q&A (no tools)0.99
    25. 85.2%1.00
    26. 82.6%1.00
    27. 69.4%1.00
    28. 60.0%1.00
    29. 67.6%1.00
    30. AIE1.00
    31. 1.00
    32. Multimodal reasoning0.98
    33. MMMU Pro1.00
    34. 76.9%1.00
    35. 73.8%1.00
    36. 52.6%1.00
    37. 44.2%0.93
    38. 49.7%0.93
    39. 1.00
    40. ★*0.74
    41. Mathematics (no tools)0.99
    42. AIME 20261.00
    43. 89.2%1.00
    44. 88.3%0.96
    45. 42.5%1.00
    46. 37.5%1.00
    47. 20.8%1.00
    48. Competitive coding1.00
    49. LiveCodeBench v61.00
    50. 80.0%1.00
    51. 77.1%1.00
    52. 52.0%1.00
    53. 44.0%1.00
    54. 29.1%1.00
    55. Scientific Knowledge (no tools)0.96
    56. GPQA Diamond1.00
    57. 84.3%1.00
    58. 82.3%1.00
    59. 58.6%1.00
    60. 43.4%1.00
    61. 42.4%1.00
    62. Agentic tool use1.00
    63. τ2-bench0.97
    64. 86.4%1.00
    65. 85.5%1.00
    66. 57.5%1.00
    67. 29.4%1.00
    68. 6.6%1.00
    69. AlEngineer0.97
    70. EUROPE1.00
    71. AlEng0.92
    72. EURO0.98
  • 4:17 #12 done9 line(s)

    shot 12·sharpness 1041.3

    1. What's new in0.99
    2. 0.96
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. Gemma 40.99
    7. AlEngineer0.97
    8. EUROPE1.00
    9. AlEr0.84
  • 4:22 #13 done8 line(s)

    shot 13·sharpness 1279.8

    1. AIE1.00
    2. 1.00
    3. 1.00
    4. Architecture Deep Dive1.00
    5. AlEngineer0.98
    6. EUROPE1.00
    7. AIEn0.88
    8. EUR1.00
  • 4:58 #14 done31 line(s)

    shot 14·sharpness 2416.4

    1. Token Embedding Layer1.00
    2. Attention1.00
    3. RMSNorm1.00
    4. DECODER BLOCK1.00
    5. Sliding window1.00
    6. 1024 tokens0.99
    7. 5:1 ratio of local to global layers0.99
    8. RMSNorm1.00
    9. 5:1 ratio0.99
    10. 4:1 ratio for the E2B model0.99
    11. Local Attn1.00
    12. or1.00
    13. Global Attn1.00
    14. Sliding context window for local1.00
    15. AIE1.00
    16. layers1.00
    17. RMSNorm1.00
    18. 1.00
    19. Global is always the final layer0.98
    20. +0.99
    21. 1.00
    22. 1.00
    23. (attending to all preceding tokens)1.00
    24. RMSNorm1.00
    25. FFNN1.00
    26. RMSNorm1.00
    27. +0.98
    28. RMSNorm1.00
    29. LM Head1.00
    30. Engineering the future of Al1.00
    31. AIEr0.89
  • 5:17 #15 done25 line(s)

    shot 15·sharpness 1627.0

    1. Regular Attention1.00
    2. Sliding Window Attention0.99
    3. (Global)1.00
    4. (local)0.99
    5. The weather in London is1.00
    6. The weather in London is1.00
    7. *★★0.61
    8. The1.00
    9. The1.00
    10. AIE1.00
    11. weather1.00
    12. weather1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. 1.00
    17. in1.00
    18. in1.00
    19. London1.00
    20. London1.00
    21. is1.00
    22. is1.00
    23. Engineering the future of Al1.00
    24. AlEng0.96
    25. EURO1.00
  • 5:43 #16 done13 line(s)

    shot 16·sharpness 1095.0

    1. Local1.00
    2. AIE1.00
    3. Local1.00
    4. 5:11.00
    5. 1.00
    6. 1.00
    7. 1.00
    8. Local1.00
    9. Local1.00
    10. Global1.00
    11. Engineering the future of Al1.00
    12. AIEn0.87
    13. EUR1.00
  • 5:58 #17 done27 line(s)

    shot 17·sharpness 2104.7

    1. Queries0.99
    2. Keys1.00
    3. Values1.00
    4. 2561.00
    5. Local1.00
    6. AIE1.00
    7. Local1.00
    8. Groups of 2 queries share the same key & value heads0.99
    9. Local1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. 5:11.00
    15. Local1.00
    16. Local1.00
    17. Global1.00
    18. Queries1.00
    19. Keys1.00
    20. Values1.00
    21. Global1.00
    22. 5121.00
    23. Groups of 8 queries share the same key & value heads0.99
    24. AlEngineer0.97
    25. EUROPE1.00
    26. AlEn0.91
    27. EUR0.99
  • 6:31 #18 done12 line(s)

    shot 18·sharpness 1757.8

    1. Grouped Query Attention (GQA)1.00
    2. Queries1.00
    3. 123456781.00
    4. AIE1.00
    5. Keys// Values0.97
    6. 1.00
    7. 0.99
    8. Query (Q)0.95
    9. Key/Value (KV)1.00
    10. AlEngineer0.97
    11. EUROPE1.00
    12. AIEr0.77
  • 7:05 #19 done29 line(s)

    shot 19·sharpness 2234.7

    1. Token Embedding Layer1.00
    2. Mixture of Experts (MoE)1.00
    3. x301.00
    4. RMSNorm1.00
    5. DECODER BLOCK1.00
    6. 1 shared router expert0.99
    7. RMSNorm1.00
    8. 128 total experts1.00
    9. Local Attn0.99
    10. or1.00
    11. Global Attn1.00
    12. 8 activated experts1.00
    13. AIE1.00
    14. Experts are small Feedforward1.00
    15. RMSNorm1.00
    16. 1.00
    17. Neural Networks (FFNN)0.99
    18. +0.98
    19. 1.00
    20. 1.00
    21. RMSNorm1.00
    22. MoE1.00
    23. RMSNorm0.99
    24. +0.97
    25. RMSNorm1.00
    26. LM Head1.00
    27. Engineering the future of Al1.00
    28. AIEr0.87
    29. EUF0.80
  • 7:21 #20 done20 line(s)

    shot 20·sharpness 1490.0

    1. MoE Layers1.00
    2. Constant Shared Expert1.00
    3. ROUTER1.00
    4. FFNN1.00
    5. SoftMax1.00
    6. AIE1.00
    7. 1.00
    8. 1.00
    9. Expert 10.95
    10. Expert 21.00
    11. Expert N0.98
    12. Expert 1271.00
    13. Expert 1281.00
    14. X0.69
    15. Active Experts1.00
    16. Inactive Experts1.00
    17. Shared Parameters0.98
    18. Engineering the future of Al0.99
    19. AlEng0.89
    20. EURO1.00
  • 8:06 #21 done28 line(s)

    shot 21·sharpness 1744.7

    1. Token Embedding Layer0.99
    2. E2B x351.00
    3. E4B x420.96
    4. RMSNorm1.00
    5. DECODER BLOCK1.00
    6. RMSNorm1.00
    7. Dense Effective1.00
    8. Local Attn1.00
    9. or1.00
    10. Global Attn1.00
    11. AIE1.00
    12. RMSNorm1.00
    13. Models1.00
    14. 1.00
    15. +0.98
    16. 0.59
    17. 0.99
    18. E2B and E4B1.00
    19. RMSNorm1.00
    20. FFNN1.00
    21. RMSNorm1.00
    22. +0.97
    23. Per Layer Embeddings1.00
    24. RMSNorm1.00
    25. LM Head0.94
    26. Engineering the future of Al0.98
    27. AlEng0.93
    28. EURO1.00
  • 8:24 #22 skipped

    shot 22·duplicate of #21

  • 8:49 #23 done39 line(s)

    shot 23·sharpness 1601.2

    1. Token Embedding Layer1.00
    2. RMSNorm1.00
    3. DECODER BLOCK0.98
    4. ID1.00
    5. Token1.00
    6. Embedding Vector (Float32)1.00
    7. RMSNorm1.00
    8. <pad>1.00
    9. [0.0012, -0.0451, 0.0892, 0.1124, ...]0.96
    10. Local Attn0.97
    11. or1.00
    12. Global Attn1.00
    13. 1.00
    14. <eos>1.00
    15. [-0.0234, 0.0122, -0.0056, 0.0671,...]0.94
    16. AIE1.00
    17. RMSNorm1.00
    18. 1.00
    19. 262,1441.00
    20. <unused216>1.00
    21. [0.0567, -0.0331, 0.0119, -0.0982,...]0.96
    22. +0.98
    23. 1.00
    24. 1.00
    25. 1.00
    26. RMSNorm1.00
    27. Embedding Size1.00
    28. E2B =1,5360.95
    29. E4B=25601.00
    30. FFNN1.00
    31. RMSNorm1.00
    32. +0.98
    33. Per Layer Embeddings0.98
    34. RMSNorm1.00
    35. LM Head1.00
    36. AlEngineer0.96
    37. EUROPE1.00
    38. AlEngi0.90
    39. EUROP1.00

Transcript

152 cues· 2,858 words· 16,523 chars

  1. 0:14 Hi, everyone.
  2. 0:15 My name is Cassidy, and I'm a researcher at Google DeepMind.
  3. 0:19 Today, I'm really excited to share with you some of the technical improvements and architecture that we have with Gemma 4.
  4. 0:28 Last week, we launched Gemma 4, which is the latest addition to our family of open source models.
  5. 0:34 Gemma 4 brought incredible improvements at a scale that has not been seen before.
  6. 0:41 We have a family of very small models with incredible performance, setting a new precedent for what's possible with small open source models.
  7. 0:51 Gemma 4 comes in four sizes.
  8. 0:54 We have two smaller effective models, which are geared towards on-device applications.
  9. 0:59 These models have been adapted and improved in order to provide incredible performance at a small scale, which are able to run locally on phones, iPads, and laptops.
  10. 1:11 We have two larger models, starting with a 26B mixture of experts model, which is the first ever GEMMA MOE.
  11. 1:19 This model has been adapted to have incredible performance while only requiring 3.9 billion active parameters.
  12. 1:26 And our largest model is our 31B Dense.
  13. 1:29 This has insane performance, a huge improvement upon what existed within GEMMA 3, at a new precedent that hasn't been seen before.
  14. 1:39 Taking a look at our larger models, our 31B and our 26B, these models have both ranked in the top six of all open source models on the LM arena.
  15. 1:54 One of the most exciting improvements and things that we've launched alongside Gemma 4 is the move to an Apache 2.0 license.
  16. 2:01 This was deliberately done in order to make our models more accessible for the everyday developer.
  17. 2:06 You should easily be able to integrate Gemma into your lifecycle of development through initial testing all the way to deployment and building within the Gemma universe.
  18. 2:18 Now let's take a little look at what each of these models are and some of the use cases that we've adapted this for.
  19. 2:25 Starting with our 31B dense model.
  20. 2:27 This is a state of the art multimodal model which has been purposely built for advanced reasoning.
  21. 2:33 This model ranked number three on the global arena for the AI leaderboard.
  22. 2:38 This is outperforming models over 20 times its size.
  23. 2:42 This is a huge improvement.
  24. 2:46 The 31B has a 256K context length, which has been purpose-built for autonomous workflows with native support for thinking, function calling, and structured JSON outputs.
  25. 2:59 We also have a slightly smaller 26B.
  26. 3:03 This 26B is the first edition of a mixture of experts model into the Gemma family.
  27. 3:09 Only requiring 3.8 billion parameters during any forward pass, this model is small and efficient.
  28. 3:16 Utilizing a total of 128 experts while only requiring eight experts during any inference, this is efficient for running while still maintaining some of the incredible performance that we saw with our 31B.
  29. 3:30 On the smaller side, we introduced two effective models.
  30. 3:34 These models are geared towards on-device applications with the additional support of audio.
  31. 3:39 These are vision, text, and audio input models, while remaining being text-only output models.
  32. 3:47 Similarly, we have our effective 2B model.
  33. 3:52 Across a variety of benchmarks, these models are incredible.
  34. 3:56 Looking at our performance across agentic capabilities, coding, multimodal, multilingual, we've truly set a new frontier for what's capable with the Gemma models.
  35. 4:06 This is significantly outperforming everything we had with the Gemma 3 family of models.
  36. 4:13 Now let's take a look at what's actually new in Gemma 4, and what have we done, and how have we actually been able to achieve this incredible performance, starting on the architecture side.
  37. 4:24 We have our standard dense model.
  38. 4:26 This is our 31B as well as our smaller effective 2B and 4B models.
  39. 4:31 We have our standard decoder block.
  40. 4:33 What we've done with GEMMA4 is we've made several improvements within attention.
  41. 4:38 We've introduced a five to one ratio of interleaving local to global layers with our smaller effective 2B having a four to one ratio.
  42. 4:47 This means that within our local layers, we have a sliding window of how many tokens we're attending to.
  43. 4:53 And lastly, with our global layers, we've now ensured that the last layer is always a global layer, meaning that our last layer is attending to all proceeding tokens.
  44. 5:03 In practice, what this looks like is our global layers are attending to every token that is proceeded within this, whereas our local models are only attending to a specific number of proceeding tokens.
  45. 5:15 In our smaller models, we have a sliding window of 512 tokens, while in our larger models, we have a sliding window of 1,024 tokens.
  46. 5:25 This sliding window has provided significant improvements in the efficiency and optimizations of our local layers while still maintaining passing through information to the preceding layers.
  47. 5:37 However, our global layers remain to be quite expensive.
  48. 5:40 Despite this interleaving of local and global layers, all of our global layers are still required to attend to all preceding tokens, which makes it quite memory intensive and expensive to run.
  49. 5:51 And this is where we've looked into introducing grouped query attention.
  50. 5:55 Within our local layers, we grouped together two queries to share the same key and value heads.

Chapters

  1. 0:00 <Untitled Chapter 1>
  2. 0:28 Introduction to the Gemma 4 model family and its four size categories
  3. 1:54 Shift to Apache 2.0 licensing for developer accessibility
  4. 2:25 Deep dive into the 31B dense reasoning and 26B mixture-of-experts (MoE) models
  5. 3:30 Overview of on-device effective models (2B and 4B) with multimodal support
  6. 4:21 Architectural updates: interleaved local/global attention and grouped query attention
  7. 6:51 Explanation of the new MoE architecture (128 experts, 8 active)
  8. 7:44 Implementation of Per Layer Embeddings (PLE) to optimize on-device memory
  9. 11:06 Multimodal advances: variable aspect ratios and resolutions for vision encoders
  10. 16:31 Audio processing enhancements via conformer architecture and audio tokenizers
  11. 18:07 Getting started: self-hosting (Hugging Face, Ollama) and cloud deployment (Vertex AI)

Open at this second