Videos _A367W_qvc8
Gemma 4 Deep Dive — Cassidy Hardin, Researcher, Google DeepMind
Scene timeline
49 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 152
- whisperx 152
- chunks
- 32
- from 152 cues
- keyframes
- 38
- kept of 49 captured
- frames with text
- 38
- 792 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 4.5 MB
- word timings on 152 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 14:11 | 1m 35s |
stt |
done | — | 2026-08-10 14:13 | 20s |
chunk |
done | — | 2026-08-10 14:13 | 0s |
text_embed |
done | — | 2026-08-10 19:50 | 1s |
keyframe |
done | — | 2026-08-10 14:13 | 1m 32s |
ocr |
done | — | 2026-08-10 14:15 | 11s |
frame_embed |
done | — | 2026-08-10 19:50 | 6s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Gemma 40.95
- AlEn0.87
-
- ★1.00
- AIE1.00
- Gemma 40.99
- ★1.00
- Google DeepMind1.00
- AIEn0.76
-
- Gemma 4 Models1.00
- Feature1.00
- E2B1.00
- E4B1.00
- 26B 4A1.00
- 31B1.00
- *★★0.57
- AIE1.00
- ★1.00
- Architecture1.00
- Dense1.00
- Dense1.00
- MoE (4B Active)1.00
- Dense1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Context Window1.00
- 128K1.00
- 128K1.00
- 256K1.00
- 256K1.00
- Modalities1.00
- Text, Vision, Audio1.00
- Text, Vision, Audio1.00
- Text, Vision0.98
- Text, Vision1.00
- Target Hardware1.00
- Pixel & Qualcomm1.00
- Pixel & Qualcomm1.00
- MacBook/ Cloud0.98
- V100 32GB1.00
- WorkOs OpenAI0.94
- Braintrust1.00
- AIEr0.88
- EUR0.96
-
- Model Performance VS Size1.00
- 14601.00
- kimi-k2.5-thinking0.94
- 14501.00
- 14401.00
- 14301.00
- AIE1.00
- 14201.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Eo Core0.66
- 14001.00
- 14101.00
- qwen3.5-27b1.00
- 13901.00
- 13801.00
- 13701.00
- 13601.00
- gpt-oss-120b0.99
- 13501.00
- 201.00
- 301.00
- 40 500.99
- 1001.00
- 2001.00
- 400 500 6000.98
- 10001.00
- Total Model Size (Billion Parameters)1.00
- AlEngineer0.96
- EUROPE1.00
- AlEn0.88
- EURC0.92
-
- AIE1.00
- ★1.00
- ★1.00
- Apache 2.0 License0.98
- AlEngineer0.97
- EUROPE1.00
- AIEn0.79
- EUR1.00
-
- Gemma 4 31B: Elite Dense Performance0.99
- A state-of-the-art dense model purpose-built for advanced reasoning.1.00
- *★★0.59
- ★1.00
- Architecture & Scale0.99
- Elite Benchmarks1.00
- AIE1.00
- ★1.00
- Wide dense model with 60 layers.0.98
- #3 spot on the global Arena Al0.99
- ★1.00
- leaderboard—rivaling models 20x its0.98
- ★1.00
- ★1.00
- ★1.00
- size.1.00
- Context Capacity1.00
- Agentic Foundation1.00
- A 256K context length with 1024 token0.99
- Purpose-built for autonomous0.99
- sliding window, and 5:1 local global layer0.99
- workflows with native support for1.00
- ratio.1.00
- thinking mode, function calling, and1.00
- structured JSON output.0.99
- AIEngineer0.95
- EUROPE1.00
- AlEng0.95
- EUROI0.92
-
- Gemma 4 26B4A: Specialized MoE Efficiency0.97
- Advanced 128-Expert architecture for hyper-granular domain intelligence.1.00
- *★*0.55
- The Expert Mix1.00
- Active Parameters1.00
- AIE1.00
- ★1.00
- Utilizes 128 Total Experts plus 1 Constant1.00
- Activates only 3.8B parameters during0.99
- ★1.00
- Shared Expert. 8 active experts are used0.99
- any forward pass1.00
- ★1.00
- during inference.1.00
- Context Capacity1.00
- KV Cache Reuse0.98
- A 256K context length with 1024 token0.98
- Shared KV Cache across 8 layers to0.99
- sliding window, and 5:1 local global layer0.99
- significantly reduce memory overhead.1.00
- ratio.1.00
- Engineering the future of Al1.00
- AIEn0.93
- EUR0.99
-
- Benchmark1.00
- Gemma 41.00
- 31B IT0.92
- Thinking1.00
- 26B A4B IT0.98
- Gemma 40.99
- Thinking0.99
- Gemma 41.00
- E4B IT0.98
- Thinking1.00
- Gemma 40.99
- Thinking1.00
- E2B IT0.96
- Gemma 30.99
- 27B IT0.96
- As of 4/2/260.96
- Arena Al (text)0.98
- 14521.00
- 14411.00
- 13651.00
- ★1.00
- ★★0.83
- MMMLU1.00
- Multilingual Q&A (no tools)0.99
- 85.2%1.00
- 82.6%1.00
- 69.4%1.00
- 60.0%1.00
- 67.6%1.00
- AIE1.00
- ★1.00
- Multimodal reasoning0.98
- MMMU Pro1.00
- 76.9%1.00
- 73.8%1.00
- 52.6%1.00
- 44.2%0.93
- 49.7%0.93
- ★1.00
- ★*0.74
- Mathematics (no tools)0.99
- AIME 20261.00
- 89.2%1.00
- 88.3%0.96
- 42.5%1.00
- 37.5%1.00
- 20.8%1.00
- Competitive coding1.00
- LiveCodeBench v61.00
- 80.0%1.00
- 77.1%1.00
- 52.0%1.00
- 44.0%1.00
- 29.1%1.00
- Scientific Knowledge (no tools)0.96
- GPQA Diamond1.00
- 84.3%1.00
- 82.3%1.00
- 58.6%1.00
- 43.4%1.00
- 42.4%1.00
- Agentic tool use1.00
- τ2-bench0.97
- 86.4%1.00
- 85.5%1.00
- 57.5%1.00
- 29.4%1.00
- 6.6%1.00
- AlEngineer0.97
- EUROPE1.00
- AlEng0.92
- EURO0.98
-
- What's new in0.99
- ★0.96
- AIE1.00
- ★1.00
- ★1.00
- Gemma 40.99
- AlEngineer0.97
- EUROPE1.00
- AlEr0.84
-
- AIE1.00
- ★1.00
- ★1.00
- Architecture Deep Dive1.00
- AlEngineer0.98
- EUROPE1.00
- AIEn0.88
- EUR1.00
-
- Token Embedding Layer1.00
- Attention1.00
- RMSNorm1.00
- DECODER BLOCK1.00
- Sliding window1.00
- 1024 tokens0.99
- 5:1 ratio of local to global layers0.99
- RMSNorm1.00
- 5:1 ratio0.99
- 4:1 ratio for the E2B model0.99
- Local Attn1.00
- or1.00
- Global Attn1.00
- Sliding context window for local1.00
- AIE1.00
- layers1.00
- RMSNorm1.00
- ★1.00
- Global is always the final layer0.98
- +0.99
- ★1.00
- ★1.00
- (attending to all preceding tokens)1.00
- RMSNorm1.00
- FFNN1.00
- RMSNorm1.00
- +0.98
- RMSNorm1.00
- LM Head1.00
- Engineering the future of Al1.00
- AIEr0.89
-
- Regular Attention1.00
- Sliding Window Attention0.99
- (Global)1.00
- (local)0.99
- The weather in London is1.00
- The weather in London is1.00
- *★★0.61
- The1.00
- The1.00
- AIE1.00
- weather1.00
- weather1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- in1.00
- in1.00
- London1.00
- London1.00
- is1.00
- is1.00
- Engineering the future of Al1.00
- AlEng0.96
- EURO1.00
-
- Local1.00
- AIE1.00
- Local1.00
- 5:11.00
- ★1.00
- ★1.00
- ★1.00
- Local1.00
- Local1.00
- Global1.00
- Engineering the future of Al1.00
- AIEn0.87
- EUR1.00
-
- Queries0.99
- Keys1.00
- Values1.00
- 2561.00
- Local1.00
- AIE1.00
- Local1.00
- Groups of 2 queries share the same key & value heads0.99
- Local1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 5:11.00
- Local1.00
- Local1.00
- Global1.00
- Queries1.00
- Keys1.00
- Values1.00
- Global1.00
- 5121.00
- Groups of 8 queries share the same key & value heads0.99
- AlEngineer0.97
- EUROPE1.00
- AlEn0.91
- EUR0.99
-
- Grouped Query Attention (GQA)1.00
- Queries1.00
- 123456781.00
- AIE1.00
- Keys// Values0.97
- ★1.00
- ★0.99
- Query (Q)0.95
- Key/Value (KV)1.00
- AlEngineer0.97
- EUROPE1.00
- AIEr0.77
-
- Token Embedding Layer1.00
- Mixture of Experts (MoE)1.00
- x301.00
- RMSNorm1.00
- DECODER BLOCK1.00
- 1 shared router expert0.99
- RMSNorm1.00
- 128 total experts1.00
- Local Attn0.99
- or1.00
- Global Attn1.00
- 8 activated experts1.00
- AIE1.00
- Experts are small Feedforward1.00
- RMSNorm1.00
- ★1.00
- Neural Networks (FFNN)0.99
- +0.98
- ★1.00
- ★1.00
- RMSNorm1.00
- MoE1.00
- RMSNorm0.99
- +0.97
- RMSNorm1.00
- LM Head1.00
- Engineering the future of Al1.00
- AIEr0.87
- EUF0.80
-
- MoE Layers1.00
- Constant Shared Expert1.00
- ROUTER1.00
- FFNN1.00
- SoftMax1.00
- AIE1.00
- ★1.00
- ★1.00
- Expert 10.95
- Expert 21.00
- Expert N0.98
- Expert 1271.00
- Expert 1281.00
- X0.69
- Active Experts1.00
- Inactive Experts1.00
- Shared Parameters0.98
- Engineering the future of Al0.99
- AlEng0.89
- EURO1.00
-
- Token Embedding Layer0.99
- E2B x351.00
- E4B x420.96
- RMSNorm1.00
- DECODER BLOCK1.00
- RMSNorm1.00
- Dense Effective1.00
- Local Attn1.00
- or1.00
- Global Attn1.00
- AIE1.00
- RMSNorm1.00
- Models1.00
- ★1.00
- +0.98
- ★0.59
- ★0.99
- E2B and E4B1.00
- RMSNorm1.00
- FFNN1.00
- RMSNorm1.00
- +0.97
- Per Layer Embeddings1.00
- RMSNorm1.00
- LM Head0.94
- Engineering the future of Al0.98
- AlEng0.93
- EURO1.00
-
- Token Embedding Layer1.00
- RMSNorm1.00
- DECODER BLOCK0.98
- ID1.00
- Token1.00
- Embedding Vector (Float32)1.00
- RMSNorm1.00
- <pad>1.00
- [0.0012, -0.0451, 0.0892, 0.1124, ...]0.96
- Local Attn0.97
- or1.00
- Global Attn1.00
- ★1.00
- <eos>1.00
- [-0.0234, 0.0122, -0.0056, 0.0671,...]0.94
- AIE1.00
- RMSNorm1.00
- ★1.00
- 262,1441.00
- <unused216>1.00
- [0.0567, -0.0331, 0.0119, -0.0982,...]0.96
- +0.98
- ★1.00
- ★1.00
- ★1.00
- RMSNorm1.00
- Embedding Size1.00
- E2B =1,5360.95
- E4B=25601.00
- FFNN1.00
- RMSNorm1.00
- +0.98
- Per Layer Embeddings0.98
- RMSNorm1.00
- LM Head1.00
- AlEngineer0.96
- EUROPE1.00
- AlEngi0.90
- EUROP1.00
Transcript
152 cues· 2,858 words· 16,523 chars
- 0:14 Hi, everyone.
- 0:15 My name is Cassidy, and I'm a researcher at Google DeepMind.
- 0:19 Today, I'm really excited to share with you some of the technical improvements and architecture that we have with Gemma 4.
- 0:28 Last week, we launched Gemma 4, which is the latest addition to our family of open source models.
- 0:34 Gemma 4 brought incredible improvements at a scale that has not been seen before.
- 0:41 We have a family of very small models with incredible performance, setting a new precedent for what's possible with small open source models.
- 0:51 Gemma 4 comes in four sizes.
- 0:54 We have two smaller effective models, which are geared towards on-device applications.
- 0:59 These models have been adapted and improved in order to provide incredible performance at a small scale, which are able to run locally on phones, iPads, and laptops.
- 1:11 We have two larger models, starting with a 26B mixture of experts model, which is the first ever GEMMA MOE.
- 1:19 This model has been adapted to have incredible performance while only requiring 3.9 billion active parameters.
- 1:26 And our largest model is our 31B Dense.
- 1:29 This has insane performance, a huge improvement upon what existed within GEMMA 3, at a new precedent that hasn't been seen before.
- 1:39 Taking a look at our larger models, our 31B and our 26B, these models have both ranked in the top six of all open source models on the LM arena.
- 1:54 One of the most exciting improvements and things that we've launched alongside Gemma 4 is the move to an Apache 2.0 license.
- 2:01 This was deliberately done in order to make our models more accessible for the everyday developer.
- 2:06 You should easily be able to integrate Gemma into your lifecycle of development through initial testing all the way to deployment and building within the Gemma universe.
- 2:18 Now let's take a little look at what each of these models are and some of the use cases that we've adapted this for.
- 2:25 Starting with our 31B dense model.
- 2:27 This is a state of the art multimodal model which has been purposely built for advanced reasoning.
- 2:33 This model ranked number three on the global arena for the AI leaderboard.
- 2:38 This is outperforming models over 20 times its size.
- 2:42 This is a huge improvement.
- 2:46 The 31B has a 256K context length, which has been purpose-built for autonomous workflows with native support for thinking, function calling, and structured JSON outputs.
- 2:59 We also have a slightly smaller 26B.
- 3:03 This 26B is the first edition of a mixture of experts model into the Gemma family.
- 3:09 Only requiring 3.8 billion parameters during any forward pass, this model is small and efficient.
- 3:16 Utilizing a total of 128 experts while only requiring eight experts during any inference, this is efficient for running while still maintaining some of the incredible performance that we saw with our 31B.
- 3:30 On the smaller side, we introduced two effective models.
- 3:34 These models are geared towards on-device applications with the additional support of audio.
- 3:39 These are vision, text, and audio input models, while remaining being text-only output models.
- 3:47 Similarly, we have our effective 2B model.
- 3:52 Across a variety of benchmarks, these models are incredible.
- 3:56 Looking at our performance across agentic capabilities, coding, multimodal, multilingual, we've truly set a new frontier for what's capable with the Gemma models.
- 4:06 This is significantly outperforming everything we had with the Gemma 3 family of models.
- 4:13 Now let's take a look at what's actually new in Gemma 4, and what have we done, and how have we actually been able to achieve this incredible performance, starting on the architecture side.
- 4:24 We have our standard dense model.
- 4:26 This is our 31B as well as our smaller effective 2B and 4B models.
- 4:31 We have our standard decoder block.
- 4:33 What we've done with GEMMA4 is we've made several improvements within attention.
- 4:38 We've introduced a five to one ratio of interleaving local to global layers with our smaller effective 2B having a four to one ratio.
- 4:47 This means that within our local layers, we have a sliding window of how many tokens we're attending to.
- 4:53 And lastly, with our global layers, we've now ensured that the last layer is always a global layer, meaning that our last layer is attending to all proceeding tokens.
- 5:03 In practice, what this looks like is our global layers are attending to every token that is proceeded within this, whereas our local models are only attending to a specific number of proceeding tokens.
- 5:15 In our smaller models, we have a sliding window of 512 tokens, while in our larger models, we have a sliding window of 1,024 tokens.
- 5:25 This sliding window has provided significant improvements in the efficiency and optimizations of our local layers while still maintaining passing through information to the preceding layers.
- 5:37 However, our global layers remain to be quite expensive.
- 5:40 Despite this interleaving of local and global layers, all of our global layers are still required to attend to all preceding tokens, which makes it quite memory intensive and expensive to run.
- 5:51 And this is where we've looked into introducing grouped query attention.
- 5:55 Within our local layers, we grouped together two queries to share the same key and value heads.
loading
Chapters
- 0:00 <Untitled Chapter 1>
- 0:28 Introduction to the Gemma 4 model family and its four size categories
- 1:54 Shift to Apache 2.0 licensing for developer accessibility
- 2:25 Deep dive into the 31B dense reasoning and 26B mixture-of-experts (MoE) models
- 3:30 Overview of on-device effective models (2B and 4B) with multimodal support
- 4:21 Architectural updates: interleaved local/global attention and grouped query attention
- 6:51 Explanation of the new MoE architecture (128 experts, 8 active)
- 7:44 Implementation of Per Layer Embeddings (PLE) to optimize on-device memory
- 11:06 Multimodal advances: variable aspect ratios and resolutions for vision encoders
- 16:31 Audio processing enhancements via conformer architecture and audio tokenizers
- 18:07 Getting started: self-hosting (Hugging Face, Ollama) and cloud deployment (Vertex AI)