Videos hacEQHHhu2Q
Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
Scene timeline
57 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 196
- whisperx 196
- chunks
- 37
- from 196 cues
- keyframes
- 31
- kept of 57 captured
- frames with text
- 31
- 608 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 7.5 MB
- word timings on 196 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:38 | 0s |
stt |
done | — | 2026-08-09 06:47 | 28s |
chunk |
done | — | 2026-08-09 06:47 | 0s |
text_embed |
done | — | 2026-08-10 19:39 | 0s |
keyframe |
done | — | 2026-08-09 06:47 | 2m 51s |
ocr |
done | — | 2026-08-09 06:50 | 12s |
frame_embed |
done | — | 2026-08-10 19:39 | 5s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.97
- Amazon AGI Lab0.99
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo0.99
- reducto1.00
- Sonar1.00
- Makers of1.00
- together.ai0.98
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- Why Large?0.99
- Tiny LMs& Agents0.99
- on Edge/Robotics1.00
- Al Engineer World's Fair 20260.99
- Cormac Brick1.00
- Principal Software Engineer, Google Al Edge1.00
-
- AlEngineer0.98
- Introduction1.00
- World'sFair1.00
- Background1.00
- PRESENTED BY1.00
- Small Models1.00
- Microsoft1.00
- > Tiny Models — off the shelf0.98
- >Tiny Models — custom0.97
- Examples1.00
- World's Fair0.97
- Engineering the future of Al0.99
-
- AlEngineer0.99
- Background1.00
- World's Fair0.96
- Google Al Edge0.98
- Al Edge Team at Google as tech lead1.00
- PRESENTED BY1.00
- Developing open source projects (LiteRT-LM,1.00
- Your App1.00
- Microsoft1.00
- LiteRT, Mediapipe). Making it easy to deploy Al1.00
- Models across edge devices.1.00
- MediaPipe1.00
- Delivering Edge Al core tech to Google products0.99
- Working with Gemma team to ensure their best0.99
- LiteRT-LM1.00
- models run well on lots of devices0.98
- LiteRT1.00
- Significant focus on small and tiny models0.99
- (FKA TensorFlow Lite)0.98
- CPU1.00
- GPU1.00
- NPU1.00
- World's Fair0.98
- TRACK 2· JULY 1,20260.97
- Robotics & World Models0.99
-
- AlEngineer0.98
- Why do edge Al?0.99
- World'sFair1.00
- Latency / UX0.95
- Privacy1.00
- Savings1.00
- Offline Use1.00
- Fast, consistent1.00
- Sensitive data stays1.00
- No cloud/token costs1.00
- Reliably available0.99
- speed1.00
- on device1.00
- World's Fair0.97
- TRACK 2· JULY 1, 20260.95
- Robotics & World Models1.00
-
- AlEngineer0.98
- Edge Al Challenges0.99
- World'sFair1.00
- DRAM Cost1.00
- Target Devices0.97
- Less Studied1.00
- Memory constraints on edge0.99
- Wide pool of varying hardware1.00
- Working on a less studied end of0.99
- hardware1.00
- (CPU, GPU, NPU)0.99
- the LLM spectrum1.00
- World's Fair0.97
- TRACK 2· JULY 1, 20260.96
- Robotics & World Models1.00
-
- AlEngineer0.97
- Small Models0.99
- World'sFair1.00
- Sometimes these are built into the OS1.00
- Minimize footprint with deep quantization1.00
- Mobile: Sometimes ship directly with app0.99
- Playbook: mostly prompting, LoRA adaptors0.99
- iOT/Robotics: Typically require ~4G+ of DDR0.98
- Somewhat robust: function calling, agent skills0.99
- Typically 0.8-2B parameters in size1.00
- World'sFair1.00
- TRACK 2· JULY 1, 20260.95
- Robotics & World Models0.99
-
- AlEngineer0.97
- Small Models1.00
- World'sFair1.00
- PRESENTED BY1.00
- Sometimes these are built into the OS0.99
- Minimize footprint with deep quantization1.00
- Microsoft1.00
- Mobile: Sometimes ship directly with app0.99
- Playbook: mostly prompting, LoRA adaptors1.00
- iOT/Robotics: Typically require ~4G+ of DDR0.99
- Somewhat robust: function calling, agent skills1.00
- Typically 0.8-2B parameters in size0.99
- World'sFair1.00
- TRACK 2· JULY 1,20260.97
- Robotics & World Models0.99
-
- AlEngineer0.99
- Gemma 4 - Smaller models with strong reasoning0.99
- World's Fair0.99
- PRESENTED BY1.00
- Microsoft1.00
- 69.4%1.00
- 60.0%1.00
- 42.4%1.00
- Gemma 41.00
- 31B0.97
- Gemma 40.98
- 26B A4B0.99
- Gemma 41.00
- E4B0.98
- Gemma 41.00
- E2B1.00
- Gemma 31.00
- 27B0.99
- Gemma 41.00
- 3180.97
- Gemma 41.00
- 26B A4B0.90
- Gemma 40.99
- E4B0.87
- Gemma 41.00
- E2B0.94
- Gemma 30.99
- 27B0.84
- MMLU Pro1.00
- GPQA Diamond1.00
- Advanced general knowledge1.00
- PhD-level scientific1.00
- and reasoning1.00
- reasoning1.00
- World's Fair0.98
- TRACK 2· JULY 1,20260.97
- Robotics & World Models0.99
-
- AlEngineer0.98
- Gemma 4 - Optimized memory footprint0.99
- World's Fair0.99
- Gemma4-E2B Model File Size and Memory Footprint0.99
- 25001.00
- 24681.00
- embedding / PLE0.96
- text embedding1.00
- Audio Encoder1.00
- 20001.00
- Vision Encoder1.00
- Other Non Model Metadata1.00
- Drafter1.00
- 15001.00
- Model weights (text only,0.98
- non-embeddings)1.00
- 11441.00
- 10000.99
- 8411.00
- 5001.00
- Model File Size (MiB)1.00
- In Memory Size (MiB,1.00
- Multimodal)0.98
- In Memory Size (MiB,1.00
- text-only)1.00
- World's Fair0.94
- TRACK 2· JULY 1, 20260.95
- Robotics & World Models1.00
-
- AlEngineer0.99
- Gemma 4 - E2B0.95
- World's Fair0.97
- (excludes MTP)1.00
- Platform + Device1.00
- Backend1.00
- Prefill (tk/s)0.99
- Decode (tk/s)0.96
- CPU1.00
- 5571.00
- 46.91.00
- Android + S26 Ultra0.99
- OpenCL1.00
- 3,8081.00
- 52.11.00
- NPU1.00
- 7,4631.00
- 48.11.00
- CPU1.00
- 5321.00
- 25.01.00
- iOS+ iPhone 17 Pro1.00
- Metal1.00
- 2,8781.00
- 56.51.00
- CPU1.00
- 2601.00
- 35.01.00
- Linux + NVIDIA1.00
- GeForce RTX 4090 & Arm 2.3 & 2.8GHz0.99
- WebGPU1.00
- 11,2341.00
- 143.41.00
- CPU1.00
- 9011.00
- 41.61.00
- macOS + MacBook Pro M41.00
- WebGPU1.00
- 7,8351.00
- 160.21.00
- CPU1.00
- 3571.00
- 10.11.00
- Windows + Intel1.00
- Raptor Lake + UHD 7701.00
- WebGPU1.00
- 4721.00
- 18.61.00
- Raspberry Pi 5 16GB0.99
- CPU1.00
- 1331.00
- 7.60.98
- Jetson Orin Nano1.00
- GPU1.00
- 1,1421.00
- 24.21.00
- Qualcomm IQ-8275 EVK1.00
- NPU1.00
- 3,7471.00
- 31.71.00
- World's Fair0.94
- TRACK 2· JULY 1, 20260.96
- Robotics & World Models1.00
-
- AlEngineer0.98
- Gemma 4 - E2B Speed0.97
- World'sFair1.00
- (excludes MTP)1.00
- Platform + Device1.00
- Backend1.00
- Prefill (tk/s)0.99
- Decode (tk/s)1.00
- CPU1.00
- 5571.00
- 46.91.00
- Android + S26 Ultra1.00
- OpenCL1.00
- 3,8081.00
- 52.11.00
- NPU1.00
- 7,4631.00
- 48.11.00
- CPU1.00
- 5321.00
- 25.01.00
- iOS+ iPhone 17 Pro0.98
- Metal1.00
- 2,8781.00
- 56.51.00
- CPU1.00
- 2601.00
- 35.01.00
- Linux + NVIDIA1.00
- GeForce RTX 4090 & Arm 2.3 & 2.8GHz0.99
- WebGPU1.00
- 11,2341.00
- 143.41.00
- CPU1.00
- 9011.00
- 41.61.00
- macOS + MacBook Pro M41.00
- WebGPU1.00
- 7,8351.00
- 160.21.00
- CPU1.00
- 3571.00
- 10.11.00
- Windows + Intel1.00
- Raptor Lake + UHD 7700.97
- WebGPU1.00
- 4721.00
- 18.61.00
- Raspberry Pi 5 16GB0.97
- CPU1.00
- 1331.00
- 7.61.00
- Jetson Orin Nano0.99
- GPU1.00
- 1,1421.00
- 24.21.00
- Qualcomm IQ-8275 EVK1.00
- NPU1.00
- 3,7471.00
- 31.71.00
- World's Fair0.95
- TRACK 2· JULY 1, 20260.96
- Robotics & World Models1.00
Transcript
196 cues· 3,631 words· 19,153 chars
- 0:12 Yeah, so yeah, a bit of a change of speed from the last two talks that we're looking at kind of higher end robots.
- 0:17 If we want for intelligence to get into lots and lots and lots of devices and not just really expensive robots, we are going to need tiny models.
- 0:27 And this talk is about what is the state of the art of tiny models at the moment?
- 0:32 What are the things they're good at?
- 0:33 And what are the things you can go start building today?
- 0:38 Okay, so firstly, a bit of background, like briefly on me and the team I work on.
- 0:42 Then we're gonna take a look at small models that you may be kind of more familiar with, just kind of explore what they can do, what they can't do yet.
- 0:51 And then kind of see, hey, why do we need even smaller models?
- 0:55 And then just looking at the state of the art of tiny models today and what you need to do to get them into a form where you can deploy them in production to do useful things.
- 1:03 And lastly, we've got a couple of examples that we can look at from work from our team.
- 1:10 Okay, so me, I've worked in Edge AI for a while.
- 1:17 These days I work as a tech lead on the AI Edge team at Google.
- 1:21 Within the team, the types of things we do are we develop kind of open source projects called like Lightro TLM, Lightro Team MediaPipe, and these make it easy to deploy AI to Edge devices.
- 1:34 We also do a lot of work delivering edge AI core technology to Google's own products, some of which would be via tiny models.
- 1:44 And then we also work with the Gemma team to ensure their models work well and run well on lots of devices.
- 1:49 And then we have a significant focus on small and tiny models, because that's what
- 1:55 That's what's useful for a lot of kind of mobile phone applications.
- 2:00 Or if we want to be able to ship a model in browser, that also has to be really, really small.
- 2:05 And generally, our kind of playbook is we develop things for first party use, like for in-house use first.
- 2:11 And then if we can figure out a way to share that via an open source package or make those tools available to the wider world, we do so.
- 2:19 And that helps kind of lots of other people build similar types of things using open source technology.
- 2:28 OK, so why do edge AI?
- 2:29 This is probably as opposed to just doing everything in the cloud.
- 2:35 It's kind of obvious, but I'll kind of go through it anyway.
- 2:37 There's kind of latency.
- 2:38 You have fast, consistent speed.
- 2:40 Privacy, data stays on the device.
- 2:42 Offline use, it's kind of reliably available.
- 2:48 That feature that you rely on in your mobile device will still work even when you don't have reception.
- 2:53 That can be very helpful.
- 2:54 And then savings, especially these days, if the alternative is to call even a faster model on the cloud, that will come at a cost, particularly if you're kind of shipping an app
- 3:07 or like a mobile phone app or something in browser where the user interaction is that kind of very, very large scale, then even though those tokens are relatively cheap, you're multiplying it by a large number and it'll add up quickly.
- 3:22 So then the main challenges then of deploying AI on the edge is the left most one is kind of new, which is DRAM cost.
- 3:32 And it's a really significant constraint that, and you'll even see some mobile phone manufacturers are putting less DRAM into their devices this year than previously.
- 3:43 You'll also see that since launch, the cost of a Raspberry Pi 3 16 gigabytes has gone up by a factor of like 2.5 X.
- 3:53 So DRAM cost is really, really significant.
- 3:56 That then casts a shadow over the rest of this talk, where in order to be able to get AI applications running on the edge, we need to really think a lot about quantization, and we also really need to think about what is the smallest possible model we can use for a given task.
- 4:15 Other challenges are, yeah, there's a wider pool of target devices.
- 4:19 And yet another challenge is, yeah, it's kind of fair to say that a lot of the research hours that go into LLMs these days are into the much larger models and MOE techniques and this types of stuff.
- 4:33 And the lower end of the LLM spectrum is a lot less studied.
- 4:38 So yeah, these are challenges of deploying the Edge.
- 4:42 Okay, so small models, and when I say small, I mean kind of typically maybe kind of one to two or one to four billion parameters.
- 4:50 You may find that these are built into the OS.
- 4:52 There's a version of a small model that ships in Android high-end phones today with AI Core.
- 4:58 There's a version that ships with Apple, with Apple Intelligence.
- 5:02 Some app vendors will ship models this size in their app.
- 5:06 We certainly work with some app vendors that do this.
- 5:11 And for like IoT and robotics, you would typically require like four, maybe four to eight gigs of DRAM in order to be able to ship this grade of model, which then is an implied cost on the device, right?
- 5:24 So then kind of restricts these models to things like laptops, mobile phones, or kind of higher end electronics and kind of puts it out of reach of maybe a lot of lower tier web browsers or the wider kind of IoT and consumer robotics market.
- 5:41 Yeah, and for smaller models, yep, developing smaller models, and we'll look in a while, we do a lot of work to minimize footprints with kind of quantization.
loading
Chapters
- 0:00 Why intelligence at scale needs tiny models
- 1:17 The Google AI Edge team and its open source stack
- 2:35 Why run on the edge at all
- 3:25 The real constraint: DRAM cost
- 4:40 Small models: 1 to 4 billion parameters
- 6:08 Shrinking Gemma to 2.9 bits per weight
- 7:36 Decode speeds across Raspberry Pi, Jetson, and NPUs
- 9:30 Try it yourself: AI Edge Gallery and a hobby robot
- 12:07 When small is still too big: tiny models
- 13:24 Off the shelf tiny models: ASR, vision, embeddings
- 14:28 Fine tuning for voice to function calling
- 17:50 In production: offline voice dictation
- 19:30 Takeaways and Q&A