read-only demo

Videos hacEQHHhu2Q

Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google

index_state ready data_status ok

AI Engineer· published 2026-07-25· 0:21:44· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:13, 1 of 1 keyframes kept
  5. Shot 4, 0:13 to 0:29, 1 of 1 keyframes kept
  6. Shot 5, 0:29 to 0:36, 1 of 1 keyframes kept
  7. Shot 6, 0:36 to 1:08, 1 of 1 keyframes kept
  8. Shot 7, 1:08 to 1:34, 1 of 1 keyframes kept
  9. Shot 8, 1:34 to 2:01, 0 of 1 keyframes kept
  10. Shot 9, 2:01 to 2:27, 0 of 1 keyframes kept
  11. Shot 10, 2:27 to 2:54, 1 of 1 keyframes kept
  12. Shot 11, 2:54 to 3:21, 0 of 1 keyframes kept
  13. Shot 12, 3:21 to 3:36, 1 of 1 keyframes kept
  14. Shot 13, 3:36 to 4:08, 0 of 1 keyframes kept
  15. Shot 14, 4:08 to 4:41, 0 of 1 keyframes kept
  16. Shot 15, 4:41 to 5:09, 1 of 1 keyframes kept
  17. Shot 16, 5:09 to 5:37, 0 of 1 keyframes kept
  18. Shot 17, 5:37 to 6:06, 1 of 1 keyframes kept
  19. Shot 18, 6:06 to 6:34, 1 of 1 keyframes kept
  20. Shot 19, 6:34 to 7:00, 1 of 1 keyframes kept
  21. Shot 20, 7:00 to 7:26, 0 of 1 keyframes kept
  22. Shot 21, 7:26 to 7:53, 1 of 1 keyframes kept
  23. Shot 22, 7:53 to 8:20, 1 of 1 keyframes kept
  24. Shot 23, 8:20 to 8:48, 0 of 1 keyframes kept
  25. Shot 24, 8:48 to 9:15, 0 of 1 keyframes kept
  26. Shot 25, 9:15 to 9:31, 1 of 1 keyframes kept
  27. Shot 26, 9:31 to 9:40, 1 of 1 keyframes kept
  28. Shot 27, 9:40 to 10:15, 0 of 1 keyframes kept
  29. Shot 28, 10:15 to 10:49, 1 of 1 keyframes kept
  30. Shot 29, 10:49 to 11:22, 0 of 1 keyframes kept
  31. Shot 30, 11:22 to 11:47, 1 of 1 keyframes kept
  32. Shot 31, 11:47 to 12:13, 0 of 1 keyframes kept
  33. Shot 32, 12:13 to 12:38, 1 of 1 keyframes kept
  34. Shot 33, 12:38 to 13:23, 1 of 1 keyframes kept
  35. Shot 34, 13:23 to 13:52, 1 of 1 keyframes kept
  36. Shot 35, 13:52 to 14:22, 0 of 1 keyframes kept
  37. Shot 36, 14:22 to 14:23, 0 of 1 keyframes kept
  38. Shot 37, 14:23 to 14:25, 0 of 1 keyframes kept
  39. Shot 38, 14:25 to 14:50, 1 of 1 keyframes kept
  40. Shot 39, 14:50 to 15:15, 0 of 1 keyframes kept
  41. Shot 40, 15:15 to 15:40, 1 of 1 keyframes kept
  42. Shot 41, 15:40 to 16:05, 0 of 1 keyframes kept
  43. Shot 42, 16:05 to 16:30, 0 of 1 keyframes kept
  44. Shot 43, 16:30 to 16:55, 1 of 1 keyframes kept
  45. Shot 44, 16:55 to 17:21, 0 of 1 keyframes kept
  46. Shot 45, 17:21 to 17:46, 0 of 1 keyframes kept
  47. Shot 46, 17:46 to 17:51, 1 of 1 keyframes kept
  48. Shot 47, 17:51 to 18:18, 1 of 1 keyframes kept
  49. Shot 48, 18:18 to 18:23, 0 of 1 keyframes kept
  50. Shot 49, 18:23 to 18:50, 0 of 1 keyframes kept
  51. Shot 50, 18:50 to 19:31, 0 of 1 keyframes kept
  52. Shot 51, 19:31 to 19:32, 0 of 1 keyframes kept
  53. Shot 52, 19:32 to 20:10, 0 of 1 keyframes kept
  54. Shot 53, 20:10 to 20:35, 1 of 1 keyframes kept
  55. Shot 54, 20:35 to 21:01, 1 of 1 keyframes kept
  56. Shot 55, 21:01 to 21:27, 1 of 1 keyframes kept
  57. Shot 56, 21:27 to 21:44, 0 of 1 keyframes kept

57 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
196
whisperx 196
chunks
37
from 196 cues
keyframes
31
kept of 57 captured
frames with text
31
608 lines read
chapters
13
from the source metadata
keyframe bytes
7.5 MB
word timings on 196 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:38 0s
stt done 2026-08-09 06:47 28s
chunk done 2026-08-09 06:47 0s
text_embed done 2026-08-10 19:39 0s
keyframe done 2026-08-09 06:47 2m 51s
ocr done 2026-08-09 06:50 12s
frame_embed done 2026-08-10 19:39 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 455.2

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 670.0

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:11 #2 done24 line(s)

    shot 2·sharpness 2742.1

    1. LAB & PLATINUM SPONSORS0.97
    2. Amazon AGI Lab0.99
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.98
    6. OpenAl0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo0.99
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. together.ai0.98
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:13 #3 done2 line(s)

    shot 3·sharpness 186.0

    1. AlEngineer0.99
    2. World's Fair1.00
  • 0:20 #4 done2 line(s)

    shot 4·sharpness 254.2

    1. AlEngineer0.99
    2. World's Fair1.00
  • 0:35 #5 done6 line(s)

    shot 5·sharpness 1891.9

    1. Why Large?0.99
    2. Tiny LMs& Agents0.99
    3. on Edge/Robotics1.00
    4. Al Engineer World's Fair 20260.99
    5. Cormac Brick1.00
    6. Principal Software Engineer, Google Al Edge1.00
  • 0:52 #6 done12 line(s)

    shot 6·sharpness 1935.8

    1. AlEngineer0.98
    2. Introduction1.00
    3. World'sFair1.00
    4. Background1.00
    5. PRESENTED BY1.00
    6. Small Models1.00
    7. Microsoft1.00
    8. > Tiny Models — off the shelf0.98
    9. >Tiny Models — custom0.97
    10. Examples1.00
    11. World's Fair0.97
    12. Engineering the future of Al0.99
  • 1:12 #7 done25 line(s)

    shot 7·sharpness 2788.4

    1. AlEngineer0.99
    2. Background1.00
    3. World's Fair0.96
    4. Google Al Edge0.98
    5. Al Edge Team at Google as tech lead1.00
    6. PRESENTED BY1.00
    7. Developing open source projects (LiteRT-LM,1.00
    8. Your App1.00
    9. Microsoft1.00
    10. LiteRT, Mediapipe). Making it easy to deploy Al1.00
    11. Models across edge devices.1.00
    12. MediaPipe1.00
    13. Delivering Edge Al core tech to Google products0.99
    14. Working with Gemma team to ensure their best0.99
    15. LiteRT-LM1.00
    16. models run well on lots of devices0.98
    17. LiteRT1.00
    18. Significant focus on small and tiny models0.99
    19. (FKA TensorFlow Lite)0.98
    20. CPU1.00
    21. GPU1.00
    22. NPU1.00
    23. World's Fair0.98
    24. TRACK 2· JULY 1,20260.97
    25. Robotics & World Models0.99
  • 1:50 #8 skipped

    shot 8·duplicate of #7

  • 2:21 #9 skipped

    shot 9·duplicate of #7

  • 2:43 #10 done16 line(s)

    shot 10·sharpness 1961.7

    1. AlEngineer0.98
    2. Why do edge Al?0.99
    3. World'sFair1.00
    4. Latency / UX0.95
    5. Privacy1.00
    6. Savings1.00
    7. Offline Use1.00
    8. Fast, consistent1.00
    9. Sensitive data stays1.00
    10. No cloud/token costs1.00
    11. Reliably available0.99
    12. speed1.00
    13. on device1.00
    14. World's Fair0.97
    15. TRACK 2· JULY 1, 20260.95
    16. Robotics & World Models1.00
  • 2:57 #11 skipped

    shot 11·duplicate of #10

  • 3:34 #12 done15 line(s)

    shot 12·sharpness 2138.4

    1. AlEngineer0.98
    2. Edge Al Challenges0.99
    3. World'sFair1.00
    4. DRAM Cost1.00
    5. Target Devices0.97
    6. Less Studied1.00
    7. Memory constraints on edge0.99
    8. Wide pool of varying hardware1.00
    9. Working on a less studied end of0.99
    10. hardware1.00
    11. (CPU, GPU, NPU)0.99
    12. the LLM spectrum1.00
    13. World's Fair0.97
    14. TRACK 2· JULY 1, 20260.96
    15. Robotics & World Models1.00
  • 3:52 #13 skipped

    shot 13·duplicate of #12

  • 4:25 #14 skipped

    shot 14·duplicate of #12

  • 4:47 #15 done13 line(s)

    shot 15·sharpness 1964.5

    1. AlEngineer0.97
    2. Small Models0.99
    3. World'sFair1.00
    4. Sometimes these are built into the OS1.00
    5. Minimize footprint with deep quantization1.00
    6. Mobile: Sometimes ship directly with app0.99
    7. Playbook: mostly prompting, LoRA adaptors0.99
    8. iOT/Robotics: Typically require ~4G+ of DDR0.98
    9. Somewhat robust: function calling, agent skills0.99
    10. Typically 0.8-2B parameters in size1.00
    11. World'sFair1.00
    12. TRACK 2· JULY 1, 20260.95
    13. Robotics & World Models0.99
  • 5:21 #16 skipped

    shot 16·duplicate of #10

  • 5:52 #17 done15 line(s)

    shot 17·sharpness 2108.6

    1. AlEngineer0.97
    2. Small Models1.00
    3. World'sFair1.00
    4. PRESENTED BY1.00
    5. Sometimes these are built into the OS0.99
    6. Minimize footprint with deep quantization1.00
    7. Microsoft1.00
    8. Mobile: Sometimes ship directly with app0.99
    9. Playbook: mostly prompting, LoRA adaptors1.00
    10. iOT/Robotics: Typically require ~4G+ of DDR0.99
    11. Somewhat robust: function calling, agent skills1.00
    12. Typically 0.8-2B parameters in size0.99
    13. World'sFair1.00
    14. TRACK 2· JULY 1,20260.97
    15. Robotics & World Models0.99
  • 6:31 #18 done37 line(s)

    shot 18·sharpness 1800.2

    1. AlEngineer0.99
    2. Gemma 4 - Smaller models with strong reasoning0.99
    3. World's Fair0.99
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. 69.4%1.00
    7. 60.0%1.00
    8. 42.4%1.00
    9. Gemma 41.00
    10. 31B0.97
    11. Gemma 40.98
    12. 26B A4B0.99
    13. Gemma 41.00
    14. E4B0.98
    15. Gemma 41.00
    16. E2B1.00
    17. Gemma 31.00
    18. 27B0.99
    19. Gemma 41.00
    20. 3180.97
    21. Gemma 41.00
    22. 26B A4B0.90
    23. Gemma 40.99
    24. E4B0.87
    25. Gemma 41.00
    26. E2B0.94
    27. Gemma 30.99
    28. 27B0.84
    29. MMLU Pro1.00
    30. GPQA Diamond1.00
    31. Advanced general knowledge1.00
    32. PhD-level scientific1.00
    33. and reasoning1.00
    34. reasoning1.00
    35. World's Fair0.98
    36. TRACK 2· JULY 1,20260.97
    37. Robotics & World Models0.99
  • 6:57 #19 done28 line(s)

    shot 19·sharpness 1759.6

    1. AlEngineer0.98
    2. Gemma 4 - Optimized memory footprint0.99
    3. World's Fair0.99
    4. Gemma4-E2B Model File Size and Memory Footprint0.99
    5. 25001.00
    6. 24681.00
    7. embedding / PLE0.96
    8. text embedding1.00
    9. Audio Encoder1.00
    10. 20001.00
    11. Vision Encoder1.00
    12. Other Non Model Metadata1.00
    13. Drafter1.00
    14. 15001.00
    15. Model weights (text only,0.98
    16. non-embeddings)1.00
    17. 11441.00
    18. 10000.99
    19. 8411.00
    20. 5001.00
    21. Model File Size (MiB)1.00
    22. In Memory Size (MiB,1.00
    23. Multimodal)0.98
    24. In Memory Size (MiB,1.00
    25. text-only)1.00
    26. World's Fair0.94
    27. TRACK 2· JULY 1, 20260.95
    28. Robotics & World Models1.00
  • 7:03 #20 skipped

    shot 20·duplicate of #19

  • 7:32 #21 done63 line(s)

    shot 21·sharpness 2795.8

    1. AlEngineer0.99
    2. Gemma 4 - E2B0.95
    3. World's Fair0.97
    4. (excludes MTP)1.00
    5. Platform + Device1.00
    6. Backend1.00
    7. Prefill (tk/s)0.99
    8. Decode (tk/s)0.96
    9. CPU1.00
    10. 5571.00
    11. 46.91.00
    12. Android + S26 Ultra0.99
    13. OpenCL1.00
    14. 3,8081.00
    15. 52.11.00
    16. NPU1.00
    17. 7,4631.00
    18. 48.11.00
    19. CPU1.00
    20. 5321.00
    21. 25.01.00
    22. iOS+ iPhone 17 Pro1.00
    23. Metal1.00
    24. 2,8781.00
    25. 56.51.00
    26. CPU1.00
    27. 2601.00
    28. 35.01.00
    29. Linux + NVIDIA1.00
    30. GeForce RTX 4090 & Arm 2.3 & 2.8GHz0.99
    31. WebGPU1.00
    32. 11,2341.00
    33. 143.41.00
    34. CPU1.00
    35. 9011.00
    36. 41.61.00
    37. macOS + MacBook Pro M41.00
    38. WebGPU1.00
    39. 7,8351.00
    40. 160.21.00
    41. CPU1.00
    42. 3571.00
    43. 10.11.00
    44. Windows + Intel1.00
    45. Raptor Lake + UHD 7701.00
    46. WebGPU1.00
    47. 4721.00
    48. 18.61.00
    49. Raspberry Pi 5 16GB0.99
    50. CPU1.00
    51. 1331.00
    52. 7.60.98
    53. Jetson Orin Nano1.00
    54. GPU1.00
    55. 1,1421.00
    56. 24.21.00
    57. Qualcomm IQ-8275 EVK1.00
    58. NPU1.00
    59. 3,7471.00
    60. 31.71.00
    61. World's Fair0.94
    62. TRACK 2· JULY 1, 20260.96
    63. Robotics & World Models1.00
  • 8:17 #22 done63 line(s)

    shot 22·sharpness 2693.8

    1. AlEngineer0.98
    2. Gemma 4 - E2B Speed0.97
    3. World'sFair1.00
    4. (excludes MTP)1.00
    5. Platform + Device1.00
    6. Backend1.00
    7. Prefill (tk/s)0.99
    8. Decode (tk/s)1.00
    9. CPU1.00
    10. 5571.00
    11. 46.91.00
    12. Android + S26 Ultra1.00
    13. OpenCL1.00
    14. 3,8081.00
    15. 52.11.00
    16. NPU1.00
    17. 7,4631.00
    18. 48.11.00
    19. CPU1.00
    20. 5321.00
    21. 25.01.00
    22. iOS+ iPhone 17 Pro0.98
    23. Metal1.00
    24. 2,8781.00
    25. 56.51.00
    26. CPU1.00
    27. 2601.00
    28. 35.01.00
    29. Linux + NVIDIA1.00
    30. GeForce RTX 4090 & Arm 2.3 & 2.8GHz0.99
    31. WebGPU1.00
    32. 11,2341.00
    33. 143.41.00
    34. CPU1.00
    35. 9011.00
    36. 41.61.00
    37. macOS + MacBook Pro M41.00
    38. WebGPU1.00
    39. 7,8351.00
    40. 160.21.00
    41. CPU1.00
    42. 3571.00
    43. 10.11.00
    44. Windows + Intel1.00
    45. Raptor Lake + UHD 7700.97
    46. WebGPU1.00
    47. 4721.00
    48. 18.61.00
    49. Raspberry Pi 5 16GB0.97
    50. CPU1.00
    51. 1331.00
    52. 7.61.00
    53. Jetson Orin Nano0.99
    54. GPU1.00
    55. 1,1421.00
    56. 24.21.00
    57. Qualcomm IQ-8275 EVK1.00
    58. NPU1.00
    59. 3,7471.00
    60. 31.71.00
    61. World's Fair0.95
    62. TRACK 2· JULY 1, 20260.96
    63. Robotics & World Models1.00
  • 8:34 #23 skipped

    shot 23·duplicate of #22

Transcript

196 cues· 3,631 words· 19,153 chars

  1. 0:12 Yeah, so yeah, a bit of a change of speed from the last two talks that we're looking at kind of higher end robots.
  2. 0:17 If we want for intelligence to get into lots and lots and lots of devices and not just really expensive robots, we are going to need tiny models.
  3. 0:27 And this talk is about what is the state of the art of tiny models at the moment?
  4. 0:32 What are the things they're good at?
  5. 0:33 And what are the things you can go start building today?
  6. 0:38 Okay, so firstly, a bit of background, like briefly on me and the team I work on.
  7. 0:42 Then we're gonna take a look at small models that you may be kind of more familiar with, just kind of explore what they can do, what they can't do yet.
  8. 0:51 And then kind of see, hey, why do we need even smaller models?
  9. 0:55 And then just looking at the state of the art of tiny models today and what you need to do to get them into a form where you can deploy them in production to do useful things.
  10. 1:03 And lastly, we've got a couple of examples that we can look at from work from our team.
  11. 1:10 Okay, so me, I've worked in Edge AI for a while.
  12. 1:17 These days I work as a tech lead on the AI Edge team at Google.
  13. 1:21 Within the team, the types of things we do are we develop kind of open source projects called like Lightro TLM, Lightro Team MediaPipe, and these make it easy to deploy AI to Edge devices.
  14. 1:34 We also do a lot of work delivering edge AI core technology to Google's own products, some of which would be via tiny models.
  15. 1:44 And then we also work with the Gemma team to ensure their models work well and run well on lots of devices.
  16. 1:49 And then we have a significant focus on small and tiny models, because that's what
  17. 1:55 That's what's useful for a lot of kind of mobile phone applications.
  18. 2:00 Or if we want to be able to ship a model in browser, that also has to be really, really small.
  19. 2:05 And generally, our kind of playbook is we develop things for first party use, like for in-house use first.
  20. 2:11 And then if we can figure out a way to share that via an open source package or make those tools available to the wider world, we do so.
  21. 2:19 And that helps kind of lots of other people build similar types of things using open source technology.
  22. 2:28 OK, so why do edge AI?
  23. 2:29 This is probably as opposed to just doing everything in the cloud.
  24. 2:35 It's kind of obvious, but I'll kind of go through it anyway.
  25. 2:37 There's kind of latency.
  26. 2:38 You have fast, consistent speed.
  27. 2:40 Privacy, data stays on the device.
  28. 2:42 Offline use, it's kind of reliably available.
  29. 2:48 That feature that you rely on in your mobile device will still work even when you don't have reception.
  30. 2:53 That can be very helpful.
  31. 2:54 And then savings, especially these days, if the alternative is to call even a faster model on the cloud, that will come at a cost, particularly if you're kind of shipping an app
  32. 3:07 or like a mobile phone app or something in browser where the user interaction is that kind of very, very large scale, then even though those tokens are relatively cheap, you're multiplying it by a large number and it'll add up quickly.
  33. 3:22 So then the main challenges then of deploying AI on the edge is the left most one is kind of new, which is DRAM cost.
  34. 3:32 And it's a really significant constraint that, and you'll even see some mobile phone manufacturers are putting less DRAM into their devices this year than previously.
  35. 3:43 You'll also see that since launch, the cost of a Raspberry Pi 3 16 gigabytes has gone up by a factor of like 2.5 X.
  36. 3:53 So DRAM cost is really, really significant.
  37. 3:56 That then casts a shadow over the rest of this talk, where in order to be able to get AI applications running on the edge, we need to really think a lot about quantization, and we also really need to think about what is the smallest possible model we can use for a given task.
  38. 4:15 Other challenges are, yeah, there's a wider pool of target devices.
  39. 4:19 And yet another challenge is, yeah, it's kind of fair to say that a lot of the research hours that go into LLMs these days are into the much larger models and MOE techniques and this types of stuff.
  40. 4:33 And the lower end of the LLM spectrum is a lot less studied.
  41. 4:38 So yeah, these are challenges of deploying the Edge.
  42. 4:42 Okay, so small models, and when I say small, I mean kind of typically maybe kind of one to two or one to four billion parameters.
  43. 4:50 You may find that these are built into the OS.
  44. 4:52 There's a version of a small model that ships in Android high-end phones today with AI Core.
  45. 4:58 There's a version that ships with Apple, with Apple Intelligence.
  46. 5:02 Some app vendors will ship models this size in their app.
  47. 5:06 We certainly work with some app vendors that do this.
  48. 5:11 And for like IoT and robotics, you would typically require like four, maybe four to eight gigs of DRAM in order to be able to ship this grade of model, which then is an implied cost on the device, right?
  49. 5:24 So then kind of restricts these models to things like laptops, mobile phones, or kind of higher end electronics and kind of puts it out of reach of maybe a lot of lower tier web browsers or the wider kind of IoT and consumer robotics market.
  50. 5:41 Yeah, and for smaller models, yep, developing smaller models, and we'll look in a while, we do a lot of work to minimize footprints with kind of quantization.

Chapters

  1. 0:00 Why intelligence at scale needs tiny models
  2. 1:17 The Google AI Edge team and its open source stack
  3. 2:35 Why run on the edge at all
  4. 3:25 The real constraint: DRAM cost
  5. 4:40 Small models: 1 to 4 billion parameters
  6. 6:08 Shrinking Gemma to 2.9 bits per weight
  7. 7:36 Decode speeds across Raspberry Pi, Jetson, and NPUs
  8. 9:30 Try it yourself: AI Edge Gallery and a hobby robot
  9. 12:07 When small is still too big: tiny models
  10. 13:24 Off the shelf tiny models: ASR, vision, embeddings
  11. 14:28 Fine tuning for voice to function calling
  12. 17:50 In production: offline voice dictation
  13. 19:30 Takeaways and Q&A

Open at this second