read-only demo

Videos -TiET_K-E_g

From 46% to 90%: Fine-Tuning Tiny LLMs for On-Device Agents — Cormac Brick, Google

index_state ready data_status ok

AI Engineer· published 2026-05-20· 0:21:00· en-US· indexed 2026-08-10 19:49

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:36, 1 of 1 keyframes kept
  5. Shot 4, 0:36 to 0:46, 1 of 1 keyframes kept
  6. Shot 5, 0:46 to 1:17, 1 of 1 keyframes kept
  7. Shot 6, 1:17 to 1:24, 0 of 1 keyframes kept
  8. Shot 7, 1:24 to 1:35, 1 of 1 keyframes kept
  9. Shot 8, 1:35 to 2:17, 1 of 1 keyframes kept
  10. Shot 9, 2:17 to 2:36, 1 of 1 keyframes kept
  11. Shot 10, 2:36 to 2:57, 1 of 1 keyframes kept
  12. Shot 11, 2:57 to 3:39, 1 of 1 keyframes kept
  13. Shot 12, 3:39 to 4:07, 1 of 1 keyframes kept
  14. Shot 13, 4:07 to 4:35, 0 of 1 keyframes kept
  15. Shot 14, 4:35 to 5:06, 1 of 1 keyframes kept
  16. Shot 15, 5:06 to 5:43, 1 of 1 keyframes kept
  17. Shot 16, 5:43 to 6:21, 0 of 1 keyframes kept
  18. Shot 17, 6:21 to 6:22, 1 of 1 keyframes kept
  19. Shot 18, 6:22 to 6:43, 1 of 1 keyframes kept
  20. Shot 19, 6:43 to 6:47, 1 of 1 keyframes kept
  21. Shot 20, 6:47 to 7:08, 1 of 1 keyframes kept
  22. Shot 21, 7:08 to 7:26, 0 of 1 keyframes kept
  23. Shot 22, 7:26 to 7:30, 1 of 1 keyframes kept
  24. Shot 23, 7:30 to 7:32, 0 of 1 keyframes kept
  25. Shot 24, 7:32 to 7:33, 1 of 1 keyframes kept
  26. Shot 25, 7:33 to 7:34, 0 of 1 keyframes kept
  27. Shot 26, 7:34 to 7:38, 0 of 1 keyframes kept
  28. Shot 27, 7:38 to 8:05, 0 of 1 keyframes kept
  29. Shot 28, 8:05 to 8:32, 0 of 1 keyframes kept
  30. Shot 29, 8:32 to 8:59, 1 of 1 keyframes kept
  31. Shot 30, 8:59 to 9:03, 1 of 1 keyframes kept
  32. Shot 31, 9:03 to 9:14, 0 of 1 keyframes kept
  33. Shot 32, 9:14 to 9:58, 0 of 1 keyframes kept
  34. Shot 33, 9:58 to 10:16, 1 of 1 keyframes kept
  35. Shot 34, 10:16 to 10:27, 0 of 1 keyframes kept
  36. Shot 35, 10:27 to 11:16, 1 of 1 keyframes kept
  37. Shot 36, 11:16 to 11:27, 1 of 1 keyframes kept
  38. Shot 37, 11:27 to 11:59, 1 of 1 keyframes kept
  39. Shot 38, 11:59 to 12:31, 0 of 1 keyframes kept
  40. Shot 39, 12:31 to 13:18, 1 of 1 keyframes kept
  41. Shot 40, 13:18 to 13:22, 1 of 1 keyframes kept
  42. Shot 41, 13:22 to 13:26, 0 of 1 keyframes kept
  43. Shot 42, 13:26 to 13:29, 1 of 1 keyframes kept
  44. Shot 43, 13:29 to 13:33, 1 of 1 keyframes kept
  45. Shot 44, 13:33 to 13:36, 0 of 1 keyframes kept
  46. Shot 45, 13:36 to 13:40, 0 of 1 keyframes kept
  47. Shot 46, 13:40 to 13:43, 0 of 1 keyframes kept
  48. Shot 47, 13:43 to 13:47, 1 of 1 keyframes kept
  49. Shot 48, 13:47 to 13:50, 0 of 1 keyframes kept
  50. Shot 49, 13:50 to 13:55, 1 of 1 keyframes kept
  51. Shot 50, 13:55 to 13:57, 0 of 1 keyframes kept
  52. Shot 51, 13:57 to 14:01, 0 of 1 keyframes kept
  53. Shot 52, 14:01 to 14:04, 0 of 1 keyframes kept
  54. Shot 53, 14:04 to 14:08, 0 of 1 keyframes kept
  55. Shot 54, 14:08 to 14:39, 1 of 1 keyframes kept
  56. Shot 55, 14:39 to 15:10, 0 of 1 keyframes kept
  57. Shot 56, 15:10 to 15:41, 0 of 1 keyframes kept
  58. Shot 57, 15:41 to 15:42, 0 of 1 keyframes kept
  59. Shot 58, 15:42 to 15:51, 0 of 1 keyframes kept
  60. Shot 59, 15:51 to 15:55, 0 of 1 keyframes kept
  61. Shot 60, 15:55 to 15:56, 1 of 1 keyframes kept
  62. Shot 61, 15:56 to 15:58, 0 of 1 keyframes kept
  63. Shot 62, 15:58 to 16:09, 0 of 1 keyframes kept
  64. Shot 63, 16:09 to 16:28, 1 of 1 keyframes kept
  65. Shot 64, 16:28 to 16:56, 1 of 1 keyframes kept
  66. Shot 65, 16:56 to 17:24, 0 of 1 keyframes kept
  67. Shot 66, 17:24 to 17:28, 1 of 1 keyframes kept
  68. Shot 67, 17:28 to 17:58, 0 of 1 keyframes kept
  69. Shot 68, 17:58 to 18:28, 1 of 1 keyframes kept
  70. Shot 69, 18:28 to 18:58, 0 of 1 keyframes kept
  71. Shot 70, 18:58 to 19:24, 1 of 1 keyframes kept
  72. Shot 71, 19:24 to 19:51, 1 of 1 keyframes kept
  73. Shot 72, 19:51 to 20:18, 1 of 1 keyframes kept
  74. Shot 73, 20:18 to 20:45, 1 of 1 keyframes kept
  75. Shot 74, 20:45 to 21:00, 1 of 1 keyframes kept

75 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
252
whisperx 252
chunks
38
from 252 cues
keyframes
43
kept of 75 captured
frames with text
43
976 lines read
chapters
19
from the source metadata
keyframe bytes
7.8 MB
word timings on 252 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 11:41 2m 03s
stt done 2026-08-10 11:43 26s
chunk done 2026-08-10 11:44 0s
text_embed done 2026-08-10 19:49 0s
keyframe done 2026-08-10 11:44 1m 53s
ocr done 2026-08-10 11:46 16s
frame_embed done 2026-08-10 19:49 8s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:19 #3 done20 line(s)

    shot 3·sharpness 645.3

    1. Google DeepMind0.98
    2. APRIL 9, 2026 / MORNING BLOCK0.96
    3. 4TH FLOOR- RUTHERFORD0.98
    4. 11:15AM1.00
    5. The agent-ready web: Simplify user actions with WebMCP1.00
    6. Tara Agyemang / Google DeepMind0.98
    7. 11:40AM0.99
    8. Build & deploy Al-powered apps0.99
    9. Paige Bailey / Google DeepMind0.98
    10. 12:00PH0.95
    11. Agentic Evaluations at Scale—For Everybody1.00
    12. Nicholas Kang, Michael Aaron / Google DeepMind1.00
    13. 12:20PM0.99
    14. Al on Android: Ask me Anything0.99
    15. Florina Muntenescu, Oli Gaymond / Google DeepMind0.99
    16. 12:40PM0.98
    17. TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM1.00
    18. Cormac Brick/ Google0.99
    19. Schedule updates may occur. Stay up to date via our app or website0.99
    20. https://al.engineer/schedule1.00
  • 0:42 #4 done12 line(s)

    shot 4·sharpness 1674.6

    1. AI ENGINEER CONFERENCE 20261.00
    2. TLMs: Tiny LLMs and0.98
    3. Agents on Edge Devices1.00
    4. AIE1.00
    5. 1.00
    6. 1.00
    7. 1.00
    8. Bringing state-of-the-art agentic skills to the1.00
    9. edge with open models0.97
    10. Cormac Brick, Principal Engineer, Google Al Edge0.99
    11. Google DeepMind1.00
    12. AIEn0.93
  • 1:13 #5 done11 line(s)

    shot 5·sharpness 1713.9

    1. Agenda1.00
    2. 00 AI Edge, SLMs & TLMs0.95
    3. 10 Agent Skills locally on Android & iOS0.98
    4. AIE1.00
    5. 20 TLM workflow1.00
    6. 1.00
    7. 1.00
    8. 30 TLMs in Action0.98
    9. Google1.00
    10. Braintrust1.00
    11. WorkOS OpenAI0.94
  • 1:18 #6 skipped

    shot 6·duplicate of #5

  • 1:27 #7 done22 line(s)

    shot 7·sharpness 1994.7

    1. Running Al on the Edge has many benefits0.98
    2. AIE1.00
    3. $0.76
    4. 1.00
    5. 1.00
    6. 1.00
    7. 1.00
    8. Latency/ UX0.96
    9. Privacy1.00
    10. Offline use0.97
    11. Savings1.00
    12. Faster, no network0.99
    13. Sensitive data not sent1.00
    14. Works without requiring1.00
    15. Lower or no1.00
    16. involved1.00
    17. off device1.00
    18. cellular1.00
    19. data-center costs1.00
    20. AlEngineer0.91
    21. EUROPE1.00
    22. AlEng0.90
  • 1:44 #8 done16 line(s)

    shot 8·sharpness 1431.7

    1. Google Al Edge0.96
    2. Your App1.00
    3. **0.84
    4. AIE1.00
    5. MediaPipe1.00
    6. 1.00
    7. 1.00
    8. LiteRT-LM1.00
    9. LiteRT1.00
    10. (FKA TensorFlow Lite)1.00
    11. CPU1.00
    12. GPU1.00
    13. NPU1.00
    14. AlEngineer0.96
    15. EUROPE1.00
    16. AIEr0.86
  • 2:34 #9 done16 line(s)

    shot 9·sharpness 2000.5

    1. Trusted at scale1.00
    2. AIE1.00
    3. 100,000+1.00
    4. 2.7 billion +0.98
    5. 1 trillion +0.99
    6. 1.00
    7. 1.00
    8. Android apps1.00
    9. Devices1.00
    10. Average daily1.00
    11. interpreter invocations in1.00
    12. Android1.00
    13. Source: Internal Android data1.00
    14. Engineering the future of Al1.00
    15. AlEng0.92
    16. EURO0.94
  • 2:44 #10 done22 line(s)

    shot 10·sharpness 1837.8

    1. ...and far beyond Android0.97
    2. ***0.61
    3. .tflite0.99
    4. AIE1.00
    5. 1.00
    6. 1.00
    7. 1.00
    8. 1.00
    9. CPU1.00
    10. GPU1.00
    11. NPU1.00
    12. i050.97
    13. Android1.00
    14. iOS0.89
    15. macOS0.97
    16. Linux1.00
    17. Windows1.00
    18. Web1.00
    19. loT0.93
    20. Engineering the future of Al0.99
    21. AlEngi0.90
    22. EURO1.00
  • 3:10 #11 done11 line(s)

    shot 11·sharpness 1925.0

    1. System level GenAl0.98
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. Gemini Nano with AICore0.97
    6. Apple Intelligence Writing0.99
    7. summarization on Android1.00
    8. tools on iOS0.97
    9. Engineering the future of Al0.99
    10. AlEng0.96
    11. EURC0.93
  • 3:45 #12 done19 line(s)

    shot 12·sharpness 1870.5

    1. System vs In-app GenAl0.98
    2. *★*0.52
    3. AIE1.00
    4. System GenAl0.97
    5. In-app GenAI0.98
    6. 1.00
    7. 1.00
    8. 1.00
    9. Most capable1.00
    10. Customized to tasks1.00
    11. highly optimized1.00
    12. Loaded with app / webpage0.99
    13. Pre-loaded with device1.00
    14. Customization & Reach1.00
    15. SLMs: 2B-4B0.99
    16. TLMs: 100M-1B0.99
    17. AlEngineer0.95
    18. EUROPE1.00
    19. AIEn0.92
  • 4:27 #13 skipped

    shot 13·duplicate of #12

  • 4:42 #14 done9 line(s)

    shot 14·sharpness 1397.3

    1. 201.00
    2. AIE1.00
    3. Agent Skills on Device0.95
    4. 1.00
    5. (systemGenAl)0.98
    6. Google1.00
    7. Engineering the future of Al1.00
    8. AlEng0.95
    9. EURD0.85
  • 5:14 #15 done79 line(s)

    shot 15·sharpness 2481.6

    1. Google Al Edge Gallery App: On-device Al in Action0.99
    2. Featuring Gemma41.00
    3. Google AI Edge Gallery0.97
    4. Gemma-4-E28-0.75
    5. Gomma-4-E20-it0.74
    6. Al Chat0.91
    7. Ask Image0.95
    8. Gemma-4-E28-it0.81
    9. in0.55
    10. D0.56
    11. Gemma-4-£20-it0.78
    12. Audio Scoribe0.91
    13. rt0.58
    14. Google Al Edge Gallery0.96
    15. me Googleplex on interactive0.95
    16. W's are in the wond0.93
    17. 1.00
    18. AIE1.00
    19. 1.00
    20. Model on GPU0.98
    21. Show thinking0.94
    22. 1.00
    23. 1.00
    24. 1.00
    25. 1.00
    26. Try Gemma 4 today1.00
    27. Agent Skils or the une cases below0.87
    28. ima 4 E28 6 E48 are herel Try them in Al Chat.0.92
    29. for the Googleplex0.99
    30. The interactive map has been shown1.00
    31. 2. Examine the wont0.89
    32. t. Analyze the request: The on0.94
    33. mood board. Write a comprehensive0.97
    34. sg brel for m ew cofebrand.0.62
    35. nthesize the aesthetic from this0.97
    36. Model on GPU0.98
    37. Explore other use cases1.00
    38. Agent Skills0.89
    39. Al Chat0.96
    40. 3. Coune the Ww:0.84
    41. 4. Verify the cout:0.91
    42. 5. Formulate the answer: Stute the0.97
    43. Strawbeery0.97
    44. -R10.80
    45. #(2)0.86
    46. -R00.75
    47. STRAWBERRY1.00
    48. Court: 1 (in strien0.84
    49. vberry) + 2 in0.86
    50. Model on GPU0.98
    51. Brand Name] - A Sensory0.99
    52. Experience1.00
    53. Brand Essence: This brand aims to0.98
    54. Design Brief: [Your Coffee0.99
    55. & Artisanal Coffee0.99
    56. positon itself as a premium, artisanal0.96
    57. coffee experience that blends the0.98
    58. warmth and ritual of traditional0.97
    59. craftsmanship with the vibrant,0.96
    60. Mountain View où je pourrais aller0.91
    61. après cet enregistrement. English: Im0.96
    62. could go after this registration.0.99
    63. faim et aidez-moi à choisir un0.98
    64. restaurant français suthentique à0.96
    65. me choose an authentic French1.00
    66. restaurant in Mountain View where 10.98
    67. natural vitality of the source. The0.99
    68. X€ View in ful scren0.85
    69. There are 3 'Rs in the word0.98
    70. aesthetic is deeply rooted in nature.0.97
    71. Type prompt..0.92
    72. Type prompt1.00
    73. Type prompt..0.93
    74. Type prompt..0.95
    75. skls0.60
    76. Google1.00
    77. Engineering the future of Al1.00
    78. AIEn0.91
    79. EUR0.98
  • 6:02 #16 skipped

    shot 16·duplicate of #15

  • 6:21 #17 done34 line(s)

    shot 17·sharpness 2262.8

    1. NEW1.00
    2. Agent Skills & Gemma 4 E2B & E4B0.95
    3. 12:300.93
    4. ('System GenAl' uses AICore when available)0.98
    5. Google Al Edge Gallery0.99
    6. Google Al Edge Gallery0.98
    7. Discover the power of n-device Al models from0.98
    8. Gemma 40.89
    9. Download1.00
    10. Try Gemma 4 today0.97
    11. AIE1.00
    12. Android1.00
    13. from1.00
    14. Agent Skls, orthe use cases below.0.72
    15. Gemma 4 E2B & E4B are herel Try them in Al Chat.0.97
    16. 1.00
    17. Play Store0.99
    18. 1.00
    19. 1.00
    20. 1.00
    21. Al Chat0.93
    22. Download1.00
    23. Agent Skills0.97
    24. fromiOS1.00
    25. App Store1.00
    26. Explore other use cases1.00
    27. View/clone1.00
    28. code on1.00
    29. Ask Image1.00
    30. Audio Scribe1.00
    31. Github1.00
    32. AlEngineer0.96
    33. EUROPE1.00
    34. AIEn0.92
  • 6:41 #18 done18 line(s)

    shot 18·sharpness 2162.9

    1. NEW1.00
    2. Agent Skills & Gemma 4 E2B & E0.96
    3. ('System GenAl' uses AICore when avail0.97
    4. Download1.00
    5. AIE1.00
    6. from1.00
    7. Android1.00
    8. Play Store0.98
    9. 1.00
    10. 1.00
    11. Download1.00
    12. fromiOS1.00
    13. App Store1.00
    14. View/clone1.00
    15. code on0.99
    16. Github1.00
    17. Engineering the future of Al0.99
    18. Al0.92
  • 6:46 #19 done11 line(s)

    shot 19·sharpness 1507.6

    1. Example Skills - Restaurant Roulette0.99
    2. What's new in Gemma 40.98
    3. AIE1.00
    4. What'snewin1.00
    5. 1.00
    6. 1.00
    7. Gemma41.00
    8. Watch onYoullube0.84
    9. Google1.00
    10. .com/watch7v=|ZVBoFOJK-Q0.92
    11. Engineering the future of Al0.99
  • 7:05 #20 done18 line(s)

    shot 20·sharpness 1576.4

    1. Example Skills - Restaurant Roulette0.99
    2. 0Gemma-4-E20-1t0.80
    3. Agent Skils0.98
    4. Model on GPU0.98
    5. AIE1.00
    6. 1.00
    7. The restaurant rulett weel for a0.90
    8. French restaurant in San Francisco0.96
    9. 1.00
    10. 1.00
    11. 1.00
    12. has been generated.1.00
    13. French spots in San Francisco0.99
    14. View in fl scren0.78
    15. Type prompt..0.94
    16. Google1.00
    17. Engineering the future of Al1.00
    18. AIEn0.83
  • 7:13 #21 skipped

    shot 21·duplicate of #19

  • 7:28 #22 done75 line(s)

    shot 22·sharpness 1568.4

    1. Dictionary1.00
    2. File0.95
    3. Edit1.00
    4. Go0.99
    5. Search1.00
    6. Window1.00
    7. Help1.00
    8. 80.81
    9. Thu Apr 9 4:54 AM0.99
    10. Displays1.00
    11. 9x490.95
    12. reposit1.00
    13. Q Search0.93
    14. Dictionary1.00
    15. Search0.98
    16. -0...17.29PM 2026-0..16.58PMM0.89
    17. creenshot0.97
    18. Screenshot1.00
    19. 2026-0...31.50PM0.99
    20. Read1.00
    21. Dictlonary0.99
    22. Thesaurus1.00
    23. Apple1.00
    24. Wikipedla0.99
    25. Directo1.00
    26. I'1l ac0.81
    27. Type a word to look up in...0.97
    28. inshot0.99
    29. .17.43PM0.94
    30. 2026-0....21.06PM0.98
    31. 2026-0...29.17AMM0.97
    32. skill s0.80
    33. /privat1.00
    34. _party/0.95
    35. Use as1.00
    36. Optimiù0.84
    37. Oxford American Writer's Thesaurus0.98
    38. New Oxford American Dictionary1.00
    39. ird0.99
    40. Showing1.00
    41. Apple Dictionary0.99
    42. AIE1.00
    43. 1.00
    44. 1.00
    45. /privat1.00
    46. party/1.00
    47. Refrest0.95
    48. Color p0.95
    49. Wikipedia1.00
    50. ird0.90
    51. 46.38PM1.00
    52. 1.00
    53. 1.00
    54. 1.00
    55. 1.00
    56. Shel0.99
    57. · I'll in0.89
    58. .04.17PM1.00
    59. Action0.98
    60. Shel1.00
    61. les/go0.85
    62. 04.33PM1.00
    63. ird_par0.94
    64. 2026-0...07.41PM0.88
    65. Screenshot1.00
    66. inshot1.00
    67. 2026-0.16.06AM 2026-0...53.02PM0.95
    68. Screenshot1.00
    69. Screenshot0.99
    70. Arrange1.00
    71. creenshot1.00
    72. -0...35.04PMPM0.86
    73. Aa1.00
    74. Engineering the future of Al0.99
    75. All0.78
  • 7:31 #23 skipped

    shot 23·duplicate of #14

Transcript

252 cues· 3,628 words· 18,759 chars

  1. 0:15 Yeah, so while we wait for it to come up, because I know we're short of time, I'm going to talk about agents on device.
  2. 0:22 So I know whoever asked the question about skills and AI core, we have an answer to that.
  3. 0:27 We've built a simple skill harness on top of AI core that you can build skills on, be able to show that.
  4. 0:31 Also going to talk about tiny LLMs, which are
  5. 0:36 We would call LLMs that are smaller than a billion parameters that are small enough to build into your app if you want to have more customization or you want to do something that isn't already available for you in AI Core.
  6. 0:46 So that's the gist.
  7. 0:48 So a quick overview of AI Edge, how we think about small language models, tiny LLMs, and system gen AI.
  8. 0:57 Then we're going to take a quick look at agent skills, which is something we can build on top of kind of system gen AI or the new models that are coming down the pipe.
  9. 1:05 And then we're going to take a quick look at tiny models.
  10. 1:10 So that's that.
  11. 1:15 Cool.
  12. 1:15 Yeah.
  13. 1:16 OK. Yeah.
  14. 1:18 Feel free.
  15. 1:21 OK.
  16. 1:22 So AI Edge, SLMs, and TLMs.
  17. 1:24 OK.
  18. 1:25 So I think Ali already covered this.
  19. 1:27 We know it's great to do things on device, latency, privacy, offline use, reliability, or savings, depending on things.
  20. 1:34 These are all motivations to do things locally.
  21. 1:38 Me, by way of intro, didn't really do this.
  22. 1:41 I'm a software engineer and kind of tech lead working on the Google AI Edge stack.
  23. 1:46 So we have MediaPipe, which is an asset some people may be familiar with.
  24. 1:50 We have LIDAR TLM, which is a LM harness that you can integrate with your app.
  25. 1:56 where you download the model and ship the model with your app.
  26. 1:59 And then we also have kind of LightRT as a runtime that supports both LightRT-LM and MediaPipe.
  27. 2:04 It's kind of formally known as TensorFlow Lite, which is a kind of cross-framework runtime for running models.
  28. 2:10 And all of that can run on CPU, GPU, or NPU, depending on the platform and depending what's best.
  29. 2:16 And you as a developer get to choose.
  30. 2:18 Yeah, it's already trusted at scale.
  31. 2:21 Like the Lido RT runtime, there's a version of that built into Android OS.
  32. 2:25 Lots of Android apps already use it.
  33. 2:27 So it does support over 2.7 billion devices, like lots and lots of daily invocations, and lots and lots of Android apps leverage this.
  34. 2:37 but also works far beyond Android as well.
  35. 2:39 So we support all of these platforms.
  36. 2:42 And for example, Gemma is available on many of these platforms.
  37. 2:46 Our team is giving another talk tomorrow, so you can hear more about Gemma performance on all of these types of platforms and how we're able to do really useful things with the latest Gemma 4 models.
  38. 2:59 But then building on Ollie's and Florina's talk, this is kind of key idea is we have system level gen AI, which is something that will be pre-installed in the system.
  39. 3:09 So there's Gemini Nano via AI Core.
  40. 3:11 This is an example of the summarization API.
  41. 3:14 Apple also has something going on with their intelligence on iOS that I probably know a lot less about.
  42. 3:20 But as a concept, right, as an app developer, when you go to build a mobile app, this is kind of one choice is there will often be some form of intelligence built into the system that you can leverage, which is, you know, highly optimized, as kind of Alia and Florina covered, that's available for use with your app.
  43. 3:40 Then so this is kind of typically like small language models like for for nano It is the Gemma for e2b and e4b are the base models for what we ship there That's really capable highly optimized preloaded with device if you can use those it's great Your app doesn't get any bigger and if it meets your use case needs it's a great place to start if you want like more
  44. 4:03 If you have a more specific task that you want to do that's kind of highly customized or something really boutique, you can use an AppGen AI.
  45. 4:12 So that's with the Lider TLM runtime.
  46. 4:15 That can be loaded with your app or even your web page.
  47. 4:19 And this offers kind of a higher degree of customization and reach.
  48. 4:23 Definitely more work.
  49. 4:24 But yeah, you have access to smaller models that can run on lots of devices and full customization.
  50. 4:32 So it's clearly a lot more work.

Chapters

  1. 0:00 Introduction to on-device agents and tiny LLMs
  2. 0:48 Overview of AI Edge, SLMs, and TLMs
  3. 0:57 Taking a look at agent skills
  4. 1:06 Taking a look at tiny models
  5. 1:24 Motivations for on-device AI (latency, privacy, offline use)
  6. 3:01 System-level GenAI (Gemini Nano via AI Core)
  7. 4:03 App-level GenAI (LiteRT-LM for custom/boutique models)
  8. 5:06 Google AI Edge Gallery app demo
  9. 6:22 Deep dive into agent skills and the skill harness
  10. 7:41 How the skill harness works (system prompts, tool calls, and JavaScript UI)
  11. 9:00 Creating and publishing your own skills
  12. 10:28 Using LiteRT-LM runtime for model deployment
  13. 12:31 Export and inference workflow (from PyTorch to deployment)
  14. 13:19 Function Gemma: Robust, small-scale function calling
  15. 14:35 Fine-tuning workflow for tiny models using synthetic data
  16. 16:01 Eloquent: A production transcription app example using tiny models
  17. 17:28 Q&A: Agent skill robustness and multi-skill calling
  18. 19:26 Q&A: LiteRT-LM file format vs. Task files
  19. 20:00 Q&A: Performance on CPU/TPU and resources

Open at this second