Videos -TiET_K-E_g
From 46% to 90%: Fine-Tuning Tiny LLMs for On-Device Agents — Cormac Brick, Google
Scene timeline
75 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 252
- whisperx 252
- chunks
- 38
- from 252 cues
- keyframes
- 43
- kept of 75 captured
- frames with text
- 43
- 976 lines read
- chapters
- 19
- from the source metadata
- keyframe bytes
- 7.8 MB
- word timings on 252 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 11:41 | 2m 03s |
stt |
done | — | 2026-08-10 11:43 | 26s |
chunk |
done | — | 2026-08-10 11:44 | 0s |
text_embed |
done | — | 2026-08-10 19:49 | 0s |
keyframe |
done | — | 2026-08-10 11:44 | 1m 53s |
ocr |
done | — | 2026-08-10 11:46 | 16s |
frame_embed |
done | — | 2026-08-10 19:49 | 8s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Google DeepMind0.98
- APRIL 9, 2026 / MORNING BLOCK0.96
- 4TH FLOOR- RUTHERFORD0.98
- 11:15AM1.00
- The agent-ready web: Simplify user actions with WebMCP1.00
- Tara Agyemang / Google DeepMind0.98
- 11:40AM0.99
- Build & deploy Al-powered apps0.99
- Paige Bailey / Google DeepMind0.98
- 12:00PH0.95
- Agentic Evaluations at Scale—For Everybody1.00
- Nicholas Kang, Michael Aaron / Google DeepMind1.00
- 12:20PM0.99
- Al on Android: Ask me Anything0.99
- Florina Muntenescu, Oli Gaymond / Google DeepMind0.99
- 12:40PM0.98
- TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM1.00
- Cormac Brick/ Google0.99
- Schedule updates may occur. Stay up to date via our app or website0.99
- https://al.engineer/schedule1.00
-
- AI ENGINEER CONFERENCE 20261.00
- TLMs: Tiny LLMs and0.98
- Agents on Edge Devices1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Bringing state-of-the-art agentic skills to the1.00
- edge with open models0.97
- Cormac Brick, Principal Engineer, Google Al Edge0.99
- Google DeepMind1.00
- AIEn0.93
-
- Agenda1.00
- 00 AI Edge, SLMs & TLMs0.95
- 10 Agent Skills locally on Android & iOS0.98
- AIE1.00
- 20 TLM workflow1.00
- ★1.00
- ★1.00
- 30 TLMs in Action0.98
- Google1.00
- Braintrust1.00
- WorkOS OpenAI0.94
-
- Running Al on the Edge has many benefits0.98
- AIE1.00
- $0.76
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Latency/ UX0.96
- Privacy1.00
- Offline use0.97
- Savings1.00
- Faster, no network0.99
- Sensitive data not sent1.00
- Works without requiring1.00
- Lower or no1.00
- involved1.00
- off device1.00
- cellular1.00
- data-center costs1.00
- AlEngineer0.91
- EUROPE1.00
- AlEng0.90
-
- Google Al Edge0.96
- Your App1.00
- **0.84
- AIE1.00
- MediaPipe1.00
- ★1.00
- ★1.00
- LiteRT-LM1.00
- LiteRT1.00
- (FKA TensorFlow Lite)1.00
- CPU1.00
- GPU1.00
- NPU1.00
- AlEngineer0.96
- EUROPE1.00
- AIEr0.86
-
- Trusted at scale1.00
- AIE1.00
- 100,000+1.00
- 2.7 billion +0.98
- 1 trillion +0.99
- ★1.00
- ★1.00
- Android apps1.00
- Devices1.00
- Average daily1.00
- interpreter invocations in1.00
- Android1.00
- Source: Internal Android data1.00
- Engineering the future of Al1.00
- AlEng0.92
- EURO0.94
-
- ...and far beyond Android0.97
- ***0.61
- .tflite0.99
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- CPU1.00
- GPU1.00
- NPU1.00
- i050.97
- Android1.00
- iOS0.89
- macOS0.97
- Linux1.00
- Windows1.00
- Web1.00
- loT0.93
- Engineering the future of Al0.99
- AlEngi0.90
- EURO1.00
-
- System level GenAl0.98
- AIE1.00
- ★1.00
- ★1.00
- Gemini Nano with AICore0.97
- Apple Intelligence Writing0.99
- summarization on Android1.00
- tools on iOS0.97
- Engineering the future of Al0.99
- AlEng0.96
- EURC0.93
-
- System vs In-app GenAl0.98
- *★*0.52
- AIE1.00
- System GenAl0.97
- In-app GenAI0.98
- ★1.00
- ★1.00
- ★1.00
- Most capable1.00
- Customized to tasks1.00
- highly optimized1.00
- Loaded with app / webpage0.99
- Pre-loaded with device1.00
- Customization & Reach1.00
- SLMs: 2B-4B0.99
- TLMs: 100M-1B0.99
- AlEngineer0.95
- EUROPE1.00
- AIEn0.92
-
- 201.00
- AIE1.00
- Agent Skills on Device0.95
- ★1.00
- (systemGenAl)0.98
- Google1.00
- Engineering the future of Al1.00
- AlEng0.95
- EURD0.85
-
- Google Al Edge Gallery App: On-device Al in Action0.99
- Featuring Gemma41.00
- Google AI Edge Gallery0.97
- Gemma-4-E28-0.75
- Gomma-4-E20-it0.74
- Al Chat0.91
- Ask Image0.95
- Gemma-4-E28-it0.81
- in0.55
- D0.56
- Gemma-4-£20-it0.78
- Audio Scoribe0.91
- rt0.58
- Google Al Edge Gallery0.96
- me Googleplex on interactive0.95
- W's are in the wond0.93
- ★1.00
- AIE1.00
- ★1.00
- Model on GPU0.98
- Show thinking0.94
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Try Gemma 4 today1.00
- Agent Skils or the une cases below0.87
- ima 4 E28 6 E48 are herel Try them in Al Chat.0.92
- for the Googleplex0.99
- The interactive map has been shown1.00
- 2. Examine the wont0.89
- t. Analyze the request: The on0.94
- mood board. Write a comprehensive0.97
- sg brel for m ew cofebrand.0.62
- nthesize the aesthetic from this0.97
- Model on GPU0.98
- Explore other use cases1.00
- Agent Skills0.89
- Al Chat0.96
- 3. Coune the Ww:0.84
- 4. Verify the cout:0.91
- 5. Formulate the answer: Stute the0.97
- Strawbeery0.97
- -R10.80
- #(2)0.86
- -R00.75
- STRAWBERRY1.00
- Court: 1 (in strien0.84
- vberry) + 2 in0.86
- Model on GPU0.98
- Brand Name] - A Sensory0.99
- Experience1.00
- Brand Essence: This brand aims to0.98
- Design Brief: [Your Coffee0.99
- & Artisanal Coffee0.99
- positon itself as a premium, artisanal0.96
- coffee experience that blends the0.98
- warmth and ritual of traditional0.97
- craftsmanship with the vibrant,0.96
- Mountain View où je pourrais aller0.91
- après cet enregistrement. English: Im0.96
- could go after this registration.0.99
- faim et aidez-moi à choisir un0.98
- restaurant français suthentique à0.96
- me choose an authentic French1.00
- restaurant in Mountain View where 10.98
- natural vitality of the source. The0.99
- X€ View in ful scren0.85
- There are 3 'Rs in the word0.98
- aesthetic is deeply rooted in nature.0.97
- Type prompt..0.92
- Type prompt1.00
- Type prompt..0.93
- Type prompt..0.95
- skls0.60
- Google1.00
- Engineering the future of Al1.00
- AIEn0.91
- EUR0.98
-
- NEW1.00
- Agent Skills & Gemma 4 E2B & E4B0.95
- 12:300.93
- ('System GenAl' uses AICore when available)0.98
- Google Al Edge Gallery0.99
- Google Al Edge Gallery0.98
- Discover the power of n-device Al models from0.98
- Gemma 40.89
- Download1.00
- Try Gemma 4 today0.97
- AIE1.00
- Android1.00
- from1.00
- Agent Skls, orthe use cases below.0.72
- Gemma 4 E2B & E4B are herel Try them in Al Chat.0.97
- ★1.00
- Play Store0.99
- ★1.00
- ★1.00
- ★1.00
- Al Chat0.93
- Download1.00
- Agent Skills0.97
- fromiOS1.00
- App Store1.00
- Explore other use cases1.00
- View/clone1.00
- code on1.00
- Ask Image1.00
- Audio Scribe1.00
- Github1.00
- AlEngineer0.96
- EUROPE1.00
- AIEn0.92
-
- NEW1.00
- Agent Skills & Gemma 4 E2B & E0.96
- ('System GenAl' uses AICore when avail0.97
- Download1.00
- AIE1.00
- from1.00
- Android1.00
- Play Store0.98
- ★1.00
- ★1.00
- Download1.00
- fromiOS1.00
- App Store1.00
- View/clone1.00
- code on0.99
- Github1.00
- Engineering the future of Al0.99
- Al0.92
-
- Example Skills - Restaurant Roulette0.99
- What's new in Gemma 40.98
- AIE1.00
- What'snewin1.00
- ★1.00
- ★1.00
- Gemma41.00
- Watch onYoullube0.84
- Google1.00
- .com/watch7v=|ZVBoFOJK-Q0.92
- Engineering the future of Al0.99
-
- Example Skills - Restaurant Roulette0.99
- 0Gemma-4-E20-1t0.80
- Agent Skils0.98
- Model on GPU0.98
- AIE1.00
- ★1.00
- The restaurant rulett weel for a0.90
- French restaurant in San Francisco0.96
- ★1.00
- ★1.00
- ★1.00
- has been generated.1.00
- French spots in San Francisco0.99
- View in fl scren0.78
- Type prompt..0.94
- Google1.00
- Engineering the future of Al1.00
- AIEn0.83
-
- Dictionary1.00
- File0.95
- Edit1.00
- Go0.99
- Search1.00
- Window1.00
- Help1.00
- 80.81
- Thu Apr 9 4:54 AM0.99
- Displays1.00
- 9x490.95
- reposit1.00
- Q Search0.93
- Dictionary1.00
- Search0.98
- -0...17.29PM 2026-0..16.58PMM0.89
- creenshot0.97
- Screenshot1.00
- 2026-0...31.50PM0.99
- Read1.00
- Dictlonary0.99
- Thesaurus1.00
- Apple1.00
- Wikipedla0.99
- Directo1.00
- I'1l ac0.81
- Type a word to look up in...0.97
- inshot0.99
- .17.43PM0.94
- 2026-0....21.06PM0.98
- 2026-0...29.17AMM0.97
- skill s0.80
- /privat1.00
- _party/0.95
- Use as1.00
- Optimiù0.84
- Oxford American Writer's Thesaurus0.98
- New Oxford American Dictionary1.00
- ird0.99
- Showing1.00
- Apple Dictionary0.99
- AIE1.00
- ★1.00
- ★1.00
- /privat1.00
- party/1.00
- Refrest0.95
- Color p0.95
- Wikipedia1.00
- ird0.90
- 46.38PM1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Shel0.99
- · I'll in0.89
- .04.17PM1.00
- Action0.98
- Shel1.00
- les/go0.85
- 04.33PM1.00
- ird_par0.94
- 2026-0...07.41PM0.88
- Screenshot1.00
- inshot1.00
- 2026-0.16.06AM 2026-0...53.02PM0.95
- Screenshot1.00
- Screenshot0.99
- Arrange1.00
- creenshot1.00
- -0...35.04PMPM0.86
- Aa1.00
- Engineering the future of Al0.99
- All0.78
Transcript
252 cues· 3,628 words· 18,759 chars
- 0:15 Yeah, so while we wait for it to come up, because I know we're short of time, I'm going to talk about agents on device.
- 0:22 So I know whoever asked the question about skills and AI core, we have an answer to that.
- 0:27 We've built a simple skill harness on top of AI core that you can build skills on, be able to show that.
- 0:31 Also going to talk about tiny LLMs, which are
- 0:36 We would call LLMs that are smaller than a billion parameters that are small enough to build into your app if you want to have more customization or you want to do something that isn't already available for you in AI Core.
- 0:46 So that's the gist.
- 0:48 So a quick overview of AI Edge, how we think about small language models, tiny LLMs, and system gen AI.
- 0:57 Then we're going to take a quick look at agent skills, which is something we can build on top of kind of system gen AI or the new models that are coming down the pipe.
- 1:05 And then we're going to take a quick look at tiny models.
- 1:10 So that's that.
- 1:15 Cool.
- 1:15 Yeah.
- 1:16 OK. Yeah.
- 1:18 Feel free.
- 1:21 OK.
- 1:22 So AI Edge, SLMs, and TLMs.
- 1:24 OK.
- 1:25 So I think Ali already covered this.
- 1:27 We know it's great to do things on device, latency, privacy, offline use, reliability, or savings, depending on things.
- 1:34 These are all motivations to do things locally.
- 1:38 Me, by way of intro, didn't really do this.
- 1:41 I'm a software engineer and kind of tech lead working on the Google AI Edge stack.
- 1:46 So we have MediaPipe, which is an asset some people may be familiar with.
- 1:50 We have LIDAR TLM, which is a LM harness that you can integrate with your app.
- 1:56 where you download the model and ship the model with your app.
- 1:59 And then we also have kind of LightRT as a runtime that supports both LightRT-LM and MediaPipe.
- 2:04 It's kind of formally known as TensorFlow Lite, which is a kind of cross-framework runtime for running models.
- 2:10 And all of that can run on CPU, GPU, or NPU, depending on the platform and depending what's best.
- 2:16 And you as a developer get to choose.
- 2:18 Yeah, it's already trusted at scale.
- 2:21 Like the Lido RT runtime, there's a version of that built into Android OS.
- 2:25 Lots of Android apps already use it.
- 2:27 So it does support over 2.7 billion devices, like lots and lots of daily invocations, and lots and lots of Android apps leverage this.
- 2:37 but also works far beyond Android as well.
- 2:39 So we support all of these platforms.
- 2:42 And for example, Gemma is available on many of these platforms.
- 2:46 Our team is giving another talk tomorrow, so you can hear more about Gemma performance on all of these types of platforms and how we're able to do really useful things with the latest Gemma 4 models.
- 2:59 But then building on Ollie's and Florina's talk, this is kind of key idea is we have system level gen AI, which is something that will be pre-installed in the system.
- 3:09 So there's Gemini Nano via AI Core.
- 3:11 This is an example of the summarization API.
- 3:14 Apple also has something going on with their intelligence on iOS that I probably know a lot less about.
- 3:20 But as a concept, right, as an app developer, when you go to build a mobile app, this is kind of one choice is there will often be some form of intelligence built into the system that you can leverage, which is, you know, highly optimized, as kind of Alia and Florina covered, that's available for use with your app.
- 3:40 Then so this is kind of typically like small language models like for for nano It is the Gemma for e2b and e4b are the base models for what we ship there That's really capable highly optimized preloaded with device if you can use those it's great Your app doesn't get any bigger and if it meets your use case needs it's a great place to start if you want like more
- 4:03 If you have a more specific task that you want to do that's kind of highly customized or something really boutique, you can use an AppGen AI.
- 4:12 So that's with the Lider TLM runtime.
- 4:15 That can be loaded with your app or even your web page.
- 4:19 And this offers kind of a higher degree of customization and reach.
- 4:23 Definitely more work.
- 4:24 But yeah, you have access to smaller models that can run on lots of devices and full customization.
- 4:32 So it's clearly a lot more work.
loading
Chapters
- 0:00 Introduction to on-device agents and tiny LLMs
- 0:48 Overview of AI Edge, SLMs, and TLMs
- 0:57 Taking a look at agent skills
- 1:06 Taking a look at tiny models
- 1:24 Motivations for on-device AI (latency, privacy, offline use)
- 3:01 System-level GenAI (Gemini Nano via AI Core)
- 4:03 App-level GenAI (LiteRT-LM for custom/boutique models)
- 5:06 Google AI Edge Gallery app demo
- 6:22 Deep dive into agent skills and the skill harness
- 7:41 How the skill harness works (system prompts, tool calls, and JavaScript UI)
- 9:00 Creating and publishing your own skills
- 10:28 Using LiteRT-LM runtime for model deployment
- 12:31 Export and inference workflow (from PyTorch to deployment)
- 13:19 Function Gemma: Robust, small-scale function calling
- 14:35 Fine-tuning workflow for tiny models using synthetic data
- 16:01 Eloquent: A production transcription app example using tiny models
- 17:28 Q&A: Agent skill robustness and multi-skill calling
- 19:26 Q&A: LiteRT-LM file format vs. Task files
- 20:00 Q&A: Performance on CPU/TPU and resources