Videos KB41dTlX1Uc
State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
Scene timeline
123 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 569
- whisperx 569
- chunks
- 80
- from 569 cues
- keyframes
- 105
- kept of 123 captured
- frames with text
- 104
- 157 lines read
- chapters
- 23
- from the source metadata
- keyframe bytes
- 13.2 MB
- word timings on 569 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 17:17 | 3m 13s |
stt |
done | — | 2026-08-10 17:21 | 53s |
chunk |
done | — | 2026-08-10 17:21 | 0s |
text_embed |
done | — | 2026-08-10 19:53 | 1s |
keyframe |
done | — | 2026-08-10 17:21 | 5m 11s |
ocr |
done | — | 2026-08-10 17:27 | 14s |
frame_embed |
done | — | 2026-08-10 19:53 | 18s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.93
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize0.94
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- roboflo0.92
- exo1.00
-
- roboflow1.00
-
- roboflow1.00
-
- roboflow1.00
-
- roboflow1.00
-
- robafton0.69
- exo1.00
-
- exo1.00
- 110.93
-
- roboflow1.00
-
- roboflow1.00
-
- exo1.00
-
- roboflow1.00
-
- roboflow1.00
-
- exo1.00
-
- roboflow1.00
-
- roboflow1.00
-
- roboflow0.83
- exo1.00
-
- exo1.00
-
- exo1.00
-
- exo1.00
Transcript
569 cues· 8,542 words· 45,642 chars
- 0:12 Can you guys hear us?
- 0:14 Sound check?
- 0:14 All right.
- 0:15 Give it up for Local AI, everyone.
- 0:23 I hope you guys are excited as we are.
- 0:25 This is the local AI summit.
- 0:28 So we're gonna be here all day talking about local AI.
- 0:31 And the reason why is we hit an inflection point this year.
- 0:35 Not only did the models get really good, but the harnesses got really good.
- 0:39 And this happened really fast.
- 0:41 It's been, I think, a struggle for anyone here to keep up.
- 0:44 I felt that most when I saw one of Andrej Karpathy's tweets.
- 0:47 In November, he tweeted that you can't really trust these coding agents alone yet.
- 0:50 You have to monitor them with an eye like a hawk.
- 0:53 Three months later, he tweets that he's struggling to keep up with the capabilities of how good this all has gone.
- 0:59 And the thing is, both times he was right.
- 1:01 This space is progressing really quickly, and it's wild how much he can do.
- 1:05 And honestly, the only thing to do is just try to use it a little bit more today than you did yesterday.
- 1:10 That's how to keep up with this space.
- 1:11 And that's exactly what you guys are doing right here.
- 1:14 And so I'm really excited about this panel.
- 1:16 The way that we use AI has also changed.
- 1:18 I'm not just using chat bots.
- 1:20 I'm not just asking simple questions.
- 1:22 When we got reasoning models, the profile of how the AI model responded changed.
- 1:28 Not only is it bursty and responding to me, but before the burst, it kind of plateaus for a bit.
- 1:33 It's reasoning.
- 1:34 It's churning on tokens that I'm not consuming.
- 1:38 And then we got agents.
- 1:39 And suddenly, I don't even want these agents to turn off.
- 1:42 I want these always on agents.
- 1:44 They can always be productive if we set them up.
- 1:47 And so we have enterprises that want to put a lot of their IP into this because it becomes more useful.
- 1:53 We have consumers with the same thing.
- 1:55 I want to give it my health data, my medical records.
- 1:58 I want to give it footage from my home camera.
- 2:00 And both enterprises and consumers, we don't want that stuff to leak.
- 2:04 And as you also have a profile of tokens continuously generating, suddenly costs matter.
- 2:11 So local is amazing for both of those things.
- 2:14 You get to make sure that you are plateaued on the costs for the tokens that you're generating.
- 2:20 And also, everything sits in that room.
- 2:22 So we have amazing demos here after these talks where everything that's being run stays on those devices.
- 2:28 It stays in this room.
- 2:29 And so that's a really nice guarantee.
- 2:31 So as we turn over the panels, first, do you guys want to introduce yourselves?
- 2:36 Yeah, should I start, please?
- 2:37 You got the mic?
- 2:38 Yeah, your mic's up.
- 2:40 So yeah, I'm Alex.
- 2:42 I'm the co-founder and CEO of ExoLabs and also the creator of Local.ai.
loading
Chapters
- 0:00 Welcome to the Local AI Summit
- 0:40 Karpathy twice right on keeping up
- 1:16 Reasoning models and always on agents
- 2:32 Panelist introductions
- 4:41 When the inflection point hit
- 6:36 GPT 4o quality in your pocket
- 7:14 Llama 405B to DeepSeek to GLM 5.2
- 8:47 The airplane accessibility story
- 10:25 Harnesses give models the real world
- 11:27 What language learns from vision
- 13:19 A multimodel world in practice
- 13:57 Coinbase: tokens up, costs flat
- 15:03 Control, sovereignty, no rug pulls
- 17:50 Small specialized models and data flywheels
- 19:45 A second headquarters inside NVIDIA
- 21:50 10x on the DGX Spark by swarming
- 24:32 Desk and data center share an architecture
- 26:11 ODS and point and click onboarding
- 27:14 Where local still falls short
- 32:33 Why finetuning as a service stalled
- 35:34 Distillation down to a submarine
- 39:42 The biggest open problems in local
- 42:01 Open source advocacy and closing