Videos QHBjufYK8TA
The State of Model Routing — NVIDIA, Cognition, OpenRouter
Scene timeline
218 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 515
- whisperx 515
- chunks
- 85
- from 515 cues
- keyframes
- 176
- kept of 218 captured
- frames with text
- 66
- 137 lines read
- chapters
- 19
- from the source metadata
- keyframe bytes
- 27.6 MB
- word timings on 515 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:16 | 0s |
stt |
done | — | 2026-08-09 00:07 | 1m 01s |
chunk |
done | — | 2026-08-09 00:08 | 0s |
text_embed |
done | — | 2026-08-10 19:35 | 2s |
keyframe |
done | — | 2026-08-09 00:09 | 9m 06s |
ocr |
done | — | 2026-08-09 00:18 | 31s |
frame_embed |
done | — | 2026-08-10 19:35 | 29s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.97
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- 3:20PM-4:05PM1.00
- Model Routing0.98
- Nader Khalil0.97
- Walden Yan1.00
- Tanay Varshney0.98
- Alex Atallah1.00
- Matthew Berman1.00
- Director of Developer Tech1.00
- Co-Founder1.00
- Principal Engineer (IC6)1.00
- Cofounder & CEO0.99
- CEO / Co-founder0.95
- NVIDIA.0.97
- Cognition1.00
- NVIDIA.0.95
- OpenRouter orwardFuture0.99
-
- Runlayor0.89
- NOR0.51
- Speaker1.00
- Tanay1.00
-
- Runlayer0.98
- Runiayer0.95
- Speaker1.00
-
- 治0.51
Transcript
515 cues· 8,334 words· 45,256 chars
- 0:12 Those have been really exciting.
- 0:13 We've tried to get a bunch of the industry leaders together to talk about some of the problems that we're facing as we try to run more on local.
- 0:21 If you guys were here for the first panel, one of the things that we talked about was model routing.
- 0:26 We firmly believe that we're in a multi-model world.
- 0:29 I think you heard this from many of the panelists.
- 0:32 Anyone who is deploying AI in production and who is doing so locally is seeing that multi-model world.
- 0:37 That's why we released these Nemotron models at NVIDIA.
- 0:40 Everything is released from the data sets to the weights with recipes so that you can customize them.
- 0:44 We do that because we know that people customizing models is going to be huge.
- 0:49 And so this panel is really exciting because we're going to talk specifically about model routing.
- 0:54 So as you are picking which model to use, how does that how essentially how does that tooling itself look?
- 1:00 Do you guys want to introduce yourselves?
- 1:03 Yeah, sure.
- 1:04 I'm Walden.
- 1:04 I'm the co-founder of Cognition.
- 1:06 We build Devon, a software engineer.
- 1:08 In addition to the product, we spend a lot of time partnering with our customers to figure out how they should deploy these models and these agents.
- 1:15 And one of the things they're constantly asking us nowadays is basically, how do I know the ROI of our models?
- 1:20 And how do I know which tasks I can actually let our engineers spend the most expensive models on versus letting them use a more cost-efficient model?
- 1:29 And so that's why we're also thinking a lot more about multi-model routing nowadays.
- 1:32 Totally.
- 1:33 I'm Carter.
- 1:34 You guys heard from me a little bit earlier, but if you weren't here, I'm a developer tech engineer at Nvidia and ultimately I spent a lot of time thinking about how to get intelligence into as many developers hands as possible and something that is continually becoming a
- 1:49 not an issue, but something that is top of mind for a lot of developers is as you use more intelligence and the frontier models get more expensive, it becomes somewhat cost prohibitive to use the best tools, what feels like the best tools, as much as you would like to use them.
- 2:04 And so this has become a recent focus is how can we still get the same desired outputs, but actually both as an individual developer, but also imagine startups and small companies
- 2:16 How can you leverage this incredible tool without totally breaking the bank?
- 2:21 I'm Tane.
- 2:22 I work on model evaluations, both in terms of its accuracies and efficiency and cost understanding of the model.
- 2:31 And then I try and understand those, implement those learnings and help build a router.
- 2:37 So it's basically...
- 2:40 My job is to understand the behavior of the model on an intimate level and then use those learnings to both improve the model and try and design a system of model that can work together with each other.
- 2:52 Totally.
- 2:54 Yeah, I love a lot of the research that you're doing at NVIDIA as we kind of see the space through.
- 2:58 I think what's really interesting is model routing itself is pretty new still.
- 3:01 And so what you'll notice is there isn't a very clear solution here.
- 3:05 That was something that came up on the first panel is that there is a lot of space for startups and for companies in the ecosystem to fill in a solution here because we're still figuring out how to best do these patterns.
- 3:15 And I think, Walden, I want to kind of ask you.
- 3:17 So Cognition just released Fusion, your guys' model router.
- 3:21 And when you guys released it, you, in your blog, said that you're actually getting better performance than Fable, than these frontier models.
- 3:28 And I feel like that was a very surprising statement to hear, because we're thinking that you're getting as good or close enough, usually, when we're running on Edge, when we're running local, in these compute-strained, smaller footprint models.
- 3:40 But you guys are getting better.
- 3:41 Can you explain how?
- 3:43 Yeah, absolutely.
- 3:44 I also want to be clear about something here.
- 3:48 We're not saying that we gap above Fable-level performance in the same way that maybe Fable-level performance gaps above other models.
- 3:54 I think actually there's this really unintuitive dynamic where smarter models actually get better and better at delegating work.
- 4:05 we had with building model routers.
- 4:07 We don't want to route people to a dumber model and then suddenly you're stuck with a model that doesn't know how to do your task and next thing you know you're switching yourself back to a smarter model anyways and now taking that expensive cost.
- 4:18 In general we think a lot of the existing model routing systems out there are probably the same ones people have been using like a year ago and so we really wanted to put out a new framework that actually lets people
- 4:28 still feel like and still have a frontier model in their system while getting all these cost benefits.
loading
Chapters
- 0:00 Welcome and the multimodel premise
- 1:16 Panel introductions
- 3:24 How Devin Fusion beats the frontier models
- 4:25 Let the frontier model plan and delegate the work
- 6:31 Jagged capabilities: no one model wins everything
- 9:42 Why naive task based routing is fragile
- 11:48 Sharing context without paying for it twice
- 13:56 Should the orchestrator be the big or the small model
- 16:01 In distribution versus out of distribution
- 19:12 Training models to collaborate
- 20:12 Flex Run and flexible model sizes
- 22:24 Lossy context and the systems you fall back to
- 26:41 How a heartbeat created the auto router
- 29:43 Routing between local and cloud
- 31:51 Compaction versus routing
- 32:55 How a small model signals it is out of its depth
- 35:00 Cache duration and self hosting economics
- 40:12 Are prompts portable across models
- 43:20 Is the router a product or plumbing