Videos FWMJQDH3iK0
Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA
Scene timeline
124 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 348
- whisperx 348
- chunks
- 74
- from 348 cues
- keyframes
- 113
- kept of 124 captured
- frames with text
- 43
- 80 lines read
- chapters
- 27
- from the source metadata
- keyframe bytes
- 11.8 MB
- word timings on 348 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:12 | 0s |
stt |
done | — | 2026-08-08 23:41 | 49s |
chunk |
done | — | 2026-08-08 23:42 | 0s |
text_embed |
done | — | 2026-08-10 19:34 | 1s |
keyframe |
done | — | 2026-08-08 23:42 | 7m 30s |
ocr |
done | — | 2026-08-08 23:50 | 16s |
frame_embed |
done | — | 2026-08-10 19:34 | 19s |
Frames, and what the machine read
-
- AIEngineer0.95
- World's Fair0.96
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- Fair1.00
-
- TSAT0.76
- 小0.97
-
- 'sFaii0.88
-
- 香0.78
-
- haring1.00
-
- mle0.58
-
- Chris0.84
- 國0.55
-
- 养业0.73
- Speaker0.98
- Cheris0.80
-
- Speaker1.00
-
- a0.58
- Speaker1.00
-
- Speaker1.00
-
- S0.65
Transcript
348 cues· 7,661 words· 40,961 chars
- 0:12 So I hope everybody had a great lunch and you got to check out some of the amazing demos that we have.
- 0:18 We're going to begin the panel, the first panel of the afternoon here where we're going to be talking about, of course, the engines that are actually powering the stuff that could remotely be used for things like local sovereign, any kind of ownership over your own artificial intelligence.
- 0:34 And of course, the engine powering those in addition to the hardware is the models themselves.
- 0:38 And so for this panel, we have excellent guests.
- 0:41 We have
- 0:42 Vincent, who is the CEO and founder of Prime Intellect.
- 0:45 We've got Lucas, the CTO of RCAI, and we've got Chris, who is the senior product research engineer on the Nematron family of models at NVIDIA.
- 0:53 Now, what's really cool about working in this industry is really cool companies like this.
- 0:58 We all get to work together.
- 0:59 And so this is one panel where we all directly get to work together on both models, infrastructure, some of the ways that we think that the direction of the industry should go.
- 1:08 And each of us kind of play a different role in that stack.
- 1:11 But I wanna leave it to you guys to introduce yourself and be able to talk about sort of the charter that you see, the problem of the stack that you guys are working on.
- 1:20 Awesome, should I kick it off?
- 1:22 Kick it off.
- 1:22 Yeah, so I'm Vincent, as you mentioned.
- 1:24 And then really the goal with Prime Intellect from the beginning was like to ensure that basically frontier intelligence will be open and accessible, not just the models, but also the full stack to train the models.
- 1:36 So kind of like this was like our motivation from the beginning.
- 1:39 And we've worked also together with a lot of gentlemen on stage.
- 1:44 Like on the one side, it's like we work with folks like Lucas and Arcee to help them train frontier open models.
- 1:51 We help also...
- 1:53 like NVIDIA on the Nemotron coalition help out there from the open malls.
- 1:56 And I think I'm actually think both like Nemotron and Trinity might be the best like two open malls right now outside of China.
- 2:05 So I think it's actually like we need to fact check that, you know, but this is actually for my, I think they might be.
- 2:12 Our marketing says that, yeah.
- 2:15 But yeah, so that's the high level.
- 2:19 My name is Lucas Atkins.
- 2:20 I'm happy to be here, and thank you for joining.
- 2:23 Very similar to Vincent, RC was founded with the idea of domain-specific owned models are going to be needed.
- 2:34 We were founded early 2023, jumping on the custom model train quite early.
- 2:41 You have all these people who are excited about AI and all the things these new generation of LLMs can do, but they're using these monolithic, very expensive closed APIs for at the time and still very narrow tasks that don't require
- 2:57 You know, at the time, it was $100 per million tokens out.
- 3:01 And through doing that, we were building on top of open models, and we were releasing a lot of our tooling in the open.
- 3:07 And we noticed that in the United States and in the West, you know, in general, we were starting to lose
- 3:15 leadership in the open model space.
- 3:17 A lot of it was coming out of China, and that's amazing.
- 3:19 I love those models.
- 3:20 We learn a lot from them.
- 3:21 We're close with a lot of the people building those.
- 3:24 But when you're working with large enterprises and companies and geopolitics gets involved, whether you like it or not, you have people that become concerned about where those models are coming from.
- 3:37 decided that we had a good group of people and we had a good group of partners like NVIDIA and like Prime Intellect where we could probably try to pre-train ourselves.
- 3:46 So last year we did that.
- 3:47 We kind of reoriented the entire company towards let's figure out how to pre-train a 400 billion parameter model in six months.
- 3:56 A lot of people said it was impossible and in many ways it was, but we figured it out and now we are an open model lab working with our wonderful partners
- 4:05 AND OUR CUSTOMERS TO BUILD WESTERN OPEN MODELS THAT ARE PERMISSIVE AND YOU CAN OWN THOSE AND CUSTOMIZE THEM OR RUN THEM WHEREVER YOU WANT.
- 4:16 AND THAT'S KIND OF WHERE WE'RE AT RIGHT NOW.
- 4:17 SO THANKS FOR HAVING ME.
- 4:19 YEAH.
- 4:19 SO I'M CHRIS.
- 4:21 I WORK AT NVIDIA AS A PRODUCT RESEARCH ENGINEER AND I SUPPORT THE NEMOTRON FAMILY OF MODELS.
- 4:26 I THINK IT'S, YOU KNOW, WE'VE TALKED A LOT ABOUT WHY WE DO NEMOTRON BUT JUST TO SAY IT A FEW MORE TIMES.
loading
Chapters
- 0:00 Welcome and why this panel is the whole stack
- 1:14 Prime Intellect: keeping the training stack open
- 2:15 Arcee: why the west was losing the open model lead
- 3:22 Pretraining a 400 billion parameter model in six months
- 4:23 Nemotron: faster models are smarter models
- 6:35 Is open source actually less trustworthy
- 7:40 Trust is not safety
- 9:48 When access stopped being guaranteed
- 10:56 Releasing the data sets alongside the weights
- 11:59 The open superintelligence stack
- 13:02 Post training as the accessible layer
- 14:07 Beating frontier models on a specific use case
- 15:11 Making your costs predictable
- 16:11 When the model and the harness blend together
- 18:15 The mismanaged genius
- 19:19 A call to action for builders
- 20:24 Who owns the data you generate
- 21:27 Owning your outputs, not just your weights
- 22:35 The open MDW license
- 23:36 Where post training unlocks new use cases
- 27:46 You do not need frontier intelligence for most tasks
- 28:53 Why efficiency has to happen in the open
- 29:55 Closed models are not the enemy
- 33:05 Predictions for the next year
- 38:22 Running everything on your laptop
- 39:29 Agent operating systems and the next Siri moment
- 41:36 Closing