Videos V-EDrhIhHzQ
Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect
Scene timeline
115 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 454
- whisperx 454
- chunks
- 83
- from 454 cues
- keyframes
- 22
- kept of 115 captured
- frames with text
- 22
- 460 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 13.1 MB
- word timings on 454 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 03:04 | 1m 19s |
stt |
done | — | 2026-08-11 03:06 | 54s |
chunk |
done | — | 2026-08-11 03:07 | 0s |
text_embed |
done | — | 2026-08-11 03:07 | 1s |
keyframe |
done | — | 2026-08-11 03:07 | 4m 16s |
ocr |
done | — | 2026-08-11 03:11 | 12s |
frame_embed |
done | — | 2026-08-11 03:11 | 3s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.97
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust1.00
- bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.94
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- together.ai0.96
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- MPRImeIntellect0.93
- TECH TALK· 20260.96
- World'sFair1.00
- PRESENTED BY0.97
- Microsoft1.00
- The Prime Intellect Stack1.00
- A deep dive into post-training with verifiers + prime-rl0.99
- WILL BROWN • HEAD OF APPLIED RESEARCH0.98
- World'sFair0.97
- Engineering the future of Al1.00
-
- AlEngineer0.99
- OVERVIEW1.00
- World'sFair1.00
- open superintelligence stack1.00
- 00.72
- verifiers for environments, prime-rl for training, renderers for tokenization, and our0.99
- Lab platform for hosted training and evaluation1.00
- FRONTIER MODEL TRAINING1.00
- INTELLECT series, customer models0.99
- LAB1.00
- hosted training, evals, inference, sandboxes0.98
- ENVIRONMENTS1.00
- verifiers + Environments Hub1.00
- PRIME-RL1.00
- open RL trainer1.00
- COMPUTE1.00
- decentralized marketplace1.00
- PRIME INTELLECT1.00
- THE STACK0.97
- 020.91
- Engineering the future of Al0.99
- World's Fair0.97
-
- AlEngineer0.98
- LAB-COOKBOOK1.00
- World'sFair1.00
- what we'll cover0.99
- follow along in the lab-cookbook repo1.00
- environments = evals0.99
- PUBLIC REPO - FOLLOW ALONG0.98
- 21.00
- verifiers v1: tasksets + harnesses0.99
- PrimeIntellect-ai /0.99
- training with prime-rl1.00
- lab-cookbook1.00
- custom algorithms1.00
- guides, configs, and environments1.00
- 5 large-scale training1.00
- for building and training on Lab0.99
- github.com/PrimeIntellect-ai/lab-cookbook1.00
- 6 Lab platform0.95
- PRIME INTELLECT1.00
- THE STACK0.96
- 030.92
- Engineering the future of Al0.98
- World's Fair0.98
-
- AlEngineer0.98
- LAB-COOKBOOK1.00
- World's Fair0.98
- what we'll cover1.00
- 00.58
- follow along in the lab-cookbook repo1.00
- PRESENTED BY1.00
- environments = evals0.99
- PUBLIC REPO - FOLLOW ALONG0.98
- Microsoft1.00
- 21.00
- verifiers v1: tasksets + harnesses1.00
- PrimeIntellect-ai /1.00
- training with prime-rl1.00
- lab-cookbook1.00
- custom algorithms1.00
- guides, configs, and environments1.00
- 5 large-scale training1.00
- for building and training on Lab0.99
- github.com/PrimeIntellect-ai/lab-cookbook1.00
- 6 Lab platform0.99
- PRIME INTELLECT1.00
- THE STACK1.00
- Engineering the future of Al0.98
- World's Fair0.99
-
- AlEngineer0.98
- THE POST-TRAINING LOOP1.00
- World's Fair0.99
- from environment to trained model1.00
- build an environment, evaluate a model on it, train with RL / SFT / OPD, deploy the0.99
- result1.00
- PRESENTED BY1.00
- Microsoft1.00
- BUILD1.00
- EVALUATE1.00
- TRAIN1.00
- DEPLOY1.00
- an environment1.00
- score a model1.00
- RL· SFT· OPD0.93
- serve the1.00
- taskset + rewards0.96
- read the rollouts0.97
- post-training1.00
- trained adapter1.00
- PRIME INTELLECT0.97
- THE STACK0.94
- Engineering the future of Al0.99
- World's Fair1.00
-
- AlEngineer0.99
- THE POST-TRAINING LOOP1.00
- World's Fair0.98
- from environment to trained model1.00
- build an environment, evaluate a model on it, train with RL / SFT / OPD, deploy the0.99
- result1.00
- BUILD1.00
- EVALUATE1.00
- TRAIN1.00
- DEPLOY1.00
- an environment1.00
- score a model1.00
- RL·SFT·OPD1.00
- serve the1.00
- taskset + rewards0.97
- read the rollouts1.00
- post-training1.00
- trained adapter1.00
- PRIME INTELLECT. THE STACK0.95
- Engineering the future of Al0.99
- World's Fair0.99
Transcript
454 cues· 9,576 words· 50,661 chars
- 0:12 Hey guys, how's it going?
- 0:14 Thanks for showing up.
- 0:14 This was a little bit of a last minute assembly.
- 0:18 A few days ago, I was talking to Swix.
- 0:21 I was like, hey, can I still do a workshop?
- 0:22 And he was like, we got one slot left.
- 0:23 It's Monday at 4.30.
- 0:24 And I was like, I'll take it.
- 0:26 And then, yeah.
- 0:28 I wanted to kind of just do a bit of an update on...
- 0:33 some of the stuff we've been building at Prime Intellect.
- 0:34 So if you don't know me, hi, I'm Will Brown.
- 0:37 I lead applied research at Prime Intellect.
- 0:39 We do a lot of stuff around every part of the kind of AI research infrastructure stack.
- 0:44 Today is gonna be about post-training, which is where I spend a lot of my time thinking and building.
- 0:49 And especially wanna be talking about the post-training tools that we build that are fully open source, the verifiers and PrimRL libraries, which kind of go hand in hand, both on the environment side and the training info side.
- 1:02 and show off some things we've been cooking over the past few months that I think is kind of the way that things have evolved as the agent use cases have gotten more complex but also kind of clearer in terms of what people want out of agents and the sorts of things that are needed to like do the sort of post training that is needed to power like the real world applications people are building nowadays.
- 1:21 And so broadly at Prime Intellect we are
- 1:25 Our goal is to make doing large scale open source AI research easier and to enable companies to train their own models and deploy them and have them improve based on the scenarios that they actually see in production in terms of use cases for applications and products and internal tasks and workflows and to give people an option to not just use the open source models that are getting quite good but to take them and make them even better on their own use cases.
- 1:53 And so we use the phrase the open superintelligence stack to describe what we mean by this.
- 1:57 And I think when we said this phrase like a year ago, it felt a little more like marketing and now it feels a little bit more like oh yeah, that's kind of what it is.
- 2:06 Like the models are getting very, very good.
- 2:08 They are superhuman in many ways at lots of things.
- 2:12 And what we want to do is give people an open toolkit that they can use to do real training with them.
- 2:19 and to have the control that they need to deploy it where they need to deploy it and customize it as much as they need to to kind of get the job done.
- 2:26 And so this is the stack that we build.
- 2:27 And it all kind of sits on top of compute.
- 2:29 So we operate a global marketplace of data centers around the world.
- 2:34 A lot of these are like quite large data centers.
- 2:36 We currently operate over 10,000 GPUs.
- 2:40 many in like hundreds or thousands within a cluster.
- 2:43 We have our primary L training framework.
- 2:45 We have environments built with the verifiers library and our environments hub platform.
- 2:50 We have our platform for research workflows that we're now calling Lab, which is an assembly of many pieces, including the environments hub, hosted training evaluations, as well as inference and sandboxes.
- 3:00 And all of this is in service of empowering and unlocking frontier model training.
- 3:06 We do this ourselves.
- 3:07 We have our intellect model series with some exciting things there coming soon.
- 3:11 And we also train models with our customers where we have lots of people we work with who their goal is to do large scale model training on their own workflows.
- 3:19 And so to do all of this, we need to give people the tools they can assemble into the pipelines, the workflows, the research that allows them to actually get the results that they need at scale with everything they need to do it.
- 3:31 And so this talk is gonna be about going deep into Verifiers and Prime RL and showing off some of these new things, but all under the umbrella of what does modern post-training look like?
- 3:41 What does it mean to kind of take a model and train it to be better at your task?
- 3:45 What are all the parts?
- 3:46 What are all the kind of gotchas?
- 3:48 And how do you orchestrate this into a system that is actually easy for people to use without needing to go build a massive research team
- 3:55 and to be able to kind of have it be accessible and sorts of things that anyone who's an AI engineer at any startup or enterprise that wants to invest in post-training can actually do.
- 4:05 And so there's a cookbook repo that is, it's kind of like an alpha release right now.
- 4:09 It's still changing a bit, but it's a preview of kind of all the stuff we've been building over the past several months.
- 4:14 And so today we'll be kind of following along that framing a good bit.
- 4:19 And so I think the first thing we'll talk about is just kind of what is an environment.
- 4:23 People talk about environment in the context of RL and think of like RL environments, but environments are more than just for RL.
loading
Chapters
- 0:00 Introduction and Overview of Prime Intellect
- 4:20 Defining the Environment in Post-Training
- 9:33 Decomposing Environments: Tasks, Harnesses, and Runtimes
- 12:46 Verifiers V1: The New Modular Pattern
- 17:46 Rewards, Metrics, and Group-Level Rewards
- 20:25 Tooling, User Simulators, and MCP Integration
- 22:00 The Interception Server Pattern
- 24:13 Trace Graphs and Handling Tokenization
- 25:35 The Renderers Library for Chat Templates
- 29:20 Primaril: Asynchronous Reinforcement Learning
- 38:02 Customizing Training Algorithms and Losses
- 42:35 The Lab Platform and Hosted Training