Videos k35LeKZEhiE
Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute
Scene timeline
48 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 122
- whisperx 122
- chunks
- 31
- from 122 cues
- keyframes
- 12
- kept of 48 captured
- frames with text
- 12
- 144 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 5.5 MB
- word timings on 122 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:31 | 0s |
stt |
done | — | 2026-08-09 14:32 | 19s |
chunk |
done | — | 2026-08-09 14:32 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 0s |
keyframe |
done | — | 2026-08-09 14:33 | 4m 13s |
ocr |
done | — | 2026-08-09 14:37 | 5s |
frame_embed |
done | — | 2026-08-10 19:42 | 2s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AIEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer1.00
- World's Fair0.99
-
- AlEngineer0.98
- World's Fair0.97
- Applied Compute1.00
- PRESENTED BY1.00
- Learning on1.00
- Microsoft1.00
- the Job: the0.98
- Future of1.00
- Post-Training1.00
- July 1, 20260.95
- World's Fair0.98
- Engineering the future of Al1.00
-
- AlEngineer0.98
- World's Fair0.97
- Post-Training Trend1.00
- Applied Compute0.98
- Controlled environment1.00
- Production environment1.00
- Curated data1.00
- Flexible data1.00
- PRESENTED BY1.00
- Easier to train (GRPO)1.00
- Harder to train (OPSD??)0.97
- Microsoft1.00
- World's Fair0.97
- Epgingfring Miatfutiurge0.71
- TRACK 9. JULY 1, 20260.92
- he2future of Al0.75
-
- AlEngineer0.96
- AppliedCompute0.99
- World'sFair1.00
- Evolution of Post-Training1.00
- 011.00
- Baby Steps - Q&A tasks0.96
- 021.00
- Grade School – Synthetic Environments0.97
- 031.00
- Internships – Bring Your Own Harness (BYOH)0.99
- 041.00
- Agentic Citizens – Autonomous, Adaptive Models1.00
- World's Fair0.94
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.99
-
- AlEngineer0.99
- Applied Compute0.98
- World's Fair0.96
- Baby Steps1.00
- Q&A Tasks0.99
- World'sFair1.00
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- Applied Compute0.98
- World'sFair1.00
- Basic Q&A Training Stack1.00
- Orchestrator1.00
- Requests1.00
- Inference engines1.00
- Proxies1.00
- Responses1.00
- Prompt1.00
- prompt + answer1.00
- Task spec:1.00
- Model1.00
- Answer1.00
- completion1.00
- Answer1.00
- endpoint1.00
- V0.88
- Weight update1.00
- Grader1.00
- Graded chats1.00
- Training Engine0.98
- - training fwd/bwd0.98
- - master weights0.99
- - optimizer state0.99
- World's Fair0.97
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- Applied Compute1.00
- World's Fair0.99
- Practice Makes Perfect1.00
- Replayable!1.00
- Outside training stack1.00
- In training stack1.00
- Requests1.00
- Responses1.00
- Model1.00
- Orchestrator1.00
- Sandbox1.00
- completion1.00
- (has task spec)1.00
- - environment state1.00
- endpoint1.00
- - tool call execution0.99
- Full task trace1.00
- Grader1.00
- World's Fair1.00
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
Transcript
122 cues· 2,778 words· 15,283 chars
- 0:12 Yeah, thank you, Jack.
- 0:14 Really grateful for the opportunity to speak here.
- 0:16 Today I'm gonna be sharing some of our frontier work on post-training and how we envision a future where agents can learn new skills on the job.
- 0:27 So over the last year or so, we've seen agents develop really strong reasoning skills
- 0:34 And they've learned to use agentic harnesses to solve longer and longer horizon tasks, which involve many turns and tool calls on complicated environment states.
- 0:46 We're seeing an increasing demand for agents that can just be deployed in a plug and play way into how enterprises use the agents.
- 0:59 So for instance, if they already have some method of
- 1:02 calling the agent to do a task, they would want to be able to train a custom model to do that task instead.
- 1:11 And that requires new ways of looking at post-training that allow you to adapt to any harness, including ones that you don't necessarily have access to the source code of.
- 1:25 So I wanted to talk about a few different levels of post-training where each one builds on top of the last.
- 1:34 One way that we kind of think of this is a framework comparing it to how humans do learning where you learn simple tasks first and you can sort of compound your understanding to more and more complicated tasks.
- 1:48 So
- 1:49 Over the last year, we've sort of, I would say, mastered or gotten a lot of reps with these simple single turn Q&A tasks and some longer horizon synthetic environment tasks.
- 2:02 But what we're increasingly seeing is we want to be able to adapt to custom harnesses and
- 2:09 be able to train directly on those instead.
- 2:12 And we kind of think of those kind of like internships where you want the model to do a specific task, but you don't necessarily know how exactly the task will play out because you don't own the harness.
- 2:26 And then finally, I would want to share some visions we have for the future of the custom model training space, where we think that there will be these kind of agentic citizens, which you can just deploy once, and they'll be able to adapt to many different types of distribution tasks and learn from their interactions.
- 2:47 So first, I just want to talk about the training setup for these simple Q&A tasks.
- 2:54 We have something that looks like this where you have an orchestrator, and the orchestrator is in charge of driving the rollouts.
- 3:02 The orchestrator holds a task spec which you can think of for now as just the simple prompt and answer, so something like a math question and a corresponding numerical answer.
- 3:13 The orchestrator will send this prompt to a model and then get an answer back.
- 3:17 Then it will send the answer to a grader and have it be graded.
- 3:20 So once all of this is done, we want to improve our model based on that interaction or maybe like a batch of interactions.
- 3:28 And the way we do that is through a training engine.
- 3:31 which takes in the graded chats and produces a wait update.
- 3:35 That wait update is then synced to some inference engines, and once those inference engines are updated, then we can start this entire process over again, where the orchestrator will have new problems to send to the model completion endpoint, and then you'll be able to get more chats, grade them, and train again.
- 3:56 The key thing to note here is that the only thing you need for improving your model is the graded chats in some format.
- 4:04 And once you have those, the training engine can compute wait updates to improve your model.
- 4:11 What's important here is that the chats are in a very specific format because we're sort of constraining everything to be inside of our training stack.
- 4:19 So in this simple setup for Q&A, you don't have anything living outside of the training stack.
- 4:24 You basically have the code of how to run the rollout and how everything is formatted.
- 4:30 So it's like in a very controlled environment.
- 4:34 However, this is kind of limited because we can only kind of do single turn tasks in this way.
- 4:40 If we want to do longer and long horizon tasks and we want to build like higher order skills into our models, we need to also increase the complexity of our environment.
- 4:51 So with synthetic environments, we have a very similar setup, but we offload a lot of the environment state outside of the training stack.
- 5:00 So you still have the same orchestrator from before, but
- 5:07 The task spec is maybe a little bit more complicated and the environment state is living outside of the training stack.
- 5:14 So the task spec might now include things like tool cost specs or like maybe an initial state for your environment like a file system.
- 5:23 And the orchestrator is now in charge of running many turns in series where maybe first it asks the model for how it wants to respond.
- 5:32 And then if the model wants to call some tools, it will then call the sandbox to actually modify the environment state or read the environment state and then return those results back to the model.
- 5:43 After all that is said and done, you get a full task trace out of this.
- 5:48 And that task trace is then sent to a grader for grading.
- 5:51 And very similar to what we had before, you'll be able to take the graded chats.
- 5:56 You'll be able to then use them to do a weight update.
- 6:01 The main thing to highlight here is that this orchestrator and sandbox setup is replayable, which is basically just saying that for any specific prompt, you can always roll back to the initial state and rerun it.
- 6:13 You can do that in parallel or you can do that in series.
- 6:16 But the reason that's important is because
- 6:19 the main sort of method that we use for reinforcement learning today is grpo and that involves comparing many rollouts for the same prompt and then comparing like relatively which one is better than the other and the training engine will then up like make an edit to the model to upweight the trajectories that were more successful and then down with the ones that were less successful
- 6:48 So some challenges that we face in this setup is that the environment is something that you want to basically use to replicate reality so that after you're done training, the improvements that you've seen actually translate to when you deploy these models into production.
- 7:05 And the main sort of problem is
loading
Chapters
- 0:00 Learning on the job
- 0:39 Custom models inside your harness
- 2:37 Deploy once and adapt
- 2:49 The RL training loop
- 4:40 Toward longer horizon tasks
- 6:48 Reward hacking in practice
- 9:06 Replicating production environments
- 9:45 Why replaying real traffic is hard
- 11:57 Non-replayability and off-policy data
- 13:41 Automated data pipelines
- 15:24 A model that learns every interaction