read-only demo

Videos k35LeKZEhiE

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:18:20· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:22, 1 of 1 keyframes kept
  5. Shot 4, 0:22 to 0:26, 1 of 1 keyframes kept
  6. Shot 5, 0:26 to 0:56, 1 of 1 keyframes kept
  7. Shot 6, 0:56 to 1:25, 0 of 1 keyframes kept
  8. Shot 7, 1:25 to 1:52, 1 of 1 keyframes kept
  9. Shot 8, 1:52 to 2:20, 0 of 1 keyframes kept
  10. Shot 9, 2:20 to 2:47, 0 of 1 keyframes kept
  11. Shot 10, 2:47 to 2:54, 1 of 1 keyframes kept
  12. Shot 11, 2:54 to 3:29, 0 of 1 keyframes kept
  13. Shot 12, 3:29 to 4:16, 1 of 1 keyframes kept
  14. Shot 13, 4:16 to 4:35, 0 of 1 keyframes kept
  15. Shot 14, 4:35 to 4:53, 0 of 1 keyframes kept
  16. Shot 15, 4:53 to 5:21, 0 of 1 keyframes kept
  17. Shot 16, 5:21 to 5:50, 0 of 1 keyframes kept
  18. Shot 17, 5:50 to 6:18, 0 of 1 keyframes kept
  19. Shot 18, 6:18 to 6:47, 1 of 1 keyframes kept
  20. Shot 19, 6:47 to 7:14, 0 of 1 keyframes kept
  21. Shot 20, 7:14 to 7:40, 0 of 1 keyframes kept
  22. Shot 21, 7:40 to 8:07, 0 of 1 keyframes kept
  23. Shot 22, 8:07 to 8:33, 0 of 1 keyframes kept
  24. Shot 23, 8:33 to 9:00, 0 of 1 keyframes kept
  25. Shot 24, 9:00 to 9:27, 0 of 1 keyframes kept
  26. Shot 25, 9:27 to 9:47, 0 of 1 keyframes kept
  27. Shot 26, 9:47 to 10:16, 0 of 1 keyframes kept
  28. Shot 27, 10:16 to 10:45, 0 of 1 keyframes kept
  29. Shot 28, 10:45 to 10:46, 0 of 1 keyframes kept
  30. Shot 29, 10:46 to 11:18, 0 of 1 keyframes kept
  31. Shot 30, 11:18 to 11:49, 0 of 1 keyframes kept
  32. Shot 31, 11:49 to 12:15, 0 of 1 keyframes kept
  33. Shot 32, 12:15 to 12:41, 0 of 1 keyframes kept
  34. Shot 33, 12:41 to 13:06, 0 of 1 keyframes kept
  35. Shot 34, 13:06 to 13:32, 0 of 1 keyframes kept
  36. Shot 35, 13:32 to 13:57, 0 of 1 keyframes kept
  37. Shot 36, 13:57 to 14:23, 0 of 1 keyframes kept
  38. Shot 37, 14:23 to 14:49, 0 of 1 keyframes kept
  39. Shot 38, 14:49 to 15:14, 0 of 1 keyframes kept
  40. Shot 39, 15:14 to 15:28, 0 of 1 keyframes kept
  41. Shot 40, 15:28 to 15:57, 0 of 1 keyframes kept
  42. Shot 41, 15:57 to 16:25, 0 of 1 keyframes kept
  43. Shot 42, 16:25 to 16:56, 0 of 1 keyframes kept
  44. Shot 43, 16:56 to 17:27, 0 of 1 keyframes kept
  45. Shot 44, 17:27 to 17:51, 1 of 1 keyframes kept
  46. Shot 45, 17:51 to 17:57, 0 of 1 keyframes kept
  47. Shot 46, 17:57 to 18:03, 1 of 1 keyframes kept
  48. Shot 47, 18:03 to 18:19, 0 of 1 keyframes kept

48 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
122
whisperx 122
chunks
31
from 122 cues
keyframes
12
kept of 48 captured
frames with text
12
144 lines read
chapters
11
from the source metadata
keyframe bytes
5.5 MB
word timings on 122 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:31 0s
stt done 2026-08-09 14:32 19s
chunk done 2026-08-09 14:32 0s
text_embed done 2026-08-10 19:42 0s
keyframe done 2026-08-09 14:33 4m 13s
ocr done 2026-08-09 14:37 5s
frame_embed done 2026-08-10 19:42 2s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 456.5

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.3

    1. AIEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2736.3

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:13 #3 done2 line(s)

    shot 3·sharpness 261.6

    1. AlEngineer1.00
    2. World's Fair0.99
  • 0:26 #4 done12 line(s)

    shot 4·sharpness 1932.9

    1. AlEngineer0.98
    2. World's Fair0.97
    3. Applied Compute1.00
    4. PRESENTED BY1.00
    5. Learning on1.00
    6. Microsoft1.00
    7. the Job: the0.98
    8. Future of1.00
    9. Post-Training1.00
    10. July 1, 20260.95
    11. World's Fair0.98
    12. Engineering the future of Al1.00
  • 0:52 #5 done16 line(s)

    shot 5·sharpness 1804.4

    1. AlEngineer0.98
    2. World's Fair0.97
    3. Post-Training Trend1.00
    4. Applied Compute0.98
    5. Controlled environment1.00
    6. Production environment1.00
    7. Curated data1.00
    8. Flexible data1.00
    9. PRESENTED BY1.00
    10. Easier to train (GRPO)1.00
    11. Harder to train (OPSD??)0.97
    12. Microsoft1.00
    13. World's Fair0.97
    14. Epgingfring Miatfutiurge0.71
    15. TRACK 9. JULY 1, 20260.92
    16. he2future of Al0.75
  • 1:13 #6 skipped

    shot 6·duplicate of #5

  • 1:47 #7 done15 line(s)

    shot 7·sharpness 2158.0

    1. AlEngineer0.96
    2. AppliedCompute0.99
    3. World'sFair1.00
    4. Evolution of Post-Training1.00
    5. 011.00
    6. Baby Steps - Q&A tasks0.96
    7. 021.00
    8. Grade School – Synthetic Environments0.97
    9. 031.00
    10. Internships – Bring Your Own Harness (BYOH)0.99
    11. 041.00
    12. Agentic Citizens – Autonomous, Adaptive Models1.00
    13. World's Fair0.94
    14. TRACK 9· JULY 1, 20260.95
    15. Posttraining & Midtraining0.99
  • 2:16 #8 skipped

    shot 8·duplicate of #7

  • 2:25 #9 skipped

    shot 9·duplicate of #7

  • 2:53 #10 done8 line(s)

    shot 10·sharpness 1467.9

    1. AlEngineer0.99
    2. Applied Compute0.98
    3. World's Fair0.96
    4. Baby Steps1.00
    5. Q&A Tasks0.99
    6. World'sFair1.00
    7. TRACK 9· JULY 1, 20260.96
    8. Posttraining & Midtraining1.00
  • 2:58 #11 skipped

    shot 11·duplicate of #7

  • 3:34 #12 done28 line(s)

    shot 12·sharpness 2246.3

    1. AlEngineer0.99
    2. Applied Compute0.98
    3. World'sFair1.00
    4. Basic Q&A Training Stack1.00
    5. Orchestrator1.00
    6. Requests1.00
    7. Inference engines1.00
    8. Proxies1.00
    9. Responses1.00
    10. Prompt1.00
    11. prompt + answer1.00
    12. Task spec:1.00
    13. Model1.00
    14. Answer1.00
    15. completion1.00
    16. Answer1.00
    17. endpoint1.00
    18. V0.88
    19. Weight update1.00
    20. Grader1.00
    21. Graded chats1.00
    22. Training Engine0.98
    23. - training fwd/bwd0.98
    24. - master weights0.99
    25. - optimizer state0.99
    26. World's Fair0.97
    27. TRACK 9· JULY 1, 20260.96
    28. Posttraining & Midtraining1.00
  • 4:20 #13 skipped

    shot 13·duplicate of #12

  • 4:42 #14 skipped

    shot 14·duplicate of #10

  • 5:18 #15 skipped

    shot 15·duplicate of #10

  • 5:36 #16 skipped

    shot 16·duplicate of #5

  • 6:12 #17 skipped

    shot 17·duplicate of #5

  • 6:30 #18 done22 line(s)

    shot 18·sharpness 2223.3

    1. AlEngineer0.99
    2. Applied Compute1.00
    3. World's Fair0.99
    4. Practice Makes Perfect1.00
    5. Replayable!1.00
    6. Outside training stack1.00
    7. In training stack1.00
    8. Requests1.00
    9. Responses1.00
    10. Model1.00
    11. Orchestrator1.00
    12. Sandbox1.00
    13. completion1.00
    14. (has task spec)1.00
    15. - environment state1.00
    16. endpoint1.00
    17. - tool call execution0.99
    18. Full task trace1.00
    19. Grader1.00
    20. World's Fair1.00
    21. TRACK 9· JULY 1, 20260.96
    22. Posttraining & Midtraining1.00
  • 7:03 #19 skipped

    shot 19·duplicate of #10

  • 7:37 #20 skipped

    shot 20·duplicate of #10

  • 7:56 #21 skipped

    shot 21·duplicate of #10

  • 8:15 #22 skipped

    shot 22·duplicate of #7

  • 8:52 #23 skipped

    shot 23·duplicate of #10

Transcript

122 cues· 2,778 words· 15,283 chars

  1. 0:12 Yeah, thank you, Jack.
  2. 0:14 Really grateful for the opportunity to speak here.
  3. 0:16 Today I'm gonna be sharing some of our frontier work on post-training and how we envision a future where agents can learn new skills on the job.
  4. 0:27 So over the last year or so, we've seen agents develop really strong reasoning skills
  5. 0:34 And they've learned to use agentic harnesses to solve longer and longer horizon tasks, which involve many turns and tool calls on complicated environment states.
  6. 0:46 We're seeing an increasing demand for agents that can just be deployed in a plug and play way into how enterprises use the agents.
  7. 0:59 So for instance, if they already have some method of
  8. 1:02 calling the agent to do a task, they would want to be able to train a custom model to do that task instead.
  9. 1:11 And that requires new ways of looking at post-training that allow you to adapt to any harness, including ones that you don't necessarily have access to the source code of.
  10. 1:25 So I wanted to talk about a few different levels of post-training where each one builds on top of the last.
  11. 1:34 One way that we kind of think of this is a framework comparing it to how humans do learning where you learn simple tasks first and you can sort of compound your understanding to more and more complicated tasks.
  12. 1:48 So
  13. 1:49 Over the last year, we've sort of, I would say, mastered or gotten a lot of reps with these simple single turn Q&A tasks and some longer horizon synthetic environment tasks.
  14. 2:02 But what we're increasingly seeing is we want to be able to adapt to custom harnesses and
  15. 2:09 be able to train directly on those instead.
  16. 2:12 And we kind of think of those kind of like internships where you want the model to do a specific task, but you don't necessarily know how exactly the task will play out because you don't own the harness.
  17. 2:26 And then finally, I would want to share some visions we have for the future of the custom model training space, where we think that there will be these kind of agentic citizens, which you can just deploy once, and they'll be able to adapt to many different types of distribution tasks and learn from their interactions.
  18. 2:47 So first, I just want to talk about the training setup for these simple Q&A tasks.
  19. 2:54 We have something that looks like this where you have an orchestrator, and the orchestrator is in charge of driving the rollouts.
  20. 3:02 The orchestrator holds a task spec which you can think of for now as just the simple prompt and answer, so something like a math question and a corresponding numerical answer.
  21. 3:13 The orchestrator will send this prompt to a model and then get an answer back.
  22. 3:17 Then it will send the answer to a grader and have it be graded.
  23. 3:20 So once all of this is done, we want to improve our model based on that interaction or maybe like a batch of interactions.
  24. 3:28 And the way we do that is through a training engine.
  25. 3:31 which takes in the graded chats and produces a wait update.
  26. 3:35 That wait update is then synced to some inference engines, and once those inference engines are updated, then we can start this entire process over again, where the orchestrator will have new problems to send to the model completion endpoint, and then you'll be able to get more chats, grade them, and train again.
  27. 3:56 The key thing to note here is that the only thing you need for improving your model is the graded chats in some format.
  28. 4:04 And once you have those, the training engine can compute wait updates to improve your model.
  29. 4:11 What's important here is that the chats are in a very specific format because we're sort of constraining everything to be inside of our training stack.
  30. 4:19 So in this simple setup for Q&A, you don't have anything living outside of the training stack.
  31. 4:24 You basically have the code of how to run the rollout and how everything is formatted.
  32. 4:30 So it's like in a very controlled environment.
  33. 4:34 However, this is kind of limited because we can only kind of do single turn tasks in this way.
  34. 4:40 If we want to do longer and long horizon tasks and we want to build like higher order skills into our models, we need to also increase the complexity of our environment.
  35. 4:51 So with synthetic environments, we have a very similar setup, but we offload a lot of the environment state outside of the training stack.
  36. 5:00 So you still have the same orchestrator from before, but
  37. 5:07 The task spec is maybe a little bit more complicated and the environment state is living outside of the training stack.
  38. 5:14 So the task spec might now include things like tool cost specs or like maybe an initial state for your environment like a file system.
  39. 5:23 And the orchestrator is now in charge of running many turns in series where maybe first it asks the model for how it wants to respond.
  40. 5:32 And then if the model wants to call some tools, it will then call the sandbox to actually modify the environment state or read the environment state and then return those results back to the model.
  41. 5:43 After all that is said and done, you get a full task trace out of this.
  42. 5:48 And that task trace is then sent to a grader for grading.
  43. 5:51 And very similar to what we had before, you'll be able to take the graded chats.
  44. 5:56 You'll be able to then use them to do a weight update.
  45. 6:01 The main thing to highlight here is that this orchestrator and sandbox setup is replayable, which is basically just saying that for any specific prompt, you can always roll back to the initial state and rerun it.
  46. 6:13 You can do that in parallel or you can do that in series.
  47. 6:16 But the reason that's important is because
  48. 6:19 the main sort of method that we use for reinforcement learning today is grpo and that involves comparing many rollouts for the same prompt and then comparing like relatively which one is better than the other and the training engine will then up like make an edit to the model to upweight the trajectories that were more successful and then down with the ones that were less successful
  49. 6:48 So some challenges that we face in this setup is that the environment is something that you want to basically use to replicate reality so that after you're done training, the improvements that you've seen actually translate to when you deploy these models into production.
  50. 7:05 And the main sort of problem is

Chapters

  1. 0:00 Learning on the job
  2. 0:39 Custom models inside your harness
  3. 2:37 Deploy once and adapt
  4. 2:49 The RL training loop
  5. 4:40 Toward longer horizon tasks
  6. 6:48 Reward hacking in practice
  7. 9:06 Replicating production environments
  8. 9:45 Why replaying real traffic is hard
  9. 11:57 Non-replayability and off-policy data
  10. 13:41 Automated data pipelines
  11. 15:24 A model that learns every interaction

Open at this second