read-only demo

Videos AQv3qRCG6Gw

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:19:26· en-US· indexed 2026-08-10 19:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:20, 1 of 1 keyframes kept
  5. Shot 4, 0:20 to 0:47, 1 of 1 keyframes kept
  6. Shot 5, 0:47 to 1:14, 1 of 1 keyframes kept
  7. Shot 6, 1:14 to 1:43, 1 of 1 keyframes kept
  8. Shot 7, 1:43 to 2:12, 0 of 1 keyframes kept
  9. Shot 8, 2:12 to 2:41, 1 of 1 keyframes kept
  10. Shot 9, 2:41 to 3:09, 1 of 1 keyframes kept
  11. Shot 10, 3:09 to 3:38, 0 of 1 keyframes kept
  12. Shot 11, 3:38 to 4:06, 0 of 1 keyframes kept
  13. Shot 12, 4:06 to 4:32, 1 of 1 keyframes kept
  14. Shot 13, 4:32 to 5:02, 1 of 1 keyframes kept
  15. Shot 14, 5:02 to 5:35, 1 of 1 keyframes kept
  16. Shot 15, 5:35 to 6:08, 1 of 1 keyframes kept
  17. Shot 16, 6:08 to 6:41, 0 of 1 keyframes kept
  18. Shot 17, 6:41 to 7:09, 0 of 1 keyframes kept
  19. Shot 18, 7:09 to 7:37, 0 of 1 keyframes kept
  20. Shot 19, 7:37 to 8:04, 1 of 1 keyframes kept
  21. Shot 20, 8:04 to 8:32, 0 of 1 keyframes kept
  22. Shot 21, 8:32 to 9:07, 1 of 1 keyframes kept
  23. Shot 22, 9:07 to 9:42, 0 of 1 keyframes kept
  24. Shot 23, 9:42 to 10:15, 0 of 1 keyframes kept
  25. Shot 24, 10:15 to 10:48, 1 of 1 keyframes kept
  26. Shot 25, 10:48 to 11:29, 1 of 1 keyframes kept
  27. Shot 26, 11:29 to 11:57, 0 of 1 keyframes kept
  28. Shot 27, 11:57 to 12:25, 0 of 1 keyframes kept
  29. Shot 28, 12:25 to 12:52, 1 of 1 keyframes kept
  30. Shot 29, 12:52 to 13:20, 0 of 1 keyframes kept
  31. Shot 30, 13:20 to 14:04, 0 of 1 keyframes kept
  32. Shot 31, 14:04 to 14:32, 1 of 1 keyframes kept
  33. Shot 32, 14:32 to 15:01, 0 of 1 keyframes kept
  34. Shot 33, 15:01 to 15:29, 1 of 1 keyframes kept
  35. Shot 34, 15:29 to 15:57, 0 of 1 keyframes kept
  36. Shot 35, 15:57 to 16:26, 1 of 1 keyframes kept
  37. Shot 36, 16:26 to 16:54, 0 of 1 keyframes kept
  38. Shot 37, 16:54 to 16:55, 1 of 1 keyframes kept
  39. Shot 38, 16:55 to 16:58, 1 of 1 keyframes kept
  40. Shot 39, 16:58 to 17:20, 0 of 1 keyframes kept
  41. Shot 40, 17:20 to 18:05, 1 of 1 keyframes kept
  42. Shot 41, 18:05 to 18:30, 1 of 1 keyframes kept
  43. Shot 42, 18:30 to 18:55, 0 of 1 keyframes kept
  44. Shot 43, 18:55 to 19:02, 0 of 1 keyframes kept
  45. Shot 44, 19:02 to 19:09, 1 of 1 keyframes kept
  46. Shot 45, 19:09 to 19:26, 0 of 1 keyframes kept

46 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
164
whisperx 164
chunks
34
from 164 cues
keyframes
26
kept of 46 captured
frames with text
26
439 lines read
chapters
11
from the source metadata
keyframe bytes
4.9 MB
word timings on 164 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:46 0s
stt done 2026-08-09 02:48 22s
chunk done 2026-08-09 02:49 0s
text_embed done 2026-08-10 19:37 1s
keyframe done 2026-08-09 02:49 2m 27s
ocr done 2026-08-09 02:51 10s
frame_embed done 2026-08-10 19:37 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:13 #3 done2 line(s)

    shot 3·sharpness 312.4

    1. AlEngineer1.00
    2. World's Fair0.98
  • 0:39 #4 done12 line(s)

    shot 4·sharpness 1710.3

    1. PRIm€Intellect0.91
    2. RESEARCH TALK - 20260.95
    3. AlEngineer0.98
    4. World'sFair1.00
    5. PRESENTED BY1.00
    6. Microsoft1.00
    7. Reinforcement Learning without0.99
    8. Verifiable Rewards1.00
    9. Scaling RL to messy, real-world agent tasks0.98
    10. WILL BROWN • HEAD OF APPLIED RESEARCH0.95
    11. World's Fair0.96
    12. Engineering the future of Al0.99
  • 1:11 #5 done13 line(s)

    shot 5·sharpness 1855.9

    1. PRIm€Intellect0.89
    2. RESEARCH TALK- 20260.94
    3. AlEngineer0.98
    4. World'sFair1.00
    5. PRESENTED BY1.00
    6. Microsoft1.00
    7. Reinforcement Learning without0.99
    8. Verifiable Rewards0.99
    9. Scaling RL to messy, real-world agent tasks0.98
    10. WILL BROWN • HEAD OF APPLIED RESEARCH0.96
    11. World'sFair1.00
    12. TRACK 9· JULY 1, 20260.95
    13. Posttraining & Midtraining1.00
  • 1:18 #6 done18 line(s)

    shot 6·sharpness 1675.5

    1. AlEngineer1.00
    2. how reinforcement learning works1.00
    3. World's Fair0.96
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. action1.00
    7. AGENT1.00
    8. ENVIRONMENT1.00
    9. model + harness0.97
    10. task + "world" + scoring0.97
    11. observation1.00
    12. reward → advantage → nudge the policy toward higher reward0.99
    13. PRIME INTELLECT1.00
    14. RESEARCH0.99
    15. 020.98
    16. World's Fair0.99
    17. TRACK 9 · JULY 1, 20260.93
    18. Posttraining & Midtraining1.00
  • 1:47 #7 skipped

    shot 7·duplicate of #6

  • 2:24 #8 done16 line(s)

    shot 8·sharpness 1548.1

    1. AlEngineer0.96
    2. how reinforcement learning works1.00
    3. World'sFair1.00
    4. action1.00
    5. AGENT1.00
    6. ENVIRONMENT1.00
    7. model + harness1.00
    8. task + "world" + scoring0.97
    9. observation1.00
    10. reward → advantage → nudge the policy toward higher reward0.98
    11. PRIME INTELLECT1.00
    12. RESEARCH0.99
    13. 020.84
    14. World'sFair1.00
    15. TRACK 9· JULY 1, 20260.94
    16. Posttraining & Midtraining0.98
  • 3:00 #9 done18 line(s)

    shot 9·sharpness 1557.3

    1. AlEngineer0.97
    2. the Prime Intellect stack1.00
    3. World'sFair1.00
    4. MODELS1.00
    5. open models, tuned to a task1.00
    6. LAB1.00
    7. hosted training, evals, inference1.00
    8. ENVIRONMENTS1.00
    9. tasksets + harnesses + verifiers0.99
    10. TRAINING1.00
    11. prime-rl, our open RL trainer1.00
    12. COMPUTE1.00
    13. GPU cluster orchestration1.00
    14. PRIME INTELLECT · RESEARCH0.97
    15. 030.98
    16. World's Fair0.99
    17. TRACK 9· JULY 1, 20260.95
    18. Posttraining & Midtraining1.00
  • 3:32 #10 skipped

    shot 10·duplicate of #8

  • 4:03 #11 skipped

    shot 11·duplicate of #9

  • 4:09 #12 done17 line(s)

    shot 12·sharpness 1426.6

    1. AlEngineer0.98
    2. what's an environment?1.00
    3. World's Fair0.96
    4. EVALS1.00
    5. SYNTHETIC DATA1.00
    6. ENVIRONMENT1.00
    7. tasks· harnesses0.97
    8. rewards1.00
    9. SFT1.00
    10. RL1.00
    11. OPD1.00
    12. GEPA1.00
    13. PRIME INTELLECT· RESEARCH0.98
    14. 040.99
    15. World's Fair0.99
    16. TRACK 9 · JULY 1, 20260.93
    17. Posttraining & Midtraining1.00
  • 4:44 #13 done17 line(s)

    shot 13·sharpness 1318.7

    1. AlEngineer0.98
    2. verifiable rewards1.00
    3. World's Fair0.97
    4. MATH1.00
    5. CODE1.00
    6. TOOL USE0.98
    7. a golden answer0.99
    8. test cases1.00
    9. final-state check1.00
    10. answer == 420.97
    11. all tests pass1.00
    12. check_state(db)1.00
    13. PRIME INTELLECT - RESEARCH0.97
    14. 050.94
    15. World's Fair0.97
    16. TRACK 9· JULY 1, 20260.96
    17. Posttraining & Midtraining1.00
  • 5:22 #14 done21 line(s)

    shot 14·sharpness 1648.2

    1. AlEngineer0.99
    2. most real tasks aren't verifiable0.98
    3. World's Fair0.98
    4. VERIFIABLE1.00
    5. REAL AGENT TASKS0.98
    6. PRESENTED BY1.00
    7. math - golden answer0.99
    8. "is this a good report?"0.98
    9. Microsoft1.00
    10. code - test cases1.00
    11. "did it book the right flight?"1.00
    12. tool use - action check0.99
    13. "is this refund handled well?"0.99
    14. clean signal ✓0.92
    15. unclear signal x1.00
    16. PRIME INTELLECT1.00
    17. RESEARCH1.00
    18. 060.96
    19. World's Fair0.99
    20. TRACK 9· JULY 1, 20260.96
    21. Posttraining & Midtraining1.00
  • 5:52 #15 done19 line(s)

    shot 15·sharpness 1946.6

    1. AlEngineer0.98
    2. making evals is hard0.99
    3. World's Fair0.97
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. TASKS DON'T SCALE BY HAND0.99
    7. OPEN-ENDED OUTPUTS1.00
    8. writing each one is expensive0.99
    9. no clean check for "good"0.98
    10. COVERAGE IS UNBOUNDED1.00
    11. REWARD HACKING1.00
    12. you can't enumerate real situations1.00
    13. agents exploit weak verifiers1.00
    14. PRIME INTELLECT1.00
    15. RESEARCH1.00
    16. 070.99
    17. World's Fair0.99
    18. TRACK 9· JULY 1, 20260.96
    19. Posttraining & Midtraining1.00
  • 6:12 #16 skipped

    shot 16·duplicate of #15

  • 7:03 #17 skipped

    shot 17·duplicate of #8

  • 7:23 #18 skipped

    shot 18·duplicate of #8

  • 7:40 #19 done17 line(s)

    shot 19·sharpness 1557.4

    1. AlEngineer0.98
    2. the goal: continual learning1.00
    3. World'sFair1.00
    4. PRODUCTION1.00
    5. NEW TASKS + WORLDS1.00
    6. real agents, real traces0.98
    7. mined from those traces1.00
    8. BETTER AGENT1.00
    9. ONLINE RL1.00
    10. redeployed to production1.00
    11. train on what actually fails1.00
    12. PRIME INTELLECT1.00
    13. RESEARCH1.00
    14. 080.98
    15. World's Fair0.99
    16. TRACK 9 · JULY 1, 20260.93
    17. Posttraining & Midtraining0.98
  • 8:13 #20 skipped

    shot 20·duplicate of #19

  • 8:49 #21 done18 line(s)

    shot 21·sharpness 1532.2

    1. AlEngineer0.98
    2. manufacturing signal1.00
    3. World'sFair0.96
    4. GROUNDING1.00
    5. JUDGES1.00
    6. SCALING SEARCH1.00
    7. from source material1.00
    8. for open-ended work0.98
    9. + test-time compute1.00
    10. docs · code · production traces0.98
    11. score vs references + rubrics1.00
    12. spend compute to sharpen both1.00
    13. PRIME INTELLECT1.00
    14. RESEARCH1.00
    15. 090.86
    16. World's Fair0.98
    17. TRACK 9· JULY 1, 20260.94
    18. Posttraining & Midtraining1.00
  • 9:25 #22 skipped

    shot 22·duplicate of #21

  • 10:02 #23 skipped

    shot 23·duplicate of #8

Transcript

164 cues· 4,023 words· 21,807 chars

  1. 0:13 Thanks all for coming to AI Engineer and checking out the post-training session.
  2. 0:17 Hopefully lots of fun stuff today and throughout the conference.
  3. 0:21 I'm Will Brown, I lead applied research at Prime Intellect, and today I want to talk about reinforcement learning without verifiable rewards.
  4. 0:28 And so many people may have been learning about RLVR over the past
  5. 0:34 year or so, year and a half as this stuff has really taken off and become the main way that we think about scaling reinforcement learning.
  6. 0:41 But often we don't actually have verifiable rewards.
  7. 0:44 And so messy real world tasks often we're kind of figuring out as we go, we're having our agents run around and we kind of in hindsight maybe can look at what they did and say like okay this was good, this was bad.
  8. 0:56 but sometimes sitting down and just like specifying, okay, this is the rule, this is the goal, is not always so straightforward.
  9. 1:02 And so this is gonna be synthesizing a lot of work we've been doing as well as from the broader research literature and some of the things we're building to kind of support extending RL into more messy real world tasks.
  10. 1:15 And so recap quickly of how reinforcement learning works.
  11. 1:18 I would imagine if you're in the post-training session here, you've probably heard a little bit about RL, but for those of you at home and for those who are kind of still just kind of looking for the crash course, generally we have an agent, which we're gonna call a model plus a harness, which we place into an environment.
  12. 1:32 And so an environment we're gonna call a task plus a world, a world you could think of as maybe it's a Docker image, maybe it's a code base,
  13. 1:40 maybe it is a collection of task-specific tools, maybe it's some skills, maybe it is a bunch of applications or browser tabs or things like this, as well as a scoring rule, verifiers or rewards, whatever you want to call them.
  14. 1:53 And the agent and the environment are going to interact in a loop, and at the end, we'll have some reward of how well the agent did in the environment for this task.
  15. 2:02 And then reinforcement learning is all about creating an advantage.
  16. 2:05 And the advantage is really about taking the reward, minusing some baseline, maybe doing some scaling, and then now you have a set of rollouts from the agent in the environment that you can use to then update the policy.
  17. 2:18 The policy here is just the model weights themselves.
  18. 2:21 And the goal here is to take a gradient which nudges the model towards getting higher reward.
  19. 2:25 And so all of the
  20. 2:26 RL stuff people talk about, whether it's GRPO or Reinforce or CISPO or any of the other new algorithms people come up with, they're all kind of in this policy gradient framework which is just about saying, okay, how do I make the model do things that have higher reward?
  21. 2:42 And so at Prime and Elect, we build a lot of tooling to power all of this at every layer.
  22. 2:46 We kind of go both from the, we start at the compute layer and do lots of large-scale GPU orchestration.
  23. 2:52 We build the PrimRL training framework, which powers all of our large-scale reinforcement learning and other algorithms running.
  24. 2:58 We build environments.
  25. 2:59 We have task sets and harnesses and verifiers as tools that you can mix and match to assemble complex worlds for agents to learn from real-world feedback.
  26. 3:11 We have a training platform called Lab which is anchored around environments where we do both hosted training and evaluations as well as inference and this is to allow people to monitor their experiments and manage their training runs and iterate on their evals and deploy these models and ultimately the models are starting generally from some open source based model and you're optimizing it for your task and the goal that we're really trying to enable is for more people to be able to become their own research lab, become
  27. 3:38 take ownership over the intelligence of their own model weights and optimize for the tasks that they care about with themselves as the experts steering the model, which means we need to make it way easier for people to do this.
  28. 3:50 Currently, for a lot of people, it's still really hard.
  29. 3:53 I think you can go, like we're all here learning more about how it works because it's hard.
  30. 3:57 We don't know how it all works, and we're figuring it out as we go in many cases, but we've spent a lot of effort and a lot of time building stuff that hopefully makes this a bit easier for people.
  31. 4:07 And so what's an environment?
  32. 4:08 An environment is tasks, harness, and rewards, but it's not just for RL.
  33. 4:13 So I think a lot of people think RL when they think environment, but environments and evals are really the same thing.
  34. 4:18 You can use these same objects for generating synthetic data, which then you could use for SFT.
  35. 4:23 You can do RL or you could do algorithms like on policy distillation.
  36. 4:26 You could do prompt optimization like JEPA.
  37. 4:27 You can use it as a kind of
  38. 4:29 scientific test bed to iterate on your agents and your harnesses.
  39. 4:33 And verifiable rewards are the easy case where we just kind of can check exactly was something done correctly or not.
  40. 4:39 And so for math, often if you have a numerical answer you can just parse this out of like a box in the answer from the model and check.
  41. 4:46 For code, maybe you want to use test cases or a linter or something like this.
  42. 4:49 For tool use, often you have some database state which you kind of know what you're expecting at the end and you can just kind of like check this deterministically.
  43. 4:56 And so these are kind of the easy cases where the reward design problem is not so difficult.
  44. 5:03 But most real-world tasks are not this verifiable.
  45. 5:06 For a lot of real agent tasks, we're having agents do things like write reports that maybe are analyzing a bunch of documents or research, or maybe ask you to do things like book flights or buy things, but there isn't always like a clean best answer here.
  46. 5:20 And there's also notions that are fuzzier of like interacting with users, like handling a refund.
  47. 5:24 Like what does it mean to handle this well?
  48. 5:27 And so here the signal is less clear.
  49. 5:29 And there's a lot of different tricks we might want to explore and techniques we want to develop to ensure that this can be done reliably and scalably.
  50. 5:36 And making evals is hard because oftentimes the benchmarks out there that we might look at in the kind of new model releases, it's like a set of a few hundred tasks that a bunch of researchers spent months kind of handcrafting and talking to experts.

Chapters

  1. 0:00 RL without verifiable rewards
  2. 1:17 How RL works, simply
  3. 2:43 The tooling that powers it
  4. 4:13 Where verifiable rewards run out
  5. 6:24 Being careful about reward design
  6. 8:04 Making RL a science
  7. 9:20 Judges and grounded question answer pairs
  8. 10:46 The reverse direction trick
  9. 14:19 Calibrating difficulty
  10. 15:08 Hunting for reward hacks
  11. 18:27 Environments as the anchor

Open at this second