read-only demo

Videos ewtOo0scUh0

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:19:11· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:32, 1 of 1 keyframes kept
  5. Shot 4, 0:32 to 1:02, 1 of 1 keyframes kept
  6. Shot 5, 1:02 to 1:32, 1 of 1 keyframes kept
  7. Shot 6, 1:32 to 2:02, 1 of 1 keyframes kept
  8. Shot 7, 2:02 to 2:32, 0 of 1 keyframes kept
  9. Shot 8, 2:32 to 3:09, 1 of 1 keyframes kept
  10. Shot 9, 3:09 to 3:45, 0 of 1 keyframes kept
  11. Shot 10, 3:45 to 3:47, 0 of 1 keyframes kept
  12. Shot 11, 3:47 to 4:13, 0 of 1 keyframes kept
  13. Shot 12, 4:13 to 4:39, 1 of 1 keyframes kept
  14. Shot 13, 4:39 to 5:05, 0 of 1 keyframes kept
  15. Shot 14, 5:05 to 5:31, 0 of 1 keyframes kept
  16. Shot 15, 5:31 to 5:57, 1 of 1 keyframes kept
  17. Shot 16, 5:57 to 6:23, 0 of 1 keyframes kept
  18. Shot 17, 6:23 to 6:50, 0 of 1 keyframes kept
  19. Shot 18, 6:50 to 6:54, 1 of 1 keyframes kept
  20. Shot 19, 6:54 to 7:20, 1 of 1 keyframes kept
  21. Shot 20, 7:20 to 7:46, 0 of 1 keyframes kept
  22. Shot 21, 7:46 to 8:12, 0 of 1 keyframes kept
  23. Shot 22, 8:12 to 8:39, 0 of 1 keyframes kept
  24. Shot 23, 8:39 to 8:44, 0 of 1 keyframes kept
  25. Shot 24, 8:44 to 8:45, 0 of 1 keyframes kept
  26. Shot 25, 8:45 to 9:05, 0 of 1 keyframes kept
  27. Shot 26, 9:05 to 9:13, 0 of 1 keyframes kept
  28. Shot 27, 9:13 to 9:40, 0 of 1 keyframes kept
  29. Shot 28, 9:40 to 10:08, 0 of 1 keyframes kept
  30. Shot 29, 10:08 to 10:35, 0 of 1 keyframes kept
  31. Shot 30, 10:35 to 11:01, 0 of 1 keyframes kept
  32. Shot 31, 11:01 to 11:27, 0 of 1 keyframes kept
  33. Shot 32, 11:27 to 11:53, 0 of 1 keyframes kept
  34. Shot 33, 11:53 to 12:01, 1 of 1 keyframes kept
  35. Shot 34, 12:01 to 12:27, 1 of 1 keyframes kept
  36. Shot 35, 12:27 to 12:43, 1 of 1 keyframes kept
  37. Shot 36, 12:43 to 13:14, 0 of 1 keyframes kept
  38. Shot 37, 13:14 to 13:45, 0 of 1 keyframes kept
  39. Shot 38, 13:45 to 13:49, 0 of 1 keyframes kept
  40. Shot 39, 13:49 to 13:51, 0 of 1 keyframes kept
  41. Shot 40, 13:51 to 14:03, 0 of 1 keyframes kept
  42. Shot 41, 14:03 to 14:33, 0 of 1 keyframes kept
  43. Shot 42, 14:33 to 15:03, 0 of 1 keyframes kept
  44. Shot 43, 15:03 to 15:33, 1 of 1 keyframes kept
  45. Shot 44, 15:33 to 16:03, 0 of 1 keyframes kept
  46. Shot 45, 16:03 to 16:07, 1 of 1 keyframes kept
  47. Shot 46, 16:07 to 16:52, 1 of 1 keyframes kept
  48. Shot 47, 16:52 to 17:20, 0 of 1 keyframes kept
  49. Shot 48, 17:20 to 17:48, 0 of 1 keyframes kept
  50. Shot 49, 17:48 to 18:17, 0 of 1 keyframes kept
  51. Shot 50, 18:17 to 18:45, 0 of 1 keyframes kept
  52. Shot 51, 18:45 to 18:54, 1 of 1 keyframes kept
  53. Shot 52, 18:54 to 19:11, 0 of 1 keyframes kept

53 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
175
whisperx 175
chunks
33
from 175 cues
keyframes
19
kept of 53 captured
frames with text
19
421 lines read
chapters
10
from the source metadata
keyframe bytes
6.1 MB
word timings on 175 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:27 0s
stt done 2026-08-09 14:21 19s
chunk done 2026-08-09 14:22 0s
text_embed done 2026-08-10 19:42 0s
keyframe done 2026-08-09 14:22 2m 31s
ocr done 2026-08-09 14:24 8s
frame_embed done 2026-08-10 19:42 3s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 457.1

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 663.6

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2741.5

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.93
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:30 #3 done11 line(s)

    shot 3·sharpness 1389.8

    1. AlEngineer0.99
    2. World's Fair0.96
    3. BESPOKE LABS1.00
    4. Data & Environment1.00
    5. Curation for1.00
    6. MAHESH1.00
    7. Post-training LLMs1.00
    8. SATHIAMOORTHY1.00
    9. CEO, BESPOKE LABS0.97
    10. Engineering the future of Al1.00
    11. World's Fair0.99
  • 0:39 #4 done18 line(s)

    shot 4·sharpness 3009.1

    1. AlEngineer0.97
    2. Who & What0.94
    3. World'sFair1.00
    4. Who we are0.98
    5. Bespoke is an applied data research lab with a mission to help enterprises0.99
    6. and frontier labs access high quality data and RL Environments for their1.00
    7. PRESENTED BY0.99
    8. post-training needs.1.00
    9. Microsoft1.00
    10. What we do1.00
    11. Last year we put out tooling for data curation (Curator)0.99
    12. We started Bespoke-Stratos which became OpenThoughts0.99
    13. Contributors to Terminal Bench benchmark0.99
    14. These days we research, build, and ship RL Envs1.00
    15. We enable post-training for enterprises (and labs)1.00
    16. B0.73
    17. Engineering the future of Al0.99
    18. World'sFair1.00
  • 1:14 #5 done19 line(s)

    shot 5·sharpness 2922.9

    1. AlEngineer0.97
    2. Who & What0.93
    3. World'sFair1.00
    4. Who we are0.98
    5. Bespoke is an applied data research lab with a mission to help enterprises0.99
    6. and frontier labs access high quality data and RL Environments for their0.99
    7. PRESENTED BY0.97
    8. post-training needs.1.00
    9. Microsoft1.00
    10. What we do1.00
    11. Last year we put out tooling for data curation (Curator)1.00
    12. We started Bespoke-Stratos which became OpenThoughts1.00
    13. Contributors to Terminal Bench benchmark1.00
    14. These days we research, build, and ship RL Envs1.00
    15. We enable post-training for enterprises (and labs)1.00
    16. B0.61
    17. TRACK 9· JUNE 30, 20260.93
    18. World'sFair1.00
    19. Data Quality0.99
  • 1:50 #6 done17 line(s)

    shot 6·sharpness 2788.9

    1. AlEngineer0.96
    2. Who & What0.92
    3. World'sFair1.00
    4. Who we are0.99
    5. Bespoke is an applied data research lab with a mission to help enterprises0.99
    6. and frontier labs access high quality data and RL Environments for their1.00
    7. post-training needs.1.00
    8. What we do1.00
    9. Last year we put out tooling for data curation (Curator)0.99
    10. We started Bespoke-Stratos which became OpenThoughts1.00
    11. Contributors to Terminal Bench benchmark1.00
    12. These days we research, build, and ship RL Envs0.99
    13. We enable post-training for enterprises (and labs)0.99
    14. B0.64
    15. TRACK 9· JUNE 30,20260.96
    16. World'sFair1.00
    17. Data Quality0.98
  • 2:23 #7 skipped

    shot 7·duplicate of #5

  • 2:57 #8 done32 line(s)

    shot 8·sharpness 3301.9

    1. AlEngineer0.99
    2. World'sFair0.98
    3. MEASURING MASSIVE MULTITASK1.00
    4. LANGUAGE UNDERSTANDING1.00
    5. Dan Hendrycks1.00
    6. Collin Burns1.00
    7. Steven Basart1.00
    8. Andy Zou1.00
    9. What do they0.99
    10. UC Berkeley0.97
    11. Columbia University1.00
    12. UChicago1.00
    13. UC Berkeley1.00
    14. know1.00
    15. Mantas Mazeika1.00
    16. Dawn Song1.00
    17. Jacob Steinhardt1.00
    18. UIUC1.00
    19. UC Berkeley0.98
    20. UC Berkeley0.96
    21. SWE-BENCH: CAN LANGUAGE MODELS RESOLVE0.98
    22. REAL-WORLD GITHUB ISSUES?0.99
    23. What can they1.00
    24. do?1.00
    25. Carlos E. Jimenez* 1,2 John Yang* 1,2 Alexander Wettig1,20.98
    26. Shunyu Yao1,2 Kexin Pei3 Ofir Press1,2 Karthik Narasimhan1,20.97
    27. 1Princeton University0.99
    28. ²Princeton Language and Intelligence0.99
    29. 3University of Chicago0.99
    30. TRACK 9· JUNE 30,20260.95
    31. World's Fair0.99
    32. Data Quality0.98
  • 3:33 #9 skipped

    shot 9·duplicate of #8

  • 3:46 #10 skipped

    shot 10·duplicate of #6

  • 4:07 #11 skipped

    shot 11·duplicate of #6

  • 4:18 #12 done9 line(s)

    shot 12·sharpness 1286.2

    1. AlEngineer0.99
    2. World'sFair1.00
    3. What's blocking autonomy for agents?1.00
    4. Reliability1.00
    5. How to improve reliability?1.00
    6. Post-training1.00
    7. TRACK 9• JUNE 30, 20260.97
    8. World'sFair1.00
    9. Data Quality1.00
  • 5:02 #13 skipped

    shot 13·duplicate of #12

  • 5:18 #14 skipped

    shot 14·duplicate of #6

  • 5:51 #15 done13 line(s)

    shot 15·sharpness 1675.8

    1. AlEngineer0.98
    2. World'sFair1.00
    3. What's blocking autonomy for agents?1.00
    4. Reliability1.00
    5. PRESENTED BY1.00
    6. How to improve reliability?1.00
    7. Microsoft1.00
    8. Post-training1.00
    9. What's the bottleneck for post-training?1.00
    10. Data (and RL Envs)0.99
    11. TRACK 9· JUNE 30, 20260.96
    12. World'sFair1.00
    13. Data Quality1.00
  • 6:17 #16 skipped

    shot 16·duplicate of #15

  • 6:44 #17 skipped

    shot 17·duplicate of #6

  • 6:50 #18 done6 line(s)

    shot 18·sharpness 733.2

    1. AlEngineer0.99
    2. World's Fair0.94
    3. OpenThoughts1.00
    4. TRACK 9• JUNE 30, 20260.97
    5. World'sFair1.00
    6. Data Quality1.00
  • 7:04 #19 done60 line(s)

    shot 19·sharpness 2776.5

    1. AlEngineer0.98
    2. World's Fair0.98
    3. 601.00
    4. AIME 20250.99
    5. 601.00
    6. LiveCodeBench1.00
    7. 601.00
    8. GPQA Diamond1.00
    9. OpenThoughts1.00
    10. 501.00
    11. 501.00
    12. 501.00
    13. (%)1.00
    14. 401.00
    15. 401.00
    16. 401.00
    17. ACcuacy0.69
    18. 301.00
    19. 301.00
    20. 301.00
    21. 201.00
    22. 201.00
    23. 201.00
    24. 101.00
    25. 101.00
    26. 101.00
    27. 1K1.00
    28. 10K 100K0.99
    29. 1M1.00
    30. 1K1.00
    31. 10K 100K0.94
    32. 1M1.00
    33. 1K1.00
    34. 10K 100K1.00
    35. 1M1.00
    36. Dataset Size1.00
    37. Dataset Size1.00
    38. Dataset Size1.00
    39. OpenThoughts31.00
    40. AM1.00
    41. LIMO1.00
    42. Nemotron Nano1.00
    43. s1.11.00
    44. Qwen-2.5-7B-Instruct1.00
    45. Eric Horvitz@erichorvitz·Apr 100.99
    46. Thanks @AlexGDimakis and colleagues for your efforts to create and share0.99
    47. "The OpenThoughts reasoning datasets are valuable0.99
    48. OpenThoughts with the community. An valuable & enabling resource.0.99
    49. artifacts for studying fine-tuning behavior, and my1.00
    50. Alex Dimakis@AlexGDimakis·Apr90.99
    51. colleagues and I have used them multiple times in our0.99
    52. Very excited that Microsoft is using our dataset OpenThoughts and1.00
    53. research." - John Schulman0.98
    54. summarizes it in this super-clever way to make reasoning much more1.00
    55. efficient.1.00
    56. Most people do not understand how verbose reasoning models can g...0.99
    57. B0.74
    58. TRACK 9·JUNE 30,20260.98
    59. World's Fair0.99
    60. Data Quality1.00
  • 7:38 #20 skipped

    shot 20·duplicate of #19

  • 8:04 #21 skipped

    shot 21·duplicate of #19

  • 8:15 #22 skipped

    shot 22·duplicate of #6

  • 8:42 #23 skipped

    shot 23·duplicate of #6

Transcript

175 cues· 2,742 words· 15,103 chars

  1. 0:12 Hey, everyone.
  2. 0:15 Today, I'll be talking about data and environment curation for post-training LLMs.
  3. 0:21 And I am Mahesh Satyamurthy.
  4. 0:24 I'm co-founder and CEO of Bespoke Labs.
  5. 0:27 And previously, I was a researcher and engineer at Google DeepMind.
  6. 0:32 So very briefly, I will tell you a little bit about Bespoke.
  7. 0:36 And after that, the talk will be mostly around open source work we have done.
  8. 0:41 So Bespoke is an applied data research lab with a mission to help enterprises and frontier labs access high quality data and RLN moments for their post-training needs.
  9. 0:53 So very briefly, what we do and what we have done is that last year we put out something called
  10. 0:59 Curator, which is a tool for curating synthetic data for post-training with basically SFT.
  11. 1:06 And right after that, actually DeepSeq landed and we started an effort to curate reasoning data.
  12. 1:12 And that's how we started something called Bespoke Stratos, which eventually formed into the project called OpenThoughts, which some of you hopefully know about.
  13. 1:22 And we have also been core contributors to Terminal Bench.
  14. 1:27 These days, we do a lot of research and build and ship RL environments.
  15. 1:32 So I was actually looking forward to the previous talk from Nick, who's also doing something similar.
  16. 1:40 And the other thing we do is we do a lot of post-training and help enterprises to get their own custom models.
  17. 1:49 That's the name of, that's how we ended up with the bespoke title for the company.
  18. 1:57 The other thing I want to kind of mention is there is this, in our industry, there are a lot of people who create data, create RL environments, and then there are the researchers who consume this.
  19. 2:09 But I feel like there is this slight mismatch and it's kind of beneficial for someone to kind of go do both at the same time.
  20. 2:18 And in fact, as you're curating data, you want to put yourself in the shoes of the researcher to see what does it take to actually move the metrics on the models.
  21. 2:30 So that's one of the motivations of how we kind of think about.
  22. 2:33 The other thing I want to kind of talk about is,
  23. 2:37 how AI has evolved.
  24. 2:39 So early on, we used to think about and evaluate models on what they know.
  25. 2:46 For example, this was a very popular benchmark on testing LLMs on various kinds of STEM, humanities, and all that knowledge.
  26. 2:56 And these days, we have all these benchmarks that test
  27. 3:00 how agents are able to do things.
  28. 3:03 We have moved on from knowing to doing, right?
  29. 3:06 So that's the idea of agents, obviously.
  30. 3:08 And one of the key principles or one of the key things about agents is that they are autonomous.
  31. 3:14 And there are, as I was saying, there are many benchmarks, including SWE Bench, Terminal Bench, and so on.
  32. 3:21 But ultimately for many people, what they care about is are these agents autonomous for long durations of time?
  33. 3:31 Ross had a great talk on long horizon, right?
  34. 3:33 So that's the goal is eventually we make these agents
  35. 3:38 autonomous for maybe a few hours or a few days or a few weeks.
  36. 3:43 And what is it that's blocking the autonomy of agents?
  37. 3:48 It's basically reliability, right?
  38. 3:50 So at some point, something falls apart, like either they call the wrong tool or they made a mistake and whatnot, right?
  39. 3:58 And what's one lever to improve reliability?
  40. 4:02 There are, of course, many.
  41. 4:04 Obviously, you can prompt your way to improving the agent's reliability, or you can update the harness, you know, the tools and whatnot.
  42. 4:14 But post-training is a very powerful tool to improve reliability, or maybe even pre-trained good models, right?
  43. 4:21 So if you think of Frontier Labs, this is one of their primary mechanisms of improving agents to get better capabilities and capabilities
  44. 4:34 various domains or for better benchmark numbers or better autonomy for longer and longer durations.
  45. 4:46 And for post-training, one of the popular techniques, as you know, is reinforcement learning.
  46. 4:51 And that's kind of something a lot of you are excited about is the notion of RL environments.
  47. 4:59 But ultimately for post-training, be it SFT or reinforcement learning, data is the bottleneck, right?
  48. 5:07 So when I talk about data, RLNs are also something I'm calling it as data.
  49. 5:11 It's just the data is now in a very different shape.
  50. 5:16 Again, here, you know, compute this kind of well-defined models, you know, good sort of models exist and the infrastructure to post train, for example, there are various providers like Fireworks, Tinker, or Slime World and whatnot.

Chapters

  1. 0:00 Standing in the researcher's shoes
  2. 1:30 Post-training at Bespoke Labs
  3. 3:13 When agents fall over on long tasks
  4. 4:44 RL environments as data
  5. 6:29 Building OpenThoughts
  6. 7:36 Finding a curation recipe
  7. 10:27 Counterintuitive lessons
  8. 13:49 A credit card compliance example
  9. 16:13 Curating reasoning data with Curator
  10. 17:16 The full curation stack

Open at this second