Videos ewtOo0scUh0
Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
Scene timeline
53 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 175
- whisperx 175
- chunks
- 33
- from 175 cues
- keyframes
- 19
- kept of 53 captured
- frames with text
- 19
- 421 lines read
- chapters
- 10
- from the source metadata
- keyframe bytes
- 6.1 MB
- word timings on 175 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:27 | 0s |
stt |
done | — | 2026-08-09 14:21 | 19s |
chunk |
done | — | 2026-08-09 14:22 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 0s |
keyframe |
done | — | 2026-08-09 14:22 | 2m 31s |
ocr |
done | — | 2026-08-09 14:24 | 8s |
frame_embed |
done | — | 2026-08-10 19:42 | 3s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.93
- OpenAI0.92
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.96
- BESPOKE LABS1.00
- Data & Environment1.00
- Curation for1.00
- MAHESH1.00
- Post-training LLMs1.00
- SATHIAMOORTHY1.00
- CEO, BESPOKE LABS0.97
- Engineering the future of Al1.00
- World's Fair0.99
-
- AlEngineer0.97
- Who & What0.94
- World'sFair1.00
- Who we are0.98
- Bespoke is an applied data research lab with a mission to help enterprises0.99
- and frontier labs access high quality data and RL Environments for their1.00
- PRESENTED BY0.99
- post-training needs.1.00
- Microsoft1.00
- What we do1.00
- Last year we put out tooling for data curation (Curator)0.99
- We started Bespoke-Stratos which became OpenThoughts0.99
- Contributors to Terminal Bench benchmark0.99
- These days we research, build, and ship RL Envs1.00
- We enable post-training for enterprises (and labs)1.00
- B0.73
- Engineering the future of Al0.99
- World'sFair1.00
-
- AlEngineer0.97
- Who & What0.93
- World'sFair1.00
- Who we are0.98
- Bespoke is an applied data research lab with a mission to help enterprises0.99
- and frontier labs access high quality data and RL Environments for their0.99
- PRESENTED BY0.97
- post-training needs.1.00
- Microsoft1.00
- What we do1.00
- Last year we put out tooling for data curation (Curator)1.00
- We started Bespoke-Stratos which became OpenThoughts1.00
- Contributors to Terminal Bench benchmark1.00
- These days we research, build, and ship RL Envs1.00
- We enable post-training for enterprises (and labs)1.00
- B0.61
- TRACK 9· JUNE 30, 20260.93
- World'sFair1.00
- Data Quality0.99
-
- AlEngineer0.96
- Who & What0.92
- World'sFair1.00
- Who we are0.99
- Bespoke is an applied data research lab with a mission to help enterprises0.99
- and frontier labs access high quality data and RL Environments for their1.00
- post-training needs.1.00
- What we do1.00
- Last year we put out tooling for data curation (Curator)0.99
- We started Bespoke-Stratos which became OpenThoughts1.00
- Contributors to Terminal Bench benchmark1.00
- These days we research, build, and ship RL Envs0.99
- We enable post-training for enterprises (and labs)0.99
- B0.64
- TRACK 9· JUNE 30,20260.96
- World'sFair1.00
- Data Quality0.98
-
- AlEngineer0.99
- World'sFair0.98
- MEASURING MASSIVE MULTITASK1.00
- LANGUAGE UNDERSTANDING1.00
- Dan Hendrycks1.00
- Collin Burns1.00
- Steven Basart1.00
- Andy Zou1.00
- What do they0.99
- UC Berkeley0.97
- Columbia University1.00
- UChicago1.00
- UC Berkeley1.00
- know1.00
- Mantas Mazeika1.00
- Dawn Song1.00
- Jacob Steinhardt1.00
- UIUC1.00
- UC Berkeley0.98
- UC Berkeley0.96
- SWE-BENCH: CAN LANGUAGE MODELS RESOLVE0.98
- REAL-WORLD GITHUB ISSUES?0.99
- What can they1.00
- do?1.00
- Carlos E. Jimenez* 1,2 John Yang* 1,2 Alexander Wettig1,20.98
- Shunyu Yao1,2 Kexin Pei3 Ofir Press1,2 Karthik Narasimhan1,20.97
- 1Princeton University0.99
- ²Princeton Language and Intelligence0.99
- 3University of Chicago0.99
- TRACK 9· JUNE 30,20260.95
- World's Fair0.99
- Data Quality0.98
-
- AlEngineer0.99
- World'sFair1.00
- What's blocking autonomy for agents?1.00
- Reliability1.00
- How to improve reliability?1.00
- Post-training1.00
- TRACK 9• JUNE 30, 20260.97
- World'sFair1.00
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- What's blocking autonomy for agents?1.00
- Reliability1.00
- PRESENTED BY1.00
- How to improve reliability?1.00
- Microsoft1.00
- Post-training1.00
- What's the bottleneck for post-training?1.00
- Data (and RL Envs)0.99
- TRACK 9· JUNE 30, 20260.96
- World'sFair1.00
- Data Quality1.00
-
- AlEngineer0.99
- World's Fair0.94
- OpenThoughts1.00
- TRACK 9• JUNE 30, 20260.97
- World'sFair1.00
- Data Quality1.00
-
- AlEngineer0.98
- World's Fair0.98
- 601.00
- AIME 20250.99
- 601.00
- LiveCodeBench1.00
- 601.00
- GPQA Diamond1.00
- OpenThoughts1.00
- 501.00
- 501.00
- 501.00
- (%)1.00
- 401.00
- 401.00
- 401.00
- ACcuacy0.69
- 301.00
- 301.00
- 301.00
- 201.00
- 201.00
- 201.00
- 101.00
- 101.00
- 101.00
- 1K1.00
- 10K 100K0.99
- 1M1.00
- 1K1.00
- 10K 100K0.94
- 1M1.00
- 1K1.00
- 10K 100K1.00
- 1M1.00
- Dataset Size1.00
- Dataset Size1.00
- Dataset Size1.00
- OpenThoughts31.00
- AM1.00
- LIMO1.00
- Nemotron Nano1.00
- s1.11.00
- Qwen-2.5-7B-Instruct1.00
- Eric Horvitz@erichorvitz·Apr 100.99
- Thanks @AlexGDimakis and colleagues for your efforts to create and share0.99
- "The OpenThoughts reasoning datasets are valuable0.99
- OpenThoughts with the community. An valuable & enabling resource.0.99
- artifacts for studying fine-tuning behavior, and my1.00
- Alex Dimakis@AlexGDimakis·Apr90.99
- colleagues and I have used them multiple times in our0.99
- Very excited that Microsoft is using our dataset OpenThoughts and1.00
- research." - John Schulman0.98
- summarizes it in this super-clever way to make reasoning much more1.00
- efficient.1.00
- Most people do not understand how verbose reasoning models can g...0.99
- B0.74
- TRACK 9·JUNE 30,20260.98
- World's Fair0.99
- Data Quality1.00
Transcript
175 cues· 2,742 words· 15,103 chars
- 0:12 Hey, everyone.
- 0:15 Today, I'll be talking about data and environment curation for post-training LLMs.
- 0:21 And I am Mahesh Satyamurthy.
- 0:24 I'm co-founder and CEO of Bespoke Labs.
- 0:27 And previously, I was a researcher and engineer at Google DeepMind.
- 0:32 So very briefly, I will tell you a little bit about Bespoke.
- 0:36 And after that, the talk will be mostly around open source work we have done.
- 0:41 So Bespoke is an applied data research lab with a mission to help enterprises and frontier labs access high quality data and RLN moments for their post-training needs.
- 0:53 So very briefly, what we do and what we have done is that last year we put out something called
- 0:59 Curator, which is a tool for curating synthetic data for post-training with basically SFT.
- 1:06 And right after that, actually DeepSeq landed and we started an effort to curate reasoning data.
- 1:12 And that's how we started something called Bespoke Stratos, which eventually formed into the project called OpenThoughts, which some of you hopefully know about.
- 1:22 And we have also been core contributors to Terminal Bench.
- 1:27 These days, we do a lot of research and build and ship RL environments.
- 1:32 So I was actually looking forward to the previous talk from Nick, who's also doing something similar.
- 1:40 And the other thing we do is we do a lot of post-training and help enterprises to get their own custom models.
- 1:49 That's the name of, that's how we ended up with the bespoke title for the company.
- 1:57 The other thing I want to kind of mention is there is this, in our industry, there are a lot of people who create data, create RL environments, and then there are the researchers who consume this.
- 2:09 But I feel like there is this slight mismatch and it's kind of beneficial for someone to kind of go do both at the same time.
- 2:18 And in fact, as you're curating data, you want to put yourself in the shoes of the researcher to see what does it take to actually move the metrics on the models.
- 2:30 So that's one of the motivations of how we kind of think about.
- 2:33 The other thing I want to kind of talk about is,
- 2:37 how AI has evolved.
- 2:39 So early on, we used to think about and evaluate models on what they know.
- 2:46 For example, this was a very popular benchmark on testing LLMs on various kinds of STEM, humanities, and all that knowledge.
- 2:56 And these days, we have all these benchmarks that test
- 3:00 how agents are able to do things.
- 3:03 We have moved on from knowing to doing, right?
- 3:06 So that's the idea of agents, obviously.
- 3:08 And one of the key principles or one of the key things about agents is that they are autonomous.
- 3:14 And there are, as I was saying, there are many benchmarks, including SWE Bench, Terminal Bench, and so on.
- 3:21 But ultimately for many people, what they care about is are these agents autonomous for long durations of time?
- 3:31 Ross had a great talk on long horizon, right?
- 3:33 So that's the goal is eventually we make these agents
- 3:38 autonomous for maybe a few hours or a few days or a few weeks.
- 3:43 And what is it that's blocking the autonomy of agents?
- 3:48 It's basically reliability, right?
- 3:50 So at some point, something falls apart, like either they call the wrong tool or they made a mistake and whatnot, right?
- 3:58 And what's one lever to improve reliability?
- 4:02 There are, of course, many.
- 4:04 Obviously, you can prompt your way to improving the agent's reliability, or you can update the harness, you know, the tools and whatnot.
- 4:14 But post-training is a very powerful tool to improve reliability, or maybe even pre-trained good models, right?
- 4:21 So if you think of Frontier Labs, this is one of their primary mechanisms of improving agents to get better capabilities and capabilities
- 4:34 various domains or for better benchmark numbers or better autonomy for longer and longer durations.
- 4:46 And for post-training, one of the popular techniques, as you know, is reinforcement learning.
- 4:51 And that's kind of something a lot of you are excited about is the notion of RL environments.
- 4:59 But ultimately for post-training, be it SFT or reinforcement learning, data is the bottleneck, right?
- 5:07 So when I talk about data, RLNs are also something I'm calling it as data.
- 5:11 It's just the data is now in a very different shape.
- 5:16 Again, here, you know, compute this kind of well-defined models, you know, good sort of models exist and the infrastructure to post train, for example, there are various providers like Fireworks, Tinker, or Slime World and whatnot.
loading
Chapters
- 0:00 Standing in the researcher's shoes
- 1:30 Post-training at Bespoke Labs
- 3:13 When agents fall over on long tasks
- 4:44 RL environments as data
- 6:29 Building OpenThoughts
- 7:36 Finding a curation recipe
- 10:27 Counterintuitive lessons
- 13:49 A credit card compliance example
- 16:13 Curating reasoning data with Curator
- 17:16 The full curation stack