Videos q4Tr-DknG2M
Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI
Scene timeline
56 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 199
- whisperx 199
- chunks
- 35
- from 199 cues
- keyframes
- 48
- kept of 56 captured
- frames with text
- 47
- 1,213 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 6.2 MB
- word timings on 199 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 02:52 | 1m 15s |
stt |
done | — | 2026-08-11 02:53 | 23s |
chunk |
done | — | 2026-08-11 02:53 | 0s |
text_embed |
done | — | 2026-08-11 02:53 | 1s |
keyframe |
done | — | 2026-08-11 02:53 | 2m 11s |
ocr |
done | — | 2026-08-11 02:56 | 20s |
frame_embed |
done | — | 2026-08-11 02:56 | 8s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.96
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.96
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- CURSOR1.00
- AIE1.00
-
- LEE ROBINSON0.99
- ML, MODEL BEHAVIOR0.99
- CURSOR1.00
- AIE0.97
-
- LEE ROBINSON0.99
- sneos0.61
- promptql0.92
- jineer0.90
- 'sFair0.99
- ML, MODEL BEHAVIOR0.98
- WoridsFair0.93
- CURSOR1.00
- Cleric0.99
- Vorld's Fair0.86
- Microsoht0.92
- AIE1.00
- Greducto0.99
- Worids Fair0.94
- World's Fair0.93
- Ravenna0.99
- Workd'sFa0.93
- eNCORD0.93
-
- World'sFair1.00
- World's Fair0.98
- World's Fai0.97
- World's Fair0.94
- World's Fair0.97
- Red Hat0.96
- World's Fain0.95
- World's Fair0.99
- mezmo*0.94
- fiddler0.97
- sigg0.79
- World's Fair0.96
- Cate0.72
- INNGEST0.97
- Snorkel0.98
- World's Fair0.96
- PRIOR1.00
- World's Fair0.97
- BAND0.98
- Keycard1.00
- AUTOMATTIC1.00
- orld's Fair0.97
- Z.AI0.95
- World's Fair0.93
- Worlds Fair0.83
- WorkOS0.97
- :neo4j0.97
- World's Fair0.98
- Google DeepMind1.00
- World's Fair0.99
- qodo0.99
- World's Fain0.94
- Cleric1.00
- Wor0.99
- A ATLASSIAN0.92
- GGRAVITEE0.97
- Meticulous1.00
- rld's Fair0.99
- World's Fair0.99
- Braintrust1.00
- Amazon AGI Lab1.00
- World's Fair0.97
- World's Fair0.96
- World's Fair0.98
- Microsoft1.00
- World's Fair0.96
- arize0.96
- Pr0.99
- reducto1.00
- Gradium0.97
- FACTORY1.00
- Microsoft1.00
- ANTHROPIC1.00
- Ddud0.56
- OpenAl0.90
- aws1.00
- OpenAl0.92
- World's Fair0.91
- bright data0.86
- Worids Fair0.88
- World's Fair0.99
- MINIMAX1.00
- World's Fair0.98
- L.AI0.90
- World'sFai1.00
- ANTHROPIC1.00
- World's Fair0.95
- d's Fair0.98
- World's Fair0.98
- DATADOG1.00
- World's Fair0.96
- Resolve.ai1.00
- CODER1.00
- d'sFai0.94
- @Airbyte0.95
- Optiver0.99
- greptle0.91
- World's Fair0.97
- builder.io1.00
- World's Fair0.98
- Ravenna1.00
- World'sFair1.00
- tKi0.96
- 's Fair0.98
- World's Fair0.97
- cognee1.00
- World'sFair1.00
- BeNCORD0.89
- I'sF0.80
- World's Far0.97
- tgus0.77
-
- Fair0.98
- ORACLE1.00
- World's Fair0.95
- arize1.00
- World's Fair0.96
- Google DeepMind1.00
- World's Fair0.89
- :neo4j0.95
- World's Fair0.90
- Z.AI0.99
- World's Fair0.85
- bright data0.99
- World's Fair0.96
- Browserba1.00
- mpute co.0.98
- World's Fair0.90
- extend1.00
- World's Fair0.96
- vast.ai1.00
- World's Fair0.95
- Ref.1.00
- World's Fair0.98
- RedHat1.00
- World's Fair0.99
- mezmo*0.97
- Worid's Fair0.96
- stigg0.92
- World's Fair0.94
- sFair0.86
- RELI0.95
- World's Fair0.98
- VAPI1.00
- World's Fair0.99
- Modal1.00
- World's Fair0.98
- promptql1.00
- World's Fair0.96
- THE VELOCITY ROOM0.97
- World's Fair0.96
- fiddler0.94
- World's Fair0.98
- Superconduc0.98
- eal0.96
- World's Fair0.96
- ZERO0.88
- World's Fair0.99
- AlEngineer0.97
- comet1.00
- World's Fair0.97
- authe0.98
- World's Fair0.95
- sFair0.91
- SOIO.IO0.90
- World's Fair0.96
- granica1.00
- World's Fair0.99
- dash00.99
- World's Fair0.99
- PRIOR1.00
- World's Fair1.00
- alOcean1.00
- World's Fair0.95
- Venmce0.79
- World's Fair0.95
- POSTMAN1.00
- World's Fair0.99
- Composio1.00
- World's Fair0.95
- sFair0.90
- Modular1.00
- Worid's Fair0.97
- MERGE1.00
- World's Fair0.97
- AUTOMATTIC1.00
- World's Fair0.98
- Buildkit0.99
- byteDB1.00
- Worid's Fair0.93
- Zed1.00
- World's Fair0.94
- Keycard1.00
- Worid's Fair0.94
- Meticulous0.99
- World's Fair0.91
- sFair0.93
- Daytona1.00
- World's Fair0.99
- twilio0.96
- World's Fair0.96
- C GRAVITEE0.93
- World's Fair0.93
- PlanetScal0.98
- nporal0.98
- World's Fair0.97
- Llamalndex0.95
- World's Fair0.98
- INNGEST0.94
- World's Fair0.96
- © BAND0.88
- World's Fair0.95
- Cleric1.00
- World's Fair0.98
- À ATLASSIAN0.95
- World's Fair0.98
- FACTORY1.00
- World's Fain0.93
- sFair0.88
- baseten1.00
- World's Fair0.98
- Snorkel1.00
- World's Fair0.93
- Z.AI0.97
- World's Fair0.97
- qodo0.95
- World's Fair0.96
- PayPal1.00
- World's Fair0.95
- Gradium1.00
- World's Fair0.96
- ANTHROPV0.99
- AGI Lab1.00
- World's Fair0.98
- Browserbase1.00
- World's Fair1.00
- :neo4j0.81
- World's Fair0.98
- Google DeepMind0.99
- World's Fair0.96
- arize1.00
- World's Fair0.96
- reducto1.00
- World's Fair0.94
- Microsoft0.95
- World's Fai0.95
- sFair0.86
- OpenAl0.98
- World's Fair0.98
- WorkOS0.88
- World's Fair0.97
- Amazon AGI Lab1.00
- World's Far0.94
- Microsoft1.00
- World's Fair0.96
- ORACLE1.00
- World's Fair0.89
- bright data0.98
- World's Fair1.00
- Google DeepM1.00
- rosoft1.00
- World's Fair0.91
- docker1.00
- World's Fair0.95
- Braintrust1.00
- World's Fair0.97
- OpenA'0.99
- rid's Fair0.96
- MINIMAX0.99
- World's Fair0.97
- Z.AI0.99
- World's Fair0.99
- ANTHROPIC1.00
- World's Fai0.96
- Fair0.86
- snyk1.00
- World's Fair0.97
- aws0.98
- World's Fair0.94
- togetherai1.00
- World'sF0.99
- Akamai0.98
- World's Fair0.94
- Unblocked0.96
- Worldin0.87
- Airbyte1.00
- Worl0.87
- ghtrun1.00
- ain0.97
- World's Fair0.97
- World's Fair0.98
- DATADOC0.95
- orld's Fair0.93
- Resolve.ai1.00
- Word's Far0.86
- World's Fa0.94
- Op1.00
- Worid's Fal0.92
- Fair1.00
- LanceDB0.97
- World's Fair0.98
- builderio0.95
- World's Fair0.97
- : Ravenna0.97
- World's Fair0.99
- CopilotKit0.99
- VERIS1.00
- Wor1.00
- World's Far0.93
- pe0.98
- Microsoft0.97
- World's Fair0.99
- TOP1.00
- World's Fair0.95
- cognee1.00
- forld's Fair0.92
- eNCORD0.93
- World's Fair0.92
- rid'sFa0.97
- ti0.66
- World's Fai0.98
- ●。。0.85
-
- AlEngineer0.98
- Amazon A1.00
- World's Fair0.99
- AIEngineer0.95
- OpenAI0.94
- World's Fair1.00
- AlEngineer1.00
- together1.00
- World's Fair0.99
-
- orkOS0.99
- WorFair0.98
- Amazon1.00
- gineer1.00
- AIEngi0.94
- I's Fair0.90
- World1.00
- VS0.95
- toge0.98
- 70.54
-
- AlEngineer0.99
- WorkOS1.00
- orld's Fair1.00
- Ama1.00
- AlEngineer0.98
- lorld's Fair0.96
- intrust1.00
- W1.00
- eer1.00
- aws1.00
- sFair1.00
-
- AlEngineer0.98
- World'sFair1.00
- Goal: build the best general Al model1.00
- AlEngineer1.00
- AGI Lab0.97
- Norld'sl0.96
- 'sFair0.93
- neer1.00
- Open1.00
- The Future of Cursor0.98
- - AlEngineer0.93
- etherai0.99
- rld'sl0.95
- Lee Robinson / ML, Model Behavior0.97
- CURSOR1.00
-
- AlEngineer0.99
- World's Fair0.97
- User feedback1.00
- Scale/improve data1.00
- Scale/improve training1.00
- Better model1.00
- AIE1.00
- WorkOs0.95
- Worl1.00
- - AlEnginer0.96
- orld'!0.75
- Bra0.91
- The Future of Cursor1.00
- al0.97
- /orl0.98
- — AIE0.90
- Lee Robinson / ML, Model Behavior0.99
- CURSOR1.00
-
- AlEngineer0.98
- World's Fair0.97
- User feedback, online metrics1.00
- High quality evals, hard training tasks1.00
- Outer loop0.99
- Inner loop:0.97
- Climb evals1.00
- Collect & create data, shape reward0.99
- Better model1.00
- d'sFair1.00
- azon1.00
- AlEngine0.98
- intru1.00
- World's1.00
- The Future of Cursor1.00
- d'sFai0.96
- ngineer1.00
- togel0.95
- Lee Robinson / ML, Model Behavior0.98
- CURSOR1.00
-
- AlEngineer0.98
- Terminal-bench 2.0 Score0.99
- Composer 2.51.00
- World'sFair1.00
- 70%1.00
- Composer 20.99
- 60%1.00
- 50%1.00
- Composer 1.50.99
- Composer 10.96
- WorkOS1.00
- Wc0.99
- 40%1.00
- Nov'251.00
- Dec'251.00
- Jan'260.99
- Feb '260.94
- Mar'261.00
- Apr'260.99
- May'261.00
- Jun '260.95
- AlEngineer1.00
- Norld's1.00
- E0.98
- The Future of Cursor1.00
- av0.94
- Ic0.75
- Lee Robinson / ML, Model Behavior0.98
- CURSOR1.00
-
- 75% CursorBench 3.1 score0.99
- AlEngineer1.00
- Fable 5 high1.00
- World'sFair1.00
- 70%1.00
- (default)1.00
- 65%1.00
- Composer 2.50.98
- 60%1.00
- GPT-5.5 medium1.00
- PRESENTED BY0.96
- 55%1.00
- Opus 4.8 high1.00
- (default)1.00
- (default)1.00
- Microsoft1.00
- 50%1.00
- Gemini 3.5 Flash1.00
- Sonnet 4.6 high1.00
- 45%1.00
- (default)1.00
- Kimi 2.60.99
- 40%1.00
- 35%1.00
- 30%1.00
- AlEngineer0.99
- $201.00
- $161.00
- $121.00
- $81.00
- $41.00
- $00.99
- rld's Fai0.94
- OR/0.96
- Average cost per task1.00
- AIEr0.77
- NM0.66
- Worl0.92
- The Future of Cursor1.00
- AlI0.60
- rl0.99
- Un1.00
- Lee Robinson / ML, Model Behavior0.98
- CURSOR1.00
-
- AlEngineer0.99
- World'sFair1.00
- Artificial Analysis Coding Agent Index1.00
- Composite average pass@1 across SWE-Bench-Pro-Hard-AA, Terminal-Bench v2, and1.00
- SWE-Atlas-QnA· Higher is better0.98
- Artificial Analysis0.98
- PRESENTED BY0.98
- Microsoft1.00
- 651.00
- 631.00
- 611.00
- 581.00
- 531.00
- 501.00
- 501.00
- 481.00
- 431.00
- Al Claude Code0.93
- Codex1.00
- Cursor CLI0.95
- Cursor CLI0.92
- Cursor CUI0.90
- Al Claude Code0.94
- Al Claude Code0.94
- Al Claude Code0.98
- Cursor CLI1.00
- G Gemini CU0.94
- Al Opus 4.70.96
- (Max)0.99
- GPT-5.50.99
- (XHigh)0.96
- Composer 2.51.00
- Fast1.00
- Al Opus 4.70.97
- (Medium)0.98
- (Medium)0.97
- GPT-5.51.00
- GLM-510.98
- 区0.62
- Kimi K2.61.00
- ar0.63
- DeepSeek V40.98
- (High)1.00
- Pro1.00
- Composer 21.00
- G1.00
- Gemini 3.10.99
- (High)0.92
- Pro0.97
- AlEngineer1.00
- rld's Fai0.93
- Amazor1.00
- AIEr0.90
- raintrust1.00
- 'ork0.97
- The Future of Cursor0.98
- AlEngineer0.98
- rld's Fair0.98
- tog0.93
- Lee Robinson/ ML, Model Behavior0.99
- CURSOR1.00
-
- AlEngineer0.99
- What to improve1.00
- World's Fair0.98
- • Bigger& smarter model0.96
- Faster training stack0.94
- PRESENTED BY0.99
- • Full pretrain to control model behavior0.97
- · Intelligent beyond coding0.96
- Microsoft1.00
- Multi-effort0.96
- • More diverse, harder data0.96
- • New evals0.95
- ·Scale RL further0.95
- Al0.94
- WorkO1.00
- Wor1.00
- — AlEngineer0.97
- orld's1.00
- Br1.00
- The Future of Cursor0.99
- av.0.93
- − AI0.71
- orl0.93
- Lee Robinson/ ML, Model Behavior0.98
- CURSOR1.00
-
- AlEngineer0.99
- World's Fair0.97
- PRESENTED BY1.00
- Improving the outer loop1.00
- Microsoft1.00
The page's on-screen-text budget of 600
lines is spent, so the last cards in this grid list fewer lines than they
hold. Narrow the page with ?frames= to read them.
Transcript
199 cues· 3,959 words· 20,860 chars
- 0:13 Please welcome to the stage the machine learning engineer, Model Behavior at Cursor, Lee Robinson.
- 0:36 All right.
- 0:36 Hey, everyone.
- 0:39 I'm excited to be here, excited to be back at AI Engineer and talk a little bit about how we're training models at Cursor.
- 0:47 So how we train the models and also how the model learns to train itself or recursive model improvement.
- 0:54 So our goal at Cursor is to build the best possible AI models, which might make sense.
- 1:01 You might have heard of this equation of if we just give the models more compute, we can get a better model out.
- 1:06 And I think this is a helpful simplification of the problem, but I want to actually click in a few layers deeper in the talk today and talk about all the different pieces that go into training these models.
- 1:16 So we can think about this loop.
- 1:19 We put a model out into the world, and then we get feedback from you all when you use the model, what goes well or places we can improve.
- 1:25 We use that to scale and improve the data that we do for the next round of training.
- 1:29 And we also then increase the amount of compute and scale up training overall to make a new model.
- 1:35 And this is a loop that can just go over and over again.
- 1:38 However, if you see my helpful snail or turtle to bunny meter down in the bottom right, it's pretty slow.
- 1:45 This is gonna be a serial process.
- 1:47 And you can only do, in this instance, one big run at a time.
- 1:51 So we wanna make this a little bit faster.
- 1:53 But I'll actually go another layer deeper and add some more color here.
- 1:57 There's actually two loops, the outer loop and the inner loop.
- 2:01 On the outer loop, we have the feedback coming in, but we also have data like online metrics.
- 2:07 So running A-B tests and seeing what users prefer a different checkpoint of a model.
- 2:11 That's going to then flow into hopefully making better high quality evals that
- 2:16 help ensure we're getting the right behaviors we want out of the model, and also being able to create much more difficult problems for the models to try to solve where we then can kind of shape the rewards that we want to get during training.
- 2:28 So we want to climb that inner loop as well.
- 2:33 We have been training models for about a year at the large scale at Cursor, and I want to talk about some of our progress so far.
- 2:41 So we put out Composer 2.5 in May, and it's now the most popular model in Cursor, which is exciting.
- 2:48 And we scaled up training here quite a bit by generating more RL environments, trying out some new methods for learning, and also just making more ambitious problems for the models to solve.
- 3:00 And the results have been pretty promising so far.
- 3:03 Like I mentioned, this is still a new effort for us.
- 3:06 ML has really been in the blood of Cursor since the start where we were training more specialized models for things like tab or code autocomplete.
- 3:13 But really in the past year we've staffed up and built a team with ambitions to train state of the art models and we made some pretty good progress just in the last 12 months.
- 3:22 People like Composer right now, I think because it is both fast and pretty smart, and also cost effective.
- 3:29 And as we've heard from other speakers today, I think there is a space in the market right now for that type of model, in addition to also having the most intelligent models in the world.
- 3:39 And we think it's important to have a good selection of both of these type of things.
- 3:42 So Composer we think is serving a good niche here.
- 3:46 And when we released it, we were honestly pretty impressed with some of the public evals.
- 3:51 It did a little better than we expected.
- 3:53 On artificial analysis, it was a pretty modest jump.
- 3:56 However, there were a lot of behaviors that we found that we really wanted to improve for the next version of the model.
- 4:02 Notably, we wanted to have a much bigger and smarter model.
- 4:05 We wanted to control every aspect of training.
- 4:07 So ideally, doing a full pre-train from scratch versus the previous open source base of Kimmy that we were using.
- 4:13 We wanted to infuse new data so that we can make the model great outside of more things than just coding, but more of a general model.
- 4:20 And then also just scale up every part of the training process.
- 4:24 More data, more compute, and really pushing RL as far as we can.
- 4:29 So first I want to talk about improving the outer loop, and then we'll drill into the inner loop.
- 4:35 If you haven't used Cursor in a while, you might think about it as this IDE or tab autocomplete thing.
- 4:42 And in reality, the vast, vast majority of our revenue today comes from agent usage.
- 4:47 And that means that all of the data inside of Cursor is also coming from agent usage, and we can use that to train better models.
- 4:55 For example, we have two different buckets of feedback.
loading
Chapters
- 0:00 <Untitled Chapter 1>
- 0:37 Introduction and recursive model improvement overview
- 1:55 The two-loop training framework (inner and outer loops)
- 2:33 Progress and success of Composer 2.5
- 4:31 Improving the outer loop with user feedback
- 5:40 Climbing the inner loop with high-quality evals
- 6:52 Solving reward hacking in public benchmarks
- 8:27 Scaling training with ambitious problems
- 9:53 New learning methods: Teacher-student textual feedback
- 11:34 Scaling compute infrastructure with SpaceX and Colossus
- 13:06 Understanding compute allocation in model training
- 15:28 Agent-based automation and research efficiency
- 18:30 The recursive future: Models training models