Videos Yk87oUPVaxU
DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve
Scene timeline
36 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 128
- whisperx 128
- chunks
- 30
- from 128 cues
- keyframes
- 12
- kept of 36 captured
- frames with text
- 12
- 207 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 4.5 MB
- word timings on 128 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:27 | 0s |
stt |
done | — | 2026-08-09 06:22 | 18s |
chunk |
done | — | 2026-08-09 06:22 | 0s |
text_embed |
done | — | 2026-08-10 19:39 | 1s |
keyframe |
done | — | 2026-08-09 06:22 | 2m 07s |
ocr |
done | — | 2026-08-09 06:24 | 5s |
frame_embed |
done | — | 2026-08-10 19:39 | 2s |
Frames, and what the machine read
-
- AIEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.98
- World's Fair0.99
-
- AlEngineer0.99
- Datacurve1.00
- World'sFair1.00
- JULY 20260.99
- DeepSWE1.00
- PRESENTED BY1.00
- Microsoft1.00
- Measuring frontier coding agents on original, long-horizon1.00
- software engineering tasks.0.99
- James1.00
- Engineer @ Datacurve0.98
- Engineering the future of Al0.98
- World's Fair0.96
-
- AlEngineer0.94
- World'sFair1.00
- Overview1.00
- DeepSWE is 113 original, long-horizon tasks0.99
- 1131.00
- 910.99
- 51.00
- PRESENTED BY1.00
- Microsoft1.00
- authored tasks0.98
- active repositories1.00
- languages1.00
- Artificial Analysis1.00
- Replaced SWE-Bench Pro with DeepSWE in the Coding1.00
- Agent Index1.00
- TRACK 8• JULY 2, 20260.96
- World's Fair0.99
- Agentic Engineering1.00
-
- AlEngineer0.98
- Datacurve1.00
- World'sFair1.00
- Datacurve builds training data for high-1.00
- ceiling domains0.98
- 01 Research and training data — coding, and adjacent high-complexity technical domains.1.00
- 021.00
- In-house platforms — built to attract and collect the highest caliber of data.0.99
- 031.00
- Research on post-training — what is data quality? How does this demonstrably move0.99
- models?1.00
- TRACK 8· JULY 2,20260.96
- World'sFair1.00
- Agentic Engineering0.99
-
- AlEngineer1.00
- Results1.00
- v1.1 · pass@1 · mini-swe-agent · updated July 1, 20260.97
- World'sFair1.00
- Frontier models diverge on DeepSWE0.99
- #1.00
- MODEL1.00
- 40.99
- PASS@11.00
- COST1.00
- 011.00
- claude-fable-5 max0.99
- 70% ±40.94
- $21.631.00
- 020.98
- gpt-5.5 xhigh1.00
- 67% ±60.94
- $7.231.00
- 030.93
- claude-opus-4.8 max1.00
- 59% ±20.97
- $13.221.00
- 040.91
- claude-sonnet-5 max1.00
- 54% ±40.97
- $26.401.00
- 050.89
- gpt-5.4 xhigh0.99
- 52% ±20.94
- $5.651.00
- 060.86
- glm-5.2 max0.99
- 44% ±20.97
- $3.921.00
- 070.97
- gemini-3.5-flash medium0.99
- 37% ±21.00
- $7.341.00
- 080.84
- kimi-k2.7-code-1.00
- 31% ±10.98
- $2.821.00
- 890.83
- claude-sonnet-4.6 high0.99
- 30% ±40.98
- $5.521.00
- 100.99
- gemini-3.1-pro high1.00
- 12% ±20.97
- $9.481.00
- TRACK 8· JULY 2, 20260.95
- World's Fair1.00
- Agentic Engineering1.00
-
- AlEngineer0.99
- Findings 1/41.00
- World'sFair1.00
- Claude is forgetful with multi-part prompts0.98
- Prompts enumerate parallel behaviors — support both sync and async — and Claude ships the0.98
- obvious branch.1.00
- PRESENTED BY0.99
- python-statemachine-state-data-scoping1.00
- langchain-request-coalescing1.00
- Microsoft1.00
- sync hook lands correctly in0.97
- invoke coalesces; batch dispatches via executor.map0.99
- BaseEngine._enter_states1.00
- x batch(["hello", "hello", "world"]] → 3 executions, not 20.95
- x AsyncEngine never receives the same hook0.98
- ≈2/31.00
- of Claude rollouts tagged MISSED_REQUIREMENT fit this one-branch-shipped0.99
- pattern.1.00
- TRACK 8· JULY 2, 20260.95
- World'sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.97
- Findings 2/41.00
- World's Fair0.97
- Claude pays close attention to its1.00
- environment1.00
- When prompt and repo disagree, Opus runs git log — and recovers the gold patch from .git history.0.99
- PRESENTED BY1.00
- Microsoft1.00
- 25%0.99
- 18%1.00
- ≈1%1.00
- 0%0.96
- claude-opus-4.61.00
- claude-opus-4.71.00
- gemini configs1.00
- gpt-5.4 · gpt-5.50.98
- share of reviewed SWE-Bench Pro passes flagged CHEATED. 87% of flags = gold commit read from .git0.99
- The gold commit ships inside SWE-Bench Pro containers. DeepSWE v1.1 deletes future refs entirely.0.99
- TRACK 8· JULY 2, 20260.95
- World's Fair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- Task shape1.00
- World's Fair0.99
- Prompts are short; solutions are long1.00
- PROMPT LENGTH1.00
- SOLUTION SIZE1.00
- 2,1581.00
- 6681.00
- PRESENTED BY1.00
- chars1.00
- lines1.00
- Microsoft1.00
- vs 4,614 on SWE-Bench Pro0.99
- vs 120 on SWE-Bench Pro0.98
- FILES TOUCHED1.00
- OUTPUT TOKENS1.00
- 71.00
- ≈2×0.99
- files1.00
- vs 5 per reference solution1.00
- vs SWE-Bench Pro runs1.00
- TRACK 8· JULY 2, 20260.95
- World's Fair0.99
- Agentic Engineering1.00
Transcript
128 cues· 2,636 words· 15,254 chars
- 0:12 Hey, everyone.
- 0:13 Can you guys hear me okay?
- 0:14 This is good.
- 0:16 Yeah, my name is James.
- 0:17 I'm one of the founding engineers at DataCurve.
- 0:20 Unfortunately, Serena's been out with a fever for the past couple of days.
- 0:24 She was supposed to be here giving this talk, so I'm just filling in in her place.
- 0:28 But I've been at DataCurve working on the research and engineering side of things, as well as DeepSuite, which is our frontier long-horizon coding benchmark, which you guys may be familiar.
- 0:41 I'll just be going over some of the most important findings about DeepSuite, a brief overview of what it is for those of you who may not know, and then going deeper into our methodology and exactly how we came about this frontier coding benchmark.
- 0:56 So DeepSuite is a long horizon software engineering benchmark comprised of 113 original software engineering tasks.
- 1:07 So this means unlike something like SuiteBench Pro, we didn't scrape this from existing PRs that have been closed.
- 1:13 There's a variety of benefits for this, namely one of them is to resist against contamination and agents being able to cheat through the course of their rollouts.
- 1:23 SweetBench Pro pulls thousands of tasks from only 40 repositories.
- 1:29 The median task per repository for us is one, so you can see across over 100 tasks, we pull from nearly 100 repositories.
- 1:38 And the language spans across TypeScript, JavaScript, Python, Rust, and Go.
- 1:42 And we have plans to add more languages later on.
- 1:47 Since its release, we've received very positive reception.
- 1:51 It's replaced SweetBench Pro in the artificial analysis coding agent index, as well as being cited by numerous frontier model labs and us helping with them in tracking their models on our benchmark as well.
- 2:05 So we've been really, really appreciative of that.
- 2:08 A bit of context about us, DataCurve works on building training data for high ceiling domains, including coding, as well as coding adjacent fields.
- 2:18 We also are trying to answer the very elusive question of what exactly makes good data, what is data quality, and how can we demonstrate that our training data, in fact, moves the needle.
- 2:32 So DeepSuite is one in a long line of initiatives that we have towards answering this question.
- 2:40 So why did we create DeepSuite?
- 2:41 Well, it's very clear that the existing benchmarks are not hitting the mark.
- 2:47 With benches like SuiteBench Pro, top models are clustering at the top.
- 2:51 It's very hard to differentiate between which one is good because they all have overlapping confidence intervals.
- 2:58 Contamination is also rampant because again, all of these tasks are mined from public PRs.
- 3:03 So all the solution tests, even the discussion around the PRs, those are all available out in the wild for these agents to access.
- 3:10 The verifiers are also very, very brittle because we're anchoring them to a specific implementation often derived from the PR that was merged in.
- 3:19 And oftentimes you also have tests that check for private helpers and functions created by the task author, which is very opinionated, right?
- 3:27 And it's not something that model should have to adhere to.
- 3:31 And finally, leakage.
- 3:32 So one thing about SweetBench Pro is for very insightful models such as Claude, they're able to directly run git log and then go through the commit hashes and cherry pick the ones out that contain the golden patches, which again, very, very serious issue.
- 3:47 So this is DeepSuite.
- 3:48 This is the updated leaderboard as of July 1st.
- 3:51 You can see, I was mentioning before the problem of differentiating, but you can see on DeepSuite here, there's a very clear difference.
- 3:59 There's a very clear performance gap between the top performing models versus, you know, at 10th place, you have Gemini 3.1 Pro.
- 4:07 Also, within the cloud and the GPT models as well, we're able to see some deviance.
- 4:12 And yeah, if you go on deepsuite.datacurve.ai, you'll also be able to see the token efficiency costs, token usage, context window, peak context, all of that stuff on the DeepSuite site as well.
- 4:26 But yeah, as of July 1st, Fable 5 is retaining the top spot on our leaderboard.
- 4:33 So the ranking information is available online.
- 4:37 I wanted to talk about some of the qualitative insights into how these different models are performing, which I think is the most interesting part.
- 4:46 Starting with the first one is we find Claude is generally a very, very thorough and exhaustive model.
- 4:52 It will try to explore everything, including go through all of the Git logs.
- 4:58 So one interesting insight was seeing that it becomes quite forgetful when it comes to multi-part prompts.
- 5:04 So when you tell it within the scope of a task, let's say, to support both synchronous and async versions of calling a hook, it will go ahead and implement the synchronous part, but it may drop the asynchronous part.
- 5:19 We observed this in roughly two out of three Cloud rollouts across all of the trials, all of the rollouts that we ran.
- 5:25 So this was definitely quite interesting because from my experiences and developers I've talked to as well, Cloud is generally very, very thorough and able to get at the developer's intent quite well.
- 5:38 Another thing about Cloud is it pays very close attention to its environment.
- 5:43 This is taken from the trials we ran ourselves independently and also from examining StreetBench Pro.
loading
Chapters
- 0:00 Introduction: the DeepSWE benchmark
- 1:03 113 original, contamination-resistant tasks
- 2:08 What makes a good benchmark
- 3:51 The leaderboard and model spread
- 5:18 Failure mode: over-scoping the task
- 7:16 Do models verify their own work?
- 8:45 Tasks authored by core contributors
- 10:15 Writing realistic, high level prompts
- 11:45 Program based verifiers and observable behavior
- 13:43 Limitations and future work
- 15:25 Reward hacking and keeping it cheating proof