read-only demo

Videos Yk87oUPVaxU

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

index_state ready data_status ok

AI Engineer· published 2026-07-26· 0:17:34· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 0:57, 1 of 1 keyframes kept
  6. Shot 5, 0:57 to 1:32, 1 of 1 keyframes kept
  7. Shot 6, 1:32 to 2:07, 0 of 1 keyframes kept
  8. Shot 7, 2:07 to 2:39, 1 of 1 keyframes kept
  9. Shot 8, 2:39 to 3:12, 0 of 1 keyframes kept
  10. Shot 9, 3:12 to 3:46, 0 of 1 keyframes kept
  11. Shot 10, 3:46 to 4:32, 1 of 1 keyframes kept
  12. Shot 11, 4:32 to 5:04, 0 of 1 keyframes kept
  13. Shot 12, 5:04 to 5:36, 1 of 1 keyframes kept
  14. Shot 13, 5:36 to 6:02, 1 of 1 keyframes kept
  15. Shot 14, 6:02 to 6:27, 0 of 1 keyframes kept
  16. Shot 15, 6:27 to 7:11, 0 of 1 keyframes kept
  17. Shot 16, 7:11 to 7:47, 0 of 1 keyframes kept
  18. Shot 17, 7:47 to 8:23, 0 of 1 keyframes kept
  19. Shot 18, 8:23 to 8:55, 0 of 1 keyframes kept
  20. Shot 19, 8:55 to 9:26, 0 of 1 keyframes kept
  21. Shot 20, 9:26 to 9:58, 0 of 1 keyframes kept
  22. Shot 21, 9:58 to 10:35, 0 of 1 keyframes kept
  23. Shot 22, 10:35 to 11:13, 0 of 1 keyframes kept
  24. Shot 23, 11:13 to 11:44, 1 of 1 keyframes kept
  25. Shot 24, 11:44 to 12:12, 0 of 1 keyframes kept
  26. Shot 25, 12:12 to 12:40, 0 of 1 keyframes kept
  27. Shot 26, 12:40 to 13:07, 0 of 1 keyframes kept
  28. Shot 27, 13:07 to 13:39, 0 of 1 keyframes kept
  29. Shot 28, 13:39 to 14:12, 0 of 1 keyframes kept
  30. Shot 29, 14:12 to 14:44, 0 of 1 keyframes kept
  31. Shot 30, 14:44 to 15:16, 0 of 1 keyframes kept
  32. Shot 31, 15:16 to 15:49, 0 of 1 keyframes kept
  33. Shot 32, 15:49 to 16:21, 0 of 1 keyframes kept
  34. Shot 33, 16:21 to 16:54, 0 of 1 keyframes kept
  35. Shot 34, 16:54 to 17:17, 1 of 1 keyframes kept
  36. Shot 35, 17:17 to 17:33, 0 of 1 keyframes kept

36 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
128
whisperx 128
chunks
30
from 128 cues
keyframes
12
kept of 36 captured
frames with text
12
207 lines read
chapters
11
from the source metadata
keyframe bytes
4.5 MB
word timings on 128 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:27 0s
stt done 2026-08-09 06:22 18s
chunk done 2026-08-09 06:22 0s
text_embed done 2026-08-10 19:39 1s
keyframe done 2026-08-09 06:22 2m 07s
ocr done 2026-08-09 06:24 5s
frame_embed done 2026-08-10 19:39 2s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 452.4

    1. AIEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 659.3

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2729.1

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:27 #3 done2 line(s)

    shot 3·sharpness 184.8

    1. AlEngineer0.98
    2. World's Fair0.99
  • 0:51 #4 done13 line(s)

    shot 4·sharpness 1603.9

    1. AlEngineer0.99
    2. Datacurve1.00
    3. World'sFair1.00
    4. JULY 20260.99
    5. DeepSWE1.00
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. Measuring frontier coding agents on original, long-horizon1.00
    9. software engineering tasks.0.99
    10. James1.00
    11. Engineer @ Datacurve0.98
    12. Engineering the future of Al0.98
    13. World's Fair0.96
  • 1:21 #5 done18 line(s)

    shot 5·sharpness 2224.4

    1. AlEngineer0.94
    2. World'sFair1.00
    3. Overview1.00
    4. DeepSWE is 113 original, long-horizon tasks0.99
    5. 1131.00
    6. 910.99
    7. 51.00
    8. PRESENTED BY1.00
    9. Microsoft1.00
    10. authored tasks0.98
    11. active repositories1.00
    12. languages1.00
    13. Artificial Analysis1.00
    14. Replaced SWE-Bench Pro with DeepSWE in the Coding1.00
    15. Agent Index1.00
    16. TRACK 8• JULY 2, 20260.96
    17. World's Fair0.99
    18. Agentic Engineering1.00
  • 2:03 #6 skipped

    shot 6·duplicate of #5

  • 2:20 #7 done14 line(s)

    shot 7·sharpness 2848.2

    1. AlEngineer0.98
    2. Datacurve1.00
    3. World'sFair1.00
    4. Datacurve builds training data for high-1.00
    5. ceiling domains0.98
    6. 01 Research and training data — coding, and adjacent high-complexity technical domains.1.00
    7. 021.00
    8. In-house platforms — built to attract and collect the highest caliber of data.0.99
    9. 031.00
    10. Research on post-training — what is data quality? How does this demonstrably move0.99
    11. models?1.00
    12. TRACK 8· JULY 2,20260.96
    13. World'sFair1.00
    14. Agentic Engineering0.99
  • 2:55 #8 skipped

    shot 8·duplicate of #7

  • 3:35 #9 skipped

    shot 9·duplicate of #7

  • 4:27 #10 done53 line(s)

    shot 10·sharpness 2302.8

    1. AlEngineer1.00
    2. Results1.00
    3. v1.1 · pass@1 · mini-swe-agent · updated July 1, 20260.97
    4. World'sFair1.00
    5. Frontier models diverge on DeepSWE0.99
    6. #1.00
    7. MODEL1.00
    8. 40.99
    9. PASS@11.00
    10. COST1.00
    11. 011.00
    12. claude-fable-5 max0.99
    13. 70% ±40.94
    14. $21.631.00
    15. 020.98
    16. gpt-5.5 xhigh1.00
    17. 67% ±60.94
    18. $7.231.00
    19. 030.93
    20. claude-opus-4.8 max1.00
    21. 59% ±20.97
    22. $13.221.00
    23. 040.91
    24. claude-sonnet-5 max1.00
    25. 54% ±40.97
    26. $26.401.00
    27. 050.89
    28. gpt-5.4 xhigh0.99
    29. 52% ±20.94
    30. $5.651.00
    31. 060.86
    32. glm-5.2 max0.99
    33. 44% ±20.97
    34. $3.921.00
    35. 070.97
    36. gemini-3.5-flash medium0.99
    37. 37% ±21.00
    38. $7.341.00
    39. 080.84
    40. kimi-k2.7-code-1.00
    41. 31% ±10.98
    42. $2.821.00
    43. 890.83
    44. claude-sonnet-4.6 high0.99
    45. 30% ±40.98
    46. $5.521.00
    47. 100.99
    48. gemini-3.1-pro high1.00
    49. 12% ±20.97
    50. $9.481.00
    51. TRACK 8· JULY 2, 20260.95
    52. World's Fair1.00
    53. Agentic Engineering1.00
  • 4:39 #11 skipped

    shot 11·duplicate of #7

  • 5:33 #12 done21 line(s)

    shot 12·sharpness 2762.9

    1. AlEngineer0.99
    2. Findings 1/41.00
    3. World'sFair1.00
    4. Claude is forgetful with multi-part prompts0.98
    5. Prompts enumerate parallel behaviors — support both sync and async — and Claude ships the0.98
    6. obvious branch.1.00
    7. PRESENTED BY0.99
    8. python-statemachine-state-data-scoping1.00
    9. langchain-request-coalescing1.00
    10. Microsoft1.00
    11. sync hook lands correctly in0.97
    12. invoke coalesces; batch dispatches via executor.map0.99
    13. BaseEngine._enter_states1.00
    14. x batch(["hello", "hello", "world"]] → 3 executions, not 20.95
    15. x AsyncEngine never receives the same hook0.98
    16. ≈2/31.00
    17. of Claude rollouts tagged MISSED_REQUIREMENT fit this one-branch-shipped0.99
    18. pattern.1.00
    19. TRACK 8· JULY 2, 20260.95
    20. World'sFair1.00
    21. Agentic Engineering1.00
  • 5:52 #13 done21 line(s)

    shot 13·sharpness 2561.9

    1. AlEngineer0.97
    2. Findings 2/41.00
    3. World's Fair0.97
    4. Claude pays close attention to its1.00
    5. environment1.00
    6. When prompt and repo disagree, Opus runs git log — and recovers the gold patch from .git history.0.99
    7. PRESENTED BY1.00
    8. Microsoft1.00
    9. 25%0.99
    10. 18%1.00
    11. ≈1%1.00
    12. 0%0.96
    13. claude-opus-4.61.00
    14. claude-opus-4.71.00
    15. gemini configs1.00
    16. gpt-5.4 · gpt-5.50.98
    17. share of reviewed SWE-Bench Pro passes flagged CHEATED. 87% of flags = gold commit read from .git0.99
    18. The gold commit ships inside SWE-Bench Pro containers. DeepSWE v1.1 deletes future refs entirely.0.99
    19. TRACK 8· JULY 2, 20260.95
    20. World's Fair1.00
    21. Agentic Engineering1.00
  • 6:10 #14 skipped

    shot 14·duplicate of #13

  • 7:06 #15 skipped

    shot 15·duplicate of #7

  • 7:40 #16 skipped

    shot 16·duplicate of #7

  • 8:09 #17 skipped

    shot 17·duplicate of #7

  • 8:51 #18 skipped

    shot 18·duplicate of #7

  • 9:23 #19 skipped

    shot 19·duplicate of #7

  • 9:54 #20 skipped

    shot 20·duplicate of #7

  • 10:31 #21 skipped

    shot 21·duplicate of #12

  • 11:08 #22 skipped

    shot 22·duplicate of #7

  • 11:28 #23 done24 line(s)

    shot 23·sharpness 2109.6

    1. AlEngineer0.98
    2. Task shape1.00
    3. World's Fair0.99
    4. Prompts are short; solutions are long1.00
    5. PROMPT LENGTH1.00
    6. SOLUTION SIZE1.00
    7. 2,1581.00
    8. 6681.00
    9. PRESENTED BY1.00
    10. chars1.00
    11. lines1.00
    12. Microsoft1.00
    13. vs 4,614 on SWE-Bench Pro0.99
    14. vs 120 on SWE-Bench Pro0.98
    15. FILES TOUCHED1.00
    16. OUTPUT TOKENS1.00
    17. 71.00
    18. ≈2×0.99
    19. files1.00
    20. vs 5 per reference solution1.00
    21. vs SWE-Bench Pro runs1.00
    22. TRACK 8· JULY 2, 20260.95
    23. World's Fair0.99
    24. Agentic Engineering1.00

Transcript

128 cues· 2,636 words· 15,254 chars

  1. 0:12 Hey, everyone.
  2. 0:13 Can you guys hear me okay?
  3. 0:14 This is good.
  4. 0:16 Yeah, my name is James.
  5. 0:17 I'm one of the founding engineers at DataCurve.
  6. 0:20 Unfortunately, Serena's been out with a fever for the past couple of days.
  7. 0:24 She was supposed to be here giving this talk, so I'm just filling in in her place.
  8. 0:28 But I've been at DataCurve working on the research and engineering side of things, as well as DeepSuite, which is our frontier long-horizon coding benchmark, which you guys may be familiar.
  9. 0:41 I'll just be going over some of the most important findings about DeepSuite, a brief overview of what it is for those of you who may not know, and then going deeper into our methodology and exactly how we came about this frontier coding benchmark.
  10. 0:56 So DeepSuite is a long horizon software engineering benchmark comprised of 113 original software engineering tasks.
  11. 1:07 So this means unlike something like SuiteBench Pro, we didn't scrape this from existing PRs that have been closed.
  12. 1:13 There's a variety of benefits for this, namely one of them is to resist against contamination and agents being able to cheat through the course of their rollouts.
  13. 1:23 SweetBench Pro pulls thousands of tasks from only 40 repositories.
  14. 1:29 The median task per repository for us is one, so you can see across over 100 tasks, we pull from nearly 100 repositories.
  15. 1:38 And the language spans across TypeScript, JavaScript, Python, Rust, and Go.
  16. 1:42 And we have plans to add more languages later on.
  17. 1:47 Since its release, we've received very positive reception.
  18. 1:51 It's replaced SweetBench Pro in the artificial analysis coding agent index, as well as being cited by numerous frontier model labs and us helping with them in tracking their models on our benchmark as well.
  19. 2:05 So we've been really, really appreciative of that.
  20. 2:08 A bit of context about us, DataCurve works on building training data for high ceiling domains, including coding, as well as coding adjacent fields.
  21. 2:18 We also are trying to answer the very elusive question of what exactly makes good data, what is data quality, and how can we demonstrate that our training data, in fact, moves the needle.
  22. 2:32 So DeepSuite is one in a long line of initiatives that we have towards answering this question.
  23. 2:40 So why did we create DeepSuite?
  24. 2:41 Well, it's very clear that the existing benchmarks are not hitting the mark.
  25. 2:47 With benches like SuiteBench Pro, top models are clustering at the top.
  26. 2:51 It's very hard to differentiate between which one is good because they all have overlapping confidence intervals.
  27. 2:58 Contamination is also rampant because again, all of these tasks are mined from public PRs.
  28. 3:03 So all the solution tests, even the discussion around the PRs, those are all available out in the wild for these agents to access.
  29. 3:10 The verifiers are also very, very brittle because we're anchoring them to a specific implementation often derived from the PR that was merged in.
  30. 3:19 And oftentimes you also have tests that check for private helpers and functions created by the task author, which is very opinionated, right?
  31. 3:27 And it's not something that model should have to adhere to.
  32. 3:31 And finally, leakage.
  33. 3:32 So one thing about SweetBench Pro is for very insightful models such as Claude, they're able to directly run git log and then go through the commit hashes and cherry pick the ones out that contain the golden patches, which again, very, very serious issue.
  34. 3:47 So this is DeepSuite.
  35. 3:48 This is the updated leaderboard as of July 1st.
  36. 3:51 You can see, I was mentioning before the problem of differentiating, but you can see on DeepSuite here, there's a very clear difference.
  37. 3:59 There's a very clear performance gap between the top performing models versus, you know, at 10th place, you have Gemini 3.1 Pro.
  38. 4:07 Also, within the cloud and the GPT models as well, we're able to see some deviance.
  39. 4:12 And yeah, if you go on deepsuite.datacurve.ai, you'll also be able to see the token efficiency costs, token usage, context window, peak context, all of that stuff on the DeepSuite site as well.
  40. 4:26 But yeah, as of July 1st, Fable 5 is retaining the top spot on our leaderboard.
  41. 4:33 So the ranking information is available online.
  42. 4:37 I wanted to talk about some of the qualitative insights into how these different models are performing, which I think is the most interesting part.
  43. 4:46 Starting with the first one is we find Claude is generally a very, very thorough and exhaustive model.
  44. 4:52 It will try to explore everything, including go through all of the Git logs.
  45. 4:58 So one interesting insight was seeing that it becomes quite forgetful when it comes to multi-part prompts.
  46. 5:04 So when you tell it within the scope of a task, let's say, to support both synchronous and async versions of calling a hook, it will go ahead and implement the synchronous part, but it may drop the asynchronous part.
  47. 5:19 We observed this in roughly two out of three Cloud rollouts across all of the trials, all of the rollouts that we ran.
  48. 5:25 So this was definitely quite interesting because from my experiences and developers I've talked to as well, Cloud is generally very, very thorough and able to get at the developer's intent quite well.
  49. 5:38 Another thing about Cloud is it pays very close attention to its environment.
  50. 5:43 This is taken from the trials we ran ourselves independently and also from examining StreetBench Pro.

Chapters

  1. 0:00 Introduction: the DeepSWE benchmark
  2. 1:03 113 original, contamination-resistant tasks
  3. 2:08 What makes a good benchmark
  4. 3:51 The leaderboard and model spread
  5. 5:18 Failure mode: over-scoping the task
  6. 7:16 Do models verify their own work?
  7. 8:45 Tasks authored by core contributors
  8. 10:15 Writing realistic, high level prompts
  9. 11:45 Program based verifiers and observable behavior
  10. 13:43 Limitations and future work
  11. 15:25 Reward hacking and keeping it cheating proof

Open at this second