read-only demo

Videos Rx8f05JI_WA

SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI

index_state ready data_status ok

AI Engineer· published 2026-07-07· 0:12:57· en-US· indexed 2026-08-11 05:11

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:46, 1 of 1 keyframes kept
  2. Shot 1, 0:46 to 1:29, 1 of 1 keyframes kept
  3. Shot 2, 1:29 to 2:01, 1 of 1 keyframes kept
  4. Shot 3, 2:01 to 2:32, 0 of 1 keyframes kept
  5. Shot 4, 2:32 to 3:09, 1 of 1 keyframes kept
  6. Shot 5, 3:09 to 3:46, 0 of 1 keyframes kept
  7. Shot 6, 3:46 to 3:49, 1 of 1 keyframes kept
  8. Shot 7, 3:49 to 3:53, 1 of 1 keyframes kept
  9. Shot 8, 3:53 to 3:57, 1 of 1 keyframes kept
  10. Shot 9, 3:57 to 3:59, 1 of 1 keyframes kept
  11. Shot 10, 3:59 to 4:05, 0 of 1 keyframes kept
  12. Shot 11, 4:05 to 4:07, 1 of 1 keyframes kept
  13. Shot 12, 4:07 to 4:15, 0 of 1 keyframes kept
  14. Shot 13, 4:15 to 4:19, 0 of 1 keyframes kept
  15. Shot 14, 4:19 to 4:23, 0 of 1 keyframes kept
  16. Shot 15, 4:23 to 4:25, 0 of 1 keyframes kept
  17. Shot 16, 4:25 to 4:31, 0 of 1 keyframes kept
  18. Shot 17, 4:31 to 4:33, 0 of 1 keyframes kept
  19. Shot 18, 4:33 to 4:40, 0 of 1 keyframes kept
  20. Shot 19, 4:40 to 4:44, 0 of 1 keyframes kept
  21. Shot 20, 4:44 to 4:48, 0 of 1 keyframes kept
  22. Shot 21, 4:48 to 4:50, 0 of 1 keyframes kept
  23. Shot 22, 4:50 to 4:56, 0 of 1 keyframes kept
  24. Shot 23, 4:56 to 4:58, 0 of 1 keyframes kept
  25. Shot 24, 4:58 to 5:06, 0 of 1 keyframes kept
  26. Shot 25, 5:06 to 5:39, 1 of 1 keyframes kept
  27. Shot 26, 5:39 to 6:12, 0 of 1 keyframes kept
  28. Shot 27, 6:12 to 6:40, 1 of 1 keyframes kept
  29. Shot 28, 6:40 to 7:07, 0 of 1 keyframes kept
  30. Shot 29, 7:07 to 7:32, 1 of 1 keyframes kept
  31. Shot 30, 7:32 to 7:57, 0 of 1 keyframes kept
  32. Shot 31, 7:57 to 8:32, 1 of 1 keyframes kept
  33. Shot 32, 8:32 to 9:07, 0 of 1 keyframes kept
  34. Shot 33, 9:07 to 9:43, 1 of 1 keyframes kept
  35. Shot 34, 9:43 to 10:19, 0 of 1 keyframes kept
  36. Shot 35, 10:19 to 10:50, 1 of 1 keyframes kept
  37. Shot 36, 10:50 to 11:22, 0 of 1 keyframes kept
  38. Shot 37, 11:22 to 11:50, 1 of 1 keyframes kept
  39. Shot 38, 11:50 to 12:18, 0 of 1 keyframes kept
  40. Shot 39, 12:18 to 12:57, 1 of 1 keyframes kept

40 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
112
whisperx 112
chunks
22
from 112 cues
keyframes
17
kept of 40 captured
frames with text
17
451 lines read
chapters
0
from the source metadata
keyframe bytes
3.5 MB
word timings on 112 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:09 1m 16s
stt done 2026-08-11 05:11 11s
chunk done 2026-08-11 05:11 0s
text_embed done 2026-08-11 05:11 1s
keyframe done 2026-08-11 05:11 23s
ocr done 2026-08-11 05:11 7s
frame_embed done 2026-08-11 05:11 3s

Frames, and what the machine read

  • 0:05 #0 done6 line(s)

    shot 0·sharpness 646.8

    1. SWE-Marathon1.00
    2. Can coding agents stay coherent over a 1 billion token budget?1.00
    3. Rishi Desai0.96
    4. abundant1.00
    5. swe-marathon.org1.00
    6. 1 /130.90
  • 0:59 #1 done18 line(s)

    shot 1·sharpness 1503.7

    1. Agents are moving to autonomous end-to-end projects1.00
    2. AI0.97
    3. Anthropic1.00
    4. OpenAl0.95
    5. anthropic.com/engineering1.00
    6. github.com/openai/parameter-golf1.00
    7. "Building a C compiler with a team of parallel0.99
    8. "Parameter Golf" — train a GPT under a tiny0.98
    9. Claudes"1.00
    10. checkpoint budget1.00
    11. Cloudflare1.00
    12. Cursor1.00
    13. blog.cloudflare.com/vinext1.00
    14. cursor.com/blog/scaling-agents1.00
    15. "How we rebuilt Next.js with Al in one week"0.97
    16. "Scaling long-running autonomous coding"0.98
    17. swe-marathon.org1.00
    18. 2 / 130.87
  • 1:36 #2 done27 line(s)

    shot 2·sharpness 1174.2

    1. From coding tasks to engineering projects0.99
    2. 20211.00
    3. Function completion1.00
    4. ~1K TOKENS0.96
    5. HumanEval1.00
    6. ~1 MIN PER TASK0.99
    7. UNLOCKED CODE LLMS0.99
    8. 20231.00
    9. Real GitHub issues1.00
    10. ~100K T0KENS0.96
    11. SWE-bench1.00
    12. ~5-15 MIN PER TASK0.97
    13. UNLOCKED CODING AGENTS0.98
    14. 20251.00
    15. Multi-step terminal tasks1.00
    16. ~1M TOKENS0.97
    17. Terminal-Bench1.00
    18. ~5-15 MIN PER TASK1.00
    19. UNLOCKED TERMINAL AGENTS1.00
    20. 20261.00
    21. Days-long agentic work0.99
    22. 1B+ T0KENS0.97
    23. SWE-Marathon1.00
    24. ~3-10 HOURS PER TASK1.00
    25. UNLOCKING AUTONOMOUS AGENTS1.00
    26. swe-marathon.org1.00
    27. 3 / 130.92
  • 2:11 #3 skipped

    shot 3·duplicate of #2

  • 2:40 #4 done13 line(s)

    shot 4·sharpness 847.1

    1. Long horizons require adversarial verification1.00
    2. Hidden tests0.99
    3. Reference parity0.99
    4. CUA checks1.00
    5. Anti-cheat1.00
    6. Fresh fixtures + replay0.99
    7. Match reference1.00
    8. Use the product1.00
    9. Catch shortcuts and1.00
    10. behavior1.00
    11. exploits1.00
    12. swe-marathon.org1.00
    13. 4/ 130.90
  • 3:14 #5 skipped

    shot 5·duplicate of #4

  • 3:49 #6 done7 line(s)

    shot 6·sharpness 646.5

    1. First benchmark with full-stack CUA-verified tasks1.00
    2. • SWE-MARATHON0.93
    3. CUA verifier grades Slack clone0.99
    4. A computer-use agent drives the Ul like a real user0.97
    5. Claude Opus 4.7 · Claude Code0.98
    6. swe-marathon.org1.00
    7. 5/ 130.92
  • 3:51 #7 done22 line(s)

    shot 7·sharpness 798.6

    1. First benchmark with full-stack CUA-verified tasks1.00
    2. CUA VERIFIER0.99
    3. PASS1.00
    4. Account created1.00
    5. Signed in as @ada1.00
    6. Welcome, Ada Lovelace0.97
    7. Check 02 / 090.98
    8. You're not in a workspace yet. Create or join one.1.00
    9. Sign up & sign in0.97
    10. Workspace name1.00
    11. Session / account0.98
    12. Slack-like layout1.00
    13. Create workapace0.99
    14. Post & persist0.97
    15. Or join with an invitation code0.99
    16. Create channels1.00
    17. Hover toolbar1.00
    18. Emoji picker0.99
    19. Reaction badges1.00
    20. Threaded replies1.00
    21. swe-marathon.org1.00
    22. 5 / 130.87
  • 3:56 #8 done33 line(s)

    shot 8·sharpness 792.4

    1. First benchmark with full-stack CUA-verified tasks0.99
    2. Analytical Engine1.00
    3. #general1.00
    4. et a topic for this channel0.99
    5. Settings0.97
    6. Channels1.00
    7. # gemeral0.91
    8. Kicking off the Analytical Engine build0.97
    9. Ada Lovelace 06:030.93
    10. Messaging1.00
    11. CUA VERIFIER0.99
    12. PASS1.00
    13. Direct messages0.98
    14. who's in?0.98
    15. Ada Lovelace 00 030.88
    16. First milestone: get auth + channels working0.97
    17. Posts persist with author + time0.98
    18. end to end.1.00
    19. Check 04 / 090.96
    20. Sign up & sign in1.00
    21. Session / account0.96
    22. Slack-like layout1.00
    23. Post & persist1.00
    24. Create channels0.98
    25. Hover toolbar1.00
    26. Emoji picker1.00
    27. Reaction badges1.00
    28. Threaded replies1.00
    29. Gada0.93
    30. Log out0.98
    31. Message #general1.00
    32. swe-marathon.org1.00
    33. 5 / 130.87
  • 3:59 #9 done29 line(s)

    shot 9·sharpness 803.7

    1. First benchmark with full-stack CUA-verified tasks0.99
    2. nalytical Engine1.00
    3. #general1.00
    4. CUA VERIFIER0.99
    5. PASS1.00
    6. Kicking off the Analytial Engine build who's in?0.96
    7. Ada Lovelace 003u0.89
    8. Create a channel1.00
    9. First milestone: get auth + channels working end to1.00
    10. Ada Lovelace 0003Pu0.90
    11. New #engineering channel1.00
    12. end.1.00
    13. Check 05 / 090.96
    14. Sign up & sign in1.00
    15. Session/account1.00
    16. Create a channel0.99
    17. Slack-like layout1.00
    18. Name (kebab-case) engineering0.99
    19. Private channel1.00
    20. Post & persist1.00
    21. Cancel0.96
    22. Create channels1.00
    23. Hover toolbar1.00
    24. Emoji picker1.00
    25. Reaction badges1.00
    26. Threaded replies1.00
    27. Message #general1.00
    28. swe-marathon.org1.00
    29. 5 / 130.93
  • 4:04 #10 skipped

    shot 10·duplicate of #8

  • 4:07 #11 done36 line(s)

    shot 11·sharpness 949.7

    1. First benchmark with full-stack CUA-verified tasks1.00
    2. #general1.00
    3. build f who's in?0.97
    4. Kicking off the Analytical Engine1.00
    5. Ada Lot0.91
    6. Thread0.97
    7. Thread1.00
    8. Ada Lovelace 06:03 Pu0.95
    9. Threaded reply1.00
    10. CUA VERIFIER0.99
    11. PASS1.00
    12. Ada Lovelace 0:03 Pu0.93
    13. First milestone: get auth +1.00
    14. Kicking off the Analytical Etngine build0.99
    15. who's in?1.00
    16. C0.77
    17. Thread panel; reply count updates1.00
    18. Check 09 / 090.98
    19. channels working end to end.1.00
    20. A1.00
    21. layer.0.97
    22. I'm in - I'ltake the realtime WebSocket0.95
    23. Ada Lovelace 06:06 Pu0.90
    24. Sign up & sign in1.00
    25. Session / account0.97
    26. Slack-like layout1.00
    27. Post & persist1.00
    28. Create channels1.00
    29. Hover toolbar1.00
    30. Emoji picker1.00
    31. Reaction badges1.00
    32. Threaded replies1.00
    33. Message #general1.00
    34. Reply...1.00
    35. swe-marathon.org1.00
    36. 5 / 130.92
  • 4:11 #12 skipped

    shot 12·duplicate of #6

  • 4:17 #13 skipped

    shot 13·duplicate of #7

  • 4:21 #14 skipped

    shot 14·duplicate of #8

  • 4:23 #15 skipped

    shot 15·duplicate of #9

  • 4:29 #16 skipped

    shot 16·duplicate of #8

  • 4:31 #17 skipped

    shot 17·duplicate of #11

  • 4:39 #18 skipped

    shot 18·duplicate of #6

  • 4:44 #19 skipped

    shot 19·duplicate of #7

  • 4:47 #20 skipped

    shot 20·duplicate of #8

  • 4:49 #21 skipped

    shot 21·duplicate of #9

  • 4:55 #22 skipped

    shot 22·duplicate of #8

  • 4:57 #23 skipped

    shot 23·duplicate of #11

Transcript

112 cues· 1,509 words· 8,998 chars

  1. 0:01 Hi, everyone.
  2. 0:02 My name is Rishi Desai.
  3. 0:04 I'm an ML engineer at Abundant AI, where we build reinforcement learning environments for Frontier Labs.
  4. 0:11 Today, I'm going to talk about SWE Marathon, a benchmark that answers a question that is starting to matter a lot more.
  5. 0:18 Can coding agents stay coherent over a billion token budget?
  6. 0:23 Can they build Slack from scratch?
  7. 0:26 Can they rewrite an entire JAX codebase in PyTorch?
  8. 0:30 Can they build a C compiler in Rust?
  9. 0:34 This is what SWE Marathon is trying to measure.
  10. 0:36 What happens when coding agents move from fixing bugs to owning entire projects end to end?
  11. 0:47 There's been a tremendous amount of interest in autonomous agent systems.
  12. 0:52 Anthropic has explored teams of agents building a C compiler.
  13. 0:57 Cloudflare rebuilt the entire Next.js on byte completely hands off with agents.
  14. 1:04 And Cursor has experimented with their days long running autonomous agent harness.
  15. 1:11 The pattern is that coding agents are being pointed at whole projects, not just GitHub issues or linear tickets.
  16. 1:20 My question is, can we turn some of these frontier lab style case studies into reproducible eval tasks?
  17. 1:32 Let's talk about the SWE benchmark lineage.
  18. 1:36 HumanEval asked whether models could write individual Python functions.
  19. 1:43 SWEBench was a big jump to real GitHub issues where agents had to inspect your repository, make a patch, and patch some unit tests.
  20. 1:54 TerminalBench pushed this even further by making each task a full environment with a verifier, so agents could use a terminal, run bash commands, inspect files, and leave behind a final container state.
  21. 2:10 SWE Marathon takes that environment plus verifier framing and stretches the horizon to project scale work.
  22. 2:19 Multi-hour trajectories and coordinated changes across many, many components.
  23. 2:23 These are literally hundreds of hours of human work compressed into a single agent rollout.
  24. 2:34 But once you make tasks this long, a big problem shows up.
  25. 2:39 Verification.
  26. 2:42 In a short benchmark, a weak test could just be considered as noise, but
  27. 2:49 In a multi-hour environment, a weak verifier becomes an attack surface.
  28. 2:55 The agent has hours, a file system, unrestricted network access potentially, and a reward signal.
  29. 3:02 So it could spend hours probing the verifier instead of actually doing the intended engineering work.
  30. 3:09 That's a big reason why Sween Marathon uses multiple independent checks.
  31. 3:16 We have hidden tests, reference parity checks, computer use agent checks for the product clone tasks, and anti-cheating tests.
  32. 3:26 We wanted independent verified channels that fail in different ways.
  33. 3:32 I'll first show you the computer use agent verification example, and then later the failure case where an agent tries to solve the C compiler task by secretly calling GCC.
  34. 3:49 You might have noticed that there are basically no full-stack product clone tasks in any Long Horizon Suite benchmark out there.
  35. 3:57 And the reason is verification.
  36. 4:01 Unit tests can pass, but the product is probably still unusable and the front end looks terrible.
  37. 4:10 Suite Verizon is the first benchmark to use a computer use agent or CUA verifier for these full-stack tasks.
  38. 4:18 For the clone Slack task, we have deterministic unit tests to check the API and the backend functionality.
  39. 4:26 But then a computer use agent uses the browser like a human.
  40. 4:31 That's what you're seeing in this GIF.
  41. 4:33 The verifier isn't reading code or calling an API directly.
  42. 4:38 It's driving the submitted Slack clone through the UI.
  43. 4:43 So it's logging in, creating channels, posting messages, reacting with emotes, and checking that the app actually works with the rubric.
  44. 4:54 The big takeaway is that full stack evals are hard because correctness is not just an API contract.
  45. 5:00 It's whether the user can actually complete the product's intended workflow.
  46. 5:09 Supreme Marathon has 20 project-scale tasks across four families.
  47. 5:15 There are library clones, full-stack product clones, ML engineering, and algorithmic tasks.
  48. 5:23 And some of these tasks even use external APIs.
  49. 5:25 For example, we have a post-train task where the agent must post-train a language model using the Tinker API.
  50. 5:34 Expert contributors from the evals community propose the tasks and reference solutions.

Open at this second