Videos Rx8f05JI_WA
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI
Scene timeline
40 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 112
- whisperx 112
- chunks
- 22
- from 112 cues
- keyframes
- 17
- kept of 40 captured
- frames with text
- 17
- 451 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 3.5 MB
- word timings on 112 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:09 | 1m 16s |
stt |
done | — | 2026-08-11 05:11 | 11s |
chunk |
done | — | 2026-08-11 05:11 | 0s |
text_embed |
done | — | 2026-08-11 05:11 | 1s |
keyframe |
done | — | 2026-08-11 05:11 | 23s |
ocr |
done | — | 2026-08-11 05:11 | 7s |
frame_embed |
done | — | 2026-08-11 05:11 | 3s |
Frames, and what the machine read
-
- SWE-Marathon1.00
- Can coding agents stay coherent over a 1 billion token budget?1.00
- Rishi Desai0.96
- abundant1.00
- swe-marathon.org1.00
- 1 /130.90
-
- Agents are moving to autonomous end-to-end projects1.00
- AI0.97
- Anthropic1.00
- OpenAl0.95
- anthropic.com/engineering1.00
- github.com/openai/parameter-golf1.00
- "Building a C compiler with a team of parallel0.99
- "Parameter Golf" — train a GPT under a tiny0.98
- Claudes"1.00
- checkpoint budget1.00
- Cloudflare1.00
- Cursor1.00
- blog.cloudflare.com/vinext1.00
- cursor.com/blog/scaling-agents1.00
- "How we rebuilt Next.js with Al in one week"0.97
- "Scaling long-running autonomous coding"0.98
- swe-marathon.org1.00
- 2 / 130.87
-
- From coding tasks to engineering projects0.99
- 20211.00
- Function completion1.00
- ~1K TOKENS0.96
- HumanEval1.00
- ~1 MIN PER TASK0.99
- UNLOCKED CODE LLMS0.99
- 20231.00
- Real GitHub issues1.00
- ~100K T0KENS0.96
- SWE-bench1.00
- ~5-15 MIN PER TASK0.97
- UNLOCKED CODING AGENTS0.98
- 20251.00
- Multi-step terminal tasks1.00
- ~1M TOKENS0.97
- Terminal-Bench1.00
- ~5-15 MIN PER TASK1.00
- UNLOCKED TERMINAL AGENTS1.00
- 20261.00
- Days-long agentic work0.99
- 1B+ T0KENS0.97
- SWE-Marathon1.00
- ~3-10 HOURS PER TASK1.00
- UNLOCKING AUTONOMOUS AGENTS1.00
- swe-marathon.org1.00
- 3 / 130.92
-
- Long horizons require adversarial verification1.00
- Hidden tests0.99
- Reference parity0.99
- CUA checks1.00
- Anti-cheat1.00
- Fresh fixtures + replay0.99
- Match reference1.00
- Use the product1.00
- Catch shortcuts and1.00
- behavior1.00
- exploits1.00
- swe-marathon.org1.00
- 4/ 130.90
-
- First benchmark with full-stack CUA-verified tasks1.00
- • SWE-MARATHON0.93
- CUA verifier grades Slack clone0.99
- A computer-use agent drives the Ul like a real user0.97
- Claude Opus 4.7 · Claude Code0.98
- swe-marathon.org1.00
- 5/ 130.92
-
- First benchmark with full-stack CUA-verified tasks1.00
- CUA VERIFIER0.99
- PASS1.00
- Account created1.00
- Signed in as @ada1.00
- Welcome, Ada Lovelace0.97
- Check 02 / 090.98
- You're not in a workspace yet. Create or join one.1.00
- Sign up & sign in0.97
- Workspace name1.00
- Session / account0.98
- Slack-like layout1.00
- Create workapace0.99
- Post & persist0.97
- Or join with an invitation code0.99
- Create channels1.00
- Hover toolbar1.00
- Emoji picker0.99
- Reaction badges1.00
- Threaded replies1.00
- swe-marathon.org1.00
- 5 / 130.87
-
- First benchmark with full-stack CUA-verified tasks0.99
- Analytical Engine1.00
- #general1.00
- et a topic for this channel0.99
- Settings0.97
- Channels1.00
- # gemeral0.91
- Kicking off the Analytical Engine build0.97
- Ada Lovelace 06:030.93
- Messaging1.00
- CUA VERIFIER0.99
- PASS1.00
- Direct messages0.98
- who's in?0.98
- Ada Lovelace 00 030.88
- First milestone: get auth + channels working0.97
- Posts persist with author + time0.98
- end to end.1.00
- Check 04 / 090.96
- Sign up & sign in1.00
- Session / account0.96
- Slack-like layout1.00
- Post & persist1.00
- Create channels0.98
- Hover toolbar1.00
- Emoji picker1.00
- Reaction badges1.00
- Threaded replies1.00
- Gada0.93
- Log out0.98
- Message #general1.00
- swe-marathon.org1.00
- 5 / 130.87
-
- First benchmark with full-stack CUA-verified tasks0.99
- nalytical Engine1.00
- #general1.00
- CUA VERIFIER0.99
- PASS1.00
- Kicking off the Analytial Engine build who's in?0.96
- Ada Lovelace 003u0.89
- Create a channel1.00
- First milestone: get auth + channels working end to1.00
- Ada Lovelace 0003Pu0.90
- New #engineering channel1.00
- end.1.00
- Check 05 / 090.96
- Sign up & sign in1.00
- Session/account1.00
- Create a channel0.99
- Slack-like layout1.00
- Name (kebab-case) engineering0.99
- Private channel1.00
- Post & persist1.00
- Cancel0.96
- Create channels1.00
- Hover toolbar1.00
- Emoji picker1.00
- Reaction badges1.00
- Threaded replies1.00
- Message #general1.00
- swe-marathon.org1.00
- 5 / 130.93
-
- First benchmark with full-stack CUA-verified tasks1.00
- #general1.00
- build f who's in?0.97
- Kicking off the Analytical Engine1.00
- Ada Lot0.91
- Thread0.97
- Thread1.00
- Ada Lovelace 06:03 Pu0.95
- Threaded reply1.00
- CUA VERIFIER0.99
- PASS1.00
- Ada Lovelace 0:03 Pu0.93
- First milestone: get auth +1.00
- Kicking off the Analytical Etngine build0.99
- who's in?1.00
- C0.77
- Thread panel; reply count updates1.00
- Check 09 / 090.98
- channels working end to end.1.00
- A1.00
- layer.0.97
- I'm in - I'ltake the realtime WebSocket0.95
- Ada Lovelace 06:06 Pu0.90
- Sign up & sign in1.00
- Session / account0.97
- Slack-like layout1.00
- Post & persist1.00
- Create channels1.00
- Hover toolbar1.00
- Emoji picker1.00
- Reaction badges1.00
- Threaded replies1.00
- Message #general1.00
- Reply...1.00
- swe-marathon.org1.00
- 5 / 130.92
Transcript
112 cues· 1,509 words· 8,998 chars
- 0:01 Hi, everyone.
- 0:02 My name is Rishi Desai.
- 0:04 I'm an ML engineer at Abundant AI, where we build reinforcement learning environments for Frontier Labs.
- 0:11 Today, I'm going to talk about SWE Marathon, a benchmark that answers a question that is starting to matter a lot more.
- 0:18 Can coding agents stay coherent over a billion token budget?
- 0:23 Can they build Slack from scratch?
- 0:26 Can they rewrite an entire JAX codebase in PyTorch?
- 0:30 Can they build a C compiler in Rust?
- 0:34 This is what SWE Marathon is trying to measure.
- 0:36 What happens when coding agents move from fixing bugs to owning entire projects end to end?
- 0:47 There's been a tremendous amount of interest in autonomous agent systems.
- 0:52 Anthropic has explored teams of agents building a C compiler.
- 0:57 Cloudflare rebuilt the entire Next.js on byte completely hands off with agents.
- 1:04 And Cursor has experimented with their days long running autonomous agent harness.
- 1:11 The pattern is that coding agents are being pointed at whole projects, not just GitHub issues or linear tickets.
- 1:20 My question is, can we turn some of these frontier lab style case studies into reproducible eval tasks?
- 1:32 Let's talk about the SWE benchmark lineage.
- 1:36 HumanEval asked whether models could write individual Python functions.
- 1:43 SWEBench was a big jump to real GitHub issues where agents had to inspect your repository, make a patch, and patch some unit tests.
- 1:54 TerminalBench pushed this even further by making each task a full environment with a verifier, so agents could use a terminal, run bash commands, inspect files, and leave behind a final container state.
- 2:10 SWE Marathon takes that environment plus verifier framing and stretches the horizon to project scale work.
- 2:19 Multi-hour trajectories and coordinated changes across many, many components.
- 2:23 These are literally hundreds of hours of human work compressed into a single agent rollout.
- 2:34 But once you make tasks this long, a big problem shows up.
- 2:39 Verification.
- 2:42 In a short benchmark, a weak test could just be considered as noise, but
- 2:49 In a multi-hour environment, a weak verifier becomes an attack surface.
- 2:55 The agent has hours, a file system, unrestricted network access potentially, and a reward signal.
- 3:02 So it could spend hours probing the verifier instead of actually doing the intended engineering work.
- 3:09 That's a big reason why Sween Marathon uses multiple independent checks.
- 3:16 We have hidden tests, reference parity checks, computer use agent checks for the product clone tasks, and anti-cheating tests.
- 3:26 We wanted independent verified channels that fail in different ways.
- 3:32 I'll first show you the computer use agent verification example, and then later the failure case where an agent tries to solve the C compiler task by secretly calling GCC.
- 3:49 You might have noticed that there are basically no full-stack product clone tasks in any Long Horizon Suite benchmark out there.
- 3:57 And the reason is verification.
- 4:01 Unit tests can pass, but the product is probably still unusable and the front end looks terrible.
- 4:10 Suite Verizon is the first benchmark to use a computer use agent or CUA verifier for these full-stack tasks.
- 4:18 For the clone Slack task, we have deterministic unit tests to check the API and the backend functionality.
- 4:26 But then a computer use agent uses the browser like a human.
- 4:31 That's what you're seeing in this GIF.
- 4:33 The verifier isn't reading code or calling an API directly.
- 4:38 It's driving the submitted Slack clone through the UI.
- 4:43 So it's logging in, creating channels, posting messages, reacting with emotes, and checking that the app actually works with the rubric.
- 4:54 The big takeaway is that full stack evals are hard because correctness is not just an API contract.
- 5:00 It's whether the user can actually complete the product's intended workflow.
- 5:09 Supreme Marathon has 20 project-scale tasks across four families.
- 5:15 There are library clones, full-stack product clones, ML engineering, and algorithmic tasks.
- 5:23 And some of these tasks even use external APIs.
- 5:25 For example, we have a post-train task where the agent must post-train a language model using the Tinker API.
- 5:34 Expert contributors from the evals community propose the tasks and reference solutions.
loading