Videos 9QebvrrY3KY
Claude for Long-Horizon Tasks — Lance Martin, Anthropic
Scene timeline
61 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 282
- whisperx 282
- chunks
- 44
- from 282 cues
- keyframes
- 16
- kept of 61 captured
- frames with text
- 15
- 239 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.3 MB
- word timings on 282 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:50 | 0s |
stt |
done | — | 2026-08-09 20:44 | 27s |
chunk |
done | — | 2026-08-09 20:45 | 0s |
text_embed |
done | — | 2026-08-10 19:43 | 0s |
keyframe |
done | — | 2026-08-09 20:45 | 3m 13s |
ocr |
done | — | 2026-08-09 20:48 | 5s |
frame_embed |
done | — | 2026-08-10 19:43 | 3s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AIEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.97
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- ANTHROP\C1.00
- World'sFair1.00
- Towards asynchronous agents1.00
- PRESENTEDBY1.00
- Microsoft1.00
- AlEnginoor0.93
- World'sFair1.00
- Lance Martin, Member of Technical Staff, Anthropic0.99
- @rlancemartin1.00
- World'sFair0.97
- Engineering the future of Al1.00
-
- AlEngineer0.98
- ANTHROP\C1.00
- World'sFair0.98
- Towards asynchronous agents1.00
- PRESENTEDBY1.00
- Microsoft1.00
- Lance Martin, Member of Technical Staff, Anthropic0.99
- @rlancemartin1.00
- Engineering the future of Al1.00
-
- AlEngineer0.99
- World'sFair1.00
- PRESENTED BY0.96
- Microsoft1.00
- ANTHROP\C1.00
- World's Fair0.94
- TRACK 1· JUNE 30, 20260.97
- Claws & Personal Agents0.99
-
- AlEngineer0.99
- Task duration (time for a human)0.99
- World'sFair1.00
- Autocomplete/Chat1.00
- Synchronous local agents1.00
- Asynchronous remote agents1.00
- 1 day0.97
- 12 hr0.99
- Claude Opus 4.61.00
- 12 hr0.99
- Claude Opus 4.51.00
- 4 hr0.88
- 4.9 hr0.93
- PRESENTEDBY1.00
- Claude Opus 41.00
- 1.7 hr0.99
- Microsoft1.00
- 1 hr0.97
- Claude Sonnet 3.71.00
- Claude Sonnet 3.51.00
- 21 min0.99
- 10 min1.00
- Claude 3 Opus1.00
- 4 min0.99
- 1 min0.98
- 20241.00
- 20251.00
- 20261.00
- ANTHROP\C0.99
- World's Fair0.90
- TRACK 1· JUNE 30, 20260.97
- Claws & Personal Agents0.99
-
- AlEngineer0.99
- World'sFair1.00
- Session1.00
- PRESENTEDBY1.00
- Microsoft1.00
- ANTHROP\C1.00
- Anthropic Engineering Blog1.00
- TRACK 1· JUNE 30, 20260.97
- Claws & Personal Agents0.99
-
- AlEngineer0.99
- World's Fair0.96
- PRESENTEDBY1.00
- Microsoft1.00
- ANTHROP\C1.00
- World's Fair0.93
- TRACK 1· JUNE 30, 20260.97
- Claws & Personal Agents0.99
Transcript
282 cues· 4,225 words· 23,361 chars
- 0:12 Good to go.
- 0:14 All right.
- 0:14 Well, take a quick sip and then let's start.
- 0:19 It is great to be here.
- 0:21 This is like my third year coming to this conference and I always really enjoy it.
- 0:25 And thank you for coming to this workshop.
- 0:26 I know there's many interesting talks.
- 0:29 Let me talk a little bit about our view of asynchronous agents at Anthropic and some things we've been up to lately.
- 0:36 So this is kind of a way I think about models and product.
- 0:39 So you can think about Cloud as a light source, and you can think about products as windows that allow the light to pass through.
- 0:46 And what's kind of interesting is over time, the window that you need to actually kind of see the light of the model kind of shifts.
- 0:54 And we've seen this over the past few years.
- 0:57 So I'm plotting here different Cloud models and their task horizon.
- 1:01 So how much autonomous work can they do over time?
- 1:05 And you might recall back in the Opus 3 days, this was kind of like 2024, models could only do maybe 10 to 20 minutes of autonomous work.
- 1:13 This is measured by meter.
- 1:15 And in that regime, only certain product surfaces made sense.
- 1:18 Things like autocomplete, things like chat, where your human is very in the loop.
- 1:21 Because the model's really only doing a very short amount of work before you're steering it.
- 1:26 Now, the past year, we saw the rise of synchronous coding agents like Cloud Code, and this is kind of a shift because then models could do maybe an hour of work.
- 1:35 So it made sense to have them run, but typically locally, where you could still steer them easily.
- 1:41 And it's kind of interesting because during this regime, I remember efforts, and I was involved in some efforts to build kind of asynchronous agents,
- 1:48 But when models can only do like an hour of work, async as an experience is kind of bad.
- 1:54 The model goes off and it like hits an error and it comes back to you over a short period of time.
- 1:59 In order to really unlock async, we needed longer task horizons.
- 2:03 And so we're starting to see that now.
- 2:06 And kind of with this shift in capability and time horizon came a shift in the API surfaces.
- 2:12 So if you look at the lower left,
- 2:15 Message API came out like two years ago.
- 2:17 It's basically prompt response.
- 2:19 It's great for building harnesses.
- 2:22 But it's a very simple API.
- 2:23 Again, you're just passing instead of messages, you get a response out.
- 2:27 There's no sense of deployment with that.
- 2:29 So you basically take messages API and you can roll your own harness.
- 2:32 You can deploy that harness and you have an agent.
- 2:34 Now, over the past year, we saw the rise of coding agents in particular.
- 2:39 So we released Asian SDK.
- 2:40 So that's basically a way to programmatically call cloud code.
- 2:43 And that's basically us giving you a harness.
- 2:47 But over the past few months since April, as we've seen longer and longer task horizons, we released a new API called Managed Agents, which basically packages both the harness as well as all the managed deployment infrastructure for you.
- 3:01 And I'll talk about some of the themes that underpin this new surface, Cloud Managed Agents, and some of the themes that
- 3:07 kind of extend beyond just managed agents broadly to think about this kind of new type of asynchronous agents, which can apply, of course, to clause and other types of kind of longer running long horizon agents.
- 3:19 So theme one is decoupling the brain from the hands.
- 3:24 So when we first set out to build managed agents, we started with the container.
- 3:28 We put the harness in the sandbox in the same container.
- 3:32 Now, the problem here is what happens if the harness dies or the container dies?
- 3:39 What we saw is we actually lose the session.
- 3:42 So, basically, this architecture is kind of tricky for long horizon agents because what can happen is your agent is running and if that container dies, you lose everything with it.
- 3:54 Also, as models get more capable, putting the credentials in the same container with the agent itself can be problematic.
loading