Videos 2aS7aKoXn64
Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software
Scene timeline
51 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 225
- whisperx 225
- chunks
- 38
- from 225 cues
- keyframes
- 23
- kept of 51 captured
- frames with text
- 23
- 411 lines read
- chapters
- 10
- from the source metadata
- keyframe bytes
- 6.8 MB
- word timings on 225 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:32 | 0s |
stt |
done | — | 2026-08-09 00:57 | 25s |
chunk |
done | — | 2026-08-09 00:57 | 0s |
text_embed |
done | — | 2026-08-10 19:36 | 1s |
keyframe |
done | — | 2026-08-09 00:58 | 2m 42s |
ocr |
done | — | 2026-08-09 01:00 | 11s |
frame_embed |
done | — | 2026-08-10 19:36 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.96
- World'sFair1.00
- DATA QUALITY · AI Engineer World's Fair 20260.98
- PRESENTED BY0.99
- Rethinking Environments for0.99
- Microsoft1.00
- Long Horizon Tasks1.00
- θ Theta0.94
- World'sFair1.00
- Engineering the future of Al0.99
-
- AlEngineer0.98
- World's Fair0.99
- Task Horizons Are Accelerating0.98
- Model1.00
- Date Released1.00
- Time Horizon0.96
- PRESENTED BY1.00
- Microsoft1.00
- Claude 3.7 Sonnet1.00
- Feb 20251.00
- 1 hr0.88
- Claude Opus 41.00
- May 20251.00
- 1.7 hrs0.99
- Claude Opus 4.50.99
- Nov 20251.00
- 4.9 hrs1.00
- Claude Opus 4.61.00
- Feb 20261.00
- 12 hrs1.00
- Source: METR, metr.org/time-horizons/1.00
- World's Fair0.99
- Engineering the future of Al1.00
-
- AlEngineer0.98
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- How do we define “long horizon"?0.97
- World's Fair0.96
- TRACK 9· JUNE 30, 20260.96
- Data Quality0.96
-
- AlEngineer0.97
- World'sFair1.00
- 1. Human Horizon0.99
- Predicted 50% time horizon: 12 hr for Claude Opus 4.60.99
- Measurements above 16 hrs are unreliable with our current task suite0.99
- 100%1.00
- Est ss te0.67
- 80%-1.00
- 60%-0.89
- 40%-1.00
- 20%-1.00
- 0%-0.99
- 1s1.00
- 4s1.00
- 15s1.00
- 1m1.00
- 4m1.00
- 15m1.00
- 1h1.00
- 4h1.00
- 16h1.00
- 64h1.00
- Estimated human time-to-complete1.00
- Source: METR, metr.org/time-horizons/0.99
- World'sFair0.99
- TRACK 9·JUNE 30,20260.97
- Data Quality0.99
-
- AlEngineer0.95
- World'sFair1.00
- 2. Model Horizon0.97
- Measuring tasks in "model units": tokens, steps, tool calls, etc.0.98
- Can be much more noisy, especially across different harnesses and models0.99
- Some models are more token-efficient than others on the same tasks (e.g. GPT-5.5 vs0.99
- Opus 4.8)1.00
- Easiest to interpret when holding these variables constant: if a task takes more tokens,0.99
- probably harder / longer-horizon0.98
- It's an imperfect metric, but still an important one to help define technical frontiers for0.99
- models: what makes a task difficult for models (e.g. compaction, memory, coherence over long0.99
- trajectories,etc.)1.00
- World'sFair0.99
- TRACK 9· JUNE 30,20260.96
- Data Quality0.97
-
- AlEngineer0.96
- World'sFair1.00
- So What's The Answer?0.99
- All of the above...1.00
- No metric is perfect or paints a fully accurate picture1.00
- What's long-horizon for a human is not necessarily long-horizon for an agent0.99
- e.g. tedious, time-intensive tasks, like formatting an Excel sheet, might take a human0.99
- many hours but might be fairly trivial for an agent (e.g. writes a Python script to fix all0.99
- formatting issues, which a financial expert is unlikely to do)0.99
- Measuring human time horizons depends a lot on the methodology and can be easy to skew1.00
- and distort0.96
- What is the quality and efficiency of human experts you're sampling from? What's0.99
- included in the measurement (reasoning/thinking for humans vs agents)? How accurate0.99
- are measurements as we push towards tasks that only the top 10% of experts can do?0.99
- Top 1%? Top 0.01%?0.99
- As human work and AI work diverges further and further, each metric tells you less and less0.99
- about the whole picture → important to look at metrics in tandem0.99
- World'sFair0.98
- TRACK 9· JUNE 30,20260.95
- Data Quality0.96
-
- AlEngineer0.98
- World'sFair1.00
- How do you measure model capabilities?1.00
- World'sFair1.00
- TRACK 9• JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.99
- World's Fair0.98
- HOW DO YOU MEASURE MODEL CAPABILITIES?0.97
- Environment Complexity — Tool Coordination0.98
- How many tools and external dependencies the agent has to coordinate1.00
- TAKEAWAYS1.00
- • Realistic tasks pull in more0.97
- Low complexity1.00
- High complexity0.99
- systems. This is a different lever1.00
- of difficulty than just running0.99
- longer1.00
- • Low end: "update only this0.98
- Python file" — one tool, one file,0.97
- 目0.99
- Grafana1.00
- nothing else to track1.00
- File1.00
- Agent1.00
- 570.67
- GitHub1.00
- • High end: the same bug fix0.97
- spans a database, Grafana, an0.99
- + many more1.00
- MCP deploy, and CI/CD — five0.94
- or six systems at once0.99
- World's Fair0.95
- TRACK 9·JUNE 30,20260.97
- Data Quality1.00
-
- AlEngineer0.99
- World's Fair0.97
- HOW DO YOU MEASURE MODEL CAPABILITIES?0.98
- Environment Complexity — Tool Coordination0.98
- How many tools and external dependencies the agent has to coordinate1.00
- TAKEAWAYS1.00
- • Realistic tasks pull in more0.97
- Low complexity1.00
- High complexity0.99
- systems. This is a different lever1.00
- of difficulty than just running0.99
- longer1.00
- • Low end: "update only this0.98
- Python file" — one tool, one file,0.97
- 目0.99
- Grafana1.00
- nothing else to track1.00
- File1.00
- Agent1.00
- 570.66
- GitHub1.00
- • High end: the same bug fix0.97
- spans a database, Grafana, an0.99
- + many more1.00
- MCP deploy, and CI/CD − five0.94
- or six systems at once0.99
- World's Fair0.99
- TRACK 9·JUNE 30,20260.98
- Data Quality1.00
-
- AlEngineer0.96
- World'sFair1.00
- HOW DO YOU MEASURE MODEL CAPABILITIES?1.00
- Environment Complexity — State Changes0.98
- The degree to which the content of the environment changes throughout the task1.00
- TAKEAWAYS1.00
- • Tasks can be made artificially long horizon by chaining together unrelated independent tasks.1.00
- However, merely making tasks take longer to complete misses the mark; a key component of adding1.00
- complexity is that an agent's initial decisions influence its options for later decisions0.99
- • Sequential complexity: a bad early query or misread dashboard cascades into every downstream1.00
- step, like a deploy and CI run that takes longer to resolve end-to-end0.99
- • Parallelizable complexity: Summarizing documentation in a large codebase can be made much1.00
- quicker by reading each file independently0.99
- World'sFair0.93
- TRACK 9· JUNE 30,20260.95
- Data Quality0.97
Transcript
225 cues· 4,565 words· 25,149 chars
- 0:12 It's great to see all of you here today.
- 0:14 We're super excited to talk about one of our favorite topics here at Theta.
- 0:18 Before we get started, we just want to introduce ourselves.
- 0:21 So hi, I'm a co-founder and CTO at Theta Software.
- 0:25 Hi, I'm Ray, and I'm a co-founder and CEO at Theta Software.
- 0:28 Prior to this, I was previously a founding engineer at DeepSilicon, where we did research into ternary models.
- 0:34 Awesome.
- 0:35 So I can get us started with the topic today.
- 0:37 We're gonna be talking about oral environments within the context of long horizon tasks.
- 0:42 And I think the most important thing for us to start with at the beginning is just talk about the trends and what long horizon actually means.
- 0:48 So, you know,
- 0:49 We all know that the horizon of which AI agents can work autonomously is accelerating really fast.
- 0:55 This is just some of the metrics that you can look at to see how this progress is really accelerating.
- 0:59 But I think it's really important to actually define what the time horizon here actually means.
- 1:03 We've gotten this data from one of the most common benchmarks out there that you've probably heard of for time horizons, previously in your Twitter feed all the time.
- 1:12 comes from meter, and meter kind of has one response or answer to this really important question about how do we actually define long horizon?
- 1:21 It's really important to understand because what we consider long horizon a year ago probably isn't really long horizon in our definition today, and what's long horizon today probably won't be long horizon in a year or two.
- 1:30 And I think that gets to our first point, which is that long horizon is really kind of a scalar metric.
- 1:36 useful for kind of measuring relative tasks, like one task might be more long rise than another, but it's really hard to define into kind of a binary category of this task is long rise and this task is not, especially as the kind of scope changes over time.
- 1:49 So I think the first way we can talk about defining this is how meter kind of looks at it, which is human horizon.
- 1:56 meaning can we use humans as a benchmark of, oh, this task takes humans a certain amount of time.
- 2:01 So if AI agents can do that, then they've reached this certain critical level of kind of a time horizon.
- 2:07 And the way Meter kind of does it is they have thresholds for the tasks they care about.
- 2:10 So they have a 50% threshold, meaning
- 2:13 If a certain model reaches a 16-hour threshold on this benchmark, that means it can achieve tasks with a 50% success rate that take a human 16 hours.
- 2:21 And there's a really rigorous methodology of how they actually measure how did it take human 16 hours, but we will kind of avoid some of those details.
- 2:29 The other way we usually think about what long horizon actually means is not with the reference of humans, but instead the reference of models.
- 2:37 So some of the relevant model units we usually care about are things like tokens, how many tokens are consumed in a trajectory, how many steps it took, how many tool calls it kind of takes.
- 2:46 And these can be really noisy, right?
- 2:47 Because I'm sure you guys have used different models, like a lot of the Codex models are seen as more token efficient than some of the Cloud models.
- 2:55 it's a pretty noisy estimate for a couple of reasons.
- 2:57 One is that which model you're using, like I just said, and different harnesses you care about have a pretty big impact on how many tokens are actually consumed on a task, right?
- 3:04 So this can be pretty hard to interpret when you're not holding variables constant.
- 3:09 You know, if a task takes GPT model 500,000 tokens, that doesn't really tell you a lot about what that task would look like for cloud models until you actually run on those cloud models.
- 3:18 But despite it being a pretty noisy metric, it's actually really useful and important for us to understand because
- 3:25 The amount of tokens that are consumed tells us a lot about how difficult a task actually is for an AI agent to kind of tackle it autonomously, right?
- 3:33 You have to deal with things like compaction over long horizons.
- 3:36 They don't really stay coherent over enough steps or trajectory length that you kind of achieve.
- 3:41 So even though it's kind of a noisy metric,
- 3:44 It can be really useful when, you know, if we look at what a GPT 5.5 model can do now, and then you use the same model generation and kind of see, oh, now can actually achieve a million trajectory based on an increased context window or improved compaction endpoint.
- 3:56 That tells us a lot about how autonomous AI agents can actually go for long periods of time in that sense.
- 4:02 And it really defines for us what the technical frontier actually means for models right now.
- 4:07 Maybe not really human adjacent, it's really hard to say how many tokens a task takes for a human, because we don't really think in tokens, but still very useful in that kind of sense.
- 4:16 So these are two different approaches we can think about, but what's actually the right way to think about this?
- 4:21 The answer is that we probably want to think about all of these.
- 4:23 And if we just look at one of these metrics in isolation, it's probably not a great way of measuring things.
- 4:28 So I went through some of the weaknesses with measuring with like model specific metrics like tokens and steps.
- 4:34 But there's also a lot of weaknesses in the other approach of kind of relying on humans.
- 4:39 You know, what's long horizon for a human isn't necessarily that difficult for a model, depending on what the actual task you care about is.
- 4:45 You know, there's a lot of tasks that are really tedious and time intensive.
loading
Chapters
- 0:00 What does long horizon mean?
- 1:13 Time horizon and the threshold metric
- 3:17 Why the metric is noisy
- 4:20 Measuring what actually matters
- 6:38 Creating tasks and environments
- 7:42 When a bad early step cascades
- 10:01 Why standardized evaluation is hard
- 11:17 Verifying from the final state
- 13:46 Judges, tools, and reused agents
- 17:45 Rubrics, QA, and careful grading