read-only demo

Videos 2aS7aKoXn64

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

index_state ready data_status ok

AI Engineer· published 2026-08-01· 0:21:14· en-US· indexed 2026-08-10 19:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:18, 1 of 1 keyframes kept
  5. Shot 4, 0:18 to 0:43, 1 of 1 keyframes kept
  6. Shot 5, 0:43 to 1:11, 1 of 1 keyframes kept
  7. Shot 6, 1:11 to 1:48, 1 of 1 keyframes kept
  8. Shot 7, 1:48 to 2:31, 1 of 1 keyframes kept
  9. Shot 8, 2:31 to 2:57, 1 of 1 keyframes kept
  10. Shot 9, 2:57 to 3:24, 0 of 1 keyframes kept
  11. Shot 10, 3:24 to 3:51, 0 of 1 keyframes kept
  12. Shot 11, 3:51 to 4:17, 0 of 1 keyframes kept
  13. Shot 12, 4:17 to 4:45, 1 of 1 keyframes kept
  14. Shot 13, 4:45 to 5:12, 0 of 1 keyframes kept
  15. Shot 14, 5:12 to 5:40, 0 of 1 keyframes kept
  16. Shot 15, 5:40 to 6:07, 0 of 1 keyframes kept
  17. Shot 16, 6:07 to 6:34, 0 of 1 keyframes kept
  18. Shot 17, 6:34 to 6:53, 1 of 1 keyframes kept
  19. Shot 18, 6:53 to 7:22, 1 of 1 keyframes kept
  20. Shot 19, 7:22 to 7:50, 1 of 1 keyframes kept
  21. Shot 20, 7:50 to 8:17, 1 of 1 keyframes kept
  22. Shot 21, 8:17 to 8:43, 0 of 1 keyframes kept
  23. Shot 22, 8:43 to 9:09, 0 of 1 keyframes kept
  24. Shot 23, 9:09 to 9:37, 0 of 1 keyframes kept
  25. Shot 24, 9:37 to 10:05, 0 of 1 keyframes kept
  26. Shot 25, 10:05 to 10:22, 1 of 1 keyframes kept
  27. Shot 26, 10:22 to 10:56, 1 of 1 keyframes kept
  28. Shot 27, 10:56 to 11:31, 0 of 1 keyframes kept
  29. Shot 28, 11:31 to 12:00, 1 of 1 keyframes kept
  30. Shot 29, 12:00 to 12:29, 0 of 1 keyframes kept
  31. Shot 30, 12:29 to 12:58, 0 of 1 keyframes kept
  32. Shot 31, 12:58 to 13:27, 0 of 1 keyframes kept
  33. Shot 32, 13:27 to 13:40, 1 of 1 keyframes kept
  34. Shot 33, 13:40 to 14:07, 0 of 1 keyframes kept
  35. Shot 34, 14:07 to 14:34, 0 of 1 keyframes kept
  36. Shot 35, 14:34 to 15:01, 1 of 1 keyframes kept
  37. Shot 36, 15:01 to 15:29, 0 of 1 keyframes kept
  38. Shot 37, 15:29 to 16:00, 0 of 1 keyframes kept
  39. Shot 38, 16:00 to 16:31, 0 of 1 keyframes kept
  40. Shot 39, 16:31 to 17:15, 0 of 1 keyframes kept
  41. Shot 40, 17:15 to 17:51, 1 of 1 keyframes kept
  42. Shot 41, 17:51 to 17:53, 1 of 1 keyframes kept
  43. Shot 42, 17:53 to 18:21, 0 of 1 keyframes kept
  44. Shot 43, 18:21 to 18:52, 1 of 1 keyframes kept
  45. Shot 44, 18:52 to 19:24, 0 of 1 keyframes kept
  46. Shot 45, 19:24 to 19:55, 0 of 1 keyframes kept
  47. Shot 46, 19:55 to 20:26, 0 of 1 keyframes kept
  48. Shot 47, 20:26 to 20:49, 1 of 1 keyframes kept
  49. Shot 48, 20:49 to 20:52, 0 of 1 keyframes kept
  50. Shot 49, 20:52 to 20:57, 0 of 1 keyframes kept
  51. Shot 50, 20:57 to 21:14, 0 of 1 keyframes kept

51 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
225
whisperx 225
chunks
38
from 225 cues
keyframes
23
kept of 51 captured
frames with text
23
411 lines read
chapters
10
from the source metadata
keyframe bytes
6.8 MB
word timings on 225 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:32 0s
stt done 2026-08-09 00:57 25s
chunk done 2026-08-09 00:57 0s
text_embed done 2026-08-10 19:36 1s
keyframe done 2026-08-09 00:58 2m 42s
ocr done 2026-08-09 01:00 11s
frame_embed done 2026-08-10 19:36 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.7

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2743.2

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:17 #3 done2 line(s)

    shot 3·sharpness 215.2

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:40 #4 done10 line(s)

    shot 4·sharpness 1470.2

    1. AlEngineer0.96
    2. World'sFair1.00
    3. DATA QUALITY · AI Engineer World's Fair 20260.98
    4. PRESENTED BY0.99
    5. Rethinking Environments for0.99
    6. Microsoft1.00
    7. Long Horizon Tasks1.00
    8. θ Theta0.94
    9. World'sFair1.00
    10. Engineering the future of Al0.99
  • 0:46 #5 done23 line(s)

    shot 5·sharpness 2492.4

    1. AlEngineer0.98
    2. World's Fair0.99
    3. Task Horizons Are Accelerating0.98
    4. Model1.00
    5. Date Released1.00
    6. Time Horizon0.96
    7. PRESENTED BY1.00
    8. Microsoft1.00
    9. Claude 3.7 Sonnet1.00
    10. Feb 20251.00
    11. 1 hr0.88
    12. Claude Opus 41.00
    13. May 20251.00
    14. 1.7 hrs0.99
    15. Claude Opus 4.50.99
    16. Nov 20251.00
    17. 4.9 hrs1.00
    18. Claude Opus 4.61.00
    19. Feb 20261.00
    20. 12 hrs1.00
    21. Source: METR, metr.org/time-horizons/1.00
    22. World's Fair0.99
    23. Engineering the future of Al1.00
  • 1:15 #6 done8 line(s)

    shot 6·sharpness 1030.1

    1. AlEngineer0.98
    2. World'sFair1.00
    3. PRESENTED BY1.00
    4. Microsoft1.00
    5. How do we define “long horizon"?0.97
    6. World's Fair0.96
    7. TRACK 9· JUNE 30, 20260.96
    8. Data Quality0.96
  • 2:26 #7 done27 line(s)

    shot 7·sharpness 1755.8

    1. AlEngineer0.97
    2. World'sFair1.00
    3. 1. Human Horizon0.99
    4. Predicted 50% time horizon: 12 hr for Claude Opus 4.60.99
    5. Measurements above 16 hrs are unreliable with our current task suite0.99
    6. 100%1.00
    7. Est ss te0.67
    8. 80%-1.00
    9. 60%-0.89
    10. 40%-1.00
    11. 20%-1.00
    12. 0%-0.99
    13. 1s1.00
    14. 4s1.00
    15. 15s1.00
    16. 1m1.00
    17. 4m1.00
    18. 15m1.00
    19. 1h1.00
    20. 4h1.00
    21. 16h1.00
    22. 64h1.00
    23. Estimated human time-to-complete1.00
    24. Source: METR, metr.org/time-horizons/0.99
    25. World'sFair0.99
    26. TRACK 9·JUNE 30,20260.97
    27. Data Quality0.99
  • 2:52 #8 done15 line(s)

    shot 8·sharpness 4144.6

    1. AlEngineer0.95
    2. World'sFair1.00
    3. 2. Model Horizon0.97
    4. Measuring tasks in "model units": tokens, steps, tool calls, etc.0.98
    5. Can be much more noisy, especially across different harnesses and models0.99
    6. Some models are more token-efficient than others on the same tasks (e.g. GPT-5.5 vs0.99
    7. Opus 4.8)1.00
    8. Easiest to interpret when holding these variables constant: if a task takes more tokens,0.99
    9. probably harder / longer-horizon0.98
    10. It's an imperfect metric, but still an important one to help define technical frontiers for0.99
    11. models: what makes a task difficult for models (e.g. compaction, memory, coherence over long0.99
    12. trajectories,etc.)1.00
    13. World'sFair0.99
    14. TRACK 9· JUNE 30,20260.96
    15. Data Quality0.97
  • 3:00 #9 skipped

    shot 9·duplicate of #8

  • 3:32 #10 skipped

    shot 10·duplicate of #8

  • 4:09 #11 skipped

    shot 11·duplicate of #8

  • 4:41 #12 done20 line(s)

    shot 12·sharpness 6150.7

    1. AlEngineer0.96
    2. World'sFair1.00
    3. So What's The Answer?0.99
    4. All of the above...1.00
    5. No metric is perfect or paints a fully accurate picture1.00
    6. What's long-horizon for a human is not necessarily long-horizon for an agent0.99
    7. e.g. tedious, time-intensive tasks, like formatting an Excel sheet, might take a human0.99
    8. many hours but might be fairly trivial for an agent (e.g. writes a Python script to fix all0.99
    9. formatting issues, which a financial expert is unlikely to do)0.99
    10. Measuring human time horizons depends a lot on the methodology and can be easy to skew1.00
    11. and distort0.96
    12. What is the quality and efficiency of human experts you're sampling from? What's0.99
    13. included in the measurement (reasoning/thinking for humans vs agents)? How accurate0.99
    14. are measurements as we push towards tasks that only the top 10% of experts can do?0.99
    15. Top 1%? Top 0.01%?0.99
    16. As human work and AI work diverges further and further, each metric tells you less and less0.99
    17. about the whole picture → important to look at metrics in tandem0.99
    18. World'sFair0.98
    19. TRACK 9· JUNE 30,20260.95
    20. Data Quality0.96
  • 5:01 #13 skipped

    shot 13·duplicate of #12

  • 5:21 #14 skipped

    shot 14·duplicate of #12

  • 5:56 #15 skipped

    shot 15·duplicate of #12

  • 6:13 #16 skipped

    shot 16·duplicate of #12

  • 6:37 #17 done6 line(s)

    shot 17·sharpness 962.6

    1. AlEngineer0.98
    2. World'sFair1.00
    3. How do you measure model capabilities?1.00
    4. World'sFair1.00
    5. TRACK 9• JUNE 30, 20260.97
    6. Data Quality1.00
  • 7:02 #18 done29 line(s)

    shot 18·sharpness 3967.6

    1. AlEngineer0.99
    2. World's Fair0.98
    3. HOW DO YOU MEASURE MODEL CAPABILITIES?0.97
    4. Environment Complexity — Tool Coordination0.98
    5. How many tools and external dependencies the agent has to coordinate1.00
    6. TAKEAWAYS1.00
    7. • Realistic tasks pull in more0.97
    8. Low complexity1.00
    9. High complexity0.99
    10. systems. This is a different lever1.00
    11. of difficulty than just running0.99
    12. longer1.00
    13. • Low end: "update only this0.98
    14. Python file" — one tool, one file,0.97
    15. 0.99
    16. Grafana1.00
    17. nothing else to track1.00
    18. File1.00
    19. Agent1.00
    20. 570.67
    21. GitHub1.00
    22. • High end: the same bug fix0.97
    23. spans a database, Grafana, an0.99
    24. + many more1.00
    25. MCP deploy, and CI/CD — five0.94
    26. or six systems at once0.99
    27. World's Fair0.95
    28. TRACK 9·JUNE 30,20260.97
    29. Data Quality1.00
  • 7:39 #19 done29 line(s)

    shot 19·sharpness 3975.3

    1. AlEngineer0.99
    2. World's Fair0.97
    3. HOW DO YOU MEASURE MODEL CAPABILITIES?0.98
    4. Environment Complexity — Tool Coordination0.98
    5. How many tools and external dependencies the agent has to coordinate1.00
    6. TAKEAWAYS1.00
    7. • Realistic tasks pull in more0.97
    8. Low complexity1.00
    9. High complexity0.99
    10. systems. This is a different lever1.00
    11. of difficulty than just running0.99
    12. longer1.00
    13. • Low end: "update only this0.98
    14. Python file" — one tool, one file,0.97
    15. 0.99
    16. Grafana1.00
    17. nothing else to track1.00
    18. File1.00
    19. Agent1.00
    20. 570.66
    21. GitHub1.00
    22. • High end: the same bug fix0.97
    23. spans a database, Grafana, an0.99
    24. + many more1.00
    25. MCP deploy, and CI/CD − five0.94
    26. or six systems at once0.99
    27. World's Fair0.99
    28. TRACK 9·JUNE 30,20260.98
    29. Data Quality1.00
  • 8:14 #20 done16 line(s)

    shot 20·sharpness 5319.9

    1. AlEngineer0.96
    2. World'sFair1.00
    3. HOW DO YOU MEASURE MODEL CAPABILITIES?1.00
    4. Environment Complexity — State Changes0.98
    5. The degree to which the content of the environment changes throughout the task1.00
    6. TAKEAWAYS1.00
    7. • Tasks can be made artificially long horizon by chaining together unrelated independent tasks.1.00
    8. However, merely making tasks take longer to complete misses the mark; a key component of adding1.00
    9. complexity is that an agent's initial decisions influence its options for later decisions0.99
    10. • Sequential complexity: a bad early query or misread dashboard cascades into every downstream1.00
    11. step, like a deploy and CI run that takes longer to resolve end-to-end0.99
    12. • Parallelizable complexity: Summarizing documentation in a large codebase can be made much1.00
    13. quicker by reading each file independently0.99
    14. World'sFair0.93
    15. TRACK 9· JUNE 30,20260.95
    16. Data Quality0.97
  • 8:22 #21 skipped

    shot 21·duplicate of #20

  • 8:51 #22 skipped

    shot 22·duplicate of #19

  • 9:26 #23 skipped

    shot 23·duplicate of #20

Transcript

225 cues· 4,565 words· 25,149 chars

  1. 0:12 It's great to see all of you here today.
  2. 0:14 We're super excited to talk about one of our favorite topics here at Theta.
  3. 0:18 Before we get started, we just want to introduce ourselves.
  4. 0:21 So hi, I'm a co-founder and CTO at Theta Software.
  5. 0:25 Hi, I'm Ray, and I'm a co-founder and CEO at Theta Software.
  6. 0:28 Prior to this, I was previously a founding engineer at DeepSilicon, where we did research into ternary models.
  7. 0:34 Awesome.
  8. 0:35 So I can get us started with the topic today.
  9. 0:37 We're gonna be talking about oral environments within the context of long horizon tasks.
  10. 0:42 And I think the most important thing for us to start with at the beginning is just talk about the trends and what long horizon actually means.
  11. 0:48 So, you know,
  12. 0:49 We all know that the horizon of which AI agents can work autonomously is accelerating really fast.
  13. 0:55 This is just some of the metrics that you can look at to see how this progress is really accelerating.
  14. 0:59 But I think it's really important to actually define what the time horizon here actually means.
  15. 1:03 We've gotten this data from one of the most common benchmarks out there that you've probably heard of for time horizons, previously in your Twitter feed all the time.
  16. 1:12 comes from meter, and meter kind of has one response or answer to this really important question about how do we actually define long horizon?
  17. 1:21 It's really important to understand because what we consider long horizon a year ago probably isn't really long horizon in our definition today, and what's long horizon today probably won't be long horizon in a year or two.
  18. 1:30 And I think that gets to our first point, which is that long horizon is really kind of a scalar metric.
  19. 1:36 useful for kind of measuring relative tasks, like one task might be more long rise than another, but it's really hard to define into kind of a binary category of this task is long rise and this task is not, especially as the kind of scope changes over time.
  20. 1:49 So I think the first way we can talk about defining this is how meter kind of looks at it, which is human horizon.
  21. 1:56 meaning can we use humans as a benchmark of, oh, this task takes humans a certain amount of time.
  22. 2:01 So if AI agents can do that, then they've reached this certain critical level of kind of a time horizon.
  23. 2:07 And the way Meter kind of does it is they have thresholds for the tasks they care about.
  24. 2:10 So they have a 50% threshold, meaning
  25. 2:13 If a certain model reaches a 16-hour threshold on this benchmark, that means it can achieve tasks with a 50% success rate that take a human 16 hours.
  26. 2:21 And there's a really rigorous methodology of how they actually measure how did it take human 16 hours, but we will kind of avoid some of those details.
  27. 2:29 The other way we usually think about what long horizon actually means is not with the reference of humans, but instead the reference of models.
  28. 2:37 So some of the relevant model units we usually care about are things like tokens, how many tokens are consumed in a trajectory, how many steps it took, how many tool calls it kind of takes.
  29. 2:46 And these can be really noisy, right?
  30. 2:47 Because I'm sure you guys have used different models, like a lot of the Codex models are seen as more token efficient than some of the Cloud models.
  31. 2:55 it's a pretty noisy estimate for a couple of reasons.
  32. 2:57 One is that which model you're using, like I just said, and different harnesses you care about have a pretty big impact on how many tokens are actually consumed on a task, right?
  33. 3:04 So this can be pretty hard to interpret when you're not holding variables constant.
  34. 3:09 You know, if a task takes GPT model 500,000 tokens, that doesn't really tell you a lot about what that task would look like for cloud models until you actually run on those cloud models.
  35. 3:18 But despite it being a pretty noisy metric, it's actually really useful and important for us to understand because
  36. 3:25 The amount of tokens that are consumed tells us a lot about how difficult a task actually is for an AI agent to kind of tackle it autonomously, right?
  37. 3:33 You have to deal with things like compaction over long horizons.
  38. 3:36 They don't really stay coherent over enough steps or trajectory length that you kind of achieve.
  39. 3:41 So even though it's kind of a noisy metric,
  40. 3:44 It can be really useful when, you know, if we look at what a GPT 5.5 model can do now, and then you use the same model generation and kind of see, oh, now can actually achieve a million trajectory based on an increased context window or improved compaction endpoint.
  41. 3:56 That tells us a lot about how autonomous AI agents can actually go for long periods of time in that sense.
  42. 4:02 And it really defines for us what the technical frontier actually means for models right now.
  43. 4:07 Maybe not really human adjacent, it's really hard to say how many tokens a task takes for a human, because we don't really think in tokens, but still very useful in that kind of sense.
  44. 4:16 So these are two different approaches we can think about, but what's actually the right way to think about this?
  45. 4:21 The answer is that we probably want to think about all of these.
  46. 4:23 And if we just look at one of these metrics in isolation, it's probably not a great way of measuring things.
  47. 4:28 So I went through some of the weaknesses with measuring with like model specific metrics like tokens and steps.
  48. 4:34 But there's also a lot of weaknesses in the other approach of kind of relying on humans.
  49. 4:39 You know, what's long horizon for a human isn't necessarily that difficult for a model, depending on what the actual task you care about is.
  50. 4:45 You know, there's a lot of tasks that are really tedious and time intensive.

Chapters

  1. 0:00 What does long horizon mean?
  2. 1:13 Time horizon and the threshold metric
  3. 3:17 Why the metric is noisy
  4. 4:20 Measuring what actually matters
  5. 6:38 Creating tasks and environments
  6. 7:42 When a bad early step cascades
  7. 10:01 Why standardized evaluation is hard
  8. 11:17 Verifying from the final state
  9. 13:46 Judges, tools, and reused agents
  10. 17:45 Rubrics, QA, and careful grading

Open at this second