read-only demo

Videos zkX03APVj0M

Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:16:32· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:21, 1 of 1 keyframes kept
  5. Shot 4, 0:21 to 0:47, 1 of 1 keyframes kept
  6. Shot 5, 0:47 to 1:13, 1 of 1 keyframes kept
  7. Shot 6, 1:13 to 1:39, 1 of 1 keyframes kept
  8. Shot 7, 1:39 to 2:06, 0 of 1 keyframes kept
  9. Shot 8, 2:06 to 2:32, 1 of 1 keyframes kept
  10. Shot 9, 2:32 to 2:58, 0 of 1 keyframes kept
  11. Shot 10, 2:58 to 3:24, 0 of 1 keyframes kept
  12. Shot 11, 3:24 to 3:50, 1 of 1 keyframes kept
  13. Shot 12, 3:50 to 4:16, 0 of 1 keyframes kept
  14. Shot 13, 4:16 to 4:43, 0 of 1 keyframes kept
  15. Shot 14, 4:43 to 5:09, 1 of 1 keyframes kept
  16. Shot 15, 5:09 to 5:35, 0 of 1 keyframes kept
  17. Shot 16, 5:35 to 6:01, 0 of 1 keyframes kept
  18. Shot 17, 6:01 to 6:27, 0 of 1 keyframes kept
  19. Shot 18, 6:27 to 6:53, 0 of 1 keyframes kept
  20. Shot 19, 6:53 to 7:20, 1 of 1 keyframes kept
  21. Shot 20, 7:20 to 7:47, 0 of 1 keyframes kept
  22. Shot 21, 7:47 to 8:13, 1 of 1 keyframes kept
  23. Shot 22, 8:13 to 8:40, 0 of 1 keyframes kept
  24. Shot 23, 8:40 to 9:06, 1 of 1 keyframes kept
  25. Shot 24, 9:06 to 9:33, 0 of 1 keyframes kept
  26. Shot 25, 9:33 to 10:00, 1 of 1 keyframes kept
  27. Shot 26, 10:00 to 10:26, 1 of 1 keyframes kept
  28. Shot 27, 10:26 to 10:53, 1 of 1 keyframes kept
  29. Shot 28, 10:53 to 11:19, 0 of 1 keyframes kept
  30. Shot 29, 11:19 to 11:46, 0 of 1 keyframes kept
  31. Shot 30, 11:46 to 12:15, 1 of 1 keyframes kept
  32. Shot 31, 12:15 to 12:45, 0 of 1 keyframes kept
  33. Shot 32, 12:45 to 13:11, 1 of 1 keyframes kept
  34. Shot 33, 13:11 to 13:38, 1 of 1 keyframes kept
  35. Shot 34, 13:38 to 14:04, 1 of 1 keyframes kept
  36. Shot 35, 14:04 to 14:30, 1 of 1 keyframes kept
  37. Shot 36, 14:30 to 14:57, 1 of 1 keyframes kept
  38. Shot 37, 14:57 to 15:23, 1 of 1 keyframes kept
  39. Shot 38, 15:23 to 15:49, 1 of 1 keyframes kept
  40. Shot 39, 15:49 to 16:15, 1 of 1 keyframes kept
  41. Shot 40, 16:15 to 16:32, 0 of 1 keyframes kept

41 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
164
whisperx 164
chunks
28
from 164 cues
keyframes
25
kept of 41 captured
frames with text
25
309 lines read
chapters
10
from the source metadata
keyframe bytes
4.3 MB
word timings on 164 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:36 0s
stt done 2026-08-09 14:46 19s
chunk done 2026-08-09 14:47 0s
text_embed done 2026-08-10 19:42 1s
keyframe done 2026-08-09 14:47 3m 44s
ocr done 2026-08-09 14:50 16s
frame_embed done 2026-08-10 19:42 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:13 #3 done2 line(s)

    shot 3·sharpness 347.8

    1. AlEngineer1.00
    2. World's Fair0.98
  • 0:29 #4 done9 line(s)

    shot 4·sharpness 1426.5

    1. July 1, 20260.98
    2. AlEngineer0.99
    3. World'sFair1.00
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. Data for running software companies1.00
    7. Emulated1.00
    8. World'sFair0.96
    9. Engineering the future of Al0.99
  • 1:10 #5 done10 line(s)

    shot 5·sharpness 1559.5

    1. July 1, 20260.98
    2. AlEngineer0.99
    3. World'sFair1.00
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. Data for running software companies1.00
    7. Emulated1.00
    8. World'sFair1.00
    9. TRACK 9· JULY 1, 20260.95
    10. Posttraining & Midtraining1.00
  • 1:19 #6 done14 line(s)

    shot 6·sharpness 1382.8

    1. Emulated1.00
    2. July 1, 20260.99
    3. AlEngineer0.99
    4. World's Fair0.97
    5. The Team1.00
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. Joseph Wang0.96
    9. Sid Patllollu1.00
    10. CEO1.00
    11. CTO1.00
    12. World's Fair0.95
    13. TRACK 9 · JULY 1, 20260.93
    14. Posttraining & Midtraining0.99
  • 1:47 #7 skipped

    shot 7·duplicate of #6

  • 2:26 #8 done9 line(s)

    shot 8·sharpness 1544.6

    1. Emulated1.00
    2. July 1, 20260.98
    3. AlEngineer0.98
    4. World'sFair1.00
    5. The model capability gap0.98
    6. is data gap1.00
    7. World's Fair0.99
    8. TRACK 9 · JULY 1, 20260.93
    9. Posttraining & Midtraining1.00
  • 2:55 #9 skipped

    shot 9·duplicate of #8

  • 3:01 #10 skipped

    shot 10·duplicate of #8

  • 3:47 #11 done28 line(s)

    shot 11·sharpness 2644.9

    1. Emulated1.00
    2. July 1, 20260.97
    3. AlEngineer0.98
    4. World'sFair1.00
    5. How do we close the gap?0.97
    6. Unspecified1.00
    7. Open-ended1.00
    8. Complex1.00
    9. Like a PM, we give agents scope and1.00
    10. We simulate issues that only appear at1.00
    11. Agents handle multiple components0.99
    12. context to discover and act on issues.1.00
    13. scale like power loss, corrupted disks,0.99
    14. with live traffic, clusters,0.99
    15. This includes company context with0.99
    16. network failures, and clock skew.1.00
    17. software-defined networks, dozens of0.99
    18. historical decisions, past incidents,1.00
    19. Agents try out multiple solutions and0.99
    20. engineering tools, pipelines, and1.00
    21. and planned work but also customer0.99
    22. we build open-ended, deterministic1.00
    23. observability stacks.0.99
    24. conversations.1.00
    25. verifiers that accept all of them.0.99
    26. World's Fair0.96
    27. TRACK 9· JULY 1, 20260.96
    28. Posttraining & Midtraining1.00
  • 4:06 #12 skipped

    shot 12·duplicate of #11

  • 4:25 #13 skipped

    shot 13·duplicate of #11

  • 4:46 #14 done46 line(s)

    shot 14·sharpness 1740.6

    1. Emulated1.00
    2. July 1, 20261.00
    3. AlEngineer0.99
    4. World's Fair0.99
    5. Service Discovery App1.00
    6. Knowledge Tools1.00
    7. Uses etcd as a store for services0.98
    8. Tickets, postmortems,1.00
    9. alerts, runbooks1.00
    10. Read/write/watch/lease workload1.00
    11. 1. Agent establishes context1.00
    12. etcd gRPC proxy1.00
    13. etcd source code1.00
    14. Directs traffic to live nodes1.00
    15. 2. Agent implements1.00
    16. observability feature0.98
    17. Read/write/watch/lease workload1.00
    18. 4. Kicks off rolling deployment0.99
    19. Live Voters1.00
    20. Plant1.00
    21. Promote-Ready Learners0.99
    22. Stale Member1.00
    23. Leader1.00
    24. Follower1.00
    25. Follower1.00
    26. Lagging1.00
    27. learner1.00
    28. Learner1.00
    29. Learner1.00
    30. Learner1.00
    31. 3. Agent cleans up stale1.00
    32. leftover from a previous1.00
    33. membership record0.99
    34. migration1.00
    35. Replication1.00
    36. stream1.00
    37. 6. Agent promotes0.99
    38. healthy nodes only0.99
    39. Note: Only one node is the learner at a1.00
    40. time. Labels are for visual consistency0.99
    41. Prometheus/Grafana1.00
    42. 5. Uses metrics to identify0.99
    43. learners ready for promotion1.00
    44. World's Fair0.97
    45. TRACK 9 · JULY 1, 20260.93
    46. Posttraining & Midtraining1.00
  • 5:27 #15 skipped

    shot 15·duplicate of #14

  • 5:55 #16 skipped

    shot 16·duplicate of #14

  • 6:12 #17 skipped

    shot 17·duplicate of #14

  • 6:30 #18 skipped

    shot 18·duplicate of #14

  • 7:07 #19 done8 line(s)

    shot 19·sharpness 1142.9

    1. Emulated1.00
    2. July 1, 20260.98
    3. AlEngineer0.98
    4. World's Fair0.97
    5. But this is not enough...0.99
    6. World's Fair0.97
    7. TRACK 9 · JULY 1, 20260.93
    8. Posttraining & Midtraining1.00
  • 7:31 #20 skipped

    shot 20·duplicate of #19

  • 7:52 #21 done9 line(s)

    shot 21·sharpness 986.8

    1. Emulated1.00
    2. July 1, 20260.98
    3. AlEngineer0.99
    4. World's Fair0.98
    5. My shiny1.00
    6. software1.00
    7. World's Fair0.96
    8. TRACK 9 · JULY 1, 20260.93
    9. Posttraining & Midtraining1.00
  • 8:32 #22 skipped

    shot 22·duplicate of #21

  • 8:51 #23 done14 line(s)

    shot 23·sharpness 1489.5

    1. Emulated1.00
    2. July 1, 20260.99
    3. AlEngineer0.99
    4. World'sFair1.00
    5. Host1.00
    6. Provisioning1.00
    7. Creates1.00
    8. Instance/Container1.00
    9. Frontend API0.99
    10. My shiny1.00
    11. software1.00
    12. World's Fair0.96
    13. TRACK 9 · JULY 1, 20260.93
    14. Posttraining & Midtraining1.00

Transcript

164 cues· 2,316 words· 13,222 chars

  1. 0:14 So I appreciate the intro.
  2. 0:15 My name is Joseph, and this is my co-founder, Sid.
  3. 0:19 Emulated is a data lab focused on increasing the reliability and autonomy of AI agents.
  4. 0:26 And if you've been at AI Engineer and you've watched the talks, seen the tracks, then there's probably one takeaway that all the talks have in common.
  5. 0:36 And it's that we're headed towards a future where agents are able to perform useful work over longer and longer horizons with little to no supervision.
  6. 0:45 So today we're going to answer some of the questions of what this means for the data and model layers.
  7. 0:52 We're going to touch on some pretty cool things, so look out for them, like how to simulate a company within a sandbox or sandboxes for multi-node systems and distributed clusters.
  8. 1:04 And if we have a little bit of time, we'll also go into some of the work that we're doing with post-training pipelines and how these new types of sandboxes are affecting post-training infra as well.
  9. 1:17 So where Sid and I come from, our backgrounds are in network infra, distributed databases, and sandbox infra.
  10. 1:25 And these are all areas where the workloads are mission critical.
  11. 1:30 We all saw a couple months ago that when something like DynamoDB goes down, so does US East 1 and half the internet.
  12. 1:39 And working on these systems, we saw a model capability gap when it came to operating and building these systems at scale and thinking about the consequences of architecture and systems design over the course of years.
  13. 1:55 Yeah, so it led to a pretty natural question, right?
  14. 1:58 For such mission critical services, why is it that my model or my agent is so proficient at handling the application layer, but struggles when it comes to reasoning through infrastructure complexities, for example, things like MVCC on a database engine, which can lead to corruption issues, which was one of the roots of the DynamoDB failure a few months ago.
  15. 2:26 So like with everything in NML, the gap in models is usually a gap in data.
  16. 2:31 Models typically are only as good as data is.
  17. 2:37 And to really highlight this point, right, model capability has never regressed whenever you introduce more high-quality data.
  18. 2:46 So with that being said, what is the data gap then?
  19. 2:48 What does data look like right now?
  20. 2:51 And how is this influencing the model capability gap here?
  21. 2:55 So if you look at any of the frontier or recent benchmarks, like SuiteBench Pro, TerminalBench, or something like Frontier Code and DeepSuite,
  22. 3:05 The tasks only operate within the codebase.
  23. 3:09 The agent is given a pretty large task and over the course of 50 to 100 turns produces a couple thousand line PR.
  24. 3:22 but it doesn't do all of the work that a human does.
  25. 3:25 It doesn't do what a PM does with talking to customers, understanding their problems, what an engineer does with trying out different approaches, performance testing them, and owning the underlying infra for the code base over the course of not just months, but years.
  26. 3:44 And this is really the gap that we're closing.
  27. 3:47 We've taken software engineering companies and we've put them into containerized environments.
  28. 3:52 So this includes organizational contexts like projects, incidents, customer conversations.
  29. 4:01 The agent also has to deal with issues that only appear at scale, like network failures between distributed nodes, data corruption, and clock skew.
  30. 4:10 And through all this, we also want the agents to reason about orchestrating through distributed clusters and also thinking about things like operational blast radius while solving live traffic.
  31. 4:21 And the result is that the task that these agents have to complete or we want the agents to learn is that environments are far more complex and long horizon than a simple code disk.
  32. 4:35 So let's just, let's bring a picture into the mix because it tends to make things more interesting.
  33. 4:42 Here's an example we've built of an etcd consensus cluster that a typical production service might rely on.
  34. 4:49 So an early environment might, tended to operate and work primarily on that little blue square entitled etcd source code in the bottom right there.
  35. 5:00 But a lot of the fun and the model capability gap that results from it is really in everything that surrounds it.
  36. 5:08 So you start with the tickets, projects, postmortems.
  37. 5:12 What are the train wrecks?
  38. 5:14 Why did they happen?
  39. 5:15 How did customers feel about them?
  40. 5:17 And oftentimes those aren't necessarily up to date.
  41. 5:21 The agent has to incorporate all that when it's reasoning through the actual change that
  42. 5:28 current environments have it make.
  43. 5:30 After it makes that change, you need to kick off rolling deployments.
  44. 5:35 Those deployment systems can oftentimes be complicated, have conflicts, may not work.
  45. 5:42 And all through that, when you're finally migrating off of from old hardware onto new hardware, you run into unforeseen problems, which you
  46. 5:52 did not, the agent has to reason through in real time, just like a human would.
  47. 5:58 You have failing nodes.
  48. 6:01 You have stale, deprecated nodes.
  49. 6:04 And while all of this is happening, the service can't go down.
  50. 6:09 Because there is a blast radius to serving live traffic.

Chapters

  1. 0:00 Useful work over longer horizons
  2. 1:20 Backgrounds in network infrastructure
  3. 2:26 How environments shape capability
  4. 3:16 Fifty to a hundred turn tasks
  5. 4:59 Why real incidents are messy
  6. 7:11 Real infrastructure isn't a code diff
  7. 7:40 Acting as an engineer inside the cloud
  8. 9:37 Deployment, cost, and scaling bars
  9. 13:29 Why it's called Emulated
  10. 15:01 Simulating full companies

Open at this second