Videos zkX03APVj0M
Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
Scene timeline
41 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 164
- whisperx 164
- chunks
- 28
- from 164 cues
- keyframes
- 25
- kept of 41 captured
- frames with text
- 25
- 309 lines read
- chapters
- 10
- from the source metadata
- keyframe bytes
- 4.3 MB
- word timings on 164 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:36 | 0s |
stt |
done | — | 2026-08-09 14:46 | 19s |
chunk |
done | — | 2026-08-09 14:47 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 1s |
keyframe |
done | — | 2026-08-09 14:47 | 3m 44s |
ocr |
done | — | 2026-08-09 14:50 | 16s |
frame_embed |
done | — | 2026-08-10 19:42 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer1.00
- World's Fair0.98
-
- July 1, 20260.98
- AlEngineer0.99
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- Data for running software companies1.00
- Emulated1.00
- World'sFair0.96
- Engineering the future of Al0.99
-
- July 1, 20260.98
- AlEngineer0.99
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- Data for running software companies1.00
- Emulated1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20260.99
- AlEngineer0.99
- World's Fair0.97
- The Team1.00
- PRESENTED BY1.00
- Microsoft1.00
- Joseph Wang0.96
- Sid Patllollu1.00
- CEO1.00
- CTO1.00
- World's Fair0.95
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining0.99
-
- Emulated1.00
- July 1, 20260.98
- AlEngineer0.98
- World'sFair1.00
- The model capability gap0.98
- is data gap1.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20260.97
- AlEngineer0.98
- World'sFair1.00
- How do we close the gap?0.97
- Unspecified1.00
- Open-ended1.00
- Complex1.00
- Like a PM, we give agents scope and1.00
- We simulate issues that only appear at1.00
- Agents handle multiple components0.99
- context to discover and act on issues.1.00
- scale like power loss, corrupted disks,0.99
- with live traffic, clusters,0.99
- This includes company context with0.99
- network failures, and clock skew.1.00
- software-defined networks, dozens of0.99
- historical decisions, past incidents,1.00
- Agents try out multiple solutions and0.99
- engineering tools, pipelines, and1.00
- and planned work but also customer0.99
- we build open-ended, deterministic1.00
- observability stacks.0.99
- conversations.1.00
- verifiers that accept all of them.0.99
- World's Fair0.96
- TRACK 9· JULY 1, 20260.96
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20261.00
- AlEngineer0.99
- World's Fair0.99
- Service Discovery App1.00
- Knowledge Tools1.00
- Uses etcd as a store for services0.98
- Tickets, postmortems,1.00
- alerts, runbooks1.00
- Read/write/watch/lease workload1.00
- 1. Agent establishes context1.00
- etcd gRPC proxy1.00
- etcd source code1.00
- Directs traffic to live nodes1.00
- 2. Agent implements1.00
- observability feature0.98
- Read/write/watch/lease workload1.00
- 4. Kicks off rolling deployment0.99
- Live Voters1.00
- Plant1.00
- Promote-Ready Learners0.99
- Stale Member1.00
- Leader1.00
- Follower1.00
- Follower1.00
- Lagging1.00
- learner1.00
- Learner1.00
- Learner1.00
- Learner1.00
- 3. Agent cleans up stale1.00
- leftover from a previous1.00
- membership record0.99
- migration1.00
- Replication1.00
- stream1.00
- 6. Agent promotes0.99
- healthy nodes only0.99
- Note: Only one node is the learner at a1.00
- time. Labels are for visual consistency0.99
- Prometheus/Grafana1.00
- 5. Uses metrics to identify0.99
- learners ready for promotion1.00
- World's Fair0.97
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20260.98
- AlEngineer0.98
- World's Fair0.97
- But this is not enough...0.99
- World's Fair0.97
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20260.98
- AlEngineer0.99
- World's Fair0.98
- My shiny1.00
- software1.00
- World's Fair0.96
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- Emulated1.00
- July 1, 20260.99
- AlEngineer0.99
- World'sFair1.00
- Host1.00
- Provisioning1.00
- Creates1.00
- Instance/Container1.00
- Frontend API0.99
- My shiny1.00
- software1.00
- World's Fair0.96
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
Transcript
164 cues· 2,316 words· 13,222 chars
- 0:14 So I appreciate the intro.
- 0:15 My name is Joseph, and this is my co-founder, Sid.
- 0:19 Emulated is a data lab focused on increasing the reliability and autonomy of AI agents.
- 0:26 And if you've been at AI Engineer and you've watched the talks, seen the tracks, then there's probably one takeaway that all the talks have in common.
- 0:36 And it's that we're headed towards a future where agents are able to perform useful work over longer and longer horizons with little to no supervision.
- 0:45 So today we're going to answer some of the questions of what this means for the data and model layers.
- 0:52 We're going to touch on some pretty cool things, so look out for them, like how to simulate a company within a sandbox or sandboxes for multi-node systems and distributed clusters.
- 1:04 And if we have a little bit of time, we'll also go into some of the work that we're doing with post-training pipelines and how these new types of sandboxes are affecting post-training infra as well.
- 1:17 So where Sid and I come from, our backgrounds are in network infra, distributed databases, and sandbox infra.
- 1:25 And these are all areas where the workloads are mission critical.
- 1:30 We all saw a couple months ago that when something like DynamoDB goes down, so does US East 1 and half the internet.
- 1:39 And working on these systems, we saw a model capability gap when it came to operating and building these systems at scale and thinking about the consequences of architecture and systems design over the course of years.
- 1:55 Yeah, so it led to a pretty natural question, right?
- 1:58 For such mission critical services, why is it that my model or my agent is so proficient at handling the application layer, but struggles when it comes to reasoning through infrastructure complexities, for example, things like MVCC on a database engine, which can lead to corruption issues, which was one of the roots of the DynamoDB failure a few months ago.
- 2:26 So like with everything in NML, the gap in models is usually a gap in data.
- 2:31 Models typically are only as good as data is.
- 2:37 And to really highlight this point, right, model capability has never regressed whenever you introduce more high-quality data.
- 2:46 So with that being said, what is the data gap then?
- 2:48 What does data look like right now?
- 2:51 And how is this influencing the model capability gap here?
- 2:55 So if you look at any of the frontier or recent benchmarks, like SuiteBench Pro, TerminalBench, or something like Frontier Code and DeepSuite,
- 3:05 The tasks only operate within the codebase.
- 3:09 The agent is given a pretty large task and over the course of 50 to 100 turns produces a couple thousand line PR.
- 3:22 but it doesn't do all of the work that a human does.
- 3:25 It doesn't do what a PM does with talking to customers, understanding their problems, what an engineer does with trying out different approaches, performance testing them, and owning the underlying infra for the code base over the course of not just months, but years.
- 3:44 And this is really the gap that we're closing.
- 3:47 We've taken software engineering companies and we've put them into containerized environments.
- 3:52 So this includes organizational contexts like projects, incidents, customer conversations.
- 4:01 The agent also has to deal with issues that only appear at scale, like network failures between distributed nodes, data corruption, and clock skew.
- 4:10 And through all this, we also want the agents to reason about orchestrating through distributed clusters and also thinking about things like operational blast radius while solving live traffic.
- 4:21 And the result is that the task that these agents have to complete or we want the agents to learn is that environments are far more complex and long horizon than a simple code disk.
- 4:35 So let's just, let's bring a picture into the mix because it tends to make things more interesting.
- 4:42 Here's an example we've built of an etcd consensus cluster that a typical production service might rely on.
- 4:49 So an early environment might, tended to operate and work primarily on that little blue square entitled etcd source code in the bottom right there.
- 5:00 But a lot of the fun and the model capability gap that results from it is really in everything that surrounds it.
- 5:08 So you start with the tickets, projects, postmortems.
- 5:12 What are the train wrecks?
- 5:14 Why did they happen?
- 5:15 How did customers feel about them?
- 5:17 And oftentimes those aren't necessarily up to date.
- 5:21 The agent has to incorporate all that when it's reasoning through the actual change that
- 5:28 current environments have it make.
- 5:30 After it makes that change, you need to kick off rolling deployments.
- 5:35 Those deployment systems can oftentimes be complicated, have conflicts, may not work.
- 5:42 And all through that, when you're finally migrating off of from old hardware onto new hardware, you run into unforeseen problems, which you
- 5:52 did not, the agent has to reason through in real time, just like a human would.
- 5:58 You have failing nodes.
- 6:01 You have stale, deprecated nodes.
- 6:04 And while all of this is happening, the service can't go down.
- 6:09 Because there is a blast radius to serving live traffic.
loading
Chapters
- 0:00 Useful work over longer horizons
- 1:20 Backgrounds in network infrastructure
- 2:26 How environments shape capability
- 3:16 Fifty to a hundred turn tasks
- 4:59 Why real incidents are messy
- 7:11 Real infrastructure isn't a code diff
- 7:40 Acting as an engineer inside the cloud
- 9:37 Deployment, cost, and scaling bars
- 13:29 Why it's called Emulated
- 15:01 Simulating full companies