Videos OqM67QG_Ikk
From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
Scene timeline
95 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 520
- whisperx 520
- chunks
- 80
- from 520 cues
- keyframes
- 53
- kept of 95 captured
- frames with text
- 53
- 1,439 lines read
- chapters
- 21
- from the source metadata
- keyframe bytes
- 11.1 MB
- word timings on 520 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 11:27 | 2m 08s |
stt |
done | — | 2026-08-10 11:29 | 49s |
chunk |
done | — | 2026-08-10 11:30 | 0s |
text_embed |
done | — | 2026-08-10 19:48 | 1s |
keyframe |
done | — | 2026-08-10 11:30 | 4m 04s |
ocr |
done | — | 2026-08-10 11:34 | 24s |
frame_embed |
done | — | 2026-08-10 19:48 | 9s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.98
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.97
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer1.00
- World's Fair0.99
-
- AlEngineer0.97
- World's Fair1.00
-
- OpenAI0.93
- AlEngineer0.98
- World'sFair1.00
- From fork() to fleet:1.00
- designing an agent1.00
- sandbox cloud1.00
- PRESENTED·BY0.97
- Microsoft1.00
- Abhishek Bhardwaj1.00
- Member of Technical Staff,1.00
- Reinforcement Learning and Agent Infrastructure1.00
- TRACK 1· JULY 1, 20260.94
- World's Fair0.96
- Sandbox & Platform Engineering0.99
-
- AlEngineer0.99
- Last years' AIE talk0.98
- World'sFair1.00
- Arrakis: how to build an Al sandbox from scratch1.00
- PRESENTEDBY1.00
- Microsoft1.00
- TRACK 1· JULY 1, 20260.95
- World'sFair1.00
- Sandbox & Platform Engineering1.00
-
- AlEngineer0.99
- WHYSANDBOXES·TRAINING1.00
- World'sFair1.00
- Why the explosion of agent sandboxes?1.00
- Reasoning models can emit a tool and its parameters. A harness executes the call;0.99
- a grader scores the result; training adjusts the weights.1.00
- PRESENTEDBY1.00
- Update weights1.00
- Microsoft1.00
- RL framework0.96
- Task1.00
- Harness1.00
- Call model1.00
- Model1.00
- training loop1.00
- calls model and tools1.00
- Response / tool call0.98
- emits tool +1.00
- parameters1.00
- Reward1.00
- Result1.00
- Execute tool call1.00
- Observation1.00
- Grader1.00
- Tool runtime1.00
- grades result1.00
- sandbox executes the call1.00
- Where and how does this run?0.99
- OpenAl0.96
- TRACK 1· JULY 1,20260.95
- World's Fair0.96
- Sandbox & Platform Engineering0.99
-
- AlEngineer0.99
- WHY SANDBOXES·INFERENCE0.98
- Why the explosion of agent sandboxes?1.00
- World's Fair0.99
- At inference time, tool calls are still model output. The harness executes0.99
- them on the model's behalf and returns the observation.1.00
- PRESENTED·BY0.95
- Microsoft1.00
- Task1.00
- Task1.00
- Harness1.00
- Call model1.00
- Model1.00
- user request + context0.97
- Result1.00
- calls model and tools1.00
- Response / tool call0.98
- returns answer or0.98
- tool call1.00
- Parse and execute1.00
- tool call_0.91
- Observation1.00
- Tool runtime1.00
- sandbox executes the call1.00
- Where and how does this run?1.00
- OpenAl0.97
- TRACK 1· JULY 1,20260.96
- World's Fair0.99
- Sandbox & Platform Engineering1.00
-
- AlEngineer0.99
- AGENT CLOUDS · LOCAL TO CLOUD0.98
- World'sFair1.00
- OpenClaw:1.00
- a peek into agent clouds1.00
- OpenClaws1.00
- Keep your agents1.00
- running 24/7.1.00
- A precision claw that keeps your Mac0.98
- slightly open so background agents can stay alive.1.00
- PRESENTED BY0.96
- Only $1990.89
- Microsoft1.00
- Pre-order now0.99
- OpenAl0.96
- TRACK 1• JULY 1, 20260.97
- World's Fair0.96
- Sandbox & Platform Engineering0.98
-
- AGENT CLOUDS · LOCAL TO CLOUD0.95
- AlEngineer0.99
- Research vs. product sandbox platforms1.00
- World's Fair0.99
- Research1.00
- Product1.00
- PRESENTEDBY1.00
- Primary optimization1.00
- Throughput1.00
- Latency1.00
- Microsoft1.00
- Why1.00
- Many parallel rollouts1.00
- People see results fast0.99
- Keep GPU compute utilized1.00
- Responsive product experience1.00
- Reliability1.00
- Avoid wasted tokens1.00
- Paramount1.00
- Avoid restarting rollouts1.00
- Unreliable products fail1.00
- Prevent generated code from gaining root0.98
- Protect the host and platform resources0.99
- Security1.00
- Protect internal resources1.00
- Isolate other production resources0.99
- Resist reward hacking0.99
- Protect user data1.00
- OpenAl0.96
- TRACK 1· JULY 1, 20260.95
- World's Fair0.99
- Sandbox & Platform Engineering0.99
-
- AlEngineer1.00
- AGENT PLATFORM1.00
- Runtime, persistence and orchestration1.00
- World's Fair0.99
- AGENT SANDBOX CLOUD0.98
- PRESENTED·BY0.96
- Microsoft1.00
- RUNTIME1.00
- PERSISTENCE1.00
- ORCHESTRATION1.00
- HOW SANDBOXES RUN0.99
- HOW AGENT STATE SURVIVES1.00
- WHERE SANDBOXES RUN0.99
- SECURE EXECUTION0.98
- SAVE SANDBOX STATE0.99
- CLUSTER SELECTION1.00
- LOW-LATENCY CREATION1.00
- LONG-RUNNING TASKS1.00
- NODE SELECTION1.00
- LOW-LATENCY TOOL EXECUTION1.00
- RELIABLE RECOVERY1.00
- LOW-LATENCY CREATES1.00
- SECURE RUNTIME· DURABLE STATE · FLEET-SCALE PLACEMENT0.97
- OpenAl0.96
- TRACK 1· JULY 1,20260.95
- World's Fair0.99
- Sandbox & Platform Engineering0.99
-
- AlEngineer0.99
- RUNTIME1.00
- World'sFair1.00
- Runtime from first principles:1.00
- linux execution model1.00
- Userspace1.00
- PRESENTED BY0.96
- Process 10.98
- Process 20.97
- Program instructions1.00
- Microsoft1.00
- Thread1.00
- Thread1.00
- mov eax, 11.00
- xor ebx, ebx1.00
- in 0x800.98
- task_struct1.00
- task_struct1.00
- ioctl0.99
- syscall0.99
- /dev/kvm1.00
- /dev/sda11.00
- Linux kernel-privileged access1.00
- Kernel gates access to hardware0.99
- OpenAl0.97
- TRACK 1· JULY 1, 20260.95
- World's Fair0.99
- Sandbox & Platform Engineering1.00
Transcript
520 cues· 7,278 words· 39,152 chars
- 0:12 Welcome, everyone.
- 0:13 Can you guys hear me OK?
- 0:15 I've been standing here for 15 minutes without saying anything, so we can start now.
- 0:20 My name is Abhishek.
- 0:21 I'm on the RL and agent infrastructure team at OpenAI.
- 0:25 What that means is we work on the infra for reinforcement learning, specifically.
- 0:30 And on the product side, we also develop infra that helps run untrusted code as part of ChatGPT, CodexWeb, securely and reliably at scale.
- 0:40 This talk is called From Fork to Fleet, Designing an Agent Sandbox Cloud.
- 0:46 I'll be very clear that there are a lot of words in the title that have OS and infra concepts, but this is a first principles talk.
- 0:53 So we'll cover what sandboxes are and why they are needed from first principles.
- 0:58 We will also try to cover design intuitions around designing an agent sandbox cloud to run sandboxes securely and reliably at scale.
- 1:07 So if some of these words don't mean anything, don't be worried.
- 1:10 We'll explain from first principles and go from there.
- 1:16 Last year, I gave a talk called How to Build an AI Sandbox from Scratch.
- 1:21 If you're interested, you can look at that talk as well.
- 1:24 Think of this as a spiritual sequel to that talk.
- 1:31 OK, so now let's forget about sandboxes or clouds for a second.
- 1:34 Let's just go back in time.
- 1:36 ChatGPT came out.
- 1:37 It's a very large pre-trained model.
- 1:40 People ask all sorts of questions, and it responds really, really well, really, really human-like answers.
- 1:45 But when people ask questions like, what is 3 plus 3, or how many R's in strawberry, sometimes it works and sometimes it doesn't.
- 1:54 It answers questions like what is 3 plus 3 quite well, because it's trained on the entire internet.
- 2:00 And apparently, people have written 3 plus 3 equal to 6 many, many times on the internet.
- 2:06 So it gets it right.
- 2:07 But people haven't asked how many hours in Strawberry enough on the internet.
- 2:11 And so it gets it wrong.
- 2:14 And so it's obvious that for anything code or math related or any problem which has a verifiable reward, which means that it can be tested whether it's true or false, the model needs something more.
- 2:26 And the key unlock was that given the model tool calling capability or a way to execute code, the model gets these verifiable reward questions around code and math correctly.
- 2:39 And that's when first principles came that if we give the models the ability to execute code, it can hill climb and be very, very good at math, code, and other domains that have verifiable rewards.
- 2:51 Cut to 2026, and we are seeing the consequences of doing that at scale.
- 2:58 So now, how can it answer what is 3 plus 3?
- 3:01 And how can it answer how many R's in strawberry?
- 3:04 Well, it can write code to do this.
- 3:06 So if you see the diagram, we have a training loop.
- 3:10 And the training loop gives it tasks or questions.
- 3:13 And then the training and the harness then parses the response of the model.
- 3:17 And the response might say, hey, execute code on my behalf.
- 3:22 The harness is responsible for executing the code.
- 3:25 And then a grader judges whether the answer is correct or not.
- 3:28 And then the training loop backdrops and changes the weights.
- 3:31 And that's how we train it to do two things.
- 3:34 We train it to call code execution or tools on certain classes of problems.
- 3:39 And secondly, we ensure that the code it executes actually solves the problem.
- 3:44 So this is why pool calling is important on the training side.
- 3:51 Now let's talk about the product side.
- 3:54 In the previous slide, we showed how the models are trained to emit code in order to solve certain tasks.
- 4:01 The agent executes the code, and we verify the reward.
- 4:04 Well, all of this is useless if we don't support this model on the product side.
- 4:08 So it's basically the same slide as before, but we don't have a training loop.
loading
Chapters
- 0:00 Introduction and motivation for AI agent sandboxes
- 1:31 Why AI models need tools and execution environments
- 3:51 Product-side challenges: Security and the need for sandboxing
- 6:44 Comparing research vs. product sandbox requirements
- 8:24 Overview of the three pillars: Runtime, Persistence, and Orchestration
- 9:05 First principles of Linux execution: System calls and security vectors
- 11:15 Evaluating fork() and exec models
- 12:06 Understanding containers: Namespaces and cgroups
- 16:26 GVisor as an application kernel alternative
- 18:29 Hardware-level virtualization (Virtual Machines)
- 20:34 How VMMs (Virtual Machine Monitors) work with KVM
- 23:16 Evolution of modern VMMs and Rust-based safety
- 24:32 What defines a "microVM"?
- 25:43 Orchestrating microVMs via APIs
- 27:16 Trade-offs of microVMs (performance vs. security)
- 30:05 The need for persistent storage in agent sandboxes
- 31:40 Use cases for persistence: Reliability, long-running tasks, and research
- 34:36 Design choices for disk snapshotting
- 36:03 First principles of Linux block storage and file systems
- 37:25 Implementing always-on vs. explicit persistence
- 41:20 Scaling and orchestrating sandboxes at fleet level