read-only demo

Videos OqM67QG_Ikk

From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI

index_state ready data_status ok

AI Engineer· published 2026-07-13· 0:44:33· en-US· indexed 2026-08-10 19:48

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:15, 1 of 1 keyframes kept
  5. Shot 4, 0:15 to 0:27, 1 of 1 keyframes kept
  6. Shot 5, 0:27 to 1:14, 1 of 1 keyframes kept
  7. Shot 6, 1:14 to 1:29, 1 of 1 keyframes kept
  8. Shot 7, 1:29 to 1:57, 1 of 1 keyframes kept
  9. Shot 8, 1:57 to 2:25, 0 of 1 keyframes kept
  10. Shot 9, 2:25 to 2:52, 0 of 1 keyframes kept
  11. Shot 10, 2:52 to 3:20, 0 of 1 keyframes kept
  12. Shot 11, 3:20 to 3:48, 0 of 1 keyframes kept
  13. Shot 12, 3:48 to 4:16, 0 of 1 keyframes kept
  14. Shot 13, 4:16 to 4:43, 0 of 1 keyframes kept
  15. Shot 14, 4:43 to 5:11, 1 of 1 keyframes kept
  16. Shot 15, 5:11 to 5:39, 0 of 1 keyframes kept
  17. Shot 16, 5:39 to 6:11, 1 of 1 keyframes kept
  18. Shot 17, 6:11 to 6:43, 0 of 1 keyframes kept
  19. Shot 18, 6:43 to 7:08, 1 of 1 keyframes kept
  20. Shot 19, 7:08 to 7:33, 0 of 1 keyframes kept
  21. Shot 20, 7:33 to 7:58, 0 of 1 keyframes kept
  22. Shot 21, 7:58 to 8:23, 0 of 1 keyframes kept
  23. Shot 22, 8:23 to 9:04, 1 of 1 keyframes kept
  24. Shot 23, 9:04 to 9:44, 1 of 1 keyframes kept
  25. Shot 24, 9:44 to 10:18, 1 of 1 keyframes kept
  26. Shot 25, 10:18 to 10:52, 0 of 1 keyframes kept
  27. Shot 26, 10:52 to 11:26, 1 of 1 keyframes kept
  28. Shot 27, 11:26 to 12:01, 0 of 1 keyframes kept
  29. Shot 28, 12:01 to 12:43, 1 of 1 keyframes kept
  30. Shot 29, 12:43 to 13:10, 1 of 1 keyframes kept
  31. Shot 30, 13:10 to 13:38, 0 of 1 keyframes kept
  32. Shot 31, 13:38 to 14:17, 1 of 1 keyframes kept
  33. Shot 32, 14:17 to 14:51, 1 of 1 keyframes kept
  34. Shot 33, 14:51 to 15:21, 1 of 1 keyframes kept
  35. Shot 34, 15:21 to 15:50, 0 of 1 keyframes kept
  36. Shot 35, 15:50 to 16:05, 1 of 1 keyframes kept
  37. Shot 36, 16:05 to 16:25, 1 of 1 keyframes kept
  38. Shot 37, 16:25 to 16:55, 1 of 1 keyframes kept
  39. Shot 38, 16:55 to 17:25, 0 of 1 keyframes kept
  40. Shot 39, 17:25 to 17:31, 1 of 1 keyframes kept
  41. Shot 40, 17:31 to 18:03, 0 of 1 keyframes kept
  42. Shot 41, 18:03 to 18:15, 0 of 1 keyframes kept
  43. Shot 42, 18:15 to 18:57, 1 of 1 keyframes kept
  44. Shot 43, 18:57 to 19:29, 1 of 1 keyframes kept
  45. Shot 44, 19:29 to 20:02, 0 of 1 keyframes kept
  46. Shot 45, 20:02 to 20:29, 1 of 1 keyframes kept
  47. Shot 46, 20:29 to 20:56, 1 of 1 keyframes kept
  48. Shot 47, 20:56 to 21:23, 0 of 1 keyframes kept
  49. Shot 48, 21:23 to 21:50, 0 of 1 keyframes kept
  50. Shot 49, 21:50 to 22:17, 0 of 1 keyframes kept
  51. Shot 50, 22:17 to 22:44, 0 of 1 keyframes kept
  52. Shot 51, 22:44 to 23:11, 0 of 1 keyframes kept
  53. Shot 52, 23:11 to 23:37, 1 of 1 keyframes kept
  54. Shot 53, 23:37 to 24:03, 0 of 1 keyframes kept
  55. Shot 54, 24:03 to 24:28, 0 of 1 keyframes kept
  56. Shot 55, 24:28 to 25:07, 1 of 1 keyframes kept
  57. Shot 56, 25:07 to 25:08, 1 of 1 keyframes kept
  58. Shot 57, 25:08 to 25:11, 0 of 1 keyframes kept
  59. Shot 58, 25:11 to 25:41, 0 of 1 keyframes kept
  60. Shot 59, 25:41 to 26:23, 1 of 1 keyframes kept
  61. Shot 60, 26:23 to 27:13, 1 of 1 keyframes kept
  62. Shot 61, 27:13 to 27:38, 1 of 1 keyframes kept
  63. Shot 62, 27:38 to 28:03, 0 of 1 keyframes kept
  64. Shot 63, 28:03 to 28:28, 0 of 1 keyframes kept
  65. Shot 64, 28:28 to 28:54, 0 of 1 keyframes kept
  66. Shot 65, 28:54 to 29:19, 0 of 1 keyframes kept
  67. Shot 66, 29:19 to 29:44, 0 of 1 keyframes kept
  68. Shot 67, 29:44 to 29:45, 1 of 1 keyframes kept
  69. Shot 68, 29:45 to 30:35, 1 of 1 keyframes kept
  70. Shot 69, 30:35 to 31:06, 1 of 1 keyframes kept
  71. Shot 70, 31:06 to 31:38, 0 of 1 keyframes kept
  72. Shot 71, 31:38 to 32:21, 1 of 1 keyframes kept
  73. Shot 72, 32:21 to 32:58, 1 of 1 keyframes kept
  74. Shot 73, 32:58 to 33:40, 1 of 1 keyframes kept
  75. Shot 74, 33:40 to 34:05, 1 of 1 keyframes kept
  76. Shot 75, 34:05 to 34:30, 0 of 1 keyframes kept
  77. Shot 76, 34:30 to 34:56, 0 of 1 keyframes kept
  78. Shot 77, 34:56 to 35:27, 1 of 1 keyframes kept
  79. Shot 78, 35:27 to 35:58, 0 of 1 keyframes kept
  80. Shot 79, 35:58 to 36:44, 1 of 1 keyframes kept
  81. Shot 80, 36:44 to 37:24, 1 of 1 keyframes kept
  82. Shot 81, 37:24 to 37:51, 1 of 1 keyframes kept
  83. Shot 82, 37:51 to 38:31, 1 of 1 keyframes kept
  84. Shot 83, 38:31 to 39:08, 1 of 1 keyframes kept
  85. Shot 84, 39:08 to 39:45, 0 of 1 keyframes kept
  86. Shot 85, 39:45 to 40:03, 1 of 1 keyframes kept
  87. Shot 86, 40:03 to 40:30, 1 of 1 keyframes kept
  88. Shot 87, 40:30 to 40:58, 0 of 1 keyframes kept
  89. Shot 88, 40:58 to 41:38, 0 of 1 keyframes kept
  90. Shot 89, 41:38 to 42:24, 1 of 1 keyframes kept
  91. Shot 90, 42:24 to 42:49, 1 of 1 keyframes kept
  92. Shot 91, 42:49 to 43:14, 0 of 1 keyframes kept
  93. Shot 92, 43:14 to 44:02, 1 of 1 keyframes kept
  94. Shot 93, 44:02 to 44:16, 1 of 1 keyframes kept
  95. Shot 94, 44:16 to 44:33, 0 of 1 keyframes kept

95 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
520
whisperx 520
chunks
80
from 520 cues
keyframes
53
kept of 95 captured
frames with text
53
1,439 lines read
chapters
21
from the source metadata
keyframe bytes
11.1 MB
word timings on 520 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 11:27 2m 08s
stt done 2026-08-10 11:29 49s
chunk done 2026-08-10 11:30 0s
text_embed done 2026-08-10 19:48 1s
keyframe done 2026-08-10 11:30 4m 04s
ocr done 2026-08-10 11:34 24s
frame_embed done 2026-08-10 19:48 9s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 451.4

    1. AlEngineer0.95
    2. World's Fair0.98
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 657.9

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2727.9

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.97
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:14 #3 done2 line(s)

    shot 3·sharpness 151.0

    1. AlEngineer1.00
    2. World's Fair0.99
  • 0:17 #4 done2 line(s)

    shot 4·sharpness 181.4

    1. AlEngineer0.97
    2. World's Fair1.00
  • 1:09 #5 done14 line(s)

    shot 5·sharpness 2415.7

    1. OpenAI0.93
    2. AlEngineer0.98
    3. World'sFair1.00
    4. From fork() to fleet:1.00
    5. designing an agent1.00
    6. sandbox cloud1.00
    7. PRESENTED·BY0.97
    8. Microsoft1.00
    9. Abhishek Bhardwaj1.00
    10. Member of Technical Staff,1.00
    11. Reinforcement Learning and Agent Infrastructure1.00
    12. TRACK 1· JULY 1, 20260.94
    13. World's Fair0.96
    14. Sandbox & Platform Engineering0.99
  • 1:20 #6 done9 line(s)

    shot 6·sharpness 1894.7

    1. AlEngineer0.99
    2. Last years' AIE talk0.98
    3. World'sFair1.00
    4. Arrakis: how to build an Al sandbox from scratch1.00
    5. PRESENTEDBY1.00
    6. Microsoft1.00
    7. TRACK 1· JULY 1, 20260.95
    8. World'sFair1.00
    9. Sandbox & Platform Engineering1.00
  • 1:43 #7 done32 line(s)

    shot 7·sharpness 2770.5

    1. AlEngineer0.99
    2. WHYSANDBOXES·TRAINING1.00
    3. World'sFair1.00
    4. Why the explosion of agent sandboxes?1.00
    5. Reasoning models can emit a tool and its parameters. A harness executes the call;0.99
    6. a grader scores the result; training adjusts the weights.1.00
    7. PRESENTEDBY1.00
    8. Update weights1.00
    9. Microsoft1.00
    10. RL framework0.96
    11. Task1.00
    12. Harness1.00
    13. Call model1.00
    14. Model1.00
    15. training loop1.00
    16. calls model and tools1.00
    17. Response / tool call0.98
    18. emits tool +1.00
    19. parameters1.00
    20. Reward1.00
    21. Result1.00
    22. Execute tool call1.00
    23. Observation1.00
    24. Grader1.00
    25. Tool runtime1.00
    26. grades result1.00
    27. sandbox executes the call1.00
    28. Where and how does this run?0.99
    29. OpenAl0.96
    30. TRACK 1· JULY 1,20260.95
    31. World's Fair0.96
    32. Sandbox & Platform Engineering0.99
  • 2:08 #8 skipped

    shot 8·duplicate of #7

  • 2:36 #9 skipped

    shot 9·duplicate of #7

  • 3:17 #10 skipped

    shot 10·duplicate of #7

  • 3:45 #11 skipped

    shot 11·duplicate of #7

  • 4:12 #12 skipped

    shot 12·duplicate of #7

  • 4:30 #13 skipped

    shot 13·duplicate of #7

  • 5:08 #14 done29 line(s)

    shot 14·sharpness 2496.6

    1. AlEngineer0.99
    2. WHY SANDBOXES·INFERENCE0.98
    3. Why the explosion of agent sandboxes?1.00
    4. World's Fair0.99
    5. At inference time, tool calls are still model output. The harness executes0.99
    6. them on the model's behalf and returns the observation.1.00
    7. PRESENTED·BY0.95
    8. Microsoft1.00
    9. Task1.00
    10. Task1.00
    11. Harness1.00
    12. Call model1.00
    13. Model1.00
    14. user request + context0.97
    15. Result1.00
    16. calls model and tools1.00
    17. Response / tool call0.98
    18. returns answer or0.98
    19. tool call1.00
    20. Parse and execute1.00
    21. tool call_0.91
    22. Observation1.00
    23. Tool runtime1.00
    24. sandbox executes the call1.00
    25. Where and how does this run?1.00
    26. OpenAl0.97
    27. TRACK 1· JULY 1,20260.96
    28. World's Fair0.99
    29. Sandbox & Platform Engineering1.00
  • 5:33 #15 skipped

    shot 15·duplicate of #14

  • 6:04 #16 done18 line(s)

    shot 16·sharpness 2051.3

    1. AlEngineer0.99
    2. AGENT CLOUDS · LOCAL TO CLOUD0.98
    3. World'sFair1.00
    4. OpenClaw:1.00
    5. a peek into agent clouds1.00
    6. OpenClaws1.00
    7. Keep your agents1.00
    8. running 24/7.1.00
    9. A precision claw that keeps your Mac0.98
    10. slightly open so background agents can stay alive.1.00
    11. PRESENTED BY0.96
    12. Only $1990.89
    13. Microsoft1.00
    14. Pre-order now0.99
    15. OpenAl0.96
    16. TRACK 1• JULY 1, 20260.97
    17. World's Fair0.96
    18. Sandbox & Platform Engineering0.98
  • 6:39 #17 skipped

    shot 17·duplicate of #16

  • 7:01 #18 done32 line(s)

    shot 18·sharpness 2013.1

    1. AGENT CLOUDS · LOCAL TO CLOUD0.95
    2. AlEngineer0.99
    3. Research vs. product sandbox platforms1.00
    4. World's Fair0.99
    5. Research1.00
    6. Product1.00
    7. PRESENTEDBY1.00
    8. Primary optimization1.00
    9. Throughput1.00
    10. Latency1.00
    11. Microsoft1.00
    12. Why1.00
    13. Many parallel rollouts1.00
    14. People see results fast0.99
    15. Keep GPU compute utilized1.00
    16. Responsive product experience1.00
    17. Reliability1.00
    18. Avoid wasted tokens1.00
    19. Paramount1.00
    20. Avoid restarting rollouts1.00
    21. Unreliable products fail1.00
    22. Prevent generated code from gaining root0.98
    23. Protect the host and platform resources0.99
    24. Security1.00
    25. Protect internal resources1.00
    26. Isolate other production resources0.99
    27. Resist reward hacking0.99
    28. Protect user data1.00
    29. OpenAl0.96
    30. TRACK 1· JULY 1, 20260.95
    31. World's Fair0.99
    32. Sandbox & Platform Engineering0.99
  • 7:28 #19 skipped

    shot 19·duplicate of #18

  • 7:53 #20 skipped

    shot 20·duplicate of #18

  • 8:11 #21 skipped

    shot 21·duplicate of #18

  • 8:28 #22 done27 line(s)

    shot 22·sharpness 2296.0

    1. AlEngineer1.00
    2. AGENT PLATFORM1.00
    3. Runtime, persistence and orchestration1.00
    4. World's Fair0.99
    5. AGENT SANDBOX CLOUD0.98
    6. PRESENTED·BY0.96
    7. Microsoft1.00
    8. RUNTIME1.00
    9. PERSISTENCE1.00
    10. ORCHESTRATION1.00
    11. HOW SANDBOXES RUN0.99
    12. HOW AGENT STATE SURVIVES1.00
    13. WHERE SANDBOXES RUN0.99
    14. SECURE EXECUTION0.98
    15. SAVE SANDBOX STATE0.99
    16. CLUSTER SELECTION1.00
    17. LOW-LATENCY CREATION1.00
    18. LONG-RUNNING TASKS1.00
    19. NODE SELECTION1.00
    20. LOW-LATENCY TOOL EXECUTION1.00
    21. RELIABLE RECOVERY1.00
    22. LOW-LATENCY CREATES1.00
    23. SECURE RUNTIME· DURABLE STATE · FLEET-SCALE PLACEMENT0.97
    24. OpenAl0.96
    25. TRACK 1· JULY 1,20260.95
    26. World's Fair0.99
    27. Sandbox & Platform Engineering0.99
  • 9:32 #23 done28 line(s)

    shot 23·sharpness 2211.0

    1. AlEngineer0.99
    2. RUNTIME1.00
    3. World'sFair1.00
    4. Runtime from first principles:1.00
    5. linux execution model1.00
    6. Userspace1.00
    7. PRESENTED BY0.96
    8. Process 10.98
    9. Process 20.97
    10. Program instructions1.00
    11. Microsoft1.00
    12. Thread1.00
    13. Thread1.00
    14. mov eax, 11.00
    15. xor ebx, ebx1.00
    16. in 0x800.98
    17. task_struct1.00
    18. task_struct1.00
    19. ioctl0.99
    20. syscall0.99
    21. /dev/kvm1.00
    22. /dev/sda11.00
    23. Linux kernel-privileged access1.00
    24. Kernel gates access to hardware0.99
    25. OpenAl0.97
    26. TRACK 1· JULY 1, 20260.95
    27. World's Fair0.99
    28. Sandbox & Platform Engineering1.00

Transcript

520 cues· 7,278 words· 39,152 chars

  1. 0:12 Welcome, everyone.
  2. 0:13 Can you guys hear me OK?
  3. 0:15 I've been standing here for 15 minutes without saying anything, so we can start now.
  4. 0:20 My name is Abhishek.
  5. 0:21 I'm on the RL and agent infrastructure team at OpenAI.
  6. 0:25 What that means is we work on the infra for reinforcement learning, specifically.
  7. 0:30 And on the product side, we also develop infra that helps run untrusted code as part of ChatGPT, CodexWeb, securely and reliably at scale.
  8. 0:40 This talk is called From Fork to Fleet, Designing an Agent Sandbox Cloud.
  9. 0:46 I'll be very clear that there are a lot of words in the title that have OS and infra concepts, but this is a first principles talk.
  10. 0:53 So we'll cover what sandboxes are and why they are needed from first principles.
  11. 0:58 We will also try to cover design intuitions around designing an agent sandbox cloud to run sandboxes securely and reliably at scale.
  12. 1:07 So if some of these words don't mean anything, don't be worried.
  13. 1:10 We'll explain from first principles and go from there.
  14. 1:16 Last year, I gave a talk called How to Build an AI Sandbox from Scratch.
  15. 1:21 If you're interested, you can look at that talk as well.
  16. 1:24 Think of this as a spiritual sequel to that talk.
  17. 1:31 OK, so now let's forget about sandboxes or clouds for a second.
  18. 1:34 Let's just go back in time.
  19. 1:36 ChatGPT came out.
  20. 1:37 It's a very large pre-trained model.
  21. 1:40 People ask all sorts of questions, and it responds really, really well, really, really human-like answers.
  22. 1:45 But when people ask questions like, what is 3 plus 3, or how many R's in strawberry, sometimes it works and sometimes it doesn't.
  23. 1:54 It answers questions like what is 3 plus 3 quite well, because it's trained on the entire internet.
  24. 2:00 And apparently, people have written 3 plus 3 equal to 6 many, many times on the internet.
  25. 2:06 So it gets it right.
  26. 2:07 But people haven't asked how many hours in Strawberry enough on the internet.
  27. 2:11 And so it gets it wrong.
  28. 2:14 And so it's obvious that for anything code or math related or any problem which has a verifiable reward, which means that it can be tested whether it's true or false, the model needs something more.
  29. 2:26 And the key unlock was that given the model tool calling capability or a way to execute code, the model gets these verifiable reward questions around code and math correctly.
  30. 2:39 And that's when first principles came that if we give the models the ability to execute code, it can hill climb and be very, very good at math, code, and other domains that have verifiable rewards.
  31. 2:51 Cut to 2026, and we are seeing the consequences of doing that at scale.
  32. 2:58 So now, how can it answer what is 3 plus 3?
  33. 3:01 And how can it answer how many R's in strawberry?
  34. 3:04 Well, it can write code to do this.
  35. 3:06 So if you see the diagram, we have a training loop.
  36. 3:10 And the training loop gives it tasks or questions.
  37. 3:13 And then the training and the harness then parses the response of the model.
  38. 3:17 And the response might say, hey, execute code on my behalf.
  39. 3:22 The harness is responsible for executing the code.
  40. 3:25 And then a grader judges whether the answer is correct or not.
  41. 3:28 And then the training loop backdrops and changes the weights.
  42. 3:31 And that's how we train it to do two things.
  43. 3:34 We train it to call code execution or tools on certain classes of problems.
  44. 3:39 And secondly, we ensure that the code it executes actually solves the problem.
  45. 3:44 So this is why pool calling is important on the training side.
  46. 3:51 Now let's talk about the product side.
  47. 3:54 In the previous slide, we showed how the models are trained to emit code in order to solve certain tasks.
  48. 4:01 The agent executes the code, and we verify the reward.
  49. 4:04 Well, all of this is useless if we don't support this model on the product side.
  50. 4:08 So it's basically the same slide as before, but we don't have a training loop.

Chapters

  1. 0:00 Introduction and motivation for AI agent sandboxes
  2. 1:31 Why AI models need tools and execution environments
  3. 3:51 Product-side challenges: Security and the need for sandboxing
  4. 6:44 Comparing research vs. product sandbox requirements
  5. 8:24 Overview of the three pillars: Runtime, Persistence, and Orchestration
  6. 9:05 First principles of Linux execution: System calls and security vectors
  7. 11:15 Evaluating fork() and exec models
  8. 12:06 Understanding containers: Namespaces and cgroups
  9. 16:26 GVisor as an application kernel alternative
  10. 18:29 Hardware-level virtualization (Virtual Machines)
  11. 20:34 How VMMs (Virtual Machine Monitors) work with KVM
  12. 23:16 Evolution of modern VMMs and Rust-based safety
  13. 24:32 What defines a "microVM"?
  14. 25:43 Orchestrating microVMs via APIs
  15. 27:16 Trade-offs of microVMs (performance vs. security)
  16. 30:05 The need for persistent storage in agent sandboxes
  17. 31:40 Use cases for persistence: Reliability, long-running tasks, and research
  18. 34:36 Design choices for disk snapshotting
  19. 36:03 First principles of Linux block storage and file systems
  20. 37:25 Implementing always-on vs. explicit persistence
  21. 41:20 Scaling and orchestrating sandboxes at fleet level

Open at this second