Videos shRR1e2HXMk
Codex, Behind the Harness — Dominik Kundel, OpenAI
Scene timeline
103 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 187
- whisperx 187
- chunks
- 37
- from 187 cues
- keyframes
- 66
- kept of 103 captured
- frames with text
- 66
- 1,698 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 13.5 MB
- word timings on 187 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 23:52 | 1m 19s |
stt |
done | — | 2026-08-10 23:53 | 22s |
chunk |
done | — | 2026-08-10 23:54 | 0s |
text_embed |
done | — | 2026-08-10 23:54 | 1s |
keyframe |
done | — | 2026-08-10 23:54 | 3m 01s |
ocr |
done | — | 2026-08-10 23:57 | 41s |
frame_embed |
done | — | 2026-08-10 23:57 | 11s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- OpenAl0.93
- AlEngineer0.98
- World'sFair1.00
- Codex: Behind the harness0.99
- PRESENTED BY1.00
- Microsoft1.00
- Dominik Kundel1.00
- Developer Experience1.00
- World's Fair0.97
- Engineering the future of Al1.00
-
- Agent loop1.00
- AlEngineer0.97
- World'sFair1.00
- ASSEMBLE CONVERSATION HISTORY0.99
- PRESENTED BY1.00
- LATEST USER INPUT0.97
- Microsoft1.00
- MODELINFERENCE1.00
- AGENT RESPONSE0.99
- TOOL CALLS0.99
- OpenAl0.96
- 5 / 550.85
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering1.00
-
- AlEngineer0.99
- World'sFair1.00
- PRESENTED BY1.00
- Do anything1.00
- Microsoft1.00
- +0.58
- Approve for me0.94
- 5.5 Extra High0.97
- Choose project1.00
- OpenAl0.95
- 6 / 550.90
- World's Fair0.96
- TRACK 8· JULY 2, 20260.94
- Agentic Engineering1.00
-
- AlEngineer0.94
- World's Fair0.98
- PlatformSolutionsResourcesOpen Source EnterprisePricing0.98
- O. Search or jump to...0.93
- Sign in0.99
- Sign up0.98
- openai/codex Public0.90
- Notifications1.00
- ¥ Fark 13.7k0.91
- Star 92.8k0.95
- () Code0.83
- Issues 5k+0.95
- à Pull requests 4290.98
- Discussions1.00
- Actions0.95
- ① Security and quality0.98
- Insights1.00
- P main0.96
- 4096 Branches1069 Tags0.97
- Q Go to file0.97
- <> Code0.93
- About0.98
- Lightweight coding agent that runs in1.00
- PRESENTED BY1.00
- bookholt-oai and viyatb-oai [codex] Reject uniowered PowerShell AST0.96
- 27122b5 - 31 minutes ago0.92
- 7,714 Commits0.96
- your terminal0.96
- .codex0.99
- core: log AGENTS.md paths as URis (#28989)0.99
- 4 days ago1.00
- Readme1.00
- Microsoft1.00
- devcontainer1.00
- build: run buildifier from just fmt (#28125)0.98
- last week0.99
- Apache-2.0 license0.99
- Ra Contributing0.92
- -github0.93
- ci restore custom Windows runner with hermetic LL.VM 0.7..0.95
- 10 hours ago1.00
- Security policy0.98
- I .vscode0.86
- Recommend Bazel VSCode extension. (#25161)1.00
- 3 weeks ago1.00
- Custom properties0.97
- A Activity0.87
- bazel1.00
- core: load AGENTS.md from foreign environments (#28958)0.97
- 4 days ago1.00
- ☆ 92.8k stars0.95
- codex-ci0.97
- cll: add package path from install context (#26189)0.98
- 3 weeks ago0.92
- ② 504 watching0.93
- ¥ 13.7k forks0.95
- codex-rs1.00
- [codex] Reject unlowered PowerShell AST regions (#24092)0.98
- 31 minutes ago1.00
- Report repository1.00
- I docs0.87
- [codex] Remove child AGENTS.md prompt experiment (#2...0.98
- 4 days ago1.00
- patches1.00
- ck restore custom Windows runner with hermetic LLVM 0.7...0.97
- 10 hours ago1.00
- Releases 8650.94
- 0.142.0 Latest0.96
- scripts1.00
- Make formatter output quiet on success(#29467)0.98
- 2 hours ago0.98
- 5 hours ago0.98
- sdk0.88
- Make formatter output quiet on success (#29467)0.99
- 2 hours ago0.99
- + 864 releases0.96
- third_party0.97
- bazel: add PowerShell o Wine test harness (#28120)0.99
- last week1.00
- Packages1.00
- tools1.00
- build: run buildifier from juast fmt (#28125)0.97
- last week0.98
- No packages published0.93
- OpenAl0.98
- 7/ 650.82
- 00.51
- World's Fair0.97
- TRACK 8· JULY 2, 20260.95
- Agentic Engineering1.00
-
- AlEngineer0.99
- World's Fair0.96
- Disclaimer:1.00
- Current snapshot1.00
- in time0.93
- OpenAl0.96
- 8 / 55 >0.78
- World's Fair0.95
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering1.00
-
- Protocols1.00
- AlEngineer0.99
- World'sFair1.00
- OpenAl0.91
- 9 / 55 30.88
- World's Fair0.94
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering1.00
-
- Protocols powering Codex1.00
- AlEngineer1.00
- World's Fair0.96
- 010.97
- 021.00
- App server1.00
- Responses APl0.98
- Handles the UI to0.98
- Handles the Agent to0.98
- Agent communication0.98
- Inference communication1.00
- OpenAl0.97
- 10 / 55 30.81
- Worid's Fair0.92
- TRACK 8• JULY 2, 20260.96
- Agentic Engineering1.00
-
- App server0.98
- AlEngineer0.97
- World's Fair0.98
- Client1.00
- Server1.00
- First-party products1.00
- Third-party integrations1.00
- Codex Desktop App1.00
- JetBrains IDEs1.00
- CODEX HARNESS0.98
- Codex TUI/CLI0.99
- JSON-RPC1.00
- JSON-RPC1.00
- VS Code (Codex)1.00
- VIA APP SERVER0.99
- Codex Web Runtime1.00
- Xcode1.00
- OpenAl0.97
- 11 / 55 30.93
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering0.99
-
- AlEngineer0.97
- World'sFair1.00
- Claude Code v2.1.850.99
- Opus 4.6 · Claude Pro0.94
- ~/Developer/codex1.00
- /plugin install codex@openai-codex0.99
- medium · /effort0.94
- OpenAl0.91
- 12 / 55 )0.89
- World's Fair0.95
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering1.00
-
- AlEngineer0.97
- World's Fair0.98
- CODEX 1010.98
- 520100%1.00
- 21.00
- 200%1.00
- 4000.99
- 3001.00
- 1001.00
- AMMOHEALTHARMS1.00
- ARMOR0.97
- 5201.00
- OpenAl0.97
- 13 / 550.90
- World's Fair0.94
- TRACK 8· JULY 2, 20260.95
- Agentic Engineering1.00
-
- Responses API0.96
- AlEngineer0.99
- World'sFair1.00
- Agentic by design1.00
- Redesign of Chat Completions1.00
- with agentic use cases in mind.1.00
- curl https://api.openai.com/v1/responses1.00
- -H "Content-Type: application/json"0.98
- -H "Authorization: Bearer $OPENAI_API_KEY"0.99
- Built-in capabilities1.00
- Includes built-in web search,0.99
- "model": "gpt-5.5",0.98
- image generation, and more.1.00
- "tools": [{"type": "web_search"}],0.99
- "input": "Build an AIEWF landing page"0.99
- Built on OpenResponses1.00
- Open protocol with a wide1.00
- range of hosting providers.1.00
- OpenAl0.96
- 14 / 55 )0.82
- World's Fair0.92
- TRACK 8· JULY 2, 20260.95
- Agentic Engineering1.00
-
- AlEngineer0.97
- World'sFair1.00
- UI (e.g. Codex app)0.98
- App server1.00
- HARNESS1.00
- Responses1.00
- LLMINFERENCE1.00
- OpenAl0.95
- 15 / 550.93
- World's Fair0.97
- TRACK 8· JULY 2, 20260.94
- Agentic Engineering1.00
-
- Building the context1.00
- AlEngineer0.98
- World's Fair0.97
- 010.98
- 021.00
- 031.00
- Size1.00
- Flexibility1.00
- Cacheability1.00
- Avoid cost1.00
- Everyone's Codex1.00
- Better cost1.00
- and confusion1.00
- looks different1.00
- and speed control1.00
- OpenAl0.98
- 18 / 550.90
- World's Fair0.96
- TRACK 8• JULY 2, 20260.95
- Agentic Engineering1.00
-
- Nano Codex1.00
- AlEngineer0.99
- Responses input1.00
- witing0.99
- World's Fair0.98
- Role:1.00
- System1.00
- Developer1.00
- User1.00
- Send the task to inspect the exact request0.99
- Find the subagent tool for inspecting tests, then tell me which tool you1.00
- would use. Do not execute it.1.00
- +0.99
- gpt-5.51.00
- 19 / 55 30.89
- 00.54
- World's Fair0.96
- TRACK 8· JULY 2, 20260.94
- Agentic Engineering1.00
-
- Nano Codex1.00
- AlEngineer0.99
- Responses input1.00
- World'sFair1.00
- Find the subagent tool for inspecting tests, then tell me0.99
- which tool you would use. Do not execute it.1.00
- Role:1.00
- 5 parts0.92
- System1.00
- Developer1.00
- User1.00
- Model instructions0.99
- Tool registry0.99
- Permissions + runtime0.97
- Available skills0.99
- 15%0.99
- Do anything1.00
- +0.99
- gpt-5.51.00
- Current task1.00
- <1%0.93
- World's Fair0.96
- TRACK 8· JULY 2, 20260.95
- Agentic Engineering0.99
-
- Nano Codex0.99
- AlEngineer0.96
- Responses input1.00
- Find the subagent tool for inspecting tests, then tell me0.98
- 6 parts0.90
- World'sFair1.00
- which tool you would use. Do not execute it.1.00
- Role:1.00
- System1.00
- Developer1.00
- User1.00
- Ran shell command0.95
- Model instructions1.00
- 66%0.97
- I'll use the test-triage skill for the test-inspection context, then search available0.99
- subagent tools.1.00
- Tool registry1.00
- PRESENTED BY0.99
- Permissions + runtime0.97
- Microsoft1.00
- Available skills1.00
- <skills_instructions> ## Skills A skill is a set of instructions provided through a SKILL.md' source. Below is the list of0.98
- 13%1.00
- skills that can be used. Each entry includes a name, description, and source locator. file' locators are on the host0.99
- filesystem, 'environment resource locators are owned by an execution environment, orchestrator resource' locators0.97
- are opaque non-filesystem resources, and 'custom resource locators use their provider's access mechanism. ###0.98
- Available skill - frontend-debugging: Debug small TypeScript and UI failures. (file: /workspace/.codex/skills/frontend-0.98
- debugging/SKILL.md) - test-triage: Interpret test output and narrow regresslions. (file: /workspace/.codex/skills/test-0.97
- triage/SKILL.md) - safe-shell: Use shell commands with narrow scope and explicit verification. (file:0.99
- /workspace/.codex/skills/safe-shell/SKILL.md) ### How to use skil - Discovery: The list above is the skills available in0.98
- this session (name + description + source locator). 'file' entries live on the host filesystem, 'environment resource0.99
- entries are owned by their execution environment, 'orchestrator resource' entries must be accessed through0.99
- skilslist' and 'skills.read', and custom resource' entries use their provider's access mechanism. - Trigger rules: If0.96
- the user names a skill (with '$SkillName' or plain text) OR the task clearly matches a skill's description shown above,0.98
- you must use that skill for that turn. Multiple mentions mean use them all. o not carry skills across turns unless re-0.99
- mentioned. - Missing/blocked: If a named skill isn't in the list or its source can't be read, say so briefly and continue0.99
- with the best fallback. - How to use a skill (progressive disclosure): 1) After deciding to use a skill, the main agent0.99
- must read its 'SKILL.md' completely before taking task actions. For a file' entry, open the listed path. For an0.97
- 'environment resource', use the filesystem of the owning environment. For an `orchestrator resource', call skills.list0.98
- Do anything1.00
- with ("authority":("kind*:"orchestrator*))', select the matching package, and pass its 'main_resource' to0.98
- +0.99
- gpt-5.51.00
- Thread history1.00
- 9%1.00
- Worild's Fair0.94
- TRACK 8· JULY 2,20260.94
- Agentic Engineering1.00
-
- Nano Codex1.00
- World'sFair1.00
- AlEngineer0.96
- which tool you would use. Do not execute it.0.99
- Find the subagent tool for inspecting tests, then tell me1.00
- Role:1.00
- Responses input1.00
- System1.00
- Developer1.00
- User1.00
- 00.60
- Ran shell command0.97
- I'll use the test-triage skill for the test-inspection context, then search available0.99
- Model instructions0.99
- subagent tools.1.00
- PRESENTED BY1.00
- I found the subagent tool I would use: spawn_agent. spawn_agent.0.99
- Microsoft0.96
- For inspecting tests without changing code, I'd choose it with agent_type:0.99
- "explorer" and a bounded message asking the subagent to inspect the0.99
- Tool registry0.99
- relevant test files and report findings.1.00
- I did not execute the subagent tool. Note: the test-triage skill file was listed0.99
- but not present at /workspace/.codex/skills/test-triage/SKILL.md, sol0.99
- Permissions + runtime0.99
- continued with the available tool discovery.1.00
- Available skills0.94
- 13%0.97
- Thread history1.00
- 9%1.00
- Do anything1.00
- +0.98
- gpt-5.51.00
- Current task1.00
- d%0.62
- World's Fair0.95
- TRACK 8· JULY 2, 20260.96
- Agentic Engineering1.00
Transcript
187 cues· 3,496 words· 18,933 chars
- 0:12 Hi, everyone.
- 0:13 We're going to start right on time because I'm going to speak basically at 2x.
- 0:17 I'm sorry I have a lot of content.
- 0:19 I'm trying to get you out here on time.
- 0:21 I want to start with a quick raise of hands.
- 0:23 So how many of you have built your own agents or are currently building your own agents?
- 0:28 Perfectly.
- 0:28 You're the right audience for this.
- 0:30 Over the next 20 minutes, I want to talk to you about a couple of different things that we're doing in the Codex harness that hopefully you can learn to apply to your own use cases or even just use the Codex harness with this in
- 0:43 in your own projects, or at bare minimum, learn what happens when you actually use codecs.
- 0:49 Since we're at iEngineer World's Fair, and we're actually on an agentic engineering track, I'm gonna stop bothering you with how does an agent work?
- 0:57 What is an agent?
- 0:59 And instead, I want to talk a bit more specifically about some key features that we have in the agent that I think are particularly interesting and are challenges you have to solve.
- 1:08 And so let's cover it from the lens of what actually happens when you send off a message.
- 1:15 Also, a quick reminder, if you're unaware, the Codex harness and everything I'm showing you is actually open source.
- 1:22 It's Apache 2 license, and the harness is written in Rust, so feel free to either learn from it, ask Codex deeper questions about what I'm covering, or fork it and make it your own.
- 1:34 Also disclaimer before we dive deep into it, this is a current state of affairs.
- 1:39 Things change so quickly.
- 1:41 You can always refer back to asking Codex what the current state is, but especially with new model releases, we often release new APIs and change sort of how the harness works.
- 1:50 So feel free to follow along as new models come out.
- 1:55 If we want to talk about how the Codex agent works, we first need to talk about what actually happens when you send off your message.
- 2:01 Namely, there's two protocols that are involved with the Codex agent.
- 2:07 The first one is what happens when you send it off in the UI and it goes to the harness.
- 2:11 We call that the app server.
- 2:12 I talked about that yesterday, so we're not going to spend too much time about it.
- 2:15 There will be a talk online that you can follow along.
- 2:18 The second part is the responses API, which handles the communication between the harness and the inference.
- 2:25 Both of these, though, are designed for an open ecosystem, meaning if you're building your own UI, you're building your own agent interface, you can actually build on top of the Codex harness using the app server protocol.
- 2:39 We use that same app server to power the Codex app, so it has really all of that functionality that you might expect from Codex, as well as we have a lot of third-party community projects that build on top of it, including Theos T3 Code or RemoteX, for example.
- 2:56 I even used that same app server to put Codex into Cloud Code.
- 2:59 So if you're a Cloud Code user and you want to leverage Codex, you can use that plugin.
- 3:05 And if you joined my talk yesterday, you saw me using that same protocol to actually put Codex into Doom, which was a fun adventure as well.
- 3:14 I mentioned the other part is the responses API.
- 3:17 So the response API was released last year as like a rethinking of the chat completions API in a more agentic world, meaning we redesigned slightly the structure, but more importantly, we added a lot of like building capabilities that are important for agents like web search, image gen, or other more complex capabilities that you will see as part of this talk.
- 3:38 We also want to make sure that this is like an open ecosystem.
- 3:42 So we worked with a lot of partners, including Ulama, LM Studio, NVIDIA, and others to codify an open responses schema and have a governance body for that so that other companies and platforms can actually build on that same responses API.
- 3:58 And you can use any responses API compatible harness model provider and actually plug it into the Codex harness.
- 4:08 So that's an overview of how these protocols work.
- 4:10 We're going from the UI to the harness with the app server protocol, and then from the harness to the LM inference using responses.
- 4:18 But what happens in the actual harness?
- 4:20 The first step, and arguably one of the most important ones, is context construction.
- 4:26 And during that, we care about three things quite a lot.
- 4:30 The first one is size.
- 4:31 We want to make sure that we don't blast through your token budgets and throw in a bunch of unnecessary content.
- 4:37 But also, the more context you have in your context, the higher it is that you have contradicting information, and it causes confusion for the model.
- 4:48 The other part is flexibility.
- 4:49 We want to make sure that regardless of many or how little skills you're using, you have a great experience, regardless of how many plugins and MCPs you install.
- 4:58 And of course, we want to make sure that things are performant and we know you're cost-sensitive, so cacheability is important as well.
- 5:06 To show you this and a couple of other things, I actually built this little nano codex here, which functions the same way.
- 5:13 It's built using the same code that is on the public repo, just turned into TypeScript.
loading