Videos bZISsg7H7DA
Your Agents Need a Save Button - Hamza Tahir, ZenML
Scene timeline
135 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 171
- whisperx 171
- chunks
- 31
- from 171 cues
- keyframes
- 89
- kept of 135 captured
- frames with text
- 89
- 5,137 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 20.3 MB
- word timings on 171 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:55 | 1m 25s |
stt |
done | — | 2026-08-11 00:56 | 18s |
chunk |
done | — | 2026-08-11 00:57 | 0s |
text_embed |
done | — | 2026-08-11 00:57 | 0s |
keyframe |
done | — | 2026-08-11 00:57 | 2m 41s |
ocr |
done | — | 2026-08-11 00:59 | 1m 33s |
frame_embed |
done | — | 2026-08-11 01:01 | 18s |
Frames, and what the machine read
-
- oKitaru0.89
- What if you could ask "what if?"0.97
- about a run that already happened?1.00
- Change one thing. Replay it. See what would have happened — no0.98
- re-running in production.1.00
-
- oKitaru0.88
- Your agents need a save button.1.00
- Hamza Tahir· Co-founder of ZenML·Building Kitaru1.00
-
- CHAPTER 1 · THE EXECUTION IS GONE0.96
- We've had the save button since the early 1980s.0.98
- So why don't our agents have them?1.00
- FILE·19841.00
- 9eS0.80
- AGENT·20260.99
- + no save0.95
- State persists. Close it, reopen it, keep working.1.00
- State lives in the process. When it exits, gone.1.00
-
- CHAPTER 10.99
- A trace is a description, held by a system that never0.99
- ran your code.0.98
- Your execution ran in your process. The telemetry went somewhere else — Braintrust, Langfuse, LangSmith.1.00
- Nothing connects the two.1.00
- OBSERVABILITY BACKEND - off to the side0.98
- run · draft-flow · completed ENDED 14:09:520.96
- YOUR PROCESS0.99
- The execution1.00
- spans·events1.00
- research1.00
- tool calls run here0.97
- ONE-WAY1.00
- write_draft1.00
- READ-ONLY1.00
- real state lives here1.00
- publish1.00
- then it's gone0.97
- done1.00
-
- CHAPTER 21.00
- Agents have many dials - simply choose0.97
- your new configuration and replay.1.00
- one replay call1.00
- {}0.99
- X0.99
- Swap the model0.99
- Mock a tool1.00
- Fail a call1.00
- Degrade1.00
- Drop a tool0.98
- retriever1.00
- overrides={model}1.00
- overrides={tool}1.00
- raise on step0.97
- swap the index1.00
- remove from set0.98
-
- CHAPTER 21.00
- Add a durable runtime layer.1.00
- Augment your traces with true checkpoints .0.99
- Claude1.00
- LangGraph1.00
- python1.00
- Pydantic1.00
- A Google Antigravity0.98
- OpenAI Agents SDK0.99
- Agent SDK1.00
- Durable runtime1.00
- checkpoints every call underneath0.99
-
- CHAPTER 3· YOUR TEST SET WROTE ITSELF0.99
- Your production runs are the ultimate test set.1.00
- Synthetic cases imagine what users might do. Real runs are what they did — real inputs, real tool responses, real0.99
- edge cases, including the ones that already went wrong.1.00
- SYNTHETIC1.00
- PRODUCTION1.00
- imagined - what users might do0.98
- real runs - what they actually did0.99
- Customer asks refund0.98
- Refund after a chargeback dispute1.00
- Check my order status0.99
- Order status - in Spanish, mid-chat0.99
- Reset my password - request0.99
- Refund request, tool timed out0.98
-
- CHAPTER 3 · THE LOOP0.97
- Close the loop and turn past runs into1.00
- full-blown evals.1.00
- Orepeat0.98
- 21.00
- 31.00
- 41.00
- Select a cohort0.98
- Replay one1.00
- Diff1.00
- Decide1.00
- change1.00
- the runs that matter0.98
- same call, one keyword1.00
- cost · decisions0.94
- divergence1.00
- ship · route · hold0.90
-
- CHAPTER 31.00
- Four steps. The methodology is the skill.0.99
- Checkpoint1.00
- the expensive, the escalated, the failed.0.99
- Replay1.00
- everything before your change, handed back from cache.0.99
- 030.98
- Diff0.98
- cost, changed decisions, where it first diverged.0.99
- 040.99
- Decide1.00
- ship what matches, route what's marginal, hold what drifts.1.00
-
- CHAPTER 4 · THIS WORKS AT SCALE0.98
- Don't believe me? Doordash and others1.00
- are doing it too.1.00
- DOORDASH1.00
- Blog0.95
- Inside DoorDash's one-click simulation and evaluation1.00
- platform for support chatbots0.99
- Joune I 20Chanan GoAas ambs0.66
- https://careersatdoordash.com/blog/doordashs-one1.00
- -click-simulation-and-evaluation-platform-for-sup1.00
- port-chatbots/1.00
-
- CHAPTER 4 · DOORDASH0.98
- 3021.00
- 21.00
- -901.00
- sims1.00
- pts1.00
- %0.97
- in 5 minutes0.93
- within target1.00
- hallucinations1.00
- where 1% of live traffic0.99
- simulated escalation rate1.00
- once the eval flywheel1.00
- took 7 hours0.97
- vs production1.00
- was in place1.00
-
- DEMO1.00
- One run → ten → ask the0.96
- agent.1.00
- A support agent in production. Every model and tool call checkpointed.0.99
-
- Chrome1.00
- File1.00
- Edit1.00
- View1.00
- History1.00
- Bookmarks1.00
- Profiles1.00
- Tab1.00
- Window1.00
- Help1.00
- 70.77
- V0.99
- A1.00
- 4分0.55
- Q0.96
- Thu 25. Jun 11:090.97
- Executions - Kitaru0.97
- 口0.74
- Your Agents Need a Save But ×0.98
- zenml-io/kitaru: Open-source ×0.99
- +0.54
- ←1.00
- →0.97
- C0.69
- github.com/zenml-io/kitaru1.00
- ☆1.00
- le0.81
- 米0.69
- Work1.00
- New Chrome available0.98
- GETTING_STARTED.md1.00
- Migrate Kitaru docs to GitBook + standalone SDK referenc...0.99
- 3 weeks ago1.00
- Justfile1.00
- Add sandbox stack and adapter command support (#423)1.00
- last week1.00
- Languages1.00
- LICENSE1.00
- Initial commit1.00
- 3 months ago1.00
- Python 98.2%0.96
- Shell 1.5%1.00
- Other 0.3%1.00
- README.md1.00
- Point Agent Harness Platform docs at the ZenML Agents g..0.99
- 2 days ago1.00
- SECURITY.md1.00
- Add security policy for vulnerability reporting0.99
- 3 months ago1.00
- lychee.toml1.00
- Point Agent Harness Platform docs at the ZenML Agents g...0.98
- 2 days ago1.00
- pyproject.toml1.00
- Release 0.18.01.00
- 1 hour ago1.00
- uv.lock1.00
- Release 0.18.01.00
- 1 hour ago1.00
- wrangler.redirect.toml1.00
- Migrate Kitaru docs to GitBook + standalone SDK referenc...0.98
- 3 weeks ago1.00
- wrangler.toml1.00
- Migrate Kitaru docs to GitBook + standalone SDK referenc...1.00
- 3 weeks ago1.00
- README1.00
- 8 Contributing0.95
- 8 Apache-2.0 license0.97
- Security1.00
- 00.95
- 三0.96
- Kitaru1.00
- The runtime layer underneath your agent stack.0.99
- Kitaru (来る, "to arrive") is a self-hosted, framework-agnostic runtime for autonomous agents — underneath the0.99
- harness your team already picked. You keep your agent SDK, your prompts, your tools, your model. Kitaru adds0.99
- durable execution: checkpoints. replav. resume. wait(). versioned deplovments. and isolated runtimes. runnina on0.99
- $10.93
-
- Chrome1.00
- File1.00
- Edit1.00
- View1.00
- History1.00
- Bookmarks1.00
- Profiles1.00
- Tab1.00
- Window1.00
- Help1.00
- V1.00
- A1.00
- G0.77
- Q0.97
- Thu 25. Jun 11:100.98
- Executions - Kitaru0.98
- x0.63
- 口0.86
- Your Agents Need a Save But ×0.97
- zenml-io/kitaru: Open-source ×0.96
- +0.57
- ←1.00
- →0.98
- C0.88
- github.com/zenml-io/kitaru1.00
- ☆1.00
- le0.81
- 米0.72
- Work1.00
- New Chrome available0.98
- pyproject.toml1.00
- Release 0.18.01.00
- 1 hour ago1.00
- uv.lock1.00
- Release 0.18.01.00
- 1 hour ago0.99
- wrangler.redirect.toml1.00
- Migrate Kitaru docs to GitBook + standalone SDK referenc...1.00
- 3 weeks ago1.00
- wrangler.toml1.00
- Migrate Kitaru docs to GitBook + standalone SDK referenc...1.00
- 3 weeks ago1.00
- README1.00
- 83 Contributing0.95
- Apache-2.0 license0.99
- Security0.97
- 三0.92
- 00.97
- Kitaru1.00
- The runtime layer underneath your agent stack.0.99
- Kitaru (来る, "to arrive") is a self-hosted, framework-agnostic runtime for autonomous agents — underneath the0.99
- harness your team already picked. You keep your agent SDK, your prompts, your tools, your model. Kitaru adds0.99
- durable execution: checkpoints, replay, resume, wait(), versioned deployments, and isolated runtimes, running on0.99
- your own infrastructure.1.00
- pypi1.00
- v0.17.11.00
- python 3.11 | 3.12 | 3.130.98
- license1.00
- Apache-2.01.00
- Docs·Quick Start·Examples·Getting Started Guide·1.00
- Roadmap0.96
- Community1.00
- kitaru0.84
- content_pipeline ACTIVE0.96
- Flows1.00
- Overview1.00
- Execution0.99
- Logs Cost Diff1.00
- PIPELI content_pipeline0.88
- 1260.89
- FATLIRES 230.62
- COT $0.0480.85
- content_pipeline1.00
- 1261.00
- ICCESS RATE0.77
- 73%1.00
- $0.0400.90
- 10.2s1.00
- $10.93
-
- Chrome1.00
- File1.00
- Edit1.00
- View1.00
- History1.00
- Bookmarks1.00
- Profiles1.00
- Tab1.00
- Window1.00
- Help1.00
- V0.99
- A1.00
- G0.74
- Q0.98
- Thu 25. Jun 11:100.99
- Executions - Kitaru0.98
- x0.64
- 口0.85
- Your Agents Need a Save But1.00
- zenml-io/kitaru: Open-source1.00
- ×0.68
- +0.55
- ←1.00
- →0.97
- C0.89
- github.com/zenml-io/kitaru1.00
- ☆1.00
- le0.81
- 米0.50
- Work1.00
- New Chrome available1.00
- 日0.59
- README1.00
- 83 Contributing0.94
- 8 Apache-2.0 license0.98
- Security1.00
- 三0.85
- 00.98
- Docs1.00
- Quick Start·0.96
- Examples1.00
- Getting Started Guide0.99
- Roadmap1.00
- Community1.00
- kitaru0.82
- Flows0.99
- content_pipeline0.99
- ACTIVE1.00
- Flows1.00
- Stacks0.92
- Overview1.00
- Execution1.00
- Logs1.00
- Cost0.99
- Diff0.98
- PIPELINE content_pipeline0.91
- : 260.77
- PATLURES 230.74
- ANE CT $0.0480.82
- content_pipeline1.00
- 1261.00
- 73%1.00
- $0.0400.89
- 10.2s0.99
- Completed0.99
- Faled0.92
- Waiting0.95
- Sleeping0.99
- Cancelled1.00
- at0.73
- at0.94
- Columns1.00
- EXECUTION0.95
- starus0.75
- DURATION0.99
- C010.54
- TORENI0.78
- CONFIO0.95
- DATE0.98
- TRIGOER0.83
- #1261.00
- COMPLETED1.00
- 12.2s0.96
- $0.0550.98
- 3.480.97
- prod0.89
- Mar 5, 2028 110.88
- schedule0.89
- #1250.92
- COMPLETED1.00
- 10.1s0.88
- 10.0370.91
- 5.080.83
- dv0.60
- Mar 5, 2020 17:270.93
- #1240.99
- COMPLETED1.00
- 18.7s0.83
- 80.0370.91
- 5.580.87
- prod0.98
- Mar 5, 2026 13:500.87
- #1230.89
- COMPLETED1.00
- 9.1s0.91
- 80.0580.93
- 4.180.96
- dev0.97
- Mar 5. 2026 12:270.91
- #1220.96
- COMPLETED0.99
- 15.6s0.87
- $0.0600.98
- $.580.75
- pred0.93
- Mer 5, 2026 13:480.87
- #1211.00
- COMPLETED1.00
- 10.1s0.93
- $0.0440.97
- 4.780.90
- prod0.98
- Mar 4, 2028 11030.93
- #1200.97
- COMPLETED1.00
- 10.8s0.95
- $0.0390.93
- 2.780.97
- dev0.78
- Mar 4, 2026 17:320.92
- #1190.99
- COMPLETED1.00
- 13.5s0.83
- $0.0390.99
- 4.480.78
- prod0.99
- Mar 4, 2026 17:110.94
- #1181.00
- COMPLETED1.00
- 11.ts0.93
- $0.0390.98
- 5.180.85
- dev0.83
- Mar 4, 2026 10:050.94
- #1170.99
- COMPLETED1.00
- 11.1s0.98
- $0.0410.95
- 4.580.88
- prod0.98
- Mar 3, 2026 13:080.92
- #1160.99
- FAILED0.97
- 1.t0.86
- $0.0140.95
- 1.780.88
- prod0.99
- Mer 2. 2020 12 280.84
- #1151.00
- CANCELLED1.00
- 2.0s0.96
- 50.0160.90
- 1.6K0.85
- prod0.98
- Mar 3, 2026 15 330.85
- #1140.84
- COMPLETED0.99
- 10.9s0.77
- 80.0430.91
- 4.180.99
- dev0.87
- Mar 3, 2026 13 390.88
- #1130.94
- COMPLETED0.98
- 1.500.74
- $0.0540.95
- 3.080.70
- dev0.61
- Mar 2, 2026 12 560.90
- Where Kitaru fits1.00
- $10.94
-
- Chrome1.00
- File1.00
- Edit1.00
- View1.00
- History1.00
- Bookmarks0.99
- Profiles1.00
- Tab1.00
- Window1.00
- Help1.00
- 口0.50
- 70.86
- V0.99
- A1.00
- G0.75
- Q0.98
- Thu 25. Jun 11:100.98
- Executions - Kitaru1.00
- ×0.87
- 口0.89
- Your Agents Need a Save But ×0.98
- zenml-io/kitaru: Open-source ×0.98
- ←1.00
- →0.97
- C0.89
- preview.demo.kitaru.zenml.io/flows/cc12f224-3790-47c8-af9b-7418cf99bda3/v/local/ex...1.00
- ☆1.00
- le0.85
- 米0.60
- Work1.00
- New Chrome available1.00
- Flows1.00
- support_copilot_flow0.99
- local0.98
- Flows1.00
- Executions1.00
- AD1.00
- support_copilot_flow· not deployed0.98
- Executions1.00
- Invoke1.00
- local0.92
- All versions1.00
- Status1.00
- Stack1.00
- Range1.00
- Q Search #num or flow...1.00
- G0.96
- Refresh1.00
- Execution1.00
- Status1.00
- Duration1.00
- Cost1.00
- ID1.00
- Author1.00
- Created ↓0.96
- #00711.00
- Completed1.00
- 1m 9s1.00
- $0.001513251.00
- 9e135198-e12e-4f32-9f76-bd4cdb0ed0371.00
- AD1.00
- admin1.00
- 25/06/2026,10:49:351.00
- …0.60
- #00701.00
- Completed1.00
- 47s1.00
- $0.001387251.00
- 7f3e3411-f73c-4c92-bd0d-daa9c70443461.00
- AD1.00
- admin1.00
- 25/06/2026,10:39:541.00
- ...0.87
- #00691.00
- Completed1.00
- 37s1.00
- $0.000986151.00
- 8e07850d-d954-4c98-8541-359cbfdf1e3b1.00
- AD1.00
- admin1.00
- 25/06/2026,10:25:591.00
- ...0.63
- #00681.00
- Failed1.00
- 6s1.00
- ea434beb-2755-474c-a6f6-185a47932b511.00
- AD1.00
- admin1.00
- 25/06/2026,10:25:481.00
- ...0.75
- #00671.00
- Completed1.00
- 35s1.00
- $0.000717351.00
- 1a48977e-0d46-4095-b9cb-b50a33abe41d1.00
- AD1.00
- admin1.00
- 25/06/2026.10:25:040.99
- #00661.00
- Completed1.00
- 36s1.00
- $0.000906551.00
- 37689f60-1b57-4a03-a163-2d41bc9b26a71.00
- AD1.00
- admin1.00
- 251.00
- $10.94
-
- ←1.00
- →0.99
- C0.94
- preview.demo.kitaru.zenml.io/flows/cc12f224-3790-47c8-af9b-7418cf99bda3/v/local/ex...1.00
- ☆1.00
- lo0.93
- 米0.86
- O0.97
- Flows1.00
- support_copilot_flow1.00
- local00.81
- #00711.00
- Execution1.00
- Logs1.00
- EXECUTIONS1.00
- List1.00
- Timeline1.00
- Filter steps...1.00
The page's on-screen-text budget of 600
lines is spent, so the last cards in this grid list fewer lines than they
hold. Narrow the page with ?frames= to read them.
Transcript
171 cues· 2,813 words· 14,686 chars
- 0:00 Have you ever looked at your agent execution and asked yourself the question, why did it do that?
- 0:06 What if it had done a different thing?
- 0:08 Would it have been cheaper?
- 0:10 Would it have been faster?
- 0:11 Well, you can do all these things if your agents have a save button.
- 0:20 We've had the save button for documents for decades now.
- 0:24 Since the 1980s, people have been used to pressing Control S, Command S,
- 0:29 or auto saving while you're working, to have a persistent state.
- 0:34 But agents, they don't have that today.
- 0:37 The only thing we have which is closest is a trace.
- 0:40 A trace gives you the emitted telemetry data of how an agent calls tools and the input and output of that state.
- 0:48 Now, while this is a good start, it is actually very disconnected from the runtime in which these agents actually execute.
- 0:55 So all the variables that are in state, all the file system that is in flight, the decisions that it makes in the code, the actual code itself,
- 1:05 All of that is lost, and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is.
- 1:16 And I think this is what's missing today in the industry, is that we don't have a clear connection between the observability spans that are emitted with ODEL and the execution.
- 1:28 And maybe at this point, you might be wondering, but why even bother?
- 1:32 Why do I need to have a save button?
- 1:35 Well, save allows you to replay.
- 1:37 You can go back in history and ask the what if question.
- 1:41 What are the types of questions you might want to ask?
- 1:43 Well, you might want to swap the model.
- 1:45 Maybe you use an open source model that is cheaper.
- 1:48 Maybe you mock a tool and you override what it returns.
- 1:52 Maybe you degrade it intentionally to see what would happen if things are wrong.
- 1:58 And these sorts of questions are only possible if you have that state.
- 2:03 And there is a category of the stack emerging which actually allows that.
- 2:09 And this sits on top of the harness, sits on top of the frameworks that allow you to create agents, and puts a durable runtime below that and augments the traces that are emitted with actually the code execution and the things that are around it to actually complete the state of the system.
- 2:31 And the good news is that once you have such a system, in production, you already have the information you need to ask those questions that are relevant to making your system better, cheaper and faster.
- 2:46 Production already has the traces.
- 2:48 It already has these state checkpoints.
- 2:53 ideally from the runtime, that can allow you to go back in time and ask those questions.
- 2:57 For example, let's take an agent example which does a customer resolution.
- 3:04 and refunds after a chargeback dispute.
- 3:06 Well, you can then see if the order status is changing languages or whether the request should have been escalated, or maybe a smaller model would have handled it if the runtime is checkpointing each of the state as it goes along.
- 3:19 Almost like an autosave, a command S, a control S in your agent.
- 3:24 And once you have that, you can even close the loop.
- 3:28 This conference is all about loops, so this is nothing different.
- 3:32 You have a cohort of runs that you think maybe matter because maybe they're too expensive, they took too long.
- 3:39 You replay a change, you diff it, you see what would have happened when you have the baseline, which you know what happened in the first place, and then you decide and you route and you ship back.
- 3:51 That's closing the loop on your evals.
- 3:54 It's basically evaluating using your production traces.
- 4:01 So it's basically evaluating using your production checkpoints.
- 4:08 So checkpoint, replay, diff, decide.
- 4:11 And this is really the methodology that I've seen and I've seen others do, which has really scaled.
- 4:18 For example, DoorDash.
- 4:19 DoorDash has a blog post on the 1st of June where they talk about having a simulated environment where they replayed customer bots and they've done what-if scenarios
- 4:31 and seeing how they could have made it better.
- 4:34 And where it used to take them hours and hours to do this, now they've reduced it to five minutes with hundreds of simulations, have 90% less hallucinations, and there's still two points within what they've seen in production.
- 4:49 So the simulations are pretty good because they're grounded in what's already happened.
- 4:54 So we're going to just walk through this in an example, and we're going to see how this works.
loading