Videos pSto5YaNGUo
The Agentic AI Engineer - Benedikt Sanftl, Mutagent
Scene timeline
128 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 256
- whisperx 256
- chunks
- 61
- from 256 cues
- keyframes
- 67
- kept of 128 captured
- frames with text
- 67
- 3,973 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 17.1 MB
- word timings on 256 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:40 | 1m 09s |
stt |
done | — | 2026-08-11 05:41 | 32s |
chunk |
done | — | 2026-08-11 05:41 | 0s |
text_embed |
done | — | 2026-08-11 05:41 | 0s |
keyframe |
done | — | 2026-08-11 05:41 | 1m 59s |
ocr |
done | — | 2026-08-11 05:43 | 1m 41s |
frame_embed |
done | — | 2026-08-11 05:45 | 11s |
Frames, and what the machine read
-
- THE MECHANISM1.00
- The Agentic Al Engineer0.99
- An Al agent is never "done." It lives in a development loop — and the speed of that loop is the whole0.99
- game.1.00
- The Agentic0.97
- Bene1.00
- Burak1.00
- 01/ 160.95
-
- MUTAGENT1.00
- THE LOOP1.00
- An agent is never "done." It lives in a loop.0.97
- Seven stages, run over and over — each pass making the agent a little better. Offline you build & sharpen it; online it0.98
- runs, and reality feeds the next pass.0.99
- Conceptualize1.00
- Optimize1.00
- Build1.00
- OFFLINE1.00
- one loop1.00
- offline build· online learn0.97
- Diagnose1.00
- Evaluate1.00
- ONLINE1.00
- Monitor1.00
- Deploy1.00
- Bene1.00
- Burak1.00
- The Agentic1.00
- 02 / 160.95
-
- MUTAGENT1.00
- THE PROBLEM1.00
- By hand, the lifecycle is one slow loop — and you are inside it.0.98
- Every improvement is a hand-run experiment — change something, generate outputs, and a human evaluates whether it helped.0.99
- Each experiment → evaluate cycle takes weeks, and the corrections never compound.0.98
- implement change0.98
- prompt - tools - logic0.91
- tweak by hand1.00
- generate samples1.00
- weeks per cycle0.99
- one slow cycle · nothing compounds0.98
- ©THE BOTTLENECK0.97
- manually evaluate1.00
- a human hand-reads a sample0.98
- HUMAN-GATED1.00
- SLOW EVEN WITH SIGNAL1.00
- CAN'T SCALE1.00
- Every change is judged by an engineer reading outputs -0.98
- Even with traces and evals, the loop still runs at review1.00
- Every new capability needs the same manual cycle — so0.98
- elow, subjective, hard to reproduce. The loop only moves0.98
- speed — root-cause → fix → confirmed result is a human0.98
- improvement is gated on human hours, not the agent.0.99
- human hours.0.99
- in that takes weeks.0.97
- Stop being the loop - start0.98
- ion; agents run the cycles.0.99
- The Agentic1.00
- Bene1.00
- Burak1.00
- 03 / 160.98
-
- MUTAGENT1.00
- THROUGHPUT1.00
- It's throughput — how many cycles fit the same window.0.97
- Each square is one development cycle. Run the loop by hand and the grid stays nearly empty; run it agentically and the same1.00
- window fills in — and every cycle can make the agent better.0.99
- HUMAN1.00
- 121.00
- weeks per cycle0.98
- cycles1.00
- AGENT1.00
- 2421.00
- minutes per cycle1.00
- cycles1.00
- same window of time →0.99
- lessmore0.97
- ame wall-clock window - the0.95
- and every one can compound on the last. Throughput is the moat.0.99
- Bene1.00
- Burak1.00
- The Agentic1.00
- 04 / 160.95
-
- MUTAGENT1.00
- THE MACHINE0.98
- The Al Engineer — the whole Al agent development lifecycle.0.98
- One orchestrator runs every stage end-to-end: work enters on the left, code PRs and agent / skill updates come out on the right.0.99
- SOURCES1.00
- LIVE·PRODUCTION1.00
- COLD START1.00
- MUTAGENT ORCHESTRATOR1.00
- end-to-end ADLC automation0.99
- Code PR1.00
- New feature or intent0.98
- deploy1.00
- EXISTING FEATURE0.98
- Bug, incident or enhancement0.99
- Define +0.98
- Design1.00
- *spec0.99
- *build1.00
- Build1.00
- Evaluate1.00
- *eval1.00
- Release1.00
- *ship1.00
- *monitor1.00
- Monitor1.00
- *diagnose1.00
- Diagnose1.00
- *optimize1.00
- Optimize1.00
- Agent update1.00
- live events / traces0.95
- Skill update1.00
- bug/incident→*diagnose-enhancement →spec/*optimize0.99
- STAGE BY STAGE1.00
- *spec1.00
- *build1.00
- *eval1.00
- *ship1.00
- *monitor1.00
- *diagnose1.00
- *optimize1.00
- Define1.00
- Build1.00
- Evalua1.00
- Release1.00
- Monitor1.00
- Diagnose1.00
- Optimize0.95
- A coding agent writes it from1.00
- Gate, ship, and verify it in0.98
- Same evals run live - catch0.94
- Cluster failures → root causes0.97
- Variants compete - only a0.96
- the spec.1.00
- production.1.00
- regressions & drift.0.99
- → new evals.0.99
- win ships.1.00
- Bene1.00
- Burak1.00
- The Agentic1.00
- 05 / 160.98
-
- MUTAGENT1.00
- THE MACHINE0.97
- The Al Engineer — the whole Al agent development lifecycle.0.98
- One orchestrator runs every stage end-to-end: work enters on the left, code PRs and agent / skill updates come out on the right.0.99
- SOURCES1.00
- LIVE1.00
- PRODUCTION0.98
- COLD START1.00
- MUTAGENT ORCHESTRATOR - end-to-Ond ADLC automation0.96
- Dotle PR0.91
- New feature or intent1.00
- deploy1.00
- EXISTING FEATURE1.00
- Bug, Incident or enhancement0.95
- Define +0.93
- Design1.00
- «spec0.65
- =buile0.80
- Build0.99
- Evaluate1.00
- *eval0.98
- Release1.00
- wship0.89
- *monitur0.90
- Monltor0.95
- *diagnose0.96
- Dlagnose0.94
- *optiniza0.89
- Optimize0.94
- Agent update1.00
- live everjts/ traces0.96
- Skill update1.00
- bug / incident → *diaghose - enhancement0.93
- fepac /*optimize0.93
- STAGE BY STAGE1.00
- *apds0.74
- *build0.88
- keval0.88
- *ship1.00
- *monitor0.98
- *diagnose1.00
- *optinize0.99
- Defane + Design0.98
- Build1.00
- Evaluate0.99
- Release1.00
- Moniter0.95
- Diagnoss0.93
- Optinize0.96
- Intent + what "good" means0.96
- A coding agent writes It from0.98
- Datnset + binary criteria →0.92
- Gate, ship. andi varify It in0.93
- Same evals run llve —catoh0.87
- Cluster failures - root csuses0.96
- Variante compete - anly a0.93
- → one signeci spec0.93
- the spec.0.96
- one success rate0.94
- production.0.95
- regressions é drift.0.95
- → new evals.0.83
- win:ships.0.87
- Bene1.00
- Burak1.00
- The Agentic1.00
- 05 / 160.93
-
- MUTAGENT1.00
- PHASE 1·CONCEPTUALIZE0.99
- Define why, designhow - into one spec.0.95
- Capture the intent and, critically, what "good" means; then shape the agent that delivers it. The signed spec is what every later0.99
- stage runs against.1.00
- business context1.00
- SPEC v10.99
- DEFINE1.00
- Why we're building it, the intent, and the bar for "good" — the acceptance0.99
- criteria the agent will be judged on.0.99
- what "good" means0.99
- intent / goal0.95
- DESIGN1.00
- The agent's shape — the routines, tools, and decision logic that turn that intent0.99
- into behaviour.1.00
- constraints1.00
- signed1.00
- The signed spec is the0.99
- er phase runs against — Build writes to it, Evaluate grades against its "good," and Optimize has to beat it.0.99
- The Agentic1.00
- Bene1.00
- Burak1.00
- 06 / 160.95
-
- MUTAGENT1.00
- PHASE 2·BUILD0.99
- Your coding agent generates the agent.1.00
- The signed spec drives a coding agent — Claude Code, Codex, Cursor, Pi, Hermes, whichever you run — that writes the agent0.99
- itself. The result is portable: the same agent runs on any harness.0.97
- YOUR CODING AGENT1.00
- AGENT1.00
- SPEC1.00
- v10.94
- →1.00
- →1.00
- Claude Code Codex Cursor0.99
- Pi0.99
- Hermes0.92
- writes the agent against the spec0.98
- RUNS ON1.00
- THE SAME AGENT RUNS ON ANY HARNESS0.98
- local . in your coding agent0.95
- cloud· managed0.97
- 米0.96
- @ weks0.72
- Claude1.00
- Pi0.99
- Hermes1.00
- Vercel1.00
- Mastra Claude Agents DeepAgents1.00
- Bene1.00
- Burak1.00
- The Agentic1.00
- 07 / 160.95
-
- MUTAGENT1.00
- PHASE 3 · EVALUATE0.92
- Make “good” measurable.0.98
- Derive a dataset of cases and a set of criteria — each a binary check, so a fail points straight at the broken dimension. Roll them0.99
- into one 0-100 success rate the loop can chase.1.00
- EVALS · THE CRITERIA0.96
- EVAL SYSTEM1.00
- GUIDED - COLD START1.00
- DATASET · PASS · FAIL PER CASE0.96
- Define criteria with the team from the spec and a few examples - no data needed yet.0.99
- DISCOVERED - FROM TRACES1.00
- CRITERIA - EACH A BINARY CHECK0.97
- Derive criteria from production traces - successes say what to keep, fallures what to check0.98
- Cites its source1.00
- PASS1.00
- for.1.00
- No fabricated facts1.00
- PASS1.00
- DATASETS · THE CASES0.96
- Correct tool used1.00
- FAIL1.00
- SYNTHESIZE - FROM GROUND TRUTH0.98
- Valid output format1.00
- PASS1.00
- Generate cases from the spec, historical exports, or known-good examples - before you have0.99
- traffic.1.00
- Policy respected1.00
- PASS1.00
- SUCCESS RATE0.99
- - FROM TRACES1.00
- 611.00
- /1000.99
- roduction runs into a representativ1.00
- ctually happen.1.00
- The Agentic1.00
- Bene1.00
- Burak1.00
- 08 / 160.98
Transcript
256 cues· 4,149 words· 22,809 chars
- 0:01 Hi, everybody.
- 0:03 Welcome to our talk, the agentic AI engineer.
- 0:06 I'm Bene, CEO and co-founder of Mutagent.
- 0:10 And I'm here with my colleague.
- 0:13 Hi, I'm Burak.
- 0:14 I'm the CTO of Mutagent.
- 0:16 And today we're basically going to talk about loops and how the agentic AI engineer works.
- 0:28 So as you're all aware of now, loops is the hot topic, how you build software in an agentic loop.
- 0:35 And we apply the same loop to the building of AI agents.
- 0:40 And as you're all aware, there's two concepts here.
- 0:44 One is
- 0:45 the offline loop where while you build, you iterate on your agent, you test it, you evaluate it, you improve it, and you go on.
- 0:54 And then you have a second loop, which we call the online loop, where once your agent is deployed to production, you monitor its traces, you diagnosis, and then you feed it back into your optimization loop to iterate and have multiple versions of your agents.
- 1:13 Yeah, up to until now, what we did was doing this loop manually.
- 1:20 It's quite slow.
- 1:22 The lifecycle is basically you have an issue, you want to change something to your agent.
- 1:29 Yeah, you implement the change.
- 1:31 You maybe vibe implement the change if you use coding agents for it.
- 1:35 Yeah, you generate some samples for this new feature issue to test it.
- 1:40 Yeah.
- 1:40 Then you look at the result, you look through the traces, how does the outcome look like?
- 1:45 Then you maybe ship it, you do AB testing and all your feedback is kind of manually, it takes very long.
- 1:52 Yeah.
- 1:52 And the bottleneck basically becomes the human review and the human
- 2:00 yeah building time and uh yeah that you can't scale especially if in your organization you are now planning to roll out hundreds of agents etc yeah and uh yeah this is why we think the agentic ai engineer is the natural next step to build agents and i'll have burak deep dive into how we improve timing and the
- 2:28 road to production reliability with the agentic engineer so yeah the key thing here is basically once you reach a certain number of agents or ai based features the human performing this loop again cannot really scale in enough time so
- 2:50 This is why doing this agentically is the key to increasing the throughput because then you can fit many more cycles into the same time window.
- 3:00 And now how that loop works is basically we have a few stages.
- 3:09 So this is when you are starting from scratch, like the current software development practices you first
- 3:18 create a spec for your agent or your skill in this case and here you need to define all the responsibilities and the functions that agents needs to handle the decisions that it has to make on certain conditions and here again this is only the definition stage
- 3:40 Once you define your agents requirements, then you can finally go into the build and build is where you then realize that spec in a specific harness or agent framework or in these days you could even build it as a cloud code or a codex agent.
- 4:01 Then comes the next step.
- 4:03 This is where you define clear evaluations to evaluate your agent's performance because these are the key metrics then where you can say, hey, my agent is functional or not.
- 4:18 Can think of essentially equivalent to unit tests for coding.
- 4:23 This is how you verify your agent works.
- 4:27 Then after evaluation, if everything looks fine, you usually have the ship basically where you deploy this agent to production.
- 4:38 Again, this can be a code update.
- 4:40 This can be a direct update on any agent platform or again, your local harness agents.
- 4:48 Then comes the online part.
- 4:51 This is where then the agent is continuously monitored for issues and based on certain trigger conditions, then you can start automatic diagnostics.
- 5:01 Again, this can be based on the volume of traces that your agent generates or weekly or daily jobs.
- 5:11 Then we go into diagnosis stage.
- 5:13 This is where you collect all the failures for your agents and do structured root cause analysis to then understand where the failures are coming from.
- 5:25 Once you understand and categorize the failures, then you can finally go on to the optimization stage.
- 5:31 This is where then you create, let's say, very specific changes or mutations for your agents to deal with the found failure modes.
- 5:43 And then the whole cycle repeats again.
- 5:46 You evaluate and if everything looks good, then you can deploy again.
- 5:52 Now we will maybe do a deep dive on each stage, what that entails.
- 6:01 So before we continue, Burak, we have two passes here.
- 6:07 Like one is the cold start path and one is basically existing features.
loading