read-only demo

Videos pSto5YaNGUo

The Agentic AI Engineer - Benedikt Sanftl, Mutagent

index_state ready data_status ok

AI Engineer· published 2026-06-29· 0:34:49· en· indexed 2026-08-11 05:45

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:26, 1 of 1 keyframes kept
  2. Shot 1, 0:26 to 1:12, 1 of 1 keyframes kept
  3. Shot 2, 1:12 to 1:42, 1 of 1 keyframes kept
  4. Shot 3, 1:42 to 2:12, 0 of 1 keyframes kept
  5. Shot 4, 2:12 to 2:38, 1 of 1 keyframes kept
  6. Shot 5, 2:38 to 3:04, 0 of 1 keyframes kept
  7. Shot 6, 3:04 to 3:33, 1 of 1 keyframes kept
  8. Shot 7, 3:33 to 4:02, 0 of 1 keyframes kept
  9. Shot 8, 4:02 to 4:31, 0 of 1 keyframes kept
  10. Shot 9, 4:31 to 5:00, 0 of 1 keyframes kept
  11. Shot 10, 5:00 to 5:30, 0 of 1 keyframes kept
  12. Shot 11, 5:30 to 5:59, 0 of 1 keyframes kept
  13. Shot 12, 5:59 to 6:00, 1 of 1 keyframes kept
  14. Shot 13, 6:00 to 6:37, 0 of 1 keyframes kept
  15. Shot 14, 6:37 to 7:14, 0 of 1 keyframes kept
  16. Shot 15, 7:14 to 7:47, 1 of 1 keyframes kept
  17. Shot 16, 7:47 to 8:20, 0 of 1 keyframes kept
  18. Shot 17, 8:20 to 8:53, 0 of 1 keyframes kept
  19. Shot 18, 8:53 to 9:21, 1 of 1 keyframes kept
  20. Shot 19, 9:21 to 9:49, 0 of 1 keyframes kept
  21. Shot 20, 9:49 to 10:17, 0 of 1 keyframes kept
  22. Shot 21, 10:17 to 10:46, 0 of 1 keyframes kept
  23. Shot 22, 10:46 to 11:14, 0 of 1 keyframes kept
  24. Shot 23, 11:14 to 11:41, 1 of 1 keyframes kept
  25. Shot 24, 11:41 to 12:08, 0 of 1 keyframes kept
  26. Shot 25, 12:08 to 12:35, 0 of 1 keyframes kept
  27. Shot 26, 12:35 to 13:03, 0 of 1 keyframes kept
  28. Shot 27, 13:03 to 13:30, 0 of 1 keyframes kept
  29. Shot 28, 13:30 to 13:57, 0 of 1 keyframes kept
  30. Shot 29, 13:57 to 14:24, 0 of 1 keyframes kept
  31. Shot 30, 14:24 to 14:51, 0 of 1 keyframes kept
  32. Shot 31, 14:51 to 15:19, 0 of 1 keyframes kept
  33. Shot 32, 15:19 to 15:22, 1 of 1 keyframes kept
  34. Shot 33, 15:22 to 15:54, 1 of 1 keyframes kept
  35. Shot 34, 15:54 to 16:26, 0 of 1 keyframes kept
  36. Shot 35, 16:26 to 16:58, 0 of 1 keyframes kept
  37. Shot 36, 16:58 to 17:28, 1 of 1 keyframes kept
  38. Shot 37, 17:28 to 17:58, 0 of 1 keyframes kept
  39. Shot 38, 17:58 to 18:28, 0 of 1 keyframes kept
  40. Shot 39, 18:28 to 18:59, 0 of 1 keyframes kept
  41. Shot 40, 18:59 to 19:02, 0 of 1 keyframes kept
  42. Shot 41, 19:02 to 19:29, 1 of 1 keyframes kept
  43. Shot 42, 19:29 to 19:56, 0 of 1 keyframes kept
  44. Shot 43, 19:56 to 20:23, 0 of 1 keyframes kept
  45. Shot 44, 20:23 to 20:51, 1 of 1 keyframes kept
  46. Shot 45, 20:51 to 21:19, 0 of 1 keyframes kept
  47. Shot 46, 21:19 to 21:47, 0 of 1 keyframes kept
  48. Shot 47, 21:47 to 22:15, 0 of 1 keyframes kept
  49. Shot 48, 22:15 to 22:45, 1 of 1 keyframes kept
  50. Shot 49, 22:45 to 23:16, 0 of 1 keyframes kept
  51. Shot 50, 23:16 to 23:44, 1 of 1 keyframes kept
  52. Shot 51, 23:44 to 24:12, 0 of 1 keyframes kept
  53. Shot 52, 24:12 to 24:41, 0 of 1 keyframes kept
  54. Shot 53, 24:41 to 25:09, 0 of 1 keyframes kept
  55. Shot 54, 25:09 to 25:38, 1 of 1 keyframes kept
  56. Shot 55, 25:38 to 26:07, 0 of 1 keyframes kept
  57. Shot 56, 26:07 to 26:36, 0 of 1 keyframes kept
  58. Shot 57, 26:36 to 27:05, 0 of 1 keyframes kept
  59. Shot 58, 27:05 to 27:22, 1 of 1 keyframes kept
  60. Shot 59, 27:22 to 27:27, 1 of 1 keyframes kept
  61. Shot 60, 27:27 to 27:32, 1 of 1 keyframes kept
  62. Shot 61, 27:32 to 27:35, 1 of 1 keyframes kept
  63. Shot 62, 27:35 to 28:06, 1 of 1 keyframes kept
  64. Shot 63, 28:06 to 28:37, 1 of 1 keyframes kept
  65. Shot 64, 28:37 to 29:01, 1 of 1 keyframes kept
  66. Shot 65, 29:01 to 29:08, 1 of 1 keyframes kept
  67. Shot 66, 29:08 to 29:10, 1 of 1 keyframes kept
  68. Shot 67, 29:10 to 29:13, 1 of 1 keyframes kept
  69. Shot 68, 29:13 to 29:15, 1 of 1 keyframes kept
  70. Shot 69, 29:15 to 29:22, 0 of 1 keyframes kept
  71. Shot 70, 29:22 to 29:24, 0 of 1 keyframes kept
  72. Shot 71, 29:24 to 29:30, 0 of 1 keyframes kept
  73. Shot 72, 29:30 to 29:33, 1 of 1 keyframes kept
  74. Shot 73, 29:33 to 29:35, 1 of 1 keyframes kept
  75. Shot 74, 29:35 to 29:37, 0 of 1 keyframes kept
  76. Shot 75, 29:37 to 29:42, 1 of 1 keyframes kept
  77. Shot 76, 29:42 to 29:44, 1 of 1 keyframes kept
  78. Shot 77, 29:44 to 29:46, 0 of 1 keyframes kept
  79. Shot 78, 29:46 to 29:48, 1 of 1 keyframes kept
  80. Shot 79, 29:48 to 29:50, 1 of 1 keyframes kept
  81. Shot 80, 29:50 to 29:51, 1 of 1 keyframes kept
  82. Shot 81, 29:51 to 29:54, 1 of 1 keyframes kept
  83. Shot 82, 29:54 to 29:55, 0 of 1 keyframes kept
  84. Shot 83, 29:55 to 30:01, 1 of 1 keyframes kept
  85. Shot 84, 30:01 to 30:07, 1 of 1 keyframes kept
  86. Shot 85, 30:07 to 30:10, 1 of 1 keyframes kept
  87. Shot 86, 30:10 to 30:12, 1 of 1 keyframes kept
  88. Shot 87, 30:12 to 30:13, 1 of 1 keyframes kept
  89. Shot 88, 30:13 to 30:28, 1 of 1 keyframes kept
  90. Shot 89, 30:28 to 30:30, 1 of 1 keyframes kept
  91. Shot 90, 30:30 to 30:45, 0 of 1 keyframes kept
  92. Shot 91, 30:45 to 30:46, 1 of 1 keyframes kept
  93. Shot 92, 30:46 to 30:48, 1 of 1 keyframes kept
  94. Shot 93, 30:48 to 30:56, 0 of 1 keyframes kept
  95. Shot 94, 30:56 to 31:03, 1 of 1 keyframes kept
  96. Shot 95, 31:03 to 31:33, 1 of 1 keyframes kept
  97. Shot 96, 31:33 to 31:45, 0 of 1 keyframes kept
  98. Shot 97, 31:45 to 31:46, 1 of 1 keyframes kept
  99. Shot 98, 31:46 to 31:48, 1 of 1 keyframes kept
  100. Shot 99, 31:48 to 31:53, 0 of 1 keyframes kept
  101. Shot 100, 31:53 to 31:55, 1 of 1 keyframes kept
  102. Shot 101, 31:55 to 31:56, 0 of 1 keyframes kept
  103. Shot 102, 31:56 to 32:01, 1 of 1 keyframes kept
  104. Shot 103, 32:01 to 32:03, 1 of 1 keyframes kept
  105. Shot 104, 32:03 to 32:09, 1 of 1 keyframes kept
  106. Shot 105, 32:09 to 32:10, 1 of 1 keyframes kept
  107. Shot 106, 32:10 to 32:12, 0 of 1 keyframes kept
  108. Shot 107, 32:12 to 32:15, 1 of 1 keyframes kept
  109. Shot 108, 32:15 to 32:20, 1 of 1 keyframes kept
  110. Shot 109, 32:20 to 32:21, 0 of 1 keyframes kept
  111. Shot 110, 32:21 to 32:24, 0 of 1 keyframes kept
  112. Shot 111, 32:24 to 32:25, 1 of 1 keyframes kept
  113. Shot 112, 32:25 to 32:27, 1 of 1 keyframes kept
  114. Shot 113, 32:27 to 32:28, 1 of 1 keyframes kept
  115. Shot 114, 32:28 to 32:30, 1 of 1 keyframes kept
  116. Shot 115, 32:30 to 32:32, 1 of 1 keyframes kept
  117. Shot 116, 32:32 to 33:03, 1 of 1 keyframes kept
  118. Shot 117, 33:03 to 33:07, 0 of 1 keyframes kept
  119. Shot 118, 33:07 to 33:08, 1 of 1 keyframes kept
  120. Shot 119, 33:08 to 33:10, 1 of 1 keyframes kept
  121. Shot 120, 33:10 to 33:12, 0 of 1 keyframes kept
  122. Shot 121, 33:12 to 33:21, 0 of 1 keyframes kept
  123. Shot 122, 33:21 to 33:22, 0 of 1 keyframes kept
  124. Shot 123, 33:22 to 33:25, 1 of 1 keyframes kept
  125. Shot 124, 33:25 to 33:28, 0 of 1 keyframes kept
  126. Shot 125, 33:28 to 34:10, 1 of 1 keyframes kept
  127. Shot 126, 34:10 to 34:15, 0 of 1 keyframes kept
  128. Shot 127, 34:15 to 34:49, 1 of 1 keyframes kept

128 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
256
whisperx 256
chunks
61
from 256 cues
keyframes
67
kept of 128 captured
frames with text
67
3,973 lines read
chapters
0
from the source metadata
keyframe bytes
17.1 MB
word timings on 256 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:40 1m 09s
stt done 2026-08-11 05:41 32s
chunk done 2026-08-11 05:41 0s
text_embed done 2026-08-11 05:41 0s
keyframe done 2026-08-11 05:41 1m 59s
ocr done 2026-08-11 05:43 1m 41s
frame_embed done 2026-08-11 05:45 11s

Frames, and what the machine read

  • 0:15 #0 done8 line(s)

    shot 0·sharpness 939.3

    1. THE MECHANISM1.00
    2. The Agentic Al Engineer0.99
    3. An Al agent is never "done." It lives in a development loop — and the speed of that loop is the whole0.99
    4. game.1.00
    5. The Agentic0.97
    6. Bene1.00
    7. Burak1.00
    8. 01/ 160.95
  • 0:58 #1 done20 line(s)

    shot 1·sharpness 1472.5

    1. MUTAGENT1.00
    2. THE LOOP1.00
    3. An agent is never "done." It lives in a loop.0.97
    4. Seven stages, run over and over — each pass making the agent a little better. Offline you build & sharpen it; online it0.98
    5. runs, and reality feeds the next pass.0.99
    6. Conceptualize1.00
    7. Optimize1.00
    8. Build1.00
    9. OFFLINE1.00
    10. one loop1.00
    11. offline build· online learn0.97
    12. Diagnose1.00
    13. Evaluate1.00
    14. ONLINE1.00
    15. Monitor1.00
    16. Deploy1.00
    17. Bene1.00
    18. Burak1.00
    19. The Agentic1.00
    20. 02 / 160.95
  • 1:35 #2 done31 line(s)

    shot 2·sharpness 2010.1

    1. MUTAGENT1.00
    2. THE PROBLEM1.00
    3. By hand, the lifecycle is one slow loop — and you are inside it.0.98
    4. Every improvement is a hand-run experiment — change something, generate outputs, and a human evaluates whether it helped.0.99
    5. Each experiment → evaluate cycle takes weeks, and the corrections never compound.0.98
    6. implement change0.98
    7. prompt - tools - logic0.91
    8. tweak by hand1.00
    9. generate samples1.00
    10. weeks per cycle0.99
    11. one slow cycle · nothing compounds0.98
    12. ©THE BOTTLENECK0.97
    13. manually evaluate1.00
    14. a human hand-reads a sample0.98
    15. HUMAN-GATED1.00
    16. SLOW EVEN WITH SIGNAL1.00
    17. CAN'T SCALE1.00
    18. Every change is judged by an engineer reading outputs -0.98
    19. Even with traces and evals, the loop still runs at review1.00
    20. Every new capability needs the same manual cycle — so0.98
    21. elow, subjective, hard to reproduce. The loop only moves0.98
    22. speed — root-cause → fix → confirmed result is a human0.98
    23. improvement is gated on human hours, not the agent.0.99
    24. human hours.0.99
    25. in that takes weeks.0.97
    26. Stop being the loop - start0.98
    27. ion; agents run the cycles.0.99
    28. The Agentic1.00
    29. Bene1.00
    30. Burak1.00
    31. 03 / 160.98
  • 2:00 #3 skipped

    shot 3·duplicate of #2

  • 2:32 #4 done21 line(s)

    shot 4·sharpness 1799.6

    1. MUTAGENT1.00
    2. THROUGHPUT1.00
    3. It's throughput — how many cycles fit the same window.0.97
    4. Each square is one development cycle. Run the loop by hand and the grid stays nearly empty; run it agentically and the same1.00
    5. window fills in — and every cycle can make the agent better.0.99
    6. HUMAN1.00
    7. 121.00
    8. weeks per cycle0.98
    9. cycles1.00
    10. AGENT1.00
    11. 2421.00
    12. minutes per cycle1.00
    13. cycles1.00
    14. same window of time →0.99
    15. lessmore0.97
    16. ame wall-clock window - the0.95
    17. and every one can compound on the last. Throughput is the moat.0.99
    18. Bene1.00
    19. Burak1.00
    20. The Agentic1.00
    21. 04 / 160.95
  • 3:01 #5 skipped

    shot 5·duplicate of #4

  • 3:27 #6 done62 line(s)

    shot 6·sharpness 2060.8

    1. MUTAGENT1.00
    2. THE MACHINE0.98
    3. The Al Engineer — the whole Al agent development lifecycle.0.98
    4. One orchestrator runs every stage end-to-end: work enters on the left, code PRs and agent / skill updates come out on the right.0.99
    5. SOURCES1.00
    6. LIVE·PRODUCTION1.00
    7. COLD START1.00
    8. MUTAGENT ORCHESTRATOR1.00
    9. end-to-end ADLC automation0.99
    10. Code PR1.00
    11. New feature or intent0.98
    12. deploy1.00
    13. EXISTING FEATURE0.98
    14. Bug, incident or enhancement0.99
    15. Define +0.98
    16. Design1.00
    17. *spec0.99
    18. *build1.00
    19. Build1.00
    20. Evaluate1.00
    21. *eval1.00
    22. Release1.00
    23. *ship1.00
    24. *monitor1.00
    25. Monitor1.00
    26. *diagnose1.00
    27. Diagnose1.00
    28. *optimize1.00
    29. Optimize1.00
    30. Agent update1.00
    31. live events / traces0.95
    32. Skill update1.00
    33. bug/incident→*diagnose-enhancement →spec/*optimize0.99
    34. STAGE BY STAGE1.00
    35. *spec1.00
    36. *build1.00
    37. *eval1.00
    38. *ship1.00
    39. *monitor1.00
    40. *diagnose1.00
    41. *optimize1.00
    42. Define1.00
    43. Build1.00
    44. Evalua1.00
    45. Release1.00
    46. Monitor1.00
    47. Diagnose1.00
    48. Optimize0.95
    49. A coding agent writes it from1.00
    50. Gate, ship, and verify it in0.98
    51. Same evals run live - catch0.94
    52. Cluster failures → root causes0.97
    53. Variants compete - only a0.96
    54. the spec.1.00
    55. production.1.00
    56. regressions & drift.0.99
    57. → new evals.0.99
    58. win ships.1.00
    59. Bene1.00
    60. Burak1.00
    61. The Agentic1.00
    62. 05 / 160.98
  • 3:48 #7 skipped

    shot 7·duplicate of #6

  • 4:22 #8 skipped

    shot 8·duplicate of #6

  • 4:54 #9 skipped

    shot 9·duplicate of #6

  • 5:12 #10 skipped

    shot 10·duplicate of #6

  • 5:39 #11 skipped

    shot 11·duplicate of #6

  • 5:59 #12 done67 line(s)

    shot 12·sharpness 1874.4

    1. MUTAGENT1.00
    2. THE MACHINE0.97
    3. The Al Engineer — the whole Al agent development lifecycle.0.98
    4. One orchestrator runs every stage end-to-end: work enters on the left, code PRs and agent / skill updates come out on the right.0.99
    5. SOURCES1.00
    6. LIVE1.00
    7. PRODUCTION0.98
    8. COLD START1.00
    9. MUTAGENT ORCHESTRATOR - end-to-Ond ADLC automation0.96
    10. Dotle PR0.91
    11. New feature or intent1.00
    12. deploy1.00
    13. EXISTING FEATURE1.00
    14. Bug, Incident or enhancement0.95
    15. Define +0.93
    16. Design1.00
    17. «spec0.65
    18. =buile0.80
    19. Build0.99
    20. Evaluate1.00
    21. *eval0.98
    22. Release1.00
    23. wship0.89
    24. *monitur0.90
    25. Monltor0.95
    26. *diagnose0.96
    27. Dlagnose0.94
    28. *optiniza0.89
    29. Optimize0.94
    30. Agent update1.00
    31. live everjts/ traces0.96
    32. Skill update1.00
    33. bug / incident → *diaghose - enhancement0.93
    34. fepac /*optimize0.93
    35. STAGE BY STAGE1.00
    36. *apds0.74
    37. *build0.88
    38. keval0.88
    39. *ship1.00
    40. *monitor0.98
    41. *diagnose1.00
    42. *optinize0.99
    43. Defane + Design0.98
    44. Build1.00
    45. Evaluate0.99
    46. Release1.00
    47. Moniter0.95
    48. Diagnoss0.93
    49. Optinize0.96
    50. Intent + what "good" means0.96
    51. A coding agent writes It from0.98
    52. Datnset + binary criteria →0.92
    53. Gate, ship. andi varify It in0.93
    54. Same evals run llve —catoh0.87
    55. Cluster failures - root csuses0.96
    56. Variante compete - anly a0.93
    57. → one signeci spec0.93
    58. the spec.0.96
    59. one success rate0.94
    60. production.0.95
    61. regressions é drift.0.95
    62. → new evals.0.83
    63. win:ships.0.87
    64. Bene1.00
    65. Burak1.00
    66. The Agentic1.00
    67. 05 / 160.93
  • 6:32 #13 skipped

    shot 13·duplicate of #6

  • 6:41 #14 skipped

    shot 14·duplicate of #6

  • 7:37 #15 done23 line(s)

    shot 15·sharpness 1617.3

    1. MUTAGENT1.00
    2. PHASE 1·CONCEPTUALIZE0.99
    3. Define why, designhow - into one spec.0.95
    4. Capture the intent and, critically, what "good" means; then shape the agent that delivers it. The signed spec is what every later0.99
    5. stage runs against.1.00
    6. business context1.00
    7. SPEC v10.99
    8. DEFINE1.00
    9. Why we're building it, the intent, and the bar for "good" — the acceptance0.99
    10. criteria the agent will be judged on.0.99
    11. what "good" means0.99
    12. intent / goal0.95
    13. DESIGN1.00
    14. The agent's shape — the routines, tools, and decision logic that turn that intent0.99
    15. into behaviour.1.00
    16. constraints1.00
    17. signed1.00
    18. The signed spec is the0.99
    19. er phase runs against — Build writes to it, Evaluate grades against its "good," and Optimize has to beat it.0.99
    20. The Agentic1.00
    21. Bene1.00
    22. Burak1.00
    23. 06 / 160.95
  • 8:10 #16 skipped

    shot 16·duplicate of #15

  • 8:43 #17 skipped

    shot 17·duplicate of #15

  • 9:15 #18 done30 line(s)

    shot 18·sharpness 1829.8

    1. MUTAGENT1.00
    2. PHASE 2·BUILD0.99
    3. Your coding agent generates the agent.1.00
    4. The signed spec drives a coding agent — Claude Code, Codex, Cursor, Pi, Hermes, whichever you run — that writes the agent0.99
    5. itself. The result is portable: the same agent runs on any harness.0.97
    6. YOUR CODING AGENT1.00
    7. AGENT1.00
    8. SPEC1.00
    9. v10.94
    10. 1.00
    11. 1.00
    12. Claude Code Codex Cursor0.99
    13. Pi0.99
    14. Hermes0.92
    15. writes the agent against the spec0.98
    16. RUNS ON1.00
    17. THE SAME AGENT RUNS ON ANY HARNESS0.98
    18. local . in your coding agent0.95
    19. cloud· managed0.97
    20. 0.96
    21. @ weks0.72
    22. Claude1.00
    23. Pi0.99
    24. Hermes1.00
    25. Vercel1.00
    26. Mastra Claude Agents DeepAgents1.00
    27. Bene1.00
    28. Burak1.00
    29. The Agentic1.00
    30. 07 / 160.95
  • 9:35 #19 skipped

    shot 19·duplicate of #18

  • 10:09 #20 skipped

    shot 20·duplicate of #18

  • 10:23 #21 skipped

    shot 21·duplicate of #18

  • 11:10 #22 skipped

    shot 22·duplicate of #18

  • 11:30 #23 done38 line(s)

    shot 23·sharpness 1607.4

    1. MUTAGENT1.00
    2. PHASE 3 · EVALUATE0.92
    3. Make “good” measurable.0.98
    4. Derive a dataset of cases and a set of criteria — each a binary check, so a fail points straight at the broken dimension. Roll them0.99
    5. into one 0-100 success rate the loop can chase.1.00
    6. EVALS · THE CRITERIA0.96
    7. EVAL SYSTEM1.00
    8. GUIDED - COLD START1.00
    9. DATASET · PASS · FAIL PER CASE0.96
    10. Define criteria with the team from the spec and a few examples - no data needed yet.0.99
    11. DISCOVERED - FROM TRACES1.00
    12. CRITERIA - EACH A BINARY CHECK0.97
    13. Derive criteria from production traces - successes say what to keep, fallures what to check0.98
    14. Cites its source1.00
    15. PASS1.00
    16. for.1.00
    17. No fabricated facts1.00
    18. PASS1.00
    19. DATASETS · THE CASES0.96
    20. Correct tool used1.00
    21. FAIL1.00
    22. SYNTHESIZE - FROM GROUND TRUTH0.98
    23. Valid output format1.00
    24. PASS1.00
    25. Generate cases from the spec, historical exports, or known-good examples - before you have0.99
    26. traffic.1.00
    27. Policy respected1.00
    28. PASS1.00
    29. SUCCESS RATE0.99
    30. - FROM TRACES1.00
    31. 611.00
    32. /1000.99
    33. roduction runs into a representativ1.00
    34. ctually happen.1.00
    35. The Agentic1.00
    36. Bene1.00
    37. Burak1.00
    38. 08 / 160.98

Transcript

256 cues· 4,149 words· 22,809 chars

  1. 0:01 Hi, everybody.
  2. 0:03 Welcome to our talk, the agentic AI engineer.
  3. 0:06 I'm Bene, CEO and co-founder of Mutagent.
  4. 0:10 And I'm here with my colleague.
  5. 0:13 Hi, I'm Burak.
  6. 0:14 I'm the CTO of Mutagent.
  7. 0:16 And today we're basically going to talk about loops and how the agentic AI engineer works.
  8. 0:28 So as you're all aware of now, loops is the hot topic, how you build software in an agentic loop.
  9. 0:35 And we apply the same loop to the building of AI agents.
  10. 0:40 And as you're all aware, there's two concepts here.
  11. 0:44 One is
  12. 0:45 the offline loop where while you build, you iterate on your agent, you test it, you evaluate it, you improve it, and you go on.
  13. 0:54 And then you have a second loop, which we call the online loop, where once your agent is deployed to production, you monitor its traces, you diagnosis, and then you feed it back into your optimization loop to iterate and have multiple versions of your agents.
  14. 1:13 Yeah, up to until now, what we did was doing this loop manually.
  15. 1:20 It's quite slow.
  16. 1:22 The lifecycle is basically you have an issue, you want to change something to your agent.
  17. 1:29 Yeah, you implement the change.
  18. 1:31 You maybe vibe implement the change if you use coding agents for it.
  19. 1:35 Yeah, you generate some samples for this new feature issue to test it.
  20. 1:40 Yeah.
  21. 1:40 Then you look at the result, you look through the traces, how does the outcome look like?
  22. 1:45 Then you maybe ship it, you do AB testing and all your feedback is kind of manually, it takes very long.
  23. 1:52 Yeah.
  24. 1:52 And the bottleneck basically becomes the human review and the human
  25. 2:00 yeah building time and uh yeah that you can't scale especially if in your organization you are now planning to roll out hundreds of agents etc yeah and uh yeah this is why we think the agentic ai engineer is the natural next step to build agents and i'll have burak deep dive into how we improve timing and the
  26. 2:28 road to production reliability with the agentic engineer so yeah the key thing here is basically once you reach a certain number of agents or ai based features the human performing this loop again cannot really scale in enough time so
  27. 2:50 This is why doing this agentically is the key to increasing the throughput because then you can fit many more cycles into the same time window.
  28. 3:00 And now how that loop works is basically we have a few stages.
  29. 3:09 So this is when you are starting from scratch, like the current software development practices you first
  30. 3:18 create a spec for your agent or your skill in this case and here you need to define all the responsibilities and the functions that agents needs to handle the decisions that it has to make on certain conditions and here again this is only the definition stage
  31. 3:40 Once you define your agents requirements, then you can finally go into the build and build is where you then realize that spec in a specific harness or agent framework or in these days you could even build it as a cloud code or a codex agent.
  32. 4:01 Then comes the next step.
  33. 4:03 This is where you define clear evaluations to evaluate your agent's performance because these are the key metrics then where you can say, hey, my agent is functional or not.
  34. 4:18 Can think of essentially equivalent to unit tests for coding.
  35. 4:23 This is how you verify your agent works.
  36. 4:27 Then after evaluation, if everything looks fine, you usually have the ship basically where you deploy this agent to production.
  37. 4:38 Again, this can be a code update.
  38. 4:40 This can be a direct update on any agent platform or again, your local harness agents.
  39. 4:48 Then comes the online part.
  40. 4:51 This is where then the agent is continuously monitored for issues and based on certain trigger conditions, then you can start automatic diagnostics.
  41. 5:01 Again, this can be based on the volume of traces that your agent generates or weekly or daily jobs.
  42. 5:11 Then we go into diagnosis stage.
  43. 5:13 This is where you collect all the failures for your agents and do structured root cause analysis to then understand where the failures are coming from.
  44. 5:25 Once you understand and categorize the failures, then you can finally go on to the optimization stage.
  45. 5:31 This is where then you create, let's say, very specific changes or mutations for your agents to deal with the found failure modes.
  46. 5:43 And then the whole cycle repeats again.
  47. 5:46 You evaluate and if everything looks good, then you can deploy again.
  48. 5:52 Now we will maybe do a deep dive on each stage, what that entails.
  49. 6:01 So before we continue, Burak, we have two passes here.
  50. 6:07 Like one is the cold start path and one is basically existing features.

Open at this second