read-only demo

Videos kZsf_Sfm7RU

The Missing Layer After Launch - Raphael Kalandadze, Wandero AI

index_state ready data_status ok

AI Engineer· published 2026-07-05· 0:19:33· en-US· indexed 2026-08-11 05:14

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:41, 1 of 1 keyframes kept
  2. Shot 1, 0:41 to 1:25, 1 of 1 keyframes kept
  3. Shot 2, 1:25 to 1:49, 1 of 1 keyframes kept
  4. Shot 3, 1:49 to 2:22, 1 of 1 keyframes kept
  5. Shot 4, 2:22 to 2:49, 1 of 1 keyframes kept
  6. Shot 5, 2:49 to 3:17, 0 of 1 keyframes kept
  7. Shot 6, 3:17 to 3:24, 1 of 1 keyframes kept
  8. Shot 7, 3:24 to 3:29, 1 of 1 keyframes kept
  9. Shot 8, 3:29 to 3:42, 1 of 1 keyframes kept
  10. Shot 9, 3:42 to 4:10, 1 of 1 keyframes kept
  11. Shot 10, 4:10 to 4:37, 0 of 1 keyframes kept
  12. Shot 11, 4:37 to 5:06, 1 of 1 keyframes kept
  13. Shot 12, 5:06 to 5:46, 1 of 1 keyframes kept
  14. Shot 13, 5:46 to 6:17, 1 of 1 keyframes kept
  15. Shot 14, 6:17 to 6:49, 0 of 1 keyframes kept
  16. Shot 15, 6:49 to 7:17, 1 of 1 keyframes kept
  17. Shot 16, 7:17 to 7:45, 0 of 1 keyframes kept
  18. Shot 17, 7:45 to 8:13, 0 of 1 keyframes kept
  19. Shot 18, 8:13 to 8:37, 1 of 1 keyframes kept
  20. Shot 19, 8:37 to 9:05, 1 of 1 keyframes kept
  21. Shot 20, 9:05 to 9:46, 1 of 1 keyframes kept
  22. Shot 21, 9:46 to 10:10, 1 of 1 keyframes kept
  23. Shot 22, 10:10 to 10:30, 1 of 1 keyframes kept
  24. Shot 23, 10:30 to 10:59, 1 of 1 keyframes kept
  25. Shot 24, 10:59 to 11:28, 0 of 1 keyframes kept
  26. Shot 25, 11:28 to 11:51, 1 of 1 keyframes kept
  27. Shot 26, 11:51 to 11:59, 1 of 1 keyframes kept
  28. Shot 27, 11:59 to 12:04, 1 of 1 keyframes kept
  29. Shot 28, 12:04 to 12:20, 1 of 1 keyframes kept
  30. Shot 29, 12:20 to 12:46, 1 of 1 keyframes kept
  31. Shot 30, 12:46 to 13:11, 0 of 1 keyframes kept
  32. Shot 31, 13:11 to 13:41, 1 of 1 keyframes kept
  33. Shot 32, 13:41 to 14:12, 0 of 1 keyframes kept
  34. Shot 33, 14:12 to 14:14, 0 of 1 keyframes kept
  35. Shot 34, 14:14 to 14:15, 1 of 1 keyframes kept
  36. Shot 35, 14:15 to 14:17, 1 of 1 keyframes kept
  37. Shot 36, 14:17 to 14:51, 1 of 1 keyframes kept
  38. Shot 37, 14:51 to 15:01, 1 of 1 keyframes kept
  39. Shot 38, 15:01 to 15:07, 0 of 1 keyframes kept
  40. Shot 39, 15:07 to 15:18, 1 of 1 keyframes kept
  41. Shot 40, 15:18 to 15:24, 1 of 1 keyframes kept
  42. Shot 41, 15:24 to 15:38, 1 of 1 keyframes kept
  43. Shot 42, 15:38 to 15:47, 1 of 1 keyframes kept
  44. Shot 43, 15:47 to 16:07, 0 of 1 keyframes kept
  45. Shot 44, 16:07 to 16:10, 0 of 1 keyframes kept
  46. Shot 45, 16:10 to 16:12, 1 of 1 keyframes kept
  47. Shot 46, 16:12 to 16:59, 1 of 1 keyframes kept
  48. Shot 47, 16:59 to 17:00, 1 of 1 keyframes kept
  49. Shot 48, 17:00 to 17:02, 1 of 1 keyframes kept
  50. Shot 49, 17:02 to 17:22, 1 of 1 keyframes kept
  51. Shot 50, 17:22 to 17:40, 0 of 1 keyframes kept
  52. Shot 51, 17:40 to 18:06, 1 of 1 keyframes kept
  53. Shot 52, 18:06 to 18:32, 1 of 1 keyframes kept
  54. Shot 53, 18:32 to 18:59, 1 of 1 keyframes kept
  55. Shot 54, 18:59 to 19:32, 1 of 1 keyframes kept

55 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
232
whisperx 232
chunks
36
from 232 cues
keyframes
42
kept of 55 captured
frames with text
42
1,036 lines read
chapters
0
from the source metadata
keyframe bytes
4.1 MB
word timings on 232 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:11 1m 05s
stt done 2026-08-11 05:12 21s
chunk done 2026-08-11 05:13 0s
text_embed done 2026-08-11 05:13 0s
keyframe done 2026-08-11 05:13 1m 00s
ocr done 2026-08-11 05:14 19s
frame_embed done 2026-08-11 05:14 7s

Frames, and what the machine read

  • 0:16 #0 done4 line(s)

    shot 0·sharpness 876.2

    1. AI ENGINEER WORLD'S FAIR 20261.00
    2. The Missing Layer1.00
    3. After Launch0.98
    4. Everything important starts after you ship.1.00
  • 1:07 #1 done9 line(s)

    shot 1·sharpness 1288.4

    1. THE EASY PART1.00
    2. We shipped it in 3 weeks.1.00
    3. 3 wks0.94
    4. ~300K1.00
    5. ~$35K1.00
    6. TO FIRST PRODUCT1.00
    7. LINES1.00
    8. SPENT1.00
    9. Shipping is fast now. That was the easy part.0.99
  • 1:37 #2 done6 line(s)

    shot 2·sharpness 936.5

    1. AFTER YOU LAUNCH0.99
    2. Monitor it?0.99
    3. Understand it?1.00
    4. Improve it?1.00
    5. Find the holes?1.00
    6. Normal software never really asked these.1.00
  • 2:18 #3 done5 line(s)

    shot 3·sharpness 1000.0

    1. WHY IT'S HARD0.98
    2. Not a few features.0.98
    3. Almost anything.0.99
    4. A normal app has 5 features and 3 buttons. An agent can do whatever you ask.1.00
    5. If it can do almost anything, almost anything can break.1.00
  • 2:41 #4 done11 line(s)

    shot 4·sharpness 1153.5

    1. THE REAL PROBLEM0.99
    2. How do you even know it's healthy?0.97
    3. thousands of conversations almost infinite tasks0.99
    4. Monitor it?0.97
    5. Understand it?1.00
    6. PRODUCT1.00
    7. AGENT1.00
    8. Improve it?1.00
    9. Find the holes?1.00
    10. You lose the feel for your oun system.0.99
    11. You lose the feel for your own system. Getting it back is the goal.0.99
  • 3:14 #5 skipped

    shot 5·duplicate of #4

  • 3:22 #6 done22 line(s)

    shot 6·sharpness 1654.3

    1. YOU'RE NOT THE ONLY ONE0.99
    2. Article1.00
    3. Harrison Chase0.99
    4. @hwchase171.00
    5. Repeat1.00
    6. Valdate the fix in production0.97
    7. 5. Online Evals0.97
    8. 2. Annotation Queues0.99
    9. Let you review and label traces0.87
    10. 4. Experiments1.00
    11. 3. Datasets1.00
    12. improve benavior0.98
    13. ncorporate examples for testing0.94
    14. You don't know what your agent0.97
    15. will do until it's in production1.00
    16. Q440.85
    17. t7 790.81
    18. 4071.00
    19. th 157K0.92
    20. When you ship traditional software to production, you have a good sense1.00
    21. of what to expect. Users click buttons. fill out forms. navigate through0.96
    22. The hard part starts after you ship — you don't know what your agent does until it's live.0.99
  • 3:28 #7 done2 line(s)

    shot 7·sharpness 498.2

    1. WHAT CHANGED1.00
    2. This is a new problem.1.00
  • 3:37 #8 done4 line(s)

    shot 8·sharpness 886.3

    1. 01·NON-DETERMINISTIC0.98
    2. Same input, different path.0.99
    3. Endless tasks.1.00
    4. You can't list it — so you can't pre-test it.0.99
  • 3:56 #9 done9 line(s)

    shot 9·sharpness 1576.8

    1. 02·INVISIBLE FAILURE0.98
    2. The failure hides itself.1.00
    3. Claude marked features "complete" — without0.99
    4. Agents reported success while the system state said1.00
    5. checking they actually worked.1.00
    6. the opposite.1.00
    7. "EFFECTIVE HARNESSES FOR LONG-RUNNING AGENTS,"ANTHROPIC0.99
    8. NORTHEASTERN RED-TEAM1.00
    9. Nothing crashes. Nothing turns red. Nothing knows.0.99
  • 4:18 #10 skipped

    shot 10·duplicate of #9

  • 4:51 #11 done5 line(s)

    shot 11·sharpness 787.6

    1. 03· HUGE TOOL SURFACE0.97
    2. One task.1.00
    3. Hundreds of tools.0.98
    4. It writes code, runs the terminal, calls other companies' services.0.99
    5. Every tool is a new, quiet way to fail.1.00
  • 5:22 #12 done5 line(s)

    shot 12·sharpness 1076.2

    1. 04· DONE ≠ HAPPY0.95
    2. "Finished" is not "helped."1.00
    3. A "technically successful" response can still fail the task.1.00
    4. HARRISON CHASE· LANGCHAIN0.99
    5. An agent can succeed and still be wrong.1.00
  • 5:49 #13 done5 line(s)

    shot 13·sharpness 1035.4

    1. THE SHIFT1.00
    2. Operating an agent is itself1.00
    3. an agentproblem.1.00
    4. "Noise or bug?" . "Trace the cause" . "Root cause or symptom?" — all reasoning tasks.0.99
    5. So I put agents on the operations.1.00
  • 6:36 #14 skipped

    shot 14·duplicate of #13

  • 7:00 #15 done18 line(s)

    shot 15·sharpness 1370.8

    1. THE CLOSED LOOP0.99
    2. Production1.00
    3. Logs &0.99
    4. Trajectory1.00
    5. Code path0.98
    6. Diagnosis1.00
    7. sessions1.00
    8. traces1.00
    9. Next fix1.00
    10. Dashboard1.00
    11. review + merge1.00
    12. Human1.00
    13. PR review0.99
    14. (agent)1.00
    15. (log-monitor)1.00
    16. Pull request1.00
    17. iterate until ready1.00
    18. The agents watch and draft. The human decides.0.98
  • 7:28 #16 skipped

    shot 16·duplicate of #15

  • 7:48 #17 skipped

    shot 17·duplicate of #15

  • 8:18 #18 done4 line(s)

    shot 18·sharpness 740.0

    1. THE PAYOFF1.00
    2. The fastest loop1.00
    3. we've ever had.1.00
    4. You feel the system improve in near-real-time.0.99
  • 8:48 #19 done11 line(s)

    shot 19·sharpness 1831.7

    1. HOW I HANDLE IT0.94
    2. Four operating agents.1.00
    3. 011.00
    4. log-monitor — detect & fix fast (reactive)0.99
    5. 021.00
    6. PR-review — gate the fix (reactive)0.97
    7. 031.00
    8. session-analyzer — feel the health (reactive)0.99
    9. 041.00
    10. QA / computer-use — go test it (proactive · roadmap)0.99
    11. One example — not a recipe.0.97
  • 9:37 #20 done31 line(s)

    shot 20·sharpness 1059.5

    1. AGENT 01·LOG-MONITOR0.99
    2. log-monitor agent0.99
    3. What it can read1.00
    4. DETECT1.00
    5. Production logs0.99
    6. pull last hour1.00
    7. filter noise0.98
    8. ignore-list0.98
    9. Agent trajectories1.00
    10. ANALYZE1.00
    11. Codebase0.99
    12. pull trajectory1.00
    13. clone repo1.00
    14. read code1.00
    15. Pull request0.99
    16. Database (read-only)1.00
    17. DECIDE1.00
    18. description1.00
    19. mermaid flow0.93
    20. does the user end up stuck?1.00
    21. Traces1.00
    22. evidence1.00
    23. labels1.00
    24. so it can really explore0.97
    25. the full signal surface -0.94
    26. FIX0.99
    27. write patch0.97
    28. open PR0.99
    29. slack alert0.98
    30. severnity0.91
    31. PR link0.99
  • 9:56 #21 done4 line(s)

    shot 21·sharpness 961.1

    1. LOG-MONITOR - THE PR IT OPENED0.98
    2. fix: guard inbound-webhook read during conversation state change1.00
    3. #20.82
    4. A clear write-up, a flow diagram, the evidence — review it at a glance.0.98
  • 10:27 #22 done26 line(s)

    shot 22·sharpness 1200.1

    1. LOG-MONITOR - SLACK0.97
    2. Our bug → it opens a PR and tells me0.99
    3. wandero APP 3.38 PM0.89
    4. Log Monitor heads-up1.00
    5. Environment1.00
    6. Period1.00
    7. production0.98
    8. Last 1h (event at 10:14 UTC)1.00
    9. 1 issue fixed (HIGH – permanent email loss)::0.95
    10. Two inbound Outlook/SharePoint emails were permanently dropped (message not ingested ond Micro0.97
    11. VARCHAR(255)→StringDutoRightTruncationErrorcrashed Outlook ingestion.- 2 occurrences0.94
    12. PR Created: #1087 - store email Message-ID / in-Reply-To as TEXT0.96
    13. Outside billing issue → just a heads-up, no PR0.96
    14. wandero APP 6:06 AM0.93
    15. Log Monitor follow-up1.00
    16. Environment1.00
    17. Period1.00
    18. production1.00
    19. Last 1h (02:06 UTC)0.98
    20. Provider billing issue - RECURRED, not resolved:0.99
    21. • porse_document hit LlamaParse 402 *exceeded maximum credits" again at 01:45 UTC (org e27df545).0.97
    22. • The earlier read was resolved after a clean streak – that is now contradicted. The credit ceiling is still being hit intermittently (11 failures / 5 orgs over 48h).0.99
    23. • Agent recovered via fallback each time - no user blocked, work persisted.0.98
    24. Next step: verify the LlamaParse (LlamaCloud) plan top-up actually landed and raise the credit limit. No code fx needed - the 402→ERROR alert is working as intended.0.98
    25. No PR opened (operational bill0.97
    26. It alerts only when it matters — and knows the difference.0.98
  • 10:44 #23 done13 line(s)

    shot 23·sharpness 854.8

    1. AGENT 02· PR-REVIEW0.98
    2. Review agent1.00
    3. ( dtifferent angle )0.96
    4. checkout the branch1.00
    5. Pull0.98
    6. run focused tests1.00
    7. verdict:1.00
    8. request1.00
    9. READY?1.00
    10. typecheck + lint0.97
    11. root cause, or symptom?1.00
    12. Human1.00
    13. review + merge0.99

Transcript

232 cues· 3,211 words· 16,803 chars

  1. 0:00 All right.
  2. 0:01 So we built an agent.
  3. 0:02 You launched it.
  4. 0:03 Everything works pretty well in the demo.
  5. 0:05 Everyone is happy.
  6. 0:06 But now let me ask you a few simple questions.
  7. 0:09 So how do you know if it's actually working out there?
  8. 0:12 How do you watch across hundreds or thousands of real conversations every day?
  9. 0:16 How do you feel or understand the health of the system?
  10. 0:20 How do you make it better?
  11. 0:21 How do you find the holes that you don't know are there yet?
  12. 0:24 And that's the thing, right?
  13. 0:25 So most of the talks about the agents end up the moment when you ship.
  14. 0:29 So we built it, it worked, the end.
  15. 0:31 But I think the shipping is the moment when the real work begins.
  16. 0:35 And somehow only a few people are talking about that, and I'm calling it a missing layer.
  17. 0:40 So let's dive into it.
  18. 0:42 So that's the world that we are living.
  19. 0:44 So you can create the whole product.
  20. 0:46 You can create a whole startup in a couple of days, in a couple of weeks.
  21. 0:49 You can write hundreds of thousands of lines of code.
  22. 0:52 You can spend a lot of tokens.
  23. 0:54 And to be honest, this is the easiest part today with the help of the latest models.
  24. 0:58 But I think the shipping is the moment when the real work begins, because you need to close the loop as soon as possible.
  25. 1:05 So after your launch, you need to have some control and understanding of the system.
  26. 1:09 And from my experience, the loop is at least as important as the product itself, sometimes even more, because the tight feedback is the one that helps you to make the product better every single day.
  27. 1:22 And that's the missing layer, and that's what the rest of this talk is all about.
  28. 1:26 So what happens after you launch?
  29. 1:28 And this is not something surprising.
  30. 1:30 We had the same questions in the classical old software.
  31. 1:33 You need to monitor what is happening.
  32. 1:35 You need to understand how it behaves.
  33. 1:37 You need to have some logs to detect the problems and fix them.
  34. 1:42 And for agentic systems, each one of those are even harder.
  35. 1:45 And sometimes they turn into something genuinely new.
  36. 1:50 So let's talk about why this is hard and why this is hard now.
  37. 1:54 So the agent is on a normal software, right?
  38. 1:56 You don't have a few features, several buttons.
  39. 1:59 You don't have a predefined flow that you can test before you go to the live.
  40. 2:06 And the coverage is endless.
  41. 2:07 So think about like Cloud Code or Codex.
  42. 2:09 They can do a giant range of stuff wherever the user needs.
  43. 2:13 And most of the agents do the same, right?
  44. 2:16 So you give the instructions and they can handle it.
  45. 2:18 And you cannot write all the conversations in advance.
  46. 2:23 So this leads to the deepest part of the problem, the part that keeps me up all night, which is you lose the feel for your own system.
  47. 2:31 So after you build the product, you need to have some kind of understanding, does it get better or worse?
  48. 2:38 So you need to monitor, understand what is happening.
  49. 2:41 And the problem is that the normal safety nets don't save you here.
  50. 2:47 And believe me, we try a few stuff.

Open at this second