Videos WLXxTaPagA8
Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD - Sumaiya Shrabony
Scene timeline
85 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 161
- whisperx 161
- chunks
- 20
- from 161 cues
- keyframes
- 78
- kept of 85 captured
- frames with text
- 47
- 902 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 12.6 MB
- word timings on 161 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 04:05 | 1m 47s |
stt |
done | — | 2026-08-11 04:06 | 11s |
chunk |
done | — | 2026-08-11 04:07 | 0s |
text_embed |
done | — | 2026-08-11 04:07 | 0s |
keyframe |
done | — | 2026-08-11 04:07 | 1m 07s |
ocr |
done | — | 2026-08-11 04:08 | 39s |
frame_embed |
done | — | 2026-08-11 04:08 | 12s |
Frames, and what the machine read
-
- AI ENGINEER WORLD'S FAIR 20260.99
- Every Solo Agent Builder1.00
- Eventually Reinvents a Worse1.00
- Version of Cl/CD0.97
- Agent demos end at the happy path. Agent operations start when the0.99
- happy path lies to you.0.99
- SUMAIYA SHRABONY1.00
- Technical Program Manager ·University of Colorado Denver0.99
- 19-skill open-source agent system·Ground Truth newsletter1.00
- github.com/safrin96/agentic-content-system1.00
-
- AI ENGINEER WORLD'S FAIR 20260.99
- Every Solo Agent Builder0.99
- Eventually Reinvents a Worse1.00
- Version of Cl/CD0.97
- Agent demos end at the happy path. Agent operations start when the0.99
- happy path lies to you.0.99
- SUMAIYA SHRABONY1.00
- Technical Program Manager · University of Colorado Denver0.98
- 19-skill open-source agent system· Ground Truth newsletter0.99
- github.com/safrin96/agentic-content-system1.00
-
- AI ENGINEER WORLD'S FAIR 20260.99
- Every Solo Agent Builder1.00
- Eventually Reinvents a Worse1.00
- Version of Cl/CD0.97
- Agent demos end at the happy path. Agent operations start when the0.99
- happy path lies to you.0.99
- SUMAIYA SHRABONY1.00
- Technical Program Manager · University of Colorado Denver0.99
- 19-skill open-source agent system· Ground Truth newsletter0.98
- github.com/safrin96/agentic-content-system1.00
-
- THE SYSTEM1.00
- A production workflow with seven0.99
- handoffs1.00
- Every handoff is a place where the system can silently lie to you.0.99
- CONTRACT BREAK1.00
- MISSED VIOLATION0.99
- FALSE PASS1.00
- Scheduler1.00
- Command1.00
- Research1.00
- →1.00
- Plan1.00
- Production Skill1.00
- →0.99
- Verifier1.00
- Reviewer1.00
- →1.00
- Output1.00
- This system runs every other Saturday. Reads from a knowledge vault. Creates research briefs.0.99
- Builds a content plan. Produces 12 pieces. Runs verifier passes, reviewer gates, deduplication.1.00
- Saves structured markdown outputs.1.00
- The content is not the point. The handoffs are the point.1.00
-
- THE SYSTEM0.99
- A production workflow with seven1.00
- handoffs1.00
- Every handoff is a place where the system can silently lie toyou.0.97
- CONTRACT BREAK0.98
- MISSED VIOLATION0.99
- FALSE PASS0.95
- Scheduler1.00
- →1.00
- Command1.00
- →1.00
- Research1.00
- →1.00
- Plan1.00
- →0.99
- Production Skill1.00
- →0.99
- Verifier1.00
- →0.99
- Reviewer1.00
- →1.00
- Output1.00
- This system runs every other Saturday. Reads from a knowledge vault. Creates research briefs.0.99
- Builds a content plan. Produces 12 pieces. Runs verifier passes, reviewer gates, deduplication.1.00
- Saves structured markdown outputs.1.00
- The content is not the point. The handoffs are the point.0.99
-
- THE SYSTEM1.00
- A production workflow with seven0.99
- handoffs1.00
- Every handoff is a place where the system can silently lie to you.0.99
- CONTRACT BREAK0.99
- MISSED VIOLATION1.00
- FALSE PASS1.00
- Scheduler1.00
- →1.00
- Command1.00
- →1.00
- Research1.00
- →1.00
- Plan1.00
- Production Skill0.98
- →0.99
- Verifier1.00
- →0.99
- Reviewer0.95
- →1.00
- Output1.00
- This system runs every other Saturday. Reads from a knowledge vault. Creates research briefs.0.99
- Builds a content plan. Produces 12 pieces. Runs verifier passes, reviewer gates, deduplication.0.99
- Saves structured markdown outputs.1.00
- The content is not the point. The handoffs are the point.0.99
-
- THE FIVE REINVENTIONS0.98
- What you'll accidentally rebuild1.00
- In roughly this order. Every time. Click a card to see the story or run1.00
- the auto-tour.1.00
- Auto-tour1.00
- SPEED1.00
- 1×0.99
- Both1.00
- Agent only0.99
- Highlight Cl/CD0.96
- 011.00
- Regression Checks1.00
- Scheduled Monitoring1.00
- You change a prompt. Something downstream breaks. You1.00
- Your cron job silently falls. You don't notice for a week. You0.97
- build a way to test if the output still matches the expected0.99
- build alerts.1.00
- shape.1.00
- ≈CI Monitoring0.96
- ≈Regression Testing0.98
-
- THE FIVE REINVENTIONS1.00
- What you'll accidentally rebuild0.99
- In roughly this order. Every time. Click a card to see the story or run1.00
- the auto-tour.1.00
- Auto-tour1.00
- SPEED1.00
- 1×0.90
- Both1.00
- Agent only1.00
- Highlight Cl/CD0.98
- 011.00
- 021.00
- Regression Checks1.00
- Scheduled Monitoring1.00
- You change a prompt. Something downstream breaks. You1.00
- Your cron job silently fails. You don't notice for a week. You1.00
- build a way to test if the output still matches the expected0.99
- build alerts.0.97
- shape.1.00
- ≈ CI Monitoring0.96
- ≈ Regression Testing1.00
- SHOW THE STORY1.00
- +0.95
-
- Wliat yuu Il acciucillally Tcvullu0.71
- In roughly this order. Every time. Click a card to see the story or run1.00
- the auto-tour.1.00
- Auto-tour1.00
- SPEED1.00
- 1x0.78
- Both1.00
- Agent only0.99
- Highlight Cl/CD0.96
- 011.00
- 021.00
- Regression Checks1.00
- Scheduled Monitoring0.98
- You change a prompt. Something downstream breaks. You0.99
- Your cron job silently fails. You don't notice for a week. You0.99
- build a way to test if the output still matches the expected0.99
- build alerts.1.00
- shape.1.00
- ≈ CI Monitoring0.96
- ≈ Regression Testing0.99
- SHOW THE STORY1.00
- +0.95
- SHOW THE STORY1.00
- +0.95
-
- 011.00
- Regression Checks1.00
- Scheduled Monitoring1.00
- You change a prompt. Something downstream breaks. You0.99
- Your cron job silently fails. You don't notice for a week. You0.98
- build a way to test if the output still matches the expected1.00
- build alerts.1.00
- shape.1.00
- = Cl Monitoring0.92
- ≈ Regression Testing0.99
- SHOW THE STORY0.99
- SHOW THE STORY1.00
- x0.93
- YOU NOTICE0.99
- A prompt tweak quietly mangles the1.00
- A0.88
- markdown structure two skills downstream.0.99
- YOU BUILD1.00
- A small harness that re-runs sample inputs0.99
- and diffs the output shape against last1.00
- week's.1.00
- ALREADY EXISTED AS0.99
- A regression test suite. With fixtures. Since0.99
- pytest was a thing.0.99
Transcript
161 cues· 1,508 words· 8,690 chars
- 0:00 Here's something nobody warns you about when you start building agents alone.
- 0:04 You think you're building prompts.
- 0:06 You think you're building skills or workflow.
- 0:10 If you build long enough, especially alone,
- 0:13 you will start building something completely different, something that looks suspiciously like CICD, except worse, because you're building it from scratch, one failure at a time.
- 0:27 I'm Sumaya.
- 0:27 I run a 19-scale cloud code agent system.
- 0:31 Writing, Research, Vault Sync, Analytics Sync, Hook, Transcript, and many more.
- 0:39 The most useful thing I learned from building this was not how to build better prompts.
- 0:45 It was recognizing the five controls I was rebuilding badly and what you can do instead.
- 0:51 Before I show you the problem, let me show you the system that taught me the problem.
- 0:56 This is the agent content system.
- 0:59 It is open source.
- 1:00 Link in the description.
- 1:02 It runs every other siren.
- 1:04 It reads from a knowledge vault, creates a research brief, builds content plan, produces 12 content placements, then runs verifier passes, reviewer gates,
- 1:15 deduplication, and finally saves the output as markdown files.
- 1:19 But here's the thing that matters for this talk, not the content.
- 1:23 What matters is that this system has seven handoffs.
- 1:27 Scheduler to command, command to research, research to content plan, content plan to production skill, production skill to verifier, verifier to reviewer, reviewer to the output pool.
- 1:40 Every single handoff is the place where the system can lie to you.
- 1:45 And if you're building the system alone, nobody catches the lies except you, usually after the damage is done.
- 1:52 Here's the pattern I want you to watch for in your own system.
- 1:56 If you build agents independently, you will rebuild these five things roughly in this order.
- 2:04 You change a prompt or a skill and something downstream breaks.
- 2:09 So you build a way to test whether the output still matches the expected stream.
- 2:14 Congratulations, you have reinvented regression testing.
- 2:19 So you set up a cron job or a scheduled task.
- 2:22 One day it silently fails, but you haven't noticed for a week.
- 2:27 So you build alerts.
- 2:29 You just reinvented CI monitoring.
- 2:32 one skill changes its output schema so three skills downstream break you decided to add a validation at the boundary because of it
- 2:42 You just reinvented contract testing.
- 2:45 An artifact looks done, but it shouldn't ship.
- 2:48 So you add a checkpoint before it goes to the ready folder.
- 2:53 You just reinvented staging environments.
- 2:55 Something goes wrong, but you cannot find out which prompt, which skill, or which handoff had the bad output.
- 3:04 So you start logging everything.
- 3:07 she just reinvented audit rails.
- 3:09 The reason the title says ORS version isn't because agents are software builds.
- 3:15 It's because you end up needing the exact same operational guarantees.
- 3:20 However, the agent systems give you none of it by default.
- 3:25 So you build them independently, the worst version without even realizing that you are building it.
- 3:33 The dangerous failure in an agent system is never a bad output.
- 3:39 A bad output is very easy to fix.
- 3:42 You glance at it and immediately you can understand it's a bad output.
- 3:46 The dangerous failure is a polished artifact that looks great at a glance.
- 3:53 However,
- 3:54 It will never pass your exit gates.
- 3:57 It uses the wrong voice pattern.
loading