read-only demo

Videos zU4EagB311U

Agents Need Feature Flags - Sachin Gupta

index_state ready data_status ok

AI Engineer· published 2026-07-18· 0:19:16· en-US· indexed 2026-08-11 00:55

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:26, 1 of 1 keyframes kept
  2. Shot 1, 0:26 to 0:53, 0 of 1 keyframes kept
  3. Shot 2, 0:53 to 1:19, 0 of 1 keyframes kept
  4. Shot 3, 1:19 to 1:54, 1 of 1 keyframes kept
  5. Shot 4, 1:54 to 2:29, 0 of 1 keyframes kept
  6. Shot 5, 2:29 to 3:04, 1 of 1 keyframes kept
  7. Shot 6, 3:04 to 3:39, 0 of 1 keyframes kept
  8. Shot 7, 3:39 to 4:07, 1 of 1 keyframes kept
  9. Shot 8, 4:07 to 4:36, 0 of 1 keyframes kept
  10. Shot 9, 4:36 to 5:10, 1 of 1 keyframes kept
  11. Shot 10, 5:10 to 5:45, 0 of 1 keyframes kept
  12. Shot 11, 5:45 to 6:08, 1 of 1 keyframes kept
  13. Shot 12, 6:08 to 6:52, 1 of 1 keyframes kept
  14. Shot 13, 6:52 to 7:31, 1 of 1 keyframes kept
  15. Shot 14, 7:31 to 7:59, 1 of 1 keyframes kept
  16. Shot 15, 7:59 to 8:28, 0 of 1 keyframes kept
  17. Shot 16, 8:28 to 9:09, 1 of 1 keyframes kept
  18. Shot 17, 9:09 to 9:28, 0 of 1 keyframes kept
  19. Shot 18, 9:28 to 9:56, 1 of 1 keyframes kept
  20. Shot 19, 9:56 to 10:24, 0 of 1 keyframes kept
  21. Shot 20, 10:24 to 10:54, 1 of 1 keyframes kept
  22. Shot 21, 10:54 to 11:24, 0 of 1 keyframes kept
  23. Shot 22, 11:24 to 11:55, 1 of 1 keyframes kept
  24. Shot 23, 11:55 to 12:27, 0 of 1 keyframes kept
  25. Shot 24, 12:27 to 13:01, 1 of 1 keyframes kept
  26. Shot 25, 13:01 to 13:34, 0 of 1 keyframes kept
  27. Shot 26, 13:34 to 14:00, 1 of 1 keyframes kept
  28. Shot 27, 14:00 to 14:27, 0 of 1 keyframes kept
  29. Shot 28, 14:27 to 14:53, 0 of 1 keyframes kept
  30. Shot 29, 14:53 to 15:19, 0 of 1 keyframes kept
  31. Shot 30, 15:19 to 15:46, 1 of 1 keyframes kept
  32. Shot 31, 15:46 to 16:14, 0 of 1 keyframes kept
  33. Shot 32, 16:14 to 16:56, 1 of 1 keyframes kept
  34. Shot 33, 16:56 to 17:46, 1 of 1 keyframes kept
  35. Shot 34, 17:46 to 18:21, 1 of 1 keyframes kept
  36. Shot 35, 18:21 to 18:57, 0 of 1 keyframes kept
  37. Shot 36, 18:57 to 19:16, 1 of 1 keyframes kept

37 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
225
whisperx 225
chunks
36
from 225 cues
keyframes
20
kept of 37 captured
frames with text
20
433 lines read
chapters
0
from the source metadata
keyframe bytes
4.1 MB
word timings on 225 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 00:52 1m 10s
stt done 2026-08-11 00:53 19s
chunk done 2026-08-11 00:54 0s
text_embed done 2026-08-11 00:54 1s
keyframe done 2026-08-11 00:54 44s
ocr done 2026-08-11 00:54 15s
frame_embed done 2026-08-11 00:55 3s

Frames, and what the machine read

  • 0:20 #0 done5 line(s)

    shot 0·sharpness 980.6

    1. AGENTS NEED FEATURE0.99
    2. FLAGS.1.00
    3. What web and mobile engineering learned the hard way, Al teams are0.98
    4. about to learn faster — and at higher cost.0.99
    5. Sachin Gupta1.00
  • 0:50 #1 skipped

    shot 1·duplicate of #0

  • 1:01 #2 skipped

    shot 2·duplicate of #0

  • 1:30 #3 done17 line(s)

    shot 3·sharpness 2239.2

    1. WHAT YOU DO0.96
    2. 21.00
    3. TODAY1.00
    4. I0.99
    5. Most Al teams ship to 100% of users on every deploy.1.00
    6. The moment your prompt change merges, every user — every request — sees the new behavior. There is no middle0.99
    7. ground.1.00
    8. 100%0.91
    9. Instantly Affected1.00
    10. of users see every behavior change the moment it ships0.98
    11. What ships globally on one deploy:1.00
    12. Prompt rewrites & system-instruction edits0.99
    13. New tool additions0.98
    14. Model swaps1.00
    15. Memory-policy changes1.00
    16. Autonomy upgrades1.00
    17. "Your 'small' prompt tweak just broke a chunk of users. You found out from a Discord screenshot."1.00
  • 2:05 #4 skipped

    shot 4·duplicate of #3

  • 3:00 #5 done27 line(s)

    shot 5·sharpness 2414.8

    1. EVIDENCE1.00
    2. I0.98
    3. This already happened. Here's what it cost.0.99
    4. 10.99
    5. 21.00
    6. Apr 2025 — Cursor "Sam"0.97
    7. Jul 2025 — Replit Agent0.96
    8. Support bot confidently cited a single-device login policy that didn't exist.0.99
    9. Day 9 of a 12-day vibe-coding experiment, mid code-freeze. Agent ignored1.00
    10. Customers compared notes on HN. CEO Michael Truell apologized publicly.0.99
    11. ALL-CAPS "no changes" instructions, dropped the production DB, then1.00
    12. generated 4,000+ fake records to conceal it.1.00
    13. Cost: public apology + session bug fix0.99
    14. Cost: weeks of rebuild0.99
    15. 31.00
    16. 41.00
    17. Nov 2025 — LangChain A2A loop0.99
    18. Apr 2026 — PocketOS0.98
    19. 4-agent A2A pipeline. Analyzer + Verifier locked in a 264-hour loop with no0.99
    20. Cursor + Claude Opus grabbed an unrelated API token from another file,0.99
    21. termination predicate. Found via billing-dashboard threshold, not the agent0.99
    22. ran a Railway GraphQL drop on prod. Co-located backups gone too.0.99
    23. system.1.00
    24. Cost: $47K · 11 days0.96
    25. Cost: 9 seconds1.00
    26. Sources: vectara/awesome-agent-failures · The Register·Fortune · Washington Post0.97
    27. 31.00
  • 3:32 #6 skipped

    shot 6·duplicate of #5

  • 4:04 #7 done17 line(s)

    shot 7·sharpness 1894.3

    1. BORROWED DISCIPLINE1.00
    2. I0.87
    3. What web and backend engineering paid for, the hard way.1.00
    4. Canary Releases1.00
    5. Segment Targeting0.99
    6. Ship to 1% first. Watch the metrics. If they hold, expand. If0.98
    7. Different behavior for different users — beta cohort, paying0.99
    8. they don't, roll back without redeploying.1.00
    9. tier, region, risk class. Not everyone sees the same thing.1.00
    10. X0.93
    11. Kill Switches0.99
    12. Rollout Monitoring1.00
    13. A pre-wired off-toggle for any feature that touches users.1.00
    14. Treat every behavior change as an experiment with its own0.99
    15. Pulled in seconds, not in deploy cycles.0.99
    16. error-rate dashboard. No dashboard, no rollout.1.00
    17. 41.00
  • 4:32 #8 skipped

    shot 8·duplicate of #7

  • 5:00 #9 done24 line(s)

    shot 9·sharpness 2381.0

    1. WHY AGENTS ARE DIFFERENT0.97
    2. I0.98
    3. Six behavior surfaces a CRUD app doesn't have.0.99
    4. 1·Prompts1.00
    5. 2·Tools0.98
    6. The system prompt is your most behavior-altering code. It changes0.99
    7. Every tool the agent can call is a newly authorized action. Tools come0.99
    8. weekly — sometimes daily — with no compile step, no type check,1.00
    9. and go, each one expanding or shifting the blast radius.0.99
    10. no test suite.1.00
    11. 3·Models0.99
    12. 4·Memory1.00
    13. Model-of-the-week swaps change personality, latency, cost, and1.00
    14. What the agent remembers across sessions silently changes behavior1.00
    15. refusal patterns overnight. Same prompt, different behavior.0.99
    16. over time — creating drift no deploy ever triggered.0.99
    17. 5·Autonomy0.99
    18. 6·Sub-agents1.00
    19. Suggest vs. auto-approve vs. auto-execute. The single largest blast-0.99
    20. Spawned children inherit the parent's flags. Or they should. Mc0.98
    21. radius dial — and the easiest to silently misconfigure.0.99
    22. systems don't enforce this — so junior agents run with full0.99
    23. permissions.1.00
    24. 51.00
  • 5:31 #10 skipped

    shot 10·duplicate of #9

  • 6:05 #11 done21 line(s)

    shot 11·sharpness 1475.3

    1. THE TAXONOMY1.00
    2. 10.90
    3. Six flag types. One playbook.0.98
    4. Prompt Variant1.00
    5. Model-Routing1.00
    6. Autonomy-Level1.00
    7. A/B prompt routing to alternate1.00
    8. Route traffic to specific models0.99
    9. Suggest /approve / execute dial0.99
    10. behaviors1.00
    11. X0.65
    12. Tool-Access1.00
    13. Memory-Policy1.00
    14. Kill Switch1.00
    15. Per-tool,per-segment authorization1.00
    16. Define persistence, expiry, and sharing0.99
    17. Pre-wired off, agent-wide & per-0.98
    18. controls1.00
    19. surface1.00
    20. Each flag maps directly to a behavior surface. None require building a new flag backend — your existing feature-flag infrastructure handl0.98
    21. 61.00
  • 6:30 #12 done20 line(s)

    shot 12·sharpness 2606.6

    1. FLAG TYPE 1 / 60.95
    2. I0.99
    3. Prompt variant flags1.00
    4. What it does1.00
    5. Routes different users — or different requests — to different0.99
    6. Mandatory When1.00
    7. system-prompt versions, on the fly, without a deploy.1.00
    8. Operating across regulatory regimes — EU GDPR, UK,0.99
    9. US state laws. One prompt cannot legally serve all three0.99
    10. jurisdictions simultaneously.1.00
    11. Example:prompt.support_agent.tone1.00
    12. Different regions demand different disclosure language,1.00
    13. beta_cohort → experimental-v3 (concise, action-first)1.00
    14. different refusal behaviors, and different data-handling0.99
    15. paid_tier → v2 (warm, expansive)0.98
    16. promises. Without prompt variant flags, you either ship the0.99
    17. strictest version everywhere (degrading experience) or risk0.99
    18. everyone else → v1 (stable, well-tested)0.98
    19. non-compliance in every market.0.98
    20. 71.00
  • 7:00 #13 done19 line(s)

    shot 13·sharpness 2185.8

    1. FLAG TYPE 2 / 60.94
    2. I0.95
    3. Tool-access flags1.00
    4. What it does1.00
    5. Authorizes — or revokes — specific tools per user segment, per0.99
    6. Mandatory When1.00
    7. request type, or per risk class. The tool exists; whether the agent1.00
    8. Agent has money-moving, data-deleting, or compliance-1.00
    9. can call it is the flag.1.00
    10. sensitive tools. Scope per customer tier, or pay the0.99
    11. AML/SOX bill.0.99
    12. Typical controls1.00
    13. Per-tool boolean —on/off for any capability0.99
    14. Also prevents: broken tool ships, prompt + send_email mass-1.00
    15. mail incidents, beta tools leaking to prod via config drift.0.99
    16. Per-segment lists — different tools for different cohorts0.98
    17. Per-risk-class gates — high-risk actions require elevated context1.00
    18. Time-of-day windows — cap costly tools during peak hours0.98
    19. 81.00
  • 7:40 #14 done21 line(s)

    shot 14·sharpness 2489.0

    1. FLAG TYPE 3 / 60.94
    2. I0.97
    3. Model-routing flags1.00
    4. What it does0.99
    5. Decides which model handles which traffic — and lets you0.99
    6. Mandatory When1.00
    7. migrate, fallback, or canary new models without code changes.1.00
    8. Healthcare (HIPAA BAA), finance (PCI), multi-jurisdiction0.99
    9. PlI. Route PHI to the right endpoint or violate the law.0.99
    10. Example:model.summarizer.route0.98
    11. When a provider has an outage — or pulls a model — your1.00
    12. high-cost segment → large frontier model (best quality)0.99
    13. fallback should be a one-flag flip, not a hot patch written1.00
    14. free tier → cheap-fast model (cost ceiling enforced)0.99
    15. during the incident.1.00
    16. on incident → fallback-stable (one-flip degraded mode)0.99
    17. Router1.00
    18. Frontier1.00
    19. Cheap-Fast1.00
    20. Fallback-Stable1.00
    21. 91.00
  • 8:03 #15 skipped

    shot 15·duplicate of #14

  • 9:00 #16 done25 line(s)

    shot 16·sharpness 2606.9

    1. FLAG TYPE 4 / 60.98
    2. I0.98
    3. Memory-policy flags0.99
    4. What it does1.00
    5. Controls what the agent remembers across sessions — what0.99
    6. Mandatory When1.00
    7. persists, what expires, what's shared, what's deletable.1.00
    8. Serving EU users. GDPR Art. 17 (right to erasure) and0.98
    9. Art. 22 (automated decisions) require per-user memory0.99
    10. control, including deletion on demand.1.00
    11. memory.retention1.00
    12. → session only / 30 days / forever0.94
    13. memory.scope1.00
    14. →per-user/per-tenant/global0.99
    15. Memory is the most under-appreciated source of behavior0.99
    16. drift. The same user, asking the same question, gets different0.99
    17. memory.write_enabled1.00
    18. → agent can persist or not0.99
    19. answers on Monday and Friday — because of what1.00
    20. memory.user_visible0.99
    21. 0.98
    22. user can inspect/delete their0.99
    23. accumulated between them.1.00
    24. own1.00
    25. 101.00
  • 9:13 #17 skipped

    shot 17·duplicate of #12

  • 9:31 #18 done37 line(s)

    shot 18·sharpness 2359.5

    1. FLAG TYPE 6 / 60.99
    2. I0.98
    3. The kill switch.0.99
    4. Pre-wired off, agent-wide and per-surface. No deploy. No restart. No code change at incident time.0.99
    5. No deploy required1.00
    6. No restart required0.99
    7. No code change required1.00
    8. Kill in seconds, not in pipelines.0.98
    9. In-flight requests respect the flag at the next1.00
    10. Pre-wired during the agent's design phase —0.99
    11. decision point.1.00
    12. not as an afterthought.0.99
    13. WITHOUT IT1.00
    14. 21.00
    15. 31.00
    16. 41.00
    17. LangChain A2A· $47K0.97
    18. PocketOS · 9 seconds0.95
    19. Replit · day 90.92
    20. OpenClaw · Meta0.96
    21. Analyzer + Verifier loop, no termination0.99
    22. Cursor + Claude wiped prod DB and co-0.99
    23. Agent ignored ALL-CAPS "no changes,"0.99
    24. Director of Alignment's agent ignored1.00
    25. predicate. Billing dashboard found it,1.00
    26. located backups with one Railway0.99
    27. deleted prod, fabricated 4,000+ fake1.00
    28. stop commands. Physical disconnect1.00
    29. not the agent.0.99
    30. mutation.1.00
    31. records.1.00
    32. required.1.00
    33. MANDATORY WHEN0.99
    34. Any high-risk Al operating in the EU after Aug 2, 2026. EU Al Act Article 14(4)(e) requires a "stop button1.00
    35. law.1.00
    36. Sources: vectara/awesome-agent-failures · TechCrunch · Fortune · The Register · EU Al Act Article 14(4)(e)0.97
    37. 121.00
  • 10:02 #19 skipped

    shot 19·duplicate of #18

  • 10:33 #20 done29 line(s)

    shot 20·sharpness 2573.0

    1. DEMO 1 / 20.89
    2. I0.96
    3. Mid-conversation flag flip0.98
    4. S ET U P April 2025: Cursor's "Sam" support bot confidently cited a single-device login policy that didn't exist. A tool-access flag is the fix.1.00
    5. SUPPORT CONVERSATION0.99
    6. agent-flags1.00
    7. admin1.00
    8. Mock data1.00
    9. Email engineering about the outage?1.00
    10. tool.send_email1.00
    11. OFF1.00
    12. Sure — drafting that email now...0.97
    13. AUDIT1.00
    14. disabled at1.00
    15. 14:32:08 UTC0.99
    16. FLAG FLIPPED0.99
    17. ·tool.send_email0.99
    18. by1.00
    19. sachin1.00
    20. AFTER FLIP. next turn, same session0.96
    21. scope1.00
    22. all in-flight conversations0.99
    23. I can draft this for you, but I'm not able to send emails right now.0.99
    24. Want me to copy the draft into your clipboard instead?1.00
    25. applies at1.00
    26. next decision point1.00
    27. active sessions1.00
    28. 121.00
    29. 131.00
  • 11:00 #21 skipped

    shot 21·duplicate of #20

  • 11:30 #22 done38 line(s)

    shot 22·sharpness 1950.3

    1. DEMO 2 / 20.90
    2. I0.95
    3. The kill switch — stopping a runaway agent mid-sentence.0.98
    4. S E T U P November 2025: A 4-agent LangChain A2A pipeline looped 11 days, $47K. Billing dashboard found it, not the agent system. Here's the 30-0.99
    5. second version.1.00
    6. TOOL CALLS / MINUTE0.97
    7. Illustrative simulation1.00
    8. agent.global_kill1.00
    9. 3001.00
    10. 2501.00
    11. EMERGENCY1.00
    12. STOP1.00
    13. 2001.00
    14. 1501.00
    15. 14:32:221.00
    16. 1001.00
    17. 501.00
    18. pressed1.00
    19. -501.00
    20. -2m0.99
    21. -1m1.00
    22. 0m1.00
    23. +10s1.00
    24. +20s1.00
    25. +25s1.00
    26. +30s1.00
    27. +45s1.00
    28. +1m1.00
    29. Mitigated in 30 seconds. No deploy.0.98
    30. T+0:001.00
    31. Runaway loop detected — 12 tool calls in 90s0.99
    32. T+0:151.00
    33. On-call Slack alert from the rate guard1.00
    34. T+0:221.00
    35. agent.global_kill = true· flag propagates0.98
    36. T+0:261.00
    37. In-flight agents see flag mid-turn, shut down gracefully1.00
    38. 141.00
  • 12:02 #23 skipped

    shot 23·duplicate of #22

Transcript

225 cues· 2,836 words· 15,835 chars

  1. 0:00 Hello everyone, I'm Sachin Gupta and I'm a backend engineer.
  2. 0:04 And today we are going to talk about agent need feature flags.
  3. 0:08 If you have been a backend engineer for any length of time, you already know these tools.
  4. 0:14 Things like canaries, segment targeting, kill switches.
  5. 0:18 Your craft has had them for over a decade and none of them is new.
  6. 0:23 The boring infrastructure that keep deploy safe is already a solved problem.
  7. 0:29 what is new is that we are shipping the most behavior changing systems we have ever built agents that send money agent that send mail agent that modify databases agent that spawn child processes and we are shipping them with none of that infrastructure we are shipping them the way web team used to ship in 2008.
  8. 0:52 Over the next few minutes, here is the plan.
  9. 0:55 I will walk you through with the six Slack types that agents specifically need.
  10. 1:00 I will show you two live demo storyboards.
  11. 1:03 The first is flipping a tool mid-conversation.
  12. 1:06 The second is stopping a runaway agent mid-sentence.
  13. 1:10 Then I will cover a rollout playbook with the numbers your team should track from day one.
  14. 1:17 Let's go.
  15. 1:20 here is the situation today the moment your prompt change merges a hundred percent of your users see the new behavior there is no canary no segment and no rollback button look at what goes out under those small all or nothing rules
  16. 1:38 We get prompt rewrite, new tool addition, model swapping, memory policy changes, autonomy upgrades, system instruction edits, and we get all of it globally and instantly.
  17. 1:51 Web teams stopped doing this back in 2012, and they stopped doing it for changes that were less risky than this.
  18. 1:59 The story that you actually hear from teams almost word for word is that it's just a small prompt peak, maybe broke a couple of a chunk of users,
  19. 2:09 And then finally people are finding it out on discard links or tech talks, or maybe another social media platform.
  20. 2:18 So that is the failure mode.
  21. 2:20 This entire talk is built around.
  22. 2:22 And let me show you that it is not hypothetical.
  23. 2:31 These are the four named incident in the last 14 months.
  24. 2:35 The first one we have is CursorSAM that happened in April of 2025, where the support bot confidently told users about a policy that never existed.
  25. 2:46 The second one we have is Replit.
  26. 2:48 This was day nine of a 12-day wipe coding experiment.
  27. 2:54 Agent did not follow the instructions and ended up deleting the production database and then fabricated over 4,000 fake users to conceal what it had done.
  28. 3:03 The third one is LandChain.
  29. 3:06 It had a four agent pipeline, researcher, analyzer, verifier, and synthesizer where two of them ran in continuous loop and costed $47,000.
  30. 3:16 The fourth one is PocketOS where a developer was using cursor and cloud.
  31. 3:21 the AI coding agent grabbed an unrelated API token from another file, treated it as authoritative, and ran a Railway GraphQL prop on the production database.
  32. 3:34 On the bottom left, you will see the sources that I used to cite it.
  33. 3:42 Web engineers learned this lesson a decade ago.
  34. 3:44 Canary releases.
  35. 3:45 You ship to a few percentage of users.
  36. 3:47 You watch the metrics.
  37. 3:49 If it works, you expand.
  38. 3:50 If it doesn't, then you roll back.
  39. 3:53 Segment targeting, different behavior for different type of users.
  40. 3:56 Kill switches, pre-wired off toggles that take effect in seconds, not in deploy cycles.
  41. 4:02 Rollout monitoring, every change has its own error rate dashboard.
  42. 4:06 None of this is new.
  43. 4:07 The tooling already exists, like LaunchDoc, UnleashLift, or maybe your homegrown flag service.
  44. 4:13 This is already a solved problem.
  45. 4:14 The discipline is already there.
  46. 4:16 We just have to apply it.
  47. 4:19 But now the problem is that web feature flags covers one thing, whether a feature is on or off.
  48. 4:25 But agent has six behavior surfaces that a CRUD app does not have.
  49. 4:30 And each one needs its own kind of flag.
  50. 4:33 And in the next slide, we are going to see that.

Open at this second