read-only demo

Videos vy7o1g2iHY8

How I deleted 95% of my agent skills and got better results — Nick Nisi, WorkOS

index_state ready data_status ok

AI Engineer· published 2026-05-30· 0:17:42· en-US· indexed 2026-08-10 19:45

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:30, 1 of 1 keyframes kept
  5. Shot 4, 0:30 to 0:58, 1 of 1 keyframes kept
  6. Shot 5, 0:58 to 1:25, 0 of 1 keyframes kept
  7. Shot 6, 1:25 to 1:53, 0 of 1 keyframes kept
  8. Shot 7, 1:53 to 2:21, 1 of 1 keyframes kept
  9. Shot 8, 2:21 to 2:49, 0 of 1 keyframes kept
  10. Shot 9, 2:49 to 3:06, 0 of 1 keyframes kept
  11. Shot 10, 3:06 to 3:32, 1 of 1 keyframes kept
  12. Shot 11, 3:32 to 3:59, 1 of 1 keyframes kept
  13. Shot 12, 3:59 to 4:25, 0 of 1 keyframes kept
  14. Shot 13, 4:25 to 4:51, 0 of 1 keyframes kept
  15. Shot 14, 4:51 to 5:18, 1 of 1 keyframes kept
  16. Shot 15, 5:18 to 5:44, 0 of 1 keyframes kept
  17. Shot 16, 5:44 to 6:10, 1 of 1 keyframes kept
  18. Shot 17, 6:10 to 6:37, 1 of 1 keyframes kept
  19. Shot 18, 6:37 to 7:07, 1 of 1 keyframes kept
  20. Shot 19, 7:07 to 7:32, 0 of 1 keyframes kept
  21. Shot 20, 7:32 to 7:58, 0 of 1 keyframes kept
  22. Shot 21, 7:58 to 8:23, 1 of 1 keyframes kept
  23. Shot 22, 8:23 to 8:49, 0 of 1 keyframes kept
  24. Shot 23, 8:49 to 9:15, 0 of 1 keyframes kept
  25. Shot 24, 9:15 to 9:40, 0 of 1 keyframes kept
  26. Shot 25, 9:40 to 10:06, 0 of 1 keyframes kept
  27. Shot 26, 10:06 to 10:45, 1 of 1 keyframes kept
  28. Shot 27, 10:45 to 11:15, 0 of 1 keyframes kept
  29. Shot 28, 11:15 to 11:45, 0 of 1 keyframes kept
  30. Shot 29, 11:45 to 12:14, 0 of 1 keyframes kept
  31. Shot 30, 12:14 to 12:44, 0 of 1 keyframes kept
  32. Shot 31, 12:44 to 13:14, 1 of 1 keyframes kept
  33. Shot 32, 13:14 to 13:48, 1 of 1 keyframes kept
  34. Shot 33, 13:48 to 14:14, 1 of 1 keyframes kept
  35. Shot 34, 14:14 to 14:39, 0 of 1 keyframes kept
  36. Shot 35, 14:39 to 15:05, 0 of 1 keyframes kept
  37. Shot 36, 15:05 to 15:31, 0 of 1 keyframes kept
  38. Shot 37, 15:31 to 15:56, 0 of 1 keyframes kept
  39. Shot 38, 15:56 to 16:22, 0 of 1 keyframes kept
  40. Shot 39, 16:22 to 16:48, 0 of 1 keyframes kept
  41. Shot 40, 16:48 to 17:13, 0 of 1 keyframes kept
  42. Shot 41, 17:13 to 17:27, 1 of 1 keyframes kept
  43. Shot 42, 17:27 to 17:42, 1 of 1 keyframes kept

43 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
233
whisperx 233
chunks
32
from 233 cues
keyframes
19
kept of 43 captured
frames with text
19
246 lines read
chapters
13
from the source metadata
keyframe bytes
4.0 MB
word timings on 233 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 01:20 1m 51s
stt done 2026-08-10 01:22 26s
chunk done 2026-08-10 01:22 0s
text_embed done 2026-08-10 19:45 0s
keyframe done 2026-08-10 01:22 1m 24s
ocr done 2026-08-10 01:24 6s
frame_embed done 2026-08-10 19:45 4s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 827.9

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:28 #3 done3 line(s)

    shot 3·sharpness 337.2

    1. The Bottleneck1.00
    2. AlEngineer0.99
    3. EUROPE1.00
  • 0:44 #4 done13 line(s)

    shot 4·sharpness 1582.4

    1. History1.00
    2. Extenslons0.99
    3. Window0.99
    4. Fri Apr 10 10:300.98
    5. Chat1.00
    6. The Bottleneck1.00
    7. Nick Nisi · WorkOS · 20+ open source repos ·8 languages0.95
    8. AIE1.00
    9. 1.00
    10. 1.00
    11. 0.99
    12. Engineering the future of Al1.00
    13. AlEngineer0.99
  • 1:04 #5 skipped

    shot 5·duplicate of #4

  • 1:39 #6 skipped

    shot 6·duplicate of #4

  • 2:10 #7 done25 line(s)

    shot 7·sharpness 2233.6

    1. View1.00
    2. Tabs1.00
    3. Bookmarks1.00
    4. History1.00
    5. Extensions0.99
    6. Window Help0.98
    7. Fri Apr 10 10:310.99
    8. Chat0.90
    9. The Bottleneck1.00
    10. Nick Nisi·WorkOS·20+ open source repos ·8 languages0.99
    11. 1.00
    12. One agent at a time doesn't scale.0.98
    13. Products built for humans don't work0.99
    14. AIE1.00
    15. Every session started with the same ten1.00
    16. for agents.1.00
    17. 1.00
    18. 1.00
    19. 1.00
    20. 1.00
    21. minutes of orientation.1.00
    22. Agents are consumers now too.1.00
    23. AlEngineer0.97
    24. EUROPE1.00
    25. AlEngineer0.98
  • 2:32 #8 skipped

    shot 8·duplicate of #4

  • 3:04 #9 skipped

    shot 9·duplicate of #3

  • 3:24 #10 done38 line(s)

    shot 10·sharpness 1671.6

    1. Flle0.91
    2. View1.00
    3. Tabs1.00
    4. Bookmarks0.99
    5. History1.00
    6. Extensions1.00
    7. Window1.00
    8. Help1.00
    9. Fri Apr 10 10:320.95
    10. 0.68
    11. Chat0.91
    12. Case· A harness for orchestrating coding agents0.97
    13. AIE1.00
    14. 1.00
    15. 1.00
    16. 1.00
    17. Writes code, runs1.00
    18. Implement1.00
    19. GATE1.00
    20. Fresh eyes,0.95
    21. Verify1.00
    22. GATE1.00
    23. Principles, code1.00
    24. Review1.00
    25. GATE1.00
    26. Evidence check,1.00
    27. Close1.00
    28. Analyzes run,1.00
    29. Retro1.00
    30. tests1.00
    31. scenario tests1.00
    32. quality1.00
    33. opens PR1.00
    34. proposes fixes1.00
    35. Revision loop on failure0.99
    36. Structured feedback, not blind retry. Budget: 2 cycles max.1.00
    37. Engineering the future of Al0.99
    38. AlEngineer0.99
  • 3:38 #11 done44 line(s)

    shot 11·sharpness 1658.7

    1. File0.88
    2. Edit1.00
    3. View1.00
    4. Tabs1.00
    5. Bookmarks1.00
    6. History1.00
    7. Extensions0.95
    8. Window1.00
    9. Help1.00
    10. Fri Apr 10 10:330.96
    11. 0.54
    12. Chat0.97
    13. Case· A harness for orchestrating coding agents0.97
    14. 1.00
    15. 0.90
    16. AIE1.00
    17. 1.00
    18. 1.00
    19. 1.00
    20. 1.00
    21. 1.00
    22. Writes code, runs1.00
    23. Implement1.00
    24. GATE1.00
    25. Fresh eyes,0.96
    26. Verify1.00
    27. GATE1.00
    28. Principles, code1.00
    29. Review1.00
    30. GATE1.00
    31. Evidence check,1.00
    32. Close1.00
    33. Analyzes run,1.00
    34. Retro1.00
    35. tests1.00
    36. scenario tests1.00
    37. quality1.00
    38. opens PR1.00
    39. proposes fixes1.00
    40. Revision loop on failure1.00
    41. Structured feedback, not blind retry. Budget: 2 cycles max.0.99
    42. AlEngineer0.97
    43. EUROPE1.00
    44. neer1.00
  • 4:04 #12 skipped

    shot 12·duplicate of #11

  • 4:28 #13 skipped

    shot 13·duplicate of #10

  • 5:05 #14 done13 line(s)

    shot 14·sharpness 1566.4

    1. Extenslons0.97
    2. Window0.97
    3. Help1.00
    4. Chat1.00
    5. "The Agent Learned to Lie"0.99
    6. WHAT THE AGENT DID0.99
    7. touch.case-tested1.00
    8. AIE1.00
    9. # Marker created. No tests run. PR went through.0.98
    10. 1.00
    11. 1.00
    12. Engineering the future of Al0.98
    13. jineer0.99
  • 5:33 #15 skipped

    shot 15·duplicate of #14

  • 5:55 #16 done22 line(s)

    shot 16·sharpness 2128.3

    1. Flle0.88
    2. View1.00
    3. Tabs1.00
    4. Bookmarks1.00
    5. History1.00
    6. Extensions0.98
    7. Fri Apr 10 10:350.98
    8. Chat0.95
    9. "The Agent Learned to Lie'0.99
    10. WHAT THE AGENT DID1.00
    11. WHAT THE GATE NOW REQUIRES1.00
    12. touch .case-tested1.00
    13. echo "$TEST_OUTPUT" | sha256sum > .case/slug/testec0.97
    14. AIE1.00
    15. # Marker created. No tests run. PR went through.0.98
    16. # Piped output. Paxsed results. Tamper-proof hash.0.98
    17. 1.00
    18. 1.00
    19. The agent stopped lying — not because I asked nicely,0.98
    20. but because lying became harder than just running the tests.0.99
    21. AlEngineer0.98
    22. EUROPE1.00
  • 6:33 #17 done20 line(s)

    shot 17·sharpness 1896.1

    1. History1.00
    2. Extensions0.97
    3. Window Help0.99
    4. Fri Apr 10 10:350.97
    5. Chat0.90
    6. "Integration Complete"1.00
    7. $ workos install1.00
    8. *★★0.51
    9. Analyzing project... TanStack Start detected1.00
    10. AIE1.00
    11. Installing @workos-inc/authkit...0.99
    12. Configuring middleware...1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. Integration Complete0.98
    17. $ pnpm build1.00
    18. ERROR: start.ts has implicit contract with bundler0.99
    19. Build failed.1.00
    20. Engineering the future of Al1.00
  • 7:03 #18 done1 line(s)

    shot 18·sharpness 276.0

    1. "Integration Complete"1.00
  • 7:22 #19 skipped

    shot 19·duplicate of #17

  • 7:55 #20 skipped

    shot 20·duplicate of #14

  • 8:03 #21 done27 line(s)

    shot 21·sharpness 2263.7

    1. Tabs1.00
    2. History1.00
    3. Extensions0.99
    4. Window Help0.99
    5. 0.70
    6. Chat1.00
    7. 10,739 Lines of Nothing1.00
    8. V1: AUTO-GENERATED GUIDES1.00
    9. V2: HUMAN-CURATED GOTCHAS0.99
    10. 10,7391.00
    11. 5531.00
    12. AIE1.00
    13. lines of scaffolding1.00
    14. lines of gotcha lists0.98
    15. 1.00
    16. 1.00
    17. 1.00
    18. × 68 min per scenario0.98
    19. ✓ 6 min per scenario0.94
    20. × 38 error-fix cycles0.98
    21. √ Zero errors0.96
    22. MORE TOKENS. WORSE RESULTS.0.97
    23. DELETED 95%. PERFORMANCE WENT UP.0.99
    24. The model already knows how to code.1.00
    25. It doesn't know where your landmines are.1.00
    26. Braintrust1.00
    27. WorkOS OpenAI0.94
  • 8:39 #22 skipped

    shot 22·duplicate of #21

  • 9:09 #23 skipped

    shot 23·duplicate of #16

Transcript

233 cues· 3,381 words· 17,254 chars

  1. 0:14 All right, good morning, everyone.
  2. 0:16 Welcome to my talk, Building AI Systems That Ship.
  3. 0:19 I'm Nick Nisi, and I work at WorkOS.
  4. 0:22 We've got a booth downstairs.
  5. 0:23 Come check us out and talk to us.
  6. 0:25 I would be happy to chat.
  7. 0:27 But let me start that over.
  8. 0:28 Hi, I'm the bottleneck.
  9. 0:31 I'm a DX engineer at WorkOS.
  10. 0:33 And I work on 20 plus repos across eight different languages.
  11. 0:39 It's all of our SDKs and open source things that we have.
  12. 0:43 And it's like AuthKit Next.js, AuthKit React, WorkOS Node, WorkOS Kotlin, WorkOS Ruby, PHP.
  13. 0:53 Everywhere so there's a lot to do across a lot of different things and I'm really good at Working on those and I've gotten really good over the last eight months of working with those via agents So I haven't written a line of code myself in probably eight months I've gotten really good at just scaling that with agents and then reviewing what they do and instructing them and getting the work done faster and better while still maintaining good quality and
  14. 1:21 But there was a big problem.
  15. 1:23 Doing that with one agent at a time across all of these repos, I'm just constantly context switching over and over and over, and it just gets harder and harder, and that's okay, but the problem is that for every one of those, there's like this little bit of setup time that I'm doing each time, which is like,
  16. 1:41 giving it 10 minutes of my time to set up and establish the problem.
  17. 1:44 Let's look at this GitHub issue.
  18. 1:45 Let's look at this linear ticket.
  19. 1:47 Let's take a look at this Slack thread and figure out what's going on and see if we can reproduce the issue and then go.
  20. 1:53 So that was a lot of my time just spent dealing with the agent, getting it basically the context that I already have, and then getting it to work on it from there.
  21. 2:03 Now, on the other side, I'm also working on products that we want to build for agents.
  22. 2:09 Because while I said I'm a developer experience engineer, the developer is still the most important in my job.
  23. 2:15 But increasingly, the pipeline to get to that developer is through agents.
  24. 2:19 And so I see the agentic experience as being equally as important, because that's how we're going to get in front of the developers.
  25. 2:27 So there's two different ways I needed to go AI native, and two different directions for that.
  26. 2:33 So on the internal side, building that, I started building this project called Case.
  27. 2:38 This is a harness.
  28. 2:39 If you've read Ryan Lepopolo's Harness Engineering, it's that.
  29. 2:44 Just kind of took those ideas and started building them.
  30. 2:47 Basically, give it a GitHub issue, a PR, a Slack thread, a linear ticket, anything, and I could just point it at it, and it could figure out the context that it needs and go.
  31. 2:59 And then it wouldn't stop until it has a PR with evidence that it actually did what I asked it to, or what the problem was, or fixed what the issue was.
  32. 3:07 But most importantly, it had to provide that evidence.
  33. 3:11 And this originally started as a Claude skill, because why not?
  34. 3:15 I thought Claude could do anything.
  35. 3:17 And it was working really well.
  36. 3:18 But as it got more complex,
  37. 3:21 the context drop became very real.
  38. 3:23 It would just start forgetting things or skipping over tasks.
  39. 3:25 And I would ask Claude, why did you do that?
  40. 3:27 He's like, oh yeah, you told me to do that.
  41. 3:29 I decided not to.
  42. 3:30 Not great.
  43. 3:31 So I rebuilt it on top of Py and using a TypeScript state machine to facilitate going through and stepping through these agents.
  44. 3:39 So it has five different agents in it.
  45. 3:41 An implementer, a verifier, a reviewer, a closer, and a retro agent.
  46. 3:46 And those are important, but they're not the most important thing.
  47. 3:48 The most important piece of case is the gates in between that.
  48. 3:52 And that's what the state machine really enforces, is the checks in between everything.
  49. 3:59 So when we implement something, we can't move on to the reviewer until the verifier verifies it.
  50. 4:05 And once the reviewer reviews it, if there's any issues, it has to send it back to the implementer to do those.

Chapters

  1. 0:00 Introduction
  2. 1:22 The challenge of context switching with agents
  3. 2:33 Introducing Case: A harness for agentic workflows
  4. 3:33 Rebuilding with a TypeScript state machine
  5. 4:45 The critical importance of evidence-based verification
  6. 5:59 Applying agentic principles to the WorkOS CLI
  7. 7:44 Lessons in documentation: Generating skills from docs
  8. 8:52 Why more data (10,000 lines) led to worse performance
  9. 9:36 The impact of using evals to measure accuracy
  10. 10:40 Key takeaway: Enforce with code, not just prompts
  11. 12:41 Treating failures as bugs in the harness system
  12. 14:39 Advice for building agentic-ready products
  13. 16:01 Final summary: Replacing trust with evidence

Open at this second