Videos vy7o1g2iHY8
How I deleted 95% of my agent skills and got better results — Nick Nisi, WorkOS
Scene timeline
43 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 233
- whisperx 233
- chunks
- 32
- from 233 cues
- keyframes
- 19
- kept of 43 captured
- frames with text
- 19
- 246 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 4.0 MB
- word timings on 233 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 01:20 | 1m 51s |
stt |
done | — | 2026-08-10 01:22 | 26s |
chunk |
done | — | 2026-08-10 01:22 | 0s |
text_embed |
done | — | 2026-08-10 19:45 | 0s |
keyframe |
done | — | 2026-08-10 01:22 | 1m 24s |
ocr |
done | — | 2026-08-10 01:24 | 6s |
frame_embed |
done | — | 2026-08-10 19:45 | 4s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- The Bottleneck1.00
- AlEngineer0.99
- EUROPE1.00
-
- History1.00
- Extenslons0.99
- Window0.99
- Fri Apr 10 10:300.98
- Chat1.00
- The Bottleneck1.00
- Nick Nisi · WorkOS · 20+ open source repos ·8 languages0.95
- AIE1.00
- ★1.00
- ★1.00
- ★0.99
- Engineering the future of Al1.00
- AlEngineer0.99
-
- View1.00
- Tabs1.00
- Bookmarks1.00
- History1.00
- Extensions0.99
- Window Help0.98
- Fri Apr 10 10:310.99
- Chat0.90
- The Bottleneck1.00
- Nick Nisi·WorkOS·20+ open source repos ·8 languages0.99
- ★1.00
- One agent at a time doesn't scale.0.98
- Products built for humans don't work0.99
- AIE1.00
- Every session started with the same ten1.00
- for agents.1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- minutes of orientation.1.00
- Agents are consumers now too.1.00
- AlEngineer0.97
- EUROPE1.00
- AlEngineer0.98
-
- Flle0.91
- View1.00
- Tabs1.00
- Bookmarks0.99
- History1.00
- Extensions1.00
- Window1.00
- Help1.00
- Fri Apr 10 10:320.95
- 国0.68
- Chat0.91
- Case· A harness for orchestrating coding agents0.97
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Writes code, runs1.00
- Implement1.00
- GATE1.00
- Fresh eyes,0.95
- Verify1.00
- GATE1.00
- Principles, code1.00
- Review1.00
- GATE1.00
- Evidence check,1.00
- Close1.00
- Analyzes run,1.00
- Retro1.00
- tests1.00
- scenario tests1.00
- quality1.00
- opens PR1.00
- proposes fixes1.00
- Revision loop on failure0.99
- Structured feedback, not blind retry. Budget: 2 cycles max.1.00
- Engineering the future of Al0.99
- AlEngineer0.99
-
- File0.88
- Edit1.00
- View1.00
- Tabs1.00
- Bookmarks1.00
- History1.00
- Extensions0.95
- Window1.00
- Help1.00
- Fri Apr 10 10:330.96
- 国0.54
- Chat0.97
- Case· A harness for orchestrating coding agents0.97
- ★1.00
- ★0.90
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Writes code, runs1.00
- Implement1.00
- GATE1.00
- Fresh eyes,0.96
- Verify1.00
- GATE1.00
- Principles, code1.00
- Review1.00
- GATE1.00
- Evidence check,1.00
- Close1.00
- Analyzes run,1.00
- Retro1.00
- tests1.00
- scenario tests1.00
- quality1.00
- opens PR1.00
- proposes fixes1.00
- Revision loop on failure1.00
- Structured feedback, not blind retry. Budget: 2 cycles max.0.99
- AlEngineer0.97
- EUROPE1.00
- neer1.00
-
- Extenslons0.97
- Window0.97
- Help1.00
- Chat1.00
- "The Agent Learned to Lie"0.99
- WHAT THE AGENT DID0.99
- touch.case-tested1.00
- AIE1.00
- # Marker created. No tests run. PR went through.0.98
- ★1.00
- ★1.00
- Engineering the future of Al0.98
- jineer0.99
-
- Flle0.88
- View1.00
- Tabs1.00
- Bookmarks1.00
- History1.00
- Extensions0.98
- Fri Apr 10 10:350.98
- Chat0.95
- "The Agent Learned to Lie'0.99
- WHAT THE AGENT DID1.00
- WHAT THE GATE NOW REQUIRES1.00
- touch .case-tested1.00
- echo "$TEST_OUTPUT" | sha256sum > .case/slug/testec0.97
- AIE1.00
- # Marker created. No tests run. PR went through.0.98
- # Piped output. Paxsed results. Tamper-proof hash.0.98
- ★1.00
- ★1.00
- The agent stopped lying — not because I asked nicely,0.98
- but because lying became harder than just running the tests.0.99
- AlEngineer0.98
- EUROPE1.00
-
- History1.00
- Extensions0.97
- Window Help0.99
- Fri Apr 10 10:350.97
- Chat0.90
- "Integration Complete"1.00
- $ workos install1.00
- *★★0.51
- Analyzing project... TanStack Start detected1.00
- AIE1.00
- Installing @workos-inc/authkit...0.99
- Configuring middleware...1.00
- ★1.00
- ★1.00
- ★1.00
- Integration Complete0.98
- $ pnpm build1.00
- ERROR: start.ts has implicit contract with bundler0.99
- Build failed.1.00
- Engineering the future of Al1.00
-
- "Integration Complete"1.00
-
- Tabs1.00
- History1.00
- Extensions0.99
- Window Help0.99
- 国0.70
- Chat1.00
- 10,739 Lines of Nothing1.00
- V1: AUTO-GENERATED GUIDES1.00
- V2: HUMAN-CURATED GOTCHAS0.99
- 10,7391.00
- 5531.00
- AIE1.00
- lines of scaffolding1.00
- lines of gotcha lists0.98
- ★1.00
- ★1.00
- ★1.00
- × 68 min per scenario0.98
- ✓ 6 min per scenario0.94
- × 38 error-fix cycles0.98
- √ Zero errors0.96
- MORE TOKENS. WORSE RESULTS.0.97
- DELETED 95%. PERFORMANCE WENT UP.0.99
- The model already knows how to code.1.00
- It doesn't know where your landmines are.1.00
- Braintrust1.00
- WorkOS OpenAI0.94
Transcript
233 cues· 3,381 words· 17,254 chars
- 0:14 All right, good morning, everyone.
- 0:16 Welcome to my talk, Building AI Systems That Ship.
- 0:19 I'm Nick Nisi, and I work at WorkOS.
- 0:22 We've got a booth downstairs.
- 0:23 Come check us out and talk to us.
- 0:25 I would be happy to chat.
- 0:27 But let me start that over.
- 0:28 Hi, I'm the bottleneck.
- 0:31 I'm a DX engineer at WorkOS.
- 0:33 And I work on 20 plus repos across eight different languages.
- 0:39 It's all of our SDKs and open source things that we have.
- 0:43 And it's like AuthKit Next.js, AuthKit React, WorkOS Node, WorkOS Kotlin, WorkOS Ruby, PHP.
- 0:53 Everywhere so there's a lot to do across a lot of different things and I'm really good at Working on those and I've gotten really good over the last eight months of working with those via agents So I haven't written a line of code myself in probably eight months I've gotten really good at just scaling that with agents and then reviewing what they do and instructing them and getting the work done faster and better while still maintaining good quality and
- 1:21 But there was a big problem.
- 1:23 Doing that with one agent at a time across all of these repos, I'm just constantly context switching over and over and over, and it just gets harder and harder, and that's okay, but the problem is that for every one of those, there's like this little bit of setup time that I'm doing each time, which is like,
- 1:41 giving it 10 minutes of my time to set up and establish the problem.
- 1:44 Let's look at this GitHub issue.
- 1:45 Let's look at this linear ticket.
- 1:47 Let's take a look at this Slack thread and figure out what's going on and see if we can reproduce the issue and then go.
- 1:53 So that was a lot of my time just spent dealing with the agent, getting it basically the context that I already have, and then getting it to work on it from there.
- 2:03 Now, on the other side, I'm also working on products that we want to build for agents.
- 2:09 Because while I said I'm a developer experience engineer, the developer is still the most important in my job.
- 2:15 But increasingly, the pipeline to get to that developer is through agents.
- 2:19 And so I see the agentic experience as being equally as important, because that's how we're going to get in front of the developers.
- 2:27 So there's two different ways I needed to go AI native, and two different directions for that.
- 2:33 So on the internal side, building that, I started building this project called Case.
- 2:38 This is a harness.
- 2:39 If you've read Ryan Lepopolo's Harness Engineering, it's that.
- 2:44 Just kind of took those ideas and started building them.
- 2:47 Basically, give it a GitHub issue, a PR, a Slack thread, a linear ticket, anything, and I could just point it at it, and it could figure out the context that it needs and go.
- 2:59 And then it wouldn't stop until it has a PR with evidence that it actually did what I asked it to, or what the problem was, or fixed what the issue was.
- 3:07 But most importantly, it had to provide that evidence.
- 3:11 And this originally started as a Claude skill, because why not?
- 3:15 I thought Claude could do anything.
- 3:17 And it was working really well.
- 3:18 But as it got more complex,
- 3:21 the context drop became very real.
- 3:23 It would just start forgetting things or skipping over tasks.
- 3:25 And I would ask Claude, why did you do that?
- 3:27 He's like, oh yeah, you told me to do that.
- 3:29 I decided not to.
- 3:30 Not great.
- 3:31 So I rebuilt it on top of Py and using a TypeScript state machine to facilitate going through and stepping through these agents.
- 3:39 So it has five different agents in it.
- 3:41 An implementer, a verifier, a reviewer, a closer, and a retro agent.
- 3:46 And those are important, but they're not the most important thing.
- 3:48 The most important piece of case is the gates in between that.
- 3:52 And that's what the state machine really enforces, is the checks in between everything.
- 3:59 So when we implement something, we can't move on to the reviewer until the verifier verifies it.
- 4:05 And once the reviewer reviews it, if there's any issues, it has to send it back to the implementer to do those.
loading
Chapters
- 0:00 Introduction
- 1:22 The challenge of context switching with agents
- 2:33 Introducing Case: A harness for agentic workflows
- 3:33 Rebuilding with a TypeScript state machine
- 4:45 The critical importance of evidence-based verification
- 5:59 Applying agentic principles to the WorkOS CLI
- 7:44 Lessons in documentation: Generating skills from docs
- 8:52 Why more data (10,000 lines) led to worse performance
- 9:36 The impact of using evals to measure accuracy
- 10:40 Key takeaway: Enforce with code, not just prompts
- 12:41 Treating failures as bugs in the harness system
- 14:39 Advice for building agentic-ready products
- 16:01 Final summary: Replacing trust with evidence