Videos 7P0elyLIxXo
What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip
Scene timeline
16 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 79
- whisperx 79
- chunks
- 13
- from 79 cues
- keyframes
- 13
- kept of 16 captured
- frames with text
- 13
- 187 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 1.4 MB
- word timings on 79 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 03:36 | 1m 20s |
stt |
done | — | 2026-08-11 03:37 | 9s |
chunk |
done | — | 2026-08-11 03:37 | 0s |
text_embed |
done | — | 2026-08-11 03:37 | 0s |
keyframe |
done | — | 2026-08-11 03:37 | 22s |
ocr |
done | — | 2026-08-11 03:38 | 5s |
frame_embed |
done | — | 2026-08-11 03:38 | 2s |
Frames, and what the machine read
-
- AI ENGINEER· TALK0.97
- What does0.94
- "done" even mean?0.99
- Get 100× more work done with agents0.98
- Lessons from Paperclip's liveness model0.99
- 01 / 13 · Title0.93
- @dotta1.00
- e Paperclip0.91
-
- THE HOOK0.99
- Agents are great at0.99
- finishing the wrong thing.0.99
- PRODUCED FAST - TRUSTED SLOWLY1.00
- PR1.00
- comment1.00
- doc1.00
- summary1.00
- approval request1.00
- The scary failures are when your agent says "done,"1.00
- and hands you something you can't trust.1.00
- 02 / 13 · Hook0.92
- @dotta1.00
- e Paperclip0.91
-
- ·THESIS0.90
- DONE ≠a status0.96
- DONE = a reliance claim:0.99
- this artifact, to this standard,0.99
- with this evidence, checked by this verifier,1.00
- with this owner holding the residual risk,1.00
- authorizing this next action.0.99
- done = artifact + scope + standard + evidence + verifier + authority + risk + next0.99
- 03 / 13 · Thesis card0.96
- @dotta1.00
- @ Paperclip0.94
-
- LEVELS OF DONE0.99
- There are levels of done.0.97
- 011.00
- 021.00
- 030.99
- 041.00
- 051.00
- 061.00
- Produced1.00
- Author-1.00
- Spec-1.00
- Accepted1.00
- Accountable1.00
- Operationally1.00
- done1.00
- done1.00
- proven1.00
- Most agent systems stop at author-done and pretend they reached accountable.0.99
- 04 / 13 · Done ladder0.93
- @dotta1.00
- Paperclip1.00
-
- WHY REVIEW BREAKS AT 1000×0.99
- 10tasks/day1.00
- 10,000tasks/day1.00
- A human catches the bad assumptions1.00
- Approval theater1.00
- Humans cannot read every line.1.00
- Humans design and operate the control system:0.98
- define standards, review exceptions, sample the process, liable for the outcome1.00
- 05 / 13 . Human review breaks0.97
- @dotta1.00
- e Paperclip0.92
-
- CONTROL SYSTEM1.00
- Statuses are an execution contract—0.99
- not decorative labels.1.00
- paperclip/AIE Talk · Board0.97
- 4 issues1.00
- ●TODO0.90
- 1IN PROGRESS1.00
- 1IN REVIEW1.00
- ● DONE0.84
- 11.00
- PAP-119441.00
- PAP-119410.98
- PAP-119391.00
- PAP-119291.00
- Cut 5-minute edit1.00
- Build slide deck1.00
- Rehearse 12-min cut0.98
- Write talk script1.00
- 06 / 13 . Paperclip control plane0.99
- @dotta1.00
- Paperclip1.00
-
- LIVENESS × ASSURANCE0.99
- Assurance without liveness becomes a review queue.0.99
- Liveness without assurance becomes fast wrongness.0.99
- A trustworthy agent system preserves both.0.99
- LIVENESS1.00
- ASSURANCE1.00
- Will the work continue to the next0.97
- Has enough been established to permit1.00
- legitimate action?1.00
- that action?0.99
- next action · blocker · wakeup · bounded loop0.95
- evidence · verifier · authority · risk ownership0.98
- DONE1.00
- The done claim is the controlled transition between the two.1.00
- 08 / 13 . Liveness × assurance0.95
- @dotta1.00
- e Paperclip0.94
-
- THE MECHANISMS1.00
- The control system.1.00
- Clear transitions1.00
- First-class blockers1.00
- Every heartbeat must leave a next action0.99
- Dependency is explicit.1.00
- Interactions & approvals0.99
- Reviewers & approvers1.00
- with an audit trail.1.00
- Enforced review by QA, CTO, etc..0.98
- Watchdogs · monitors · recovery0.98
- Evidence ≠ liveness1.00
- Ensure the /goal.1.00
- Comments, docs, work products are evidence.0.99
- Child issues + plan decomposition1.00
- Break work into smaller units.1.00
- 07 / 13 · Paperclip mechanisms0.98
- @dotta1.00
- e Paperclip0.96
-
- THREE INVARIANTS1.00
- Productive work continues0.99
- 011.00
- Only real blockers stop work1.00
- 021.00
- Infinite loops are bounded1.00
- 031.00
- 09 / 13 . The three invariants0.99
- @dotta1.00
- e Paperclip0.96
-
- THE CODE MOMENT0.98
- {0.95
- "artifact": "talk-draft.md",0.99
- "scope": "5-15 minute AIE talk",1.00
- paperclip/issue0.99
- "standard": "acceptance criteria + source accuracy",0.99
- PAP-119411.00
- DONE0.99
- "evidence": ["source check", "timing", "QA read"],0.98
- Build branded slide deck0.99
- "verifier": "QA",1.00
- todo + / in_progress + in_review done0.93
- parentie0.90
- AlE talk - PAP-119260.88
- "authority": "CTO",0.97
- blockedy0.96
- — none open0.94
- "residualRisk": "speaker-specific polish",1.00
- nextAction0.90
- unblocks PAP-119360.98
- "allowedNextAction": "build slides"1.00
- }0.96
- Treat done as an object, not a boolean.0.99
- 10 / 13 · Done-claim JSON0.98
- @dotta1.00
- Paperclip0.99
-
- STEAL THIS CHECKLIST1.00
- Define done levels before agents start.1.00
- Require evidence, not just confidence.1.00
- Separate verifier from author.1.00
- Route by risk and reversibility.0.99
- Keep blockers explicit.0.99
- Bound recovery loops.1.00
- Manage allowed next actions.1.00
- 11 / 13 · Steal-this checklist0.98
- @dotta1.00
- e Paperclip0.95
-
- STEAL THIS CHECKLIST1.00
- Define done levels before agents start.1.00
- Require evidence, not just confidence.1.00
- Separate verifier from author.1.00
- Route by risk and reversibility.0.99
- Keep blockers explicit.0.99
- Bound recovery loops.1.00
- Manage allowed next actions.1.00
- 11 / 13 · Steal-this checklist0.98
- @dotta1.00
- e Paperclip0.95
-
- THE POINT1.00
- Make "done" specific enough that0.99
- someone else can safely build on it.0.98
- The output multiplier comes after the trust protocol.1.00
- 12 / 13 · The point0.94
- @dotta1.00
- e Paperclip0.92
Transcript
79 cues· 1,323 words· 7,310 chars
- 0:00 An agent opens a pull request.
- 0:01 It passes the tests.
- 0:03 It updates the documentation.
- 0:04 It closes the issue and comments, looks done to me.
- 0:07 But is it actually done?
- 0:09 Is it done enough to merge?
- 0:10 Is it done enough to deploy?
- 0:12 Is it done enough to announce to your customers?
- 0:15 These are fundamentally different operational claims, and most agent systems just flatten it to a single green check mark.
- 0:21 I'm Dota.
- 0:22 I'm the creator of Paperclip, and I'm going to give you some hard-earned lessons that we've learned in creating Paperclip's liveness model.
- 0:28 What does done even mean?
- 0:30 Here's the thing.
- 0:31 Programming is solved, and agents can now produce more code and documentation faster than any human can ever verify.
- 0:39 And this actually gives us a new failure mode, is that agents can actually create more work than humans have time to verify.
- 0:45 So we need a way to verify that our agents are done more than just letting them check a checkbox.
- 0:50 Done doesn't mean that an agent just changed the status of a task being done.
- 0:54 Saying that something is done is actually a bundle of claims.
- 0:58 You're saying that an artifact was produced, that you have evidence that the task is actually complete, and you have a rubric in which you can verify against.
- 1:07 You know exactly who the owner is for the next step, and you know exactly what the next step is.
- 1:14 There's different levels to how done something is.
- 1:16 The producer might claim something is complete, but you need to have a reviewer, another party that looks at it and finds no obvious issues.
- 1:23 You want to verify and make sure that the evidence actually meets a specified standard.
- 1:28 You want to make sure that a person who is authorized to approve it actually approves that the work is done.
- 1:34 And you want to make sure that there's somebody who actually stands behind the decision.
- 1:38 And ideally what you want is that the outcome has actually survived real world conditions.
- 1:45 Because exhaustive human verification fails at high volume.
- 1:49 You might be able to verify a few tasks per day, but essentially if you have humans verifying all the tasks and they have to sign off on it, eventually what you just get is a form of verification theater.
- 2:03 What you need is a protocol for defining how tasks actually progress through a system.
- 2:09 You want to make sure that tasks are always kept moving, but they don't get stuck into invalid states.
- 2:16 You need a control plane that actually has the execution of the tasks being tied to specific contracts and constraints about what the system will do and what agents it will hand off your next task to.
- 2:30 Because really what you're trying to play against is this idea around keeping work moving, but also having it verified.
- 2:37 When a task has been reviewed by a human, you get the assurance that it's correct.
- 2:41 But having a human verify it means that the task is dead in its tracks.
- 2:46 You also want to keep liveness.
- 2:48 Liveness means that the work is continuing with no blockers.
- 2:52 And you're always trying to keep these two things in balance.
- 2:54 If you have tasks that are completely alive with no approvals, then what you get is this classic AI slop because you're producing a lot of things with kind of no quality control and it's worse than creating nothing after a long period of time.
- 3:08 But if you have pure review, then you have this enormous review queue where humans can't actually review it by hand anyway.
- 3:17 These agents will be creating far more than you can ever actually review.
- 3:22 And so we have to find a way to tease apart the bundle of claims that are involved in saying a task is done.
- 3:27 With Paperclip, we have a number of mechanisms to keep this going.
- 3:31 You might think that you can easily just write a for loop over your task manager and have your agents work, but quickly you'll find that falls apart.
- 3:39 As soon as you start integrating task dependency trees, blockers, multiple agents, idipotent checkouts, like locks on checkouts, you find that this tension between liveness and verification actually gets quite complicated.
- 3:55 There's really three invariants that are extremely important when you're thinking about what you want out of a control plane for your agentic work.
- 4:03 You want to ensure that productive work continues.
- 4:06 You want to make sure that only real blockers stop work.
- 4:09 And you want to make sure that infinite loops are bounded.
- 4:12 In Paperclip, we have built a number of mechanisms to deal with this problem.
- 4:17 So for example, every time you have a task, there's clear transitions to what the next state could be.
loading