read-only demo

Videos 7P0elyLIxXo

What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip

index_state ready data_status ok

AI Engineer· published 2026-07-12· 0:07:13· en-US· indexed 2026-08-11 03:38

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:30, 1 of 1 keyframes kept
  2. Shot 1, 0:30 to 0:49, 1 of 1 keyframes kept
  3. Shot 2, 0:49 to 1:13, 1 of 1 keyframes kept
  4. Shot 3, 1:13 to 1:44, 1 of 1 keyframes kept
  5. Shot 4, 1:44 to 2:03, 1 of 1 keyframes kept
  6. Shot 5, 2:03 to 2:29, 1 of 1 keyframes kept
  7. Shot 6, 2:29 to 2:58, 1 of 1 keyframes kept
  8. Shot 7, 2:58 to 3:27, 0 of 1 keyframes kept
  9. Shot 8, 3:27 to 3:54, 1 of 1 keyframes kept
  10. Shot 9, 3:54 to 4:13, 1 of 1 keyframes kept
  11. Shot 10, 4:13 to 4:44, 0 of 1 keyframes kept
  12. Shot 11, 4:44 to 5:14, 0 of 1 keyframes kept
  13. Shot 12, 5:14 to 5:55, 1 of 1 keyframes kept
  14. Shot 13, 5:55 to 6:22, 1 of 1 keyframes kept
  15. Shot 14, 6:22 to 6:48, 1 of 1 keyframes kept
  16. Shot 15, 6:48 to 7:13, 1 of 1 keyframes kept

16 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
79
whisperx 79
chunks
13
from 79 cues
keyframes
13
kept of 16 captured
frames with text
13
187 lines read
chapters
0
from the source metadata
keyframe bytes
1.4 MB
word timings on 79 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 03:36 1m 20s
stt done 2026-08-11 03:37 9s
chunk done 2026-08-11 03:37 0s
text_embed done 2026-08-11 03:37 0s
keyframe done 2026-08-11 03:37 22s
ocr done 2026-08-11 03:38 5s
frame_embed done 2026-08-11 03:38 2s

Frames, and what the machine read

  • 0:15 #0 done8 line(s)

    shot 0·sharpness 1050.8

    1. AI ENGINEER· TALK0.97
    2. What does0.94
    3. "done" even mean?0.99
    4. Get 100× more work done with agents0.98
    5. Lessons from Paperclip's liveness model0.99
    6. 01 / 13 · Title0.93
    7. @dotta1.00
    8. e Paperclip0.91
  • 0:47 #1 done14 line(s)

    shot 1·sharpness 1089.5

    1. THE HOOK0.99
    2. Agents are great at0.99
    3. finishing the wrong thing.0.99
    4. PRODUCED FAST - TRUSTED SLOWLY1.00
    5. PR1.00
    6. comment1.00
    7. doc1.00
    8. summary1.00
    9. approval request1.00
    10. The scary failures are when your agent says "done,"1.00
    11. and hands you something you can't trust.1.00
    12. 02 / 13 · Hook0.92
    13. @dotta1.00
    14. e Paperclip0.91
  • 1:10 #2 done11 line(s)

    shot 2·sharpness 1211.4

    1. ·THESIS0.90
    2. DONE ≠a status0.96
    3. DONE = a reliance claim:0.99
    4. this artifact, to this standard,0.99
    5. with this evidence, checked by this verifier,1.00
    6. with this owner holding the residual risk,1.00
    7. authorizing this next action.0.99
    8. done = artifact + scope + standard + evidence + verifier + authority + risk + next0.99
    9. 03 / 13 · Thesis card0.96
    10. @dotta1.00
    11. @ Paperclip0.94
  • 1:32 #3 done21 line(s)

    shot 3·sharpness 1021.6

    1. LEVELS OF DONE0.99
    2. There are levels of done.0.97
    3. 011.00
    4. 021.00
    5. 030.99
    6. 041.00
    7. 051.00
    8. 061.00
    9. Produced1.00
    10. Author-1.00
    11. Spec-1.00
    12. Accepted1.00
    13. Accountable1.00
    14. Operationally1.00
    15. done1.00
    16. done1.00
    17. proven1.00
    18. Most agent systems stop at author-done and pretend they reached accountable.0.99
    19. 04 / 13 · Done ladder0.93
    20. @dotta1.00
    21. Paperclip1.00
  • 1:54 #4 done11 line(s)

    shot 4·sharpness 1453.7

    1. WHY REVIEW BREAKS AT 1000×0.99
    2. 10tasks/day1.00
    3. 10,000tasks/day1.00
    4. A human catches the bad assumptions1.00
    5. Approval theater1.00
    6. Humans cannot read every line.1.00
    7. Humans design and operate the control system:0.98
    8. define standards, review exceptions, sample the process, liable for the outcome1.00
    9. 05 / 13 . Human review breaks0.97
    10. @dotta1.00
    11. e Paperclip0.92
  • 2:26 #5 done21 line(s)

    shot 5·sharpness 885.9

    1. CONTROL SYSTEM1.00
    2. Statuses are an execution contract—0.99
    3. not decorative labels.1.00
    4. paperclip/AIE Talk · Board0.97
    5. 4 issues1.00
    6. ●TODO0.90
    7. 1IN PROGRESS1.00
    8. 1IN REVIEW1.00
    9. ● DONE0.84
    10. 11.00
    11. PAP-119441.00
    12. PAP-119410.98
    13. PAP-119391.00
    14. PAP-119291.00
    15. Cut 5-minute edit1.00
    16. Build slide deck1.00
    17. Rehearse 12-min cut0.98
    18. Write talk script1.00
    19. 06 / 13 . Paperclip control plane0.99
    20. @dotta1.00
    21. Paperclip1.00
  • 2:35 #6 done17 line(s)

    shot 6·sharpness 1881.7

    1. LIVENESS × ASSURANCE0.99
    2. Assurance without liveness becomes a review queue.0.99
    3. Liveness without assurance becomes fast wrongness.0.99
    4. A trustworthy agent system preserves both.0.99
    5. LIVENESS1.00
    6. ASSURANCE1.00
    7. Will the work continue to the next0.97
    8. Has enough been established to permit1.00
    9. legitimate action?1.00
    10. that action?0.99
    11. next action · blocker · wakeup · bounded loop0.95
    12. evidence · verifier · authority · risk ownership0.98
    13. DONE1.00
    14. The done claim is the controlled transition between the two.1.00
    15. 08 / 13 . Liveness × assurance0.95
    16. @dotta1.00
    17. e Paperclip0.94
  • 3:20 #7 skipped

    shot 7·duplicate of #6

  • 3:46 #8 done19 line(s)

    shot 8·sharpness 1189.5

    1. THE MECHANISMS1.00
    2. The control system.1.00
    3. Clear transitions1.00
    4. First-class blockers1.00
    5. Every heartbeat must leave a next action0.99
    6. Dependency is explicit.1.00
    7. Interactions & approvals0.99
    8. Reviewers & approvers1.00
    9. with an audit trail.1.00
    10. Enforced review by QA, CTO, etc..0.98
    11. Watchdogs · monitors · recovery0.98
    12. Evidence ≠ liveness1.00
    13. Ensure the /goal.1.00
    14. Comments, docs, work products are evidence.0.99
    15. Child issues + plan decomposition1.00
    16. Break work into smaller units.1.00
    17. 07 / 13 · Paperclip mechanisms0.98
    18. @dotta1.00
    19. e Paperclip0.96
  • 3:58 #9 done10 line(s)

    shot 9·sharpness 844.9

    1. THREE INVARIANTS1.00
    2. Productive work continues0.99
    3. 011.00
    4. Only real blockers stop work1.00
    5. 021.00
    6. Infinite loops are bounded1.00
    7. 031.00
    8. 09 / 13 . The three invariants0.99
    9. @dotta1.00
    10. e Paperclip0.96
  • 4:34 #10 skipped

    shot 10·duplicate of #8

  • 5:08 #11 skipped

    shot 11·duplicate of #8

  • 5:31 #12 done26 line(s)

    shot 12·sharpness 1718.3

    1. THE CODE MOMENT0.98
    2. {0.95
    3. "artifact": "talk-draft.md",0.99
    4. "scope": "5-15 minute AIE talk",1.00
    5. paperclip/issue0.99
    6. "standard": "acceptance criteria + source accuracy",0.99
    7. PAP-119411.00
    8. DONE0.99
    9. "evidence": ["source check", "timing", "QA read"],0.98
    10. Build branded slide deck0.99
    11. "verifier": "QA",1.00
    12. todo + / in_progress + in_review done0.93
    13. parentie0.90
    14. AlE talk - PAP-119260.88
    15. "authority": "CTO",0.97
    16. blockedy0.96
    17. — none open0.94
    18. "residualRisk": "speaker-specific polish",1.00
    19. nextAction0.90
    20. unblocks PAP-119360.98
    21. "allowedNextAction": "build slides"1.00
    22. }0.96
    23. Treat done as an object, not a boolean.0.99
    24. 10 / 13 · Done-claim JSON0.98
    25. @dotta1.00
    26. Paperclip0.99
  • 6:03 #13 done11 line(s)

    shot 13·sharpness 1605.2

    1. STEAL THIS CHECKLIST1.00
    2. Define done levels before agents start.1.00
    3. Require evidence, not just confidence.1.00
    4. Separate verifier from author.1.00
    5. Route by risk and reversibility.0.99
    6. Keep blockers explicit.0.99
    7. Bound recovery loops.1.00
    8. Manage allowed next actions.1.00
    9. 11 / 13 · Steal-this checklist0.98
    10. @dotta1.00
    11. e Paperclip0.95
  • 6:35 #14 done11 line(s)

    shot 14·sharpness 1601.1

    1. STEAL THIS CHECKLIST1.00
    2. Define done levels before agents start.1.00
    3. Require evidence, not just confidence.1.00
    4. Separate verifier from author.1.00
    5. Route by risk and reversibility.0.99
    6. Keep blockers explicit.0.99
    7. Bound recovery loops.1.00
    8. Manage allowed next actions.1.00
    9. 11 / 13 · Steal-this checklist0.98
    10. @dotta1.00
    11. e Paperclip0.95
  • 7:10 #15 done7 line(s)

    shot 15·sharpness 1085.8

    1. THE POINT1.00
    2. Make "done" specific enough that0.99
    3. someone else can safely build on it.0.98
    4. The output multiplier comes after the trust protocol.1.00
    5. 12 / 13 · The point0.94
    6. @dotta1.00
    7. e Paperclip0.92

Transcript

79 cues· 1,323 words· 7,310 chars

  1. 0:00 An agent opens a pull request.
  2. 0:01 It passes the tests.
  3. 0:03 It updates the documentation.
  4. 0:04 It closes the issue and comments, looks done to me.
  5. 0:07 But is it actually done?
  6. 0:09 Is it done enough to merge?
  7. 0:10 Is it done enough to deploy?
  8. 0:12 Is it done enough to announce to your customers?
  9. 0:15 These are fundamentally different operational claims, and most agent systems just flatten it to a single green check mark.
  10. 0:21 I'm Dota.
  11. 0:22 I'm the creator of Paperclip, and I'm going to give you some hard-earned lessons that we've learned in creating Paperclip's liveness model.
  12. 0:28 What does done even mean?
  13. 0:30 Here's the thing.
  14. 0:31 Programming is solved, and agents can now produce more code and documentation faster than any human can ever verify.
  15. 0:39 And this actually gives us a new failure mode, is that agents can actually create more work than humans have time to verify.
  16. 0:45 So we need a way to verify that our agents are done more than just letting them check a checkbox.
  17. 0:50 Done doesn't mean that an agent just changed the status of a task being done.
  18. 0:54 Saying that something is done is actually a bundle of claims.
  19. 0:58 You're saying that an artifact was produced, that you have evidence that the task is actually complete, and you have a rubric in which you can verify against.
  20. 1:07 You know exactly who the owner is for the next step, and you know exactly what the next step is.
  21. 1:14 There's different levels to how done something is.
  22. 1:16 The producer might claim something is complete, but you need to have a reviewer, another party that looks at it and finds no obvious issues.
  23. 1:23 You want to verify and make sure that the evidence actually meets a specified standard.
  24. 1:28 You want to make sure that a person who is authorized to approve it actually approves that the work is done.
  25. 1:34 And you want to make sure that there's somebody who actually stands behind the decision.
  26. 1:38 And ideally what you want is that the outcome has actually survived real world conditions.
  27. 1:45 Because exhaustive human verification fails at high volume.
  28. 1:49 You might be able to verify a few tasks per day, but essentially if you have humans verifying all the tasks and they have to sign off on it, eventually what you just get is a form of verification theater.
  29. 2:03 What you need is a protocol for defining how tasks actually progress through a system.
  30. 2:09 You want to make sure that tasks are always kept moving, but they don't get stuck into invalid states.
  31. 2:16 You need a control plane that actually has the execution of the tasks being tied to specific contracts and constraints about what the system will do and what agents it will hand off your next task to.
  32. 2:30 Because really what you're trying to play against is this idea around keeping work moving, but also having it verified.
  33. 2:37 When a task has been reviewed by a human, you get the assurance that it's correct.
  34. 2:41 But having a human verify it means that the task is dead in its tracks.
  35. 2:46 You also want to keep liveness.
  36. 2:48 Liveness means that the work is continuing with no blockers.
  37. 2:52 And you're always trying to keep these two things in balance.
  38. 2:54 If you have tasks that are completely alive with no approvals, then what you get is this classic AI slop because you're producing a lot of things with kind of no quality control and it's worse than creating nothing after a long period of time.
  39. 3:08 But if you have pure review, then you have this enormous review queue where humans can't actually review it by hand anyway.
  40. 3:17 These agents will be creating far more than you can ever actually review.
  41. 3:22 And so we have to find a way to tease apart the bundle of claims that are involved in saying a task is done.
  42. 3:27 With Paperclip, we have a number of mechanisms to keep this going.
  43. 3:31 You might think that you can easily just write a for loop over your task manager and have your agents work, but quickly you'll find that falls apart.
  44. 3:39 As soon as you start integrating task dependency trees, blockers, multiple agents, idipotent checkouts, like locks on checkouts, you find that this tension between liveness and verification actually gets quite complicated.
  45. 3:55 There's really three invariants that are extremely important when you're thinking about what you want out of a control plane for your agentic work.
  46. 4:03 You want to ensure that productive work continues.
  47. 4:06 You want to make sure that only real blockers stop work.
  48. 4:09 And you want to make sure that infinite loops are bounded.
  49. 4:12 In Paperclip, we have built a number of mechanisms to deal with this problem.
  50. 4:17 So for example, every time you have a task, there's clear transitions to what the next state could be.

Open at this second