read-only demo

Videos ow1we5PzK-o

The Multi-Agent Architecture That Actually Ships — Luke Alvoeiro, Factory

index_state ready data_status ok

AI Engineer· published 2026-05-06· 0:18:30· en-US· indexed 2026-08-10 19:44

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:16, 1 of 1 keyframes kept
  4. Shot 3, 0:16 to 1:03, 1 of 1 keyframes kept
  5. Shot 4, 1:03 to 1:49, 1 of 1 keyframes kept
  6. Shot 5, 1:49 to 2:06, 1 of 1 keyframes kept
  7. Shot 6, 2:06 to 2:43, 1 of 1 keyframes kept
  8. Shot 7, 2:43 to 3:19, 0 of 1 keyframes kept
  9. Shot 8, 3:19 to 4:03, 1 of 1 keyframes kept
  10. Shot 9, 4:03 to 4:33, 1 of 1 keyframes kept
  11. Shot 10, 4:33 to 5:04, 1 of 1 keyframes kept
  12. Shot 11, 5:04 to 5:34, 0 of 1 keyframes kept
  13. Shot 12, 5:34 to 6:05, 0 of 1 keyframes kept
  14. Shot 13, 6:05 to 6:30, 1 of 1 keyframes kept
  15. Shot 14, 6:30 to 6:56, 1 of 1 keyframes kept
  16. Shot 15, 6:56 to 7:32, 1 of 1 keyframes kept
  17. Shot 16, 7:32 to 8:07, 0 of 1 keyframes kept
  18. Shot 17, 8:07 to 8:48, 1 of 1 keyframes kept
  19. Shot 18, 8:48 to 9:15, 1 of 1 keyframes kept
  20. Shot 19, 9:15 to 10:05, 1 of 1 keyframes kept
  21. Shot 20, 10:05 to 10:26, 1 of 1 keyframes kept
  22. Shot 21, 10:26 to 10:31, 1 of 1 keyframes kept
  23. Shot 22, 10:31 to 10:35, 1 of 1 keyframes kept
  24. Shot 23, 10:35 to 10:40, 1 of 1 keyframes kept
  25. Shot 24, 10:40 to 10:50, 1 of 1 keyframes kept
  26. Shot 25, 10:50 to 10:55, 0 of 1 keyframes kept
  27. Shot 26, 10:55 to 10:59, 0 of 1 keyframes kept
  28. Shot 27, 10:59 to 11:04, 0 of 1 keyframes kept
  29. Shot 28, 11:04 to 11:14, 0 of 1 keyframes kept
  30. Shot 29, 11:14 to 11:19, 0 of 1 keyframes kept
  31. Shot 30, 11:19 to 11:21, 0 of 1 keyframes kept
  32. Shot 31, 11:21 to 11:52, 1 of 1 keyframes kept
  33. Shot 32, 11:52 to 12:24, 0 of 1 keyframes kept
  34. Shot 33, 12:24 to 13:02, 1 of 1 keyframes kept
  35. Shot 34, 13:02 to 13:30, 1 of 1 keyframes kept
  36. Shot 35, 13:30 to 13:58, 0 of 1 keyframes kept
  37. Shot 36, 13:58 to 14:21, 1 of 1 keyframes kept
  38. Shot 37, 14:21 to 14:32, 1 of 1 keyframes kept
  39. Shot 38, 14:32 to 15:21, 1 of 1 keyframes kept
  40. Shot 39, 15:21 to 15:49, 1 of 1 keyframes kept
  41. Shot 40, 15:49 to 16:32, 1 of 1 keyframes kept
  42. Shot 41, 16:32 to 16:35, 1 of 1 keyframes kept
  43. Shot 42, 16:35 to 17:02, 1 of 1 keyframes kept
  44. Shot 43, 17:02 to 17:29, 0 of 1 keyframes kept
  45. Shot 44, 17:29 to 17:56, 1 of 1 keyframes kept
  46. Shot 45, 17:56 to 18:06, 1 of 1 keyframes kept
  47. Shot 46, 18:06 to 18:16, 1 of 1 keyframes kept
  48. Shot 47, 18:16 to 18:29, 1 of 1 keyframes kept
  49. Shot 48, 18:29 to 18:30, 0 of 1 keyframes kept

49 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
197
whisperx 197
chunks
31
from 197 cues
keyframes
35
kept of 49 captured
frames with text
35
963 lines read
chapters
11
from the source metadata
keyframe bytes
3.9 MB
word timings on 197 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 00:52 1m 53s
stt done 2026-08-10 00:54 19s
chunk done 2026-08-10 00:55 0s
text_embed done 2026-08-10 19:44 1s
keyframe done 2026-08-10 00:55 1m 31s
ocr done 2026-08-10 00:56 20s
frame_embed done 2026-08-10 19:44 6s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:53 #3 done17 line(s)

    shot 3·sharpness 1765.2

    1. S0.94
    2. gineer0.97
    3. UROPE0.92
    4. Assembling agent teams that solve1.00
    5. problems 15x harder than single0.97
    6. rosoft1.00
    7. agents can1.00
    8. AlEngineer1.00
    9. Luke Alvoeiro - Factory0.97
    10. EUROPE1.00
    11. PRESENTED BY1.00
    12. Google DeepMind1.00
    13. AlEngineer0.99
    14. # Braintrust0.95
    15. WorkOS1.00
    16. OpenAI0.93
    17. EUROPE1.00
  • 1:43 #4 done30 line(s)

    shot 4·sharpness 2572.0

    1. EUROPE1.00
    2. AGENT TEAMS1.00
    3. neer1.00
    4. together.0.99
    5. The bottleneck is no longer intelligence0.99
    6. er.de1.00
    7. jineer0.98
    8. The best engineers can only0.99
    9. IPE0.95
    10. focus on a couple things at a0.99
    11. time. They have a backlog of 500.99
    12. features but can only drive a1.00
    13. rize1.00
    14. neer1.00
    15. few forward per day. Today's0.99
    16. models are smart enough to0.98
    17. build all 50.0.97
    18. ngineer1.00
    19. What if the human decides1.00
    20. UDFLARE1.00
    21. EUROPE1.00
    22. what to build, and the0.96
    23. system figures outhow?0.99
    24. AI Engineer - Factory1.00
    25. 20261.00
    26. AlEngineer0.97
    27. # Braintrust0.94
    28. WorkOS1.00
    29. OpenAl0.92
    30. EUROPE1.00
  • 1:51 #5 done47 line(s)

    shot 5·sharpness 1893.3

    1. EUROPE1.00
    2. ineer1.00
    3. togethe1.00
    4. AGENT TEAMS0.99
    5. OPE1.00
    6. Five multi-agent strategies1.00
    7. ger.c0.88
    8. AlEngineer0.99
    9. Q.0.71
    10. EUROPE1.00
    11. Delegation1.00
    12. Creator-Verifier1.00
    13. Direct Communication0.99
    14. One agent spawns another for a1.00
    15. One agent builds, a different1.00
    16. Agents talking peer-to-peer1.00
    17. subtask. The simplest1.00
    18. agent checks. Separation of0.99
    19. without a coordinator. Hard1.00
    20. to1.00
    21. arize0.99
    22. multi-agent pattern.1.00
    23. concerns removes sunk-cost bias.1.00
    24. get right: state fragments.1.00
    25. gin0.99
    26. OPE1.00
    27. Q20.58
    28. (o))0.85
    29. Negotiation1.00
    30. Broadcast1.00
    31. Two agents coordinate over1.00
    32. One agent sends status updates1.00
    33. shared resources. Best when1.00
    34. and shared context to many.1.00
    35. AlEngineer0.98
    36. there's a possible win-win.0.98
    37. Critical for coherence.0.99
    38. OUDFLA0.97
    39. EUROPE1.00
    40. AI Engineer - Factory1.00
    41. 20261.00
    42. 21.00
    43. AlEngineer0.99
    44. Braintrust1.00
    45. WorkOS1.00
    46. OpenAl0.94
    47. EUROPE1.00
  • 2:18 #6 done28 line(s)

    shot 6·sharpness 1853.7

    1. AGENT TEAMS0.99
    2. Five multi-agent strategies1.00
    3. Q.0.72
    4. Delegation1.00
    5. Creator-Verifier1.00
    6. Direct Communication1.00
    7. One agent spawns another for a1.00
    8. One agent builds, a different1.00
    9. Agents talking peer-to-peer0.99
    10. subtask. The simplest1.00
    11. agent checks. Separation of0.99
    12. without a coordinator. Hard to1.00
    13. multi-agent pattern.0.98
    14. concerns removes sunk-cost bias.0.98
    15. get right: state fragments.1.00
    16. Q20.71
    17. ((o))0.91
    18. Negotiation1.00
    19. Broadcast1.00
    20. Two agents coordinate over0.97
    21. One agent sends status updates1.00
    22. shared resources. Best when1.00
    23. and shared context to many.0.99
    24. there's a possible win-win.0.98
    25. Critical for coherence.1.00
    26. AI Engineer - Factory1.00
    27. 20261.00
    28. 21.00
  • 2:58 #7 skipped

    shot 7·duplicate of #6

  • 3:29 #8 done37 line(s)

    shot 8·sharpness 1953.9

    1. AlEn0.96
    2. AGENT TEAMS1.00
    3. Five multi-agent strategies1.00
    4. Q0.95
    5. Delegation1.00
    6. Creator-Verifier1.00
    7. Direct Communication0.99
    8. One agent spawns another for a1.00
    9. One agent builds, a different1.00
    10. Agents talking peer-to-peer1.00
    11. subtask. The simplest0.98
    12. agent checks. Separation of1.00
    13. without a coordinator. Hard1.00
    14. to1.00
    15. multi-agent pattern.1.00
    16. concerns removes sunk-cost bias.1.00
    17. get right: state fragments.1.00
    18. AlEngineer0.98
    19. 820.65
    20. (o))0.85
    21. EUROPE1.00
    22. Negotiation1.00
    23. Broadcast1.00
    24. Two agents coordinate over1.00
    25. One agent sends status updates1.00
    26. PRESENTED BY0.99
    27. shared resources. Best when1.00
    28. and shared context to many.1.00
    29. there's a possible win-win.0.98
    30. Critical for coherence.0.99
    31. Google DeepMind0.99
    32. AI Engineer - Factory1.00
    33. 20261.00
    34. 21.00
    35. Google DeepMind1.00
    36. AlEngineer0.97
    37. EUROPE1.00
  • 4:30 #9 done32 line(s)

    shot 9·sharpness 3560.6

    1. fust0.89
    2. AlEngineer0.96
    3. EUROPE1.00
    4. INTRODUCING MISSIONS1.00
    5. e04j0.88
    6. Missions -a system that combines delegation,1.00
    7. er1.00
    8. creator-verifier, broadcast, and negotiation into a single1.00
    9. workflow1.00
    10. kel0.94
    11. HOW TO USE THEM0.99
    12. er1.00
    13. NTRY1.00
    14. Describe a software1.00
    15. Scope it through1.00
    16. Approve the plan0.98
    17. Missions handles1.00
    18. goal.1.00
    19. conversation.1.00
    20. execution1.00
    21. ar1.00
    22. gineer1.00
    23. OPE1.00
    24. A mission is not a single agent session that runs for a long time.0.99
    25. It's an ecosystem of agents coordinating through structured handoffs and0.99
    26. shared state.1.00
    27. AI Engineer - Factory0.99
    28. 20261.00
    29. 31.00
    30. Google DeepMind0.99
    31. AlEngineer0.99
    32. EUROPE1.00
  • 4:43 #10 done19 line(s)

    shot 10·sharpness 1214.7

    1. INTRODUCING MISSIONS1.00
    2. Three-role architecture1.00
    3. THE ORCHESTRATOR1.00
    4. Plans features, milestones, and the1.00
    5. validation contract1.00
    6. CHILD1.00
    7. CHILD1.00
    8. Workers1.00
    9. Validators1.00
    10. Fresh context per feature.1.00
    11. Adversarial verification. Have1.00
    12. Implement, commit via git, hand off.1.00
    13. never seen the code before.1.00
    14. A key takeaway1.00
    15. The validation contract defines what "done" means before any1.00
    16. code is written0.99
    17. AI Engineer - Factory1.00
    18. 20261.00
    19. 41.00
  • 5:13 #11 skipped

    shot 11·duplicate of #10

  • 5:44 #12 skipped

    shot 12·duplicate of #10

  • 6:20 #13 done6 line(s)

    shot 13·sharpness 453.9

    1. HOW THEY WORK0.99
    2. The validation loop0.98
    3. Tests written after implementation don't catch bugs. They confirm decisions.0.99
    4. AI Engineer - Factory0.99
    5. 20261.00
    6. 51.00
  • 6:50 #14 done10 line(s)

    shot 14·sharpness 842.3

    1. HOW THEY WORK0.98
    2. The validation loop0.98
    3. Tests written after implementation don't catch bugs. They confirm decisions.1.00
    4. PLANNING PHASE1.00
    5. Validation Contract0.99
    6. Written by orchestrator during planning, before any code. Hundreds of assertions define correctness1.00
    7. independently of implementation.1.00
    8. AI Engineer - Factory0.98
    9. 20261.00
    10. 51.00
  • 7:00 #15 done25 line(s)

    shot 15·sharpness 2620.2

    1. HOW THEY WORK1.00
    2. The validation loop1.00
    3. Tests written after implementation don't catch bugs. They confirm decisions.1.00
    4. PLANNING PHASE1.00
    5. Validation Contract0.99
    6. Written by orchestrator during planning, before any code. Hundreds of assertions define correctness0.99
    7. independently of implementation.0.99
    8. After each milestone...0.99
    9. </>0.90
    10. </>0.88
    11. Scrutiny Validator1.00
    12. User-Testing Validator0.99
    13. Runs tests, typechecking,0.98
    14. Acts like a QA engineer.1.00
    15. linting. Spawns code1.00
    16. Launches the app,1.00
    17. review agents for each1.00
    18. navigates via computer-use1.00
    19. completed feature.1.00
    20. & verifies flows1.00
    21. end-to-end.1.00
    22. Neither validator has ever seen the code. Validation is adversarial by design.1.00
    23. AI Engineer - Factory0.98
    24. 20261.00
    25. 51.00
  • 8:00 #16 skipped

    shot 16·duplicate of #15

  • 8:32 #17 done19 line(s)

    shot 17·sharpness 1134.3

    1. HOW THEY WORK0.99
    2. Structured handoffs1.00
    3. How agents stay coherent over days, not just minutes.0.99
    4. New Task1.00
    5. Every worker reports1.00
    6. What was implemented1.00
    7. Execute1.00
    8. What was left undone1.00
    9. Encode Skill0.96
    10. Commands run + exit codes1.00
    11. CONTINUOUS1.00
    12. Issues discovered1.00
    13. LEARNING1.00
    14. Whether procedures were followed0.98
    15. Observe1.00
    16. Learn1.00
    17. AI Engineer - Factory0.98
    18. 20261.00
    19. 61.00
  • 8:51 #18 done26 line(s)

    shot 18·sharpness 1378.8

    1. lind0.77
    2. AlEngineer0.96
    3. EUROPE1.00
    4. HOW THEY WORK1.00
    5. toner.ai0.93
    6. Structured handoffs1.00
    7. How agents stay coherent over days, not just minutes.1.00
    8. ev1.00
    9. Every worker reports0.99
    10. What was implemented0.98
    11. What was left undone1.00
    12. Commands run + exit codes1.00
    13. CONTINUOUS1.00
    14. Issues discovered1.00
    15. LEARNING1.00
    16. Whether procedures were followed0.99
    17. ARE1.00
    18. EURC0.97
    19. AI Engineer - Factory0.99
    20. 20261.00
    21. 61.00
    22. AlEngineer0.99
    23. Braintrust1.00
    24. WorkOS1.00
    25. OpenAl0.93
    26. EUROPE1.00
  • 9:59 #19 done40 line(s)

    shot 19·sharpness 1668.6

    1. AlEngineer0.89
    2. EUROPE1.00
    3. HOW THEY WORK1.00
    4. Why serial beats parallel (mostly)0.99
    5. OpenAl0.98
    6. Engineer1.00
    7. EUROPE1.00
    8. The next question, how should this execute...1.00
    9. Aici0.72
    10. AlEngineer0.95
    11. WHAT PEOPLE EXPECT1.00
    12. WHAT ACTUALLY WORKS0.96
    13. EUROPE1.00
    14. Parallelism1.00
    15. Serial Execution0.99
    16. Agents conflict and step on each other1.00
    17. Features execute one at a time. Each worker0.98
    18. ngineer1.00
    19. EUROPE0.99
    20. CNRY0.81
    21. Duplicate work and inconsistent architecture1.00
    22. Coordination overhead eats speed gains1.00
    23. Every conflict burns tokens0.99
    24. inherits the full codebase from the last0.99
    25. Parallelism is reserved for work that can't1.00
    26. through git1.00
    27. ENTED BY1.00
    28. conflict: codebase exploration, API research,1.00
    29. documentation reads, and validation reviews.0.98
    30. DeepMind1.00
    31. Slower on paper. But for multi-day runs,0.99
    32. IEngineer0.97
    33. correctness compounds.0.98
    34. EUROPE0.99
    35. AI Engineer - Factory0.99
    36. 20261.00
    37. 71.00
    38. Engineering the future of Al1.00
    39. AlEngineer0.98
    40. EUROPE1.00
  • 10:12 #20 done22 line(s)

    shot 20·sharpness 1617.3

    1. HOW THEY WORK0.97
    2. Why serial beats parallel (mostly)1.00
    3. The next question, how should this execute...0.99
    4. WHAT PEOPLE EXPECT1.00
    5. WHAT ACTUALLY WORKS1.00
    6. Parallelism1.00
    7. Serial Execution1.00
    8. Agents conflict and step on each other1.00
    9. Features execute one at a time. Each worker1.00
    10. Duplicate work and inconsistent architecture0.99
    11. inherits the full codebase from the last1.00
    12. through git0.99
    13. Coordination overhead eats speed gains0.99
    14. Every conflict burns tokens1.00
    15. Parallelism is reserved for work that can't1.00
    16. conflict: codebase exploration, API research,0.99
    17. documentation reads, and validation reviews.1.00
    18. Slower on paper. But for multi-day runs,0.98
    19. correctness compounds.1.00
    20. AI Engineer - Factory0.99
    21. 20261.00
    22. 71.00
  • 10:31 #21 done72 line(s)

    shot 21·sharpness 994.3

    1. USING MISSIONS1.00
    2. ∴Droid0.98
    3. Mission Control0.99
    4. Mission Control1.00
    5. -/Development/note-tracker1.00
    6. TIME 56m 54s0.99
    7. Input 324.0K0.97
    8. Cached 16.8M0.99
    9. Output 111.0K0.96
    10. RUNNING1.00
    11. 3/17 [+6]0.91
    12. A dedicated view for multi-day autonomous1.00
    13. Active Feature scrutiny-validator-app-shell0.99
    14. Features1.00
    15. 3/111.00
    16. work. Monitor, redirect, or close your0.99
    17. skill scrutiny-validator0.98
    18. bootstrap-tauri-vorkspace0.98
    19. laptop and come back tomorrow.1.00
    20. ailestone app-shell0.99
    21. ✓ menu-bar-shell-and-lifecycle0.95
    22. Preconditions1.00
    23. scrutiny-validator-app-shell0.99
    24. All implementation features for milestone "app-shell" are complete0.99
    25. categories-and-tagging-experience0.99
    26. Expected Behavior1.00
    27. due-date-parsing-and-note-metadata0.98
    28. Validators pass (test, typecheck, lint)0.99
    29. I0.92
    30. google-calendar-auth-and-secure-storage1.00
    31. Reviev subagents spawned for each feature0.98
    32. -3 more0.92
    33. Findings synthesized into scrutiny report1.00
    34. Progress Log1.00
    35. 1-9 of 180.98
    36. Description1.00
    37. Scrutiny validation for milestone "app-shell". Runs test suite, typecheck,0.99
    38. <1m ago0.89
    39. #95fb1b2d started [scrutiny-validator-app-shell]0.99
    40. and lint. Spawns review subagents for each completed feature. Synthesizes0.99
    41. <1n ago0.85
    42. Milestone validation: app-shell1.00
    43. findings. Always returns to orchestrator.0.98
    44. <1m ago0.95
    45. #d1b6b4a7 completed [quick-note-save-and-main-list]0.99
    46. W0.96
    47. 5m ago0.92
    48. 5m ago0.90
    49. #afc827a8 conpleted [menu-bar-shell-and-lifecycle] √0.97
    50. #d1b6b4a7 started [quick-note-save-and-main-list]0.99
    51. 13m ago0.98
    52. #afc827a8 started [menu-bar-shell-and-lifecycle]1.00
    53. 21= ago0.87
    54. 13m ago0.96
    55. #1963e245 conpleted [bootstrap-tauri-workspace]0.99
    56. #1963e245 started[bootstrap-tauri-workspace]0.99
    57. 21m ago1.00
    58. Run started: Mission artifacts authored, repo scaffold..0.99
    59. Active Worker #1 scrutiny-validator-app-shell1.00
    60. Duration 7s1.00
    61. ## Your Assigned Featurejson { "id": "scrutiny-validator-app-shell", "description": "Scrutiny validation for milestone \"app-shell\". Ru0.98
    62. ns test suite, typecheck, and lint. Spawns reviev subagents for each completed feature. Synthesizes findings. Always returns to orchestrato...0.98
    63. F Features1.00
    64. Workers0.98
    65. M Models0.98
    66. P Pause0.99
    67. D Mission Dir0.95
    68. Ctr1+T Back To Orchestrator0.99
    69. Monitor from a bird's eye view1.00
    70. AI Engineer - Factory0.99
    71. 20261.00
    72. 81.00
  • 10:35 #22 done61 line(s)

    shot 22·sharpness 811.0

    1. USING MISSIONS1.00
    2. ∴Droid0.97
    3. Mission Control1.00
    4. "Mission Control0.97
    5. -/Development/note-tracker0.99
    6. TIME 56m 54s0.98
    7. Input: 324.0K0.94
    8. Cached 16.8M0.98
    9. Output 111.0K0.97
    10. Workers(4)0.99
    11. A dedicated view for multi-day autonomous1.00
    12. Al1 (4) | Active (1) | Completed [3) | Failad (8)0.91
    13. work. Monitor, redirect, or close your0.98
    14. laptop and come back tomorrow.0.98
    15. Session1.00
    16. Start0.99
    17. Duration1.00
    18. Status1.00
    19. Input1.00
    20. Cached1.00
    21. Output0.90
    22. Feature1.00
    23. 95fb1b2d1.00
    24. Running0.97
    25. scrutiny-validator-app-shell1.00
    26. √20.92
    27. afc827a81.00
    28. 1963e2451.00
    29. d1b6b4a71.00
    30. 15:020.96
    31. 14:461.00
    32. 14:541.00
    33. 5m 47s0.93
    34. 8m 1s0.97
    35. Success1.00
    36. Success1.00
    37. Success1.00
    38. 541.00
    39. 480.98
    40. 680.97
    41. 5.0M0.94
    42. 2.0M0.97
    43. 7.3M0.97
    44. 21.BK0.99
    45. 15.1K0.99
    46. 18.BK0.99
    47. quick-note-save-and-main-list1.00
    48. bootstrap-tauri-workspace1.00
    49. menu-bar-shell-and-lifecycle1.00
    50. I0.80
    51. Enter1.00
    52. ↑↓ Select0.86
    53. Enter Viev0.98
    54. T Filter0.98
    55. Ese Back0.96
    56. F Features0.98
    57. M Models0.98
    58. Inspect all the workers0.98
    59. AI Engineer - Factory0.98
    60. 20261.00
    61. 81.00
  • 10:40 #23 done50 line(s)

    shot 23·sharpness 889.4

    1. USING MISSIONS1.00
    2. 0.51
    3. ∴Droid0.98
    4. Mission Control1.00
    5. ∴Mission Control0.92
    6. -/Development/note-tracker1.00
    7. TIME 56m 54s0.98
    8. Input 324.0K0.94
    9. Cached 16.8M1.00
    10. Output 111.0K0.97
    11. Worker Session0.99
    12. A dedicated view for multi-day autonomous1.00
    13. Feature: quick-note-save-and-main-list1.00
    14. Session:d1b6b4a7-30b2-488f-a1ca-9dea374bb8b80.99
    15. work. Monitor, redirect, or close your0.97
    16. laptop and come back tomorrow.1.00
    17. Model: Opus 4.6 [High)0.95
    18. Status: Success Duration: 5m 47s0.98
    19. Tokens: Input 40 Cached 2.0M Output 18.0K0.98
    20. ↑ 93 more above0.93
    21. Plan 1. [completed] Run baseline tests (pnpm test + cargo test)2. [completed] Add SQLite dependency and create note persistence layer in Rust...1.00
    22. →T0ODO List Updated0.92
    23. I0.91
    24. Execute cd /Users/luke/Development/note-tracker && git status0.99
    25. →On branch main Changes not staged for commit: (use "git add <file>... to update what will be committed) (use "git restore <file>...* to disc0.97
    26. ard changes in working directory) modified: src-tauri/Cargo.lock modified: src-tauri/Cargo.toml modified: src-tauri/src/lib.rs modified: src/__0.99
    27. tests../App.test.tsx modified: src/_.tests./QuickNote.test.tsx modified: src/components/MainApp.tsx modified: src/components/QuickNote.tsx mo..0.97
    28. Execute cd /Users/luke/Development/note-tracker && git add -A && git commit -m "Add note persistence vith SQLite and quick-note save path - Ad..0.99
    29. 644 src-tauri/src/db.rs create mode 188644 src/lib/tauri.ts create0.98
    30. -[main bc131a1] Add note persistence with SQLite and quick-note sav0.98
    31. M1.00
    32. 8644 src/lib/types.ts [Process exited with code 8]0.98
    33. 11 files changed, 619 insertions(+), 27 deletions(-) create mode 1800.98
    34. Execute cd /Users/luke/Development/note-tracker && git status && git rev-parse HEAD0.99
    35. →On branch main nothing to comit, working tree clean bc131a15bdedf4c13d7a85d55744c231fcddfb4d [Process exited with code 8]0.99
    36. ∴Clean working tree. Now let me call EndFeatureRun:0.99
    37. End Feature Run success1.00
    38. →{"recorded":true,"nextAction":"continue","message":"Another worker will continue with next feature, IMPORTANT: Your session is nou complete.0.99
    39. Do not make any further tool calls or continue working. End your turn immediately."}0.99
    40. ↑↓ Scroll0.86
    41. [] Density0.95
    42. g Top0.94
    43. G Bottom0.89
    44. s Interrupt/Chat0.99
    45. h Handoff0.95
    46. Ese Back0.94
    47. Read the handoffs1.00
    48. AI Engineer - Factory0.99
    49. 20261.00
    50. 81.00

Transcript

197 cues· 2,923 words· 17,019 chars

  1. 0:15 Hi everyone, my name is Luke, and my goal is that 20 minutes from now, you'll be able to assemble agent teams that can complete tasks orders of magnitude harder than what you can complete with a single agent today.
  2. 0:27 A little bit about me.
  3. 0:30 I come from a background in dev tools.
  4. 0:32 About two and a half years ago, I started a project at Block, which is where I was working at the time, and that project evolved into Goose.
  5. 0:40 Goose is now one of the leading coding agents that is open source, and it recently was donated to the Agentic AI Foundation.
  6. 0:50 So it's been really cool to see.
  7. 0:53 Nowadays I work at Factory, where I lead our core agent harness.
  8. 0:57 And Factory's mission is to bring autonomy to the entire software development life cycle.
  9. 1:04 So I want to start off with a claim.
  10. 1:06 The bottleneck in software engineering nowadays is not intelligence.
  11. 1:10 It's now limited by human attention.
  12. 1:13 Even the best engineers can only complete a couple of tasks at a time.
  13. 1:17 They may have a backlog of 50 features, but they can only drive a few forward per day because every task requires their attention, every commit needs their review.
  14. 1:26 Today's models are smart enough to figure out all 50 of these tasks, but there's not enough bandwidth to supervise their implementation.
  15. 1:36 So we kept asking ourselves, what if a human decides what to build and then a system figures out how to do so?
  16. 1:43 An agent could just work for hours, for days, and you come back to finish work.
  17. 1:47 So that's what I'm here to talk about.
  18. 1:50 When you start researching multi-agent frameworks and systems, you quickly realize that the field's a bit of a mess.
  19. 1:56 Everyone has their own framework, their own terminology, their own opinions of what works and doesn't work.
  20. 2:02 And so I want to propose a simple taxonomy.
  21. 2:05 There's five frontier multi-agent frameworks.
  22. 2:07 One is delegation.
  23. 2:09 This is where one agent spawns another agent, and the parent agent may say, go figure out the database schema, and then gets a response back.
  24. 2:17 This is the simplest form of multi-agent communication as what most people implement first.
  25. 2:23 You have, you know, sub-agents and coding tools are the most common example.
  26. 2:28 The other one is creator verifier, right, where one agent builds something and then you have another agent that checks that work.
  27. 2:35 And the key here is a separation of concerns.
  28. 2:38 The agent that implemented the code has sunk cost bias.
  29. 2:43 It wants that code to work.
  30. 2:45 A fresh agent with fresh context is way more likely to find issues, and this is why we do code review as humans as well.
  31. 2:52 Another one is direct communication.
  32. 2:54 This is when agents communicate without a central coordinator.
  33. 2:57 It's kind of like DMing each other.
  34. 3:00 It's hard to get right, though, because state fragments across conversations without that coordinator, and there's no single source of truth.
  35. 3:09 The next one is negotiation.
  36. 3:11 Negotiation is when agents communicate, but over a shared resource.
  37. 3:16 So that might be they want to use the same API.
  38. 3:19 They want to modify the same portion of the code base.
  39. 3:23 But negotiation doesn't need to be adversarial.
  40. 3:25 In fact, the best use case is when there's net positive sum trading.
  41. 3:30 And that's when agents have a potential win-win situation while interacting.
  42. 3:36 And then the last one is broadcast.
  43. 3:38 And that is when one agent sends information to many.
  44. 3:41 Think of it like status updates, new context that applies to everyone, new shared constraints.
  45. 3:48 It's a bit less flashy than the other ones, but it's critical for maintaining coherence over long-running tasks.
  46. 3:56 And so when you have all of these different building blocks, how do you assemble that into a system that can run for many days?
  47. 4:03 So missions is our answer.
  48. 4:05 It's a system that combines four of those, delegation, creator-verifier, broadcast and negotiation,
  49. 4:13 into a single workflow.
  50. 4:14 You describe a goal.

Chapters

  1. 0:00 Introduction to multi-agent systems and the bottleneck of human attention
  2. 1:50 Taxonomy of five frontier multi-agent frameworks
  3. 4:04 Introducing 'Missions': The three-role architecture (Orchestrator, Workers, Validators)
  4. 6:34 The importance of validation contracts for consistent quality
  5. 8:09 Maintaining long-term context through structured handoffs
  6. 9:17 The case for serial execution over parallel execution
  7. 10:30 Mission control: Monitoring agent progress
  8. 11:22 Strategic model selection per role ('Droid whispering')
  9. 13:06 Production data analysis: Building a Slack clone
  10. 14:34 Designing systems that improve with each model generation
  11. 15:51 Conclusion: The shifting economics of software engineering

Open at this second