Videos ow1we5PzK-o
The Multi-Agent Architecture That Actually Ships — Luke Alvoeiro, Factory
Scene timeline
49 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 197
- whisperx 197
- chunks
- 31
- from 197 cues
- keyframes
- 35
- kept of 49 captured
- frames with text
- 35
- 963 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 3.9 MB
- word timings on 197 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 00:52 | 1m 53s |
stt |
done | — | 2026-08-10 00:54 | 19s |
chunk |
done | — | 2026-08-10 00:55 | 0s |
text_embed |
done | — | 2026-08-10 19:44 | 1s |
keyframe |
done | — | 2026-08-10 00:55 | 1m 31s |
ocr |
done | — | 2026-08-10 00:56 | 20s |
frame_embed |
done | — | 2026-08-10 19:44 | 6s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- S0.94
- gineer0.97
- UROPE0.92
- Assembling agent teams that solve1.00
- problems 15x harder than single0.97
- rosoft1.00
- agents can1.00
- AlEngineer1.00
- Luke Alvoeiro - Factory0.97
- EUROPE1.00
- PRESENTED BY1.00
- Google DeepMind1.00
- AlEngineer0.99
- # Braintrust0.95
- WorkOS1.00
- OpenAI0.93
- EUROPE1.00
-
- EUROPE1.00
- AGENT TEAMS1.00
- neer1.00
- together.0.99
- The bottleneck is no longer intelligence0.99
- er.de1.00
- jineer0.98
- The best engineers can only0.99
- IPE0.95
- focus on a couple things at a0.99
- time. They have a backlog of 500.99
- features but can only drive a1.00
- rize1.00
- neer1.00
- few forward per day. Today's0.99
- models are smart enough to0.98
- build all 50.0.97
- ngineer1.00
- What if the human decides1.00
- UDFLARE1.00
- EUROPE1.00
- what to build, and the0.96
- system figures outhow?0.99
- AI Engineer - Factory1.00
- 20261.00
- AlEngineer0.97
- # Braintrust0.94
- WorkOS1.00
- OpenAl0.92
- EUROPE1.00
-
- EUROPE1.00
- ineer1.00
- togethe1.00
- AGENT TEAMS0.99
- OPE1.00
- Five multi-agent strategies1.00
- ger.c0.88
- AlEngineer0.99
- Q.0.71
- EUROPE1.00
- Delegation1.00
- Creator-Verifier1.00
- Direct Communication0.99
- One agent spawns another for a1.00
- One agent builds, a different1.00
- Agents talking peer-to-peer1.00
- subtask. The simplest1.00
- agent checks. Separation of0.99
- without a coordinator. Hard1.00
- to1.00
- arize0.99
- multi-agent pattern.1.00
- concerns removes sunk-cost bias.1.00
- get right: state fragments.1.00
- gin0.99
- OPE1.00
- Q20.58
- (o))0.85
- Negotiation1.00
- Broadcast1.00
- Two agents coordinate over1.00
- One agent sends status updates1.00
- shared resources. Best when1.00
- and shared context to many.1.00
- AlEngineer0.98
- there's a possible win-win.0.98
- Critical for coherence.0.99
- OUDFLA0.97
- EUROPE1.00
- AI Engineer - Factory1.00
- 20261.00
- 21.00
- AlEngineer0.99
- Braintrust1.00
- WorkOS1.00
- OpenAl0.94
- EUROPE1.00
-
- AGENT TEAMS0.99
- Five multi-agent strategies1.00
- Q.0.72
- Delegation1.00
- Creator-Verifier1.00
- Direct Communication1.00
- One agent spawns another for a1.00
- One agent builds, a different1.00
- Agents talking peer-to-peer0.99
- subtask. The simplest1.00
- agent checks. Separation of0.99
- without a coordinator. Hard to1.00
- multi-agent pattern.0.98
- concerns removes sunk-cost bias.0.98
- get right: state fragments.1.00
- Q20.71
- ((o))0.91
- Negotiation1.00
- Broadcast1.00
- Two agents coordinate over0.97
- One agent sends status updates1.00
- shared resources. Best when1.00
- and shared context to many.0.99
- there's a possible win-win.0.98
- Critical for coherence.1.00
- AI Engineer - Factory1.00
- 20261.00
- 21.00
-
- AlEn0.96
- AGENT TEAMS1.00
- Five multi-agent strategies1.00
- Q0.95
- Delegation1.00
- Creator-Verifier1.00
- Direct Communication0.99
- One agent spawns another for a1.00
- One agent builds, a different1.00
- Agents talking peer-to-peer1.00
- subtask. The simplest0.98
- agent checks. Separation of1.00
- without a coordinator. Hard1.00
- to1.00
- multi-agent pattern.1.00
- concerns removes sunk-cost bias.1.00
- get right: state fragments.1.00
- AlEngineer0.98
- 820.65
- (o))0.85
- EUROPE1.00
- Negotiation1.00
- Broadcast1.00
- Two agents coordinate over1.00
- One agent sends status updates1.00
- PRESENTED BY0.99
- shared resources. Best when1.00
- and shared context to many.1.00
- there's a possible win-win.0.98
- Critical for coherence.0.99
- Google DeepMind0.99
- AI Engineer - Factory1.00
- 20261.00
- 21.00
- Google DeepMind1.00
- AlEngineer0.97
- EUROPE1.00
-
- fust0.89
- AlEngineer0.96
- EUROPE1.00
- INTRODUCING MISSIONS1.00
- e04j0.88
- Missions -a system that combines delegation,1.00
- er1.00
- creator-verifier, broadcast, and negotiation into a single1.00
- workflow1.00
- kel0.94
- HOW TO USE THEM0.99
- er1.00
- NTRY1.00
- Describe a software1.00
- Scope it through1.00
- Approve the plan0.98
- Missions handles1.00
- goal.1.00
- conversation.1.00
- execution1.00
- ar1.00
- gineer1.00
- OPE1.00
- A mission is not a single agent session that runs for a long time.0.99
- It's an ecosystem of agents coordinating through structured handoffs and0.99
- shared state.1.00
- AI Engineer - Factory0.99
- 20261.00
- 31.00
- Google DeepMind0.99
- AlEngineer0.99
- EUROPE1.00
-
- INTRODUCING MISSIONS1.00
- Three-role architecture1.00
- THE ORCHESTRATOR1.00
- Plans features, milestones, and the1.00
- validation contract1.00
- CHILD1.00
- CHILD1.00
- Workers1.00
- Validators1.00
- Fresh context per feature.1.00
- Adversarial verification. Have1.00
- Implement, commit via git, hand off.1.00
- never seen the code before.1.00
- A key takeaway1.00
- The validation contract defines what "done" means before any1.00
- code is written0.99
- AI Engineer - Factory1.00
- 20261.00
- 41.00
-
- HOW THEY WORK0.99
- The validation loop0.98
- Tests written after implementation don't catch bugs. They confirm decisions.0.99
- AI Engineer - Factory0.99
- 20261.00
- 51.00
-
- HOW THEY WORK0.98
- The validation loop0.98
- Tests written after implementation don't catch bugs. They confirm decisions.1.00
- PLANNING PHASE1.00
- Validation Contract0.99
- Written by orchestrator during planning, before any code. Hundreds of assertions define correctness1.00
- independently of implementation.1.00
- AI Engineer - Factory0.98
- 20261.00
- 51.00
-
- HOW THEY WORK1.00
- The validation loop1.00
- Tests written after implementation don't catch bugs. They confirm decisions.1.00
- PLANNING PHASE1.00
- Validation Contract0.99
- Written by orchestrator during planning, before any code. Hundreds of assertions define correctness0.99
- independently of implementation.0.99
- After each milestone...0.99
- </>0.90
- </>0.88
- Scrutiny Validator1.00
- User-Testing Validator0.99
- Runs tests, typechecking,0.98
- Acts like a QA engineer.1.00
- linting. Spawns code1.00
- Launches the app,1.00
- review agents for each1.00
- navigates via computer-use1.00
- completed feature.1.00
- & verifies flows1.00
- end-to-end.1.00
- Neither validator has ever seen the code. Validation is adversarial by design.1.00
- AI Engineer - Factory0.98
- 20261.00
- 51.00
-
- HOW THEY WORK0.99
- Structured handoffs1.00
- How agents stay coherent over days, not just minutes.0.99
- New Task1.00
- Every worker reports1.00
- What was implemented1.00
- Execute1.00
- What was left undone1.00
- Encode Skill0.96
- Commands run + exit codes1.00
- CONTINUOUS1.00
- Issues discovered1.00
- LEARNING1.00
- Whether procedures were followed0.98
- Observe1.00
- Learn1.00
- AI Engineer - Factory0.98
- 20261.00
- 61.00
-
- lind0.77
- AlEngineer0.96
- EUROPE1.00
- HOW THEY WORK1.00
- toner.ai0.93
- Structured handoffs1.00
- How agents stay coherent over days, not just minutes.1.00
- ev1.00
- Every worker reports0.99
- What was implemented0.98
- What was left undone1.00
- Commands run + exit codes1.00
- CONTINUOUS1.00
- Issues discovered1.00
- LEARNING1.00
- Whether procedures were followed0.99
- ARE1.00
- EURC0.97
- AI Engineer - Factory0.99
- 20261.00
- 61.00
- AlEngineer0.99
- Braintrust1.00
- WorkOS1.00
- OpenAl0.93
- EUROPE1.00
-
- AlEngineer0.89
- EUROPE1.00
- HOW THEY WORK1.00
- Why serial beats parallel (mostly)0.99
- OpenAl0.98
- Engineer1.00
- EUROPE1.00
- The next question, how should this execute...1.00
- Aici0.72
- AlEngineer0.95
- WHAT PEOPLE EXPECT1.00
- WHAT ACTUALLY WORKS0.96
- EUROPE1.00
- Parallelism1.00
- Serial Execution0.99
- Agents conflict and step on each other1.00
- Features execute one at a time. Each worker0.98
- ngineer1.00
- EUROPE0.99
- CNRY0.81
- Duplicate work and inconsistent architecture1.00
- Coordination overhead eats speed gains1.00
- Every conflict burns tokens0.99
- inherits the full codebase from the last0.99
- Parallelism is reserved for work that can't1.00
- through git1.00
- ENTED BY1.00
- conflict: codebase exploration, API research,1.00
- documentation reads, and validation reviews.0.98
- DeepMind1.00
- Slower on paper. But for multi-day runs,0.99
- IEngineer0.97
- correctness compounds.0.98
- EUROPE0.99
- AI Engineer - Factory0.99
- 20261.00
- 71.00
- Engineering the future of Al1.00
- AlEngineer0.98
- EUROPE1.00
-
- HOW THEY WORK0.97
- Why serial beats parallel (mostly)1.00
- The next question, how should this execute...0.99
- WHAT PEOPLE EXPECT1.00
- WHAT ACTUALLY WORKS1.00
- Parallelism1.00
- Serial Execution1.00
- Agents conflict and step on each other1.00
- Features execute one at a time. Each worker1.00
- Duplicate work and inconsistent architecture0.99
- inherits the full codebase from the last1.00
- through git0.99
- Coordination overhead eats speed gains0.99
- Every conflict burns tokens1.00
- Parallelism is reserved for work that can't1.00
- conflict: codebase exploration, API research,0.99
- documentation reads, and validation reviews.1.00
- Slower on paper. But for multi-day runs,0.98
- correctness compounds.1.00
- AI Engineer - Factory0.99
- 20261.00
- 71.00
-
- USING MISSIONS1.00
- ∴Droid0.98
- Mission Control0.99
- Mission Control1.00
- -/Development/note-tracker1.00
- TIME 56m 54s0.99
- Input 324.0K0.97
- Cached 16.8M0.99
- Output 111.0K0.96
- RUNNING1.00
- 3/17 [+6]0.91
- A dedicated view for multi-day autonomous1.00
- Active Feature scrutiny-validator-app-shell0.99
- Features1.00
- 3/111.00
- work. Monitor, redirect, or close your0.99
- skill scrutiny-validator0.98
- bootstrap-tauri-vorkspace0.98
- laptop and come back tomorrow.1.00
- ailestone app-shell0.99
- ✓ menu-bar-shell-and-lifecycle0.95
- Preconditions1.00
- scrutiny-validator-app-shell0.99
- All implementation features for milestone "app-shell" are complete0.99
- categories-and-tagging-experience0.99
- Expected Behavior1.00
- due-date-parsing-and-note-metadata0.98
- Validators pass (test, typecheck, lint)0.99
- I0.92
- google-calendar-auth-and-secure-storage1.00
- Reviev subagents spawned for each feature0.98
- -3 more0.92
- Findings synthesized into scrutiny report1.00
- Progress Log1.00
- 1-9 of 180.98
- Description1.00
- Scrutiny validation for milestone "app-shell". Runs test suite, typecheck,0.99
- <1m ago0.89
- #95fb1b2d started [scrutiny-validator-app-shell]0.99
- and lint. Spawns review subagents for each completed feature. Synthesizes0.99
- <1n ago0.85
- Milestone validation: app-shell1.00
- findings. Always returns to orchestrator.0.98
- <1m ago0.95
- #d1b6b4a7 completed [quick-note-save-and-main-list]0.99
- W0.96
- 5m ago0.92
- 5m ago0.90
- #afc827a8 conpleted [menu-bar-shell-and-lifecycle] √0.97
- #d1b6b4a7 started [quick-note-save-and-main-list]0.99
- 13m ago0.98
- #afc827a8 started [menu-bar-shell-and-lifecycle]1.00
- 21= ago0.87
- 13m ago0.96
- #1963e245 conpleted [bootstrap-tauri-workspace]0.99
- #1963e245 started[bootstrap-tauri-workspace]0.99
- 21m ago1.00
- Run started: Mission artifacts authored, repo scaffold..0.99
- Active Worker #1 scrutiny-validator-app-shell1.00
- Duration 7s1.00
- ## Your Assigned Featurejson { "id": "scrutiny-validator-app-shell", "description": "Scrutiny validation for milestone \"app-shell\". Ru0.98
- ns test suite, typecheck, and lint. Spawns reviev subagents for each completed feature. Synthesizes findings. Always returns to orchestrato...0.98
- F Features1.00
- Workers0.98
- M Models0.98
- P Pause0.99
- D Mission Dir0.95
- Ctr1+T Back To Orchestrator0.99
- Monitor from a bird's eye view1.00
- AI Engineer - Factory0.99
- 20261.00
- 81.00
-
- USING MISSIONS1.00
- ∴Droid0.97
- Mission Control1.00
- "Mission Control0.97
- -/Development/note-tracker0.99
- TIME 56m 54s0.98
- Input: 324.0K0.94
- Cached 16.8M0.98
- Output 111.0K0.97
- Workers(4)0.99
- A dedicated view for multi-day autonomous1.00
- Al1 (4) | Active (1) | Completed [3) | Failad (8)0.91
- work. Monitor, redirect, or close your0.98
- laptop and come back tomorrow.0.98
- Session1.00
- Start0.99
- Duration1.00
- Status1.00
- Input1.00
- Cached1.00
- Output0.90
- Feature1.00
- 95fb1b2d1.00
- Running0.97
- scrutiny-validator-app-shell1.00
- √20.92
- afc827a81.00
- 1963e2451.00
- d1b6b4a71.00
- 15:020.96
- 14:461.00
- 14:541.00
- 5m 47s0.93
- 8m 1s0.97
- Success1.00
- Success1.00
- Success1.00
- 541.00
- 480.98
- 680.97
- 5.0M0.94
- 2.0M0.97
- 7.3M0.97
- 21.BK0.99
- 15.1K0.99
- 18.BK0.99
- quick-note-save-and-main-list1.00
- bootstrap-tauri-workspace1.00
- menu-bar-shell-and-lifecycle1.00
- I0.80
- Enter1.00
- ↑↓ Select0.86
- Enter Viev0.98
- T Filter0.98
- Ese Back0.96
- F Features0.98
- M Models0.98
- Inspect all the workers0.98
- AI Engineer - Factory0.98
- 20261.00
- 81.00
-
- USING MISSIONS1.00
- □0.51
- ∴Droid0.98
- Mission Control1.00
- ∴Mission Control0.92
- -/Development/note-tracker1.00
- TIME 56m 54s0.98
- Input 324.0K0.94
- Cached 16.8M1.00
- Output 111.0K0.97
- Worker Session0.99
- A dedicated view for multi-day autonomous1.00
- Feature: quick-note-save-and-main-list1.00
- Session:d1b6b4a7-30b2-488f-a1ca-9dea374bb8b80.99
- work. Monitor, redirect, or close your0.97
- laptop and come back tomorrow.1.00
- Model: Opus 4.6 [High)0.95
- Status: Success Duration: 5m 47s0.98
- Tokens: Input 40 Cached 2.0M Output 18.0K0.98
- ↑ 93 more above0.93
- Plan 1. [completed] Run baseline tests (pnpm test + cargo test)2. [completed] Add SQLite dependency and create note persistence layer in Rust...1.00
- →T0ODO List Updated0.92
- I0.91
- Execute cd /Users/luke/Development/note-tracker && git status0.99
- →On branch main Changes not staged for commit: (use "git add <file>... to update what will be committed) (use "git restore <file>...* to disc0.97
- ard changes in working directory) modified: src-tauri/Cargo.lock modified: src-tauri/Cargo.toml modified: src-tauri/src/lib.rs modified: src/__0.99
- tests../App.test.tsx modified: src/_.tests./QuickNote.test.tsx modified: src/components/MainApp.tsx modified: src/components/QuickNote.tsx mo..0.97
- Execute cd /Users/luke/Development/note-tracker && git add -A && git commit -m "Add note persistence vith SQLite and quick-note save path - Ad..0.99
- 644 src-tauri/src/db.rs create mode 188644 src/lib/tauri.ts create0.98
- -[main bc131a1] Add note persistence with SQLite and quick-note sav0.98
- M1.00
- 8644 src/lib/types.ts [Process exited with code 8]0.98
- 11 files changed, 619 insertions(+), 27 deletions(-) create mode 1800.98
- Execute cd /Users/luke/Development/note-tracker && git status && git rev-parse HEAD0.99
- →On branch main nothing to comit, working tree clean bc131a15bdedf4c13d7a85d55744c231fcddfb4d [Process exited with code 8]0.99
- ∴Clean working tree. Now let me call EndFeatureRun:0.99
- End Feature Run success1.00
- →{"recorded":true,"nextAction":"continue","message":"Another worker will continue with next feature, IMPORTANT: Your session is nou complete.0.99
- Do not make any further tool calls or continue working. End your turn immediately."}0.99
- ↑↓ Scroll0.86
- [] Density0.95
- g Top0.94
- G Bottom0.89
- s Interrupt/Chat0.99
- h Handoff0.95
- Ese Back0.94
- Read the handoffs1.00
- AI Engineer - Factory0.99
- 20261.00
- 81.00
Transcript
197 cues· 2,923 words· 17,019 chars
- 0:15 Hi everyone, my name is Luke, and my goal is that 20 minutes from now, you'll be able to assemble agent teams that can complete tasks orders of magnitude harder than what you can complete with a single agent today.
- 0:27 A little bit about me.
- 0:30 I come from a background in dev tools.
- 0:32 About two and a half years ago, I started a project at Block, which is where I was working at the time, and that project evolved into Goose.
- 0:40 Goose is now one of the leading coding agents that is open source, and it recently was donated to the Agentic AI Foundation.
- 0:50 So it's been really cool to see.
- 0:53 Nowadays I work at Factory, where I lead our core agent harness.
- 0:57 And Factory's mission is to bring autonomy to the entire software development life cycle.
- 1:04 So I want to start off with a claim.
- 1:06 The bottleneck in software engineering nowadays is not intelligence.
- 1:10 It's now limited by human attention.
- 1:13 Even the best engineers can only complete a couple of tasks at a time.
- 1:17 They may have a backlog of 50 features, but they can only drive a few forward per day because every task requires their attention, every commit needs their review.
- 1:26 Today's models are smart enough to figure out all 50 of these tasks, but there's not enough bandwidth to supervise their implementation.
- 1:36 So we kept asking ourselves, what if a human decides what to build and then a system figures out how to do so?
- 1:43 An agent could just work for hours, for days, and you come back to finish work.
- 1:47 So that's what I'm here to talk about.
- 1:50 When you start researching multi-agent frameworks and systems, you quickly realize that the field's a bit of a mess.
- 1:56 Everyone has their own framework, their own terminology, their own opinions of what works and doesn't work.
- 2:02 And so I want to propose a simple taxonomy.
- 2:05 There's five frontier multi-agent frameworks.
- 2:07 One is delegation.
- 2:09 This is where one agent spawns another agent, and the parent agent may say, go figure out the database schema, and then gets a response back.
- 2:17 This is the simplest form of multi-agent communication as what most people implement first.
- 2:23 You have, you know, sub-agents and coding tools are the most common example.
- 2:28 The other one is creator verifier, right, where one agent builds something and then you have another agent that checks that work.
- 2:35 And the key here is a separation of concerns.
- 2:38 The agent that implemented the code has sunk cost bias.
- 2:43 It wants that code to work.
- 2:45 A fresh agent with fresh context is way more likely to find issues, and this is why we do code review as humans as well.
- 2:52 Another one is direct communication.
- 2:54 This is when agents communicate without a central coordinator.
- 2:57 It's kind of like DMing each other.
- 3:00 It's hard to get right, though, because state fragments across conversations without that coordinator, and there's no single source of truth.
- 3:09 The next one is negotiation.
- 3:11 Negotiation is when agents communicate, but over a shared resource.
- 3:16 So that might be they want to use the same API.
- 3:19 They want to modify the same portion of the code base.
- 3:23 But negotiation doesn't need to be adversarial.
- 3:25 In fact, the best use case is when there's net positive sum trading.
- 3:30 And that's when agents have a potential win-win situation while interacting.
- 3:36 And then the last one is broadcast.
- 3:38 And that is when one agent sends information to many.
- 3:41 Think of it like status updates, new context that applies to everyone, new shared constraints.
- 3:48 It's a bit less flashy than the other ones, but it's critical for maintaining coherence over long-running tasks.
- 3:56 And so when you have all of these different building blocks, how do you assemble that into a system that can run for many days?
- 4:03 So missions is our answer.
- 4:05 It's a system that combines four of those, delegation, creator-verifier, broadcast and negotiation,
- 4:13 into a single workflow.
- 4:14 You describe a goal.
loading
Chapters
- 0:00 Introduction to multi-agent systems and the bottleneck of human attention
- 1:50 Taxonomy of five frontier multi-agent frameworks
- 4:04 Introducing 'Missions': The three-role architecture (Orchestrator, Workers, Validators)
- 6:34 The importance of validation contracts for consistent quality
- 8:09 Maintaining long-term context through structured handoffs
- 9:17 The case for serial execution over parallel execution
- 10:30 Mission control: Monitoring agent progress
- 11:22 Strategic model selection per role ('Droid whispering')
- 13:06 Production data analysis: Building a Slack clone
- 14:34 Designing systems that improve with each model generation
- 15:51 Conclusion: The shifting economics of software engineering