Videos 2czYyrTzILg
From Chaos to Choreography: Multi-Agent Orchestration Patterns That Actually Work — Sandipan Bhaumik
Scene timeline
53 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 252
- whisperx 252
- chunks
- 50
- from 252 cues
- keyframes
- 36
- kept of 53 captured
- frames with text
- 36
- 546 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 4.7 MB
- word timings on 252 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 13:59 | 1m 23s |
stt |
done | — | 2026-08-10 14:00 | 28s |
chunk |
done | — | 2026-08-10 14:01 | 0s |
text_embed |
done | — | 2026-08-10 19:50 | 1s |
keyframe |
done | — | 2026-08-10 14:01 | 1m 12s |
ocr |
done | — | 2026-08-10 14:02 | 11s |
frame_embed |
done | — | 2026-08-10 19:50 | 6s |
Frames, and what the machine read
-
- MULTI-AGENT1.00
- ORCHESTRATION1.00
- PATTERNS1.00
- Sandipan Bhaumik0.98
- Data & Al Tech Lead0.98
- Databricks0.98
- April 20260.99
-
- • PROBLEM0.97
- War story, why complexity explodes0.99
- • PATTERNS0.97
- AGENDA1.00
- Coordination1.00
- State management (immutable)1.00
- Failure recovery1.00
- • IMPLEMENTATION0.98
- Production architecture1.00
- April 20260.95
-
- ONE AGENT0.98
- WORKS1.00
- Input1.00
- Output1.00
- One agent is a feature1.00
-
- FIVE AGENTS1.00
- A1.00
- ARE CHAOS1.00
- C0.99
- B1.00
- F1.00
- Five agents is a1.00
- distributed1.00
- E0.99
- systems problem1.00
-
- WAR STORY0.96
- THE RACE CONDITION1.00
- Incorrect1.00
- Risk1.00
- Ratings1.00
- Incorrectly1.00
- flagged1.00
- customers1.00
- Week 2:1.00
- Week 3 - 5:0.92
- 5 Agents Live in1.00
- Firefighting1.00
- Production1.00
- Delayed1.00
- Value1.00
-
- WAR STORY1.00
- THE BUG0.96
- 6801.00
- write score1.00
- read score1.00
- (T= 0ms)0.95
- (T=500ms)1.00
- A1.00
- B1.00
- stale data0.97
- Credit Score1.00
- Database1.00
- Cache1.00
- Risk Assessment1.00
- Calculation Agent1.00
- Agent1.00
-
- WAR STORY1.00
- THE BUG1.00
- Agent B is1.00
- reading1.00
- stale1.00
- 6801.00
- version of0.97
- write score1.00
- X1.00
- read score1.00
- (T= 0ms)0.98
- (T=500ms)1.00
- A1.00
- B1.00
- the data -0.99
- stale data0.99
- that leads1.00
- Credit Score1.00
- Database1.00
- Cache1.00
- Risk Assessment1.00
- to wrong1.00
- Calculation Agent1.00
- Agent1.00
- decisions1.00
-
- THE PROBLEM WASN'T THE LLM0.99
- We built a distributed1.00
- system...1.00
- ...without1.00
- distributed system thinking0.99
-
- THE1.00
- COMPLEXITY1.00
- Nu nts0.86
- CURVE1.00
- 5 agents = 25x comnlex0.93
- Coordination complexity1.00
-
- THE1.00
- COMPLEXITY1.00
- Nu nts0.82
- CURVE1.00
- 5 agents = 25x comnlex0.98
- Coordination complexity1.00
-
- PATTERNS1.00
-
- THE FUNDAMENTAL QUESTION1.00
- How Should Your Agents Coordinate?1.00
- VS0.99
- CHOREOGRAPHY1.00
- ORCHESTRATION1.00
- Decentralized. Event-driven. Autonomous1.00
- Centralized. Coordinated. Controlled0.98
-
- CHOREOGRAPHY1.00
- EVENT-DRIVEN COORDINATION1.00
- Research1.00
- Analysis1.00
- Report1.00
- Agent1.00
- Agent1.00
- Agent1.00
- Publish 'research_completed'0.99
- Subscribe & Consume0.99
- Publish'analysis_ready'1.00
- Subscribe & Consume1.00
- Message Bus0.96
-
- CHOREOGRAPHY1.00
- Complex1.00
- Dependencies1.00
- WHEN TO USE AND AVOID0.99
- Need1.00
- Loosely1.00
- High Agent1.00
- Centralized1.00
- Coupled1.00
- Autonomy1.00
- Rollback1.00
- Workflows1.00
- needed1.00
- Weak1.00
- Observability1.00
- Frequently1.00
- Strong1.00
- Added Agents1.00
- Observability1.00
- Infrastructure1.00
- Debugging is1.00
- TE0.70
- Difficult1.00
-
- ORCHESTRATION1.00
- CENTRALIZED COORDINATION1.00
- Step 10.95
- Agent A1.00
- Agent B1.00
- Step 2 (parallel)1.00
- Orchestrator1.00
- Step 2 (parallel)1.00
- Agent C1.00
- 601lf0.52
- Step 30.94
- Agent D0.99
-
- ORCHESTRATION1.00
- CENTRALIZED COORDINATION1.00
- Step 10.94
- Agent A0.99
- Agent B1.00
- Step 2 (parallel)0.98
- Orchestrator1.00
- Step 2 (parallel)1.00
- Agent C1.00
- Step 30.94
- Agent D0.99
-
- ORCHESTRATION1.00
- High1.00
- Autonomy1.00
- WHEN TO USE AND AVOID1.00
- Required1.00
- Can't Tolerate0.97
- Complex1.00
- Rollback1.00
- Bottleneck1.00
- Dependencies1.00
- Capability1.00
- Required0.99
- Constantly1.00
- Changing1.00
- Need1.00
- Stable1.00
- Workflows1.00
- Centralized1.00
- Workflows1.00
- Control1.00
- Need1.00
- Distributed1.00
- Scaling1.00
-
- DECISION1.00
- MATRIX1.00
- HH0.61
- CHOREOGRAPHY1.00
- HYBRID1.00
- Event-driven1.00
- Distributed1.00
- microservices1.00
- transactions1.00
- Autmy0.77
- →MO70.84
- SIMPLE1.00
- FULL ORCHESTRATION0.99
- ORCHESTRATION1.00
- Linear Workflows0.98
- Enterprise Workflows1.00
- SIMPLE1.00
- Workflow complexity1.00
- COMPLEX1.00
-
- Q0.94
- Q0.96
- STATE1.00
- MANAGEMENT1.00
Transcript
252 cues· 4,013 words· 23,681 chars
- 0:00 Hi everyone, I'm Sandy.
- 0:01 I've spent 18 years building data systems, a major part of it focusing on building and scaling distributed data systems in the cloud.
- 0:09 I've done it for multi-tenant systems for software and SaaS companies, and then for scaling data and AI platforms in regulated industries like financial services and healthcare.
- 0:19 I've learned a great deal about production grade distributed systems while I have been working at AWS and now in Databricks.
- 0:26 For the last two years, I've been deploying multi-agent AI systems in production and I have watched brilliant engineers make the same mistakes over and over.
- 0:36 They think adding more agents is just like adding more features.
- 0:40 It's not.
- 0:41 It's building a distributed system.
- 0:43 And today I'm going to show you the patterns that actually work when you make that transition.
- 0:48 These are lessons that I have learned working in the trenches.
- 0:51 And today I'm here to share it with you.
- 0:53 Here's what we are covering today.
- 0:55 First, the problem.
- 0:56 I'll share you a very basic production war story about race conditions and why complexity explodes when you go from one agent to five agents.
- 1:08 I'll talk about the patterns, choreography and orchestration patterns for coordination of agents.
- 1:13 I'll talk about state management, talk about failure recovery and how we can design for failure in production systems.
- 1:21 And then I'll share how a production grade architecture would look like in as simple way possible.
- 1:28 And I'll also show you an example on how we build this on Databricks.
- 1:31 So let's dive into it.
- 1:32 You see, one agent works beautifully.
- 1:35 You have got your LLM, some prompts, maybe a retrieval augmented generation pipeline, maybe some tool calls.
- 1:43 It demos great.
- 1:45 Leadership loves it.
- 1:46 You feel happy and your team is happy.
- 1:49 And then product comes back with a request that changes everything.
- 1:54 They want five more agents.
- 1:55 And here's what happens.
- 1:57 You think, okay, I know how to build agents and I will add five more.
- 2:01 Except now you have coordination problems.
- 2:04 Agent A produces data that agent B needs.
- 2:08 Agent C is waiting on both agent A and agent B.
- 2:11 agent b just updated the shared state that agent b was reading and agent e just crashed and took down this entire workflow this is no longer an ai problem this is a distributed system problem and most of you didn't sign up to be distributed systems engineer
- 2:30 let me tell you about a production deployment where this went very wrong we built a credit decisioning system for a financial services company the first agent credit score calculation worked perfectly it worked great in demos two weeks in production zero issues then we added four more agents income verification risk assessment fraud detection and final approval
- 2:52 We deployed all five.
- 2:53 In three days time, we started seeing weird approvals.
- 2:57 20% of the decisions had incorrect risk ratings.
- 3:01 Customers who should have been flagged were getting approved.
- 3:04 The business team was panicking.
- 3:06 It took us two days to find out what was happening.
- 3:09 One credit score agent calculated a score of 750 and wrote to the database.
- 3:14 The risk assessment agent, on the other hand, read from the database 500 milliseconds later and got a score of 680 for the same customer.
- 3:24 Why did it happen?
- 3:25 Because we had a caching layer for customer records.
- 3:28 The write to PostgreSQL succeeded, but the cache was not invalidated.
- 3:33 The risk agent read from the cache and it got stale data.
- 3:40 used it used the wrong score and made the wrong decision this is a classic distributive systems problem we had caching layer between the agents and the database cache invalidation failed and the agent was reading stale values the race condition wasn't in the database it was in the architecture multiple agents shared cache no coordination on cache invalidation
- 4:06 this took us quite a while to find the pattern it created delays in delivery and led to wrong decisions and here's the lesson we learned the problem was of course not with the model the problem wasn't with the prompts the problem was we built a distributed system without distributed system thinking and that's what kills multi-agent projects not bad ai but bad architecture
- 4:32 now i will show you the architecture that works we will also look into a production grade architecture but first let's understand why this complexity explodes so quickly now when you move from a one agent system to a multi-agent let's say five agent systems it doesn't get just five times harder it gets 25 times more complex coordination complexity grows exponentially one agent has got zero coordination problems
- 4:59 Two agents have got at least one connection.
- 5:01 Five agents have got at least 10 potential connections and coordination.
loading