Videos ObTPqBGsEbA
The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
Scene timeline
85 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 403
- whisperx 403
- chunks
- 64
- from 403 cues
- keyframes
- 58
- kept of 85 captured
- frames with text
- 58
- 1,391 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 10.2 MB
- word timings on 403 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 17:01 | 1m 46s |
stt |
done | — | 2026-08-10 17:03 | 40s |
chunk |
done | — | 2026-08-10 17:04 | 0s |
text_embed |
done | — | 2026-08-10 19:52 | 1s |
keyframe |
done | — | 2026-08-10 17:04 | 3m 06s |
ocr |
done | — | 2026-08-10 17:07 | 29s |
frame_embed |
done | — | 2026-08-10 19:52 | 10s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- The Producti1.00
- AlEngineer0.99
- EUROPE1.00
- AI Playbook1.00
- Learnings from deploying agents in enterprises0.98
- SWIACC0.66
- Sandipan Bhaumik1.00
- Data & Al Tech Lead0.99
- Databricks1.00
- AlEngineer0.99
- EUROPE1.00
-
- The Producti0.97
- AlEngineer0.99
- EUROPE1.00
- AI Playbook1.00
- Learnings from deploying agents in enterprises1.00
- Sandipan Bhaumik0.98
- Data & Al Tech Lead0.96
- Databricks1.00
- AlEngineer0.99
- EUROPE1.00
-
- The Producti0.97
- AlEngineer0.99
- EUROPE1.00
- AI Playbook1.00
- Learnings from deploying agents in enterprises0.99
- EUIACC0.80
- Sandipan Bhaumik1.00
- Data & Al Tech Lead0.96
- Databricks1.00
- AlEngineer0.99
- EUROPE1.00
-
- THEPROBLEM1.00
- The pattern you already know0.99
- *★★0.63
- AIE1.00
- ★1.00
- Weeks 1-40.96
- Weeks 4–80.95
- Weeks 8-120.99
- Week 141.00
- Month 61.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Pick models.0.98
- Build features.1.00
- Demo to leaders.0.99
- "Why is Al1.00
- $$in projects0.99
- Looks great.0.96
- Sign-off.Ship.1.00
- b*s*-ing us?"0.99
- failed last year.0.97
- Sound familiar? You're not here because you haven't seen this. You're here because you want to stop it.1.00
- AIEngin0.95
- EUROPE1.00
- Engineering the future of Al0.99
- AlEngineer1.00
- 20261.00
-
- THE INSIGHT1.00
- The AI is the easy part1.00
- You can't debug what you can't see.1.00
- 011.00
- *★*0.57
- Observability gap1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- You can't improve what you can't measure.1.00
- 021.00
- Evaluation gap1.00
- You can't trust what you can't explain.1.00
- 031.00
- Governance gap1.00
- AIEngin0.93
- EUROPE1.00
- AlEngineer0.96
- AlEngineer0.97
- 20260.99
- EUROPE1.00
-
- THE INSIGHT1.00
- The AI is the easy part0.98
- You can't debug what you can't see.1.00
- 011.00
- Observability gap1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- You can't improve what you can't measure.1.00
- 021.00
- Evaluation gap1.00
- You can't trust what you can't explain.0.99
- 031.00
- Governance gap1.00
- AlEngin0.95
- EUROPE1.00
- Engineering the future of Al1.00
- AlEngineer1.00
- 20261.00
-
- THE FRAMEWORK1.00
- The Five Pillars of Production AI0.98
- 011.00
- 021.00
- 031.00
- 041.00
- 051.00
- *★★0.61
- AIE1.00
- ★1.00
- Evaluation1.00
- Observability1.00
- Data Foundation1.00
- Orchestration1.00
- Governance1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Define success before1.00
- See everything, always1.00
- Question + Tracking1.00
- Patterns that scale1.00
- What keeps you in0.96
- code1.00
- data1.00
- production1.00
- Each pillar enables the next. Skip one and the whole thing breaks in production.1.00
- AlEngin0.93
- EUROPE1.00
- Engineering the future of Al0.99
- AlEngineer0.98
- 20260.99
-
- THE FR0.94
- The Five Pillars of Pr0.98
- AlEngineer0.99
- 011.00
- 021.00
- 031.00
- EUROPE1.00
- Evaluation1.00
- Observability1.00
- Data Foundatic0.97
- Define success before1.00
- See everything, always1.00
- Question + Trackir0.99
- code1.00
- data1.00
- Each pillar enables the next.0.98
- AlEngineer0.99
- EUROPE1.00
-
- THE FR0.92
- The Five Pillars of Pr0.99
- AlEngineer0.98
- 011.00
- 021.00
- 031.00
- EUROPE1.00
- Evaluation1.00
- Observability1.00
- Data Foundatic0.98
- Define success before1.00
- See everything, always1.00
- Question + Trackir0.99
- data1.00
- Each pillar enables the next. Skip one and the0.97
- AlEngineer0.99
- EUROPE1.00
-
- THE FRAMEWORK1.00
- The Five Pillars of Production AI0.99
- 011.00
- 020.99
- 031.00
- 041.00
- 051.00
- ★1.00
- AIE1.00
- ★1.00
- Evaluation1.00
- Observability1.00
- Data Foundation1.00
- Orchestration1.00
- Governance1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Define success before0.98
- See everything, always0.99
- Question + Tracking1.00
- Patterns that scale1.00
- What keeps you in1.00
- code1.00
- data1.00
- production1.00
- Each pillar enables the next. Skip one and the whole thing breaks in production.0.99
- AIEngin0.94
- EUROPE1.00
- AlEngineer0.97
- AlEngineer1.00
- 20261.00
- EUROPE1.00
-
- PILLAR 010.96
- Define success with numbers1.00
- 11.00
- Not 'accurate' — but 87% on fee disputes, <2% false positives, 60% deflection0.99
- Evaluation1.00
- rate. Numbers make model selection obvious.0.99
- First1.00
- AIE1.00
- ★1.00
- ★0.97
- Build test cases from real data1.00
- 21.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Your eval suite is0.99
- Pull 200+ real conversations from production logs. Anonymise them. Every case1.00
- needs pass/fail criteria — not 'looks good'.0.99
- your specification.1.00
- Without it, you're1.00
- guessing.1.00
- Wire automated grading1.00
- 31.00
- Evaluation pipeline that runs on every code change. Dashboard showing metrics.1.00
- You measure success before you have anything to measure.1.00
- AlEngin0.95
- EUROPE1.00
- AlEngineer0.97
- AlEngineer0.96
- 20261.00
- EUROPE1.00
-
- PILLAR 011.00
- Define success with numbers1.00
- 11.00
- Not 'accurate' — but 87% on fee disputes, <2% false positives, 60% deflection0.99
- Evaluation1.00
- rate. Numbers make model selection obvious.0.99
- First1.00
- *★★0.58
- AIE1.00
- ★1.00
- ★1.00
- Build test cases from real data1.00
- 21.00
- ★1.00
- ★1.00
- Your eval suite is1.00
- Pull 200+ real conversations from production logs. Anonymise them. Every case0.99
- needs pass/fail criteria — not 'looks good'.0.98
- your specification.1.00
- Without it, you're1.00
- guessing.1.00
- Wire automated grading1.00
- 31.00
- Evaluation pipeline that runs on every code change. Dashboard showing metrics.1.00
- You measure success before you have anything to measure.1.00
- AlEngin0.95
- EUROPE1.00
- Engineering the future of Al1.00
- AlEngineer1.00
- 20261.00
-
- PILLAR O0.98
- Three ray ers of evaluat0.99
- AlEngineer0.99
- Layer 1– Deterministic0.96
- EUROPE1.00
- validation, Response length bounds0.98
- PIl detection (NER + regex), Output format0.98
- Layer 2-Semantic0.96
- Correctness & groundedness, LLM-as-a-Judge, Non-0.97
- determinism fix: run each test 3x - flag variance0.96
- above threshold1.00
- EMIACC0.82
- Layer 3-Behavioural0.98
- Did it call the right tools, in the right order? Did it0.98
- escalate when confidence was low? Did it stay1.00
- within scope?1.00
- AlEngineer0.99
- EUROPE1.00
Transcript
403 cues· 6,084 words· 33,835 chars
- 0:15 All right, thank you for joining my session.
- 0:20 Thank you, mate.
- 0:22 I'm Sandy.
- 0:23 I'm a technical lead for data and AI at Databricks.
- 0:28 Prior to working in Databricks, I worked in Amazon Web Services for five years as a principal architect for data and AI.
- 0:35 In the past few years, I've worked extensively building and scaling data and AI platforms using distributed systems and technology.
- 0:43 And in the past couple of years, specifically, I've been working with customers trying to figure out what we do with this new AI technology.
- 0:52 When I say new AI, AI has been here for a long time, but we all started experimenting quite exponentially in the past couple of years, right?
- 1:03 I have learned a great deal of lessons from building demos and how to take those demos to production, working with different customers in B2B software industries and then regulated industries like financial services.
- 1:18 So in this session, I want to share a playbook, a framework that I put together from lessons that I have learned working in the trenches.
- 1:29 that you can take and apply on when you think about how to put your AI systems into production.
- 1:36 And I think this session is nicely placed in the afternoon because what you can do now is in this framework, you can feed the different knowledge, the knowledge that you have gathered attending these different sessions throughout the day and see where they fit in each of these elements in the framework.
- 1:52 So when I started two years ago, this is the pattern I noticed in every customer conversation, right?
- 1:58 So everyone wanted to do something with AI.
- 2:01 There was immense pressure from the top to do something, to build a demo.
- 2:06 And every conversation started with, let's choose the model, right?
- 2:09 And it was nobody's fault because the market was like that.
- 2:11 We were talking about models.
- 2:12 The models were new technology for us, right?
- 2:15 And every conversation started, shall we use GPT?
- 2:18 Shall we use cloud?
- 2:20 You know, there was huge debate within organizations.
- 2:23 Then you would choose a model, you'll build some features, offsets over what features to build for that application.
- 2:31 You would build that in a controlled environment, so predictable data sets, you know, limited scenarios, and then it looked great as a demo.
- 2:40 And then leadership would get happy, they would sign it off, and they'll put it into an environment, in a production environment.
- 2:48 Then after a few weeks, people would start asking questions that what the hell is AI doing, right?
- 2:56 Why is it not answering the questions the way we expected it to answer when we were doing the demos?
- 3:02 It would result in not only no realization in return on investment, but also loss of money and effort in building these demos that can never scale to production
- 3:15 Throughout these meetings, I gathered three insights that connect to everything that we are talking about when thinking about taking AI to production.
- 3:26 The first one is the observability gap.
- 3:28 When we use AI and put it into production, if we can't see what it is actually doing, if we can't trace every decision that it's making, it's no use in production.
- 3:39 Second is the evaluation gap.
- 3:41 A lot of these conversations that we were doing, we were not actually thinking about what is that one thing that we are measuring.
- 3:48 Yes, we talk about accuracy, we talk about latency, we talk about groundedness, but we were not defining what is that exact thing that matters to the business and how can we build a system that can continuously measure that.
- 4:03 Whether it's improving, whether it's not improving, what is that system that we need to build?
- 4:08 And that was that evaluation gap that I noticed.
- 4:10 And the third is the governance gap.
- 4:12 We were not actually thinking, what happens when AI fails in production?
- 4:16 Who's accountable?
- 4:17 Who do I go to when something happens at 3 AM in the morning?
- 4:22 Who needs to own the data assets that feed some AI responses?
- 4:27 What happens if AI
- 4:30 you know, talks nonsense to a customer, right?
- 4:35 What happens, right?
- 4:36 So there is no accountability, no governance around it.
- 4:39 And these three insights led me to build a framework on how
- 4:46 I think AI should be taken to production, and this has been implemented across multiple customer organizations, and I think this is something that you can pick up from here.
- 4:55 These are the five pillars, and these are absolutely what you need to think about even before starting a project, right?
- 5:01 Then you start, build them gradually, preferably in sequence, but in real life, I know that this sequence don't work, but these are the pillars that you have to know about and you have to think about when start building.
- 5:13 First one is evaluation.
loading