Videos T0HhO4YtTfE
AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
Scene timeline
63 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 247
- whisperx 247
- chunks
- 51
- from 247 cues
- keyframes
- 42
- kept of 63 captured
- frames with text
- 42
- 338 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 7.8 MB
- word timings on 247 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 07:48 | 1m 24s |
stt |
done | — | 2026-08-11 07:49 | 33s |
chunk |
done | — | 2026-08-11 07:50 | 0s |
text_embed |
done | — | 2026-08-11 07:50 | 1s |
keyframe |
done | — | 2026-08-11 07:50 | 1m 19s |
ocr |
done | — | 2026-08-11 07:51 | 13s |
frame_embed |
done | — | 2026-08-11 07:51 | 7s |
Frames, and what the machine read
-
- MongoDB1.00
-
- AIEWF261.00
- AI System Design:1.00
- From Idea to Production1.00
- Apoorva Joshi1.00
- Staff AI/ML Developer Advocate @MongoDB1.00
-
- So, how will we “Vibe Code” in prod?0.99
- Forget the code exists, but1.00
- Erik Schulntz, Anthropic1.00
- NOT that the product exists!0.98
- Vibe coding in prod talk1.00
- World's Fair0.97
- d'sFair1.00
- IN THE NEAR FUTURE...0.99
- ...the person who communicates best becomes the0.99
- os0.99
- programmer.1.00
- "If you can communicate, you can program."1.00
- Fair1.00
- Sean Grove, OpenAI1.00
- The New Code talk0.99
-
- Specs are the new code0.98
- Evaluation criteria1.00
- System design1.00
- Product requirements1.00
-
- Product1.00
- System Design1.00
- Evaluation1.00
- Production1.00
- Requirements1.00
- and Monitoring1.00
- Readiness1.00
- Identify the business1.00
- Identify data sources1.00
- Define system0.99
- Optimize for accuracy0.99
- problem1.00
- and retrieval techniques0.99
- guardrails1.00
- Identify your1.00
- Select system1.00
- Define offline0.95
- Optimize for cost and0.98
- constraints1.00
- architecture and tech1.00
- evaluation metrics1.00
- latency1.00
- stack1.00
- Define the role of AI1.00
- Determine the UX and1.00
- Define online1.00
- Optimize for reliability1.00
- feedback loops1.00
- evaluation metrics1.00
- Define success metrics1.00
-
- Health insurance claims review0.99
-
- Background1.00
- Background1.00
- Health insurers and health funds around the world employ medical reviewers to assess0.99
- whether a requested treatment, procedure, or medication is covered under a patient's0.99
- policy. This requires cross-referencing clinical documentation, coverage policies, clinical1.00
- guidelines, and patient claims history. It is one of the most administratively intensive0.99
- roles in healthcare operations.1.00
- What we are building1.00
- We are building an internal claims review system for a fictitious health insurance0.99
- company called MDB Health. This system is used by medical reviewers to assess and1.00
- adjudicate claims. The external-facing system through which healthcare providers1.00
- submit claims and receive feedback on outcomes is out of scope.1.00
-
- Product1.00
- requirements1.00
-
- Business problem1.00
- Medical reviewers at MDB Health spend an average of 2 days1.00
- processing claim review requests — 4 times the industry standard0.99
- for non-urgent cases and 12 times the industry standard for0.99
- urgent ones. Delays at this scale postpone patient care,0.99
- particularly for time-sensitive treatments.1.00
- User-specific1.00
- Current state1.00
- Measurable1.00
- Solution-agnostic1.00
- Focused1.00
-
- Business constraints1.00
- All denial decisions must include a human reviewer.1.00
- 2.0.99
- Reviewers must be able to override any system recommendation.1.00
- 3.0.99
- Complex cases, as defined by MDB Health's internal clinical guidelines, must be reviewed1.00
- by a senior physician.0.97
- 4.0.97
- Patient data can only be processed within MDB Health's approved cloud environment.1.00
- 5.0.98
- Only models available on MDB Health's approved cloud provider are permitted.0.99
- 6.0.89
- All decisions must be auditable and traceable to a specific source document.1.00
-
- Performance constraints1.00
- 1. P95 response time for an automated recommendation must be under 5 minutes.0.99
- 2.0.98
- System must handle up to 5,000 requests per day.0.99
- 3.0.97
- Denial decisions must be communicated to the healthcare provider within 24 hours.0.99
- 4.0.85
- Monthly model (embedding, LLM) costs must not exceed $10,000.0.99
- 5.0.98
- 99.9% uptime SLA required1.00
-
- Define the role of AI1.00
- Role1.00
- Answer1.00
- Critical / Complementary1.00
- Complementary1.00
- Reactive / Proactive1.00
- Reactive1.00
- Level of autonomy1.00
- Semi-autonomous1.00
-
- Define success metrics1.00
- Reduce the average processing time for urgent claim review1.00
- requests from 2 days to1 hour within 90 days of launch.1.00
- Specific1.00
- Measurable1.00
- Achievable1.00
- Relevant1.00
- Time-bound1.00
-
- Data and retrieval1.00
-
- Inventory your data sources0.99
- What data does your AI need? Where does it reside?1.00
- Data1.00
- Where does it live?1.00
- Accessed as1.00
- Clinical guidelines1.00
- Confluence1.00
- PDF1.00
- Internal coverage policies1.00
- Confluence1.00
- PDF1.00
- Patients' claims history1.00
- MongoDB1.00
- BSON1.00
-
- Data updates1.00
- How often do the data sources update?0.99
- Data1.00
- Where does it live?1.00
- Clinical guidelines1.00
- Annually1.00
- Internal coverage policies1.00
- Quarterly1.00
- Patients' claims history0.99
- Hourly1.00
-
- Determine data processing needs1.00
- Data source1.00
- Data processing0.98
- Clinical guidelines0.99
- Chunking, embedding, metadata extraction1.00
- Internal coverage policies1.00
- Chunking, embedding, metadata extraction0.99
- Patients' claims history0.99
- Sensitive data handling1.00
Transcript
247 cues· 4,063 words· 23,652 chars
- 0:02 Hi, everyone.
- 0:03 I'm Apoorva, and I'm a data scientist turned developer advocate currently at MongoDB.
- 0:08 I spent the first years of my career building machine learning applications for various cybersecurity use cases, and I now use that applied machine learning knowledge to help AI builders successfully build AI applications with MongoDB and Voyage AI.
- 0:25 In this talk, you'll learn how to think through building AI systems end-to-end from idea to production.
- 0:31 We'll take a real-world use case and walk through all the steps of designing it.
- 0:36 Hopefully, in the end, you'll walk away with a repeatable framework that you can apply to any AI system you design.
- 0:46 Now, you might be thinking, in the age of AI, do we even need to think about what we build?
- 0:52 Just wipe code it and ship it, right?
- 0:54 But that's actually where I want to start.
- 0:57 Now, here's the problem with that.
- 1:01 Wipe coding works great when you're building for fun, the stakes are low, and you can easily eyeball whether the output of what you're building is right.
- 1:09 But the moment you're building something real, something other people depend on, something with real consequences, just ship it is actually kind of dangerous.
- 1:19 And it's not just me saying it.
- 1:20 Folks from Anthropic and OpenAI who are so bullish on AI coding are saying the same thing.
- 1:25 These are actually quotes from their talks over the past few months.
- 1:30 Specs are the new code.
- 1:32 The art is in defining the product requirements, the system design, and evaluation criteria so you can be confident that your AI coding buddies are building the right thing.
- 1:42 And the rest of the talk is just about that.
- 1:44 How do we do this well?
- 1:49 It's useful to think about it as a framework, four phases.
- 1:53 You start with product requirements.
- 1:55 What are you actually building, for whom, and what are the constraints?
- 2:00 Then system design, the data, the architecture, the patterns that actually help you meet these requirements.
- 2:07 Then comes evaluation and monitoring.
- 2:09 How do you know what's being built actually works before and also after you ship?
- 2:16 And finally, how do you optimize for cost, not just for accuracy, but also for cost, latency, and reliability before you ship and or as you find gaps in production?
- 2:30 Instead of talking abstractly through a framework, I thought why not apply it to a real world use case so you can see how exactly each decision flows from one stage of the framework to the next.
- 2:43 So let's take the example of a health insurance claims review system.
- 2:49 So adjudicating health insurance claims has historically been an extremely manual process where human medical reviewers have to cross-reference clinical guidelines, insurance coverage policies, patient history and stuff to decide whether or not a medical treatment procedure or medication is covered by a patient's health insurance policy.
- 3:12 Now, if you live somewhere that has free health care, even there, there's things that are covered and things that are not by your health care system.
- 3:21 So you're still subject to some version of this process.
- 3:27 Now, our goal with building the system is to see if we can improve or simplify the experience for medical reviewers with the help of AI.
- 3:38 So let's see what the product requirements stage of the framework for this application might look like.
- 3:45 So first thing to do is quantify the business problem.
- 3:49 The business problem should focus on something specific, clearly state who the users of the application are, the current state of things, and quantify the user's pain point.
- 4:02 It should not prescribe what the system is going to be, whether it's going to be an agent, a multi-agent system, something else.
- 4:09 That'll come later.
- 4:14 So here's what the business problem for our claims review system could look like.
- 4:18 Medical reviewers at MDB Health spend an average of two days processing claim review requests, which is four times the industry standard for non-urgent cases and 12 times the industry standard for urgent ones.
- 4:31 Delays at this scale postpone patient care, particularly for time-sensitive treatment.
- 4:38 So as you can see, this business problem is user specific.
- 4:41 It says this is meant for medical reviewers.
- 4:44 It states the current state, which is like this manual process that leads to them spending two days processing claim reviews.
- 4:54 It's measurable because it states some baselines.
- 4:58 It's solution agnostic.
- 4:59 It doesn't really tell you what the system should look like.
- 5:03 And it's very focused on a specific problem.
- 5:11 next you want to gather any business constraints for the application so any regulatory compliance requirements any constraints on data leaving your organization procurement constraints meaning are certain vendors not approved for use within the organization all of these would be good to know before you even start designing the system
- 5:33 So for our application, say these are some of the business constraints.
- 5:36 Patient data has to stay within the approved cloud environment.
loading
Chapters
- 0:00 Introduction to AI framework
- 1:53 Defining product requirements
- 3:38 Setting AI success metrics
- 8:35 Data strategy and retrieval
- 12:12 System architecture design
- 14:10 Applying design patterns
- 17:30 User experience and feedback
- 19:37 Tech stack considerations
- 20:56 Evaluation and monitoring
- 24:00 Monitoring in production
- 25:12 Optimizing for performance
- 27:17 Key takeaways and conclusion