Videos vljxQZfJ9wY
Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
Scene timeline
18 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 117
- whisperx 117
- chunks
- 15
- from 117 cues
- keyframes
- 16
- kept of 18 captured
- frames with text
- 16
- 259 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.2 MB
- word timings on 117 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 10:11 | 1m 12s |
stt |
done | — | 2026-08-11 10:12 | 9s |
chunk |
done | — | 2026-08-11 10:12 | 0s |
text_embed |
done | — | 2026-08-11 10:12 | 0s |
keyframe |
done | — | 2026-08-11 10:12 | 19s |
ocr |
done | — | 2026-08-11 10:12 | 5s |
frame_embed |
done | — | 2026-08-11 10:13 | 3s |
Frames, and what the machine read
-
- Nishant Gupta1.00
- Production Evals0.99
- for Agentic Systems0.99
- Measuring reliability beyond0.99
- accuracy. Building evaluation systems1.00
- for autonomousAl workflows.1.00
- Nishant Gupta1.00
- Tech Lead @ Meta0.99
-
- AI systems evolved faster than1.00
- Nishant Gupta1.00
- our evaluation methods1.00
- The Illusion0.99
- The Reality1.00
- 100%1.00
- Invisible Failure1.00
- Modes1.00
- 75%1.00
- Degraded Production1.00
- Behavior1.00
- 50%1.00
- Unpredictable User1.00
- Reliability Gaps0.99
- 90%0.99
- 25%1.00
- Benchmark Accuracy1.00
- 0%1.00
- T-01.00
- T+10ms1.00
- T+50ms1.00
- T+100ms1.00
-
- The Paradigm Shift: Output vs. Behavior1.00
- Nishant Gupta0.98
- Traditional LLM Evaluation0.99
- Agent Evaluation1.00
- Goal1.00
- Output Accuracy1.00
- →0.99
- Workflow Behavior1.00
- Environment1.00
- Static Datasets1.00
- →1.00
- Dynamic Contexts1.00
- Execution1.00
- Single-path Processing1.00
- →1.00
- Multi-path & Tool Dependent1.00
- Failure Mode0.98
- Hallucination1.00
- →1.00
- Cascading Workflow Failure1.00
-
- The Anatomy of Agentic Failure0.99
- Nishant Gupta1.00
- Apex: Coordination0.99
- (Multi-agent conflicts)1.00
- Tool Usage1.00
- (Wrong API execution)1.00
- Planning1.00
- (Bad task decomposition)1.00
- Reasoning1.00
- (Wrong decisions)1.00
- Foundation:1.00
- Memory (Incorrect retrieval)1.00
- & Safety (Unsafe actions)1.00
-
- Think like an SRE: Accuracy gives way to Reliab0.98
- Nishant Gupta1.00
- Task Success1.00
- Human1.00
- Tool Success1.00
- Satisfaction1.00
- System1.00
- Reliabiliity0.97
- Planning1.00
- Accuracy1.00
- Safety1.00
- Quality1.00
- Cost1.00
- Latency1.00
-
- The Evaluation Signal Hierarchy1.00
- Nishant Gupta1.00
- Production Telemetry1.00
- Low volume, maximum1.00
- signal value.0.96
- Scenario Evals1.00
- Operational1.00
- Volume1.00
- Targeted workflows.1.00
- Value1.00
- Benchmarks1.00
- High volume, low operational value.1.00
- The foundation, not the destination.1.00
-
- Offline Evals: Scenario-Driven Simulation1.00
- Nishant Gupta1.00
- Agent Sandbox0.97
- Discrete Outputs1.00
- Test Runner1.00
- Simulated1.00
- Execution1.00
- State1.00
- Metrics1.00
- Tools1.00
- Steps1.00
- Update1.00
- Completion Rate1.00
- 98.5%1.00
- Tool Correctness0.97
- 100%1.00
- Plan Quality0.98
- High1.00
- Simulated Cost1.00
- $0.051.00
- Scenario-driven, not prompt-driven.1.00
-
- Online Evals: The Production Stream1.00
- Nishant Gupta0.98
- Interactions1.00
- User1.00
- →D→0.88
- Evaluation1.00
- Gateway1.00
- Metadata & Telemetry1.00
- Analytics1.00
- Database1.00
- 目0.79
- Production is your largest evaluation dataset.1.00
- Every interaction is signal.1.00
-
- Human-in-the-Loop Calibration1.00
- Nishant Gupta0.99
- Review Node1.00
- Automated1.00
- Alert1.00
- Correctness1.00
- Usefulness1.00
- Core Model0.99
- Trust1.00
- Safety1.00
- Humans are evaluators, not merely fallback systems.1.00
-
- The Silent Killer: Agent Drift1.00
- Nishant Gupta1.00
- Degradation Curve1.00
- 100%1.00
- Model Updates0.98
- Prompt Changes0.97
- sys ity0.86
- 80%1.00
- Tool/API Changes1.00
- 60%1.00
- User Behavior Shifts1.00
- 40%1.00
- Dropping success rates1.00
- 20%1.00
- Rising escalation rates1.00
- Spiking tool errors1.00
- 0%1.00
- Time1.00
-
- Observability is the Prerequisite0.99
- Nishant Gupta1.00
- The Trace Waterfall0.98
- Live Metrics Dashboard1.00
- Latency1.00
- ↑0.97
- User Prompt (35m - 70ms)0.99
- 345 ms1.00
- Planner Iteration (7tm - 35ms)0.94
- Retries1.00
- 70.85
- Vector DB Lookup - 8ms)0.98
- 2.5%0.87
- API A - 45ms0.93
- Step Costs1.00
- 70.99
- Parallel API Tool Calls0.98
- API B - 38ms0.98
- $0.0141.00
- API C - 52ms0.95
- Memory Usage1.00
- JT0.97
- 480 MB1.00
- 10ms0.94
- 20ms0.89
- 30ms1.00
- 40ms0.97
- 45ms1.00
- 50ms1.00
- 60ms1.00
- 70ms1.00
- 80ms0.88
- 90ms1.00
- 10ms1.00
- 11ms1.00
- 12ms1.00
- 13ms1.00
- Time1.00
- "You cannot evaluate what you cannot observe."0.99
-
- The Continuous Evaluation Loop0.99
- Nishant Gupta0.97
- A1.00
- Online Telemetry1.00
- detects drift/errors1.00
- Evaluation is an0.96
- Offline Scenario Evals1.00
- always-running1.00
- validate system updates1.00
- D1.00
- B1.00
- Triggers HITL review0.99
- before pushing back to1.00
- service, not a0.97
- for edge cases1.00
- Production1.00
- testing phase.1.00
- C0.95
- Human feedback feeds1.00
- into Offline Datasets0.99
-
- The Reliability Scorecard1.00
- Nishant Gupta1.00
- Engineering Metric0.99
- Business & Operational Impact0.99
- Task Completion1.00
- Business Outcome0.99
- Tool Success1.00
- Operational Reliability0.99
- Escalation Rate1.00
- Human Burden1.00
- Safety Violations0.99
- Risk Exposure0.97
- Latency1.00
- User Experience1.00
- Cost per Task0.99
- SystemScalability1.00
- Recovery Rate1.00
- System Resilience1.00
-
- TheAgentic Control Plane1.00
- Reference Architecture1.00
- Nishant Gupta1.00
- Control Plane1.00
- Scenario1.00
- Tracing &0.96
- Telemetry1.00
- HITL1.00
- Observability0.98
- Stream Analytics1.00
- Simulators1.00
- Calibration UI1.00
- (0ffline)0.97
- Agent1.00
- External1.00
- LLM1.00
- Orchestrator0.99
- Tools1.00
- Execution Plane1.00
-
- Architectural Imperatives0.99
- Nishant Gupta1.00
- 1. Offline benchmarks are necessary but insufficient.1.00
- 2. Agentic systems must be evaluated as full workflows.0.99
- 3. Production telemetry is the ultimate evaluation signal.0.99
- 4. Reliability always supersedes raw model accuracy.1.00
- 5. Evals are no longer tests; they are core infrastructure.0.99
- "You can't improve what you don't continuously evaluate.'0.99
-
- Nishant Gupta1.00
- THANK YOU0.97
Transcript
117 cues· 1,135 words· 7,748 chars
- 0:03 Hey everyone, my name is Nishant Gupta and I'm a software engineering tech lead at Meta, working on building the training and inference infrastructure for the Meta Super Dangerous Lab and the infrastructure organization.
- 0:17 Today we are going to be talking about production events for our GenTech systems.
- 0:21 When most people hear the word evaluation, they think about benchmarks.
- 0:25 A model scores 90% on a benchmark.
- 0:28 A new version scores 92%.
- 0:30 The team celebrates.
- 0:31 But agentic systems have fundamentally changed what the evaluation means.
- 0:35 Today, the systems don't simply generate answers.
- 0:38 They plan.
- 0:39 They call tools.
- 0:40 They retrieve information.
- 0:41 They execute workflows.
- 0:42 They interact with the production infrastructure.
- 0:45 The question is no longer, did the model generate the right answer?
- 0:49 The question is, did the system behave correctly?
- 0:51 Today, I would like to discuss how evaluation is evolving from model benchmarking into production infrastructure.
- 1:02 This is the problem almost every AI organization is encountering today.
- 1:06 Offline benchmarks continue improving, yet production reliability often remains unpredictable.
- 1:12 Why is that?
- 1:13 because benchmarks measure model capability.
- 1:15 Production measures system behavior.
- 1:18 A benchmark doesn't capture tool failure, API outage, context changes, user variability, long running workflows.
- 1:25 And as systems become more autonomous, the gap between the benchmark performance and production performance grows.
- 1:30 The result is what many teams experience today.
- 1:34 High benchmark scores, as you can see, but unreliable production behavior.
- 1:41 Traditional LLM evaluation focus on outputs.
- 1:44 But we should ask the question, did the model produce a correct answer?
- 1:48 Agentic systems force us to ask a different question.
- 1:51 Did the system behave correctly?
- 1:53 Behavior includes planning quality, rule usage, execution, workflow execution, recovery from failures, decision making.
- 2:01 In other words, we are moving from evaluating answers to evaluating workflows.
- 2:06 And that requires fundamentally different evaluation architectures.
- 2:10 Many teams still think hallucinations are the primary AI failure modes.
- 2:15 In production, they are often just one category.
- 2:17 Agentech systems introduce an entire hierarchy of failure modes.
- 2:21 At the very foundation, the memory failures, retrieval failures, safety failures.
- 2:26 As you go up, you have to think about reasoning mistakes, poor planning, incorrect tool execution.
- 2:31 At the highest layer, you have to think about multi-agent coordination failures.
- 2:35 And this is why evaluating only model output misses the most production risk we observe.
- 2:43 One of the most useful mindset shifts is to stop thinking like researchers and start thinking like a SRE or a production engineer.
- 2:51 SREs don't measure success using accuracy.
- 2:53 They measure reliability, availability, latency, cost recovery.
- 2:57 And agentic systems require the same approach.
- 2:59 The goal is not maximizing the benchmark scores.
- 3:02 The goal is to maximize dependable outcomes.
- 3:05 Reliability becomes the North Star metric.
- 3:07 Accuracy becomes the only input.
- 3:14 In this pyramid is how I think, what I think about modern AI evaluation systems.
- 3:19 At the bottom, you can see their benchmarks.
- 3:21 They're useful, they're scalable, they're repeatable, but their operational value is limited.
loading