read-only demo

Videos vljxQZfJ9wY

Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs

index_state ready data_status ok

AI Engineer· published 2026-06-25· 0:08:12· en-US· indexed 2026-08-11 10:13

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:30, 1 of 1 keyframes kept
  2. Shot 1, 0:30 to 1:00, 0 of 1 keyframes kept
  3. Shot 2, 1:00 to 1:39, 1 of 1 keyframes kept
  4. Shot 3, 1:39 to 2:09, 1 of 1 keyframes kept
  5. Shot 4, 2:09 to 2:41, 1 of 1 keyframes kept
  6. Shot 5, 2:41 to 3:10, 1 of 1 keyframes kept
  7. Shot 6, 3:10 to 3:41, 1 of 1 keyframes kept
  8. Shot 7, 3:41 to 3:42, 1 of 1 keyframes kept
  9. Shot 8, 3:42 to 4:13, 0 of 1 keyframes kept
  10. Shot 9, 4:13 to 4:37, 1 of 1 keyframes kept
  11. Shot 10, 4:37 to 5:04, 1 of 1 keyframes kept
  12. Shot 11, 5:04 to 5:33, 1 of 1 keyframes kept
  13. Shot 12, 5:33 to 6:09, 1 of 1 keyframes kept
  14. Shot 13, 6:09 to 6:36, 1 of 1 keyframes kept
  15. Shot 14, 6:36 to 7:10, 1 of 1 keyframes kept
  16. Shot 15, 7:10 to 7:38, 1 of 1 keyframes kept
  17. Shot 16, 7:38 to 8:11, 1 of 1 keyframes kept
  18. Shot 17, 8:11 to 8:11, 1 of 1 keyframes kept

18 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
117
whisperx 117
chunks
15
from 117 cues
keyframes
16
kept of 18 captured
frames with text
16
259 lines read
chapters
0
from the source metadata
keyframe bytes
2.2 MB
word timings on 117 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 10:11 1m 12s
stt done 2026-08-11 10:12 9s
chunk done 2026-08-11 10:12 0s
text_embed done 2026-08-11 10:12 0s
keyframe done 2026-08-11 10:12 19s
ocr done 2026-08-11 10:12 5s
frame_embed done 2026-08-11 10:13 3s

Frames, and what the machine read

  • 0:20 #0 done8 line(s)

    shot 0·sharpness 2757.0

    1. Nishant Gupta1.00
    2. Production Evals0.99
    3. for Agentic Systems0.99
    4. Measuring reliability beyond0.99
    5. accuracy. Building evaluation systems1.00
    6. for autonomousAl workflows.1.00
    7. Nishant Gupta1.00
    8. Tech Lead @ Meta0.99
  • 0:39 #1 skipped

    shot 1·duplicate of #0

  • 1:20 #2 done22 line(s)

    shot 2·sharpness 1822.4

    1. AI systems evolved faster than1.00
    2. Nishant Gupta1.00
    3. our evaluation methods1.00
    4. The Illusion0.99
    5. The Reality1.00
    6. 100%1.00
    7. Invisible Failure1.00
    8. Modes1.00
    9. 75%1.00
    10. Degraded Production1.00
    11. Behavior1.00
    12. 50%1.00
    13. Unpredictable User1.00
    14. Reliability Gaps0.99
    15. 90%0.99
    16. 25%1.00
    17. Benchmark Accuracy1.00
    18. 0%1.00
    19. T-01.00
    20. T+10ms1.00
    21. T+50ms1.00
    22. T+100ms1.00
  • 1:57 #3 done20 line(s)

    shot 3·sharpness 2113.7

    1. The Paradigm Shift: Output vs. Behavior1.00
    2. Nishant Gupta0.98
    3. Traditional LLM Evaluation0.99
    4. Agent Evaluation1.00
    5. Goal1.00
    6. Output Accuracy1.00
    7. 0.99
    8. Workflow Behavior1.00
    9. Environment1.00
    10. Static Datasets1.00
    11. 1.00
    12. Dynamic Contexts1.00
    13. Execution1.00
    14. Single-path Processing1.00
    15. 1.00
    16. Multi-path & Tool Dependent1.00
    17. Failure Mode0.98
    18. Hallucination1.00
    19. 1.00
    20. Cascading Workflow Failure1.00
  • 2:34 #4 done13 line(s)

    shot 4·sharpness 1921.8

    1. The Anatomy of Agentic Failure0.99
    2. Nishant Gupta1.00
    3. Apex: Coordination0.99
    4. (Multi-agent conflicts)1.00
    5. Tool Usage1.00
    6. (Wrong API execution)1.00
    7. Planning1.00
    8. (Bad task decomposition)1.00
    9. Reasoning1.00
    10. (Wrong decisions)1.00
    11. Foundation:1.00
    12. Memory (Incorrect retrieval)1.00
    13. & Safety (Unsafe actions)1.00
  • 3:01 #5 done14 line(s)

    shot 5·sharpness 1757.0

    1. Think like an SRE: Accuracy gives way to Reliab0.98
    2. Nishant Gupta1.00
    3. Task Success1.00
    4. Human1.00
    5. Tool Success1.00
    6. Satisfaction1.00
    7. System1.00
    8. Reliabiliity0.97
    9. Planning1.00
    10. Accuracy1.00
    11. Safety1.00
    12. Quality1.00
    13. Cost1.00
    14. Latency1.00
  • 3:37 #6 done13 line(s)

    shot 6·sharpness 1821.9

    1. The Evaluation Signal Hierarchy1.00
    2. Nishant Gupta1.00
    3. Production Telemetry1.00
    4. Low volume, maximum1.00
    5. signal value.0.96
    6. Scenario Evals1.00
    7. Operational1.00
    8. Volume1.00
    9. Targeted workflows.1.00
    10. Value1.00
    11. Benchmarks1.00
    12. High volume, low operational value.1.00
    13. The foundation, not the destination.1.00
  • 3:41 #7 done21 line(s)

    shot 7·sharpness 2128.9

    1. Offline Evals: Scenario-Driven Simulation1.00
    2. Nishant Gupta1.00
    3. Agent Sandbox0.97
    4. Discrete Outputs1.00
    5. Test Runner1.00
    6. Simulated1.00
    7. Execution1.00
    8. State1.00
    9. Metrics1.00
    10. Tools1.00
    11. Steps1.00
    12. Update1.00
    13. Completion Rate1.00
    14. 98.5%1.00
    15. Tool Correctness0.97
    16. 100%1.00
    17. Plan Quality0.98
    18. High1.00
    19. Simulated Cost1.00
    20. $0.051.00
    21. Scenario-driven, not prompt-driven.1.00
  • 4:06 #8 skipped

    shot 8·duplicate of #7

  • 4:29 #9 done13 line(s)

    shot 9·sharpness 1940.7

    1. Online Evals: The Production Stream1.00
    2. Nishant Gupta0.98
    3. Interactions1.00
    4. User1.00
    5. →D→0.88
    6. Evaluation1.00
    7. Gateway1.00
    8. Metadata & Telemetry1.00
    9. Analytics1.00
    10. Database1.00
    11. 0.79
    12. Production is your largest evaluation dataset.1.00
    13. Every interaction is signal.1.00
  • 4:50 #10 done11 line(s)

    shot 10·sharpness 3058.3

    1. Human-in-the-Loop Calibration1.00
    2. Nishant Gupta0.99
    3. Review Node1.00
    4. Automated1.00
    5. Alert1.00
    6. Correctness1.00
    7. Usefulness1.00
    8. Core Model0.99
    9. Trust1.00
    10. Safety1.00
    11. Humans are evaluators, not merely fallback systems.1.00
  • 5:30 #11 done18 line(s)

    shot 11·sharpness 2169.4

    1. The Silent Killer: Agent Drift1.00
    2. Nishant Gupta1.00
    3. Degradation Curve1.00
    4. 100%1.00
    5. Model Updates0.98
    6. Prompt Changes0.97
    7. sys ity0.86
    8. 80%1.00
    9. Tool/API Changes1.00
    10. 60%1.00
    11. User Behavior Shifts1.00
    12. 40%1.00
    13. Dropping success rates1.00
    14. 20%1.00
    15. Rising escalation rates1.00
    16. Spiking tool errors1.00
    17. 0%1.00
    18. Time1.00
  • 6:04 #12 done39 line(s)

    shot 12·sharpness 2061.2

    1. Observability is the Prerequisite0.99
    2. Nishant Gupta1.00
    3. The Trace Waterfall0.98
    4. Live Metrics Dashboard1.00
    5. Latency1.00
    6. 0.97
    7. User Prompt (35m - 70ms)0.99
    8. 345 ms1.00
    9. Planner Iteration (7tm - 35ms)0.94
    10. Retries1.00
    11. 70.85
    12. Vector DB Lookup - 8ms)0.98
    13. 2.5%0.87
    14. API A - 45ms0.93
    15. Step Costs1.00
    16. 70.99
    17. Parallel API Tool Calls0.98
    18. API B - 38ms0.98
    19. $0.0141.00
    20. API C - 52ms0.95
    21. Memory Usage1.00
    22. JT0.97
    23. 480 MB1.00
    24. 10ms0.94
    25. 20ms0.89
    26. 30ms1.00
    27. 40ms0.97
    28. 45ms1.00
    29. 50ms1.00
    30. 60ms1.00
    31. 70ms1.00
    32. 80ms0.88
    33. 90ms1.00
    34. 10ms1.00
    35. 11ms1.00
    36. 12ms1.00
    37. 13ms1.00
    38. Time1.00
    39. "You cannot evaluate what you cannot observe."0.99
  • 6:30 #13 done20 line(s)

    shot 13·sharpness 2455.6

    1. The Continuous Evaluation Loop0.99
    2. Nishant Gupta0.97
    3. A1.00
    4. Online Telemetry1.00
    5. detects drift/errors1.00
    6. Evaluation is an0.96
    7. Offline Scenario Evals1.00
    8. always-running1.00
    9. validate system updates1.00
    10. D1.00
    11. B1.00
    12. Triggers HITL review0.99
    13. before pushing back to1.00
    14. service, not a0.97
    15. for edge cases1.00
    16. Production1.00
    17. testing phase.1.00
    18. C0.95
    19. Human feedback feeds1.00
    20. into Offline Datasets0.99
  • 6:50 #14 done18 line(s)

    shot 14·sharpness 3251.6

    1. The Reliability Scorecard1.00
    2. Nishant Gupta1.00
    3. Engineering Metric0.99
    4. Business & Operational Impact0.99
    5. Task Completion1.00
    6. Business Outcome0.99
    7. Tool Success1.00
    8. Operational Reliability0.99
    9. Escalation Rate1.00
    10. Human Burden1.00
    11. Safety Violations0.99
    12. Risk Exposure0.97
    13. Latency1.00
    14. User Experience1.00
    15. Cost per Task0.99
    16. SystemScalability1.00
    17. Recovery Rate1.00
    18. System Resilience1.00
  • 7:24 #15 done19 line(s)

    shot 15·sharpness 2471.8

    1. TheAgentic Control Plane1.00
    2. Reference Architecture1.00
    3. Nishant Gupta1.00
    4. Control Plane1.00
    5. Scenario1.00
    6. Tracing &0.96
    7. Telemetry1.00
    8. HITL1.00
    9. Observability0.98
    10. Stream Analytics1.00
    11. Simulators1.00
    12. Calibration UI1.00
    13. (0ffline)0.97
    14. Agent1.00
    15. External1.00
    16. LLM1.00
    17. Orchestrator0.99
    18. Tools1.00
    19. Execution Plane1.00
  • 7:51 #16 done8 line(s)

    shot 16·sharpness 3361.8

    1. Architectural Imperatives0.99
    2. Nishant Gupta1.00
    3. 1. Offline benchmarks are necessary but insufficient.1.00
    4. 2. Agentic systems must be evaluated as full workflows.0.99
    5. 3. Production telemetry is the ultimate evaluation signal.0.99
    6. 4. Reliability always supersedes raw model accuracy.1.00
    7. 5. Evals are no longer tests; they are core infrastructure.0.99
    8. "You can't improve what you don't continuously evaluate.'0.99
  • 8:11 #17 done2 line(s)

    shot 17·sharpness 168.1

    1. Nishant Gupta1.00
    2. THANK YOU0.97

Transcript

117 cues· 1,135 words· 7,748 chars

  1. 0:03 Hey everyone, my name is Nishant Gupta and I'm a software engineering tech lead at Meta, working on building the training and inference infrastructure for the Meta Super Dangerous Lab and the infrastructure organization.
  2. 0:17 Today we are going to be talking about production events for our GenTech systems.
  3. 0:21 When most people hear the word evaluation, they think about benchmarks.
  4. 0:25 A model scores 90% on a benchmark.
  5. 0:28 A new version scores 92%.
  6. 0:30 The team celebrates.
  7. 0:31 But agentic systems have fundamentally changed what the evaluation means.
  8. 0:35 Today, the systems don't simply generate answers.
  9. 0:38 They plan.
  10. 0:39 They call tools.
  11. 0:40 They retrieve information.
  12. 0:41 They execute workflows.
  13. 0:42 They interact with the production infrastructure.
  14. 0:45 The question is no longer, did the model generate the right answer?
  15. 0:49 The question is, did the system behave correctly?
  16. 0:51 Today, I would like to discuss how evaluation is evolving from model benchmarking into production infrastructure.
  17. 1:02 This is the problem almost every AI organization is encountering today.
  18. 1:06 Offline benchmarks continue improving, yet production reliability often remains unpredictable.
  19. 1:12 Why is that?
  20. 1:13 because benchmarks measure model capability.
  21. 1:15 Production measures system behavior.
  22. 1:18 A benchmark doesn't capture tool failure, API outage, context changes, user variability, long running workflows.
  23. 1:25 And as systems become more autonomous, the gap between the benchmark performance and production performance grows.
  24. 1:30 The result is what many teams experience today.
  25. 1:34 High benchmark scores, as you can see, but unreliable production behavior.
  26. 1:41 Traditional LLM evaluation focus on outputs.
  27. 1:44 But we should ask the question, did the model produce a correct answer?
  28. 1:48 Agentic systems force us to ask a different question.
  29. 1:51 Did the system behave correctly?
  30. 1:53 Behavior includes planning quality, rule usage, execution, workflow execution, recovery from failures, decision making.
  31. 2:01 In other words, we are moving from evaluating answers to evaluating workflows.
  32. 2:06 And that requires fundamentally different evaluation architectures.
  33. 2:10 Many teams still think hallucinations are the primary AI failure modes.
  34. 2:15 In production, they are often just one category.
  35. 2:17 Agentech systems introduce an entire hierarchy of failure modes.
  36. 2:21 At the very foundation, the memory failures, retrieval failures, safety failures.
  37. 2:26 As you go up, you have to think about reasoning mistakes, poor planning, incorrect tool execution.
  38. 2:31 At the highest layer, you have to think about multi-agent coordination failures.
  39. 2:35 And this is why evaluating only model output misses the most production risk we observe.
  40. 2:43 One of the most useful mindset shifts is to stop thinking like researchers and start thinking like a SRE or a production engineer.
  41. 2:51 SREs don't measure success using accuracy.
  42. 2:53 They measure reliability, availability, latency, cost recovery.
  43. 2:57 And agentic systems require the same approach.
  44. 2:59 The goal is not maximizing the benchmark scores.
  45. 3:02 The goal is to maximize dependable outcomes.
  46. 3:05 Reliability becomes the North Star metric.
  47. 3:07 Accuracy becomes the only input.
  48. 3:14 In this pyramid is how I think, what I think about modern AI evaluation systems.
  49. 3:19 At the bottom, you can see their benchmarks.
  50. 3:21 They're useful, they're scalable, they're repeatable, but their operational value is limited.

Open at this second