Videos APh1Vx0oLmQ
Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
Scene timeline
17 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 115
- whisperx 115
- chunks
- 13
- from 115 cues
- keyframes
- 16
- kept of 17 captured
- frames with text
- 16
- 456 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.9 MB
- word timings on 115 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:38 | 1m 02s |
stt |
done | — | 2026-08-11 05:39 | 8s |
chunk |
done | — | 2026-08-11 05:39 | 0s |
text_embed |
done | — | 2026-08-11 05:39 | 0s |
keyframe |
done | — | 2026-08-11 05:39 | 18s |
ocr |
done | — | 2026-08-11 05:40 | 7s |
frame_embed |
done | — | 2026-08-11 05:40 | 3s |
Frames, and what the machine read
-
- BUILDING DETERMINISTIC1.00
- INFRASTRUCTURE FOR0.99
- Nishant Gupta0.99
- DETERMINISTIC BOUNDARY [DB-S00]0.99
- NON-DETERMINISTIC1.00
- ENTROPY REDUCTION: 99.9%1.00
- PREDICTABILITY INDEX: 1.01.00
- SYSTEM VARIANCE: <0.01%0.99
- AIAGENTS1.00
- S0°0.82
- CONTROL ENCLOSURE [CE-99.99%]0.99
- THE EMERGING CONTROL PLANE1.00
- FOR AUTONOMOUS AI SYSTEMS0.99
- SIGNAL1.00
- RECTIFICATION1.00
- SIGNAL1.00
- SYNCHRONIZATION1.00
- STATE1.00
- SYNCHRONIZATION1.00
- 180.53
- 92-0.70
- 90°0.88
- RIGID FRAMEWORK1.00
- CONTROL ENCLOSURE1.00
- Nishant Gupta1.00
- [RF-AIZ]0.96
- [CE-99.99%]1.00
- Tech Lead @ Meta1.00
-
- Infrastructure was built for0.98
- Nishant Gupta1.00
- predictable microservices.1.00
- THE GREAT MISMATCH1.00
- Traditional1.00
- Autonomous1.00
- FRICTION POINTS &0.98
- Microservices1.00
- INCOMPATIBILITY ZONES0.98
- Al Agents0.96
- Stateless1.00
- →0.97
- Stateful1.00
- [STATE:0]0.99
- →0.89
- [STATE:PERSIST]0.99
- Deterministic1.00
- Probabilistic1.00
- [PATH:FIXED]0.98
- X0.87
- [PATH: DYNAMIC]0.96
- Request-response1.00
- Multi-step workflows1.00
- [FLOW: SYNC]0.96
- [FLOW:ASYNC]0.99
- Millisecond execution1.00
- Long-running1.00
- [TIME:<100MS]0.97
- [TIME: >MIN/HR]0.97
-
- Demos optimize for capability.1.00
- Nishant Gupta1.00
- Production demands reliability.1.00
- [STRUCTURAL INTEGRITY: CRITICAL]0.99
- Capabilities0.99
- Prompt0.94
- LM0.97
- 50.210.94
- Reliability1.00
- KISO0.88
- Core1.00
- [L0Y38 0N]0.60
- Reliability Core0.99
- [LOAD BEARING]0.99
- [LOAD BEARING]0.99
- [STRUCTURAL INTEGRITY: CRITICAL]0.99
- [ISOMETRIC VIEW: SYSTEM CROSS-SECTION]0.99
-
- Real production failures originate1.00
- the infrastructure, not the model.1.00
- Diagnostic Failure Tree0.99
- 1091.00
- Logic1.00
- Recursive1.00
- Workflow1.00
- 0.130.93
- reasoning loops0.99
- deadlocks1.00
- DCCC0.73
- 0330.84
- 1501.00
- 1001.00
- Stochastic1.00
- Action1.00
- Tool1.00
- Retry1.00
- Cost1.00
- Model Output1.00
- 0.30.92
- hallucinations1.00
- F391.00
- storms1.00
- explosions1.00
- 5551.00
- 1590.95
- State1.00
- Context1.00
- Memory1.00
- 0.JS0.71
- drift1.00
- poisoning1.00
- 5340.86
-
- The anatomy of an1.00
- [CRIT1.00
- Nishant Gupta0.99
- Exponential GPU0.98
- +85S%0.91
- agent retry storm.1.00
- Compute Spike0.98
- SPIKE1.00
- +4338%1.00
- SATURATION1.00
- RESOURCE1.00
- Step 4:1.00
- +5559%1.00
- COST1.00
- Recursive1.00
- EXPLOSION1.00
- reasoning loop1.00
- 4780.99
- 1001.00
- 5551.00
- locks up.1.00
- hallucinates1.00
- invalid API1.00
- parameters.1.00
- Step 1: Agent1.00
- Step 2:0.95
- Tool returns1.00
- error code.1.00
- attempts to fix0.98
- Step 3: Agent0.98
- hallucinates new1.00
- error but1.00
- RECURSION1.00
- LOOP1.00
- * INVALID0.94
- +0.13#0.83
- PARAMS1.00
- GPU c L OPda)0.52
- Step 5:0.96
- The1.00
- Consequence1.00
- invalid parameter.1.00
- 01'00.85
- 1601.00
- 1001.00
- Step 1: Agent1.00
- Step 3: Agent0.99
- Step 4:0.99
- hallucinates1.00
- invalid API1.00
- Step 2:1.00
- Tool returns1.00
- attempts to fix0.97
- error but0.99
- reasoning loop0.98
- Recursive1.00
- Step 5:1.00
- The1.00
- parameters.1.00
- error code.1.00
- hallucinates new1.00
- locks up.1.00
- Time (ms)1.00
- Consequence1.00
- invalid parameter.1.00
- 5180.89
- DATA SPIKE EVENT1.00
- [CRITICAL FAILURE PATH]0.99
- [TIMELINE AXIS0.98
-
- The platform decides.1.00
- Nishant Gupta0.99
- The model merely proposes.1.00
- Validate1.00
- Validate1.00
- Aligned1.00
- Output1.00
- Stochastic1.00
- Safe1.00
- Core1.00
- Execution1.00
- System1.00
- Compliance1.00
- Deterministic1.00
- Filter1.00
- Wrapper1.00
- Route1.00
- [ANALYSIS: STRUCTURAL INTEGRITY]1.00
- [FLOW: DETERMINISTIC CONTROL0.99
-
- The Agent Control Plane1.00
- ARCHITEC1.00
- Nishant Gupta1.00
- is the new infrastructure layer.1.00
- Agents require1.00
- an operating0.98
- system.1.00
- [Contzol Plane Blueprint]0.99
- Agent1.00
- Applications1.00
- Just as Kubernetes1.00
- Research1.00
- Ceding0.95
- Interaction1.00
- Passeosers1.00
- became the1.00
- control plane for1.00
- Memory Coordinator1.00
- Orchestration Engine1.00
- orchestrating1.00
- INDEX0.97
- containers, a1.00
- dedicated Agent1.00
- THE AGENT0.98
- Control Plane is0.99
- CONTROL PLANE0.97
- Safety Policy Node0.98
- Compute Scheduler1.00
- emerging to1.00
- !A!!!0.64
- govern the runtime1.00
- execution of1.00
- 区0.73
- VALIDATE & FILTER1.00
- RESOURCE OPTIKIZKTION0.97
- autonomous Al.1.00
- [ANALYSIIS: HIGH]0.99
- Foundational LLMs & Data APls0.99
- [COST 0ARAP: 10%]0.95
- 31.00
- [FLOW: DETERMINISTIC GOVERNANCE]0.99
-
- Logs are dead. Autonomous workflows1.00
- require multidimensional observability.1.00
- Nishant Gupta1.00
- Agent Trace Timeline0.98
- [Telemetry Sidebar]1.00
- T+0.1s0.92
- T+0.2s0.99
- T+8.1s1.00
- T+0.5s0.99
- T+1.0s1.00
- T+1.2s1.00
- [FLON: REASONING]0.99
- [FLON: REASONING]0.99
- LLM Decisions0.99
- Track 1:1.00
- Call Tool A1.00
- Decision:1.00
- Call Tool A0.99
- Decision:1.00
- Query Memory1.00
- Decision:1.00
- Query Memory1.00
- Decision:1.00
- Call Tool A1.00
- Decision:1.00
- [DATA: ACTIVE]0.99
- Track 2:1.00
- Plan:0.90
- Plan:1.00
- Plan:0.96
- Plans1.00
- Orchestration1.00
- Execute0.99
- Step 10.96
- Execute1.00
- Step 10.98
- Validate1.00
- Output1.00
- [FLOR: DRXTING]0.95
- Track 3:1.00
- Tool Calls &1.00
- Execution1.00
- [DATA: ACTIVE]0.99
- Tool A Executing1.00
- API Response1.00
- Received1.00
- [FLOW: REASONING]1.00
- [BATA: ACTIVE]0.96
- Track 4:1.00
- Read:1.00
- Write:1.00
- Memory Access1.00
- Context Block X0.99
- New State Y1.00
- [FLOW: REASONING]0.98
- [DATA: ACTIVE]0.97
- [04TA: ACTIVE]0.98
- Track 5:0.99
- State Transitions1.00
- Processing1.00
- State:1.00
- Waiting for Tool1.00
- State:1.00
- Updated1.00
- State:1.00
- Updated1.00
- State:1.00
- T+8.1s0.95
- T+0.1s0.95
- event1.00
- T+0.5s0.96
- T+1.0s0.98
- T+1.0s1.00
- 41.00
- [ANALYSIS: FLOW TRACING]1.00
- [DATA: CONTINUOUS STATE]0.99
-
- Shared memory coordinates1.00
- Nishant Gupta0.99
- chaotic multi-agent consistency1.00
- Shared Memory Architecture1.00
- Agent A1.00
- Agent B0.97
- State Store1.00
- [-25-3.00, 00, 5291]0.97
- Stale memory0.97
- Context drift1.00
- Agent C0.97
- Conflicting facts1.00
-
- Autonomous systems demand0.99
- Nishant Gupta1.00
- layered containment boundaries1.00
- Ring 5: Audit Layer0.99
- Ring 2: Tool Permissions0.99
- Ring 3: Policy Engine1.00
- BLOCKED1.00
- Model Output1.00
- Ring 4: Human Approval1.00
- Ring 5: Audit Layer0.99
-
- Human oversight is an escalation path,1.00
- not an operational bottleneck.1.00
- Workflow Routing Diagram1.00
- Decision Gate1.00
- Risk / Confidence Check0.97
- Confidence Score: >95X1.00
- Anomaly Detection: Low1.00
- Automated Agent Tasks1.00
- Automated Agent Tasks1.00
- 45°1.00
- 45°1.00
- Confidence Score: >95%0.98
- Anonaly Detection: Low0.99
- Edge Case1.00
- Approved State /0.97
- Routing1.00
- Resolution1.00
- Human Approval Node1.00
- Manual Review Required0.99
- [ANALYSIS: ESCALATION PATHWAYS]0.98
-
- Inference at scale fundamentally1.00
- becomes a cluster scheduling problem.1.00
- Elastic Scaling Graph1.00
- 3001.00
- Variable reasoning1.00
- depth1.00
- Long-running1.00
- workflows1.00
- 2001.00
- COME UMUUITS0.72
- Resource Contention1.00
- & GPU Limits1.00
- 1001.00
- Bursty Demand1.00
- 1001.00
- 2001.00
- 3001.00
- TIME (sec)0.96
-
- The Reliability Rosetta Stone.1.00
- 830.85
- Distributed Systems Pattern1.00
- Agent Equivalent1.00
- Circuit Breakers1.00
- Tool Isolation0.97
- [PATTERN: ISOLATION]0.98
- [PATTERN: ISOLATION]0.98
- Rate Limiting1.00
- Agent Limits0.96
- [LINIT_ENFORCEMENT]1.00
- [LINIT_ENFORCEMENT]1.00
- Retries1.00
- Controlled Recovery1.00
- [RECOVERY_LOGIC]1.00
- [RECOVERY_LOGIC]1.00
- Quotas1.00
- Cost Governance1.00
- [RESOURCE_CAPS]1.00
- [RESOURCE_CAPS]1.00
- Observability1.00
- Agent Tracing1.00
- [TRACEABILITY_SYSTEM]1.00
- [TRACEABILITY_SYSTEM]1.00
- 1250.81
- 1251.00
- [ANALYSIS: RELIABILITY MAPPING]0.99
-
- 中0.51
- Competitive advantage has shifted fron0.99
- prompt engineering to systems engineering.1.00
- THE PARADIGM SHIFT0.99
- PROMPTS1.00
- Compound Growth1.00
- Trajectory1.00
- MODELS1.00
- Commoditization1.00
- Piateau1.00
- =>=>>0.63
- Commoditization1.00
- Plateau1.00
- INFRASTRUCTURE1.00
- 2020.75
- 20231.00
- 20241.00
- 2025+1.00
-
- Al agents are distributed systems.1.00
- FINAL DIREC0.99
- Nishant Gupta1.00
- Treat them accordingly.1.00
- The future1.00
- of Al won't1.00
- ..80.54
- be won by1.00
- CORE TAKEAWAYS0.99
- better1.00
- . Models are inherently stochastic;0.99
- prompts.1.00
- Control Plane1.00
- infrastructure must be strictly.1.00
- Blueprint1.00
- Poltoy Engines0.94
- Agest Execution0.93
- Reliability is no longer a model1.00
- It will be0.99
- .9.860.76
- problem—it is an infrastructure.0.99
- ececution0.94
- Agant1.00
- won by1.00
- Multidimensional observability is1.00
- mandatory, not optional.0.99
- better1.00
- ongite0.87
- State0.99
- • Agent control planes are the0.97
- systems.0.99
- Stute0.89
- required runtime layer for1.00
- production Al.1.00
- 2020.67
- 125°0.91
-
- Nishant Gupta1.00
- THANK YOU0.98
Transcript
115 cues· 1,025 words· 6,976 chars
- 0:03 Hey, everyone.
- 0:04 My name is Dishant Gupta.
- 0:06 I'm a software engineering tech leader at Meta, working on building the training and infrastructure.
- 0:12 And today, we're going to be talking about building deterministic infrastructure for non-deterministic AI agents.
- 0:19 So most of the conversations around AI over the last few years has been focused on models.
- 0:23 Bigger models, more parameters, better reasoning.
- 0:26 But as organizations move from chatbots to autonomous agents, a different problem emerges.
- 0:31 The challenge is no longer intelligence.
- 0:33 The challenge is reliability.
- 0:36 At Meta and across the industry, we are seeing agents move beyond answering questions and beginning to plan tool calls, coordinate workflows, and make decisions that affect production systems.
- 0:46 These systems are fundamentally probabilistic.
- 0:50 Infrastructure is not allowed to be.
- 0:52 Today, I want to discuss this topic in more detail.
- 0:58 The modern cloud infrastructure evolved around a set of assumptions.
- 1:01 Most of the requests are short-lived.
- 1:03 Services are deterministic, more or less.
- 1:06 Execution paths are known.
- 1:08 Failures are bounded.
- 1:10 However, autonomous AI agents violate nearly every one of those assumptions.
- 1:14 They are stateful.
- 1:14 They are long-running.
- 1:15 They make decisions dynamically.
- 1:17 They may execute different workflows for the same inputs.
- 1:21 This is what I call the great mismatch.
- 1:23 We're trying to run autonomous systems on infrastructure that was designed for deterministic workflows.
- 1:30 This is probably the most important mind shift.
- 1:33 Most AI demos showcase capability.
- 1:36 But can it solve a problem?
- 1:37 Can it use a tool?
- 1:38 Can it complete a workflow?
- 1:40 Production systems have a different objective.
- 1:43 Can it do it reliably?
- 1:45 Can it do it 10,000 times, 100,000 times, million times?
- 1:49 Can it recover from failures?
- 1:50 Can it operate safely?
- 1:51 Can it do it at an acceptable cost with an acceptable latency, with an acceptable outcome?
- 1:57 The majority of the engineering effort moves below the model layer into orchestration, monitoring, safety, evaluation, and recovery systems.
- 2:06 When people hear AI failures, they immediately think hallucinations.
- 2:10 In reality, hallucinations are often the least interesting for their mood.
- 2:14 What we see instead are infrastructure failures, recursive reasoning loops, over-floated logs, retry amplification, context corruption, memory poisoning, cost explosions.
- 2:25 The model makes mistakes, but however the infrastructure turns that mistake into an outage.
- 2:29 That's the real challenge.
- 2:33 So as this slide shows a pattern that distributed system engineers will probably recognize immediately.
- 2:38 An agent calls a tool incorrectly.
- 2:40 The tool returns an error.
- 2:42 Instead of recovering, the agent generates a slightly different but still invalid request.
- 2:48 The cycle repeats.
- 2:49 Each retry consumes more compute, reasoning depth increases, GPU consumption rises, eventually you get exponential resource growth.
- 2:57 What started as a minor API error became a compute incident.
- 3:01 This is why uncontrolled retries are one of the biggest risks in agentic systems.
loading