Videos u1yaOeEX4e8
Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase
Scene timeline
45 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 164
- whisperx 164
- chunks
- 33
- from 164 cues
- keyframes
- 24
- kept of 45 captured
- frames with text
- 24
- 643 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 5.3 MB
- word timings on 164 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 05:50 | 1m 26s |
stt |
done | — | 2026-08-10 05:51 | 18s |
chunk |
done | — | 2026-08-10 05:52 | 0s |
text_embed |
done | — | 2026-08-10 19:48 | 0s |
keyframe |
done | — | 2026-08-10 05:52 | 1m 52s |
ocr |
done | — | 2026-08-10 05:54 | 16s |
frame_embed |
done | — | 2026-08-10 19:48 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.97
- World'sFair0.98
- AI ENGINEER WORLD'S FAIR 20261.00
- GRAPHS TRACK1.00
- Learned Execution Graphs1.00
- PRESENTED BY1.00
- Microsoft1.00
- for Real-Time Anomaly Detection &1.00
- Drift Classification in APls0.99
- Turn every request's trace into an attributed execution DAG. Learn the distribution of normal graphs. Score0.99
- anomalies and classify drift — in near real time, at production throughput.1.00
- Ritvik Pandya1.00
- Engineering Lead · J.P. Morgan Chase - CIB Payments0.98
- Views are my own. Examples & figures generalized.1.00
- LEARNED EXECUTION GRAPHS0.97
- AIE WORLD'S FAIR 26 GRAPHS TRACK0.97
- 01/121.00
- Engineering the future of Al0.99
- World'sFair1.00
-
- AlEngineer0.99
- World'sFair1.00
- 011.00
- FRAMING1.00
- A different kind of graph1.00
- PRESENTED BY1.00
- WHAT THE TRACK MEANS BY"GRAPH"0.99
- THIS TALK1.00
- Persistent / property graph0.99
- Execution graph (a trace)1.00
- Microsoft1.00
- Knowledge graphs, GraphRAG, entity links0.98
- spans1.00
- One execution DAG per request, built from its1.00
- Lives in a database (Neo4j, ..)0.97
- Lives in the request path for milliseconds0.99
- Queried for what is related to what0.99
- Describes how this call actually ran1.00
- Edges are facts; the graph is the data1.00
- Edges are causality; the graph is behavior1.00
- Same word, opposite lifecycle. One is a store you query. The other is a signal you have ~40 ms to read before it's gone.0.99
- Why the graph, not just metrics? For "is it slow?" you don't need it — but where, cause vs. symptom, and what kind of change are structural,0.99
- and a number can't answer them.1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR '26 GRAPHS TRACK0.97
- 02 / 120.86
- TRACK 5·JULY 2,20260.96
- Graphs1.00
- World's Fair0.96
-
- AlEngineer0.99
- World'sFair1.00
- 021.00
- DEFINITION1.00
- From span tree to execution DAG1.00
- OpenTelemetry gives you a span tree — one parent per span. We enrich it into an execution DAG: add span links, async0.99
- producerconsumer edges, shared-resource and join causality. ("Notify" re-converges via a constructed join edge.)0.99
- fan-out1.00
- join edge1.00
- Fraud Score1.00
- ml-svc1.00
- Edge GW1.00
- AuthN/z0.94
- Orchestrator1.00
- Ledger Write0.99
- Notify1.00
- ingress1.00
- token-svc1.00
- payments-core1.00
- cockroachdb1.00
- response1.00
- FX Rate1.00
- ext-api1.00
- G = (V, E)0.96
- V = spans (operations)0.99
- E = parent-child + links + async + join1.00
- node feats: svc,1.00
- op, status, duration1.00
- edge feats: sync/async, lag0.99
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR26GRAPHS TRACK0.97
- 03-/-120.97
- TRACK 5· JULY 2,20260.97
- Graphs1.00
- World'sFair1.00
-
- AlEngineer0.99
- World's Fair0.97
- 031.00
- WHY A DAG1.00
- Acyclicity is load-bearing, not cosmetic0.99
- Topological order0.99
- Causal direction0.99
- Ordered propagation1.00
- A DAG has a valid linearization → you0.99
- Edges point parentchild. That lets the1.00
- Process spans in topological order -0.98
- can compute the critical path and1.00
- detector separate a root cause1.00
- each node aggregates its full upstream1.00
- attribute end-to-end latency to the0.99
- (upstream) from its symptoms1.00
- context in one causal pass, instead of0.99
- exact span that owns it.0.99
- (everything downstream of it).1.00
- iterating rounds and over-smoothing the1.00
- whole graph.1.00
- THE HONEST CAVEAT - RETRIES & LOOPS0.98
- Retries, polling, and sagas can introduce cycles. Restore the DAG before modeling: unroll repeats into distinct nodes (call#1, call#2) or1.00
- type the edge as ret ry/as ync. A retry storm then shows up as a structural feature — exactly the signal you want.0.98
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR'26 GRAPHS TRACK0.99
- 04 / 120.89
- TRACK 5·JULY 2,20260.97
- Graphs1.00
- World'sFair1.00
-
- AlEngineer0.98
- World'sFair1.00
- 041.00
- THE MODEL1.00
- "Learned" = model the distribution, not rules0.99
- Hand-written thresholds break on every deploy. Instead, learn per-node baselines and the graph's normal structure from1.00
- telemetry — no labels, nothing to train. Three statistical layers, cheapest first:0.99
- TIER1.00
- METHOD1.00
- LEARNS1.00
- CATCHES1.00
- GRANULAR1.00
- COST1.00
- ROLE1.00
- Tier 00.98
- graph-hash + timing baseline0.97
- topology + latency normal1.00
- rare / off-timing shape0.97
- √ by signature -us0.95
- gate1.00
- Tier 10.93
- deviation ratios + structural compare0.99
- per-node baselines + topology0.98
- change1.00
- slow node + structural1.00
- √ per node0.97
- <1ms1.00
- attribute1.00
- Tier 20.99
- per-client EMA + KL divergence0.99
- client graph profiles1.00
- behavioral drift + cause1.00
- √ per client0.94
- ~ms0.99
- classify1.00
- All three run in microseconds to milliseconds — no GPU, no inference, no black box. (A learned graph autoencoder is a natural1.00
- extension; for payments, statistics wins on latency and interpretability.)1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR'26 GRAPHS TRACK0.98
- 05 / 120.89
- TRACK 5·JULY 2,20260.99
- Graphs1.00
- World's Fair0.98
-
- AlEngineer0.97
- World's Fair0.98
- 051.00
- THE APPROACH1.00
- From graph to localized cause1.00
- No model to train — learn per-node baselines from normal traffic. A request's deviation from them is the score; the worst node1.00
- is the cause.1.00
- PRESENTED BY1.00
- Represent1.00
- Baseline1.00
- Deviate1.00
- Localize0.95
- Threshold1.00
- trace → DAG0.96
- per-node stats1.00
- obs / baseline1.00
- argmax ratio1.00
- FP/day budget0.97
- Microsoft1.00
- THE BASELINES - learned, not trained0.99
- TWO CHECKS ON ONE GRAPH0.99
- Per-node statistical baselines — latency (μ, σ, p99), dependency set,0.99
- execution frequency. Mined from normal traffic. No labels, no neural1.00
- deviation ratio → bottleneck node attribute1.00
- net.1.00
- baseline(n) = {latency u/o, deps, freq)0.96
- structure check →missing/ reordered /new structural0.97
- deviation(n) = observed / baseline - score0.97
- cause = argmax(deviation)- localize0.97
- or a mandatory one that's gone.0.99
- One finds a node that's slow; the other finds a node that shouldn't be there -0.98
- Threshold = a false-positives-per-day budget, not a latency guess. Baselines recalibrate as normal re-learns — gated by the drift classifier and0.99
- the deploy log, so an expected change doesn't page.1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR '26 · GRAPHS TRACK0.96
- 06 / 120.93
- TRACK 5· JULY 2, 20260.93
- Graphs1.00
- World's Fair0.99
-
- AlEngineer0.98
- World's Fair0.97
- 061.00
- ANOMALY DETECTION0.97
- Score the graph, then localize the span0.98
- The same payment DAG — now the external FX dependency degrades. Each node's observed latency over its baseline: the offending0.99
- span stands an order of magnitude above normal while the rest sit at ~1×.0.99
- PRESENTED BY0.98
- Fraud Score1.00
- Microsoft1.00
- ml-svc0.99
- Edge GW1.00
- AuthN/z0.94
- Orchestrator1.00
- Ledger Write1.00
- Notify1.00
- ingress1.00
- token-svc1.00
- payments-core1.00
- cockroachdb1.00
- response1.00
- FX Rate1.00
- ext-api1.00
- PER-NODE DEVIATION1.00
- (observed / baseline)1.00
- edge-gw 1.0×0.99
- auth 1.1×0.99
- orchestrator 1.0×1.00
- fraud 1.2×1.00
- ledger 1.1×0.99
- notify 1.0×1.00
- fx-rate 38×0.97
- fx-rate z-score > threshold0.97
- ANOMALY, localized to fx-rate1.00
- end-to-end p99 barely moved1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR'260.96
- GRAPHS TRACK1.00
- 07 / 120.93
- TRACK 5·JULY 2,20260.96
- Graphs1.00
- World's Fair0.99
-
- AlEngineer0.98
- World's Fair0.96
- 071.00
- EVIDENCE1.00
- One real experiment beats ten illustrations0.98
- ILLUSTRATIVE result shape - reproduce on the public benchmark, then replace with your measured run. No production0.99
- data needed.1.00
- SETUP1.00
- RESULTS1.00
- (report these)1.00
- Benchmark1.00
- OpenTelemetry + DeathStarBench0.99
- Detection 86% of injected incidents0.99
- Train on1.00
- ~1.9M normal traces- 7-day window0.98
- Deviation scoring0.4 ms median0.99
- Inject1.00
- FX latency· missing span retry storm0.97
- Localization top-1 83% of detected0.96
- False alerts 4 / service / day0.98
- post-deploy topology· traffic shift0.98
- CPU sat.1.00
- Structural + KL check < 2 ms (no GPU)0.96
- Illustrative: "On ~1.9M synthetic traces across 6 injected failure classes, per-node deviation flagged 86% of incidents at 0.4 ms median latency and localized the0.99
- responsible node in 83% of detected cases — no GPU, no inference step."0.98
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR'26 GRAPHS TRACK0.98
- 08 / 120.88
- TRACK 5·JULY 2,20260.99
- Graphs1.00
- World's Fair0.96
-
- AlEngineer0.97
- World'sFair1.00
- 08 · THE PIVOT0.93
- ANOMALY - a graph off the normal cloud1.00
- DRIFT - the normal cloud itself shifts1.00
- reference1.00
- current1.00
- instantaneous1.00
- a single request is wrong0.99
- over time "normal" has changed0.99
- An anomaly detector retrained on drift learns the drift as normal. You must detect the shift itself — and say what kind it is.0.99
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR '26 GRAPHS TRACK0.97
- 09/121.00
- TRACK 5· JULY 2,20260.96
- Graphs1.00
- World'sFair1.00
-
- AlEngineer0.97
- World'sFair1.00
- 08· THE PIVOT0.96
- ANOMALY - a graph off the normal cloud0.99
- DRIFT - the normal cloud itself shifts1.00
- reference1.00
- current1.00
- instantaneous1.00
- a single request is wrong0.98
- over time "normal" has changed0.99
- An anomaly detector retrained on drift learns the drift as normal. You must detect the shift itself — and say what kind it is.0.99
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR 26 GRAPHS TRACK0.97
- 09/121.00
- TRACK 5· JULY 2,20260.96
- AlEngine0.93
- Graphs1.00
- World'sFair1.00
-
- AlEngineer0.98
- World'sFair1.00
- 091.00
- DRIFT - ACTION0.93
- Name the drift — the class picks the fix0.98
- DRIFT CLASS1.00
- what changed1.00
- VERDICT1.00
- ACTION1.00
- Structural1.00
- Expected /0.96
- Deploy-correlated + health-gated → controlled0.98
- topology Δ-0.91
- new nodes / edges1.00
- Investigate1.00
- rebaseline. Nothing shipped → page.0.99
- timing Δ - shape stable0.95
- Performance1.00
- Mitigate1.00
- Scale, shed load, or break the slow circuit.1.00
- Covariate1.00
- Absorb1.00
- Traffic shift, not a fault — update the baseline, don't0.98
- input mix Δ - system fine0.98
- page.1.00
- Concept1.00
- same inputs - new graphs0.97
- Regression1.00
- Flag for rollback; trigger model retrain.0.98
- Detection without classification is just a louder alarm. Responding to a KIND — not a magnitude — is what makes automation safe.1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR '26 GRAPHS TRACK0.97
- 10 / 120.95
- TRACK 5· JULY 2, 20260.95
- AlEngine0.97
- Graphs1.00
- World's Fair0.99
-
- AlEngineer0.99
- World'sFair1.00
- 091.00
- DRIFT - ACTION0.95
- Name the drift — the class picks the fix0.98
- DRIFT CLASS1.00
- what changed1.00
- VERDICT1.00
- ACTION1.00
- PRESENTED BY1.00
- Structural1.00
- Expected /0.97
- Deploy-correlated + health-gated → controlled0.98
- Microsoft1.00
- topology0.99
- new nodes / edges1.00
- Investigate1.00
- rebaseline. Nothing shipped → page.0.98
- Performance1.00
- timing Δ - shape stable0.98
- Mitigate1.00
- Scale, shed load, or break the slow circuit.1.00
- Covariate1.00
- input mix Δ - system fine0.98
- Absorb1.00
- page.1.00
- Traffic shift, not a fault — update the baseline, don't0.98
- Concept1.00
- same inputs - new graphs0.98
- Regression1.00
- Flag for rollback; trigger model retrain.0.99
- Detection without classification is just a louder alarm. Responding to a KIND — not a magnitude — is what makes automation safe.1.00
- LEARNED EXECUTION GRAPHS1.00
- AIE WORLD'S FAIR 26 GRAPHS TRACK0.97
- 10 / 120.95
- TRACK 5·JULY 2,20260.96
- Graphs1.00
- World's Fair0.99
Transcript
164 cues· 2,388 words· 12,566 chars
- 0:13 Hi, thanks.
- 0:17 And hope everyone is out of the lunch coma and will survive this talk.
- 0:23 So yeah, myself, Ritvik, I lead the payments team in JPMorgan.
- 0:30 And today I'll be talking about execution graphs
- 0:35 how these graphs can help to detect any anomaly and drifts.
- 0:41 Also, how we can automate a few things around that.
- 0:45 And at the same time, if we can reduce the manual detection work and going on that side.
- 1:04 hear about graph, there are persistence graph and property graphs, which Neo4j and other products, we use for them.
- 1:16 We query those graphs.
- 1:19 and get the answers out of it.
- 1:21 What I'm talking about today is execution graph.
- 1:24 It's short-lived graph.
- 1:27 And idea here is holistically try to identify how the request processing happens and if there is any deviation on that and how to detect that and how to fix that.
- 1:41 So here is a simple example.
- 1:44 Say we have set of applications.
- 1:48 You have one edge layer.
- 1:50 the first layer where a request comes in.
- 1:53 And then you have some gateways.
- 1:56 If gate is there, you have ingress layer on top of it.
- 2:00 Then authentication authorization happens.
- 2:03 After that, there is some orchestration layer and a few other systems which could be called in parallel.
- 2:10 Once everything is done, you are notifying your client that what's the update on that request, right?
- 2:18 Here, the idea is representing this overall request processing as DAG.
- 2:25 And using DAG simplifies most of the things here.
- 2:30 One, now you know that in what order service execution will be happening, right?
- 2:37 So that's one of the things.
- 2:38 The other thing is you know the context that at what node, what context will be there.
- 2:46 and what will be passed to the next node.
- 2:49 In that way, it will be very ordered and simplified, simply can be represented.
- 2:57 There are a few other use cases could be there in terms of retries and the loops, et cetera.
- 3:06 The idea here is every loop to put in the graph as a separate entity itself.
- 3:15 In that way, it could be tracked easily.
- 3:23 How we can make this system more reliable at the same time not using most of the resources, right?
- 3:33 So in the tier one check, or it's your first check, it's like,
- 3:38 going to airport and just boarding passes, someone is looking at the boarding pass and let you go.
- 3:46 So now if you know the baseline of your request execution end to end, if everything looks good, you don't need to go to the tier two or next tier of check, right?
- 3:59 If you find that there is some delay, so now you need to check that what changed here.
- 4:07 the drift here could be because of the structural change.
- 4:10 So if any new node or new step added, which you are not aware of, that could be one of the thing, or one of the step which is removed, that could be another reason, right?
- 4:23 Once you know about that, then further analysis could be done in terms of KL deviations or divergence.
- 4:36 Exponential MA.
- 4:37 So in simpler terms, if you know that client A's request is taking this much time normally and client B's request might take more time than the client A because of, say, one client is local to you and one client is, you know, the request is coming from outside and there are a few more checks needs to be done.
- 5:01 So in that case, the baseline will change
- 5:05 client to client.
- 5:06 And now you know what your threshold is and how you can reduce the noise of such alerts.
- 5:17 So here, the idea is very simple.
- 5:22 First, you represent the entire request processing as DAG.
- 5:29 You come up with the baseline.
- 5:31 You find out the deviation.
- 5:33 And then you try to find out where exactly the issue is.
loading
Chapters
- 0:00 Execution graphs for anomaly and drift detection
- 1:07 What a short lived execution graph is
- 3:28 Tiered checks and per client baselines
- 5:23 The method: baseline, deviation, localize, act
- 6:16 Localizing a slow node, and how the system is trained
- 7:33 Anomaly versus drift
- 8:55 The three kinds of drift: structural, volume, covariate
- 12:46 The pipeline: from telemetry to gradual rollout
- 13:54 Hot path versus recon, and worked examples
- 15:21 Tuning it: delayed events, sampling, cold starts
- 17:09 Results and lessons