Videos Ib5t2RLtxvM
From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
Scene timeline
45 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 262
- whisperx 262
- chunks
- 36
- from 262 cues
- keyframes
- 32
- kept of 45 captured
- frames with text
- 32
- 673 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 5.2 MB
- word timings on 262 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:40 | 0s |
stt |
done | — | 2026-08-09 06:52 | 19s |
chunk |
done | — | 2026-08-09 06:53 | 0s |
text_embed |
done | — | 2026-08-10 19:39 | 1s |
keyframe |
done | — | 2026-08-09 06:53 | 2m 30s |
ocr |
done | — | 2026-08-09 06:55 | 15s |
frame_embed |
done | — | 2026-08-10 19:39 | 5s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.97
- Snorkel0.97
- World'sFair1.00
- The Frontier Al Data Lab1.00
- THE NEXT ERA OF AGENT EVALUATION1.00
- From agent traces to agent1.00
- PRESENTED BY1.00
- Microsoft1.00
- simulations1.00
- Building repeatable, production-like benchmarks for Al agents.1.00
- AlEngineer0.98
- Rustem Feyzkhanov1.00
- World'sFair0.99
- Al Platform Engineering, Snorkel Al1.00
- World'sFair1.00
- Engineering the future of Al0.97
-
- Snorkel o w NNome0.61
- AlEngineer0.98
- THE TAKEAWAYS1.00
- World's Fair0.96
- Three things to take away0.97
- PRESENTED BY1.00
- Every company needs its1.00
- It has to be as close to1.00
- It has to live in the agent1.00
- Microsoft1.00
- own benchmark1.00
- production as possible1.00
- lifecycle1.00
- The only reliable way to evaluate,1.00
- Mimic production surfaces — tools,0.98
- Populated by production traces and0.99
- improve, and release your agents.0.99
- schemas, APls, policies, and0.99
- gated by a Cl pipeline — so every0.98
- Public benchmarks are priors, not1.00
- workflows — without standing up0.98
- change has to pass it before it ships.1.00
- verdicts on your workload.1.00
- full production for every test.0.99
- © 2026 Snorkel Al0.98
- 121.00
- World's Fair0.98
- Engineering the future of Al0.99
-
- Snorkel efow Ne0.63
- AlEngineer0.98
- WHY OFFLINE EVALS MATTER0.99
- World's Fair0.98
- Why Snorkel Al is giving this talk1.00
- Snorkel Al builds expert-authored datasets, simulation0.99
- What we've learned building benchmarks at scale1.00
- environments, and evaluation pipelines for frontier Al0.99
- models and production agents. We treat benchmark0.99
- High-volume simulation runs1.00
- Secure sandbox execution0.99
- construction as an engineering discipline.1.00
- What we build1.00
- Expert-authored1.00
- Simulation1.00
- datasets1.00
- environments1.00
- Expert-authored environments1.00
- Dense rewards + verifier design0.99
- Evaluation pipelines0.99
- pipelines1.00
- Multi-layer quality1.00
- Secure sandbox1.00
- Frontier-lab1.00
- Production-like task1.00
- Release-relevant evaluation1.00
- execution1.00
- collaboration1.00
- construction1.00
- Millions of agent simulations per month1.00
- This talk shares the practical patterns we've learned building private benchmarks at scale.0.99
- © 2026 Snorkel Al0.96
- 131.00
- World's Fair0.99
- TRACK 5· JULY 1, 20260.92
- Evals1.00
-
- Snorkel Tefwow NOe0.65
- AlEngineer1.00
- WHY OFFLINE EVALS MATTER0.98
- World's Fair0.99
- A trace is evidence, not an experiment0.99
- Production users1.00
- Example production trace1.00
- user "I was charged twice for April. Fix it now."0.99
- Live agent run1.00
- agent1.00
- → billing_api.lookup_invoice(user_id)1.00
- Production traces1.00
- - invoice status: disputed0.99
- agent1.00
- What drifts between runs1.00
- → attempts refund workflow0.98
- user intent1.00
- DB / API state0.99
- final “I've adjusted your bill."0.97
- tool versions1.00
- context / history1.00
- observed failure0.99
- should have escalated the dispute to a specialist1.00
- wall-clock time1.00
- Great for finding real failures. Weak for comparing variants — it happened once, under one state of the world.0.99
- © 2026 Snorkel Al0.97
- 140.99
- World's Fair0.99
- TRACK 5·JULY 1,20260.97
- Evals1.00
-
- Snorkel tefotw NOe0.62
- AlEngineer0.98
- WHY OFFLINE EVALS MATTER0.99
- World's Fair0.99
- Offline simulation turns traces into repeatable experiments0.99
- Production traces0.99
- extract / cluster0.98
- Representative tasks0.98
- observability·Arize Ax0.95
- the failures worth fixing1.00
- Agent config A0.98
- model·prompt·tools-harness0.97
- Offline simulation0.98
- Comparable metrics1.00
- Agent config B1.00
- model·prompt·tools-harness0.96
- same task·same environment0.97
- benchmark1.00
- score1.00
- success rate1.00
- cost per success0.97
- same seed·same state·same0.97
- latency1.00
- Agent config C1.00
- tools1.00
- retries- trace quality0.98
- model·prompt·tools-harness0.97
- Offline and bounded — so the same task can run thousands of times in parallel, and every configuration is judged on the0.99
- same footing.1.00
- © 2026 Snorkel Al0.96
- 150.99
- World's Fair0.98
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- Snorkel e forow NODe0.61
- AlEngineer0.99
- WHY OFFLINE EVALS MATTER0.99
- World's Fair0.98
- Public benchmarks are priors, not verdicts0.99
- Public benchmark1.00
- Company benchmark1.00
- orientation - a prior0.93
- release-relevant measurement0.99
- SWE-bench1.00
- Your codebase + Cl1.00
- Can it fix GitHub issues?0.99
- Can it change your service safely?1.00
- WebArena/BrowseComp1.00
- Your tools + workflows0.98
- Can it browse and research?1.00
- Can it drive your CRM, DB, and APIs?0.99
- GPQA / MMLU0.96
- Your policies1.00
- Can it answer hard questions?0.98
- Can it follow your operating constraints?1.00
- Terminal-Bench1.00
- Your production-like sandbox1.00
- Can it solve terminal tasks?0.98
- Can it complete your end-to-end task?0.99
- Use public benchmarks to orient. Use your benchmark to ship — your release-relevant distribution, frozen into repeatable1.00
- tasks.1.00
- © 2026 Snorkel Al0.96
- 161.00
- World's Fair0.99
- TRACK 5·JULY 1,20260.98
- Evals1.00
-
- Snorkel oefoow NDoue0.61
- AlEngineer0.98
- WHY OFFLINE EVALS MATTER0.98
- World'sFair1.00
- Benchmark the whole agent stack1.00
- Agent behavior1.00
- Agent configuration — what you tune0.98
- Model + thinking level1.00
- configurable1.00
- evaluation layer0.98
- Observer /0.99
- Prompt / policy1.00
- configurable0.99
- Verifiers1.00
- Scoring1.00
- Harness / context / memory0.99
- configurable1.00
- Traces1.00
- Metrics1.00
- Skills / tools0.99
- configurable1.00
- Observes and scores1.00
- Runtime environment1.00
- the run - it is not0.97
- part of the stack.1.00
- MCP servers / APIs / DBs0.97
- Environment data1.00
- We benchmark configurations, not models.1.00
- © 2026 Snorkel Al0.97
- 181.00
- World's Fair0.99
- TRACK 5· JULY 1, 20260.93
- Evals1.00
-
- Snrel fwNe0.58
- AlEngineer0.98
- WHY OFFLINE EVALS MATTER0.99
- World's Fair0.99
- The simulation flywheel1.00
- Simulate1.00
- Measure1.00
- Diagnose1.00
- Generate data1.00
- Improve agent1.00
- PRESENTED BY1.00
- harder benchmark tasks1.00
- Microsoft1.00
- Make it work1.00
- Make it right0.99
- Make it fast0.99
- Short term1.00
- Mid term1.00
- Long term1.00
- Model selection1.00
- Regression testing0.97
- Cost tuning0.96
- Baseline pass rate0.96
- Release gates1.00
- Latency tuning0.99
- Trace debugging1.00
- Prompt / tool / harness iteration1.00
- RL / post-training0.97
- Evaluation becomes a data engine, not just a scorecard.0.99
- © 2026 Snorkel Al0.98
- 191.00
- World's Fair0.99
- TRACK 5· JULY 1,20260.94
- Evals1.00
-
- AlEngineer0.97
- World'sFair1.00
- 10.99
- Anatomy of a benchmark task1.00
- The pieces that make a task runnable, solvable, and scorable — including the1.00
- two pieces teams often forget: oracle solutions and metadata.1.00
- © 2026 Snorkel Al0.99
- 1100.85
- World'sFair0.99
- TRACK 5·JULY 1,20260.96
- Evals1.00
-
- Snrkel oo Ns ue0.51
- AlEngineer1.00
- ANATOMY OF A BENCHMARK TASK0.99
- World'sFair1.00
- What a simulation benchmark contains0.99
- Agent run evaluates candidate agent1.00
- Oracle validation validates task + env + verifiers0.99
- ↓1.00
- Task spec0.99
- Task spec0.99
- Environment copy A1.00
- Environment1.00
- Agent1.00
- files·DB·APls·tools0.97
- Oracle solution1.00
- copy B1.00
- ↓0.89
- same setup1.00
- ↓0.93
- Trace + final state + artifacts0.99
- ↓read-only1.00
- Oracle trace + final state + artifacts1.00
- ↓read-only0.99
- Verifiers read env + trace0.99
- Same verifiers read env + trace1.00
- ↓1.00
- Reward·metrics·labels0.99
- Must pass reliably1.00
- The agent is evaluated by verifiers. The oracle proves the task is solvable and the verifiers are correct.0.99
- © 2026 Snorkel Al0.97
- |110.93
- World's Fair0.99
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- Snorkel tefww NO0.66
- AlEngineer0.99
- ANATOMY OF A BENCHMARK TASK0.99
- World's Fair0.98
- Anatomy of a benchmark task0.98
- Agent-visible1.00
- simulation_task/1.00
- Task spec and the environment it runs in.1.00
- instruction.md1.00
- # agent-visible goal0.99
- environment/1.00
- # runtime surface0.99
- Verifier-only1.00
- Dockerfile1.00
- Oracle and checks — hidden from the agent.0.99
- services/1.00
- Metadata1.00
- fixtures/1.00
- Resource limits, tags, and stored baselines.1.00
- oracle/ # known-good solution0.98
- solve.sh0.96
- Harbor-style task abstraction1.00
- verifiers/ # verifier-only checks0.99
- A task is a packaged artifact — spec + environment + oracle + verifiers +0.99
- test_outputs.py1.00
- metadata — not just a prompt and a test.0.99
- trace_checks.py1.00
- portable·runnable·repeatable·comparable0.97
- metadata.toml # resources, tags0.99
- baselines/1.00
- model_results.json0.98
- © 2026 Snorkel Al0.99
- 1120.93
- World's Fair0.99
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- Snorkel Tefotoe Ne0.65
- AlEngineer0.98
- BENCHMARK TASK ENVIRONMENT1.00
- World's Fair0.96
- The environment is mini-production1.00
- AGENT-VISIBLEZONE1.00
- VERIFIER-ONLY ZONE0.98
- SIMULATION SANDBOX · network disabled / controlled0.98
- Oracle solution1.00
- Simulated user1.00
- Hidden checks1.00
- Tests0.95
- Agent1.00
- Walled off from the1.00
- Database1.00
- API service0.99
- Files / fixtures0.98
- MCP / tools0.99
- agent — this is what0.96
- prevents leakage and1.00
- Mocks preserve production contracts, not production data0.99
- cheating.1.00
- Small enough to run in parallel. Realistic enough to preserve production contracts.1.00
- © 2026 Snorkel Al0.98
- |140.89
- World's Fair0.99
- TRACK 5·JULY 1,20260.97
- Evals1.00
-
- Snorkel tefooe Ne0.65
- AlEngineer0.98
- BENCHMARK TASK ENVIRONMENT1.00
- World's Fair0.96
- Simulation environment patterns1.00
- m0.81
- DO0.85
- Multi-container1.00
- MCP servers1.00
- Simulated user0.99
- App + DB + APl as separate services.0.99
- Tools over the same protocol as prod.0.99
- An interactive counterpart that adapts.1.00
- 品0.93
- 0000.95
- □=0.59
- Skills1.00
- Multi-step1.00
- Production-like APls0.99
- Reusable, packaged capabilities.1.00
- Shared state evolving over time.0.98
- Same contract, mocked locally.0.99
- Choose patterns based on what your production agent actually touches.1.00
- © 2026 Snorkel Al0.98
- |150.94
- World's Fair0.96
- TRACK 5· JULY 1,20260.96
- Evals0.95
-
- Snorkel tefwtw NOe0.62
- AlEngineer0.98
- BENCHMARK TASK VERIFIERS1.00
- World's Fair0.98
- Don't just grade outputs — simulate worlds0.99
- Static grading1.00
- World simulation1.00
- PRESENTED BY1.00
- Agent1.00
- answer1.00
- Grader1.00
- → score0.96
- Agent1.00
- observation / response0.99
- action / tool call0.96
- Simulated1.00
- world1.00
- Microsoft1.00
- state changes — consequences0.96
- world = DB state - API responses · user replies - files / artifacts · tool outputs0.94
- Verifier1.00
- reads final state + trace + artifacts1.00
- One shot. A single score against a key.0.99
- Don't just grade the output. Simulate the world the agent operates in.0.99
- © 2026 Snorkel Al0.98
- 1170.86
- World's Fair0.98
- TRACK 5· JULY 1,20260.96
- Evals1.00
-
- Snorkel wewow Desue0.65
- AlEngineer1.00
- BENCHMARK TASK VERIFIERS1.00
- World's Fair0.99
- Verifiers define what success means1.00
- Signal1.00
- Programmatic /1.00
- LLM judge1.00
- Human / SME0.97
- deterministic0.99
- Final output1.00
- Primary1.00
- Optional1.00
- Optional1.00
- PRESENTED BY0.98
- Trace quality1.00
- Primary1.00
- Optional1.00
- Microsoft1.00
- Tool calls1.00
- Primary1.00
- Optional1.00
- Process / planning0.99
- Primary1.00
- Optional1.00
- Error recovery1.00
- Primary1.00
- Optional1.00
- Constraint following1.00
- Primary1.00
- Optional1.00
- Example- billing dispute: route to specialist · no refund tool call - policy passes - clear response0.97
- © 2026 Snorkel Al0.97
- |180.90
- World's Fair0.99
- TRACK 5· JULY 1,20260.96
- Evals1.00
Transcript
262 cues· 3,002 words· 16,983 chars
- 0:13 Thanks, everyone, for coming.
- 0:15 And I know this is the last session before lunch.
- 0:17 So thanks for staying here.
- 0:19 Let's make it smooth and with good vibes, just as Dad said.
- 0:22 And thanks, Dad, for introduction and for inviting me.
- 0:26 So my name is Rostam.
- 0:27 I'm leading AI platform team at Snorkel.
- 0:31 And today, I want to tell you how to turn agent traces into agent simulations and why this becomes the next stage for agent evaluations.
- 0:42 So three main things that I want you to take away from my talk is every company needs a benchmark.
- 0:49 It's the only way to reliably evaluate, release, and improve your agents.
- 0:54 It has to be as close to production as possible.
- 0:57 It has to mimic your real tools, real API services, policies, and workflows.
- 1:04 And finally, it has to be part of your agenda lifecycle.
- 1:09 It's not a static benchmark.
- 1:11 It's a constantly populated data set from your production traces.
- 1:17 So why is Snorkel AI giving this talk?
- 1:21 We are a data as a service company.
- 1:23 And we are basically selling benchmarks.
- 1:27 And we're producing benchmarks at scale.
- 1:29 And for us, benchmark construction is an engineering discipline.
- 1:33 We run millions of agent simulations per month.
- 1:35 And we learned how to do environment build and scale.
- 1:41 working using both agents and subject matter experts to build reliable benchmarks that are close to production and specific domains.
- 1:52 So a lot of the time when people say about agent evaluation, they're focused on traces.
- 1:56 And traces are very useful.
- 1:58 The usual traces, like you can see an example on the screen, basically shows, OK, here is the input prompt.
- 2:04 Here are the actions that agent took.
- 2:06 And here is the agent output.
- 2:09 And then evaluation can analyze it
- 2:11 and say, OK, was agent successful or not?
- 2:13 Was there any edge case?
- 2:16 So it is useful to find failures in production, but it's hard to test different variants.
- 2:22 You can run A-B testing, and that's one way of checking different agent configurations.
- 2:28 But it's hard to make sure that everything is repeatable, because you will get different database state, different tool versions, and so on.
- 2:35 So never fully compare apples to apples.
- 2:38 Offline simulation turns traces into repeatable experiments.
- 2:42 Now you take production traces, you construct tasks, and then you can run simulation benchmark with different agent configuration offline.
- 2:53 And you can compare agents using different metrics, not just success rate, but cost, latency, and retries.
- 3:00 And you can run those at parallel.
- 3:05 But you can ask, OK, but why do we need it?
- 3:07 We already have public benchmarks.
- 3:09 The challenge with public benchmarks is that usually they are focused on the very specific domains.
- 3:14 For example, is focused on fixing GitHub issues.
- 3:18 will focus on agent running in terminal.
- 3:21 And will focus on computer use agent.
- 3:24 In your case, you want your benchmark to be focused on your company's domain, both from perspective of use cases and in terms of tooling that your agent has.
- 3:34 whether it follows the policies that your company uses, and whether you get full production environment.
- 3:41 Basically, public benchmark is useful to orient and build your prior, but your private benchmark is useful to ship.
- 3:52 And a lot of the time, public benchmarks, they're specifically focused on pass rate.
- 3:57 Every time you see a new model release, you see performance like pass rate on different benchmarks.
loading
Chapters
- 0:00 Introduction: why Snorkel builds agent benchmarks
- 1:17 Benchmark construction and testing agents in production
- 3:08 The limits of public benchmarks
- 4:01 Why you need a private benchmark
- 4:52 Environments, tools, and evaluators
- 6:34 Anatomy of a simulation task
- 7:37 Task formats: instruction files and Oracle data
- 9:20 Multistep, long horizon simulations
- 10:54 Verifiers, LLM as a judge, and reward hacking
- 13:02 A CI pipeline for agents
- 15:19 Connecting traces, experiments, and benchmarks
- 16:51 Q&A