read-only demo

Videos Ib5t2RLtxvM

From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

index_state ready data_status ok

AI Engineer· published 2026-07-25· 0:20:23· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:42, 1 of 1 keyframes kept
  5. Shot 4, 0:42 to 1:17, 1 of 1 keyframes kept
  6. Shot 5, 1:17 to 1:51, 1 of 1 keyframes kept
  7. Shot 6, 1:51 to 2:37, 1 of 1 keyframes kept
  8. Shot 7, 2:37 to 3:04, 1 of 1 keyframes kept
  9. Shot 8, 3:04 to 3:51, 1 of 1 keyframes kept
  10. Shot 9, 3:51 to 4:32, 0 of 1 keyframes kept
  11. Shot 10, 4:32 to 5:11, 1 of 1 keyframes kept
  12. Shot 11, 5:11 to 5:47, 1 of 1 keyframes kept
  13. Shot 12, 5:47 to 6:23, 0 of 1 keyframes kept
  14. Shot 13, 6:23 to 6:32, 1 of 1 keyframes kept
  15. Shot 14, 6:32 to 7:03, 1 of 1 keyframes kept
  16. Shot 15, 7:03 to 7:35, 0 of 1 keyframes kept
  17. Shot 16, 7:35 to 8:24, 1 of 1 keyframes kept
  18. Shot 17, 8:24 to 8:31, 0 of 1 keyframes kept
  19. Shot 18, 8:31 to 9:17, 1 of 1 keyframes kept
  20. Shot 19, 9:17 to 9:49, 1 of 1 keyframes kept
  21. Shot 20, 9:49 to 10:20, 0 of 1 keyframes kept
  22. Shot 21, 10:20 to 10:24, 0 of 1 keyframes kept
  23. Shot 22, 10:24 to 11:00, 1 of 1 keyframes kept
  24. Shot 23, 11:00 to 11:27, 1 of 1 keyframes kept
  25. Shot 24, 11:27 to 11:53, 0 of 1 keyframes kept
  26. Shot 25, 11:53 to 12:05, 0 of 1 keyframes kept
  27. Shot 26, 12:05 to 12:35, 0 of 1 keyframes kept
  28. Shot 27, 12:35 to 13:05, 0 of 1 keyframes kept
  29. Shot 28, 13:05 to 13:30, 1 of 1 keyframes kept
  30. Shot 29, 13:30 to 13:56, 0 of 1 keyframes kept
  31. Shot 30, 13:56 to 14:24, 1 of 1 keyframes kept
  32. Shot 31, 14:24 to 15:09, 1 of 1 keyframes kept
  33. Shot 32, 15:09 to 15:38, 1 of 1 keyframes kept
  34. Shot 33, 15:38 to 16:08, 0 of 1 keyframes kept
  35. Shot 34, 16:08 to 16:28, 1 of 1 keyframes kept
  36. Shot 35, 16:28 to 16:45, 1 of 1 keyframes kept
  37. Shot 36, 16:45 to 17:10, 1 of 1 keyframes kept
  38. Shot 37, 17:10 to 17:35, 1 of 1 keyframes kept
  39. Shot 38, 17:35 to 18:00, 1 of 1 keyframes kept
  40. Shot 39, 18:00 to 18:25, 1 of 1 keyframes kept
  41. Shot 40, 18:25 to 18:51, 1 of 1 keyframes kept
  42. Shot 41, 18:51 to 19:16, 1 of 1 keyframes kept
  43. Shot 42, 19:16 to 19:41, 1 of 1 keyframes kept
  44. Shot 43, 19:41 to 20:06, 1 of 1 keyframes kept
  45. Shot 44, 20:06 to 20:23, 0 of 1 keyframes kept

45 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
262
whisperx 262
chunks
36
from 262 cues
keyframes
32
kept of 45 captured
frames with text
32
673 lines read
chapters
12
from the source metadata
keyframe bytes
5.2 MB
word timings on 262 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:40 0s
stt done 2026-08-09 06:52 19s
chunk done 2026-08-09 06:53 0s
text_embed done 2026-08-10 19:39 1s
keyframe done 2026-08-09 06:53 2m 30s
ocr done 2026-08-09 06:55 15s
frame_embed done 2026-08-10 19:39 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:38 #3 done16 line(s)

    shot 3·sharpness 2005.9

    1. AlEngineer0.97
    2. Snorkel0.97
    3. World'sFair1.00
    4. The Frontier Al Data Lab1.00
    5. THE NEXT ERA OF AGENT EVALUATION1.00
    6. From agent traces to agent1.00
    7. PRESENTED BY1.00
    8. Microsoft1.00
    9. simulations1.00
    10. Building repeatable, production-like benchmarks for Al agents.1.00
    11. AlEngineer0.98
    12. Rustem Feyzkhanov1.00
    13. World'sFair0.99
    14. Al Platform Engineering, Snorkel Al1.00
    15. World'sFair1.00
    16. Engineering the future of Al0.97
  • 0:46 #4 done28 line(s)

    shot 4·sharpness 3160.0

    1. Snorkel o w NNome0.61
    2. AlEngineer0.98
    3. THE TAKEAWAYS1.00
    4. World's Fair0.96
    5. Three things to take away0.97
    6. PRESENTED BY1.00
    7. Every company needs its1.00
    8. It has to be as close to1.00
    9. It has to live in the agent1.00
    10. Microsoft1.00
    11. own benchmark1.00
    12. production as possible1.00
    13. lifecycle1.00
    14. The only reliable way to evaluate,1.00
    15. Mimic production surfaces — tools,0.98
    16. Populated by production traces and0.99
    17. improve, and release your agents.0.99
    18. schemas, APls, policies, and0.99
    19. gated by a Cl pipeline — so every0.98
    20. Public benchmarks are priors, not1.00
    21. workflows — without standing up0.98
    22. change has to pass it before it ships.1.00
    23. verdicts on your workload.1.00
    24. full production for every test.0.99
    25. © 2026 Snorkel Al0.98
    26. 121.00
    27. World's Fair0.98
    28. Engineering the future of Al0.99
  • 1:47 #5 done36 line(s)

    shot 5·sharpness 3189.3

    1. Snorkel efow Ne0.63
    2. AlEngineer0.98
    3. WHY OFFLINE EVALS MATTER0.99
    4. World's Fair0.98
    5. Why Snorkel Al is giving this talk1.00
    6. Snorkel Al builds expert-authored datasets, simulation0.99
    7. What we've learned building benchmarks at scale1.00
    8. environments, and evaluation pipelines for frontier Al0.99
    9. models and production agents. We treat benchmark0.99
    10. High-volume simulation runs1.00
    11. Secure sandbox execution0.99
    12. construction as an engineering discipline.1.00
    13. What we build1.00
    14. Expert-authored1.00
    15. Simulation1.00
    16. datasets1.00
    17. environments1.00
    18. Expert-authored environments1.00
    19. Dense rewards + verifier design0.99
    20. Evaluation pipelines0.99
    21. pipelines1.00
    22. Multi-layer quality1.00
    23. Secure sandbox1.00
    24. Frontier-lab1.00
    25. Production-like task1.00
    26. Release-relevant evaluation1.00
    27. execution1.00
    28. collaboration1.00
    29. construction1.00
    30. Millions of agent simulations per month1.00
    31. This talk shares the practical patterns we've learned building private benchmarks at scale.0.99
    32. © 2026 Snorkel Al0.96
    33. 131.00
    34. World's Fair0.99
    35. TRACK 5· JULY 1, 20260.92
    36. Evals1.00
  • 2:23 #6 done30 line(s)

    shot 6·sharpness 2426.2

    1. Snorkel Tefwow NOe0.65
    2. AlEngineer1.00
    3. WHY OFFLINE EVALS MATTER0.98
    4. World's Fair0.99
    5. A trace is evidence, not an experiment0.99
    6. Production users1.00
    7. Example production trace1.00
    8. user "I was charged twice for April. Fix it now."0.99
    9. Live agent run1.00
    10. agent1.00
    11. → billing_api.lookup_invoice(user_id)1.00
    12. Production traces1.00
    13. - invoice status: disputed0.99
    14. agent1.00
    15. What drifts between runs1.00
    16. → attempts refund workflow0.98
    17. user intent1.00
    18. DB / API state0.99
    19. final “I've adjusted your bill."0.97
    20. tool versions1.00
    21. context / history1.00
    22. observed failure0.99
    23. should have escalated the dispute to a specialist1.00
    24. wall-clock time1.00
    25. Great for finding real failures. Weak for comparing variants — it happened once, under one state of the world.0.99
    26. © 2026 Snorkel Al0.97
    27. 140.99
    28. World's Fair0.99
    29. TRACK 5·JULY 1,20260.97
    30. Evals1.00
  • 3:01 #7 done34 line(s)

    shot 7·sharpness 2809.0

    1. Snorkel tefotw NOe0.62
    2. AlEngineer0.98
    3. WHY OFFLINE EVALS MATTER0.99
    4. World's Fair0.99
    5. Offline simulation turns traces into repeatable experiments0.99
    6. Production traces0.99
    7. extract / cluster0.98
    8. Representative tasks0.98
    9. observability·Arize Ax0.95
    10. the failures worth fixing1.00
    11. Agent config A0.98
    12. model·prompt·tools-harness0.97
    13. Offline simulation0.98
    14. Comparable metrics1.00
    15. Agent config B1.00
    16. model·prompt·tools-harness0.96
    17. same task·same environment0.97
    18. benchmark1.00
    19. score1.00
    20. success rate1.00
    21. cost per success0.97
    22. same seed·same state·same0.97
    23. latency1.00
    24. Agent config C1.00
    25. tools1.00
    26. retries- trace quality0.98
    27. model·prompt·tools-harness0.97
    28. Offline and bounded — so the same task can run thousands of times in parallel, and every configuration is judged on the0.99
    29. same footing.1.00
    30. © 2026 Snorkel Al0.96
    31. 150.99
    32. World's Fair0.98
    33. TRACK 5· JULY 1, 20260.95
    34. Evals1.00
  • 3:14 #8 done32 line(s)

    shot 8·sharpness 2591.0

    1. Snorkel e forow NODe0.61
    2. AlEngineer0.99
    3. WHY OFFLINE EVALS MATTER0.99
    4. World's Fair0.98
    5. Public benchmarks are priors, not verdicts0.99
    6. Public benchmark1.00
    7. Company benchmark1.00
    8. orientation - a prior0.93
    9. release-relevant measurement0.99
    10. SWE-bench1.00
    11. Your codebase + Cl1.00
    12. Can it fix GitHub issues?0.99
    13. Can it change your service safely?1.00
    14. WebArena/BrowseComp1.00
    15. Your tools + workflows0.98
    16. Can it browse and research?1.00
    17. Can it drive your CRM, DB, and APIs?0.99
    18. GPQA / MMLU0.96
    19. Your policies1.00
    20. Can it answer hard questions?0.98
    21. Can it follow your operating constraints?1.00
    22. Terminal-Bench1.00
    23. Your production-like sandbox1.00
    24. Can it solve terminal tasks?0.98
    25. Can it complete your end-to-end task?0.99
    26. Use public benchmarks to orient. Use your benchmark to ship — your release-relevant distribution, frozen into repeatable1.00
    27. tasks.1.00
    28. © 2026 Snorkel Al0.96
    29. 161.00
    30. World's Fair0.99
    31. TRACK 5·JULY 1,20260.98
    32. Evals1.00
  • 4:08 #9 skipped

    shot 9·duplicate of #8

  • 5:03 #10 done33 line(s)

    shot 10·sharpness 2197.3

    1. Snorkel oefoow NDoue0.61
    2. AlEngineer0.98
    3. WHY OFFLINE EVALS MATTER0.98
    4. World'sFair1.00
    5. Benchmark the whole agent stack1.00
    6. Agent behavior1.00
    7. Agent configuration — what you tune0.98
    8. Model + thinking level1.00
    9. configurable1.00
    10. evaluation layer0.98
    11. Observer /0.99
    12. Prompt / policy1.00
    13. configurable0.99
    14. Verifiers1.00
    15. Scoring1.00
    16. Harness / context / memory0.99
    17. configurable1.00
    18. Traces1.00
    19. Metrics1.00
    20. Skills / tools0.99
    21. configurable1.00
    22. Observes and scores1.00
    23. Runtime environment1.00
    24. the run - it is not0.97
    25. part of the stack.1.00
    26. MCP servers / APIs / DBs0.97
    27. Environment data1.00
    28. We benchmark configurations, not models.1.00
    29. © 2026 Snorkel Al0.97
    30. 181.00
    31. World's Fair0.99
    32. TRACK 5· JULY 1, 20260.93
    33. Evals1.00
  • 5:36 #11 done34 line(s)

    shot 11·sharpness 2229.2

    1. Snrel fwNe0.58
    2. AlEngineer0.98
    3. WHY OFFLINE EVALS MATTER0.99
    4. World's Fair0.99
    5. The simulation flywheel1.00
    6. Simulate1.00
    7. Measure1.00
    8. Diagnose1.00
    9. Generate data1.00
    10. Improve agent1.00
    11. PRESENTED BY1.00
    12. harder benchmark tasks1.00
    13. Microsoft1.00
    14. Make it work1.00
    15. Make it right0.99
    16. Make it fast0.99
    17. Short term1.00
    18. Mid term1.00
    19. Long term1.00
    20. Model selection1.00
    21. Regression testing0.97
    22. Cost tuning0.96
    23. Baseline pass rate0.96
    24. Release gates1.00
    25. Latency tuning0.99
    26. Trace debugging1.00
    27. Prompt / tool / harness iteration1.00
    28. RL / post-training0.97
    29. Evaluation becomes a data engine, not just a scorecard.0.99
    30. © 2026 Snorkel Al0.98
    31. 191.00
    32. World's Fair0.99
    33. TRACK 5· JULY 1,20260.94
    34. Evals1.00
  • 6:12 #12 skipped

    shot 12·duplicate of #11

  • 6:28 #13 done11 line(s)

    shot 13·sharpness 1091.9

    1. AlEngineer0.97
    2. World'sFair1.00
    3. 10.99
    4. Anatomy of a benchmark task1.00
    5. The pieces that make a task runnable, solvable, and scorable — including the1.00
    6. two pieces teams often forget: oracle solutions and metadata.1.00
    7. © 2026 Snorkel Al0.99
    8. 1100.85
    9. World'sFair0.99
    10. TRACK 5·JULY 1,20260.96
    11. Evals1.00
  • 7:00 #14 done34 line(s)

    shot 14·sharpness 2335.2

    1. Snrkel oo Ns ue0.51
    2. AlEngineer1.00
    3. ANATOMY OF A BENCHMARK TASK0.99
    4. World'sFair1.00
    5. What a simulation benchmark contains0.99
    6. Agent run evaluates candidate agent1.00
    7. Oracle validation validates task + env + verifiers0.99
    8. 1.00
    9. Task spec0.99
    10. Task spec0.99
    11. Environment copy A1.00
    12. Environment1.00
    13. Agent1.00
    14. files·DB·APls·tools0.97
    15. Oracle solution1.00
    16. copy B1.00
    17. 0.89
    18. same setup1.00
    19. 0.93
    20. Trace + final state + artifacts0.99
    21. ↓read-only1.00
    22. Oracle trace + final state + artifacts1.00
    23. ↓read-only0.99
    24. Verifiers read env + trace0.99
    25. Same verifiers read env + trace1.00
    26. 1.00
    27. Reward·metrics·labels0.99
    28. Must pass reliably1.00
    29. The agent is evaluated by verifiers. The oracle proves the task is solvable and the verifiers are correct.0.99
    30. © 2026 Snorkel Al0.97
    31. |110.93
    32. World's Fair0.99
    33. TRACK 5· JULY 1, 20260.95
    34. Evals1.00
  • 7:31 #15 skipped

    shot 15·duplicate of #14

  • 8:13 #16 done36 line(s)

    shot 16·sharpness 2314.4

    1. Snorkel tefww NO0.66
    2. AlEngineer0.99
    3. ANATOMY OF A BENCHMARK TASK0.99
    4. World's Fair0.98
    5. Anatomy of a benchmark task0.98
    6. Agent-visible1.00
    7. simulation_task/1.00
    8. Task spec and the environment it runs in.1.00
    9. instruction.md1.00
    10. # agent-visible goal0.99
    11. environment/1.00
    12. # runtime surface0.99
    13. Verifier-only1.00
    14. Dockerfile1.00
    15. Oracle and checks — hidden from the agent.0.99
    16. services/1.00
    17. Metadata1.00
    18. fixtures/1.00
    19. Resource limits, tags, and stored baselines.1.00
    20. oracle/ # known-good solution0.98
    21. solve.sh0.96
    22. Harbor-style task abstraction1.00
    23. verifiers/ # verifier-only checks0.99
    24. A task is a packaged artifact — spec + environment + oracle + verifiers +0.99
    25. test_outputs.py1.00
    26. metadata — not just a prompt and a test.0.99
    27. trace_checks.py1.00
    28. portable·runnable·repeatable·comparable0.97
    29. metadata.toml # resources, tags0.99
    30. baselines/1.00
    31. model_results.json0.98
    32. © 2026 Snorkel Al0.99
    33. 1120.93
    34. World's Fair0.99
    35. TRACK 5· JULY 1, 20260.95
    36. Evals1.00
  • 8:30 #17 skipped

    shot 17·duplicate of #13

  • 8:54 #18 done28 line(s)

    shot 18·sharpness 2044.0

    1. Snorkel Tefotoe Ne0.65
    2. AlEngineer0.98
    3. BENCHMARK TASK ENVIRONMENT1.00
    4. World's Fair0.96
    5. The environment is mini-production1.00
    6. AGENT-VISIBLEZONE1.00
    7. VERIFIER-ONLY ZONE0.98
    8. SIMULATION SANDBOX · network disabled / controlled0.98
    9. Oracle solution1.00
    10. Simulated user1.00
    11. Hidden checks1.00
    12. Tests0.95
    13. Agent1.00
    14. Walled off from the1.00
    15. Database1.00
    16. API service0.99
    17. Files / fixtures0.98
    18. MCP / tools0.99
    19. agent — this is what0.96
    20. prevents leakage and1.00
    21. Mocks preserve production contracts, not production data0.99
    22. cheating.1.00
    23. Small enough to run in parallel. Realistic enough to preserve production contracts.1.00
    24. © 2026 Snorkel Al0.98
    25. |140.89
    26. World's Fair0.99
    27. TRACK 5·JULY 1,20260.97
    28. Evals1.00
  • 9:33 #19 done28 line(s)

    shot 19·sharpness 2141.2

    1. Snorkel tefooe Ne0.65
    2. AlEngineer0.98
    3. BENCHMARK TASK ENVIRONMENT1.00
    4. World's Fair0.96
    5. Simulation environment patterns1.00
    6. m0.81
    7. DO0.85
    8. Multi-container1.00
    9. MCP servers1.00
    10. Simulated user0.99
    11. App + DB + APl as separate services.0.99
    12. Tools over the same protocol as prod.0.99
    13. An interactive counterpart that adapts.1.00
    14. 0.93
    15. 0000.95
    16. □=0.59
    17. Skills1.00
    18. Multi-step1.00
    19. Production-like APls0.99
    20. Reusable, packaged capabilities.1.00
    21. Shared state evolving over time.0.98
    22. Same contract, mocked locally.0.99
    23. Choose patterns based on what your production agent actually touches.1.00
    24. © 2026 Snorkel Al0.98
    25. |150.94
    26. World's Fair0.96
    27. TRACK 5· JULY 1,20260.96
    28. Evals0.95
  • 10:01 #20 skipped

    shot 20·duplicate of #19

  • 10:22 #21 skipped

    shot 21·duplicate of #13

  • 10:35 #22 done29 line(s)

    shot 22·sharpness 2161.0

    1. Snorkel tefwtw NOe0.62
    2. AlEngineer0.98
    3. BENCHMARK TASK VERIFIERS1.00
    4. World's Fair0.98
    5. Don't just grade outputs — simulate worlds0.99
    6. Static grading1.00
    7. World simulation1.00
    8. PRESENTED BY1.00
    9. Agent1.00
    10. answer1.00
    11. Grader1.00
    12. → score0.96
    13. Agent1.00
    14. observation / response0.99
    15. action / tool call0.96
    16. Simulated1.00
    17. world1.00
    18. Microsoft1.00
    19. state changes — consequences0.96
    20. world = DB state - API responses · user replies - files / artifacts · tool outputs0.94
    21. Verifier1.00
    22. reads final state + trace + artifacts1.00
    23. One shot. A single score against a key.0.99
    24. Don't just grade the output. Simulate the world the agent operates in.0.99
    25. © 2026 Snorkel Al0.98
    26. 1170.86
    27. World's Fair0.98
    28. TRACK 5· JULY 1,20260.96
    29. Evals1.00
  • 11:11 #23 done37 line(s)

    shot 23·sharpness 2375.5

    1. Snorkel wewow Desue0.65
    2. AlEngineer1.00
    3. BENCHMARK TASK VERIFIERS1.00
    4. World's Fair0.99
    5. Verifiers define what success means1.00
    6. Signal1.00
    7. Programmatic /1.00
    8. LLM judge1.00
    9. Human / SME0.97
    10. deterministic0.99
    11. Final output1.00
    12. Primary1.00
    13. Optional1.00
    14. Optional1.00
    15. PRESENTED BY0.98
    16. Trace quality1.00
    17. Primary1.00
    18. Optional1.00
    19. Microsoft1.00
    20. Tool calls1.00
    21. Primary1.00
    22. Optional1.00
    23. Process / planning0.99
    24. Primary1.00
    25. Optional1.00
    26. Error recovery1.00
    27. Primary1.00
    28. Optional1.00
    29. Constraint following1.00
    30. Primary1.00
    31. Optional1.00
    32. Example- billing dispute: route to specialist · no refund tool call - policy passes - clear response0.97
    33. © 2026 Snorkel Al0.97
    34. |180.90
    35. World's Fair0.99
    36. TRACK 5· JULY 1,20260.96
    37. Evals1.00

Transcript

262 cues· 3,002 words· 16,983 chars

  1. 0:13 Thanks, everyone, for coming.
  2. 0:15 And I know this is the last session before lunch.
  3. 0:17 So thanks for staying here.
  4. 0:19 Let's make it smooth and with good vibes, just as Dad said.
  5. 0:22 And thanks, Dad, for introduction and for inviting me.
  6. 0:26 So my name is Rostam.
  7. 0:27 I'm leading AI platform team at Snorkel.
  8. 0:31 And today, I want to tell you how to turn agent traces into agent simulations and why this becomes the next stage for agent evaluations.
  9. 0:42 So three main things that I want you to take away from my talk is every company needs a benchmark.
  10. 0:49 It's the only way to reliably evaluate, release, and improve your agents.
  11. 0:54 It has to be as close to production as possible.
  12. 0:57 It has to mimic your real tools, real API services, policies, and workflows.
  13. 1:04 And finally, it has to be part of your agenda lifecycle.
  14. 1:09 It's not a static benchmark.
  15. 1:11 It's a constantly populated data set from your production traces.
  16. 1:17 So why is Snorkel AI giving this talk?
  17. 1:21 We are a data as a service company.
  18. 1:23 And we are basically selling benchmarks.
  19. 1:27 And we're producing benchmarks at scale.
  20. 1:29 And for us, benchmark construction is an engineering discipline.
  21. 1:33 We run millions of agent simulations per month.
  22. 1:35 And we learned how to do environment build and scale.
  23. 1:41 working using both agents and subject matter experts to build reliable benchmarks that are close to production and specific domains.
  24. 1:52 So a lot of the time when people say about agent evaluation, they're focused on traces.
  25. 1:56 And traces are very useful.
  26. 1:58 The usual traces, like you can see an example on the screen, basically shows, OK, here is the input prompt.
  27. 2:04 Here are the actions that agent took.
  28. 2:06 And here is the agent output.
  29. 2:09 And then evaluation can analyze it
  30. 2:11 and say, OK, was agent successful or not?
  31. 2:13 Was there any edge case?
  32. 2:16 So it is useful to find failures in production, but it's hard to test different variants.
  33. 2:22 You can run A-B testing, and that's one way of checking different agent configurations.
  34. 2:28 But it's hard to make sure that everything is repeatable, because you will get different database state, different tool versions, and so on.
  35. 2:35 So never fully compare apples to apples.
  36. 2:38 Offline simulation turns traces into repeatable experiments.
  37. 2:42 Now you take production traces, you construct tasks, and then you can run simulation benchmark with different agent configuration offline.
  38. 2:53 And you can compare agents using different metrics, not just success rate, but cost, latency, and retries.
  39. 3:00 And you can run those at parallel.
  40. 3:05 But you can ask, OK, but why do we need it?
  41. 3:07 We already have public benchmarks.
  42. 3:09 The challenge with public benchmarks is that usually they are focused on the very specific domains.
  43. 3:14 For example, is focused on fixing GitHub issues.
  44. 3:18 will focus on agent running in terminal.
  45. 3:21 And will focus on computer use agent.
  46. 3:24 In your case, you want your benchmark to be focused on your company's domain, both from perspective of use cases and in terms of tooling that your agent has.
  47. 3:34 whether it follows the policies that your company uses, and whether you get full production environment.
  48. 3:41 Basically, public benchmark is useful to orient and build your prior, but your private benchmark is useful to ship.
  49. 3:52 And a lot of the time, public benchmarks, they're specifically focused on pass rate.
  50. 3:57 Every time you see a new model release, you see performance like pass rate on different benchmarks.

Chapters

  1. 0:00 Introduction: why Snorkel builds agent benchmarks
  2. 1:17 Benchmark construction and testing agents in production
  3. 3:08 The limits of public benchmarks
  4. 4:01 Why you need a private benchmark
  5. 4:52 Environments, tools, and evaluators
  6. 6:34 Anatomy of a simulation task
  7. 7:37 Task formats: instruction files and Oracle data
  8. 9:20 Multistep, long horizon simulations
  9. 10:54 Verifiers, LLM as a judge, and reward hacking
  10. 13:02 A CI pipeline for agents
  11. 15:19 Connecting traces, experiments, and benchmarks
  12. 16:51 Q&A

Open at this second