Videos q2JrUKBMf0w
The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
Scene timeline
17 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 69
- whisperx 69
- chunks
- 10
- from 69 cues
- keyframes
- 15
- kept of 17 captured
- frames with text
- 15
- 460 lines read
- chapters
- 4
- from the source metadata
- keyframe bytes
- 2.2 MB
- word timings on 69 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:53 | 0s |
stt |
done | — | 2026-08-09 07:19 | 7s |
chunk |
done | — | 2026-08-09 07:19 | 0s |
text_embed |
done | — | 2026-08-10 19:40 | 0s |
keyframe |
done | — | 2026-08-09 07:19 | 49s |
ocr |
done | — | 2026-08-09 07:20 | 8s |
frame_embed |
done | — | 2026-08-10 19:40 | 3s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab1.00
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.97
- OpenAl0.94
- 70.60
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust1.00
- bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- together.ai0.98
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEnginee0.98
- Oper0.99
- ir1.00
- World's1.00
- AlEn0.77
- ai1.00
- Wor'0.98
- Akan1.00
- arize1.00
- AIEnginee0.95
- E口0.52
- World's1.00
-
- AlEngineer0.98
- penAl0.95
- Id's Fair0.98
- -AlEngineer0.96
- orld's Fair0.99
- mai0.96
- DATADOG1.00
- I'sFair0.99
-
- AlEngineer0.98
- OpenAI0.93
- Vorld's Fair0.96
- AlEngineer0.98
- World's Fair0.98
- arize1.00
- Akamai1.00
- AlEngineer0.97
- DATADOG1.00
- vorld's Fair0.97
-
- AlEngineer0.99
- Evals are growing quickly0.98
- World'sFair1.00
- 100M1.00
- reddit1.00
- Booking.com1.00
- Uber1.00
- evals run every month1.00
- Tripadvisor1.00
- AATLASSIAN0.99
- owayfair0.94
- 3,800+1.00
- eval jobs run by the top Al teams1.00
- DOORDASH1.00
- klaviyo"0.98
- PagerDuty1.00
- 12.31.00
- Typeform1.00
- priceline1.00
- Handshake1.00
- eval jobs the average team runs0.97
- AlEngineer-0.94
- World'sF1.00
- Λ arize0.89
- 02 / 100.91
- air1.00
- Akama1.00
- Evals Track0.96
- AlEngineer0.97
- orld's F0.92
- Apara Dhinakaran / Co-founder & Chief Product Officer0.98
- arize0.98
-
- Evals were going to solve everything0.99
- AlEngineer1.00
- World'sFair1.00
- Kevin Weil1.00
- Writing evals is going to become a core skill for1.00
- product managers.1.00
- Mike Krieger1.00
- Writing evals is probably the most important thing we1.00
- can teach people.1.00
- We added evals1.00
- So they'll catch the failures, right?0.99
- Garry Tan0.95
- Evals are emerging as the real moat for Al startups.1.00
- Greg Brockman1.00
- Al0.82
- AlEngineer -0.94
- evals are surprisingly often all you need1.00
- They'll catch the failures, right?1.00
- World'sF1.00
- arize0.97
- -air0.84
- arize0.82
- Akama0.98
- Evals Track0.96
- Engineer1.00
- DOG0.82
- forld'sF0.95
- Apama Dhinakaran / Co-founder & Chief Product Officer0.97
- arize1.00
-
- AlEngineer0.98
- Agents got more complex0.99
- World's Fair0.99
- Run sub-agents1.00
- over long horizons1.00
- Research & code on0.98
- their own1.00
- Operate a computer1.00
- Reason1.00
- step-by-step1.00
- Use tools & browse1.00
- the web1.00
- Call functions & run1.00
- code1.00
- Answer a prompt1.00
- GFT-40.62
- interpreter0.97
- Workflows0.99
- -air0.90
- Open/0.91
- 20230.95
- 20241.00
- 20251.00
- 20261.00
- TIME1.00
- arize0.96
- AlEngineer -0.95
- er.r0.81
- harlae0.71
- orld'sF1.00
- Evals Track1.00
- Fair0.90
- DATAD0.99
- Apama Dhinakaran / Co-founder & Chief Product Officer0.98
- arize1.00
-
- copilot-prod0.98
- Trace: c969d7d3... caa33138 Session:488250.96
- AlEngineer0.99
- Failures got more complex1.00
- Signal 27 Evals & Metrics 50.98
- Status OK Cost $0.364943 Start Time 6/26/2026, 11:51:26 AM0.97
- World's Fair0.97
- Arize Default0.96
- Q0.98
- Trace Tree 411.00
- Search by span name1.00
- Timeline Agent Graph1.00
- AGENT ROUTER-TRAC0.98
- Jalbreak Detection: saf0.94
- We saw this with our own agent Alyx.1.00
- 00 Alyx0.77
- Traces 4.13k Spans 52.92k0.98
- 80.98
- ROUTER-TRACE_AGENT1.00
- 1 eval 62.83s0.98
- Attributes1.00
- Input / Out0.94
- one_alyx_agent1.00
- Q0.98
- Search1.00
- 62.54s1.00
- Context1.00
- seen late or lost in long run1.00
- 4.13k traces1.00
- Orchestrator - Runlteration 10.99
- 3.1s, 30,229 tk, $0.1051580.99
- input 1 item0.98
- value 3 items0.98
- Tool wrong tool, bad args, stale schema1.00
- Status1.00
- Kind1.00
- AGENT1.00
- get_trace_data1.00
- 0.15s1.00
- GQL. query GqISpanRecords..0.95
- input_context 70.98
- type1.00
- question1.00
- Execution1.00
- sandbox ≠ prod; deps, timeouts0.97
- G0.60
- AGENT1.00
- 0.08s1.00
- trace_ids Arra0.98
- 51.00
- Permission blocked, over-permitted, unsafe0.99
- Recovery1.00
- Loops1.00
- Model-harness fit0.98
- repeats errors, false success1.00
- Repeats the same step1.00
- differs across models1.00
- AGENT1.00
- AGENT1.00
- AGENT1.00
- AGENT1.00
- AGENT1.00
- AGENT1.00
- Orchestrator - Runlteration 30.99
- 3s, 31,525 tk, $0.0126240.99
- Orchestrator - Runlteration 20.99
- 2.69s, 31,116 tk, $0.0120540.99
- Orchestrator - Runlteration 40.99
- 0.01s0.95
- grep_ison0.90
- 0.02s1.00
- iq0.98
- • tracing_page0.92
- • span_ids Arra0.97
- highlighted_cc0.88
- • dataset 5 ite0.92
- [0]1.00
- [0]1.00
- page1.00
- project_id1.00
- view1.00
- G0.59
- AGENT1.00
- 1.89s, 31,700 tk, $0.0117391.00
- startTime1.00
- rld'sFai1.00
- AlEngineer0.95
- Λ arize0.88
- 81.00
- Trajectory1.00
- False Done claims success midway0.99
- ran, but not effectively1.00
- G0.52
- G0.53
- G0.56
- AGENT1.00
- AGENT1.00
- AGENT0.98
- AGENT1.00
- AGENT1.00
- AGENT1.00
- AGENT1.00
- 40.80
- Orchestrator - Runlteration 50.99
- 2.13s, 32,071 tk, $0.0121140.99
- 1.7s, 32,236 tk, $0.0116190.95
- Orchestrator - Runlteration 61.00
- Orchestrator - Runlteration 70.98
- 0.01s0.99
- 0.01s1.00
- 0.01s0.99
- iq0.97
- grep_jison0.93
- iq0.97
- • calumnVisib0.89
- endTime1.00
- subQuery1.00
- name1.00
- input_val1.00
- timeZone1.00
- queryFilte0.99
- attributes0.99
- output_v1.00
- otart tim0.91
- oget0.99
- porlze0.63
- Wol0.93
- Evals Track1.00
- AlEngi0.99
- rld'sF0.97
- Aparma Dhinakaran/ Co-founder & Chief Product Officer0.99
- arize1.00
-
- AlEngineer0.98
- World'sFair1.00
- The best way to evaluate an1.00
- agent is another agent1.00
- Engineer1.00
- Op1.00
- Id'sFair0.98
- Λ arize0.89
- AIE1.00
- gether0.94
- Worl1.00
- Evals Track1.00
- Engineer1.00
- Id'sFa0.97
- D/0.93
- Aparma Dhinakaran / Co-founder & Chief Product Officer0.96
- arize1.00
-
- AlEngineer1.00
- Different Evals for Different Failures0.99
- World's Fair0.97
- Deterministic code1.00
- Prompt hallucinations1.00
- Trajectory & harness failures1.00
- PRESENTED BY1.00
- Code Evaluators0.99
- Classic LLM-as-Judge1.00
- Agent-as-Judge1.00
- Microsoft1.00
- Unit tests, assertions, rules.1.00
- Score a known, well-defined failure.0.98
- Another agent reviews the whole trajectory.1.00
- input1.00
- output1.00
- criteria1.00
- traces0.97
- 0000.65
- flesystem1.00
- skills0.99
- valid_json (output)0.97
- / status == 2000.94
- / has_required_fields()0.97
- LLM1.00
- tools1.00
- Agent1.00
- subagents1.00
- no_pii_leaked(text)0.97
- AlEngine0.97
- PAsS 4 / 4 assertions - 0.02s0.90
- nAl0.98
- World's1.00
- label·score·explanation0.98
- insights1.00
- failure modes0.99
- evidence1.00
- Λarize0.92
- eer1.00
- sFai0.88
- Akal0.97
- Evals Track1.00
- AlEngine0.93
- ADOC0.94
- World's1.00
- Apama Dhinakaran / Co-founder & Chief Product Officer0.97
- arize1.00
-
- AlEngineer0.98
- Arize Signal1.00
- World's Fair0.96
- copilot-prod0.97
- Jun 19 - Jun 261.00
- 4 Live0.87
- Online Evaluators1.00
- Signal 271.00
- Evals & Metrics 50.97
- Traces1.00
- Spans1.00
- Sessions Agent Graph Agent Path1.00
- Signal Issues Active0.95
- arize-ai/arize0.98
- Viewing Signa's last 30 days of issues0.98
- ..0.78
- PRESENTED BY1.00
- Issues Found1.00
- Issues Over Time1.00
- Entity1.00
- Count1.00
- Microsoft1.00
- 271.00
- ROUTER-HOME_PAGE1.00
- run_mapping_fix_sub_agent0.98
- playground_agent1.00
- search_agent1.00
- Jun 90.98
- Jun 100.99
- Jun 140.99
- Jun 151.00
- Jun 161.00
- Jun 170.95
- Jun 190.99
- Jun 220.99
- Jun 240.99
- Jun 271.00
- Jun 281.00
- Jun 291.00
- Jun 301.00
- Detected Issues0.98
- todo_update retry loop on already-in-progress todo causes stream cancellation0.99
- Mute1.00
- Mark as fixed1.00
- #1 todo_update retry loop on already-in-0.99
- todo_update1.00
- 6/17|2026, 6:24 AM0.97
- progress todo causes stream cancellation0.99
- A'ROUTER-PROMPT_OPTIMIZATION' session0.95
- Next Steps Add to Dataset1.00
- + Create Evaluator0.99
- Open PR0.99
- ended with 'status_code=ERROR' and..0.96
- 6 matching traces1.00
- Overview1.00
- Al0.79
- World's1.00
- AlEngineer0.99
- #2 ROUTER-HOME_PAGE generates over-0.99
- verbose responses causing 20-75s latency0.99
- item cancels the parent stream.0.99
- A ROUTER-PROMPT_OPTIMIZATION session ended with status_code=ERROR and status_message=stream_cancelled . The export contains 100.99
- todo_update spans across the batch, and this is the same pattern previously identified where a todo_update retry loop on an already-in-progress0.99
- 83.57s end-to-end. The elevated latency in this run.0.98
- Reconfirmed: a ROUTER-HOME_PAGE session took0.99
- Evidence1.00
- • trace 1a9d4cd945573e097b4471e3dc7046cd — ROUTER-PROMPT_OPTIMIZATION, ERROR, stream_cancelled. User was working on prompt0.99
- arize1.00
- 33 matching traces1.00
- integration (partial message visible in attributes).0.99
- Proposed fix1.00
- Fa1.00
- arize0.81
- Akam1.00
- Evals Track1.00
- AlEngineer0.96
- DOG1.00
- Vorld's1.00
- Apara Dhinakaran/ Co-founder & Chief Product Officer0.98
- arize1.00
-
- AlEngineer1.00
- World's Fair0.98
- Λ arize0.92
- PRESENTED BY1.00
- Microsoft1.00
- Trace. Eval. Learn0.97
- -AlEngineer0.99
- rld's Fai0.96
- Aparna Dhinakaran1.00
- @aparnadhinak1.00
- tarlze0.69
- Oso0.68
- Evals Track1.00
- ite1.00
- gineer1.00
- I'sFai0.95
- Aparna Dhinakaran / Co-founder & Chief Product Officer0.97
- arize1.00
-
- AlEngineer0.97
- OpenAl0.92
- rld's Fair0.96
- AlEngineer0.99
- osoft1.00
- orld's Fair0.98
- arize1.00
- ngineer1.00
- Buildkite1.00
- I'sFair0.94
Transcript
69 cues· 968 words· 5,196 chars
- 0:12 Awesome.
- 0:13 Well, hey, everyone.
- 0:14 My name is Aparna, one of the founders of Arise.
- 0:16 We work with some amazing teams to help them build evals.
- 0:21 And we have an incredible lineup of talks for you all today at the evals track.
- 0:26 It's happening in room 2005.
- 0:29 And there's going to be amazing speakers from Turnbench and Uber and Snorkel all happening after this.
- 0:35 But today, I'm here to talk to you about the future of evals.
- 0:38 evals have gone from the new skill that every PM and every AI engineer has to learn to the thing that every serious AI team is betting on.
- 0:50 We've been really fortunate to get to work with some of the best AI teams in the world.
- 0:54 So we get a front row seat into not just what's happening when they're building their actual agents and before they actually ship, but actually the evals the teams are running on their live production agent via their traces.
- 1:09 A little bit of some stats for you guys.
- 1:10 We run over 100 million evals every month.
- 1:13 The average team runs about 12 different eval jobs, with the top teams running over 3,800 different evaluators.
- 1:22 And offline evals, online evals, they each have their own place, but today what I'm actually gonna talk to you about is the teams that are running evals on their traces.
- 1:32 This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning loops.
- 1:43 And the industry kind of agrees.
- 1:45 I mean, all the CPOs of Anthropic, OpenAI, all, you know, GDB, you have Gary Tan saying, evals are everything you need.
- 1:53 And the whole industry kind of agrees.
- 1:55 So we added evals, they catch all the failures, right?
- 2:00 Here's the problem.
- 2:01 While we were building all of these first gen evals, the thing that we were actually evaluating has changed underneath us.
- 2:09 In 2023, it was about just answering a prompt.
- 2:12 In 2024, we started to see all the frontier models.
- 2:16 They've added tool calls, they've added reasoning, they've added deep research.
- 2:21 Now what we have is teams running loops on real world data with sub-agents kicked off on long horizon tasks.
- 2:30 Every one of these was actually a massive jump in complexity and we didn't just make the problem harder, we actually got a fundamentally different type of problem.
- 2:41 What that meant is that as these systems got more complex, so did the way that they actually fail.
- 2:46 We're really lucky because we have our own agent that we've built, Alex, that lives in our UI.
- 2:50 And we kind of get to feel this pain ourselves.
- 2:54 Every time the Frontier Labs added new functionality, we added it to our agent.
- 2:58 And now Alex has much longer memory.
- 3:03 It has the ability to create dynamic UIs.
- 3:05 It can go search across an enormous volume of traces.
- 3:09 But we also realized that it would forget context.
- 3:12 It wouldn't know when something was done.
- 3:14 Sometimes it would just get stuck in these loops.
- 3:17 And the key thing here is that the classical LLM as a judge evals that probably many of you have written in this room just weren't enough for us to be able to catch all the types of failures that we were experiencing.
- 3:32 I mean, it's just fundamentally different, right?
- 3:34 You have a deterministic flow and now what we have is literally every time a user interacted with Alex, it would create a new UI.
- 3:42 That's a fundamentally different trajectory.
- 3:45 So this led to our really big revelation.
- 3:48 What if the best way to evaluate an agent was actually with an agent?
- 3:54 Doesn't mean that all of the ways that we did evals with deterministic evals, with LLM as a judge, classic evals, doesn't matter anymore, but it just means that we have a different type of tool to solve a different type of problem.
- 4:07 Agent as a judge is about adaptive dynamic analysis.
- 4:12 LLM as a judge just gives you a fixed rubric with these fixed scores.
- 4:16 It's what everyone's doing, but when your agent's doing
- 4:19 completely different trajectories every time a user puts in data, it just means that you need a fundamentally different type of eval.
- 4:27 My take is that most teams today are doing the first two, but the future of evals is actually having all three.
- 4:35 And today I'm actually excited to share we've released Agent as a Judge to help our teams on their eval journey.
loading