Videos 3z2uT5aDx_Y
Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
Scene timeline
85 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 253
- whisperx 253
- chunks
- 66
- from 253 cues
- keyframes
- 28
- kept of 85 captured
- frames with text
- 28
- 612 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.7 MB
- word timings on 253 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 00:02 | 1m 40s |
stt |
done | — | 2026-08-10 00:04 | 37s |
chunk |
done | — | 2026-08-10 00:04 | 0s |
text_embed |
done | — | 2026-08-10 19:43 | 0s |
keyframe |
done | — | 2026-08-10 00:04 | 1m 27s |
ocr |
done | — | 2026-08-10 00:06 | 15s |
frame_embed |
done | — | 2026-08-10 19:43 | 5s |
Frames, and what the machine read
-
- Build Evals that0.99
- actuallymatter1.00
- Nick Ung0.95
- Making customer support agent eval more consequential and actionable.1.00
-
- Build Evals that0.99
- actually matter0.97
- AkshaySharma1.00
- Making customer support agent eval more consequential and actionable.1.00
-
- Agenda1.00
- Offline Evaluations1.00
- Online Evaluations1.00
- Eval Harness1.00
- Future Steps1.00
- Nick Ung1.00
-
- Anatomy of Al Agent Eval Flywheel0.99
- Development1.00
- Production1.00
- Agent Engineering1.00
- Offline Eval1.00
- Yes1.00
- Al Agent0.99
- Context management1.00
- Simulated1.00
- Can we0.95
- RAG pipeline1.00
- Tools definition1.00
- Graph orchestration1.00
- LLM as a judge0.98
- User LLM1.00
- conversation1.00
- launch?1.00
- System prompt0.98
- Online Eval1.00
- Nick Ung1.00
- No1.00
- LLM as a judge1.00
- Langsmith traces0.99
- Continual Learning1.00
- Human annotator1.00
- lyft1.00
-
- Why do Eval Fail?0.97
- 21.00
- 31.00
- Numbers don't gate0.98
- Judges are too noisy to trust1.00
- No owner of the0.97
- anything1.00
- consequence1.00
- If moving a score costs nothing, no one0.99
- People quietly stop believing the metric.1.00
- When something regresses, nobody is1.00
- defends it.0.97
- on the hook.0.99
- What to do?: Make eval results both trustworthy and consequential. Automation is the last step, not the first.0.99
- NickUng1.00
-
- Offline Eval0.99
- π0.69
- Building τ2-Bench for Agentic Application1.00
- Agent Domain Policy1.00
- As a telecom agenl, you can help0.92
- users with technical support.0.98
- The curent time is 202502-250.94
- 12:08:00 EST.0.98
- Agent1.00
- Agent1.00
- Tools1.00
- Agent1.00
- DB1.00
- date_of_birth = "1985-06-15°0.94
- phone_number = "555-123-2002"0.97
- World1.00
- You mobile data is not working0.99
- User Instruction0.98
- [device]1.00
- speed on your phone.0.98
- User1.00
- User1.00
- Tools0.94
- User1.00
- DB1.00
- data_erabled = true0.92
- that interact with a database, and is tasked with resolving the user's request via Tool-Agent-User0.99
- Figure 1: Supporting dual-control environment in τ2-bench. The agent have access to a set of tools0.99
- Nick Ung1.00
- (TAU) interactions while adhering to the domain policy. To test real-world scenarios, the user is0.99
- simulated by another AI agent given a scenario-based instruction and a set of tools that interact with1.00
- its own database. The simulated user can be regarded as handling an easier version of the TAU1.00
- interaction in a dual format (Tool-User-Agent), where it only need to follow instructions but does not0.99
- need to reason about solutions for the task.1.00
- Image taken from: t2-Bench: Evaluating Conversational Agents in a Dual-Control Environment (arXiv.org)0.99
-
- Offline Eval1.00
- π0.67
- Offline Eval - simulator0.99
- # al Aosist orfline Simlatioe0.81
- simulation:0.98
- 1d: driver_cancel_fee_loyalty_concession0.94
- intest: earnings.cancel_fee_dispute0.86
- description:0.95
- Driver disputes missing cancel fee. Rider canceled withi0.97
- wiedow (ne fee owed by policy). but driver is long-tenur0.89
- (So,o,Oo,S,a,OS_T,0.69
- User: ..0.94
- trajectoryτ0.97
- a_T)0.98
- # Initial state0.87
- Assistant: ...0.99
- world_state:0.89
- drivar:0.93
- loyalty segnest: top,5,pt0.63
- tonure_years: 6.20.93
- tier: lox0.82
- User:...1.00
- Tool: ...0.97
- Assistant: ...0.95
- prior_concesssona_90d:0.66
- cancelnd_by: rider0.97
- seconds_te_cancel: 470.63
- Langgraph Agent0.99
- policy__fe_d: fale0.65
- # Simvlated user0.85
- user_persona:0.95
- archetype: loyal_frustrated_lux_driver0.97
- framing: fairness_not_money0.96
- sentinent: frustrated_but_loyal0.93
- opening_nessage: >0.89
- years driving Lux and I gut stiffed on a cancel fee?0.93
- Rider hailed aftar I drove 2 miles. This ian't right.0.87
- Evaluator1.00
- LLM Judge0.99
- Assertion1.00
- Code1.00
- concession_granted: true1.00
- is_escalated: false1.00
- concession_amount_usd: $101.00
- turns_to_resolution: {max: 6 }0.98
- Nick Ung1.00
- lyft1.00
Transcript
253 cues· 4,822 words· 26,732 chars
- 0:05 Hi everyone, my name is Nick and I'm here with Akshay to give a talk about evals.
- 0:14 We are from Lyft and we've been building Lyft customer support AI agent for a year or two now and gave a lot of thoughts about how to build eval that actually matters and scale our multi-AI agent system.
- 0:34 Just a little bit of quick introductions.
- 0:36 My name is Nick.
- 0:37 I'm a data science manager.
- 0:39 I've been at Lyft for six years, a long time Lyfter.
- 0:43 Really excited to talk to you a little bit more about eBumps.
- 0:50 Hi everyone, I'm Akshay.
- 0:52 I'm in Nick's team and we've been working together on customer support agents for Lyft, improving the hardness, improving the evals, things like that.
- 1:04 And I've been at Lyft for almost four years now and very excited to be here and talk about building evals that actually matter.
- 1:16 super excited to be here and super honored to be to be on the online track for ai engineer warfare yeah and let's dive in for the agenda of today i will talk primarily focus on eval
- 1:35 We will start by sharing how we think about the end-to-end pipeline for our evaluations for building customer support AI agent system.
- 1:46 We'll go into detail into each component more deeply as we go.
- 1:51 We'll start by talking about offline evaluations, online evaluations, eval harness, as well as what we are planning to build going forward.
- 2:02 All right, let's try this.
- 2:04 I want to quickly explain the
- 2:09 you know, high level system of how we think about evaluation system for AI agents.
- 2:15 So here you can see we have the development phase and the production phase.
- 2:21 So during development, if you're building agents, you should be very familiar with, you know, managing contacts, building RAC pipeline to give your agents educational contacts, defining your tool, building your agentic graph, as far as writing a system prompt.
- 2:39 So once all of that engineering process is done, you have an AI agent.
- 2:46 The way we think about this is before we launch this AI agent to productions, we want to go through a rigorous offline evaluation process to make sure that this agent actually has sufficient performance before we launch this to live users.
- 3:04 So coming from data science and machine learning background,
- 3:09 We've been building machine learning model for a while.
- 3:14 And I think the way that we think about agent development is very similar to building machine learning models as well.
- 3:21 If we are running offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well or any other agentic applications.
- 3:36 But what I think, you know, offline evaluation, how that is different than traditional machine learning model is that, you know, we typically we're building specifically for customer support AI use case, we're building an agent that's multi-turn.
- 3:52 So for offline evaluation, there will be a component of simulated conversations.
- 3:58 So you typically want to have a
- 4:02 data sets, synthetic data set that's representative of your production's traffic, have or use the LLM that plays out the complete multi-turn simulated conversation, as well as having a grader, such as the LLM as a judge, to be able to evaluate how good that interaction was.
- 4:23 And then we have a launch gate, right?
- 4:26 We want to make sure that we have certain criteria on our offline eval, and we're meeting that criteria before we decide to launch this AI agent to productions.
- 4:38 And so the real imperative here really is that we don't want to use our live user as test data for our AI agents.
- 4:47 And I think in any cases, that is not good practice.
- 4:52 So we really want to emphasize the importance of having an offline evaluation process.
- 4:58 So once the agent hit productions, we also have an online evaluation pipeline as well.
- 5:04 We have our own favorite tracing tools to trace all the executions and contacts that the AI agent used to respond to a real user in a production environment.
- 5:17 We have our online grader as well that grades how well our AI agent is doing in productions.
- 5:23 And as far as having a human in the loop pipeline to do error analysis, identify failure mode, and feedback that insights to the development teams to continuously improve our AI agents.
- 5:39 I want to quickly go over, I think, three of the most common reasons why we think evaluation typically fail for different teams.
- 5:49 So the first reason is that the grader that we create, the scores that we create, needs to be meaningfully gating something.
- 6:00 what we really emphasized only in the previous slide, that we need to have a launch gate.
- 6:05 If your LLM as a judge is just floating out there, there's a score, but no one is really using that score as a meaningful gate for your development and productions environment, then that LLM as a judge is not available.
- 6:22 We've also seen a lot of
- 6:26 mishap people have when they're creating their LMS judge.
- 6:31 There's a lot of different opinions out there in terms of how do you create a good LMS judge.
- 6:37 And typically, and unfortunately also very early on in our journey, the LMS judge that we created are
- 6:46 very noisy, too generic.
- 6:48 It will output a score but people don't really believe in what the Allen judge is doing or they don't think the Allen judge insight is actionable.
- 7:00 And finally, I think when something regresses in productions, we need to have clear mechanism to be able to catch that regression, as well as identify clear owners to be able to take actions on the insights of our graders and regression gate.
- 7:20 Very cool.
loading