read-only demo

Videos 3z2uT5aDx_Y

Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft

index_state ready data_status ok

AI Engineer· published 2026-07-19· 0:37:44· en-US· indexed 2026-08-10 19:43

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:36, 1 of 1 keyframes kept
  2. Shot 1, 0:36 to 1:13, 1 of 1 keyframes kept
  3. Shot 2, 1:13 to 2:01, 1 of 1 keyframes kept
  4. Shot 3, 2:01 to 2:28, 1 of 1 keyframes kept
  5. Shot 4, 2:28 to 2:55, 0 of 1 keyframes kept
  6. Shot 5, 2:55 to 3:22, 0 of 1 keyframes kept
  7. Shot 6, 3:22 to 3:50, 0 of 1 keyframes kept
  8. Shot 7, 3:50 to 4:17, 0 of 1 keyframes kept
  9. Shot 8, 4:17 to 4:44, 0 of 1 keyframes kept
  10. Shot 9, 4:44 to 5:11, 0 of 1 keyframes kept
  11. Shot 10, 5:11 to 5:38, 0 of 1 keyframes kept
  12. Shot 11, 5:38 to 6:00, 1 of 1 keyframes kept
  13. Shot 12, 6:00 to 6:12, 0 of 1 keyframes kept
  14. Shot 13, 6:12 to 6:45, 0 of 1 keyframes kept
  15. Shot 14, 6:45 to 7:18, 0 of 1 keyframes kept
  16. Shot 15, 7:18 to 7:44, 1 of 1 keyframes kept
  17. Shot 16, 7:44 to 8:10, 0 of 1 keyframes kept
  18. Shot 17, 8:10 to 8:36, 0 of 1 keyframes kept
  19. Shot 18, 8:36 to 9:02, 1 of 1 keyframes kept
  20. Shot 19, 9:02 to 9:28, 0 of 1 keyframes kept
  21. Shot 20, 9:28 to 9:54, 0 of 1 keyframes kept
  22. Shot 21, 9:54 to 10:19, 0 of 1 keyframes kept
  23. Shot 22, 10:19 to 10:45, 0 of 1 keyframes kept
  24. Shot 23, 10:45 to 11:11, 0 of 1 keyframes kept
  25. Shot 24, 11:11 to 11:42, 1 of 1 keyframes kept
  26. Shot 25, 11:42 to 12:14, 0 of 1 keyframes kept
  27. Shot 26, 12:14 to 12:45, 0 of 1 keyframes kept
  28. Shot 27, 12:45 to 13:11, 1 of 1 keyframes kept
  29. Shot 28, 13:11 to 13:37, 0 of 1 keyframes kept
  30. Shot 29, 13:37 to 14:03, 0 of 1 keyframes kept
  31. Shot 30, 14:03 to 14:29, 0 of 1 keyframes kept
  32. Shot 31, 14:29 to 14:48, 1 of 1 keyframes kept
  33. Shot 32, 14:48 to 14:51, 0 of 1 keyframes kept
  34. Shot 33, 14:51 to 15:04, 0 of 1 keyframes kept
  35. Shot 34, 15:04 to 15:15, 1 of 1 keyframes kept
  36. Shot 35, 15:15 to 15:18, 0 of 1 keyframes kept
  37. Shot 36, 15:18 to 15:32, 0 of 1 keyframes kept
  38. Shot 37, 15:32 to 15:38, 0 of 1 keyframes kept
  39. Shot 38, 15:38 to 15:46, 0 of 1 keyframes kept
  40. Shot 39, 15:46 to 16:21, 0 of 1 keyframes kept
  41. Shot 40, 16:21 to 16:57, 0 of 1 keyframes kept
  42. Shot 41, 16:57 to 17:30, 1 of 1 keyframes kept
  43. Shot 42, 17:30 to 17:58, 1 of 1 keyframes kept
  44. Shot 43, 17:58 to 18:26, 1 of 1 keyframes kept
  45. Shot 44, 18:26 to 19:01, 1 of 1 keyframes kept
  46. Shot 45, 19:01 to 19:36, 0 of 1 keyframes kept
  47. Shot 46, 19:36 to 20:07, 1 of 1 keyframes kept
  48. Shot 47, 20:07 to 20:37, 0 of 1 keyframes kept
  49. Shot 48, 20:37 to 21:08, 0 of 1 keyframes kept
  50. Shot 49, 21:08 to 21:35, 1 of 1 keyframes kept
  51. Shot 50, 21:35 to 22:02, 0 of 1 keyframes kept
  52. Shot 51, 22:02 to 22:29, 0 of 1 keyframes kept
  53. Shot 52, 22:29 to 22:56, 1 of 1 keyframes kept
  54. Shot 53, 22:56 to 23:24, 0 of 1 keyframes kept
  55. Shot 54, 23:24 to 23:52, 1 of 1 keyframes kept
  56. Shot 55, 23:52 to 24:20, 0 of 1 keyframes kept
  57. Shot 56, 24:20 to 24:48, 0 of 1 keyframes kept
  58. Shot 57, 24:48 to 25:16, 1 of 1 keyframes kept
  59. Shot 58, 25:16 to 25:43, 0 of 1 keyframes kept
  60. Shot 59, 25:43 to 26:15, 1 of 1 keyframes kept
  61. Shot 60, 26:15 to 26:46, 0 of 1 keyframes kept
  62. Shot 61, 26:46 to 27:17, 0 of 1 keyframes kept
  63. Shot 62, 27:17 to 27:48, 1 of 1 keyframes kept
  64. Shot 63, 27:48 to 28:16, 0 of 1 keyframes kept
  65. Shot 64, 28:16 to 28:44, 0 of 1 keyframes kept
  66. Shot 65, 28:44 to 29:11, 1 of 1 keyframes kept
  67. Shot 66, 29:11 to 29:52, 0 of 1 keyframes kept
  68. Shot 67, 29:52 to 30:17, 1 of 1 keyframes kept
  69. Shot 68, 30:17 to 30:42, 0 of 1 keyframes kept
  70. Shot 69, 30:42 to 31:09, 1 of 1 keyframes kept
  71. Shot 70, 31:09 to 31:36, 0 of 1 keyframes kept
  72. Shot 71, 31:36 to 32:02, 0 of 1 keyframes kept
  73. Shot 72, 32:02 to 32:29, 0 of 1 keyframes kept
  74. Shot 73, 32:29 to 33:02, 1 of 1 keyframes kept
  75. Shot 74, 33:02 to 33:36, 0 of 1 keyframes kept
  76. Shot 75, 33:36 to 34:09, 0 of 1 keyframes kept
  77. Shot 76, 34:09 to 34:35, 1 of 1 keyframes kept
  78. Shot 77, 34:35 to 35:02, 0 of 1 keyframes kept
  79. Shot 78, 35:02 to 35:28, 0 of 1 keyframes kept
  80. Shot 79, 35:28 to 35:55, 0 of 1 keyframes kept
  81. Shot 80, 35:55 to 36:21, 0 of 1 keyframes kept
  82. Shot 81, 36:21 to 36:48, 0 of 1 keyframes kept
  83. Shot 82, 36:48 to 36:49, 1 of 1 keyframes kept
  84. Shot 83, 36:49 to 36:54, 0 of 1 keyframes kept
  85. Shot 84, 36:54 to 37:44, 0 of 1 keyframes kept

85 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
253
whisperx 253
chunks
66
from 253 cues
keyframes
28
kept of 85 captured
frames with text
28
612 lines read
chapters
0
from the source metadata
keyframe bytes
6.7 MB
word timings on 253 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 00:02 1m 40s
stt done 2026-08-10 00:04 37s
chunk done 2026-08-10 00:04 0s
text_embed done 2026-08-10 19:43 0s
keyframe done 2026-08-10 00:04 1m 27s
ocr done 2026-08-10 00:06 15s
frame_embed done 2026-08-10 19:43 5s

Frames, and what the machine read

  • 0:32 #0 done4 line(s)

    shot 0·sharpness 842.4

    1. Build Evals that0.99
    2. actuallymatter1.00
    3. Nick Ung0.95
    4. Making customer support agent eval more consequential and actionable.1.00
  • 1:05 #1 done4 line(s)

    shot 1·sharpness 844.5

    1. Build Evals that0.99
    2. actually matter0.97
    3. AkshaySharma1.00
    4. Making customer support agent eval more consequential and actionable.1.00
  • 1:19 #2 done6 line(s)

    shot 2·sharpness 830.5

    1. Agenda1.00
    2. Offline Evaluations1.00
    3. Online Evaluations1.00
    4. Eval Harness1.00
    5. Future Steps1.00
    6. Nick Ung1.00
  • 2:17 #3 done26 line(s)

    shot 3·sharpness 1523.3

    1. Anatomy of Al Agent Eval Flywheel0.99
    2. Development1.00
    3. Production1.00
    4. Agent Engineering1.00
    5. Offline Eval1.00
    6. Yes1.00
    7. Al Agent0.99
    8. Context management1.00
    9. Simulated1.00
    10. Can we0.95
    11. RAG pipeline1.00
    12. Tools definition1.00
    13. Graph orchestration1.00
    14. LLM as a judge0.98
    15. User LLM1.00
    16. conversation1.00
    17. launch?1.00
    18. System prompt0.98
    19. Online Eval1.00
    20. Nick Ung1.00
    21. No1.00
    22. LLM as a judge1.00
    23. Langsmith traces0.99
    24. Continual Learning1.00
    25. Human annotator1.00
    26. lyft1.00
  • 2:52 #4 skipped

    shot 4·duplicate of #3

  • 3:19 #5 skipped

    shot 5·duplicate of #3

  • 3:41 #6 skipped

    shot 6·duplicate of #3

  • 4:03 #7 skipped

    shot 7·duplicate of #3

  • 4:20 #8 skipped

    shot 8·duplicate of #3

  • 4:57 #9 skipped

    shot 9·duplicate of #3

  • 5:29 #10 skipped

    shot 10·duplicate of #3

  • 5:43 #11 done15 line(s)

    shot 11·sharpness 1077.7

    1. Why do Eval Fail?0.97
    2. 21.00
    3. 31.00
    4. Numbers don't gate0.98
    5. Judges are too noisy to trust1.00
    6. No owner of the0.97
    7. anything1.00
    8. consequence1.00
    9. If moving a score costs nothing, no one0.99
    10. People quietly stop believing the metric.1.00
    11. When something regresses, nobody is1.00
    12. defends it.0.97
    13. on the hook.0.99
    14. What to do?: Make eval results both trustworthy and consequential. Automation is the last step, not the first.0.99
    15. NickUng1.00
  • 6:09 #12 skipped

    shot 12·duplicate of #3

  • 6:35 #13 skipped

    shot 13·duplicate of #11

  • 6:59 #14 skipped

    shot 14·duplicate of #11

  • 7:29 #15 done35 line(s)

    shot 15·sharpness 2144.5

    1. Offline Eval0.99
    2. π0.69
    3. Building τ2-Bench for Agentic Application1.00
    4. Agent Domain Policy1.00
    5. As a telecom agenl, you can help0.92
    6. users with technical support.0.98
    7. The curent time is 202502-250.94
    8. 12:08:00 EST.0.98
    9. Agent1.00
    10. Agent1.00
    11. Tools1.00
    12. Agent1.00
    13. DB1.00
    14. date_of_birth = "1985-06-15°0.94
    15. phone_number = "555-123-2002"0.97
    16. World1.00
    17. You mobile data is not working0.99
    18. User Instruction0.98
    19. [device]1.00
    20. speed on your phone.0.98
    21. User1.00
    22. User1.00
    23. Tools0.94
    24. User1.00
    25. DB1.00
    26. data_erabled = true0.92
    27. that interact with a database, and is tasked with resolving the user's request via Tool-Agent-User0.99
    28. Figure 1: Supporting dual-control environment in τ2-bench. The agent have access to a set of tools0.99
    29. Nick Ung1.00
    30. (TAU) interactions while adhering to the domain policy. To test real-world scenarios, the user is0.99
    31. simulated by another AI agent given a scenario-based instruction and a set of tools that interact with1.00
    32. its own database. The simulated user can be regarded as handling an easier version of the TAU1.00
    33. interaction in a dual format (Tool-User-Agent), where it only need to follow instructions but does not0.99
    34. need to reason about solutions for the task.1.00
    35. Image taken from: t2-Bench: Evaluating Conversational Agents in a Dual-Control Environment (arXiv.org)0.99
  • 7:50 #16 skipped

    shot 16·duplicate of #15

  • 8:33 #17 skipped

    shot 17·duplicate of #15

  • 8:52 #18 done47 line(s)

    shot 18·sharpness 1259.4

    1. Offline Eval1.00
    2. π0.67
    3. Offline Eval - simulator0.99
    4. # al Aosist orfline Simlatioe0.81
    5. simulation:0.98
    6. 1d: driver_cancel_fee_loyalty_concession0.94
    7. intest: earnings.cancel_fee_dispute0.86
    8. description:0.95
    9. Driver disputes missing cancel fee. Rider canceled withi0.97
    10. wiedow (ne fee owed by policy). but driver is long-tenur0.89
    11. (So,o,Oo,S,a,OS_T,0.69
    12. User: ..0.94
    13. trajectoryτ0.97
    14. a_T)0.98
    15. # Initial state0.87
    16. Assistant: ...0.99
    17. world_state:0.89
    18. drivar:0.93
    19. loyalty segnest: top,5,pt0.63
    20. tonure_years: 6.20.93
    21. tier: lox0.82
    22. User:...1.00
    23. Tool: ...0.97
    24. Assistant: ...0.95
    25. prior_concesssona_90d:0.66
    26. cancelnd_by: rider0.97
    27. seconds_te_cancel: 470.63
    28. Langgraph Agent0.99
    29. policy__fe_d: fale0.65
    30. # Simvlated user0.85
    31. user_persona:0.95
    32. archetype: loyal_frustrated_lux_driver0.97
    33. framing: fairness_not_money0.96
    34. sentinent: frustrated_but_loyal0.93
    35. opening_nessage: >0.89
    36. years driving Lux and I gut stiffed on a cancel fee?0.93
    37. Rider hailed aftar I drove 2 miles. This ian't right.0.87
    38. Evaluator1.00
    39. LLM Judge0.99
    40. Assertion1.00
    41. Code1.00
    42. concession_granted: true1.00
    43. is_escalated: false1.00
    44. concession_amount_usd: $101.00
    45. turns_to_resolution: {max: 6 }0.98
    46. Nick Ung1.00
    47. lyft1.00
  • 9:15 #19 skipped

    shot 19·duplicate of #18

  • 9:38 #20 skipped

    shot 20·duplicate of #18

  • 10:09 #21 skipped

    shot 21·duplicate of #18

  • 10:27 #22 skipped

    shot 22·duplicate of #18

  • 11:08 #23 skipped

    shot 23·duplicate of #18

Transcript

253 cues· 4,822 words· 26,732 chars

  1. 0:05 Hi everyone, my name is Nick and I'm here with Akshay to give a talk about evals.
  2. 0:14 We are from Lyft and we've been building Lyft customer support AI agent for a year or two now and gave a lot of thoughts about how to build eval that actually matters and scale our multi-AI agent system.
  3. 0:34 Just a little bit of quick introductions.
  4. 0:36 My name is Nick.
  5. 0:37 I'm a data science manager.
  6. 0:39 I've been at Lyft for six years, a long time Lyfter.
  7. 0:43 Really excited to talk to you a little bit more about eBumps.
  8. 0:50 Hi everyone, I'm Akshay.
  9. 0:52 I'm in Nick's team and we've been working together on customer support agents for Lyft, improving the hardness, improving the evals, things like that.
  10. 1:04 And I've been at Lyft for almost four years now and very excited to be here and talk about building evals that actually matter.
  11. 1:16 super excited to be here and super honored to be to be on the online track for ai engineer warfare yeah and let's dive in for the agenda of today i will talk primarily focus on eval
  12. 1:35 We will start by sharing how we think about the end-to-end pipeline for our evaluations for building customer support AI agent system.
  13. 1:46 We'll go into detail into each component more deeply as we go.
  14. 1:51 We'll start by talking about offline evaluations, online evaluations, eval harness, as well as what we are planning to build going forward.
  15. 2:02 All right, let's try this.
  16. 2:04 I want to quickly explain the
  17. 2:09 you know, high level system of how we think about evaluation system for AI agents.
  18. 2:15 So here you can see we have the development phase and the production phase.
  19. 2:21 So during development, if you're building agents, you should be very familiar with, you know, managing contacts, building RAC pipeline to give your agents educational contacts, defining your tool, building your agentic graph, as far as writing a system prompt.
  20. 2:39 So once all of that engineering process is done, you have an AI agent.
  21. 2:46 The way we think about this is before we launch this AI agent to productions, we want to go through a rigorous offline evaluation process to make sure that this agent actually has sufficient performance before we launch this to live users.
  22. 3:04 So coming from data science and machine learning background,
  23. 3:09 We've been building machine learning model for a while.
  24. 3:14 And I think the way that we think about agent development is very similar to building machine learning models as well.
  25. 3:21 If we are running offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well or any other agentic applications.
  26. 3:36 But what I think, you know, offline evaluation, how that is different than traditional machine learning model is that, you know, we typically we're building specifically for customer support AI use case, we're building an agent that's multi-turn.
  27. 3:52 So for offline evaluation, there will be a component of simulated conversations.
  28. 3:58 So you typically want to have a
  29. 4:02 data sets, synthetic data set that's representative of your production's traffic, have or use the LLM that plays out the complete multi-turn simulated conversation, as well as having a grader, such as the LLM as a judge, to be able to evaluate how good that interaction was.
  30. 4:23 And then we have a launch gate, right?
  31. 4:26 We want to make sure that we have certain criteria on our offline eval, and we're meeting that criteria before we decide to launch this AI agent to productions.
  32. 4:38 And so the real imperative here really is that we don't want to use our live user as test data for our AI agents.
  33. 4:47 And I think in any cases, that is not good practice.
  34. 4:52 So we really want to emphasize the importance of having an offline evaluation process.
  35. 4:58 So once the agent hit productions, we also have an online evaluation pipeline as well.
  36. 5:04 We have our own favorite tracing tools to trace all the executions and contacts that the AI agent used to respond to a real user in a production environment.
  37. 5:17 We have our online grader as well that grades how well our AI agent is doing in productions.
  38. 5:23 And as far as having a human in the loop pipeline to do error analysis, identify failure mode, and feedback that insights to the development teams to continuously improve our AI agents.
  39. 5:39 I want to quickly go over, I think, three of the most common reasons why we think evaluation typically fail for different teams.
  40. 5:49 So the first reason is that the grader that we create, the scores that we create, needs to be meaningfully gating something.
  41. 6:00 what we really emphasized only in the previous slide, that we need to have a launch gate.
  42. 6:05 If your LLM as a judge is just floating out there, there's a score, but no one is really using that score as a meaningful gate for your development and productions environment, then that LLM as a judge is not available.
  43. 6:22 We've also seen a lot of
  44. 6:26 mishap people have when they're creating their LMS judge.
  45. 6:31 There's a lot of different opinions out there in terms of how do you create a good LMS judge.
  46. 6:37 And typically, and unfortunately also very early on in our journey, the LMS judge that we created are
  47. 6:46 very noisy, too generic.
  48. 6:48 It will output a score but people don't really believe in what the Allen judge is doing or they don't think the Allen judge insight is actionable.
  49. 7:00 And finally, I think when something regresses in productions, we need to have clear mechanism to be able to catch that regression, as well as identify clear owners to be able to take actions on the insights of our graders and regression gate.
  50. 7:20 Very cool.

Open at this second