read-only demo

Videos FB-MLPhL9Ms

The maturity phases of running evals — Phil Hetzel, Braintrust

index_state ready data_status ok

AI Engineer· published 2026-05-27· 0:18:33· en-US· indexed 2026-08-10 19:55

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:42, 1 of 1 keyframes kept
  5. Shot 4, 0:42 to 1:11, 1 of 1 keyframes kept
  6. Shot 5, 1:11 to 1:45, 1 of 1 keyframes kept
  7. Shot 6, 1:45 to 2:20, 1 of 1 keyframes kept
  8. Shot 7, 2:20 to 2:45, 1 of 1 keyframes kept
  9. Shot 8, 2:45 to 3:11, 0 of 1 keyframes kept
  10. Shot 9, 3:11 to 3:39, 1 of 1 keyframes kept
  11. Shot 10, 3:39 to 4:07, 0 of 1 keyframes kept
  12. Shot 11, 4:07 to 4:35, 1 of 1 keyframes kept
  13. Shot 12, 4:35 to 5:03, 0 of 1 keyframes kept
  14. Shot 13, 5:03 to 5:31, 0 of 1 keyframes kept
  15. Shot 14, 5:31 to 5:59, 0 of 1 keyframes kept
  16. Shot 15, 5:59 to 6:27, 1 of 1 keyframes kept
  17. Shot 16, 6:27 to 6:55, 0 of 1 keyframes kept
  18. Shot 17, 6:55 to 7:24, 1 of 1 keyframes kept
  19. Shot 18, 7:24 to 7:53, 1 of 1 keyframes kept
  20. Shot 19, 7:53 to 8:22, 0 of 1 keyframes kept
  21. Shot 20, 8:22 to 8:52, 0 of 1 keyframes kept
  22. Shot 21, 8:52 to 9:07, 1 of 1 keyframes kept
  23. Shot 22, 9:07 to 9:42, 1 of 1 keyframes kept
  24. Shot 23, 9:42 to 10:10, 1 of 1 keyframes kept
  25. Shot 24, 10:10 to 10:37, 0 of 1 keyframes kept
  26. Shot 25, 10:37 to 11:05, 0 of 1 keyframes kept
  27. Shot 26, 11:05 to 11:32, 0 of 1 keyframes kept
  28. Shot 27, 11:32 to 12:00, 0 of 1 keyframes kept
  29. Shot 28, 12:00 to 12:01, 1 of 1 keyframes kept
  30. Shot 29, 12:01 to 12:38, 1 of 1 keyframes kept
  31. Shot 30, 12:38 to 12:53, 0 of 1 keyframes kept
  32. Shot 31, 12:53 to 13:18, 1 of 1 keyframes kept
  33. Shot 32, 13:18 to 13:43, 0 of 1 keyframes kept
  34. Shot 33, 13:43 to 14:08, 0 of 1 keyframes kept
  35. Shot 34, 14:08 to 14:33, 0 of 1 keyframes kept
  36. Shot 35, 14:33 to 14:58, 1 of 1 keyframes kept
  37. Shot 36, 14:58 to 15:00, 1 of 1 keyframes kept
  38. Shot 37, 15:00 to 15:16, 1 of 1 keyframes kept
  39. Shot 38, 15:16 to 15:45, 1 of 1 keyframes kept
  40. Shot 39, 15:45 to 16:13, 0 of 1 keyframes kept
  41. Shot 40, 16:13 to 16:41, 0 of 1 keyframes kept
  42. Shot 41, 16:41 to 16:45, 1 of 1 keyframes kept
  43. Shot 42, 16:45 to 16:46, 1 of 1 keyframes kept
  44. Shot 43, 16:46 to 17:15, 1 of 1 keyframes kept
  45. Shot 44, 17:15 to 17:45, 0 of 1 keyframes kept
  46. Shot 45, 17:45 to 18:14, 0 of 1 keyframes kept
  47. Shot 46, 18:14 to 18:18, 1 of 1 keyframes kept
  48. Shot 47, 18:18 to 18:32, 1 of 1 keyframes kept
  49. Shot 48, 18:32 to 18:33, 0 of 1 keyframes kept

49 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
191
whisperx 191
chunks
32
from 191 cues
keyframes
28
kept of 49 captured
frames with text
28
767 lines read
chapters
0
from the source metadata
keyframe bytes
5.1 MB
word timings on 191 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 18:23 1m 58s
stt done 2026-08-10 18:25 20s
chunk done 2026-08-10 18:25 0s
text_embed done 2026-08-10 19:55 0s
keyframe done 2026-08-10 18:25 1m 31s
ocr done 2026-08-10 18:27 13s
frame_embed done 2026-08-10 19:55 5s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:31 #3 done16 line(s)

    shot 3·sharpness 1289.7

    1. # Braintrust0.95
    2. AIE BT Expo Session t T..0.97
    3. Join Meeting1.00
    4. ***0.67
    5. AIE1.00
    6. 0.94
    7. 1.00
    8. 1.00
    9. 1.00
    10. The maturity phases of1.00
    11. running evals1.00
    12. Ship quality AI0.97
    13. braintrust.dev1.00
    14. Braintrust1.00
    15. Engineering the future of Al1.00
    16. AlEngir0.98
  • 0:51 #4 done17 line(s)

    shot 4·sharpness 1645.8

    1. Agenda1.00
    2. AIE BT Expo Session t T..0.98
    3. Join Meeting0.98
    4. Intro1.00
    5. Overview1.00
    6. *★★0.61
    7. AIE1.00
    8. 0.96
    9. 1.00
    10. 1.00
    11. 1.00
    12. Different stages of eval platform builds0.99
    13. What's next1.00
    14. Confidential1.00
    15. Braintrust1.00
    16. WorkOS OpenAI0.92
    17. AlEngir0.93
  • 1:41 #5 done21 line(s)

    shot 5·sharpness 3399.4

    1. #Braintrust0.97
    2. AIE BT Expo Session t T..0.94
    3. Join Meeting0.98
    4. Phil Hetzel1.00
    5. Head of Solution Engineering, Braintrust0.99
    6. 1.00
    7. - Twelve years in consulting/implementation0.99
    8. AIE1.00
    9. Former leader of Slalom's global Databricks0.99
    10. business unit0.98
    11. 1.00
    12. 1.00
    13. Likes to play chess (poorly) and spend time0.98
    14. with wife and dachshund (Pistol Pete,0.99
    15. pictured)1.00
    16. The maturity phases of1.00
    17. running evals1.00
    18. Ship quality AI1.00
    19. braintrust.dev1.00
    20. AlEngineer0.97
    21. EUROPE1.00
  • 2:12 #6 done19 line(s)

    shot 6·sharpness 3389.9

    1. #Braintrust0.97
    2. AIE BT Expo Session t: ..0.95
    3. Join Meeting1.00
    4. Phil Hetzel1.00
    5. Head of Solution Engineering, Braintrust0.99
    6. - Twelve years in consulting/implementation0.99
    7. AIE1.00
    8. Former leader of Slalom's global Databricks1.00
    9. business unit1.00
    10. 1.00
    11. 0.99
    12. Likes to play chess (poorly) and spend time0.96
    13. with wife and dachshund (Pistol Pete,1.00
    14. pictured)1.00
    15. The maturity phases of1.00
    16. running evals1.00
    17. Ship quality AI0.99
    18. braintrust.dev1.00
    19. Engineering the future of Al0.99
  • 2:40 #7 done40 line(s)

    shot 7·sharpness 3545.4

    1. What is Braintrust?1.00
    2. Agents fail in un0.99
    3. AIE BT Expo Session t T..0.98
    4. Evals and observ0.99
    5. before users do.1.00
    6. How do you know your AI feature works?1.00
    7. AI IN YOUR APP1.00
    8. SCORES1.00
    9. Eval test your AI with real data and score the results.1.00
    10. AI0.85
    11. 98% Toxicity1.00
    12. You can determine whether the results improve or hurt1.00
    13. 83% Accuracy0.99
    14. 1.00
    15. 1.00
    16. AIE1.00
    17. 1.00
    18. performance.1.00
    19. ×0.58
    20. 74% Hallucination0.99
    21. 1.00
    22. 1.00
    23. Are bad responses reaching users?1.00
    24. 1.00
    25. 1.00
    26. Production monitoring tracks live model responses and1.00
    27. alerts you when quality drops or incorrect outputs1.00
    28. increase.1.00
    29. Can your team improve quality without1.00
    30. ×0.95
    31. ×0.75
    32. guesswork?1.00
    33. PROMPT A1.00
    34. PROMPT B1.00
    35. PROMPT C1.00
    36. Loop builds and refines scorers to measure the specific0.99
    37. quality metrics that matter for your application.0.99
    38. Confidential1.00
    39. Engineering the future of Al1.00
    40. AlEng0.98
  • 2:53 #8 skipped

    shot 8·duplicate of #7

  • 3:22 #9 done29 line(s)

    shot 9·sharpness 3924.7

    1. Grounding the discussion1.00
    2. AIE BT Expo Session t ..0.96
    3. North star ideas1.00
    4. An eval is:1.00
    5. Why do we do evals?1.00
    6. Task (the thing you are0.99
    7. Wholly in service to agent quality.1.00
    8. evaluating)1.00
    9. 1.00
    10. 1.00
    11. AIE1.00
    12. Are evals unit tests?1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. No, evals are rerunning production on known inputs1.00
    17. would initiate an agent1.00
    18. Data (an array of inputs that1.00
    19. Do eval results need to be perfect?0.99
    20. invocation)1.00
    21. No, and it'd be surprising if they were0.98
    22. Scorers (the functions that0.99
    23. What should we eval?1.00
    24. judge the result and0.99
    25. Start here: how would a humanjudge quality?1.00
    26. execution of a Task)0.97
    27. Confidential1.00
    28. AlEngineer0.97
    29. EUROPE1.00
  • 3:48 #10 skipped

    shot 10·duplicate of #9

  • 4:31 #11 done26 line(s)

    shot 11·sharpness 3859.6

    1. Grounding the discussion1.00
    2. North star ideas1.00
    3. An eval is:0.96
    4. Why do we do evals?1.00
    5. Task (the thing you are1.00
    6. Wholly in service to agent quality.1.00
    7. evaluating)1.00
    8. AIE1.00
    9. Are evals unit tests?1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. No, evals are rerunning production on known inputs0.99
    14. Data (an array of inputs that1.00
    15. would initiate an agent1.00
    16. Do eval results need to be perfect?0.99
    17. invocation)1.00
    18. No, and it'd be surprising if they were0.98
    19. Scorers (the functions that0.99
    20. What should we eval?1.00
    21. judge the result and1.00
    22. Start here: how would a human judge quality?0.99
    23. execution of a Task)0.97
    24. Confidential1.00
    25. Engineering the future of Al0.99
    26. AlEngir0.97
  • 4:54 #12 skipped

    shot 12·duplicate of #11

  • 5:09 #13 skipped

    shot 13·duplicate of #11

  • 5:42 #14 skipped

    shot 14·duplicate of #11

  • 6:18 #15 done23 line(s)

    shot 15·sharpness 2992.2

    1. Your eval technique0.98
    2. will mature0.97
    3. Level 0 - Just get started0.99
    4. Eval techniques will grow with the0.99
    5. level of maturity of the team building0.99
    6. agents1.00
    7. AIE1.00
    8. The more complex your agent, the1.00
    9. Level 1 - Measure to manage1.00
    10. more vectors for failure. The more0.98
    11. 1.00
    12. 1.00
    13. vectors for failure, the more complex1.00
    14. your scoring mechanisms1.00
    15. We'll focus only on evals and not the0.99
    16. platform surrounding evals in this1.00
    17. Level 2 - Accounting for1.00
    18. session1.00
    19. complexity1.00
    20. Level 3 - Advanced evals0.99
    21. Confidential1.00
    22. Engineering the future of Al1.00
    23. AlEng0.97
  • 6:30 #16 skipped

    shot 16·duplicate of #15

  • 7:20 #17 done29 line(s)

    shot 17·sharpness 3838.0

    1. Level 0 - Just get0.97
    2. started1.00
    3. •It's not wrong to start with vibes0.99
    4. Task1.00
    5. input1.00
    6. Data1.00
    7. Have human annotators ask "is this0.99
    8. Probably something simple, a1.00
    9. A short human generated or0.99
    10. good or bad?"1.00
    11. prompt or workflow1.00
    12. synthesized dataset1.00
    13. 1.00
    14. You should document why your1.00
    15. AIE1.00
    16. human annotators are choosing good0.99
    17. 1.00
    18. or bad1.00
    19. output1.00
    20. 1.00
    21. 1.00
    22. These these will eventually1.00
    23. become your scorers1.00
    24. Scorers1.00
    25. Is this good or bad? Human1.00
    26. annotated0.95
    27. Confidential1.00
    28. Google DeepMind1.00
    29. Al Engir0.95
  • 7:50 #18 done30 line(s)

    shot 18·sharpness 3933.3

    1. Level 0 - Just get0.95
    2. started1.00
    3. It's not wrong to start with vibes1.00
    4. Task1.00
    5. input1.00
    6. Data1.00
    7. Have human annotators ask "is this0.99
    8. Probably something simple, a1.00
    9. A short human generated or1.00
    10. good or bad?"1.00
    11. prompt or workflow1.00
    12. synthesized dataset0.97
    13. 1.00
    14. You should document why your0.99
    15. AIE1.00
    16. human annotators are choosing good0.99
    17. 1.00
    18. or bad1.00
    19. output1.00
    20. 1.00
    21. 1.00
    22. 1.00
    23. These these will eventually0.99
    24. become your scorers1.00
    25. Scorers1.00
    26. Is this good or bad? Human1.00
    27. annotated0.97
    28. Confidential1.00
    29. Braintrust1.00
    30. WorkOS OpenAI0.96
  • 8:02 #19 skipped

    shot 19·duplicate of #17

  • 8:26 #20 skipped

    shot 20·duplicate of #17

  • 9:04 #21 done15 line(s)

    shot 21·sharpness 2662.4

    1. Level 0 - Just get0.97
    2. started1.00
    3. AIE1.00
    4. Trace1.00
    5. 1.00
    6. 1.00
    7. Justification1.00
    8. Human grader1.00
    9. LLM as judge1.00
    10. LLM / coding0.98
    11. scorers1.00
    12. agent1.00
    13. Confidential1.00
    14. AlEngineer0.97
    15. EUROPE1.00
  • 9:15 #22 done67 line(s)

    shot 22·sharpness 2378.2

    1. Level 0 - Just get0.96
    2. started1.00
    3. Braintrust Demos I8r-customer-service0.98
    4. Experiments0.98
    5. •18r-sim-eval-2026-01-25-06dcf6d10.96
    6. Diff注Revlew0.84
    7. Share0.98
    8. Comparisons0.99
    9. Alt experiment rows view0.97
    10. X0.59
    11. 4388:3540.99
    12. Related Tag % Score Find0.91
    13. è share0.92
    14. 18r-sin-oval-2026-01-25-40.87
    15. rE Trace0.83
    16. Timeline0.93
    17. Thread0.92
    18. Views0.99
    19. Prompt0.92
    20. fame0.87
    21. CONVERSATION1.00
    22. Ad Justitication0.89
    23. 1.00
    24. 1.00
    25. Dataset0.94
    26. 1.00
    27. AIE1.00
    28. L8rCustomerServiceDataset0.97
    29. © eval0.84
    30. eval0.71
    31. Would you like me to pause this payment plan for you?0.99
    32. 1.00
    33. 1.00
    34. None0.87
    35. © eval0.87
    36. 1.00
    37. 1.00
    38. 1.00
    39. % AI0.74
    40. Scorers and distribution0.99
    41. ① eval0.82
    42. ©eval0.86
    43. © eval0.91
    44. (Turns: , ume: frustrated, Goat pause m Best Buy order plan (Personality: direct)0.87
    45. FINAL RESPONISE0.95
    46. % Trace scorers0.99
    47. 0 eval0.85
    48. © eval0.84
    49. USER: I want more favorable payment terms0.99
    50. ASSISTANT: To help you with more favorable payment terms,I can ook into your current0.95
    51. installment plans and see what we can modily. We might be able to pause a payment,0.99
    52. eval0.95
    53. reschedule it, or even explore other options based on your situation.0.99
    54. I0.58
    55. ©eval0.85
    56. Would you lke me to check your active installment plans?0.97
    57. USER: Yes, check my active installment plans.1.00
    58. % GoalAchievement0.94
    59. Braintrust1.00
    60. ASSISTANT: It loks ike you don' curently have any active installment plans. you have any0.88
    61. payment concerns or if there's anything else I can help you with, pleae let me know!0.98
    62. % QualityCheck0.96
    63. USER: I want to pause my Best Buy order plan, not just check installment plans.0.98
    64. QUERES 250.97
    65. Confidential1.00
    66. Engineering the future of Al0.99
    67. AlEng0.99
  • 9:56 #23 done32 line(s)

    shot 23·sharpness 4155.5

    1. Level 1 - Measure to1.00
    2. manage1.00
    3. You have the justifications for what1.00
    4. Task1.00
    5. input1.00
    6. Data1.00
    7. makes a result/execution good or1.00
    8. An agent1.00
    9. Input examples from production1.00
    10. bad; use it.1.00
    11. traces1.00
    12. 1.00
    13. •For subjective failure modes, use0.98
    14. AIE1.00
    15. LLMs.1.00
    16. 1.00
    17. 1.00
    18. For objective failure modes, use1.00
    19. input, expected1.00
    20. 1.00
    21. 1.00
    22. code.1.00
    23. output1.00
    24. The team should be using real1.00
    25. examples from production at this1.00
    26. stage as eval inputs.1.00
    27. Scorers1.00
    28. Built from failure modes, LLM as1.00
    29. automated moman0.93
    30. judge or deterministic code0.98
    31. Confidential1.00
    32. Engineering the future of Al1.00

Transcript

191 cues· 2,825 words· 15,488 chars

  1. 0:15 Um, it's always a challenge to be a presenter directly after lunch because that's typically when the energy level goes from right around here to around here.
  2. 0:24 But I'm gonna try to make this session worth your while, uh, today.
  3. 0:27 We've got 18 very quick minutes, uh, together.
  4. 0:31 And, uh, during that time, I'm gonna be talking about the, uh, different maturity levels that I see people go through as they perform evals for their agents.
  5. 0:41 Before we get into that, just roughly, quick agenda today.
  6. 0:46 I'll explain a little bit about myself, the company that I work for.
  7. 0:49 We'll spend most of the time today on more theoretical concepts, not product concepts.
  8. 0:54 And then we'll talk about where I think this field is going in the future.
  9. 1:00 I'll also make sure to leave enough time, hopefully a couple minutes, for questions as well.
  10. 1:05 I didn't over-prepare the content in hopes that we could have a little bit more of a discussion at the end of this.
  11. 1:11 First of all, this is me.
  12. 1:13 My name is Phil Hetzel.
  13. 1:15 I lead solutions engineering for a company called Braintrust.
  14. 1:18 Effectively, what that means is that it is me and my team's job to make sure that people are getting the most value out of the platform as quickly as possible.
  15. 1:28 Prior to Braintrust, I spent 12 years in consulting and systems implementation.
  16. 1:33 First four years with KPMG, last eight years in consulting with a company called Slalom Consulting.
  17. 1:40 And with Slalom, I led their global Databricks business unit.
  18. 1:43 And I noticed that a lot of my customers were prolific at creating generative AI proofs of concepts.
  19. 1:50 They were not as prolific at bringing those proofs of concepts to production.
  20. 1:53 So I started using Braintrust first as a user.
  21. 1:57 because I wanted to help bridge that gap for my customers.
  22. 2:01 And I liked the product so much that I ended up joining the company, and I've been here for about a year.
  23. 2:07 Outside of work, I like to play chess, but I'm not very good at it.
  24. 2:11 And I like to spend time with my wife and my dachshund.
  25. 2:14 His name's Pistol Pete.
  26. 2:16 He's the one in brown, not the one in black.
  27. 2:20 What is Braintrust, the company that I work for?
  28. 2:23 Braintrust is an agent quality company.
  29. 2:26 I guess two of the main ways that we contribute to agent quality are evals and observability, which we consider to be very much the same problem from a systems perspective.
  30. 2:39 evals of course being the thing that you're doing in order to gain confidence in your agent as you want to bring it to production, and then observability being the practice of once that agent is in production, remaining confident in it.
  31. 2:56 It's a growing space, it's a very fast moving space.
  32. 2:59 And when you build an eval's platform, you really have to grow with the technology, the underlying technology as it changes.
  33. 3:06 So it's a very fun place to be in.
  34. 3:11 Let me give a quick overview of the problem.
  35. 3:14 We talked a little bit about why we do evals in the first place.
  36. 3:18 How many of you all are doing evals today, hopefully, as you build?
  37. 3:23 Every single hand should be up.
  38. 3:24 And certainly, when I give this talk next year at this conference, all of you are going to come back, of course, to this session, and every hand is going to be up.
  39. 3:32 Eval is very important.
  40. 3:33 The reason why we do evals is wholly in service to agent quality.
  41. 3:37 That's the most important thing.
  42. 3:39 We want to make sure that our agents are doing what we expect when confronted with real usage and real users.
  43. 3:48 This is really important from a risk perspective and a brand perspective.
  44. 3:52 We don't want the reputational risk of an agent being unkind or unhelpful to a customer.
  45. 3:59 We don't want the systems risk of an agent costing us too much money as it operates.
  46. 4:05 And there could even be compliance and legal risks if your agent goes too far off the rails.
  47. 4:10 So evals are both a defense against those types of risks,
  48. 4:14 But they're also, uh, they can play offense with evals in knowing with each tweak that you make to your agent, how it's improving and how much it's improving your application.
  49. 4:26 Um, a couple of primitives here, evals are not unit tests where- whereas unit tests are very exhaustive in- in how you perform them.
  50. 4:35 With evals, you want to make sure that you start very high level with the failure modes of your agent.

Open at this second