Videos FB-MLPhL9Ms
The maturity phases of running evals — Phil Hetzel, Braintrust
Scene timeline
49 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 191
- whisperx 191
- chunks
- 32
- from 191 cues
- keyframes
- 28
- kept of 49 captured
- frames with text
- 28
- 767 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.1 MB
- word timings on 191 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 18:23 | 1m 58s |
stt |
done | — | 2026-08-10 18:25 | 20s |
chunk |
done | — | 2026-08-10 18:25 | 0s |
text_embed |
done | — | 2026-08-10 19:55 | 0s |
keyframe |
done | — | 2026-08-10 18:25 | 1m 31s |
ocr |
done | — | 2026-08-10 18:27 | 13s |
frame_embed |
done | — | 2026-08-10 19:55 | 5s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- # Braintrust0.95
- AIE BT Expo Session t T..0.97
- Join Meeting1.00
- ***0.67
- AIE1.00
- ★0.94
- ★1.00
- ★1.00
- ★1.00
- The maturity phases of1.00
- running evals1.00
- Ship quality AI0.97
- braintrust.dev1.00
- Braintrust1.00
- Engineering the future of Al1.00
- AlEngir0.98
-
- Agenda1.00
- AIE BT Expo Session t T..0.98
- Join Meeting0.98
- Intro1.00
- Overview1.00
- *★★0.61
- AIE1.00
- ★0.96
- ★1.00
- ★1.00
- ★1.00
- Different stages of eval platform builds0.99
- What's next1.00
- Confidential1.00
- Braintrust1.00
- WorkOS OpenAI0.92
- AlEngir0.93
-
- #Braintrust0.97
- AIE BT Expo Session t T..0.94
- Join Meeting0.98
- Phil Hetzel1.00
- Head of Solution Engineering, Braintrust0.99
- ★1.00
- - Twelve years in consulting/implementation0.99
- AIE1.00
- Former leader of Slalom's global Databricks0.99
- business unit0.98
- ★1.00
- ★1.00
- Likes to play chess (poorly) and spend time0.98
- with wife and dachshund (Pistol Pete,0.99
- pictured)1.00
- The maturity phases of1.00
- running evals1.00
- Ship quality AI1.00
- braintrust.dev1.00
- AlEngineer0.97
- EUROPE1.00
-
- #Braintrust0.97
- AIE BT Expo Session t: ..0.95
- Join Meeting1.00
- Phil Hetzel1.00
- Head of Solution Engineering, Braintrust0.99
- - Twelve years in consulting/implementation0.99
- AIE1.00
- Former leader of Slalom's global Databricks1.00
- business unit1.00
- ★1.00
- ★0.99
- Likes to play chess (poorly) and spend time0.96
- with wife and dachshund (Pistol Pete,1.00
- pictured)1.00
- The maturity phases of1.00
- running evals1.00
- Ship quality AI0.99
- braintrust.dev1.00
- Engineering the future of Al0.99
-
- What is Braintrust?1.00
- Agents fail in un0.99
- AIE BT Expo Session t T..0.98
- Evals and observ0.99
- before users do.1.00
- How do you know your AI feature works?1.00
- AI IN YOUR APP1.00
- SCORES1.00
- Eval test your AI with real data and score the results.1.00
- AI0.85
- 98% Toxicity1.00
- You can determine whether the results improve or hurt1.00
- 83% Accuracy0.99
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- performance.1.00
- ×0.58
- 74% Hallucination0.99
- ★1.00
- ★1.00
- Are bad responses reaching users?1.00
- ★1.00
- ★1.00
- Production monitoring tracks live model responses and1.00
- alerts you when quality drops or incorrect outputs1.00
- increase.1.00
- Can your team improve quality without1.00
- ×0.95
- ×0.75
- guesswork?1.00
- PROMPT A1.00
- PROMPT B1.00
- PROMPT C1.00
- Loop builds and refines scorers to measure the specific0.99
- quality metrics that matter for your application.0.99
- Confidential1.00
- Engineering the future of Al1.00
- AlEng0.98
-
- Grounding the discussion1.00
- AIE BT Expo Session t ..0.96
- North star ideas1.00
- An eval is:1.00
- Why do we do evals?1.00
- Task (the thing you are0.99
- Wholly in service to agent quality.1.00
- evaluating)1.00
- ★1.00
- ★1.00
- AIE1.00
- Are evals unit tests?1.00
- ★1.00
- ★1.00
- ★1.00
- No, evals are rerunning production on known inputs1.00
- would initiate an agent1.00
- Data (an array of inputs that1.00
- Do eval results need to be perfect?0.99
- invocation)1.00
- No, and it'd be surprising if they were0.98
- Scorers (the functions that0.99
- What should we eval?1.00
- judge the result and0.99
- Start here: how would a humanjudge quality?1.00
- execution of a Task)0.97
- Confidential1.00
- AlEngineer0.97
- EUROPE1.00
-
- Grounding the discussion1.00
- North star ideas1.00
- An eval is:0.96
- Why do we do evals?1.00
- Task (the thing you are1.00
- Wholly in service to agent quality.1.00
- evaluating)1.00
- AIE1.00
- Are evals unit tests?1.00
- ★1.00
- ★1.00
- ★1.00
- No, evals are rerunning production on known inputs0.99
- Data (an array of inputs that1.00
- would initiate an agent1.00
- Do eval results need to be perfect?0.99
- invocation)1.00
- No, and it'd be surprising if they were0.98
- Scorers (the functions that0.99
- What should we eval?1.00
- judge the result and1.00
- Start here: how would a human judge quality?0.99
- execution of a Task)0.97
- Confidential1.00
- Engineering the future of Al0.99
- AlEngir0.97
-
- Your eval technique0.98
- will mature0.97
- Level 0 - Just get started0.99
- Eval techniques will grow with the0.99
- level of maturity of the team building0.99
- agents1.00
- AIE1.00
- The more complex your agent, the1.00
- Level 1 - Measure to manage1.00
- more vectors for failure. The more0.98
- ★1.00
- ★1.00
- vectors for failure, the more complex1.00
- your scoring mechanisms1.00
- We'll focus only on evals and not the0.99
- platform surrounding evals in this1.00
- Level 2 - Accounting for1.00
- session1.00
- complexity1.00
- Level 3 - Advanced evals0.99
- Confidential1.00
- Engineering the future of Al1.00
- AlEng0.97
-
- Level 0 - Just get0.97
- started1.00
- •It's not wrong to start with vibes0.99
- Task1.00
- input1.00
- Data1.00
- Have human annotators ask "is this0.99
- Probably something simple, a1.00
- A short human generated or0.99
- good or bad?"1.00
- prompt or workflow1.00
- synthesized dataset1.00
- ★1.00
- You should document why your1.00
- AIE1.00
- human annotators are choosing good0.99
- ★1.00
- or bad1.00
- output1.00
- ★1.00
- ★1.00
- These these will eventually1.00
- become your scorers1.00
- Scorers1.00
- Is this good or bad? Human1.00
- annotated0.95
- Confidential1.00
- Google DeepMind1.00
- Al Engir0.95
-
- Level 0 - Just get0.95
- started1.00
- It's not wrong to start with vibes1.00
- Task1.00
- input1.00
- Data1.00
- Have human annotators ask "is this0.99
- Probably something simple, a1.00
- A short human generated or1.00
- good or bad?"1.00
- prompt or workflow1.00
- synthesized dataset0.97
- ★1.00
- You should document why your0.99
- AIE1.00
- human annotators are choosing good0.99
- ★1.00
- or bad1.00
- output1.00
- ★1.00
- ★1.00
- ★1.00
- These these will eventually0.99
- become your scorers1.00
- Scorers1.00
- Is this good or bad? Human1.00
- annotated0.97
- Confidential1.00
- Braintrust1.00
- WorkOS OpenAI0.96
-
- Level 0 - Just get0.97
- started1.00
- AIE1.00
- Trace1.00
- ★1.00
- ★1.00
- Justification1.00
- Human grader1.00
- LLM as judge1.00
- LLM / coding0.98
- scorers1.00
- agent1.00
- Confidential1.00
- AlEngineer0.97
- EUROPE1.00
-
- Level 0 - Just get0.96
- started1.00
- Braintrust Demos I8r-customer-service0.98
- Experiments0.98
- •18r-sim-eval-2026-01-25-06dcf6d10.96
- Diff注Revlew0.84
- Share0.98
- Comparisons0.99
- Alt experiment rows view0.97
- X0.59
- 4388:3540.99
- Related Tag % Score Find0.91
- è share0.92
- 18r-sin-oval-2026-01-25-40.87
- rE Trace0.83
- Timeline0.93
- Thread0.92
- Views0.99
- Prompt0.92
- fame0.87
- CONVERSATION1.00
- Ad Justitication0.89
- ★1.00
- ★1.00
- Dataset0.94
- ★1.00
- AIE1.00
- L8rCustomerServiceDataset0.97
- © eval0.84
- eval0.71
- Would you like me to pause this payment plan for you?0.99
- ★1.00
- ★1.00
- None0.87
- © eval0.87
- ★1.00
- ★1.00
- ★1.00
- % AI0.74
- Scorers and distribution0.99
- ① eval0.82
- ©eval0.86
- © eval0.91
- (Turns: , ume: frustrated, Goat pause m Best Buy order plan (Personality: direct)0.87
- FINAL RESPONISE0.95
- % Trace scorers0.99
- 0 eval0.85
- © eval0.84
- USER: I want more favorable payment terms0.99
- ASSISTANT: To help you with more favorable payment terms,I can ook into your current0.95
- installment plans and see what we can modily. We might be able to pause a payment,0.99
- eval0.95
- reschedule it, or even explore other options based on your situation.0.99
- I0.58
- ©eval0.85
- Would you lke me to check your active installment plans?0.97
- USER: Yes, check my active installment plans.1.00
- % GoalAchievement0.94
- Braintrust1.00
- ASSISTANT: It loks ike you don' curently have any active installment plans. you have any0.88
- payment concerns or if there's anything else I can help you with, pleae let me know!0.98
- % QualityCheck0.96
- USER: I want to pause my Best Buy order plan, not just check installment plans.0.98
- QUERES 250.97
- Confidential1.00
- Engineering the future of Al0.99
- AlEng0.99
-
- Level 1 - Measure to1.00
- manage1.00
- You have the justifications for what1.00
- Task1.00
- input1.00
- Data1.00
- makes a result/execution good or1.00
- An agent1.00
- Input examples from production1.00
- bad; use it.1.00
- traces1.00
- ★1.00
- •For subjective failure modes, use0.98
- AIE1.00
- LLMs.1.00
- ★1.00
- ★1.00
- For objective failure modes, use1.00
- input, expected1.00
- ★1.00
- ★1.00
- code.1.00
- output1.00
- The team should be using real1.00
- examples from production at this1.00
- stage as eval inputs.1.00
- Scorers1.00
- Built from failure modes, LLM as1.00
- automated moman0.93
- judge or deterministic code0.98
- Confidential1.00
- Engineering the future of Al1.00
Transcript
191 cues· 2,825 words· 15,488 chars
- 0:15 Um, it's always a challenge to be a presenter directly after lunch because that's typically when the energy level goes from right around here to around here.
- 0:24 But I'm gonna try to make this session worth your while, uh, today.
- 0:27 We've got 18 very quick minutes, uh, together.
- 0:31 And, uh, during that time, I'm gonna be talking about the, uh, different maturity levels that I see people go through as they perform evals for their agents.
- 0:41 Before we get into that, just roughly, quick agenda today.
- 0:46 I'll explain a little bit about myself, the company that I work for.
- 0:49 We'll spend most of the time today on more theoretical concepts, not product concepts.
- 0:54 And then we'll talk about where I think this field is going in the future.
- 1:00 I'll also make sure to leave enough time, hopefully a couple minutes, for questions as well.
- 1:05 I didn't over-prepare the content in hopes that we could have a little bit more of a discussion at the end of this.
- 1:11 First of all, this is me.
- 1:13 My name is Phil Hetzel.
- 1:15 I lead solutions engineering for a company called Braintrust.
- 1:18 Effectively, what that means is that it is me and my team's job to make sure that people are getting the most value out of the platform as quickly as possible.
- 1:28 Prior to Braintrust, I spent 12 years in consulting and systems implementation.
- 1:33 First four years with KPMG, last eight years in consulting with a company called Slalom Consulting.
- 1:40 And with Slalom, I led their global Databricks business unit.
- 1:43 And I noticed that a lot of my customers were prolific at creating generative AI proofs of concepts.
- 1:50 They were not as prolific at bringing those proofs of concepts to production.
- 1:53 So I started using Braintrust first as a user.
- 1:57 because I wanted to help bridge that gap for my customers.
- 2:01 And I liked the product so much that I ended up joining the company, and I've been here for about a year.
- 2:07 Outside of work, I like to play chess, but I'm not very good at it.
- 2:11 And I like to spend time with my wife and my dachshund.
- 2:14 His name's Pistol Pete.
- 2:16 He's the one in brown, not the one in black.
- 2:20 What is Braintrust, the company that I work for?
- 2:23 Braintrust is an agent quality company.
- 2:26 I guess two of the main ways that we contribute to agent quality are evals and observability, which we consider to be very much the same problem from a systems perspective.
- 2:39 evals of course being the thing that you're doing in order to gain confidence in your agent as you want to bring it to production, and then observability being the practice of once that agent is in production, remaining confident in it.
- 2:56 It's a growing space, it's a very fast moving space.
- 2:59 And when you build an eval's platform, you really have to grow with the technology, the underlying technology as it changes.
- 3:06 So it's a very fun place to be in.
- 3:11 Let me give a quick overview of the problem.
- 3:14 We talked a little bit about why we do evals in the first place.
- 3:18 How many of you all are doing evals today, hopefully, as you build?
- 3:23 Every single hand should be up.
- 3:24 And certainly, when I give this talk next year at this conference, all of you are going to come back, of course, to this session, and every hand is going to be up.
- 3:32 Eval is very important.
- 3:33 The reason why we do evals is wholly in service to agent quality.
- 3:37 That's the most important thing.
- 3:39 We want to make sure that our agents are doing what we expect when confronted with real usage and real users.
- 3:48 This is really important from a risk perspective and a brand perspective.
- 3:52 We don't want the reputational risk of an agent being unkind or unhelpful to a customer.
- 3:59 We don't want the systems risk of an agent costing us too much money as it operates.
- 4:05 And there could even be compliance and legal risks if your agent goes too far off the rails.
- 4:10 So evals are both a defense against those types of risks,
- 4:14 But they're also, uh, they can play offense with evals in knowing with each tweak that you make to your agent, how it's improving and how much it's improving your application.
- 4:26 Um, a couple of primitives here, evals are not unit tests where- whereas unit tests are very exhaustive in- in how you perform them.
- 4:35 With evals, you want to make sure that you start very high level with the failure modes of your agent.
loading