Videos xyL2Ltkh-SA
How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads
Scene timeline
47 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 216
- whisperx 216
- chunks
- 34
- from 216 cues
- keyframes
- 23
- kept of 47 captured
- frames with text
- 22
- 431 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.6 MB
- word timings on 216 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:48 | 0s |
stt |
done | — | 2026-08-09 07:08 | 21s |
chunk |
done | — | 2026-08-09 07:08 | 0s |
text_embed |
done | — | 2026-08-10 19:40 | 0s |
keyframe |
done | — | 2026-08-09 07:08 | 2m 28s |
ocr |
done | — | 2026-08-09 07:11 | 13s |
frame_embed |
done | — | 2026-08-10 19:40 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.99
-
- Ask Oemini0.97
- AWork0.88
- Finish update I0.95
- AlEngineer0.97
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- Model Whisperers1.00
- Evals for Production Agents1.00
- *all opinions are my own and not of my employer1.00
- Engineering the future of Al0.98
- World'sFair1.00
-
- Ask Oemini0.97
- Work0.96
- Finish update 10.96
- AlEngineer0.98
- World'sFair1.00
- PRESENTED BY0.99
- Microsoft1.00
- Building an AI agent is hard1.00
- Engineering the future of Al0.99
- World'sFair1.00
-
- Ask Oemini0.96
- docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
- OWork0.84
- Finish update i0.93
- AlEngineer0.96
- World'sFair1.00
- Agent foundation1.00
- PRESENTED BY1.00
- Microsoft1.00
- A focused set of strong, optimized LLM-friendly tools1.00
- provides a good foundation1.00
- Once your tools are optimized (important), an independent1.00
- critique agent with a remediation loop can fill more gaps1.00
- by providing a self-correction mechanism1.00
- Once your base structure is defined, a strong eval provides0.98
- a method for proving value and ablation - an essential tool0.99
- for climbing the quality ladder1.00
- TRACK 5· JULY 1,20260.97
- World'sFair1.00
- Evals1.00
-
- Ask Oemini0.95
- docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
- OWork0.82
- Finish update !0.91
- AlEngineer0.97
- World'sFair1.00
- Agent foundation1.00
- A focused set of strong, optimized LLM-friendly tools1.00
- provides a good foundation1.00
- Once your tools are optimized (important), an independent1.00
- critique agent with a remediation loop can fill more gaps0.98
- by providing a self-correction mechanism1.00
- Once your base structure is defined, a strong eval provides0.99
- a method for proving value and ablation - an essential tool1.00
- for climbing the quality ladder1.00
- World'sFair1.00
- TRACK 5· JULY 1,20260.96
- Evals1.00
-
- Ask Gemini0.99
- docs.google.com/presentation/d/lesaioTRybTIWMFD2JM7UUL_g7HCfLxbidbE6GET7k/edit?resourcekey=0-ihtkJI0R17HJiBOCBECQeg&slide=id.g311...☆0.91
- Work0.89
- Finish update I0.94
- AlEngineer0.98
- World'sFair1.00
- Reliability = f(agent capability, guardrails, evals)1.00
- Understanding what agents do in the0.99
- Defining good1.00
- real world1.00
- Evals allows us to understand and0.99
- Gen AI outputs are non-deterministic0.99
- improve how modes behave in the real0.99
- world by defining what "good" looks0.96
- We can't guarantee how an agent will0.99
- like.1.00
- behave in the wild without measuring0.99
- it at scale.1.00
- To build evals that actually scale,1.00
- they must be strict and measurable.0.99
- World'sFair1.00
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- Ask Gemini0.93
- .com/presentation/d/lesaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJI0Rt7HJiBOCBECQeg&slide=id.g3f..0.97
- Work0.93
- Finish update I0.97
- AlEngineer0.99
- World's Fair0.99
- Vibing can be good for you (early on!)1.00
- PRESENTED BY1.00
- 801.00
- Microsoft1.00
- Intuition based validations can be1.00
- 701.00
- good even if non-scalable0.99
- Pass te%0.74
- 601.00
- Prompt tweaks can have large1.00
- 501.00
- performance gains0.97
- 401.00
- hillclimb in a targeted way0.99
- Learn the failure patterns and1.00
- 301.00
- The Eval Rollercoaster1.00
- Jumping to scaled raters too early led0.99
- to many ups/downs in the eval process1.00
- and slowed us down1.00
- World's Fair0.98
- TRACK 5· JULY 1,20260.95
- Evals1.00
-
- Ask Gemini0.95
- docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
- 日0.79
- Work0.94
- Finish update I0.97
- AlEngineer0.99
- World's Fair0.98
- Start early, start small0.98
- PRESENTED BY0.99
- HOW EVALS GET DONE0.97
- Microsoft1.00
- Don't need to build a massive golden1.00
- WRITING1.00
- set on Day 1; start with a few core0.99
- THE EVALS0.98
- tasks1.00
- Test the negatives - checking if the0.99
- model didn't do something bad is just0.99
- as critical as checking if it did the0.97
- task1.00
- HUMANS ARGUING1.00
- OVER THE RUBRIC1.00
- World's Fair0.99
- TRACK 5·JULY 1,20260.98
- Evals1.00
-
- Ask Gemini0.94
- 日0.57
- work0.94
- Finish update I0.96
- AlEngineer0.99
- World's Fair0.99
- Get outside help0.98
- How do we evaluate0.98
- We use another1.00
- the model?1.00
- model to grade it.1.00
- But how do we1.00
- It's just models grading0.98
- evaluate that model?1.00
- models all the way down.0.99
- World's Fair1.00
- TRACK 5 · JULY 1, 20260.92
- Evals1.00
-
- Ask Gemini0.98
- docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f0..0.98
- ☆0.56
- Work0.89
- Finish update I0.96
- AlEngineer0.99
- World's Fair0.97
- Working with scaled raters1.00
- Clear rubric template with trusted0.99
- examples1.00
- Human-human agreement should be1.00
- Corporate needs-you to find the differences0.97
- strong1.00
- between this picture and this picture.0.99
- Get explanations along with verdict0.99
- O0.57
- Categorical input1.00
- pass/fail1.00
- Multi-output1.00
- is it accurate?1.00
- is it brand safe?0.98
- Explanations can be used to1.00
- make the agent better1.00
- They're the same picture.0.98
- World's Fair0.97
- TRACK 5·JULY 1,20260.97
- Evals1.00
Transcript
216 cues· 3,345 words· 17,780 chars
- 0:12 Hi, everyone.
- 0:15 Sounds like everybody came back from lunch, so I hope everybody is recharged and not sleepy at all.
- 0:20 It's always interesting to do a talk right after lunch, because you never know, it's a mixed crowd.
- 0:24 But we're very happy to be here, happy to see you all.
- 0:27 Our talk is gonna be about evals, of course.
- 0:29 We're in the evals track.
- 0:30 We're gonna talk you through what are some things that worked for us while we were building evals, especially for YouTube ads.
- 0:36 We work on the YouTube ads team as part of the, we do image and video models for YouTube ads.
- 0:42 So building an agent is hard.
- 0:44 I think anybody who is here in the audience probably has built an agent as a side project or as part of production systems.
- 0:50 It's a very hard thing to do.
- 0:52 It's laborious.
- 0:52 It takes a lot of time.
- 0:55 Making it reliable is harder.
- 0:57 So having it do things that you actually want it to do in production, understanding the different kind of things that it can play with, how it's going to react when you launch it to your end users, that's always a very hard thing to do, which is why evals are a pretty handy way to manage that.
- 1:15 Yeah, and then basically the first step when you're doing this is, of course, you need to have your agent foundation.
- 1:23 So when you're building your agent, you will want to have a focused and strong set of LLM-friendly tools to give your agent a very good foundation.
- 1:35 So yeah, I would say it's important to first optimize these tools and make sure they're the best they can be before just jumping onto
- 1:44 larger agent evals.
- 1:46 So once your tools are optimized, you can also take some other steps like making an independent critique agent with a remediation loop.
- 1:57 And this can fill more gaps as far as having a self-correction mechanism and filling those gaps where maybe your base toolset has limitations.
- 2:09 And then once your base structure is defined, you can have
- 2:13 You can then go to having an eval and having a strong eval is very important as this gives you like a way of proving the value of changes you make as well as running ablation experiments on any changes you make.
- 2:27 So I would say this is a very essential tool for climbing the quality ladder.
- 2:31 But again, it's very important to have that good foundation to begin with.
- 2:39 So yeah, the reliability of your agent is basically a function of the capabilities of the agent, the guardrails and the evals.
- 2:52 So understanding what your agents do in the real world, basically generative AI outputs, as I'm sure you're all familiar, are not exactly deterministic, right?
- 3:04 So it can often fail in certain areas or one time it can succeed, one time it can fail.
- 3:11 So we can't really guarantee how it will behave in the wild.
- 3:13 And for some use cases, this is extremely important, right?
- 3:17 And we need a way to measure at scale and make sure that
- 3:21 It is getting the output we want despite the non-determinism of these models.
- 3:28 So we need to define what's good here.
- 3:31 And evals allow us to basically understand and improve how the models behave in the real world by defining what good looks like.
- 3:43 It's basically just setting this is our target output.
- 3:49 So to build evals that actually scale, they really need to be strict and measurable.
- 3:58 And so an interesting thing here that I think might be somewhat counterintuitive is that early on vibing can actually be kind of good for you.
- 4:08 And what I mean here by vibing is basically doing things that are not exactly scalable to begin with.
- 4:15 So when you're first starting out, it may be that, you know, you could
- 4:23 take a track of basically just going ahead and making the super comprehensive eval.
- 4:28 But we found it actually works better to first do intuition-based approach, where you kind of first see the capabilities and look at the outputs.
- 4:37 And at this stage, it's pretty easy to tell what the issues actually are.
- 4:42 So even though this is non-scalable, it will still give you a very good idea of when you change this, what happens.
- 4:50 And it allows you to more quickly iterate as well.
- 4:54 So at this stage, prompt tweets can also have large performance gains.
- 4:58 You can make a radical change to the architecture.
- 5:01 And your eval is not hindering you in this way.
- 5:04 So it's a very good way.
- 5:06 Kind of like an early stage company of just first doing something, making more radical changes quickly.
- 5:13 So yeah, this way, I think you can also get very familiar with what you're building, what the failure patterns are.
loading