read-only demo

Videos xyL2Ltkh-SA

How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads

index_state ready data_status ok

AI Engineer· published 2026-07-24· 0:19:29· en-US· indexed 2026-08-10 19:40

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:20, 1 of 1 keyframes kept
  5. Shot 4, 0:20 to 0:42, 1 of 1 keyframes kept
  6. Shot 5, 0:42 to 1:12, 1 of 1 keyframes kept
  7. Shot 6, 1:12 to 1:40, 1 of 1 keyframes kept
  8. Shot 7, 1:40 to 2:08, 1 of 1 keyframes kept
  9. Shot 8, 2:08 to 2:36, 0 of 1 keyframes kept
  10. Shot 9, 2:36 to 3:03, 1 of 1 keyframes kept
  11. Shot 10, 3:03 to 3:30, 0 of 1 keyframes kept
  12. Shot 11, 3:30 to 3:56, 0 of 1 keyframes kept
  13. Shot 12, 3:56 to 4:26, 0 of 1 keyframes kept
  14. Shot 13, 4:26 to 4:56, 0 of 1 keyframes kept
  15. Shot 14, 4:56 to 5:25, 1 of 1 keyframes kept
  16. Shot 15, 5:25 to 5:55, 0 of 1 keyframes kept
  17. Shot 16, 5:55 to 6:26, 1 of 1 keyframes kept
  18. Shot 17, 6:26 to 6:56, 0 of 1 keyframes kept
  19. Shot 18, 6:56 to 7:29, 1 of 1 keyframes kept
  20. Shot 19, 7:29 to 7:55, 1 of 1 keyframes kept
  21. Shot 20, 7:55 to 8:20, 0 of 1 keyframes kept
  22. Shot 21, 8:20 to 8:46, 0 of 1 keyframes kept
  23. Shot 22, 8:46 to 9:11, 0 of 1 keyframes kept
  24. Shot 23, 9:11 to 9:37, 0 of 1 keyframes kept
  25. Shot 24, 9:37 to 10:05, 0 of 1 keyframes kept
  26. Shot 25, 10:05 to 10:33, 1 of 1 keyframes kept
  27. Shot 26, 10:33 to 11:00, 0 of 1 keyframes kept
  28. Shot 27, 11:00 to 11:28, 1 of 1 keyframes kept
  29. Shot 28, 11:28 to 11:56, 0 of 1 keyframes kept
  30. Shot 29, 11:56 to 12:20, 1 of 1 keyframes kept
  31. Shot 30, 12:20 to 12:58, 1 of 1 keyframes kept
  32. Shot 31, 12:58 to 13:24, 1 of 1 keyframes kept
  33. Shot 32, 13:24 to 13:49, 0 of 1 keyframes kept
  34. Shot 33, 13:49 to 14:14, 0 of 1 keyframes kept
  35. Shot 34, 14:14 to 14:46, 1 of 1 keyframes kept
  36. Shot 35, 14:46 to 15:19, 0 of 1 keyframes kept
  37. Shot 36, 15:19 to 15:51, 1 of 1 keyframes kept
  38. Shot 37, 15:51 to 16:21, 0 of 1 keyframes kept
  39. Shot 38, 16:21 to 16:52, 0 of 1 keyframes kept
  40. Shot 39, 16:52 to 17:23, 0 of 1 keyframes kept
  41. Shot 40, 17:23 to 17:57, 0 of 1 keyframes kept
  42. Shot 41, 17:57 to 18:03, 0 of 1 keyframes kept
  43. Shot 42, 18:03 to 18:37, 1 of 1 keyframes kept
  44. Shot 43, 18:37 to 19:12, 1 of 1 keyframes kept
  45. Shot 44, 19:12 to 19:13, 0 of 1 keyframes kept
  46. Shot 45, 19:13 to 19:27, 0 of 1 keyframes kept
  47. Shot 46, 19:27 to 19:28, 1 of 1 keyframes kept

47 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
216
whisperx 216
chunks
34
from 216 cues
keyframes
23
kept of 47 captured
frames with text
22
431 lines read
chapters
0
from the source metadata
keyframe bytes
5.6 MB
word timings on 216 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:48 0s
stt done 2026-08-09 07:08 21s
chunk done 2026-08-09 07:08 0s
text_embed done 2026-08-10 19:40 0s
keyframe done 2026-08-09 07:08 2m 28s
ocr done 2026-08-09 07:11 13s
frame_embed done 2026-08-10 19:40 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:19 #3 done2 line(s)

    shot 3·sharpness 283.5

    1. AlEngineer0.99
    2. World's Fair0.99
  • 0:35 #4 done12 line(s)

    shot 4·sharpness 1697.8

    1. Ask Oemini0.97
    2. AWork0.88
    3. Finish update I0.95
    4. AlEngineer0.97
    5. World'sFair1.00
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. Model Whisperers1.00
    9. Evals for Production Agents1.00
    10. *all opinions are my own and not of my employer1.00
    11. Engineering the future of Al0.98
    12. World'sFair1.00
  • 0:45 #5 done10 line(s)

    shot 5·sharpness 1558.5

    1. Ask Oemini0.97
    2. Work0.96
    3. Finish update 10.96
    4. AlEngineer0.98
    5. World'sFair1.00
    6. PRESENTED BY0.99
    7. Microsoft1.00
    8. Building an AI agent is hard1.00
    9. Engineering the future of Al0.99
    10. World'sFair1.00
  • 1:16 #6 done20 line(s)

    shot 6·sharpness 2115.7

    1. Ask Oemini0.96
    2. docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
    3. OWork0.84
    4. Finish update i0.93
    5. AlEngineer0.96
    6. World'sFair1.00
    7. Agent foundation1.00
    8. PRESENTED BY1.00
    9. Microsoft1.00
    10. A focused set of strong, optimized LLM-friendly tools1.00
    11. provides a good foundation1.00
    12. Once your tools are optimized (important), an independent1.00
    13. critique agent with a remediation loop can fill more gaps1.00
    14. by providing a self-correction mechanism1.00
    15. Once your base structure is defined, a strong eval provides0.98
    16. a method for proving value and ablation - an essential tool0.99
    17. for climbing the quality ladder1.00
    18. TRACK 5· JULY 1,20260.97
    19. World'sFair1.00
    20. Evals1.00
  • 1:46 #7 done18 line(s)

    shot 7·sharpness 1984.6

    1. Ask Oemini0.95
    2. docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
    3. OWork0.82
    4. Finish update !0.91
    5. AlEngineer0.97
    6. World'sFair1.00
    7. Agent foundation1.00
    8. A focused set of strong, optimized LLM-friendly tools1.00
    9. provides a good foundation1.00
    10. Once your tools are optimized (important), an independent1.00
    11. critique agent with a remediation loop can fill more gaps0.98
    12. by providing a self-correction mechanism1.00
    13. Once your base structure is defined, a strong eval provides0.99
    14. a method for proving value and ablation - an essential tool1.00
    15. for climbing the quality ladder1.00
    16. World'sFair1.00
    17. TRACK 5· JULY 1,20260.96
    18. Evals1.00
  • 2:22 #8 skipped

    shot 8·duplicate of #7

  • 3:00 #9 done23 line(s)

    shot 9·sharpness 1995.2

    1. Ask Gemini0.99
    2. docs.google.com/presentation/d/lesaioTRybTIWMFD2JM7UUL_g7HCfLxbidbE6GET7k/edit?resourcekey=0-ihtkJI0R17HJiBOCBECQeg&slide=id.g311...☆0.91
    3. Work0.89
    4. Finish update I0.94
    5. AlEngineer0.98
    6. World'sFair1.00
    7. Reliability = f(agent capability, guardrails, evals)1.00
    8. Understanding what agents do in the0.99
    9. Defining good1.00
    10. real world1.00
    11. Evals allows us to understand and0.99
    12. Gen AI outputs are non-deterministic0.99
    13. improve how modes behave in the real0.99
    14. world by defining what "good" looks0.96
    15. We can't guarantee how an agent will0.99
    16. like.1.00
    17. behave in the wild without measuring0.99
    18. it at scale.1.00
    19. To build evals that actually scale,1.00
    20. they must be strict and measurable.0.99
    21. World'sFair1.00
    22. TRACK 5· JULY 1, 20260.96
    23. Evals1.00
  • 3:06 #10 skipped

    shot 10·duplicate of #9

  • 3:35 #11 skipped

    shot 11·duplicate of #7

  • 4:11 #12 skipped

    shot 12·duplicate of #7

  • 4:46 #13 skipped

    shot 13·duplicate of #7

  • 5:22 #14 done29 line(s)

    shot 14·sharpness 2186.2

    1. Ask Gemini0.93
    2. .com/presentation/d/lesaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJI0Rt7HJiBOCBECQeg&slide=id.g3f..0.97
    3. Work0.93
    4. Finish update I0.97
    5. AlEngineer0.99
    6. World's Fair0.99
    7. Vibing can be good for you (early on!)1.00
    8. PRESENTED BY1.00
    9. 801.00
    10. Microsoft1.00
    11. Intuition based validations can be1.00
    12. 701.00
    13. good even if non-scalable0.99
    14. Pass te%0.74
    15. 601.00
    16. Prompt tweaks can have large1.00
    17. 501.00
    18. performance gains0.97
    19. 401.00
    20. hillclimb in a targeted way0.99
    21. Learn the failure patterns and1.00
    22. 301.00
    23. The Eval Rollercoaster1.00
    24. Jumping to scaled raters too early led0.99
    25. to many ups/downs in the eval process1.00
    26. and slowed us down1.00
    27. World's Fair0.98
    28. TRACK 5· JULY 1,20260.95
    29. Evals1.00
  • 5:37 #15 skipped

    shot 15·duplicate of #7

  • 6:10 #16 done25 line(s)

    shot 16·sharpness 1960.4

    1. Ask Gemini0.95
    2. docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f..0.97
    3. 0.79
    4. Work0.94
    5. Finish update I0.97
    6. AlEngineer0.99
    7. World's Fair0.98
    8. Start early, start small0.98
    9. PRESENTED BY0.99
    10. HOW EVALS GET DONE0.97
    11. Microsoft1.00
    12. Don't need to build a massive golden1.00
    13. WRITING1.00
    14. set on Day 1; start with a few core0.99
    15. THE EVALS0.98
    16. tasks1.00
    17. Test the negatives - checking if the0.99
    18. model didn't do something bad is just0.99
    19. as critical as checking if it did the0.97
    20. task1.00
    21. HUMANS ARGUING1.00
    22. OVER THE RUBRIC1.00
    23. World's Fair0.99
    24. TRACK 5·JULY 1,20260.98
    25. Evals1.00
  • 6:29 #17 skipped

    shot 17·duplicate of #16

  • 7:07 #18 done18 line(s)

    shot 18·sharpness 1986.7

    1. Ask Gemini0.94
    2. 0.57
    3. work0.94
    4. Finish update I0.96
    5. AlEngineer0.99
    6. World's Fair0.99
    7. Get outside help0.98
    8. How do we evaluate0.98
    9. We use another1.00
    10. the model?1.00
    11. model to grade it.1.00
    12. But how do we1.00
    13. It's just models grading0.98
    14. evaluate that model?1.00
    15. models all the way down.0.99
    16. World's Fair1.00
    17. TRACK 5 · JULY 1, 20260.92
    18. Evals1.00
  • 7:37 #19 done27 line(s)

    shot 19·sharpness 2013.4

    1. Ask Gemini0.98
    2. docs.google.com/presentation/d/1esaioTRybTihWMFD2JM7UUL_g7HCfLxbidbE6GET7fk/edit?resourcekey=0-ihtkJIORt7HJiBOCBECQeg&slide=id.g3f0..0.98
    3. 0.56
    4. Work0.89
    5. Finish update I0.96
    6. AlEngineer0.99
    7. World's Fair0.97
    8. Working with scaled raters1.00
    9. Clear rubric template with trusted0.99
    10. examples1.00
    11. Human-human agreement should be1.00
    12. Corporate needs-you to find the differences0.97
    13. strong1.00
    14. between this picture and this picture.0.99
    15. Get explanations along with verdict0.99
    16. O0.57
    17. Categorical input1.00
    18. pass/fail1.00
    19. Multi-output1.00
    20. is it accurate?1.00
    21. is it brand safe?0.98
    22. Explanations can be used to1.00
    23. make the agent better1.00
    24. They're the same picture.0.98
    25. World's Fair0.97
    26. TRACK 5·JULY 1,20260.97
    27. Evals1.00
  • 8:17 #20 skipped

    shot 20·duplicate of #19

  • 8:33 #21 skipped

    shot 21·duplicate of #19

  • 8:51 #22 skipped

    shot 22·duplicate of #19

  • 9:24 #23 skipped

    shot 23·duplicate of #19

Transcript

216 cues· 3,345 words· 17,780 chars

  1. 0:12 Hi, everyone.
  2. 0:15 Sounds like everybody came back from lunch, so I hope everybody is recharged and not sleepy at all.
  3. 0:20 It's always interesting to do a talk right after lunch, because you never know, it's a mixed crowd.
  4. 0:24 But we're very happy to be here, happy to see you all.
  5. 0:27 Our talk is gonna be about evals, of course.
  6. 0:29 We're in the evals track.
  7. 0:30 We're gonna talk you through what are some things that worked for us while we were building evals, especially for YouTube ads.
  8. 0:36 We work on the YouTube ads team as part of the, we do image and video models for YouTube ads.
  9. 0:42 So building an agent is hard.
  10. 0:44 I think anybody who is here in the audience probably has built an agent as a side project or as part of production systems.
  11. 0:50 It's a very hard thing to do.
  12. 0:52 It's laborious.
  13. 0:52 It takes a lot of time.
  14. 0:55 Making it reliable is harder.
  15. 0:57 So having it do things that you actually want it to do in production, understanding the different kind of things that it can play with, how it's going to react when you launch it to your end users, that's always a very hard thing to do, which is why evals are a pretty handy way to manage that.
  16. 1:15 Yeah, and then basically the first step when you're doing this is, of course, you need to have your agent foundation.
  17. 1:23 So when you're building your agent, you will want to have a focused and strong set of LLM-friendly tools to give your agent a very good foundation.
  18. 1:35 So yeah, I would say it's important to first optimize these tools and make sure they're the best they can be before just jumping onto
  19. 1:44 larger agent evals.
  20. 1:46 So once your tools are optimized, you can also take some other steps like making an independent critique agent with a remediation loop.
  21. 1:57 And this can fill more gaps as far as having a self-correction mechanism and filling those gaps where maybe your base toolset has limitations.
  22. 2:09 And then once your base structure is defined, you can have
  23. 2:13 You can then go to having an eval and having a strong eval is very important as this gives you like a way of proving the value of changes you make as well as running ablation experiments on any changes you make.
  24. 2:27 So I would say this is a very essential tool for climbing the quality ladder.
  25. 2:31 But again, it's very important to have that good foundation to begin with.
  26. 2:39 So yeah, the reliability of your agent is basically a function of the capabilities of the agent, the guardrails and the evals.
  27. 2:52 So understanding what your agents do in the real world, basically generative AI outputs, as I'm sure you're all familiar, are not exactly deterministic, right?
  28. 3:04 So it can often fail in certain areas or one time it can succeed, one time it can fail.
  29. 3:11 So we can't really guarantee how it will behave in the wild.
  30. 3:13 And for some use cases, this is extremely important, right?
  31. 3:17 And we need a way to measure at scale and make sure that
  32. 3:21 It is getting the output we want despite the non-determinism of these models.
  33. 3:28 So we need to define what's good here.
  34. 3:31 And evals allow us to basically understand and improve how the models behave in the real world by defining what good looks like.
  35. 3:43 It's basically just setting this is our target output.
  36. 3:49 So to build evals that actually scale, they really need to be strict and measurable.
  37. 3:58 And so an interesting thing here that I think might be somewhat counterintuitive is that early on vibing can actually be kind of good for you.
  38. 4:08 And what I mean here by vibing is basically doing things that are not exactly scalable to begin with.
  39. 4:15 So when you're first starting out, it may be that, you know, you could
  40. 4:23 take a track of basically just going ahead and making the super comprehensive eval.
  41. 4:28 But we found it actually works better to first do intuition-based approach, where you kind of first see the capabilities and look at the outputs.
  42. 4:37 And at this stage, it's pretty easy to tell what the issues actually are.
  43. 4:42 So even though this is non-scalable, it will still give you a very good idea of when you change this, what happens.
  44. 4:50 And it allows you to more quickly iterate as well.
  45. 4:54 So at this stage, prompt tweets can also have large performance gains.
  46. 4:58 You can make a radical change to the architecture.
  47. 5:01 And your eval is not hindering you in this way.
  48. 5:04 So it's a very good way.
  49. 5:06 Kind of like an early stage company of just first doing something, making more radical changes quickly.
  50. 5:13 So yeah, this way, I think you can also get very familiar with what you're building, what the failure patterns are.

Open at this second