read-only demo

Videos Ubwb6NzegyA

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

index_state ready data_status ok

AI Engineer· published 2026-05-25· 0:20:02· en-US· indexed 2026-08-10 19:56

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 0:34, 1 of 1 keyframes kept
  6. Shot 5, 0:34 to 1:09, 1 of 1 keyframes kept
  7. Shot 6, 1:09 to 1:40, 1 of 1 keyframes kept
  8. Shot 7, 1:40 to 2:05, 1 of 1 keyframes kept
  9. Shot 8, 2:05 to 2:07, 1 of 1 keyframes kept
  10. Shot 9, 2:07 to 2:52, 1 of 1 keyframes kept
  11. Shot 10, 2:52 to 3:23, 1 of 1 keyframes kept
  12. Shot 11, 3:23 to 3:53, 0 of 1 keyframes kept
  13. Shot 12, 3:53 to 4:20, 1 of 1 keyframes kept
  14. Shot 13, 4:20 to 4:47, 0 of 1 keyframes kept
  15. Shot 14, 4:47 to 5:18, 0 of 1 keyframes kept
  16. Shot 15, 5:18 to 5:48, 0 of 1 keyframes kept
  17. Shot 16, 5:48 to 5:55, 0 of 1 keyframes kept
  18. Shot 17, 5:55 to 6:27, 1 of 1 keyframes kept
  19. Shot 18, 6:27 to 6:59, 0 of 1 keyframes kept
  20. Shot 19, 6:59 to 7:45, 1 of 1 keyframes kept
  21. Shot 20, 7:45 to 7:48, 1 of 1 keyframes kept
  22. Shot 21, 7:48 to 8:23, 0 of 1 keyframes kept
  23. Shot 22, 8:23 to 8:59, 1 of 1 keyframes kept
  24. Shot 23, 8:59 to 9:34, 0 of 1 keyframes kept
  25. Shot 24, 9:34 to 9:44, 1 of 1 keyframes kept
  26. Shot 25, 9:44 to 10:05, 1 of 1 keyframes kept
  27. Shot 26, 10:05 to 10:15, 0 of 1 keyframes kept
  28. Shot 27, 10:15 to 10:36, 0 of 1 keyframes kept
  29. Shot 28, 10:36 to 10:46, 0 of 1 keyframes kept
  30. Shot 29, 10:46 to 10:53, 0 of 1 keyframes kept
  31. Shot 30, 10:53 to 11:23, 1 of 1 keyframes kept
  32. Shot 31, 11:23 to 11:54, 0 of 1 keyframes kept
  33. Shot 32, 11:54 to 12:00, 1 of 1 keyframes kept
  34. Shot 33, 12:00 to 12:07, 1 of 1 keyframes kept
  35. Shot 34, 12:07 to 12:20, 1 of 1 keyframes kept
  36. Shot 35, 12:20 to 12:29, 0 of 1 keyframes kept
  37. Shot 36, 12:29 to 12:36, 1 of 1 keyframes kept
  38. Shot 37, 12:36 to 12:42, 0 of 1 keyframes kept
  39. Shot 38, 12:42 to 13:10, 0 of 1 keyframes kept
  40. Shot 39, 13:10 to 13:37, 1 of 1 keyframes kept
  41. Shot 40, 13:37 to 14:05, 0 of 1 keyframes kept
  42. Shot 41, 14:05 to 14:32, 0 of 1 keyframes kept
  43. Shot 42, 14:32 to 15:00, 0 of 1 keyframes kept
  44. Shot 43, 15:00 to 15:31, 1 of 1 keyframes kept
  45. Shot 44, 15:31 to 16:02, 0 of 1 keyframes kept
  46. Shot 45, 16:02 to 16:33, 0 of 1 keyframes kept
  47. Shot 46, 16:33 to 16:37, 0 of 1 keyframes kept
  48. Shot 47, 16:37 to 16:38, 0 of 1 keyframes kept
  49. Shot 48, 16:38 to 16:52, 1 of 1 keyframes kept
  50. Shot 49, 16:52 to 17:23, 0 of 1 keyframes kept
  51. Shot 50, 17:23 to 17:56, 0 of 1 keyframes kept
  52. Shot 51, 17:56 to 18:28, 1 of 1 keyframes kept
  53. Shot 52, 18:28 to 19:00, 0 of 1 keyframes kept
  54. Shot 53, 19:00 to 19:32, 0 of 1 keyframes kept
  55. Shot 54, 19:32 to 19:47, 1 of 1 keyframes kept
  56. Shot 55, 19:47 to 20:01, 1 of 1 keyframes kept
  57. Shot 56, 20:01 to 20:02, 0 of 1 keyframes kept

57 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
234
whisperx 234
chunks
34
from 234 cues
keyframes
29
kept of 57 captured
frames with text
29
533 lines read
chapters
0
from the source metadata
keyframe bytes
5.9 MB
word timings on 234 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 19:14 1m 27s
stt done 2026-08-10 19:15 23s
chunk done 2026-08-10 19:15 0s
text_embed done 2026-08-10 19:56 0s
keyframe done 2026-08-10 19:15 1m 41s
ocr done 2026-08-10 19:17 11s
frame_embed done 2026-08-10 19:56 5s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:26 #3 done3 line(s)

    shot 3·sharpness 290.5

    1. Agentic evaluation at scale1.00
    2. - for everybody0.97
    3. AlEng0.93
  • 0:31 #4 done9 line(s)

    shot 4·sharpness 1684.6

    1. Agentic evaluation at scale1.00
    2. AIE1.00
    3. – for everybody0.96
    4. 0.99
    5. Al Engineer Europe 20260.99
    6. kaggle1.00
    7. DeepMind0.98
    8. Google DeepMind1.00
    9. AlEng0.95
  • 1:05 #5 done13 line(s)

    shot 5·sharpness 1497.5

    1. Who we are0.99
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. 1.00
    7. Nicholas Kang1.00
    8. Michael Aaron0.97
    9. Product Manager1.00
    10. Software Engineer0.98
    11. Kaggle @ AIE0.91
    12. Braintrust1.00
    13. WorkOS OpenAI0.98
  • 1:12 #6 done24 line(s)

    shot 6·sharpness 2171.7

    1. Whatis Kaggle?1.00
    2. World's largest online community of Al/ML practitioners, researchers and enthusiasts.1.00
    3. Over 30 million people have registered on Kaggle to enter competitions, solve machine learning0.99
    4. problems, explore open datasets, publish cutting-edge models, learn, and share data science knowledge.0.99
    5. AIE1.00
    6. 1.00
    7. 1.00
    8. 1.00
    9. 1.00
    10. KAGGLEHIGHLIGHTS1.00
    11. 30M+0.99
    12. 500+1.00
    13. 5K+0.94
    14. 1.4M+1.00
    15. 470K+0.99
    16. Kaggle Members1.00
    17. Featured Competitions &0.99
    18. Al evaluations1.00
    19. Public Notebooks1.00
    20. Public Datasets1.00
    21. Hackathons1.00
    22. Kaggle AIE0.94
    23. Braintrust1.00
    24. WorkOS OpenAI0.93
  • 1:48 #7 done11 line(s)

    shot 7·sharpness 1953.7

    1. Agenda for today1.00
    2. 1. Al evals today are kinda broken0.99
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. 1.00
    7. 1.00
    8. 2. Kaggle is trying to fix it, but it's tough!0.99
    9. AlEngineer0.96
    10. EUROPE1.00
    11. AIEn0.91
  • 2:06 #8 done8 line(s)

    shot 8·sharpness 1443.6

    1. 011.00
    2. AIE1.00
    3. 1.00
    4. 0.99
    5. Al evals today are kinda broken1.00
    6. AlEngineer0.97
    7. EUROPE1.00
    8. AIEn0.92
  • 2:47 #9 done14 line(s)

    shot 9·sharpness 2690.1

    1. Evals are scattered, decentralized, and get stale fast1.00
    2. AIE0.99
    3. arXiv0.99
    4. Allabs0.96
    5. 1.00
    6. 1.00
    7. GitHub1.00
    8. Most benchmarks today live in Github repos, Arxiv papers, and in the confines of Al labs0.99
    9. Understanding landscape is a full time job to keep track of the latest & greatest coming out of1.00
    10. Arxiv1.00
    11. Once leaderboards are published, they often don't get updated by the original publishers0.98
    12. anymore1.00
    13. Engineering the future of Al0.99
    14. AIEn0.92
  • 3:01 #10 done40 line(s)

    shot 10·sharpness 3021.7

    1. Evals aren't always transparent, accessible, and verifiable1.00
    2. Introducing GPT-5.41.00
    3. When labs report their results, we often only see the1.00
    4. reported results1.00
    5. AIE1.00
    6. 1.00
    7. 1.00
    8. GPT-5.41.00
    9. GPT-5.3-Codex1.00
    10. GPT-5.21.00
    11. But how were the benchmarks set up? What1.00
    12. 1.00
    13. 1.00
    14. GDPval (wins or ties)0.99
    15. 83.0%1.00
    16. 70.9%1.00
    17. 70.9%1.00
    18. configs were used on the models? What are the0.98
    19. SWE-Bench Pro (Public)1.00
    20. 57.7%1.00
    21. 56.8%1.00
    22. 55.6%1.00
    23. benchmarks actually testing?0.99
    24. OSWorld-Verified0.98
    25. 75.0%0.97
    26. 74.0%*0.92
    27. 47.3%1.00
    28. Toolathlon0.99
    29. 54.6%1.00
    30. 51.9%0.99
    31. 46.3%1.00
    32. Sometimes, different labs publish different1.00
    33. BrowseComp1.00
    34. 82.7%1.00
    35. 77.3%0.98
    36. 65.8%0.93
    37. results for their competitors on the same0.97
    38. benchmarks1.00
    39. Engineering the future of Al0.99
    40. AIEn0.88
  • 3:29 #11 skipped

    shot 11·duplicate of #7

  • 3:56 #12 done22 line(s)

    shot 12·sharpness 2558.8

    1. Most benchmarks are created by Al researchers1.00
    2. All of the world and1.00
    3. its knowledge1.00
    4. AIE1.00
    5. We expect Al to impact most of1.00
    6. 1.00
    7. humanity1.00
    8. 1.00
    9. 1.00
    10. Technical1.00
    11. However, most benchmarks are1.00
    12. professionals1.00
    13. created by Al researchers1.00
    14. Yes, they hire experts, but there is a0.99
    15. Al researchers1.00
    16. long tail of important knowledge to0.99
    17. go after if we want Al to benefit all of0.99
    18. humanity1.00
    19. Very illustrative1.00
    20. AlEngineer0.98
    21. EUROPE1.00
    22. AIEn0.92
  • 4:29 #13 skipped

    shot 13·duplicate of #7

  • 5:08 #14 skipped

    shot 14·duplicate of #7

  • 5:45 #15 skipped

    shot 15·duplicate of #7

  • 5:54 #16 skipped

    shot 16·duplicate of #8

  • 5:59 #17 done26 line(s)

    shot 17·sharpness 2550.9

    1. Kaggle is trying to solve these problems1.00
    2. 1.00
    3. Agent Exams1.00
    4. AIE1.00
    5. Hackathons1.00
    6. An experimental MVP, where we let0.99
    7. 1.00
    8. 1.00
    9. 1.00
    10. 1.00
    11. Create and host your own1.00
    12. hackathons in minutes1.00
    13. 0.85
    14. 0.56
    15. you test your agent with a 1-liner1.00
    16. prompt & publish it results on a0.98
    17. leaderboard1.00
    18. Game Arena1.00
    19. Watch top Al models compete0.98
    20. Benchmarks1.00
    21. in PvP games like Chess, Poker,0.99
    22. Build and run your own evals with us0.98
    23. etc.1.00
    24. AlEngineer0.96
    25. EUROPE1.00
    26. AlEn0.89
  • 6:34 #18 skipped

    shot 18·duplicate of #17

  • 7:26 #19 done23 line(s)

    shot 19·sharpness 2259.4

    1. Hackathons1.00
    2. Channel the community's energy0.99
    3. Measuring Progress Toward AGI - Cognitive0.98
    4. & expertise to solving a problem1.00
    5. Abilities1.00
    6. 1.00
    7. 1.00
    8. AIE1.00
    9. Putting guardrails around the0.99
    10. problem and providing a clear0.99
    11. Overview1.00
    12. 1.00
    13. problem statement to inspire0.98
    14. 1.00
    15. 1.00
    16. them1.00
    17. Discussions to enable innovation1.00
    18. to flourish0.98
    19. Results open sourced for the1.00
    20. benefit of everybody1.00
    21. Kaggle @ AIE0.92
    22. Engineering the future of Al0.99
    23. AIEn0.90
  • 7:45 #20 done44 line(s)

    shot 20·sharpness 3037.7

    1. But running a hackathon isn't all that easy0.99
    2. Hackathon Judging1.00
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. Producing a clear problem statement and1.00
    7. are going to come and many may interpret it0.99
    8. evaluation rubric is hard! Thousands of participants0.99
    9. differently than you intend1.00
    10. How: Use /agent to loop through submissions and vet against a defined rubric - do manual0.96
    11. Goal: remove invalid and low quality submissions that don't meet the basics of our rubric0.99
    12. spotcheck within each round to ensure qualty and repeat multiple times until we have a reasonab!0.96
    13. # submissions left0.96
    14. Round 1: Initial LLM eligibility screening1.00
    15. Contributors: Yao Yan Martyna Plomecka1.00
    16. 1.00
    17. Est time: 1 week (week of Apr 20)0.98
    18. 1.00
    19. 1.00
    20. 1.00
    21. Enabling participants to do their best work1.00
    22. requires that you provide them the right tools0.99
    23. Goal: Fiterto top -20 submissions per judge0.96
    24. Round 2: First human review0.99
    25. Contributors: All judges0.96
    26. (e.g., how they can host their dataset, use Al models,0.98
    27. Methodology: Provide a 0-10 score for each row in the rubric and get a weighted sum0.98
    28. and share their work in a clear way that can be easily0.98
    29. How: Each judge will receive N submissions, grade it per the methodology, and retum it to us. The0.97
    30. The organizing team then aggregates all the submissions together and takes out the top -200.99
    31. accessed and understood by others)1.00
    32. Est time: 2 weeks (week of Apr 27 & May 4)0.99
    33. Round 3: Final human review0.99
    34. Contributors:Al judges0.95
    35. Judging still requires human experts and0.99
    36. Methodology: Each judge reviews up to 20 submissions and scores them with the same rubric in0.99
    37. Goal: Select shortlist of grand prize winners (4x) and track winners (10x) in total0.99
    38. coordination + alignment is not trivial!1.00
    39. round 21.00
    40. Est time: 2 weeks (week of May 11 and 18)0.99
    41. Kaggle @ AIE0.88
    42. | Proprietary & Confidential0.96
    43. Engineering the future of Al1.00
    44. AIEn0.82
  • 8:09 #21 skipped

    shot 21·duplicate of #19

  • 8:45 #22 done46 line(s)

    shot 22·sharpness 3019.0

    1. But running a hackathon isn't all that easy1.00
    2. Hackathon Judging0.99
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. Producing a clear problem statement and1.00
    7. are going to come and many may interpret it1.00
    8. evaluation rubric is hard! Thousands of participants0.99
    9. differently than you intend1.00
    10. How: Use /agent to loop through submissions and vet against a defined rubric - do manual0.97
    11. Goal: remove invalid and low quality submissions that don't meet the basics of our rubric0.99
    12. spotcheck within each round to ensure qualty and repeat multiple times until we have a reasonabi0.96
    13. # submissions left0.96
    14. Round 1: Initial LLM eligibility screening0.98
    15. Contributors: Yao Yan Martyna Plomecka0.99
    16. 1.00
    17. Est time: 1 week (week of Apr 20)0.98
    18. 1.00
    19. 1.00
    20. 1.00
    21. Enabling participants to do their best work1.00
    22. requires that you provide them the right tools0.99
    23. Goal: Fite to top -20 submissions per judge0.95
    24. Contributors: All judges0.94
    25. Round 2: First human review1.00
    26. (e.g., how they can host their dataset, use Al models,0.98
    27. Methodology: Provide a 0-10 score for each row in the rubric and get a weighted sum0.98
    28. and share their work in a clear way that can be easily0.98
    29. How: Each judge will receive N submissions, grade it per the methodology, and retum it to us. The0.98
    30. The organizing team then aggregates all the submissions together and takes out the top -200.99
    31. accessed and understood by others)1.00
    32. Est time: 2 weeks (week of Apr 27 & May 4)0.98
    33. Round 3: Final human review1.00
    34. Contributors:Al judges0.96
    35. Judging still requires human experts and0.99
    36. Methodology: Each judge reviews up to 20 submissions and scores them with the same rubric in0.99
    37. Goal: Select shortlist of grand prize winners (4x) and track winners (10x) in total0.99
    38. coordination + alignment is not trivial!1.00
    39. round 21.00
    40. Est time: 2 weeks (week of May 11 and 18)0.99
    41. Kaggle (@ AIE0.91
    42. 131.00
    43. Google | Proprietary & Confidential0.95
    44. AlEngineer0.96
    45. EUROPE1.00
    46. AIEn0.86
  • 9:30 #23 skipped

    shot 23·duplicate of #20

Transcript

234 cues· 3,888 words· 21,061 chars

  1. 0:15 All right.
  2. 0:15 Hi, everybody.
  3. 0:17 Let me just try to stand straight so I don't have to crouch over.
  4. 0:20 Thank you all for coming.
  5. 0:22 This is our talk on energetic evaluations at scale for everybody.
  6. 0:26 I hope everyone's in the right room.
  7. 0:28 And if you are, thank you for coming.
  8. 0:30 We were expecting like 20 people.
  9. 0:31 So this is like way more than what we expected.
  10. 0:35 So all right, who are we?
  11. 0:36 So I'm Nick.
  12. 0:36 I'm a product manager in Kaggle Benchmarks.
  13. 0:39 And I basically run and build our benchmarks platform alongside a couple of our engineers.
  14. 0:44 And I also focus on our agentic eval solutions.
  15. 0:47 I'm originally from Singapore, but I live in the San Francisco Bay Area.
  16. 0:50 And so I flew in to do this talk and attend all the great talks out in this conference today.
  17. 0:56 And hi, I'm Michael.
  18. 0:58 I'm a software engineer on Kaggle.
  19. 0:59 I've been working at Google for about a third of the time that I've been alive, and Kaggle for about half of that.
  20. 1:04 So yeah, but mostly working on evaluations and benchmarks for Kaggle at the moment.
  21. 1:09 Has anyone here heard of Kaggle?
  22. 1:11 Put your hands up if you have.
  23. 1:13 Okay, great.
  24. 1:13 So a lot of people know us for competitions, but we don't just use that.
  25. 1:17 We're the world's largest AI ML community of 30-plus million users, and we've been working a lot in the gen AI eval space over the course of the past two years.
  26. 1:26 And we think there are lots of interesting problems in the industry that not many people are trying to solve, and we feel like we're positioned well to solve them.
  27. 1:34 and we want to share more of the work that we've been doing and also invite contributions if you want to get involved in the space
  28. 1:41 So we have a simple agenda for today.
  29. 1:42 First is AI evals today are kind of broken, and we'll talk about why.
  30. 1:46 And then step two is we're trying to solve it, not saying we're the all cure and we have everything kind of sorted out.
  31. 1:53 We'll talk about what we're trying to do, the challenges we're running into, and also maybe it might inspire some of you in terms of how you think you might be able to help contribute to this very important problem that we're trying to solve for.
  32. 2:05 With that, I'll jump into the first section.
  33. 2:08 So first problem, evals are scattered, decentralized, and get stale fast.
  34. 2:12 I don't know how many of you have tried to keep track of AI benchmarks, but basically 10 plus of them drop every single day.
  35. 2:20 And the best way to find out what they are is go to archive and spend hours scrolling through them, reading every paper.
  36. 2:26 That doesn't make sense.
  37. 2:27 We don't think it makes sense.
  38. 2:28 I can't even do it, even though it's my full-time job.
  39. 2:31 And I think what happens after that paper gets published is that, you know, you see some of the leaderboards in the papers and what happens after that?
  40. 2:40 They just get stale.
  41. 2:41 The authors move on to the next best benchmark because they just want to publish lots of papers, no fault of their own, but these leaderboards no longer become relevant as time goes on.
  42. 2:53 The second issue is that evals aren't always transparent, accessible, and verifiable.
  43. 2:58 I'm sure many of us have seen these charts on these model publisher notes when they release a new model.
  44. 3:04 But what's the problem with that?
  45. 3:06 We don't actually know how these benchmarks are set up.
  46. 3:09 It's a lot of configurations you could use for the models themselves, and also how the benchmark is orchestrated and facilitated.
  47. 3:16 And we don't always know what's actually being tested here.
  48. 3:20 I'll give you one real anecdote, which is that we had published a benchmark with one of these AI labs, and another competing AI lab came to us and said, hey, we don't like the results of this particular benchmark you published.
  49. 3:33 Let's run it on our own.
  50. 3:35 And so they ran it, and then they published it with much higher, much better results.

Open at this second