Videos Ubwb6NzegyA
Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
Scene timeline
57 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 234
- whisperx 234
- chunks
- 34
- from 234 cues
- keyframes
- 29
- kept of 57 captured
- frames with text
- 29
- 533 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.9 MB
- word timings on 234 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 19:14 | 1m 27s |
stt |
done | — | 2026-08-10 19:15 | 23s |
chunk |
done | — | 2026-08-10 19:15 | 0s |
text_embed |
done | — | 2026-08-10 19:56 | 0s |
keyframe |
done | — | 2026-08-10 19:15 | 1m 41s |
ocr |
done | — | 2026-08-10 19:17 | 11s |
frame_embed |
done | — | 2026-08-10 19:56 | 5s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Agentic evaluation at scale1.00
- - for everybody0.97
- AlEng0.93
-
- Agentic evaluation at scale1.00
- AIE1.00
- – for everybody0.96
- ★0.99
- Al Engineer Europe 20260.99
- kaggle1.00
- DeepMind0.98
- Google DeepMind1.00
- AlEng0.95
-
- Who we are0.99
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Nicholas Kang1.00
- Michael Aaron0.97
- Product Manager1.00
- Software Engineer0.98
- Kaggle @ AIE0.91
- Braintrust1.00
- WorkOS OpenAI0.98
-
- Whatis Kaggle?1.00
- World's largest online community of Al/ML practitioners, researchers and enthusiasts.1.00
- Over 30 million people have registered on Kaggle to enter competitions, solve machine learning0.99
- problems, explore open datasets, publish cutting-edge models, learn, and share data science knowledge.0.99
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- KAGGLEHIGHLIGHTS1.00
- 30M+0.99
- 500+1.00
- 5K+0.94
- 1.4M+1.00
- 470K+0.99
- Kaggle Members1.00
- Featured Competitions &0.99
- Al evaluations1.00
- Public Notebooks1.00
- Public Datasets1.00
- Hackathons1.00
- Kaggle AIE0.94
- Braintrust1.00
- WorkOS OpenAI0.93
-
- Agenda for today1.00
- 1. Al evals today are kinda broken0.99
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 2. Kaggle is trying to fix it, but it's tough!0.99
- AlEngineer0.96
- EUROPE1.00
- AIEn0.91
-
- 011.00
- AIE1.00
- ★1.00
- ★0.99
- Al evals today are kinda broken1.00
- AlEngineer0.97
- EUROPE1.00
- AIEn0.92
-
- Evals are scattered, decentralized, and get stale fast1.00
- AIE0.99
- arXiv0.99
- Allabs0.96
- ★1.00
- ★1.00
- GitHub1.00
- Most benchmarks today live in Github repos, Arxiv papers, and in the confines of Al labs0.99
- Understanding landscape is a full time job to keep track of the latest & greatest coming out of1.00
- Arxiv1.00
- Once leaderboards are published, they often don't get updated by the original publishers0.98
- anymore1.00
- Engineering the future of Al0.99
- AIEn0.92
-
- Evals aren't always transparent, accessible, and verifiable1.00
- Introducing GPT-5.41.00
- When labs report their results, we often only see the1.00
- reported results1.00
- AIE1.00
- ★1.00
- ★1.00
- GPT-5.41.00
- GPT-5.3-Codex1.00
- GPT-5.21.00
- But how were the benchmarks set up? What1.00
- ★1.00
- ★1.00
- GDPval (wins or ties)0.99
- 83.0%1.00
- 70.9%1.00
- 70.9%1.00
- configs were used on the models? What are the0.98
- SWE-Bench Pro (Public)1.00
- 57.7%1.00
- 56.8%1.00
- 55.6%1.00
- benchmarks actually testing?0.99
- OSWorld-Verified0.98
- 75.0%0.97
- 74.0%*0.92
- 47.3%1.00
- Toolathlon0.99
- 54.6%1.00
- 51.9%0.99
- 46.3%1.00
- Sometimes, different labs publish different1.00
- BrowseComp1.00
- 82.7%1.00
- 77.3%0.98
- 65.8%0.93
- results for their competitors on the same0.97
- benchmarks1.00
- Engineering the future of Al0.99
- AIEn0.88
-
- Most benchmarks are created by Al researchers1.00
- All of the world and1.00
- its knowledge1.00
- AIE1.00
- We expect Al to impact most of1.00
- ★1.00
- humanity1.00
- ★1.00
- ★1.00
- Technical1.00
- However, most benchmarks are1.00
- professionals1.00
- created by Al researchers1.00
- Yes, they hire experts, but there is a0.99
- Al researchers1.00
- long tail of important knowledge to0.99
- go after if we want Al to benefit all of0.99
- humanity1.00
- Very illustrative1.00
- AlEngineer0.98
- EUROPE1.00
- AIEn0.92
-
- Kaggle is trying to solve these problems1.00
- ★1.00
- Agent Exams1.00
- AIE1.00
- Hackathons1.00
- An experimental MVP, where we let0.99
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Create and host your own1.00
- hackathons in minutes1.00
- 口0.85
- 区0.56
- you test your agent with a 1-liner1.00
- prompt & publish it results on a0.98
- leaderboard1.00
- Game Arena1.00
- Watch top Al models compete0.98
- Benchmarks1.00
- in PvP games like Chess, Poker,0.99
- Build and run your own evals with us0.98
- etc.1.00
- AlEngineer0.96
- EUROPE1.00
- AlEn0.89
-
- Hackathons1.00
- Channel the community's energy0.99
- Measuring Progress Toward AGI - Cognitive0.98
- & expertise to solving a problem1.00
- Abilities1.00
- ★1.00
- ★1.00
- AIE1.00
- Putting guardrails around the0.99
- problem and providing a clear0.99
- Overview1.00
- ★1.00
- problem statement to inspire0.98
- ★1.00
- ★1.00
- them1.00
- Discussions to enable innovation1.00
- to flourish0.98
- Results open sourced for the1.00
- benefit of everybody1.00
- Kaggle @ AIE0.92
- Engineering the future of Al0.99
- AIEn0.90
-
- But running a hackathon isn't all that easy0.99
- Hackathon Judging1.00
- AIE1.00
- ★1.00
- ★1.00
- Producing a clear problem statement and1.00
- are going to come and many may interpret it0.99
- evaluation rubric is hard! Thousands of participants0.99
- differently than you intend1.00
- How: Use /agent to loop through submissions and vet against a defined rubric - do manual0.96
- Goal: remove invalid and low quality submissions that don't meet the basics of our rubric0.99
- spotcheck within each round to ensure qualty and repeat multiple times until we have a reasonab!0.96
- # submissions left0.96
- Round 1: Initial LLM eligibility screening1.00
- Contributors: Yao Yan Martyna Plomecka1.00
- ★1.00
- Est time: 1 week (week of Apr 20)0.98
- ★1.00
- ★1.00
- ★1.00
- Enabling participants to do their best work1.00
- requires that you provide them the right tools0.99
- Goal: Fiterto top -20 submissions per judge0.96
- Round 2: First human review0.99
- Contributors: All judges0.96
- (e.g., how they can host their dataset, use Al models,0.98
- Methodology: Provide a 0-10 score for each row in the rubric and get a weighted sum0.98
- and share their work in a clear way that can be easily0.98
- How: Each judge will receive N submissions, grade it per the methodology, and retum it to us. The0.97
- The organizing team then aggregates all the submissions together and takes out the top -200.99
- accessed and understood by others)1.00
- Est time: 2 weeks (week of Apr 27 & May 4)0.99
- Round 3: Final human review0.99
- Contributors:Al judges0.95
- Judging still requires human experts and0.99
- Methodology: Each judge reviews up to 20 submissions and scores them with the same rubric in0.99
- Goal: Select shortlist of grand prize winners (4x) and track winners (10x) in total0.99
- coordination + alignment is not trivial!1.00
- round 21.00
- Est time: 2 weeks (week of May 11 and 18)0.99
- Kaggle @ AIE0.88
- | Proprietary & Confidential0.96
- Engineering the future of Al1.00
- AIEn0.82
-
- But running a hackathon isn't all that easy1.00
- Hackathon Judging0.99
- AIE1.00
- ★1.00
- ★1.00
- Producing a clear problem statement and1.00
- are going to come and many may interpret it1.00
- evaluation rubric is hard! Thousands of participants0.99
- differently than you intend1.00
- How: Use /agent to loop through submissions and vet against a defined rubric - do manual0.97
- Goal: remove invalid and low quality submissions that don't meet the basics of our rubric0.99
- spotcheck within each round to ensure qualty and repeat multiple times until we have a reasonabi0.96
- # submissions left0.96
- Round 1: Initial LLM eligibility screening0.98
- Contributors: Yao Yan Martyna Plomecka0.99
- ★1.00
- Est time: 1 week (week of Apr 20)0.98
- ★1.00
- ★1.00
- ★1.00
- Enabling participants to do their best work1.00
- requires that you provide them the right tools0.99
- Goal: Fite to top -20 submissions per judge0.95
- Contributors: All judges0.94
- Round 2: First human review1.00
- (e.g., how they can host their dataset, use Al models,0.98
- Methodology: Provide a 0-10 score for each row in the rubric and get a weighted sum0.98
- and share their work in a clear way that can be easily0.98
- How: Each judge will receive N submissions, grade it per the methodology, and retum it to us. The0.98
- The organizing team then aggregates all the submissions together and takes out the top -200.99
- accessed and understood by others)1.00
- Est time: 2 weeks (week of Apr 27 & May 4)0.98
- Round 3: Final human review1.00
- Contributors:Al judges0.96
- Judging still requires human experts and0.99
- Methodology: Each judge reviews up to 20 submissions and scores them with the same rubric in0.99
- Goal: Select shortlist of grand prize winners (4x) and track winners (10x) in total0.99
- coordination + alignment is not trivial!1.00
- round 21.00
- Est time: 2 weeks (week of May 11 and 18)0.99
- Kaggle (@ AIE0.91
- 131.00
- Google | Proprietary & Confidential0.95
- AlEngineer0.96
- EUROPE1.00
- AIEn0.86
Transcript
234 cues· 3,888 words· 21,061 chars
- 0:15 All right.
- 0:15 Hi, everybody.
- 0:17 Let me just try to stand straight so I don't have to crouch over.
- 0:20 Thank you all for coming.
- 0:22 This is our talk on energetic evaluations at scale for everybody.
- 0:26 I hope everyone's in the right room.
- 0:28 And if you are, thank you for coming.
- 0:30 We were expecting like 20 people.
- 0:31 So this is like way more than what we expected.
- 0:35 So all right, who are we?
- 0:36 So I'm Nick.
- 0:36 I'm a product manager in Kaggle Benchmarks.
- 0:39 And I basically run and build our benchmarks platform alongside a couple of our engineers.
- 0:44 And I also focus on our agentic eval solutions.
- 0:47 I'm originally from Singapore, but I live in the San Francisco Bay Area.
- 0:50 And so I flew in to do this talk and attend all the great talks out in this conference today.
- 0:56 And hi, I'm Michael.
- 0:58 I'm a software engineer on Kaggle.
- 0:59 I've been working at Google for about a third of the time that I've been alive, and Kaggle for about half of that.
- 1:04 So yeah, but mostly working on evaluations and benchmarks for Kaggle at the moment.
- 1:09 Has anyone here heard of Kaggle?
- 1:11 Put your hands up if you have.
- 1:13 Okay, great.
- 1:13 So a lot of people know us for competitions, but we don't just use that.
- 1:17 We're the world's largest AI ML community of 30-plus million users, and we've been working a lot in the gen AI eval space over the course of the past two years.
- 1:26 And we think there are lots of interesting problems in the industry that not many people are trying to solve, and we feel like we're positioned well to solve them.
- 1:34 and we want to share more of the work that we've been doing and also invite contributions if you want to get involved in the space
- 1:41 So we have a simple agenda for today.
- 1:42 First is AI evals today are kind of broken, and we'll talk about why.
- 1:46 And then step two is we're trying to solve it, not saying we're the all cure and we have everything kind of sorted out.
- 1:53 We'll talk about what we're trying to do, the challenges we're running into, and also maybe it might inspire some of you in terms of how you think you might be able to help contribute to this very important problem that we're trying to solve for.
- 2:05 With that, I'll jump into the first section.
- 2:08 So first problem, evals are scattered, decentralized, and get stale fast.
- 2:12 I don't know how many of you have tried to keep track of AI benchmarks, but basically 10 plus of them drop every single day.
- 2:20 And the best way to find out what they are is go to archive and spend hours scrolling through them, reading every paper.
- 2:26 That doesn't make sense.
- 2:27 We don't think it makes sense.
- 2:28 I can't even do it, even though it's my full-time job.
- 2:31 And I think what happens after that paper gets published is that, you know, you see some of the leaderboards in the papers and what happens after that?
- 2:40 They just get stale.
- 2:41 The authors move on to the next best benchmark because they just want to publish lots of papers, no fault of their own, but these leaderboards no longer become relevant as time goes on.
- 2:53 The second issue is that evals aren't always transparent, accessible, and verifiable.
- 2:58 I'm sure many of us have seen these charts on these model publisher notes when they release a new model.
- 3:04 But what's the problem with that?
- 3:06 We don't actually know how these benchmarks are set up.
- 3:09 It's a lot of configurations you could use for the models themselves, and also how the benchmark is orchestrated and facilitated.
- 3:16 And we don't always know what's actually being tested here.
- 3:20 I'll give you one real anecdote, which is that we had published a benchmark with one of these AI labs, and another competing AI lab came to us and said, hey, we don't like the results of this particular benchmark you published.
- 3:33 Let's run it on our own.
- 3:35 And so they ran it, and then they published it with much higher, much better results.
loading