Videos -npY6XjM8CQ
When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI
Scene timeline
49 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 166
- whisperx 166
- chunks
- 29
- from 166 cues
- keyframes
- 32
- kept of 49 captured
- frames with text
- 32
- 569 lines read
- chapters
- 10
- from the source metadata
- keyframe bytes
- 6.4 MB
- word timings on 166 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:27 | 0s |
stt |
done | — | 2026-08-09 00:48 | 18s |
chunk |
done | — | 2026-08-09 00:49 | 0s |
text_embed |
done | — | 2026-08-10 19:36 | 0s |
keyframe |
done | — | 2026-08-09 00:49 | 2m 19s |
ocr |
done | — | 2026-08-09 00:51 | 14s |
frame_embed |
done | — | 2026-08-10 19:36 | 6s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- World's Fair0.96
-
- Model Release Hype Cycle1.00
- SWE-bench Verified0.99
- AlEngineer0.99
- With thinking0.94
- World's Fair0.97
- 1.1.00
- Model launch: Announcement post, optional chart crime0.99
- 74.90.99
- 52 is smaller than 690.96
- 2.1.00
- People actually use it1.00
- 69.10.98
- 3081.00
- 3.1.00
- If it's a disappointment, allegations of "benchmaxxing" arise0.98
- PRESENTED BY1.00
- TechCrunch1.00
- Microsoft1.00
- Meta's benchmarks for its new Al models are a bit misleading1.00
- Meta appears to have used an unreleased, custom version of one of its new flagship Al0.99
- models, Maverick, to boost a benchmark score.1.00
- Apr 6, 20250.97
- davewolfs·5h ago0.99
- Why is this test contradicting the feedback being given here? I am confused. This at first glance says0.99
- Maverick is literally Topgun. But the feedback seems to be very different.0.99
- 1Reply0.98
- Award1.00
- Share0.99
- surge1.00
- Engineering the future of Al0.99
- World's Fair0.99
-
- AlEngineer0.98
- World'sFair1.00
- Why does benchmaxxing happen?1.00
- Incentives1.00
- Why are traditional benchmarks1.00
- Poor/under-funded1.00
- PRESENTED BY1.00
- such unreliable indicators of1.00
- methodologies1.00
- Microsoft1.00
- real-world value?1.00
- Is this intrinsic to benchmarking?0.99
- No1.00
- Will we ever know love which0.97
- Yes1.00
- models are the best?1.00
- surge1.00
- TRACK 9• JUNE 30, 20260.95
- World's Fair0.96
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- SurgeAl0.97
- We're the biggest post-training data1.00
- PRESENTED BY1.00
- + evals supplier across the industry,1.00
- Microsoft1.00
- working with the frontier labs0.99
- Hit $1B+ annual revenue with $0 VC Funding0.99
- Nick Heiner1.00
- Lead of RL Envs @ Surge0.99
- surge1.00
- TRACK 9· JUNE 30, 20260.96
- World'sFair1.00
- Data Quality1.00
-
- Benchmarks ≠ Reality0.99
- AlEngineer0.97
- World's Fair0.97
- Tech - OpenAl0.96
- Which company has top Al model end of June? (Style0.98
- Control On)0.99
- Past1.00
- Jun 301.00
- Jul 310.91
- Anthropic 94%OpenAl 3.2%Google 2.5% Meta <1%0.97
- Polymarket1.00
- 100%1.00
- 75%1.00
- 50%1.00
- 25%1.00
- $1,787,646 Vol. Jun 30, 20260.96
- 1H 6H 1D 1W 1M ALL 可0.91
- Anthropic1.00
- $60,546 Vol.0.97
- 94%64%0.93
- Buy Yes 95e0.96
- Buy No 7€0.99
- OpenAl0.94
- $72,006 Vol.1.00
- 3%36%0.95
- Buy Yes 3.9€0.97
- Buy No 97.5€0.99
- Google1.00
- $44,076 Vol.0.97
- 2%38%1.00
- Buy Yes 2.9€0.96
- Buy No 98.0€0.96
- surge1.00
- TRACK 9· JUNE 30, 20260.96
- World's Fair0.97
- Data Quality0.97
-
- AlEngineer0.96
- World'sFair0.98
- "The issue with ... the LMArena stuff, is that they're ... quite easily gameable .. It was0.97
- trivial for our team to tune a version of Llama 4 Maverick that would sit way at the top."0.99
- 3surge0.92
- TRACK 9·JUNE 30,20260.98
- World'sFair1.00
- Data Quality0.98
-
- Gwern:1.00
- AlEngineer0.96
- World'sFair1.00
- "It can be easily gamed. The users are self-selected, and they0.98
- have zero incentive to be honest ... you should take being #1 on0.99
- LMArena seriously as useful and important news - as a reason0.99
- not to use a model."0.99
- "It's past time for LMArena people to sit down and have some0.99
- thorough reflection on whether it is still worth running at all, and1.00
- at what point they are doing more harm than good. No0.99
- benchmark lives forever and it is normal and healthy to shut0.99
- them down at some point after having been saturated, but some0.98
- manage to live a lot longer than they should have"1.00
- surge1.00
- TRACK 9· JUNE 30,20260.96
- World'sFair1.00
- Data Quality0.97
-
- Incentives1.00
- AlEngineer0.99
- World'sFair1.00
- surge1.00
- TRACK 9· JUNE 30, 20260.97
- World's Fair0.96
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- Why Are Benchmarks Unreliable Indicators of1.00
- Real-World Value?1.00
- Price1.00
- Contamination1.00
- Ambition1.00
- Taste1.00
- Operational Ability1.00
- surge1.00
- TRACK 9· JUNE 30, 20260.96
- World'sFair0.99
- Data Quality0.99
-
- Price1.00
- AlEngineer0.98
- World'sFair1.00
- 1000 tasks * 60 hours per task using $500k/year SWEs = $15M up-front, $5M on-going1.00
- Workaround1.00
- => Impact0.97
- Use public data that's already in the1.00
- Contamination1.00
- training corpus1.00
- Use lots of Al assistance0.99
- The whole effort becomes kinda circular0.99
- Very small sample size1.00
- Noisy, very sensitive to contingencies1.00
- Use cheap labor1.00
- Low task quality; lack of domain0.99
- expertise1.00
- surge1.00
- TRACK 9• JUNE 30, 20260.97
- World's Fair0.97
- Data Quality1.00
-
- Contamination1.00
- AlEngineer0.98
- World's Fair0.96
- surge1.00
- TRACK 9· JUNE 30, 20260.97
- World'sFair1.00
- Data Quality1.00
-
- In v5.3, NDDataRef mask propagation fails when one of the1.00
- operand does not have a mask #149780.98
- SWE-Bench Verified instance0.98
- AlEngineer0.98
- ID astropy_astropy-14995 (PR)0.98
- World'sFair1.00
- prompt Opus 4.80.98
- with this first part0.99
- and it'll recite this0.97
- second part0.99
- surge1.00
- TRACK 9• JUNE 30, 20260.94
- World'sFair1.00
- Data Quality1.00
-
- Fixed #30191-- Only selected referenced fields during cascade deletion. #110870.98
- SWE-Bench Verified instance0.98
- AlEngineer0.98
- ID django__django-110870.98
- World'sFair1.00
- prompt Opus 4.81.00
- with this first part0.99
- and it'll recite0.93
- this second part0.99
- surge1.00
- TRACK 9· JUNE 30,20260.95
- World'sFair1.00
- Data Quality0.96
-
- Reward Hacking1.00
- AlEngineer0.98
- World'sFair1.00
- Prompt: Please write an0.99
- Sample Response1.00
- 80-word summary of the1.00
- Renewable energy plays a crucial part in reducing0.99
- importance of renewable1.00
- carbon emissions rapidly, sustainability.0.99
- PRESENTED BY1.00
- energy in reducing carbon1.00
- a greener future, harmony.1.00
- Clean energy sources like tidal and geothermal create0.99
- emissions, ensuring to1.00
- Microsoft1.00
- Harnessing energy from the sun and wind decreases1.00
- completely avoid any words1.00
- reliance heavily, infrastructure.1.00
- that contain the letter 'o.'1.00
- This transition helps mitigate climate change impacts0.99
- Additionally, use a sentence1.00
- severely, catastrophe.1.00
- structure such that every1.00
- Investing in renewable energy creates a cleaner1.00
- environment steadily, prosperity.1.00
- sentence ends with a noun.0.99
- A sustainable energy mix is essential urgently,1.00
- priority.1.00
- Renewable energy can help achieve a greener planet0.99
- gradually, sanctuary.1.00
- surge1.00
- TRACK 9· JUNE 30,20260.96
- World's Fair0.98
- Data Quality1.00
-
- IFEval1.00
- AlEngineer0.99
- World'sFair1.00
- Family1.00
- ~Count1.00
- Example1.00
- word / sentence / paragraph0.98
- ~901.00
- "at least 500 words" / "less than 171.00
- count bound0.99
- sentences"1.00
- required keyword(s)1.00
- ~701.00
- "include the keywords 'atlantis' and1.00
- 'constable'"0.96
- no-comma / punctuation ban0.99
- ~701.00
- "Do not use any commas in your0.98
- response."1.00
- All-lowercase or all-uppercase1.00
- ~1001.00
- "in all lowercase letters. No capital0.98
- letters are allowed"0.99
- letter-frequency / lipogram0.99
- ~321.00
- "the letter t should appear at most0.98
- once"1.00
- surge1.00
- TRACK 9• JUNE 30, 20260.97
- World's Fair0.99
- Data Quality1.00
Transcript
166 cues· 2,916 words· 16,146 chars
- 0:12 Let's get started.
- 0:14 When will the benchmarking plague end?
- 0:19 In the tech industry, we love a hype cycle, and in AI, we really love a hype cycle.
- 0:25 And the way we do that is when a model comes out, there's a big announcement, there's a lot of benchmarks cited.
- 0:32 Sometimes, to keep things interesting, we do a little chart crime, and then people actually go and use it.
- 0:38 And if the expectations aren't met by the reality, then we have allegations of benchmarking.
- 0:45 Benchmarking, of course, being when labs are training too hard on benchmarks in a way that deviates from what people actually care about.
- 0:55 So the existence of that term indicates that we have a sense that benchmarks don't always equal reality.
- 1:00 And so in this talk, we're gonna figure out why does benchmarking happen?
- 1:05 Why are traditional benchmarks not always accurate reflections of real-world value?
- 1:10 Is this intrinsic to all benchmarks, and will we ever know which models are best?
- 1:15 And the answers are incentives, poor methodologies, no, and yes.
- 1:19 All right, that was my talk.
- 1:20 Thank you so much for coming.
- 1:23 Actually, it looks like I have a few extra minutes, so let's move on.
- 1:27 I have a few extra slides we'll go through.
- 1:31 So we have a sense that benchmarks don't equal reality, but the industry is dominated by a lot of popular but very bad benchmarks.
- 1:39 So there's millions of dollars on prediction markets being wagered on LM Arena outcomes, even as we have industry leaders openly bragging about gaming LM Arena.
- 1:52 And you have thought leaders like Wern saying it can be easily gamed.
- 1:56 It's past time for the LM Arena people to sit down and think about whether they're doing more harm than good.
- 2:02 Andrej Karpathy had a similar observation when he noticed that the models that he thought were best were not lining up with what LM Arena was ranking.
- 2:11 He said, unfortunately, the teams are not getting better models overall, but better LM Arena models, whatever that is.
- 2:17 Possibly something with a lot of nested lists, bullet points, and emojis.
- 2:21 So why does this happen?
- 2:22 That industry insiders are telling us that this benchmark is not useful, but it still gets a lot of play.
- 2:30 The problem is that AI is aimed at everyone in the world, is something everyone in the world can use, and so everyone needs some tool to figure out which models are best, and benchmarks are what we have for that.
- 2:43 But if you can't, if you don't have the ability to assess if a benchmark is good,
- 2:47 what you do have is the ability to assess what's popular.
- 2:50 And this creates this avalanche, this feedback effect where the conversation is very much driven by incumbency and marketing and less by real-world value.
- 2:59 And even myself, unless I actually look at a benchmark in a fair amount of detail, I don't have an opinion on it.
- 3:05 So it's a very challenging problem.
- 3:08 So what are the things that benchmarks do that lead to these problems?
- 3:14 There are a handful of key anti-patterns that we're gonna go through.
- 3:19 The first is price.
- 3:20 Let's say you want to make an agentic coding benchmark, which these days is a very popular thing to want to do, and you want 1,000 tasks in your benchmark.
- 3:28 Each task takes 60 hours to make.
- 3:30 Each software engineer in your workforce costs half a million a year.
- 3:34 That's $15 million to make your benchmark.
- 3:38 And if you think that over time, about a third of those tasks are gonna get washed away every year due to models getting better, that's $5 million to replace them.
- 3:48 So that puts you out of budget for most projects.
- 3:52 So then people turn to a variety of workarounds that have their own problems.
- 3:56 One of which is trying to use a lot of AI assistance, which ultimately does not really work.
- 4:02 Like you can't push the frontier forward from within the frontier.
- 4:05 You need to inject that external human expertise.
- 4:09 And it needs to be good expertise.
- 4:12 If you try to use cheap labor, you're gonna get what you're paid for and the whole result is not gonna be that useful.
- 4:19 At Surge, one of our differentiators has long been that we are not trying to minimize cost, we are trying to maximize quality.
- 4:27 And part of that means paying a lot of money for good workers.
- 4:31 We've always believed that, but especially in 2026, models are just beyond the point where you can make do with anything less than the best workers.
- 4:41 Contamination is often thought of as when labs are explicitly training on the test set.
loading
Chapters
- 0:00 The benchmark versus reality gap
- 0:55 Why the word benchmaxxing exists
- 2:38 Reading a benchmark fairly
- 3:14 Antipattern: broken tasks
- 4:41 Antipattern: contamination
- 5:57 Antipattern: reward hacking
- 6:23 Misaligned prompts and verifiers
- 10:40 Benchmaxxing as a two way street
- 13:14 Domain expertise and getting it right
- 15:47 Human eval and a higher standard