read-only demo

Videos -npY6XjM8CQ

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

index_state ready data_status ok

AI Engineer· published 2026-08-02· 0:17:24· en-US· indexed 2026-08-10 19:36

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:28, 1 of 1 keyframes kept
  5. Shot 4, 0:28 to 0:53, 1 of 1 keyframes kept
  6. Shot 5, 0:53 to 1:19, 1 of 1 keyframes kept
  7. Shot 6, 1:19 to 1:30, 1 of 1 keyframes kept
  8. Shot 7, 1:30 to 1:44, 1 of 1 keyframes kept
  9. Shot 8, 1:44 to 1:51, 1 of 1 keyframes kept
  10. Shot 9, 1:51 to 2:01, 1 of 1 keyframes kept
  11. Shot 10, 2:01 to 2:47, 0 of 1 keyframes kept
  12. Shot 11, 2:47 to 3:08, 1 of 1 keyframes kept
  13. Shot 12, 3:08 to 3:17, 1 of 1 keyframes kept
  14. Shot 13, 3:17 to 3:45, 0 of 1 keyframes kept
  15. Shot 14, 3:45 to 4:12, 1 of 1 keyframes kept
  16. Shot 15, 4:12 to 4:39, 0 of 1 keyframes kept
  17. Shot 16, 4:39 to 5:01, 1 of 1 keyframes kept
  18. Shot 17, 5:01 to 5:16, 1 of 1 keyframes kept
  19. Shot 18, 5:16 to 5:19, 1 of 1 keyframes kept
  20. Shot 19, 5:19 to 5:53, 0 of 1 keyframes kept
  21. Shot 20, 5:53 to 6:22, 1 of 1 keyframes kept
  22. Shot 21, 6:22 to 7:10, 0 of 1 keyframes kept
  23. Shot 22, 7:10 to 7:57, 0 of 1 keyframes kept
  24. Shot 23, 7:57 to 8:32, 1 of 1 keyframes kept
  25. Shot 24, 8:32 to 8:48, 1 of 1 keyframes kept
  26. Shot 25, 8:48 to 8:58, 0 of 1 keyframes kept
  27. Shot 26, 8:58 to 9:05, 0 of 1 keyframes kept
  28. Shot 27, 9:05 to 9:33, 1 of 1 keyframes kept
  29. Shot 28, 9:33 to 9:45, 1 of 1 keyframes kept
  30. Shot 29, 9:45 to 10:05, 0 of 1 keyframes kept
  31. Shot 30, 10:05 to 10:12, 0 of 1 keyframes kept
  32. Shot 31, 10:12 to 10:33, 1 of 1 keyframes kept
  33. Shot 32, 10:33 to 10:44, 0 of 1 keyframes kept
  34. Shot 33, 10:44 to 10:50, 1 of 1 keyframes kept
  35. Shot 34, 10:50 to 11:29, 1 of 1 keyframes kept
  36. Shot 35, 11:29 to 12:01, 1 of 1 keyframes kept
  37. Shot 36, 12:01 to 12:23, 1 of 1 keyframes kept
  38. Shot 37, 12:23 to 13:09, 1 of 1 keyframes kept
  39. Shot 38, 13:09 to 13:16, 1 of 1 keyframes kept
  40. Shot 39, 13:16 to 13:45, 0 of 1 keyframes kept
  41. Shot 40, 13:45 to 14:14, 0 of 1 keyframes kept
  42. Shot 41, 14:14 to 14:42, 1 of 1 keyframes kept
  43. Shot 42, 14:42 to 15:11, 0 of 1 keyframes kept
  44. Shot 43, 15:11 to 15:39, 1 of 1 keyframes kept
  45. Shot 44, 15:39 to 16:08, 1 of 1 keyframes kept
  46. Shot 45, 16:08 to 16:37, 0 of 1 keyframes kept
  47. Shot 46, 16:37 to 16:54, 0 of 1 keyframes kept
  48. Shot 47, 16:54 to 17:07, 1 of 1 keyframes kept
  49. Shot 48, 17:07 to 17:24, 0 of 1 keyframes kept

49 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
166
whisperx 166
chunks
29
from 166 cues
keyframes
32
kept of 49 captured
frames with text
32
569 lines read
chapters
10
from the source metadata
keyframe bytes
6.4 MB
word timings on 166 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:27 0s
stt done 2026-08-09 00:48 18s
chunk done 2026-08-09 00:49 0s
text_embed done 2026-08-10 19:36 0s
keyframe done 2026-08-09 00:49 2m 19s
ocr done 2026-08-09 00:51 14s
frame_embed done 2026-08-10 19:36 6s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.7

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2743.2

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:14 #3 done1 line(s)

    shot 3·sharpness 357.0

    1. World's Fair0.96
  • 0:36 #4 done31 line(s)

    shot 4·sharpness 3815.0

    1. Model Release Hype Cycle1.00
    2. SWE-bench Verified0.99
    3. AlEngineer0.99
    4. With thinking0.94
    5. World's Fair0.97
    6. 1.1.00
    7. Model launch: Announcement post, optional chart crime0.99
    8. 74.90.99
    9. 52 is smaller than 690.96
    10. 2.1.00
    11. People actually use it1.00
    12. 69.10.98
    13. 3081.00
    14. 3.1.00
    15. If it's a disappointment, allegations of "benchmaxxing" arise0.98
    16. PRESENTED BY1.00
    17. TechCrunch1.00
    18. Microsoft1.00
    19. Meta's benchmarks for its new Al models are a bit misleading1.00
    20. Meta appears to have used an unreleased, custom version of one of its new flagship Al0.99
    21. models, Maverick, to boost a benchmark score.1.00
    22. Apr 6, 20250.97
    23. davewolfs·5h ago0.99
    24. Why is this test contradicting the feedback being given here? I am confused. This at first glance says0.99
    25. Maverick is literally Topgun. But the feedback seems to be very different.0.99
    26. 1Reply0.98
    27. Award1.00
    28. Share0.99
    29. surge1.00
    30. Engineering the future of Al0.99
    31. World's Fair0.99
  • 1:16 #5 done20 line(s)

    shot 5·sharpness 3628.4

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Why does benchmaxxing happen?1.00
    4. Incentives1.00
    5. Why are traditional benchmarks1.00
    6. Poor/under-funded1.00
    7. PRESENTED BY1.00
    8. such unreliable indicators of1.00
    9. methodologies1.00
    10. Microsoft1.00
    11. real-world value?1.00
    12. Is this intrinsic to benchmarking?0.99
    13. No1.00
    14. Will we ever know love which0.97
    15. Yes1.00
    16. models are the best?1.00
    17. surge1.00
    18. TRACK 9• JUNE 30, 20260.95
    19. World's Fair0.96
    20. Data Quality1.00
  • 1:26 #6 done15 line(s)

    shot 6·sharpness 3189.9

    1. AlEngineer0.98
    2. World'sFair1.00
    3. SurgeAl0.97
    4. We're the biggest post-training data1.00
    5. PRESENTED BY1.00
    6. + evals supplier across the industry,1.00
    7. Microsoft1.00
    8. working with the frontier labs0.99
    9. Hit $1B+ annual revenue with $0 VC Funding0.99
    10. Nick Heiner1.00
    11. Lead of RL Envs @ Surge0.99
    12. surge1.00
    13. TRACK 9· JUNE 30, 20260.96
    14. World'sFair1.00
    15. Data Quality1.00
  • 1:40 #7 done36 line(s)

    shot 7·sharpness 1770.7

    1. Benchmarks ≠ Reality0.99
    2. AlEngineer0.97
    3. World's Fair0.97
    4. Tech - OpenAl0.96
    5. Which company has top Al model end of June? (Style0.98
    6. Control On)0.99
    7. Past1.00
    8. Jun 301.00
    9. Jul 310.91
    10. Anthropic 94%OpenAl 3.2%Google 2.5% Meta <1%0.97
    11. Polymarket1.00
    12. 100%1.00
    13. 75%1.00
    14. 50%1.00
    15. 25%1.00
    16. $1,787,646 Vol. Jun 30, 20260.96
    17. 1H 6H 1D 1W 1M ALL 可0.91
    18. Anthropic1.00
    19. $60,546 Vol.0.97
    20. 94%64%0.93
    21. Buy Yes 95e0.96
    22. Buy No 7€0.99
    23. OpenAl0.94
    24. $72,006 Vol.1.00
    25. 3%36%0.95
    26. Buy Yes 3.9€0.97
    27. Buy No 97.5€0.99
    28. Google1.00
    29. $44,076 Vol.0.97
    30. 2%38%1.00
    31. Buy Yes 2.9€0.96
    32. Buy No 98.0€0.96
    33. surge1.00
    34. TRACK 9· JUNE 30, 20260.96
    35. World's Fair0.97
    36. Data Quality0.97
  • 1:47 #8 done8 line(s)

    shot 8·sharpness 2257.3

    1. AlEngineer0.96
    2. World'sFair0.98
    3. "The issue with ... the LMArena stuff, is that they're ... quite easily gameable .. It was0.97
    4. trivial for our team to tune a version of Llama 4 Maverick that would sit way at the top."0.99
    5. 3surge0.92
    6. TRACK 9·JUNE 30,20260.98
    7. World'sFair1.00
    8. Data Quality0.98
  • 1:58 #9 done17 line(s)

    shot 9·sharpness 3980.2

    1. Gwern:1.00
    2. AlEngineer0.96
    3. World'sFair1.00
    4. "It can be easily gamed. The users are self-selected, and they0.98
    5. have zero incentive to be honest ... you should take being #1 on0.99
    6. LMArena seriously as useful and important news - as a reason0.99
    7. not to use a model."0.99
    8. "It's past time for LMArena people to sit down and have some0.99
    9. thorough reflection on whether it is still worth running at all, and1.00
    10. at what point they are doing more harm than good. No0.99
    11. benchmark lives forever and it is normal and healthy to shut0.99
    12. them down at some point after having been saturated, but some0.98
    13. manage to live a lot longer than they should have"1.00
    14. surge1.00
    15. TRACK 9· JUNE 30,20260.96
    16. World'sFair1.00
    17. Data Quality0.97
  • 2:24 #10 skipped

    shot 10·duplicate of #9

  • 2:59 #11 done7 line(s)

    shot 11·sharpness 1509.2

    1. Incentives1.00
    2. AlEngineer0.99
    3. World'sFair1.00
    4. surge1.00
    5. TRACK 9· JUNE 30, 20260.97
    6. World's Fair0.96
    7. Data Quality1.00
  • 3:14 #12 done13 line(s)

    shot 12·sharpness 2455.8

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Why Are Benchmarks Unreliable Indicators of1.00
    4. Real-World Value?1.00
    5. Price1.00
    6. Contamination1.00
    7. Ambition1.00
    8. Taste1.00
    9. Operational Ability1.00
    10. surge1.00
    11. TRACK 9· JUNE 30, 20260.96
    12. World'sFair0.99
    13. Data Quality0.99
  • 3:36 #13 skipped

    shot 13·duplicate of #9

  • 4:06 #14 done20 line(s)

    shot 14·sharpness 4004.8

    1. Price1.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. 1000 tasks * 60 hours per task using $500k/year SWEs = $15M up-front, $5M on-going1.00
    5. Workaround1.00
    6. => Impact0.97
    7. Use public data that's already in the1.00
    8. Contamination1.00
    9. training corpus1.00
    10. Use lots of Al assistance0.99
    11. The whole effort becomes kinda circular0.99
    12. Very small sample size1.00
    13. Noisy, very sensitive to contingencies1.00
    14. Use cheap labor1.00
    15. Low task quality; lack of domain0.99
    16. expertise1.00
    17. surge1.00
    18. TRACK 9• JUNE 30, 20260.97
    19. World's Fair0.97
    20. Data Quality1.00
  • 4:25 #15 skipped

    shot 15·duplicate of #9

  • 4:41 #16 done7 line(s)

    shot 16·sharpness 2739.1

    1. Contamination1.00
    2. AlEngineer0.98
    3. World's Fair0.96
    4. surge1.00
    5. TRACK 9· JUNE 30, 20260.97
    6. World'sFair1.00
    7. Data Quality1.00
  • 5:10 #17 done14 line(s)

    shot 17·sharpness 2142.2

    1. In v5.3, NDDataRef mask propagation fails when one of the1.00
    2. operand does not have a mask #149780.98
    3. SWE-Bench Verified instance0.98
    4. AlEngineer0.98
    5. ID astropy_astropy-14995 (PR)0.98
    6. World'sFair1.00
    7. prompt Opus 4.80.98
    8. with this first part0.99
    9. and it'll recite this0.97
    10. second part0.99
    11. surge1.00
    12. TRACK 9• JUNE 30, 20260.94
    13. World'sFair1.00
    14. Data Quality1.00
  • 5:16 #18 done13 line(s)

    shot 18·sharpness 1836.1

    1. Fixed #30191-- Only selected referenced fields during cascade deletion. #110870.98
    2. SWE-Bench Verified instance0.98
    3. AlEngineer0.98
    4. ID django__django-110870.98
    5. World'sFair1.00
    6. prompt Opus 4.81.00
    7. with this first part0.99
    8. and it'll recite0.93
    9. this second part0.99
    10. surge1.00
    11. TRACK 9· JUNE 30,20260.95
    12. World'sFair1.00
    13. Data Quality0.96
  • 5:46 #19 skipped

    shot 19·duplicate of #5

  • 6:13 #20 done34 line(s)

    shot 20·sharpness 6647.9

    1. Reward Hacking1.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. Prompt: Please write an0.99
    5. Sample Response1.00
    6. 80-word summary of the1.00
    7. Renewable energy plays a crucial part in reducing0.99
    8. importance of renewable1.00
    9. carbon emissions rapidly, sustainability.0.99
    10. PRESENTED BY1.00
    11. energy in reducing carbon1.00
    12. a greener future, harmony.1.00
    13. Clean energy sources like tidal and geothermal create0.99
    14. emissions, ensuring to1.00
    15. Microsoft1.00
    16. Harnessing energy from the sun and wind decreases1.00
    17. completely avoid any words1.00
    18. reliance heavily, infrastructure.1.00
    19. that contain the letter 'o.'1.00
    20. This transition helps mitigate climate change impacts0.99
    21. Additionally, use a sentence1.00
    22. severely, catastrophe.1.00
    23. structure such that every1.00
    24. Investing in renewable energy creates a cleaner1.00
    25. environment steadily, prosperity.1.00
    26. sentence ends with a noun.0.99
    27. A sustainable energy mix is essential urgently,1.00
    28. priority.1.00
    29. Renewable energy can help achieve a greener planet0.99
    30. gradually, sanctuary.1.00
    31. surge1.00
    32. TRACK 9· JUNE 30,20260.96
    33. World's Fair0.98
    34. Data Quality1.00
  • 7:05 #21 skipped

    shot 21·duplicate of #9

  • 7:16 #22 skipped

    shot 22·duplicate of #5

  • 8:14 #23 done31 line(s)

    shot 23·sharpness 5033.7

    1. IFEval1.00
    2. AlEngineer0.99
    3. World'sFair1.00
    4. Family1.00
    5. ~Count1.00
    6. Example1.00
    7. word / sentence / paragraph0.98
    8. ~901.00
    9. "at least 500 words" / "less than 171.00
    10. count bound0.99
    11. sentences"1.00
    12. required keyword(s)1.00
    13. ~701.00
    14. "include the keywords 'atlantis' and1.00
    15. 'constable'"0.96
    16. no-comma / punctuation ban0.99
    17. ~701.00
    18. "Do not use any commas in your0.98
    19. response."1.00
    20. All-lowercase or all-uppercase1.00
    21. ~1001.00
    22. "in all lowercase letters. No capital0.98
    23. letters are allowed"0.99
    24. letter-frequency / lipogram0.99
    25. ~321.00
    26. "the letter t should appear at most0.98
    27. once"1.00
    28. surge1.00
    29. TRACK 9• JUNE 30, 20260.97
    30. World's Fair0.99
    31. Data Quality1.00

Transcript

166 cues· 2,916 words· 16,146 chars

  1. 0:12 Let's get started.
  2. 0:14 When will the benchmarking plague end?
  3. 0:19 In the tech industry, we love a hype cycle, and in AI, we really love a hype cycle.
  4. 0:25 And the way we do that is when a model comes out, there's a big announcement, there's a lot of benchmarks cited.
  5. 0:32 Sometimes, to keep things interesting, we do a little chart crime, and then people actually go and use it.
  6. 0:38 And if the expectations aren't met by the reality, then we have allegations of benchmarking.
  7. 0:45 Benchmarking, of course, being when labs are training too hard on benchmarks in a way that deviates from what people actually care about.
  8. 0:55 So the existence of that term indicates that we have a sense that benchmarks don't always equal reality.
  9. 1:00 And so in this talk, we're gonna figure out why does benchmarking happen?
  10. 1:05 Why are traditional benchmarks not always accurate reflections of real-world value?
  11. 1:10 Is this intrinsic to all benchmarks, and will we ever know which models are best?
  12. 1:15 And the answers are incentives, poor methodologies, no, and yes.
  13. 1:19 All right, that was my talk.
  14. 1:20 Thank you so much for coming.
  15. 1:23 Actually, it looks like I have a few extra minutes, so let's move on.
  16. 1:27 I have a few extra slides we'll go through.
  17. 1:31 So we have a sense that benchmarks don't equal reality, but the industry is dominated by a lot of popular but very bad benchmarks.
  18. 1:39 So there's millions of dollars on prediction markets being wagered on LM Arena outcomes, even as we have industry leaders openly bragging about gaming LM Arena.
  19. 1:52 And you have thought leaders like Wern saying it can be easily gamed.
  20. 1:56 It's past time for the LM Arena people to sit down and think about whether they're doing more harm than good.
  21. 2:02 Andrej Karpathy had a similar observation when he noticed that the models that he thought were best were not lining up with what LM Arena was ranking.
  22. 2:11 He said, unfortunately, the teams are not getting better models overall, but better LM Arena models, whatever that is.
  23. 2:17 Possibly something with a lot of nested lists, bullet points, and emojis.
  24. 2:21 So why does this happen?
  25. 2:22 That industry insiders are telling us that this benchmark is not useful, but it still gets a lot of play.
  26. 2:30 The problem is that AI is aimed at everyone in the world, is something everyone in the world can use, and so everyone needs some tool to figure out which models are best, and benchmarks are what we have for that.
  27. 2:43 But if you can't, if you don't have the ability to assess if a benchmark is good,
  28. 2:47 what you do have is the ability to assess what's popular.
  29. 2:50 And this creates this avalanche, this feedback effect where the conversation is very much driven by incumbency and marketing and less by real-world value.
  30. 2:59 And even myself, unless I actually look at a benchmark in a fair amount of detail, I don't have an opinion on it.
  31. 3:05 So it's a very challenging problem.
  32. 3:08 So what are the things that benchmarks do that lead to these problems?
  33. 3:14 There are a handful of key anti-patterns that we're gonna go through.
  34. 3:19 The first is price.
  35. 3:20 Let's say you want to make an agentic coding benchmark, which these days is a very popular thing to want to do, and you want 1,000 tasks in your benchmark.
  36. 3:28 Each task takes 60 hours to make.
  37. 3:30 Each software engineer in your workforce costs half a million a year.
  38. 3:34 That's $15 million to make your benchmark.
  39. 3:38 And if you think that over time, about a third of those tasks are gonna get washed away every year due to models getting better, that's $5 million to replace them.
  40. 3:48 So that puts you out of budget for most projects.
  41. 3:52 So then people turn to a variety of workarounds that have their own problems.
  42. 3:56 One of which is trying to use a lot of AI assistance, which ultimately does not really work.
  43. 4:02 Like you can't push the frontier forward from within the frontier.
  44. 4:05 You need to inject that external human expertise.
  45. 4:09 And it needs to be good expertise.
  46. 4:12 If you try to use cheap labor, you're gonna get what you're paid for and the whole result is not gonna be that useful.
  47. 4:19 At Surge, one of our differentiators has long been that we are not trying to minimize cost, we are trying to maximize quality.
  48. 4:27 And part of that means paying a lot of money for good workers.
  49. 4:31 We've always believed that, but especially in 2026, models are just beyond the point where you can make do with anything less than the best workers.
  50. 4:41 Contamination is often thought of as when labs are explicitly training on the test set.

Chapters

  1. 0:00 The benchmark versus reality gap
  2. 0:55 Why the word benchmaxxing exists
  3. 2:38 Reading a benchmark fairly
  4. 3:14 Antipattern: broken tasks
  5. 4:41 Antipattern: contamination
  6. 5:57 Antipattern: reward hacking
  7. 6:23 Misaligned prompts and verifiers
  8. 10:40 Benchmaxxing as a two way street
  9. 13:14 Domain expertise and getting it right
  10. 15:47 Human eval and a higher standard

Open at this second