read-only demo

Videos jWq-aZIU0kM

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:12:48· en-US· indexed 2026-08-10 19:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:27, 1 of 1 keyframes kept
  5. Shot 4, 0:27 to 1:05, 1 of 1 keyframes kept
  6. Shot 5, 1:05 to 1:22, 1 of 1 keyframes kept
  7. Shot 6, 1:22 to 1:46, 1 of 1 keyframes kept
  8. Shot 7, 1:46 to 1:56, 1 of 1 keyframes kept
  9. Shot 8, 1:56 to 2:29, 1 of 1 keyframes kept
  10. Shot 9, 2:29 to 3:02, 1 of 1 keyframes kept
  11. Shot 10, 3:02 to 3:35, 0 of 1 keyframes kept
  12. Shot 11, 3:35 to 4:03, 1 of 1 keyframes kept
  13. Shot 12, 4:03 to 4:50, 1 of 1 keyframes kept
  14. Shot 13, 4:50 to 5:16, 0 of 1 keyframes kept
  15. Shot 14, 5:16 to 5:59, 1 of 1 keyframes kept
  16. Shot 15, 5:59 to 6:29, 1 of 1 keyframes kept
  17. Shot 16, 6:29 to 6:58, 0 of 1 keyframes kept
  18. Shot 17, 6:58 to 7:24, 1 of 1 keyframes kept
  19. Shot 18, 7:24 to 7:51, 0 of 1 keyframes kept
  20. Shot 19, 7:51 to 8:07, 1 of 1 keyframes kept
  21. Shot 20, 8:07 to 8:33, 1 of 1 keyframes kept
  22. Shot 21, 8:33 to 8:59, 1 of 1 keyframes kept
  23. Shot 22, 8:59 to 9:25, 1 of 1 keyframes kept
  24. Shot 23, 9:25 to 9:51, 0 of 1 keyframes kept
  25. Shot 24, 9:51 to 10:17, 1 of 1 keyframes kept
  26. Shot 25, 10:17 to 10:43, 1 of 1 keyframes kept
  27. Shot 26, 10:43 to 11:08, 0 of 1 keyframes kept
  28. Shot 27, 11:08 to 11:34, 0 of 1 keyframes kept
  29. Shot 28, 11:34 to 12:00, 1 of 1 keyframes kept
  30. Shot 29, 12:00 to 12:24, 1 of 1 keyframes kept
  31. Shot 30, 12:24 to 12:31, 1 of 1 keyframes kept
  32. Shot 31, 12:31 to 12:48, 0 of 1 keyframes kept

32 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
121
whisperx 121
chunks
23
from 121 cues
keyframes
24
kept of 32 captured
frames with text
24
405 lines read
chapters
11
from the source metadata
keyframe bytes
3.4 MB
word timings on 121 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:43 0s
stt done 2026-08-09 02:45 11s
chunk done 2026-08-09 02:46 0s
text_embed done 2026-08-10 19:37 0s
keyframe done 2026-08-09 02:46 1m 36s
ocr done 2026-08-09 02:47 7s
frame_embed done 2026-08-10 19:37 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 455.2

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 670.0

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:11 #2 done24 line(s)

    shot 2·sharpness 2742.1

    1. LAB & PLATINUM SPONSORS0.97
    2. Amazon AGI Lab0.99
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.98
    6. OpenAl0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo0.99
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. together.ai0.98
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:14 #3 done2 line(s)

    shot 3·sharpness 379.1

    1. AlEngineer0.99
    2. World's Fair0.98
  • 1:00 #4 done13 line(s)

    shot 4·sharpness 2010.6

    1. AlEngineer0.98
    2. G2i1.00
    3. World's Fair0.96
    4. Ali Khial1.00
    5. PRESENTED BY1.00
    6. Director of Al/ML @ G2i1.00
    7. Microsoft1.00
    8. Software Engineer @1.00
    9. 50+ abandoned side projects1.00
    10. G2i. All rights reserved.1.00
    11. World's Fair0.99
    12. TRACK 9 · JULY 1, 20260.94
    13. Posttraining & Midtraining1.00
  • 1:10 #5 done11 line(s)

    shot 5·sharpness 2005.2

    1. AlEngineer0.98
    2. G2i1.00
    3. World'sFair1.00
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. A SWE-bench pro task instructions1.00
    7. G2i. All rights reserved.1.00
    8. 11.00
    9. World'sFair1.00
    10. TRACK 9 · JULY 1, 20260.93
    11. Posttraining & Midtraining1.00
  • 1:38 #6 done14 line(s)

    shot 6·sharpness 1615.2

    1. AlEngineer0.99
    2. G2i1.00
    3. World'sFair1.00
    4. No one1.00
    5. writes1.00
    6. prompts like0.99
    7. WHAT?1.00
    8. this. ever!1.00
    9. What the pro*pt ?!?1.00
    10. G2i. All rights reserved.1.00
    11. 21.00
    12. World's Fair0.99
    13. TRACK 9 · JULY 1, 20260.93
    14. Posttraining & Midtraining1.00
  • 1:51 #7 done10 line(s)

    shot 7·sharpness 1275.7

    1. AlEngineer0.98
    2. G2i1.00
    3. World'sFair1.00
    4. What are1.00
    5. benchmarks ?0.96
    6. G2i. All rights reserved.1.00
    7. 31.00
    8. World's Fair0.99
    9. TRACK 9 · JULY 1, 20260.93
    10. Posttraining & Midtraining1.00
  • 2:25 #8 done20 line(s)

    shot 8·sharpness 2368.4

    1. AlEngineer0.99
    2. G2i1.00
    3. World's Fair0.98
    4. Graders1.00
    5. LLM-as-judge1.00
    6. What are1.00
    7. Long Horizon1.00
    8. Verfiers1.00
    9. benchmarks ?0.98
    10. N-gram1.00
    11. Fail-to-pass1.00
    12. Benchmaxxing1.00
    13. Evals1.00
    14. Pass-to-pass1.00
    15. Hill Climbing0.99
    16. G2i. All rights reserved.1.00
    17. 31.00
    18. World's Fair0.97
    19. TRACK 9 • JULY 1, 20260.93
    20. Posttraining & Midtraining1.00
  • 2:58 #9 done20 line(s)

    shot 9·sharpness 2017.5

    1. AlEngineer0.98
    2. G2i1.00
    3. World's Fair0.99
    4. Prompts/1.00
    5. Models /0.94
    6. Solutions1.00
    7. Instructions1.00
    8. Agents1.00
    9. Trajectories1.00
    10. Scores1.00
    11. Metadata1.00
    12. Verfiers /0.96
    13. Rubrics1.00
    14. Harness1.00
    15. A simplistic view of a benchmark scaffold1.00
    16. G2i. All rights reserved.1.00
    17. 41.00
    18. World's Fair0.97
    19. TRACK 9 · JULY 1, 20260.93
    20. Posttraining & Midtraining1.00
  • 3:06 #10 skipped

    shot 10·duplicate of #9

  • 3:49 #11 done15 line(s)

    shot 11·sharpness 2169.5

    1. AlEngineer0.98
    2. G2i1.00
    3. World's Fair0.98
    4. When the instrument0.98
    5. is measuring the1.00
    6. 481.6 words / instruction1.00
    7. wrong things1.00
    8. ~2 pages / task0.99
    9. Unrealistic1.00
    10. Instructions1.00
    11. G2i. All rights reserved.1.00
    12. 61.00
    13. World's Fair0.97
    14. TRACK 9 • JULY 1, 20260.93
    15. Posttraining & Midtraining1.00
  • 4:44 #12 done21 line(s)

    shot 12·sharpness 1842.2

    1. AlEngineer0.98
    2. G2i1.00
    3. World'sFair1.00
    4. Leaky prompts1.00
    5. ## Steps to Reproduce0.99
    6. 1. Include an expression with ((regexp.match("foo")}} in Lib/utis/parse.0.97
    7. 1. Run the tests defined1.00
    8. parse_test.go (e.g. TestMatch or TestMatchers)0.97
    9. 1. Notice that compilation fails with errors like undefined: Matcher or undefined: regexpMatcher.0.99
    10. Type: Interface1.00
    11. Name: Matcher0.99
    12. Path: lib/utils/parse/parse.go0.99
    13. Input: in string (for method Match)0.98
    14. Output: bool (indicating if the input matches)0.98
    15. Source: SWE-Bench Pro task0.98
    16. instance_gravitational_teleport-1330415d33a27594c948a36d9d7701f4962291.00
    17. e9f1.00
    18. G2i. All rights reserved.1.00
    19. World'sFair1.00
    20. TRACK 9· JULY 1, 20260.95
    21. Posttraining & Midtraining1.00
  • 5:10 #13 skipped

    shot 13·duplicate of #12

  • 5:54 #14 done33 line(s)

    shot 14·sharpness 2008.9

    1. AlEngineer0.99
    2. G2i1.00
    3. World'sFair1.00
    4. When the1.00
    5. PRESENTED BY1.00
    6. False positive rate0.99
    7. Verifier accepted a wrong implementation1.00
    8. Microsoft1.00
    9. instrument is1.00
    10. SWE-Bench Pro1.00
    11. 8.5%1.00
    12. DeepSWE1.00
    13. 0.3%1.00
    14. 0%1.00
    15. 5%1.00
    16. 10%1.00
    17. miscalibrated1.00
    18. False negative rate0.99
    19. Verifier rejected a correct implementation0.99
    20. SWE-Bench Pro1.00
    21. 24.0%1.00
    22. Weak Verifiers1.00
    23. DeepSWE1.00
    24. 0%1.00
    25. 1.1%1.00
    26. 15%1.00
    27. 30%1.00
    28. Source: https://deepswe.datacurve.ai/blog/deepswe1.00
    29. G2i. All rights reserved.1.00
    30. 81.00
    31. World'sFair1.00
    32. TRACK 9· JULY 1, 20260.95
    33. Posttraining & Midtraining1.00
  • 6:11 #15 done24 line(s)

    shot 15·sharpness 1484.0

    1. AlEngineer0.99
    2. G2i1.00
    3. World's Fair0.99
    4. title: "string literal",1.00
    5. in:1.00
    6. foo0.99
    7. out:1.00
    8. &regexpMatcher{re:0.99
    9. regexp.MustCompile(`^foo$`)},0.97
    10. PRESENTED BY0.97
    11. Microsoft1.00
    12. cmp1.00
    13. AllowUnexported1.00
    14. regexpMatcher{},1.00
    15. prefixSuffixMatcher{},1.00
    16. notMatcher{},1.00
    17. regexp.Regexp{},1.00
    18. ),0.70
    19. SWE-Bench Pro test patch excerpts1.00
    20. G2i. All rights reserved.1.00
    21. 91.00
    22. World's Fair1.00
    23. TRACK 9 · JULY 1, 20260.93
    24. Posttraining & Midtraining1.00
  • 6:46 #16 skipped

    shot 16·duplicate of #15

  • 7:01 #17 done25 line(s)

    shot 17·sharpness 1867.8

    1. AlEngineer0.98
    2. G2i1.00
    3. World's Fair0.99
    4. When the instrument0.99
    5. Standard harnessStrict harness1.00
    6. changes what is being1.00
    7. 100% SWE-bench Multilingual score0.98
    8. 801.00
    9. -9.1%1.00
    10. -7.5%60.93
    11. -0.3%0.99
    12. 601.00
    13. measured1.00
    14. 401.00
    15. 201.00
    16. Reward Hacking1.00
    17. Opus 4.8 Max1.00
    18. Composer 2.51.00
    19. Opus 4.6 Max0.96
    20. Source: https://cursor.com/blog/reward-hacking-coding-benchmarks1.00
    21. G2i. All rights reserved.0.99
    22. 101.00
    23. World's Fair1.00
    24. TRACK 9· JULY 1, 20260.95
    25. Posttraining &Midtraining0.98
  • 7:40 #18 skipped

    shot 18·duplicate of #17

  • 8:05 #19 done62 line(s)

    shot 19·sharpness 1757.0

    1. AlEngineer0.99
    2. G2i1.00
    3. World'sFair1.00
    4. MODEL1.00
    5. STANDARD1.00
    6. STRICT1.00
    7. Δ0.57
    8. Composer 2.51.00
    9. 74.74%1.00
    10. 54.04%1.00
    11. +20.71.00
    12. 21.00
    13. Opus 4.8 (max)0.96
    14. 87.14%1.00
    15. 73.03%1.00
    16. +14.11.00
    17. 31.00
    18. Opus 4.8 (xhigh)1.00
    19. 84.67%1.00
    20. 70.86%1.00
    21. +13.81.00
    22. 41.00
    23. Opus 4.8 (medium)1.00
    24. 76.80%1.00
    25. 67.72%1.00
    26. +9.11.00
    27. 51.00
    28. Opus 4.8 (high)1.00
    29. 78.93%1.00
    30. 69.88%1.00
    31. +9.11.00
    32. 61.00
    33. GPT-5.4 (high)1.00
    34. 60.00%1.00
    35. 53.40%1.00
    36. +6.61.00
    37. 71.00
    38. Opus 4.8 (low)1.00
    39. 71.75%1.00
    40. 66.07%1.00
    41. +5.71.00
    42. 81.00
    43. Opus 4.7 (max)1.00
    44. 69.88%1.00
    45. 64.71%1.00
    46. +5.21.00
    47. 91.00
    48. Opus 4.7 (xhigh)1.00
    49. 67.99%1.00
    50. 62.86%1.00
    51. +5.11.00
    52. 101.00
    53. Opus 4.7 (high)0.98
    54. 67.65%1.00
    55. 62.72%1.00
    56. +4.91.00
    57. Source: https://cursor.com/blog/reward-hacking-coding-benchmarks0.99
    58. G2i. All rights reserved.1.00
    59. 111.00
    60. World's Fair1.00
    61. TRACK 9· JULY 1, 20260.95
    62. Posttraining & Midtraining0.99
  • 8:15 #20 done9 line(s)

    shot 20·sharpness 1249.4

    1. AlEngineer0.98
    2. G2i1.00
    3. World'sFair1.00
    4. This quality gap ⇒ A trust gap.0.99
    5. G2i. All rights reserved.1.00
    6. 121.00
    7. World's Fair0.99
    8. TRACK 9 · JULY 1, 20260.92
    9. Posttraining & Midtraining1.00
  • 8:49 #21 done11 line(s)

    shot 21·sharpness 1302.7

    1. AlEngineer0.99
    2. G2i1.00
    3. World's Fair0.99
    4. Closing the gap1.00
    5. Principles of better1.00
    6. benchmarks1.00
    7. G2i. All rights reserved.1.00
    8. 131.00
    9. World's Fair0.99
    10. TRACK 9 · JULY 1, 20260.93
    11. Posttraining & Midtraining1.00
  • 9:22 #22 done11 line(s)

    shot 22·sharpness 1303.8

    1. AlEngineer0.98
    2. G2i1.00
    3. World'sFair1.00
    4. 1. Human Instructions0.99
    5. Authored by humans.1.00
    6. Reviewed by humans.1.00
    7. G2i. All rights reserved.1.00
    8. 141.00
    9. World's Fair0.99
    10. TRACK 9 · JULY 1, 20260.93
    11. Posttraining & Midtraining1.00
  • 9:43 #23 skipped

    shot 23·duplicate of #22

Transcript

121 cues· 1,623 words· 8,602 chars

  1. 0:12 Hello, everyone.
  2. 0:14 This is the last talk of this session, so hopefully it's going to be short.
  3. 0:17 And now that you guys had to go through a long day, so try to keep it short and light for you all.
  4. 0:23 I'm going to present myself.
  5. 0:25 I'm Ali.
  6. 0:26 I'm the director of AIML at G2I.
  7. 0:30 I have zero experience in ML, so I don't know why they put the ML in my title.
  8. 0:34 I'm a software engineer at heart, and to prove that I have more than 50 abandoned side projects in my machine, so you can know.
  9. 0:42 So I'm going to make a disclaimer that the title of the presentation is a little bit misleading.
  10. 0:49 As I was working on it, I realized that it would be better if I presented my journey into benchmarks and what I learned instead of trying to find a dichotomy of the bad, the ugly, and the good.
  11. 1:02 So let's start with, I wanna grab your attention and I invite you to look at this.
  12. 1:10 This beautiful three screenshots are a single prompt on one of the benchmark tasks.
  13. 1:16 As I was looking at it, I was like, how can an engineer write a task like this?
  14. 1:20 So I said, yeah, it's impossible.
  15. 1:23 No one writes prompts like these ever, but I wanted to double check with my engineers.
  16. 1:28 So I took three of our best engineers.
  17. 1:30 I showed them the prompt and I said, would you ever write a prompt like this?
  18. 1:34 And the answer was.
  19. 1:39 they're right they shouldn't and so at that point i'm i was like what is the what are benchmarks anyway i needed to take a step back i needed to look more i need to understand and so as i was researching i faced a wall of keywords graders long horizon verifiers bench benchmarks and and a lot of jargon so
  20. 2:06 I was like, either this is too complicated or there's a lot of jargon and a lot of words to work through here.
  21. 2:15 So I worked through it, worked with my team.
  22. 2:20 I have a lot of good researchers in the team.
  23. 2:22 And we kind of like simplified to the most basics.
  24. 2:30 The way I see it is that it starts as a prompt or an instruction.
  25. 2:34 That prompt is fed to models and agents.
  26. 2:40 Agents provide solutions.
  27. 2:42 Those solutions are verified and graded through verifiers and rubrics.
  28. 2:48 All of that is wrapped in a harness that's preventing it from the external factors.
  29. 2:56 And if it all goes good, we have
  30. 3:00 trajectories, scores, and metadata that we can use to basically rank models.
  31. 3:13 And so the equation is simple.
  32. 3:15 If prompts and instructions are great, and verifiers and rubrics are doing their job while the harness is preventing or creating an environment that is good for a benchmark,
  33. 3:28 we should have amazing results.
  34. 3:31 But that's not the reality.
  35. 3:33 So what went wrong?
  36. 3:37 So the first thing is, when looking deeper in benchmarks, most of the instructions are unrealistic.
  37. 3:45 I did a quick research on SweetBench Pro, and there's 481 words per instruction in average.
  38. 3:55 That's a two pager per task.
  39. 3:58 That is not how people write prompts.
  40. 4:01 And to illustrate more of that, I took a couple examples here.
  41. 4:07 The first one I called the leaky prompt.
  42. 4:10 It's a goal task that's basically trying to match in some rejects and doing tests on some rejects.
  43. 4:18 So in the first screenshot here, the instruction is pointing directly to the test file, which basically means that the LLM has all the ingredient it needs to go and find that test file and implement based on that.
  44. 4:31 The second one is,
  45. 4:34 is even worse.
  46. 4:35 It's basically providing a complete interface of the implementation, basically locking the LLM from any kind of creativity, and it's forcing it to do it that way.
  47. 4:46 So that's the leaky prompt.
  48. 4:49 The second example, it's the not economically valuable prompt.
  49. 4:54 This is from Sweet Marathon.
  50. 4:57 And this prompt is well formed.

Chapters

  1. 0:00 The good, the bad, and the ugly
  2. 1:27 Testing with our best engineers
  3. 2:30 A benchmark as a spec
  4. 3:37 When instructions are too ambiguous
  5. 4:44 Tests that check the wrong thing
  6. 6:12 Good answers marked wrong
  7. 7:03 Models learning to game the test
  8. 8:08 The quality gap leaderboards hide
  9. 9:03 Precise where it matters
  10. 10:47 Keeping a private held out set
  11. 11:13 Principles for benchmarks worth trusting

Open at this second