Videos jWq-aZIU0kM
Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i
Scene timeline
32 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 121
- whisperx 121
- chunks
- 23
- from 121 cues
- keyframes
- 24
- kept of 32 captured
- frames with text
- 24
- 405 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 3.4 MB
- word timings on 121 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:43 | 0s |
stt |
done | — | 2026-08-09 02:45 | 11s |
chunk |
done | — | 2026-08-09 02:46 | 0s |
text_embed |
done | — | 2026-08-10 19:37 | 0s |
keyframe |
done | — | 2026-08-09 02:46 | 1m 36s |
ocr |
done | — | 2026-08-09 02:47 | 7s |
frame_embed |
done | — | 2026-08-10 19:37 | 4s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.96
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.97
- Amazon AGI Lab0.99
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo0.99
- reducto1.00
- Sonar1.00
- Makers of1.00
- together.ai0.98
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.99
- World's Fair0.98
-
- AlEngineer0.98
- G2i1.00
- World's Fair0.96
- Ali Khial1.00
- PRESENTED BY1.00
- Director of Al/ML @ G2i1.00
- Microsoft1.00
- Software Engineer @1.00
- 50+ abandoned side projects1.00
- G2i. All rights reserved.1.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.94
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- A SWE-bench pro task instructions1.00
- G2i. All rights reserved.1.00
- 11.00
- World'sFair1.00
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- G2i1.00
- World'sFair1.00
- No one1.00
- writes1.00
- prompts like0.99
- WHAT?1.00
- this. ever!1.00
- What the pro*pt ?!?1.00
- G2i. All rights reserved.1.00
- 21.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World'sFair1.00
- What are1.00
- benchmarks ?0.96
- G2i. All rights reserved.1.00
- 31.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- G2i1.00
- World's Fair0.98
- Graders1.00
- LLM-as-judge1.00
- What are1.00
- Long Horizon1.00
- Verfiers1.00
- benchmarks ?0.98
- N-gram1.00
- Fail-to-pass1.00
- Benchmaxxing1.00
- Evals1.00
- Pass-to-pass1.00
- Hill Climbing0.99
- G2i. All rights reserved.1.00
- 31.00
- World's Fair0.97
- TRACK 9 • JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World's Fair0.99
- Prompts/1.00
- Models /0.94
- Solutions1.00
- Instructions1.00
- Agents1.00
- Trajectories1.00
- Scores1.00
- Metadata1.00
- Verfiers /0.96
- Rubrics1.00
- Harness1.00
- A simplistic view of a benchmark scaffold1.00
- G2i. All rights reserved.1.00
- 41.00
- World's Fair0.97
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World's Fair0.98
- When the instrument0.98
- is measuring the1.00
- 481.6 words / instruction1.00
- wrong things1.00
- ~2 pages / task0.99
- Unrealistic1.00
- Instructions1.00
- G2i. All rights reserved.1.00
- 61.00
- World's Fair0.97
- TRACK 9 • JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World'sFair1.00
- Leaky prompts1.00
- ## Steps to Reproduce0.99
- 1. Include an expression with ((regexp.match("foo")}} in Lib/utis/parse.0.97
- 1. Run the tests defined1.00
- parse_test.go (e.g. TestMatch or TestMatchers)0.97
- 1. Notice that compilation fails with errors like undefined: Matcher or undefined: regexpMatcher.0.99
- Type: Interface1.00
- Name: Matcher0.99
- Path: lib/utils/parse/parse.go0.99
- Input: in string (for method Match)0.98
- Output: bool (indicating if the input matches)0.98
- Source: SWE-Bench Pro task0.98
- instance_gravitational_teleport-1330415d33a27594c948a36d9d7701f4962291.00
- e9f1.00
- G2i. All rights reserved.1.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- G2i1.00
- World'sFair1.00
- When the1.00
- PRESENTED BY1.00
- False positive rate0.99
- Verifier accepted a wrong implementation1.00
- Microsoft1.00
- instrument is1.00
- SWE-Bench Pro1.00
- 8.5%1.00
- DeepSWE1.00
- 0.3%1.00
- 0%1.00
- 5%1.00
- 10%1.00
- miscalibrated1.00
- False negative rate0.99
- Verifier rejected a correct implementation0.99
- SWE-Bench Pro1.00
- 24.0%1.00
- Weak Verifiers1.00
- DeepSWE1.00
- 0%1.00
- 1.1%1.00
- 15%1.00
- 30%1.00
- Source: https://deepswe.datacurve.ai/blog/deepswe1.00
- G2i. All rights reserved.1.00
- 81.00
- World'sFair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- G2i1.00
- World's Fair0.99
- title: "string literal",1.00
- in:1.00
- foo0.99
- out:1.00
- ®expMatcher{re:0.99
- regexp.MustCompile(`^foo$`)},0.97
- PRESENTED BY0.97
- Microsoft1.00
- cmp1.00
- AllowUnexported1.00
- regexpMatcher{},1.00
- prefixSuffixMatcher{},1.00
- notMatcher{},1.00
- regexp.Regexp{},1.00
- ),0.70
- SWE-Bench Pro test patch excerpts1.00
- G2i. All rights reserved.1.00
- 91.00
- World's Fair1.00
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World's Fair0.99
- When the instrument0.99
- Standard harnessStrict harness1.00
- changes what is being1.00
- 100% SWE-bench Multilingual score0.98
- 801.00
- -9.1%1.00
- -7.5%60.93
- -0.3%0.99
- 601.00
- measured1.00
- 401.00
- 201.00
- Reward Hacking1.00
- Opus 4.8 Max1.00
- Composer 2.51.00
- Opus 4.6 Max0.96
- Source: https://cursor.com/blog/reward-hacking-coding-benchmarks1.00
- G2i. All rights reserved.0.99
- 101.00
- World's Fair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining &Midtraining0.98
-
- AlEngineer0.99
- G2i1.00
- World'sFair1.00
- MODEL1.00
- STANDARD1.00
- STRICT1.00
- Δ0.57
- Composer 2.51.00
- 74.74%1.00
- 54.04%1.00
- +20.71.00
- 21.00
- Opus 4.8 (max)0.96
- 87.14%1.00
- 73.03%1.00
- +14.11.00
- 31.00
- Opus 4.8 (xhigh)1.00
- 84.67%1.00
- 70.86%1.00
- +13.81.00
- 41.00
- Opus 4.8 (medium)1.00
- 76.80%1.00
- 67.72%1.00
- +9.11.00
- 51.00
- Opus 4.8 (high)1.00
- 78.93%1.00
- 69.88%1.00
- +9.11.00
- 61.00
- GPT-5.4 (high)1.00
- 60.00%1.00
- 53.40%1.00
- +6.61.00
- 71.00
- Opus 4.8 (low)1.00
- 71.75%1.00
- 66.07%1.00
- +5.71.00
- 81.00
- Opus 4.7 (max)1.00
- 69.88%1.00
- 64.71%1.00
- +5.21.00
- 91.00
- Opus 4.7 (xhigh)1.00
- 67.99%1.00
- 62.86%1.00
- +5.11.00
- 101.00
- Opus 4.7 (high)0.98
- 67.65%1.00
- 62.72%1.00
- +4.91.00
- Source: https://cursor.com/blog/reward-hacking-coding-benchmarks0.99
- G2i. All rights reserved.1.00
- 111.00
- World's Fair1.00
- TRACK 9· JULY 1, 20260.95
- Posttraining & Midtraining0.99
-
- AlEngineer0.98
- G2i1.00
- World'sFair1.00
- This quality gap ⇒ A trust gap.0.99
- G2i. All rights reserved.1.00
- 121.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.92
- Posttraining & Midtraining1.00
-
- AlEngineer0.99
- G2i1.00
- World's Fair0.99
- Closing the gap1.00
- Principles of better1.00
- benchmarks1.00
- G2i. All rights reserved.1.00
- 131.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
-
- AlEngineer0.98
- G2i1.00
- World'sFair1.00
- 1. Human Instructions0.99
- Authored by humans.1.00
- Reviewed by humans.1.00
- G2i. All rights reserved.1.00
- 141.00
- World's Fair0.99
- TRACK 9 · JULY 1, 20260.93
- Posttraining & Midtraining1.00
Transcript
121 cues· 1,623 words· 8,602 chars
- 0:12 Hello, everyone.
- 0:14 This is the last talk of this session, so hopefully it's going to be short.
- 0:17 And now that you guys had to go through a long day, so try to keep it short and light for you all.
- 0:23 I'm going to present myself.
- 0:25 I'm Ali.
- 0:26 I'm the director of AIML at G2I.
- 0:30 I have zero experience in ML, so I don't know why they put the ML in my title.
- 0:34 I'm a software engineer at heart, and to prove that I have more than 50 abandoned side projects in my machine, so you can know.
- 0:42 So I'm going to make a disclaimer that the title of the presentation is a little bit misleading.
- 0:49 As I was working on it, I realized that it would be better if I presented my journey into benchmarks and what I learned instead of trying to find a dichotomy of the bad, the ugly, and the good.
- 1:02 So let's start with, I wanna grab your attention and I invite you to look at this.
- 1:10 This beautiful three screenshots are a single prompt on one of the benchmark tasks.
- 1:16 As I was looking at it, I was like, how can an engineer write a task like this?
- 1:20 So I said, yeah, it's impossible.
- 1:23 No one writes prompts like these ever, but I wanted to double check with my engineers.
- 1:28 So I took three of our best engineers.
- 1:30 I showed them the prompt and I said, would you ever write a prompt like this?
- 1:34 And the answer was.
- 1:39 they're right they shouldn't and so at that point i'm i was like what is the what are benchmarks anyway i needed to take a step back i needed to look more i need to understand and so as i was researching i faced a wall of keywords graders long horizon verifiers bench benchmarks and and a lot of jargon so
- 2:06 I was like, either this is too complicated or there's a lot of jargon and a lot of words to work through here.
- 2:15 So I worked through it, worked with my team.
- 2:20 I have a lot of good researchers in the team.
- 2:22 And we kind of like simplified to the most basics.
- 2:30 The way I see it is that it starts as a prompt or an instruction.
- 2:34 That prompt is fed to models and agents.
- 2:40 Agents provide solutions.
- 2:42 Those solutions are verified and graded through verifiers and rubrics.
- 2:48 All of that is wrapped in a harness that's preventing it from the external factors.
- 2:56 And if it all goes good, we have
- 3:00 trajectories, scores, and metadata that we can use to basically rank models.
- 3:13 And so the equation is simple.
- 3:15 If prompts and instructions are great, and verifiers and rubrics are doing their job while the harness is preventing or creating an environment that is good for a benchmark,
- 3:28 we should have amazing results.
- 3:31 But that's not the reality.
- 3:33 So what went wrong?
- 3:37 So the first thing is, when looking deeper in benchmarks, most of the instructions are unrealistic.
- 3:45 I did a quick research on SweetBench Pro, and there's 481 words per instruction in average.
- 3:55 That's a two pager per task.
- 3:58 That is not how people write prompts.
- 4:01 And to illustrate more of that, I took a couple examples here.
- 4:07 The first one I called the leaky prompt.
- 4:10 It's a goal task that's basically trying to match in some rejects and doing tests on some rejects.
- 4:18 So in the first screenshot here, the instruction is pointing directly to the test file, which basically means that the LLM has all the ingredient it needs to go and find that test file and implement based on that.
- 4:31 The second one is,
- 4:34 is even worse.
- 4:35 It's basically providing a complete interface of the implementation, basically locking the LLM from any kind of creativity, and it's forcing it to do it that way.
- 4:46 So that's the leaky prompt.
- 4:49 The second example, it's the not economically valuable prompt.
- 4:54 This is from Sweet Marathon.
- 4:57 And this prompt is well formed.
loading
Chapters
- 0:00 The good, the bad, and the ugly
- 1:27 Testing with our best engineers
- 2:30 A benchmark as a spec
- 3:37 When instructions are too ambiguous
- 4:44 Tests that check the wrong thing
- 6:12 Good answers marked wrong
- 7:03 Models learning to game the test
- 8:08 The quality gap leaderboards hide
- 9:03 Precise where it matters
- 10:47 Keeping a private held out set
- 11:13 Principles for benchmarks worth trusting