Videos O3FEoMYvUf8
Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers
Scene timeline
54 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 199
- whisperx 199
- chunks
- 42
- from 199 cues
- keyframes
- 34
- kept of 54 captured
- frames with text
- 34
- 648 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.2 MB
- word timings on 199 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 03:01 | 1m 17s |
stt |
done | — | 2026-08-11 03:02 | 24s |
chunk |
done | — | 2026-08-11 03:03 | 0s |
text_embed |
done | — | 2026-08-11 03:03 | 1s |
keyframe |
done | — | 2026-08-11 03:03 | 1m 08s |
ocr |
done | — | 2026-08-11 03:04 | 18s |
frame_embed |
done | — | 2026-08-11 03:04 | 6s |
Frames, and what the machine read
-
- AlEngineer1.00
- World'sFair1.00
- Stop evaluating models like1.00
- it's the 50s.1.00
- A century of science measuring minds, applied0.99
- toLLMs.1.00
- What can we borrow from modern psychometrics and1.00
- measurement theory?1.00
- Alejandro Vidal·[email protected]1.00
- Stop evaluating models like it's the 50s. Alejandro Vidal. @dobleio0.98
-
- AlEngineer0.99
- World'sFair1.00
- A benchmark is one number. It shouldn't be.0.99
- Gemma 4 31B0.97
- o4-mini (high)0.97
- DeepSeek-R1 (0528)1.00
- GLM-4.71.00
- Grok41.00
- GPT-5 (high)0.99
- Gemini 2.5 Pro1.00
- GPT-5.5 Pro (xhigh)1.00
- Gemini 3 Pro0.97
- Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.97
-
- World'sFair0.99
- A benchmark is one number. It shouldn't be.1.00
- That sum only holds under one assumption: all items weigh the same.0.99
- Gemma 4 31B0.97
- o4-mini(high)1.00
- DeepSeek-R1 (0528)1.00
- GLM-4.71.00
- Grok 40.93
- GPT-5 (high)0.99
- Gemini 2.5 Pro1.00
- GPT-5.5 Pro (xhigh)1.00
- Gemini 3 Pro1.00
- Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
-
- AlEngineer1.00
- World's Fair1.00
- A benchmark is one number. It shouldn't be.0.99
- Now model each item with a curve, an Item Response Function. Each item has a0.99
- difficulty b: P of a correct answer crosses one half exactly at ability equal to b.1.00
- Gemma 4 31B0.98
- o4-mini (high)1.00
- DeepSeek-R1 (0528)1.00
- GLM-4.71.00
- Grok 41.00
- GPT-5 (high)1.00
- Gemini 2.5 Pro1.00
- GPT-5.5 Pro (xhigh)1.00
- Gemini 3 Pro1.00
- on o0.74
- P=11.00
- P=0.51.00
- qod0.63
- = -2.050.98
- LLM's intelligence1.00
- Stop evaluating models like it's the 50s0.99
- Alejandro Vidal1.00
- @dobleio1.00
-
- AlEngineer1.00
- World's Fair1.00
- A benchmark is one number. It shouldn't be.0.99
- Each item has a difficulty b; each model a skill level theta. Both live on one0.99
- b~θ~N0.91
- shared scale.1.00
- Gemma 4 31B0.97
- -1.741.00
- o4-mini (high)1.00
- -1.150.99
- DeepSeek-R1 (0528)1.00
- -0.921.00
- GLM-4.71.00
- -0.011.00
- Grok 40.99
- +0.181.00
- GPT-5 (high)1.00
- +0.511.00
- Gemini 2.5 Pro1.00
- = +0.640.97
- GPT-5.5 Pro (xhigh)1.00
- = +1.200.99
- Gemini 3 Pro1.00
- = +1.290.97
- P=11.00
- P = 99%0.98
- ot weor0.75
- P=0.51.00
- qood0.73
- P=00.99
- b = -1.230.98
- LLM's intelligence1.00
- θ = 1.200.93
- Stop evaluating models like it's the 50s0.99
- Alejandro Vidal1.00
- @dobleiPtem 1f13aeaea4 . b = -1.230.95
- GPT-5.5 Pro (xhigh) . θ = 1.20: ✓ correct0.92
-
- AlEngineer1.00
- World'sFair1.00
- IRT modeling: a visual introduction.1.00
- 2PL1.00
- Steeper a concentrates more information at b (the shadow).0.99
- Flat ≈ noise.0.98
- 1.01.00
- a discrimination+0.811.00
- 0.51.00
- b difficulty1.00
- +0.001.00
- P(cret)0.72
- Fisher information I(θ) = a².P(1-P)0.95
- 0.01.00
- 00.98
- +21.00
- +41.00
- θ ability0.96
- Stop evaluating models like it's the 50s . Alejandro Vidal . @dobleio0.96
-
- AlEngineer0.99
- World'sFair1.00
- Estimating ability: every answer bends the likelihood.0.98
- Which θ makes this answer most probable?0.99
- item curves P1(0)0.94
- reading Grok 4's answers → = argmax L(0)0.97
- responses - Grok 40.97
- item c1045fe954· a=1.81· b=-0.96 ·0.94
- item 549afc0fe1· a=1.87. b=0.97· x0.94
- 40.68
- relative likelihood R(θ) = L(θ)/L(θ)0.96
- θ = +0.020.88
- -21.00
- 21.00
- abilityθ→0.99
- Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
-
- AlEnginee0.99
- World'sFair1.00
- Stop evaluating models like it's the 50s · Alejandro Vidal· @dobleio0.96
-
- AlEngineer1.00
- World'sFair1.00
- Same score, different ability.0.99
- likelihood L(θ) - normalized0.95
- SWE-bench Verified· 337 tas0.97
- ê +0.800.95
- ê +1.510.95
- model1.00
- score1.00
- Claude Opus 4.10.96
- 245/0.97
- +0.801.00
- Gemini 3 Pro1.00
- 247+1.511.00
- +2√0.80
- +0.711.00
- 40.79
- 05% likelihood interval . R(0) ≥ 0.150.97
- 0.51.00
- 11.00
- 1.51.00
- 21.00
- ability θ →0.96
- Stop evaluating models like it's the 50s. Alejandro Vidal . @dobleio0.96
-
- AlEngineer0.99
- World'sFair0.98
- SKILL1.00
- /aipsychometrics:audit-items1.00
- Audit your benchmark.0.98
- Stop evaluating models like it's the 50s. Alejandro Vidal @dobleio0.98
-
- AlEngineer1.00
- World'sFair0.99
- SKILL1.00
- /aipsychometrics:audit-items1.00
- Audit your benchmark.1.00
- Stop evaluating models like it's the 50s . Alejandro Vidal . @dobleio0.96
-
- AlEngineer1.00
- World's Fair0.96
- Audit your benchmark. The bad items raise their hand.1.00
- Rank every item by its discrimination a.1.00
- SimpleQA-verified1.00
- SWE-bench Verified0.96
- GPQA Diamond1.00
- AIME1.00
- n = 1000 items1.00
- n = 484 items0.99
- n = 198 items1.00
- n = 45 items0.97
- SimpleQA-verified· a ∈ [-2, -1.5)0.96
- sbility θ - . P(correct) ↑0.85
- -20.77
- 1 item · a < 0 · stronger models do worse (flag0.96
- 21.00
- 21.00
- -21.00
- only if CI excludes 0)0.99
- item discrimination a→1.00
- item discrimination a→0.98
- item discrimination a→0.98
- Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
-
- AlEngineer0.94
- World'sFair0.96
- Are the flagged items actually broken?0.98
- Q: Where was the first B.A.S.S. Bassmaster Tournament held?0.99
- gold (answer key): Lake Mead ×0.97
- correct: Beaver Lake, Arkansas ✓0.96
- Q: What is the total number of passengers that died when KLM Flight 4805 and Pan Am Flight 1736 collided?0.99
- gold (answer key): 583 (total people killed: passengers + crew) ×0.98
- correct: 560 passengers √0.96
- SimpleQA-verified1.00
- Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.97
Transcript
199 cues· 3,781 words· 20,149 chars
- 0:02 Hi everyone, I'm Alejandro Vidal, the founder of Mind Makers and my background is psychology and computer science, which is kind of weird but for today is going to be extremely helpful because I'm going to show you how you can borrow ideas from psychology and psychometrics to improve the way that you are evaluating models right now.
- 0:21 because at this moment the state in the industry is counting the number of right answers.
- 0:27 That actually has a name, it's classical test theory, and we have by far better tools to do that.
- 0:33 So it makes sense to borrow ideas from IQ tests and related stuff so we can apply them to LLMs.
- 0:41 Let me start with a very simple example here.
- 0:44 We are using real data from epoch.ai.
- 0:47 If you don't know them, their project is amazing and they have quite open data sets so you can actually use them.
- 0:55 And here we have a random selection of models with a real benchmark.
- 1:00 As you can see here, we have an accuracy for each one of them.
- 1:03 That's the current state of the art.
- 1:05 So if we split its bar into different questions, each one of them is going to be a different question or a different item.
- 1:14 I'm going to use item for the same idea of question.
- 1:19 In psychometrics, we use item instead of question.
- 1:23 If you sum all together, we are using a very strong assumption.
- 1:27 We are saying that every question is equally important.
- 1:30 They should weigh the same, which is kind of insane if you think about that.
- 1:34 We have better questions, more complicated questions that maybe we should pay more attention to.
- 1:42 And also we can have questions that are mislabeled or something like that.
- 1:46 So we are going to improve this.
- 1:48 What are we going to do is we are going to use each item, each column here is going to be one item and we are going to treat them as individual variables.
- 1:58 So we are going to have this matrix here.
- 2:00 As you can see here on the top right corner we have difficult questions for weaker models and on the other side we have
- 2:10 very easy questions for strong models.
- 2:12 So, it makes sense that we observe this pattern, okay?
- 2:16 But we are going to estimate for each question, for each item, a difficulty level.
- 2:21 That is going to be called B.
- 2:24 The B parameter is going to be the difficulty of each one of them and we are going to create a function for each question.
- 2:32 that function maps the LLM intelligence to the probability of getting that answer right.
- 2:40 So, very easy items are going to be here and extremely complicated items are going to be there.
- 2:46 As you can see here, B is the point that crosses 50% chance in that curve, which is going to be useful later.
- 2:56 Also, B is going to be distributed by a normal distribution, which is going to be also helpful to use that for interpretation, okay?
- 3:04 So, with that in mind, we can actually estimate also theta.
- 3:09 Theta is going to be the level of intelligence for each model.
- 3:13 That's going to be that dot, that black dot.
- 3:15 So, as you can see on the right side of each dot,
- 3:19 mostly all the questions are going to be read, which makes sense if they are extremely complicated or more complicated than the level of intelligence of that model, the model is going to fail them, okay?
- 3:32 So we are going to model that way.
- 3:34 So for example here, if I click on this button, I'm going to see that GPT-5.5 here is going to be able to answer that question because GPT-5 has a theta value of 1.2.
- 3:48 and the difficulty of that item is minus 1.2 so the probability of the right answer is 99 okay so with that in mind what are we doing right here is actually calibrating each question each item so we are going to improve a lot our estimations we're going to improve also our confident intervals and many other properties this the other thing that i want to explain here is theta and b is going to be a pair of numbers
- 4:16 that are distributed with normal distributions, so we can actually interpret them.
- 4:20 For example, an item of b equals 0 means that it's going to be average.
- 4:26 Half of the models in my dataset are going to be able to answer that question 50% of the time.
- 4:33 So that's going to be extremely helpful because right now, to evaluate
- 4:38 benchmarks, we need that reference compared with other models.
- 4:41 So it's better to have one by default with this methodology.
- 4:46 Actually, this model is called item response theory, which is the evolution of classical test theory.
- 4:54 So on top of that, I'm going to have another parameter.
- 4:57 It's going to be the slope, the discrimination of that item.
- 5:01 So high discrimination are going to have steeper functions.
- 5:06 Also, we can have
loading