read-only demo

Videos O3FEoMYvUf8

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

index_state ready data_status ok

AI Engineer· published 2026-07-13· 0:23:34· en-US· indexed 2026-08-11 03:04

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:40, 1 of 1 keyframes kept
  2. Shot 1, 0:40 to 1:07, 1 of 1 keyframes kept
  3. Shot 2, 1:07 to 1:33, 1 of 1 keyframes kept
  4. Shot 3, 1:33 to 1:59, 0 of 1 keyframes kept
  5. Shot 4, 1:59 to 2:26, 0 of 1 keyframes kept
  6. Shot 5, 2:26 to 2:52, 1 of 1 keyframes kept
  7. Shot 6, 2:52 to 3:23, 0 of 1 keyframes kept
  8. Shot 7, 3:23 to 3:53, 1 of 1 keyframes kept
  9. Shot 8, 3:53 to 4:24, 0 of 1 keyframes kept
  10. Shot 9, 4:24 to 4:54, 0 of 1 keyframes kept
  11. Shot 10, 4:54 to 5:31, 1 of 1 keyframes kept
  12. Shot 11, 5:31 to 6:05, 1 of 1 keyframes kept
  13. Shot 12, 6:05 to 6:39, 0 of 1 keyframes kept
  14. Shot 13, 6:39 to 6:42, 1 of 1 keyframes kept
  15. Shot 14, 6:42 to 7:09, 1 of 1 keyframes kept
  16. Shot 15, 7:09 to 7:36, 0 of 1 keyframes kept
  17. Shot 16, 7:36 to 8:03, 0 of 1 keyframes kept
  18. Shot 17, 8:03 to 8:18, 0 of 1 keyframes kept
  19. Shot 18, 8:18 to 8:19, 1 of 1 keyframes kept
  20. Shot 19, 8:19 to 8:43, 1 of 1 keyframes kept
  21. Shot 20, 8:43 to 9:27, 1 of 1 keyframes kept
  22. Shot 21, 9:27 to 9:32, 0 of 1 keyframes kept
  23. Shot 22, 9:32 to 9:59, 1 of 1 keyframes kept
  24. Shot 23, 9:59 to 10:27, 0 of 1 keyframes kept
  25. Shot 24, 10:27 to 10:28, 0 of 1 keyframes kept
  26. Shot 25, 10:28 to 11:12, 0 of 1 keyframes kept
  27. Shot 26, 11:12 to 11:41, 1 of 1 keyframes kept
  28. Shot 27, 11:41 to 12:10, 1 of 1 keyframes kept
  29. Shot 28, 12:10 to 12:40, 0 of 1 keyframes kept
  30. Shot 29, 12:40 to 13:09, 0 of 1 keyframes kept
  31. Shot 30, 13:09 to 13:38, 1 of 1 keyframes kept
  32. Shot 31, 13:38 to 13:40, 1 of 1 keyframes kept
  33. Shot 32, 13:40 to 14:02, 1 of 1 keyframes kept
  34. Shot 33, 14:02 to 14:32, 1 of 1 keyframes kept
  35. Shot 34, 14:32 to 15:01, 1 of 1 keyframes kept
  36. Shot 35, 15:01 to 15:31, 1 of 1 keyframes kept
  37. Shot 36, 15:31 to 16:01, 0 of 1 keyframes kept
  38. Shot 37, 16:01 to 16:30, 0 of 1 keyframes kept
  39. Shot 38, 16:30 to 16:31, 1 of 1 keyframes kept
  40. Shot 39, 16:31 to 16:59, 1 of 1 keyframes kept
  41. Shot 40, 16:59 to 17:46, 1 of 1 keyframes kept
  42. Shot 41, 17:46 to 18:29, 1 of 1 keyframes kept
  43. Shot 42, 18:29 to 18:31, 1 of 1 keyframes kept
  44. Shot 43, 18:31 to 18:44, 1 of 1 keyframes kept
  45. Shot 44, 18:44 to 19:29, 1 of 1 keyframes kept
  46. Shot 45, 19:29 to 20:00, 1 of 1 keyframes kept
  47. Shot 46, 20:00 to 20:02, 1 of 1 keyframes kept
  48. Shot 47, 20:02 to 20:25, 1 of 1 keyframes kept
  49. Shot 48, 20:25 to 20:58, 1 of 1 keyframes kept
  50. Shot 49, 20:58 to 21:31, 0 of 1 keyframes kept
  51. Shot 50, 21:31 to 22:04, 0 of 1 keyframes kept
  52. Shot 51, 22:04 to 22:34, 1 of 1 keyframes kept
  53. Shot 52, 22:34 to 23:04, 0 of 1 keyframes kept
  54. Shot 53, 23:04 to 23:34, 1 of 1 keyframes kept

54 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
199
whisperx 199
chunks
42
from 199 cues
keyframes
34
kept of 54 captured
frames with text
34
648 lines read
chapters
0
from the source metadata
keyframe bytes
5.2 MB
word timings on 199 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 03:01 1m 17s
stt done 2026-08-11 03:02 24s
chunk done 2026-08-11 03:03 0s
text_embed done 2026-08-11 03:03 1s
keyframe done 2026-08-11 03:03 1m 08s
ocr done 2026-08-11 03:04 18s
frame_embed done 2026-08-11 03:04 6s

Frames, and what the machine read

  • 0:35 #0 done10 line(s)

    shot 0·sharpness 1192.1

    1. AlEngineer1.00
    2. World'sFair1.00
    3. Stop evaluating models like1.00
    4. it's the 50s.1.00
    5. A century of science measuring minds, applied0.99
    6. toLLMs.1.00
    7. What can we borrow from modern psychometrics and1.00
    8. measurement theory?1.00
    9. Alejandro Vidal·[email protected]1.00
    10. Stop evaluating models like it's the 50s. Alejandro Vidal. @dobleio0.98
  • 0:46 #1 done13 line(s)

    shot 1·sharpness 1445.2

    1. AlEngineer0.99
    2. World'sFair1.00
    3. A benchmark is one number. It shouldn't be.0.99
    4. Gemma 4 31B0.97
    5. o4-mini (high)0.97
    6. DeepSeek-R1 (0528)1.00
    7. GLM-4.71.00
    8. Grok41.00
    9. GPT-5 (high)0.99
    10. Gemini 2.5 Pro1.00
    11. GPT-5.5 Pro (xhigh)1.00
    12. Gemini 3 Pro0.97
    13. Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.97
  • 1:12 #2 done13 line(s)

    shot 2·sharpness 2534.6

    1. World'sFair0.99
    2. A benchmark is one number. It shouldn't be.1.00
    3. That sum only holds under one assumption: all items weigh the same.0.99
    4. Gemma 4 31B0.97
    5. o4-mini(high)1.00
    6. DeepSeek-R1 (0528)1.00
    7. GLM-4.71.00
    8. Grok 40.93
    9. GPT-5 (high)0.99
    10. Gemini 2.5 Pro1.00
    11. GPT-5.5 Pro (xhigh)1.00
    12. Gemini 3 Pro1.00
    13. Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
  • 1:54 #3 skipped

    shot 3·duplicate of #1

  • 2:08 #4 skipped

    shot 4·duplicate of #1

  • 2:42 #5 done23 line(s)

    shot 5·sharpness 2118.3

    1. AlEngineer1.00
    2. World's Fair1.00
    3. A benchmark is one number. It shouldn't be.0.99
    4. Now model each item with a curve, an Item Response Function. Each item has a0.99
    5. difficulty b: P of a correct answer crosses one half exactly at ability equal to b.1.00
    6. Gemma 4 31B0.98
    7. o4-mini (high)1.00
    8. DeepSeek-R1 (0528)1.00
    9. GLM-4.71.00
    10. Grok 41.00
    11. GPT-5 (high)1.00
    12. Gemini 2.5 Pro1.00
    13. GPT-5.5 Pro (xhigh)1.00
    14. Gemini 3 Pro1.00
    15. on o0.74
    16. P=11.00
    17. P=0.51.00
    18. qod0.63
    19. = -2.050.98
    20. LLM's intelligence1.00
    21. Stop evaluating models like it's the 50s0.99
    22. Alejandro Vidal1.00
    23. @dobleio1.00
  • 2:59 #6 skipped

    shot 6·duplicate of #5

  • 3:44 #7 done37 line(s)

    shot 7·sharpness 2360.8

    1. AlEngineer1.00
    2. World's Fair1.00
    3. A benchmark is one number. It shouldn't be.0.99
    4. Each item has a difficulty b; each model a skill level theta. Both live on one0.99
    5. b~θ~N0.91
    6. shared scale.1.00
    7. Gemma 4 31B0.97
    8. -1.741.00
    9. o4-mini (high)1.00
    10. -1.150.99
    11. DeepSeek-R1 (0528)1.00
    12. -0.921.00
    13. GLM-4.71.00
    14. -0.011.00
    15. Grok 40.99
    16. +0.181.00
    17. GPT-5 (high)1.00
    18. +0.511.00
    19. Gemini 2.5 Pro1.00
    20. = +0.640.97
    21. GPT-5.5 Pro (xhigh)1.00
    22. = +1.200.99
    23. Gemini 3 Pro1.00
    24. = +1.290.97
    25. P=11.00
    26. P = 99%0.98
    27. ot weor0.75
    28. P=0.51.00
    29. qood0.73
    30. P=00.99
    31. b = -1.230.98
    32. LLM's intelligence1.00
    33. θ = 1.200.93
    34. Stop evaluating models like it's the 50s0.99
    35. Alejandro Vidal1.00
    36. @dobleiPtem 1f13aeaea4 . b = -1.230.95
    37. GPT-5.5 Pro (xhigh) . θ = 1.20: ✓ correct0.92
  • 4:20 #8 skipped

    shot 8·duplicate of #5

  • 4:50 #9 skipped

    shot 9·duplicate of #5

  • 5:23 #10 done19 line(s)

    shot 10·sharpness 939.5

    1. AlEngineer1.00
    2. World'sFair1.00
    3. IRT modeling: a visual introduction.1.00
    4. 2PL1.00
    5. Steeper a concentrates more information at b (the shadow).0.99
    6. Flat ≈ noise.0.98
    7. 1.01.00
    8. a discrimination+0.811.00
    9. 0.51.00
    10. b difficulty1.00
    11. +0.001.00
    12. P(cret)0.72
    13. Fisher information I(θ) = a².P(1-P)0.95
    14. 0.01.00
    15. 00.98
    16. +21.00
    17. +41.00
    18. θ ability0.96
    19. Stop evaluating models like it's the 50s . Alejandro Vidal . @dobleio0.96
  • 6:01 #11 done16 line(s)

    shot 11·sharpness 1616.2

    1. AlEngineer0.99
    2. World'sFair1.00
    3. Estimating ability: every answer bends the likelihood.0.98
    4. Which θ makes this answer most probable?0.99
    5. item curves P1(0)0.94
    6. reading Grok 4's answers → = argmax L(0)0.97
    7. responses - Grok 40.97
    8. item c1045fe954· a=1.81· b=-0.96 ·0.94
    9. item 549afc0fe1· a=1.87. b=0.97· x0.94
    10. 40.68
    11. relative likelihood R(θ) = L(θ)/L(θ)0.96
    12. θ = +0.020.88
    13. -21.00
    14. 21.00
    15. abilityθ→0.99
    16. Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
  • 6:28 #12 skipped

    shot 12·duplicate of #11

  • 6:41 #13 done3 line(s)

    shot 13·sharpness 186.3

    1. AlEnginee0.99
    2. World'sFair1.00
    3. Stop evaluating models like it's the 50s · Alejandro Vidal· @dobleio0.96
  • 6:56 #14 done24 line(s)

    shot 14·sharpness 992.8

    1. AlEngineer1.00
    2. World'sFair1.00
    3. Same score, different ability.0.99
    4. likelihood L(θ) - normalized0.95
    5. SWE-bench Verified· 337 tas0.97
    6. ê +0.800.95
    7. ê +1.510.95
    8. model1.00
    9. score1.00
    10. Claude Opus 4.10.96
    11. 245/0.97
    12. +0.801.00
    13. Gemini 3 Pro1.00
    14. 247+1.511.00
    15. +2√0.80
    16. +0.711.00
    17. 40.79
    18. 05% likelihood interval . R(0) ≥ 0.150.97
    19. 0.51.00
    20. 11.00
    21. 1.51.00
    22. 21.00
    23. ability θ →0.96
    24. Stop evaluating models like it's the 50s. Alejandro Vidal . @dobleio0.96
  • 7:28 #15 skipped

    shot 15·duplicate of #14

  • 8:00 #16 skipped

    shot 16·duplicate of #14

  • 8:06 #17 skipped

    shot 17·duplicate of #14

  • 8:19 #18 done6 line(s)

    shot 18·sharpness 665.5

    1. AlEngineer0.99
    2. World'sFair0.98
    3. SKILL1.00
    4. /aipsychometrics:audit-items1.00
    5. Audit your benchmark.0.98
    6. Stop evaluating models like it's the 50s. Alejandro Vidal @dobleio0.98
  • 8:36 #19 done6 line(s)

    shot 19·sharpness 746.1

    1. AlEngineer1.00
    2. World'sFair0.99
    3. SKILL1.00
    4. /aipsychometrics:audit-items1.00
    5. Audit your benchmark.1.00
    6. Stop evaluating models like it's the 50s . Alejandro Vidal . @dobleio0.96
  • 9:09 #20 done24 line(s)

    shot 20·sharpness 1502.5

    1. AlEngineer1.00
    2. World's Fair0.96
    3. Audit your benchmark. The bad items raise their hand.1.00
    4. Rank every item by its discrimination a.1.00
    5. SimpleQA-verified1.00
    6. SWE-bench Verified0.96
    7. GPQA Diamond1.00
    8. AIME1.00
    9. n = 1000 items1.00
    10. n = 484 items0.99
    11. n = 198 items1.00
    12. n = 45 items0.97
    13. SimpleQA-verified· a ∈ [-2, -1.5)0.96
    14. sbility θ - . P(correct) ↑0.85
    15. -20.77
    16. 1 item · a < 0 · stronger models do worse (flag0.96
    17. 21.00
    18. 21.00
    19. -21.00
    20. only if CI excludes 0)0.99
    21. item discrimination a→1.00
    22. item discrimination a→0.98
    23. item discrimination a→0.98
    24. Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.98
  • 9:31 #21 skipped

    shot 21·duplicate of #13

  • 9:38 #22 done11 line(s)

    shot 22·sharpness 1807.4

    1. AlEngineer0.94
    2. World'sFair0.96
    3. Are the flagged items actually broken?0.98
    4. Q: Where was the first B.A.S.S. Bassmaster Tournament held?0.99
    5. gold (answer key): Lake Mead ×0.97
    6. correct: Beaver Lake, Arkansas ✓0.96
    7. Q: What is the total number of passengers that died when KLM Flight 4805 and Pan Am Flight 1736 collided?0.99
    8. gold (answer key): 583 (total people killed: passengers + crew) ×0.98
    9. correct: 560 passengers √0.96
    10. SimpleQA-verified1.00
    11. Stop evaluating models like it's the 50s. Alejandro Vidal· @dobleio0.97
  • 10:08 #23 skipped

    shot 23·duplicate of #22

Transcript

199 cues· 3,781 words· 20,149 chars

  1. 0:02 Hi everyone, I'm Alejandro Vidal, the founder of Mind Makers and my background is psychology and computer science, which is kind of weird but for today is going to be extremely helpful because I'm going to show you how you can borrow ideas from psychology and psychometrics to improve the way that you are evaluating models right now.
  2. 0:21 because at this moment the state in the industry is counting the number of right answers.
  3. 0:27 That actually has a name, it's classical test theory, and we have by far better tools to do that.
  4. 0:33 So it makes sense to borrow ideas from IQ tests and related stuff so we can apply them to LLMs.
  5. 0:41 Let me start with a very simple example here.
  6. 0:44 We are using real data from epoch.ai.
  7. 0:47 If you don't know them, their project is amazing and they have quite open data sets so you can actually use them.
  8. 0:55 And here we have a random selection of models with a real benchmark.
  9. 1:00 As you can see here, we have an accuracy for each one of them.
  10. 1:03 That's the current state of the art.
  11. 1:05 So if we split its bar into different questions, each one of them is going to be a different question or a different item.
  12. 1:14 I'm going to use item for the same idea of question.
  13. 1:19 In psychometrics, we use item instead of question.
  14. 1:23 If you sum all together, we are using a very strong assumption.
  15. 1:27 We are saying that every question is equally important.
  16. 1:30 They should weigh the same, which is kind of insane if you think about that.
  17. 1:34 We have better questions, more complicated questions that maybe we should pay more attention to.
  18. 1:42 And also we can have questions that are mislabeled or something like that.
  19. 1:46 So we are going to improve this.
  20. 1:48 What are we going to do is we are going to use each item, each column here is going to be one item and we are going to treat them as individual variables.
  21. 1:58 So we are going to have this matrix here.
  22. 2:00 As you can see here on the top right corner we have difficult questions for weaker models and on the other side we have
  23. 2:10 very easy questions for strong models.
  24. 2:12 So, it makes sense that we observe this pattern, okay?
  25. 2:16 But we are going to estimate for each question, for each item, a difficulty level.
  26. 2:21 That is going to be called B.
  27. 2:24 The B parameter is going to be the difficulty of each one of them and we are going to create a function for each question.
  28. 2:32 that function maps the LLM intelligence to the probability of getting that answer right.
  29. 2:40 So, very easy items are going to be here and extremely complicated items are going to be there.
  30. 2:46 As you can see here, B is the point that crosses 50% chance in that curve, which is going to be useful later.
  31. 2:56 Also, B is going to be distributed by a normal distribution, which is going to be also helpful to use that for interpretation, okay?
  32. 3:04 So, with that in mind, we can actually estimate also theta.
  33. 3:09 Theta is going to be the level of intelligence for each model.
  34. 3:13 That's going to be that dot, that black dot.
  35. 3:15 So, as you can see on the right side of each dot,
  36. 3:19 mostly all the questions are going to be read, which makes sense if they are extremely complicated or more complicated than the level of intelligence of that model, the model is going to fail them, okay?
  37. 3:32 So we are going to model that way.
  38. 3:34 So for example here, if I click on this button, I'm going to see that GPT-5.5 here is going to be able to answer that question because GPT-5 has a theta value of 1.2.
  39. 3:48 and the difficulty of that item is minus 1.2 so the probability of the right answer is 99 okay so with that in mind what are we doing right here is actually calibrating each question each item so we are going to improve a lot our estimations we're going to improve also our confident intervals and many other properties this the other thing that i want to explain here is theta and b is going to be a pair of numbers
  40. 4:16 that are distributed with normal distributions, so we can actually interpret them.
  41. 4:20 For example, an item of b equals 0 means that it's going to be average.
  42. 4:26 Half of the models in my dataset are going to be able to answer that question 50% of the time.
  43. 4:33 So that's going to be extremely helpful because right now, to evaluate
  44. 4:38 benchmarks, we need that reference compared with other models.
  45. 4:41 So it's better to have one by default with this methodology.
  46. 4:46 Actually, this model is called item response theory, which is the evolution of classical test theory.
  47. 4:54 So on top of that, I'm going to have another parameter.
  48. 4:57 It's going to be the slope, the discrimination of that item.
  49. 5:01 So high discrimination are going to have steeper functions.
  50. 5:06 Also, we can have

Open at this second