read-only demo

Videos b_PmGocP4rc

Evaling Video Slop — Maor Bril, Character.ai

index_state ready data_status ok

AI Engineer· published 2026-07-25· 0:23:13· en-US· indexed 2026-08-10 19:40

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:21, 1 of 1 keyframes kept
  5. Shot 4, 0:21 to 0:23, 1 of 1 keyframes kept
  6. Shot 5, 0:23 to 0:49, 1 of 1 keyframes kept
  7. Shot 6, 0:49 to 1:15, 1 of 1 keyframes kept
  8. Shot 7, 1:15 to 1:44, 1 of 1 keyframes kept
  9. Shot 8, 1:44 to 2:13, 1 of 1 keyframes kept
  10. Shot 9, 2:13 to 2:38, 1 of 1 keyframes kept
  11. Shot 10, 2:38 to 3:05, 1 of 1 keyframes kept
  12. Shot 11, 3:05 to 3:31, 0 of 1 keyframes kept
  13. Shot 12, 3:31 to 3:58, 0 of 1 keyframes kept
  14. Shot 13, 3:58 to 4:34, 1 of 1 keyframes kept
  15. Shot 14, 4:34 to 4:36, 1 of 1 keyframes kept
  16. Shot 15, 4:36 to 5:11, 1 of 1 keyframes kept
  17. Shot 16, 5:11 to 5:46, 1 of 1 keyframes kept
  18. Shot 17, 5:46 to 6:00, 1 of 1 keyframes kept
  19. Shot 18, 6:00 to 6:28, 1 of 1 keyframes kept
  20. Shot 19, 6:28 to 6:57, 1 of 1 keyframes kept
  21. Shot 20, 6:57 to 7:28, 1 of 1 keyframes kept
  22. Shot 21, 7:28 to 8:07, 1 of 1 keyframes kept
  23. Shot 22, 8:07 to 8:41, 1 of 1 keyframes kept
  24. Shot 23, 8:41 to 9:15, 0 of 1 keyframes kept
  25. Shot 24, 9:15 to 9:41, 1 of 1 keyframes kept
  26. Shot 25, 9:41 to 10:06, 0 of 1 keyframes kept
  27. Shot 26, 10:06 to 10:31, 1 of 1 keyframes kept
  28. Shot 27, 10:31 to 11:19, 1 of 1 keyframes kept
  29. Shot 28, 11:19 to 11:46, 1 of 1 keyframes kept
  30. Shot 29, 11:46 to 12:12, 0 of 1 keyframes kept
  31. Shot 30, 12:12 to 12:39, 0 of 1 keyframes kept
  32. Shot 31, 12:39 to 13:06, 0 of 1 keyframes kept
  33. Shot 32, 13:06 to 13:31, 1 of 1 keyframes kept
  34. Shot 33, 13:31 to 13:56, 0 of 1 keyframes kept
  35. Shot 34, 13:56 to 14:46, 1 of 1 keyframes kept
  36. Shot 35, 14:46 to 15:32, 1 of 1 keyframes kept
  37. Shot 36, 15:32 to 16:01, 1 of 1 keyframes kept
  38. Shot 37, 16:01 to 16:30, 1 of 1 keyframes kept
  39. Shot 38, 16:30 to 16:59, 1 of 1 keyframes kept
  40. Shot 39, 16:59 to 17:28, 1 of 1 keyframes kept
  41. Shot 40, 17:28 to 17:33, 1 of 1 keyframes kept
  42. Shot 41, 17:33 to 17:59, 1 of 1 keyframes kept
  43. Shot 42, 17:59 to 18:26, 1 of 1 keyframes kept
  44. Shot 43, 18:26 to 18:53, 1 of 1 keyframes kept
  45. Shot 44, 18:53 to 19:20, 1 of 1 keyframes kept
  46. Shot 45, 19:20 to 19:47, 1 of 1 keyframes kept
  47. Shot 46, 19:47 to 20:14, 1 of 1 keyframes kept
  48. Shot 47, 20:14 to 20:41, 1 of 1 keyframes kept
  49. Shot 48, 20:41 to 21:08, 1 of 1 keyframes kept
  50. Shot 49, 21:08 to 21:35, 1 of 1 keyframes kept
  51. Shot 50, 21:35 to 22:02, 1 of 1 keyframes kept
  52. Shot 51, 22:02 to 22:29, 1 of 1 keyframes kept
  53. Shot 52, 22:29 to 22:55, 1 of 1 keyframes kept
  54. Shot 53, 22:55 to 23:12, 0 of 1 keyframes kept

54 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
234
whisperx 234
chunks
40
from 234 cues
keyframes
45
kept of 54 captured
frames with text
45
551 lines read
chapters
11
from the source metadata
keyframe bytes
6.2 MB
word timings on 234 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:43 0s
stt done 2026-08-09 06:57 23s
chunk done 2026-08-09 06:58 0s
text_embed done 2026-08-10 19:39 1s
keyframe done 2026-08-09 06:58 3m 02s
ocr done 2026-08-09 07:01 15s
frame_embed done 2026-08-10 19:39 8s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:19 #3 done2 line(s)

    shot 3·sharpness 294.9

    1. AlEngineer0.99
    2. World's Fair0.98
  • 0:23 #4 done17 line(s)

    shot 4·sharpness 1227.5

    1. AlEngineer0.97
    2. World'sFair1.00
    3. AI ENGINEER WORLD'S FAIR1.00
    4. 20261.00
    5. Evals for AI Video0.98
    6. REC1.00
    7. 00:151.00
    8. How I stopped trusting CLIP scores and trained a reward model0.99
    9. that actually watches the movie.0.98
    10. quality score: undefined1.00
    11. Maor Bril· Character.Al0.99
    12. Judge Judy1.00
    13. open source under Character.AI0.98
    14. EVALS FOR AI VIDEO1.00
    15. 01 / 200.88
    16. World'sFair1.00
    17. Engineering the future of Al1.00
  • 0:36 #5 done12 line(s)

    shot 5·sharpness 1290.7

    1. AlEngineer0.97
    2. World'sFair1.00
    3. Al video got amazing.1.00
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. Judging it didn't.1.00
    7. We can make the footage. We still can't agree on whether it's any good.1.00
    8. EVALS FOR AI VIDEO1.00
    9. 02 /200.92
    10. Thoranate0.56
    11. World'sFair0.99
    12. Engineering the future of Al0.98
  • 1:12 #6 done13 line(s)

    shot 6·sharpness 1108.2

    1. AlEngineer0.97
    2. World'sFair1.00
    3. Al video got amazing.1.00
    4. PRESENTED BY1.00
    5. Microsoft1.00
    6. Judging it didn't.0.98
    7. We can make the footage. We still can't agree on whether it's any good.0.99
    8. EVALS FOR AI VIDEO1.00
    9. 02 /200.92
    10. dinatc0.59
    11. World'sFair1.00
    12. TRACK 5· JULY 1,20260.97
    13. Evals1.00
  • 1:19 #7 done24 line(s)

    shot 7·sharpness 1700.1

    1. AlEngineer1.00
    2. 011.00
    3. THE PROBLEM1.00
    4. World's Fair1.00
    5. The hard part was never making it.1.00
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. Generating video is0.99
    9. Generating good video is1.00
    10. So a human watches1.00
    11. basically free.1.00
    12. not.1.00
    13. everything.1.00
    14. Type a prompt, get 400 clips before0.99
    15. 'Good' is taste, timing, and a0.99
    16. That human is the bottleneck. (Hi. It's0.99
    17. lunch.1.00
    18. director's gut.1.00
    19. me.)1.00
    20. EVALS FOR AI VIDEO1.00
    21. 03 / 200.97
    22. World's Fair0.96
    23. TRACK 5· JULY 1, 20260.95
    24. Evals1.00
  • 1:50 #8 done22 line(s)

    shot 8·sharpness 1569.0

    1. AlEngineer0.99
    2. 011.00
    3. THE PROBLEM1.00
    4. World's Fair0.98
    5. The hard part was never making it.1.00
    6. Generating video is1.00
    7. Generating good video is1.00
    8. So a human watches0.98
    9. basically free.1.00
    10. not.1.00
    11. everything.1.00
    12. Type a prompt, get 400 clips before0.98
    13. 'Good' is taste, timing, and a0.99
    14. That human is the bottleneck. (Hi. It's1.00
    15. lunch.1.00
    16. director's gut.1.00
    17. me.)1.00
    18. EVALS FOR AI VIDEO0.97
    19. 03 / 200.96
    20. World's Fair0.97
    21. TRACK 5 · JULY 1, 20260.91
    22. Evals1.00
  • 2:25 #9 done21 line(s)

    shot 9·sharpness 1466.0

    1. AlEngineer0.97
    2. 1020.87
    3. THE GAP1.00
    4. World's Fair0.97
    5. We're grading video with text-era tools.1.00
    6. WHAT WE REACH FOR0.98
    7. CLIPScore1.00
    8. LPIPS1.00
    9. All born in a world of single frames and1.00
    10. captions.1.00
    11. FVD1.00
    12. FID1.00
    13. Great at pixels and prompt-matching -0.99
    14. prompt adherence1.00
    15. frame consistency1.00
    16. the easy 20%.0.97
    17. EVALS FOR AI VIDEO0.97
    18. 04 / 200.93
    19. World's Fair0.97
    20. TRACK 5· JULY 1, 20260.95
    21. Evals1.00
  • 2:56 #10 done22 line(s)

    shot 10·sharpness 1720.0

    1. AlEngineer0.99
    2. 021.00
    3. THE GAP1.00
    4. World's Fair0.97
    5. They see frames. They can't see the story.0.99
    6. WHAT THEY CATCH1.00
    7. 0.65
    8. WHAT THEY MISS1.00
    9. Per-frame sharpness0.98
    10. X Does it tell the story you meant?0.97
    11. X Does it obey physics?0.95
    12. ✓Frame-to-frame consistency0.98
    13. × Character stays the same person0.95
    14. ✓ Prompt / caption match0.96
    15. X Pacing that actually lands0.95
    16. X Audio synced to the picture0.99
    17. EVALS FOR AI VIDEO0.97
    18. 05 / 200.95
    19. oordinatat0.57
    20. World'sFair1.00
    21. TRACK 5· JULY 1, 20260.95
    22. Evals1.00
  • 3:26 #11 skipped

    shot 11·duplicate of #10

  • 3:35 #12 skipped

    shot 12·duplicate of #10

  • 4:06 #13 done20 line(s)

    shot 13·sharpness 1411.7

    1. AlEngineer0.99
    2. 021.00
    3. THE GAP0.95
    4. World's Fair0.97
    5. LLM judges help — as far as you point them.0.98
    6. Only as good as the axis you0.99
    7. "is it consistent?"0.97
    8. name and the prompt you1.00
    9. write.1.00
    10. "does it match the prompt?"0.99
    11. LLM judge1.00
    12. Ask two judges the same1.00
    13. question, get two moods.0.99
    14. "is it...god?”_(ツ)__0.85
    15. EVALS FOR AI VIDEO0.98
    16. 06 / 200.91
    17. Taorcinat0.58
    18. World's Fair0.95
    19. TRACK 5· JULY 1,20260.97
    20. Evals1.00
  • 4:35 #14 done18 line(s)

    shot 14·sharpness 1732.6

    1. AlEngineer0.97
    2. 031.00
    3. OUR ANSWER0.99
    4. World'sFair1.00
    5. But offline eval was never the point.0.98
    6. The goal: make good video cheap, consistent, and everywhere.0.99
    7. drag the gate back here0.99
    8. GENERATE1.00
    9. EVAL (too late)1.00
    10. Move the high-impact gates first.0.99
    11. Calibrate them on real signals — engagement performance across the platform + human annotation (yes — the Friday0.99
    12. when everyone scores videos).0.99
    13. EVALS FOR AI VIDEO0.97
    14. 08 /200.95
    15. Toordnatts0.55
    16. World'sFair1.00
    17. TRACK 5· JULY 1, 20260.95
    18. Evals0.99
  • 4:43 #15 done22 line(s)

    shot 15·sharpness 1067.6

    1. AlEngineer0.99
    2. 1030.90
    3. OUR ANSWER0.99
    4. World's Fair0.97
    5. So we built Judge Judy.1.00
    6. An open-source video-eval harness. (Yes, that's the real name.)1.00
    7. Existing tools0.98
    8. Judge Judy1.00
    9. Scores1.00
    10. CLIP1.00
    11. . LPIps0.65
    12. VLM judges1.00
    13. multi-axis scoring1.00
    14. per axis, per clip1.00
    15. calibration loop: scores checked against human annotations0.99
    16. still 100% offline0.98
    17. EVALS FOR AI VIDEO0.98
    18. 07 / 200.92
    19. ioordinatc0.54
    20. World's Fair0.97
    21. TRACK 5·JULY 1,20260.98
    22. Evals1.00
  • 5:38 #16 done25 line(s)

    shot 16·sharpness 1198.4

    1. AlEngineer0.99
    2. 031.00
    3. OUR ANSWER1.00
    4. World's Fair0.98
    5. So we built Judge Judy.0.97
    6. An open-source video-eval harness. (Yes, that's the real name.)1.00
    7. PRESENTED BY1.00
    8. Microsoft1.00
    9. Existing tools0.98
    10. Judge Judy1.00
    11. Scores1.00
    12. CLIP1.00
    13. .LpIps0.68
    14. VLM judges1.00
    15. multi-axis scoring1.00
    16. per axis, per clip0.99
    17. G0.51
    18. calibration loop: scores checked against human annotations0.99
    19. still 100% offline0.98
    20. EVALS FOR AI VIDEO0.96
    21. 07 / 200.93
    22. hardinatnat0.58
    23. World's Fair0.99
    24. TRACK 5·JULY 1,20260.97
    25. Evals1.00
  • 5:51 #17 done20 line(s)

    shot 17·sharpness 1880.1

    1. AlEngineer0.98
    2. 1030.84
    3. OUR ANSWER0.99
    4. World'sFair0.97
    5. But offline eval was never the point.0.99
    6. PRESENTED BY1.00
    7. The goal: make good video cheap, consistent, and everywhere.0.99
    8. Microsoft1.00
    9. drag the gate back here0.99
    10. GENERATE1.00
    11. EVAL (too late)1.00
    12. Move the high-impact gates first.0.99
    13. Calibrate them on real signals — engagement performance across the platform + human annotation (yes — the Friday0.99
    14. when everyone scores videos).0.98
    15. EVALS FOR AI VIDEO1.00
    16. 08 /200.94
    17. hrdinatat0.55
    18. World's Fair0.98
    19. TRACK 5· JULY 1, 20260.95
    20. Evals1.00
  • 6:11 #18 done18 line(s)

    shot 18·sharpness 1586.8

    1. AlEngineer0.97
    2. 041.00
    3. CLOSE THE LOOP1.00
    4. World'sFair1.00
    5. Catch drift where it's cheap to fix.1.00
    6. PRESENTED BY1.00
    7. IMAGE - VIDEO0.97
    8. LONG-FORM1.00
    9. (2-4 MIN)0.99
    10. Microsoft1.00
    11. Character drifts between keyframes. Fix it here - not1.00
    12. Regenerate one 6-second shot - not the whole film.1.00
    13. three minutes later.1.00
    14. EVALS FOR AI VIDEO0.97
    15. 09 /200.95
    16. World's Fair0.98
    17. TRACK 5· JULY 1, 20260.95
    18. Evals1.00
  • 6:48 #19 done16 line(s)

    shot 19·sharpness 1440.3

    1. AlEngineer0.98
    2. 041.00
    3. CLOSE THE LOOP1.00
    4. World'sFair1.00
    5. Catch drift where it's cheap to fix.1.00
    6. IMAGE - VIDEO0.97
    7. LONG-FORM1.00
    8. (2-4 MIN)1.00
    9. Character drifts between keyframes. Fix it here - not1.00
    10. Regenerate one 6-second shot - not the whole film.1.00
    11. three minutes later.1.00
    12. EVALS FOR AI VIDEO0.99
    13. 09 /200.95
    14. World'sFair1.00
    15. TRACK 5· JULY 1,20260.97
    16. Evals1.00
  • 7:03 #20 done24 line(s)

    shot 20·sharpness 1442.5

    1. AlEngineer0.99
    2. 041.00
    3. CLOSE THE LOOP0.99
    4. World's Fair0.97
    5. Some axes only exist across time.1.00
    6. STORY1.00
    7. PACING1.00
    8. SOUND1.00
    9. coherent1.00
    10. slop1.00
    11. uneven rhythm = bad pacing0.98
    12. the slam lands here0.99
    13. Does the story hold — or melt into1.00
    14. Do the cuts land — or drag and0.98
    15. Does the slam hit on the exact0.99
    16. slop?1.00
    17. rush?1.00
    18. frame?1.00
    19. None of these live in a single frame — judge them across time, at the shot level.0.99
    20. EVALS FOR AI VIDEO0.98
    21. 10 / 200.91
    22. World's Fair0.99
    23. TRACK 5· JULY 1,20260.94
    24. Evals1.00
  • 7:40 #21 done20 line(s)

    shot 21·sharpness 1589.0

    1. AlEngineer0.98
    2. 051.00
    3. MAKE IT FAST1.00
    4. World'sFair1.00
    5. Online means fast. Judge Judy wasn't.0.98
    6. You can't run a committee of models1.00
    7. on every 6-second shot in a0.99
    8. generation loop.1.00
    9. So we distilled the committee into a1.00
    10. A stack of heavy tools0.98
    11. One small model1.00
    12. single reward model.0.98
    13. accurate·slow·moody1.00
    14. one score, fast enough to1.00
    15. gate1.00
    16. EVALS FOR AI VIDEO0.97
    17. 11 / 200.93
    18. World's Fair0.95
    19. TRACK 5· JULY 1, 20260.95
    20. Evals1.00
  • 8:14 #22 done17 line(s)

    shot 22·sharpness 1190.9

    1. AlEngineer0.97
    2. 051.00
    3. MAKE IT FAST1.00
    4. World'sFair1.00
    5. Small model. On purpose.1.00
    6. 1 smal1 VLM0.93
    7. ~3s0.99
    8. Qwen3-VL-8B - vision-language, 8B params0.99
    9. to score a 15-second clip0.99
    10. needs eyes → vision-language0.99
    11. needs speed → it lives at the gate0.98
    12. WeA/B'dabiggermodel:marginally better, significantly slower. Not worth it.0.99
    13. EVALS FOR AI VIDEO1.00
    14. 12 /200.98
    15. World'sFair0.99
    16. TRACK 5· JULY 1,20260.96
    17. Evals1.00
  • 9:05 #23 skipped

    shot 23·duplicate of #22

Transcript

234 cues· 3,257 words· 16,956 chars

  1. 0:13 So, hi, I'm Eorah.
  2. 0:14 I've been with Character for a bit over two years, and we'll talk about AI slop, right?
  3. 0:21 I think that, you know,
  4. 0:24 When we look at video generations as a whole, we have two parallel tracks.
  5. 0:31 One is the video generation, which became insanely good from models like Kling and Seedance and Veo and Sora, we still remember Sora.
  6. 0:43 But the part that got left behind is how we evaluate the quality of the video that was generated.
  7. 0:50 generated, right?
  8. 0:51 So on the one hand, we still kind of squint at it and decide whether or not it's good.
  9. 0:58 But on the other hand, we know that the generation has gotten a lot better.
  10. 1:02 And when we look at X or whatever social you're consuming your content on, there are a lot of guides on how to create amazing videos with this model or that.
  11. 1:16 So the hard part was never how to make video.
  12. 1:19 The hard part was how do we generate
  13. 1:24 good enough video, and how do we judge if the video is good enough?
  14. 1:28 So now we've gone to a world where the generation of video is basically free, right?
  15. 1:34 Free, especially when you compare it to how much studios would charge.
  16. 1:39 But the problem is the most, the grand majority of videos that is generated is not that good, right?
  17. 1:45 We have a lot of hallucinations, like a third limb,
  18. 1:49 opening and closing the door at the same time, hovering, physics, et cetera.
  19. 1:54 So unfortunately, in order to get high-quality content, we need a human to judge.
  20. 2:00 And I don't know when was the last time you've seen how someone is creating these long-form generated video.
  21. 2:09 It's usually a lot of shorter generations and a lot of editing.
  22. 2:13 The problem is because we're using a lot of the tools that we built for the text era, for the image era, for videos, right?
  23. 2:21 We're using things like Clip Score, which is great to judge a single frame.
  24. 2:28 Things like...
  25. 2:33 IPS will help us kind of detect the drift between frames.
  26. 2:36 But we don't have, I mean, but the problem is when you kind of combine all these together, all these tools, they're good at watching the individual frames.
  27. 2:44 They're good at checking this, this, this, this, does this specific frame, does it match the prompt that generated it?
  28. 2:53 It will check consistency between frames, and it will check whether or not it matched the prompt that drove it.
  29. 2:59 But what it won't do, it doesn't tell you if you told a story that you meant to tell.
  30. 3:07 If you think about what is video, video is a storytelling medium.
  31. 3:11 Video is just another form on how we tell a story for any type of story.
  32. 3:17 So one of the things we have to look at
  33. 3:19 Does it tell the actual story?
  34. 3:21 Does the physics make sense?
  35. 3:22 Like, for example, if we want a video of a character walking downstairs, does it actually walk or hover?
  36. 3:30 Does the character say the same character across multiple shots?
  37. 3:35 Does the pacing make sense?
  38. 3:37 For example, people take time going from one place to another.
  39. 3:41 We need to make sure that the pacing makes sense as well.
  40. 3:44 And especially when we add audio, we want to make sure that the audio is kind of synced with the imagery.
  41. 3:49 For example, if someone is slamming a door, we want that sound of the door being slammed to be exactly when the door is actually being slammed.
  42. 3:59 Now, the next iteration we all went to a while ago, we started using LLM as a judge for everything, and we have amazing foundational models that we just throw videos at them.
  43. 4:11 The problem with them is that, A, they're slow, B, they're only as good as your prompts, and multiple people will prompt multiple ways, and the same model may respond in a very, very different way.
  44. 4:22 And sometimes the prompt we use, like, is it consistent?
  45. 4:26 Does this match the prompt?
  46. 4:28 But then the question we really care about, is it good?
  47. 4:31 And the answer varies.
  48. 4:34 So, oops, sorry about that.
  49. 4:36 So our first iteration is like, let's take all these things and build a repeatable
  50. 4:42 benchmark on how we test video that we can rerun over and over and over again.

Chapters

  1. 0:00 Introduction: evaluating AI generated video
  2. 1:19 Why video generation drifts between frames
  3. 3:14 Story and sound: what a clip has to get right
  4. 4:43 LLM as a judge, and catching drift early
  5. 7:01 Story and sound failure modes
  6. 8:28 Small model vs bigger model as judge
  7. 9:20 Don't score, compare: pairwise preference
  8. 10:47 When the judge scores vibe over substance
  9. 11:53 Pairing real footage to train a quality detector
  10. 13:27 Self verification in the generation loop
  11. 15:05 Q&A

Open at this second