read-only demo

Videos lCBf9slCanI

Ending AI Slop — Thais Castello Branco, Taste Labs

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:16:30· en-US· indexed 2026-08-10 19:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 0:46, 1 of 1 keyframes kept
  6. Shot 5, 0:46 to 0:51, 1 of 1 keyframes kept
  7. Shot 6, 0:51 to 1:04, 1 of 1 keyframes kept
  8. Shot 7, 1:04 to 1:32, 1 of 1 keyframes kept
  9. Shot 8, 1:32 to 2:00, 1 of 1 keyframes kept
  10. Shot 9, 2:00 to 2:41, 1 of 1 keyframes kept
  11. Shot 10, 2:41 to 3:04, 1 of 1 keyframes kept
  12. Shot 11, 3:04 to 3:23, 1 of 1 keyframes kept
  13. Shot 12, 3:23 to 3:30, 1 of 1 keyframes kept
  14. Shot 13, 3:30 to 4:00, 1 of 1 keyframes kept
  15. Shot 14, 4:00 to 4:29, 0 of 1 keyframes kept
  16. Shot 15, 4:29 to 4:55, 1 of 1 keyframes kept
  17. Shot 16, 4:55 to 5:20, 0 of 1 keyframes kept
  18. Shot 17, 5:20 to 5:46, 0 of 1 keyframes kept
  19. Shot 18, 5:46 to 6:22, 1 of 1 keyframes kept
  20. Shot 19, 6:22 to 6:59, 0 of 1 keyframes kept
  21. Shot 20, 6:59 to 7:29, 0 of 1 keyframes kept
  22. Shot 21, 7:29 to 7:56, 1 of 1 keyframes kept
  23. Shot 22, 7:56 to 8:23, 0 of 1 keyframes kept
  24. Shot 23, 8:23 to 8:50, 0 of 1 keyframes kept
  25. Shot 24, 8:50 to 9:15, 1 of 1 keyframes kept
  26. Shot 25, 9:15 to 9:46, 1 of 1 keyframes kept
  27. Shot 26, 9:46 to 10:22, 1 of 1 keyframes kept
  28. Shot 27, 10:22 to 10:48, 1 of 1 keyframes kept
  29. Shot 28, 10:48 to 11:13, 1 of 1 keyframes kept
  30. Shot 29, 11:13 to 11:38, 0 of 1 keyframes kept
  31. Shot 30, 11:38 to 12:04, 1 of 1 keyframes kept
  32. Shot 31, 12:04 to 12:30, 0 of 1 keyframes kept
  33. Shot 32, 12:30 to 12:55, 0 of 1 keyframes kept
  34. Shot 33, 12:55 to 13:21, 1 of 1 keyframes kept
  35. Shot 34, 13:21 to 13:47, 0 of 1 keyframes kept
  36. Shot 35, 13:47 to 14:12, 0 of 1 keyframes kept
  37. Shot 36, 14:12 to 14:38, 0 of 1 keyframes kept
  38. Shot 37, 14:38 to 15:04, 0 of 1 keyframes kept
  39. Shot 38, 15:04 to 15:29, 0 of 1 keyframes kept
  40. Shot 39, 15:29 to 16:09, 1 of 1 keyframes kept
  41. Shot 40, 16:09 to 16:13, 1 of 1 keyframes kept
  42. Shot 41, 16:13 to 16:29, 0 of 1 keyframes kept

42 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
181
whisperx 181
chunks
29
from 181 cues
keyframes
26
kept of 42 captured
frames with text
26
323 lines read
chapters
10
from the source metadata
keyframe bytes
4.3 MB
word timings on 181 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:40 0s
stt done 2026-08-09 02:42 19s
chunk done 2026-08-09 02:42 0s
text_embed done 2026-08-10 19:37 1s
keyframe done 2026-08-09 02:42 2m 05s
ocr done 2026-08-09 02:44 7s
frame_embed done 2026-08-10 19:37 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.7

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2743.2

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:20 #3 done2 line(s)

    shot 3·sharpness 289.9

    1. AlEngineer1.00
    2. World's Fair0.99
  • 0:43 #4 done2 line(s)

    shot 4·sharpness 289.5

    1. AlEngineer0.99
    2. World's Fair0.99
  • 0:48 #5 done2 line(s)

    shot 5·sharpness 286.8

    1. AlEngineer1.00
    2. World's Fair1.00
  • 1:02 #6 done20 line(s)

    shot 6·sharpness 1735.3

    1. AlEngineer0.98
    2. World's Fair0.99
    3. Decoding subjective domains1.00
    4. Taste Labs -Ai Engineer0.96
    5. PRESENTED BY0.97
    6. Microsoft1.00
    7. t0.76
    8. da0.79
    9. PARALLEL P0.98
    10. PARALLEL1.00
    11. ARALLEL1.00
    12. PARALLEL P0.95
    13. PARALLEL1.00
    14. R PARALLEL PAR0.99
    15. R PARALLEL0.99
    16. R PARAL0.99
    17. with Thais Castello Branco, CEO0.98
    18. tastelabs.com1.00
    19. d's Fair0.94
    20. Engineering the future of Al0.99
  • 1:07 #7 done17 line(s)

    shot 7·sharpness 1386.2

    1. AlEngineer0.98
    2. World's Fair0.95
    3. Taste Labs1.00
    4. PRESENTED BY1.00
    5. Design me a elegant and interactive website for1.00
    6. my company that wants to invite artist to join0.98
    7. Microsoft1.00
    8. 173M1.00
    9. Bulld!0.75
    10. For training models1.00
    11. ForAgents/Apps1.00
    12. We work with the top frontier labs to give their0.98
    13. We work with app layer, coding agents and0.99
    14. models taste, vision and design capabilities1.00
    15. creative technology companies.1.00
    16. 'sFair0.98
    17. Engineering the future of Al1.00
  • 1:46 #8 done18 line(s)

    shot 8·sharpness 1315.3

    1. AlEngineer0.99
    2. World's Fair0.96
    3. Taste Labs1.00
    4. PRESENTED BY1.00
    5. Design me a elegant and interactive website for0.99
    6. my company that wants to invite artist to join0.98
    7. Microsoft1.00
    8. 173M1.00
    9. Bulld!0.76
    10. For training models1.00
    11. ForAgents/Apps1.00
    12. We work with the top frontier labs to give their0.98
    13. We work with app layer, coding agents and1.00
    14. models taste, vision and design capabilities0.99
    15. creative technology companies.1.00
    16. air1.00
    17. TRACK 9· JUNE 30, 20260.97
    18. Data Quality1.00
  • 2:25 #9 done7 line(s)

    shot 9·sharpness 1140.6

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Why are models so much worse1.00
    4. at subjective domains?1.00
    5. air1.00
    6. TRACK 9• JUNE 30, 20260.96
    7. Data Quality1.00
  • 2:57 #10 done12 line(s)

    shot 10·sharpness 1199.8

    1. AlEngineer0.98
    2. World'sFair1.00
    3. We tend to treat "good at code" as a fact0.98
    4. It decomposes.0.98
    5. about models. It's really a fact about code.0.98
    6. Code has three properties almost nothing0.99
    7. else has at once:0.95
    8. Itverifies.1.00
    9. It executes.0.99
    10. air1.00
    11. TRACK 9• JUNE 30, 20260.94
    12. Data Quality1.00
  • 3:17 #11 done7 line(s)

    shot 11·sharpness 749.9

    1. AlEngineer0.99
    2. World'sFair1.00
    3. Capability follows measurability.0.99
    4. The mean is not optimal.0.98
    5. air1.00
    6. TRACK 9• JUNE 30, 20260.96
    7. Data Quality1.00
  • 3:27 #12 done10 line(s)

    shot 12·sharpness 1454.3

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Jason Wei put a name to this with what he0.98
    4. calls Verifier's Law:1.00
    5. The ease of training Al to solve a0.99
    6. task is proportional to how0.99
    7. verifiable that task is.0.98
    8. air1.00
    9. TRACK 9• JUNE 30, 20260.93
    10. Data Quality1.00
  • 3:34 #13 done25 line(s)

    shot 13·sharpness 1154.5

    1. AlEngineer0.99
    2. World'sFair0.96
    3. Transform from fuzzy into1.00
    4. more verifiable.1.00
    5. < RELATIBILITY >0.99
    6. 56%1.00
    7. < ORIGINALITY >0.98
    8. 098%0.90
    9. 251.00
    10. 501.00
    11. 751.00
    12. 60.98
    13. COMPOSITION >0.98
    14. 41%1.00
    15. vr1.00
    16. WOW FACTOR >0.99
    17. 22%1.00
    18. 601.00
    19. COMMON1.00
    20. 12%1.00
    21. < PRESENCE >0.99
    22. 56%1.00
    23. air1.00
    24. TRACK 9· JUNE 30, 20260.97
    25. Data Quality1.00
  • 4:06 #14 skipped

    shot 14·duplicate of #13

  • 4:44 #15 done15 line(s)

    shot 15·sharpness 1037.7

    1. AlEngineer0.99
    2. World's Fair0.96
    3. Example:Brand adherence1.00
    4. @reducto.ai0.94
    5. Document work1.00
    6. starts here.0.98
    7. Colors1.00
    8. → → Decompose → →0.95
    9. Typography1.00
    10. Aa1.00
    11. Aa1.00
    12. Buttons & icons0.96
    13. air1.00
    14. TRACK 9· JUNE 30, 20260.96
    15. Data Quality1.00
  • 5:17 #16 skipped

    shot 16·duplicate of #15

  • 5:23 #17 skipped

    shot 17·duplicate of #15

  • 6:11 #18 done24 line(s)

    shot 18·sharpness 1164.3

    1. AlEngineer0.99
    2. World'sFair1.00
    3. Example: Environment0.99
    4. Task1.00
    5. Agent1.00
    6. Output1.00
    7. Grading1.00
    8. Trainer1.00
    9. Context + Prompt0.99
    10. Against ground truth1.00
    11. PRESENTED BY1.00
    12. reducto.ai0.98
    13. Microsoft1.00
    14. Document work1.00
    15. Preditable pricing0.94
    16. starts here.1.00
    17. for every volume0.99
    18. 4991.00
    19. Custom1.00
    20. Aa1.00
    21. Aa0.86
    22. air1.00
    23. TRACK 9• JUNE 30, 20260.97
    24. Data Quality1.00
  • 6:30 #19 skipped

    shot 19·duplicate of #18

  • 7:05 #20 skipped

    shot 20·duplicate of #11

  • 7:40 #21 done8 line(s)

    shot 21·sharpness 1307.8

    1. AlEngineer0.99
    2. World's Fair0.97
    3. Why does it feel like slop?0.99
    4. is the0.89
    5. Sof Tct ffor a0.79
    6. air1.00
    7. TRACK 9 • JUNE 30, 20260.96
    8. Data Quality1.00
  • 8:15 #22 skipped

    shot 22·duplicate of #21

  • 8:47 #23 skipped

    shot 23·duplicate of #21

Transcript

181 cues· 3,240 words· 17,807 chars

  1. 0:13 Hello, everyone.
  2. 0:15 It's great to meet you all.
  3. 0:17 I'm Thais.
  4. 0:17 I'm the founder of Taste Labs.
  5. 0:20 For those of you who don't know us, we came out of stealth a few weeks ago, and our whole mission is basically how do we end AI slop?
  6. 0:27 And we believe that to really solve this problem, we have to first
  7. 0:32 decompose and understand subjective domains.
  8. 0:34 I think, as probably all of you know, AI has gotten quite good at things like coding and math.
  9. 0:39 But it's still super behind on things like design, creative writing, personality, emotional intelligence.
  10. 0:44 And to understand these domains, I think we have to take a little bit of a different approach than we do with objective ones.
  11. 0:50 So our idea is like, how do we become this data and infrastructure layer to really understand these problems and to become the solution for them across the stack?
  12. 0:59 So from the foundation model layer all the way to how do we build solutions for agents as well?
  13. 1:04 So we work primarily in two ways.
  14. 1:07 We work with the top frontier labs on how do we evaluate, benchmark their models, understand where they're breaking, understand how we can fix them, and how do we determine also which problem is better fixed through each method?
  15. 1:18 So what things should be turned into RL environments?
  16. 1:21 Which things should be turned into post-training data problems?
  17. 1:24 But then we also go and work with a lot of agent and application layer companies on what are things that we actually don't believe should be solved at the foundation model layer.
  18. 1:32 And that might be better solved through methods like context or understanding user intent, right?
  19. 1:36 We're basically betting on a world where suddenly you're gonna have billions of people creating that are not necessarily experts.
  20. 1:43 So this understanding of user intent and context is equally as important as how do we get these models to improve.
  21. 1:49 For today, I'm going to focus on the model training part.
  22. 1:51 For those that are here tomorrow, I'll also be giving a chat on the design track where I'll cover more on what we're doing on the agent side of the house.
  23. 1:58 But there we go.
  24. 2:01 Okay.
  25. 2:01 So most of the world is subjective, as I was mentioning.
  26. 2:04 A lot of the world is subjective, right?
  27. 2:05 If we talk about these domains of writing, even workflows within companies, right, of sales, marketing, a lot of the times there's this multitude of answers.
  28. 2:14 There's not one clear right answer, and it's very hard to define what great even means.
  29. 2:18 I think at the end of the day, those are why these domains are so difficult.
  30. 2:23 And we oftentimes, I would say, forget to mention why.
  31. 2:29 We treat, for example, the fact that code is verifiable and measurable as something that is a property about models.
  32. 2:37 And models are great at coding
  33. 2:40 because we've made them great at coding.
  34. 2:42 But realistically, it's actually a fact about code.
  35. 2:44 Code is something that decomposes, it verifies, it executes, and so it makes it a lot easier for us to be able to train on these domains.
  36. 2:51 For something like design or writing, how do you decompose it?
  37. 2:54 How do you verify it?
  38. 2:55 How do you judge if it's actually good?
  39. 2:57 So that's why they become so difficult.
  40. 2:59 So there's two characteristics that I want to touch on today on why fundamentally subjective domains are harder.
  41. 3:05 One is that capability follows measurability.
  42. 3:08 So if we can solve the measurability problem, or at least part of it, then we can solve a big portion of these domains.
  43. 3:14 The second, which I'll touch on later, is basically this collapse to the mean and why the mean is not necessarily optimal in subjective domains.
  44. 3:25 OK.
  45. 3:32 So to really start solving this problem, we have to turn something that feels fuzzy, like if I ask you what is great design, into something that is more verifiable.
  46. 3:40 So there's a few questions here, right?
  47. 3:42 Because if I ask you this of what is great design, you could ask yourself, OK,
  48. 3:48 Do you mean great for which type of person, for which type of taste, for which situation?
  49. 3:52 The same slide could be amazing, for example, if you are a startup and completely inappropriate if you are a finance firm.
  50. 3:59 So it's contextual, first of all.

Chapters

  1. 0:00 What ending AI slop means
  2. 1:06 Working with frontier labs on taste
  3. 2:33 The verifiable to preference spectrum
  4. 3:38 Decomposing design to grade it
  5. 5:30 Verifying whether something is on brand
  6. 7:40 Collapse to the mean
  7. 9:18 Shifting preference toward verifiable
  8. 11:25 Building a preference vector
  9. 13:05 Human QA tied to specific choices
  10. 15:00 Training taste in subjective domains

Open at this second