Videos lCBf9slCanI
Ending AI Slop — Thais Castello Branco, Taste Labs
Scene timeline
42 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 181
- whisperx 181
- chunks
- 29
- from 181 cues
- keyframes
- 26
- kept of 42 captured
- frames with text
- 26
- 323 lines read
- chapters
- 10
- from the source metadata
- keyframe bytes
- 4.3 MB
- word timings on 181 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:40 | 0s |
stt |
done | — | 2026-08-09 02:42 | 19s |
chunk |
done | — | 2026-08-09 02:42 | 0s |
text_embed |
done | — | 2026-08-10 19:37 | 1s |
keyframe |
done | — | 2026-08-09 02:42 | 2m 05s |
ocr |
done | — | 2026-08-09 02:44 | 7s |
frame_embed |
done | — | 2026-08-10 19:37 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer1.00
- World's Fair0.99
-
- AlEngineer0.99
- World's Fair0.99
-
- AlEngineer1.00
- World's Fair1.00
-
- AlEngineer0.98
- World's Fair0.99
- Decoding subjective domains1.00
- Taste Labs -Ai Engineer0.96
- PRESENTED BY0.97
- Microsoft1.00
- t0.76
- da0.79
- PARALLEL P0.98
- PARALLEL1.00
- ARALLEL1.00
- PARALLEL P0.95
- PARALLEL1.00
- R PARALLEL PAR0.99
- R PARALLEL0.99
- R PARAL0.99
- with Thais Castello Branco, CEO0.98
- tastelabs.com1.00
- d's Fair0.94
- Engineering the future of Al0.99
-
- AlEngineer0.98
- World's Fair0.95
- Taste Labs1.00
- PRESENTED BY1.00
- Design me a elegant and interactive website for1.00
- my company that wants to invite artist to join0.98
- Microsoft1.00
- 173M1.00
- Bulld!0.75
- For training models1.00
- ForAgents/Apps1.00
- We work with the top frontier labs to give their0.98
- We work with app layer, coding agents and0.99
- models taste, vision and design capabilities1.00
- creative technology companies.1.00
- 'sFair0.98
- Engineering the future of Al1.00
-
- AlEngineer0.99
- World's Fair0.96
- Taste Labs1.00
- PRESENTED BY1.00
- Design me a elegant and interactive website for0.99
- my company that wants to invite artist to join0.98
- Microsoft1.00
- 173M1.00
- Bulld!0.76
- For training models1.00
- ForAgents/Apps1.00
- We work with the top frontier labs to give their0.98
- We work with app layer, coding agents and1.00
- models taste, vision and design capabilities0.99
- creative technology companies.1.00
- air1.00
- TRACK 9· JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- Why are models so much worse1.00
- at subjective domains?1.00
- air1.00
- TRACK 9• JUNE 30, 20260.96
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- We tend to treat "good at code" as a fact0.98
- It decomposes.0.98
- about models. It's really a fact about code.0.98
- Code has three properties almost nothing0.99
- else has at once:0.95
- Itverifies.1.00
- It executes.0.99
- air1.00
- TRACK 9• JUNE 30, 20260.94
- Data Quality1.00
-
- AlEngineer0.99
- World'sFair1.00
- Capability follows measurability.0.99
- The mean is not optimal.0.98
- air1.00
- TRACK 9• JUNE 30, 20260.96
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- Jason Wei put a name to this with what he0.98
- calls Verifier's Law:1.00
- The ease of training Al to solve a0.99
- task is proportional to how0.99
- verifiable that task is.0.98
- air1.00
- TRACK 9• JUNE 30, 20260.93
- Data Quality1.00
-
- AlEngineer0.99
- World'sFair0.96
- Transform from fuzzy into1.00
- more verifiable.1.00
- < RELATIBILITY >0.99
- 56%1.00
- < ORIGINALITY >0.98
- 098%0.90
- 251.00
- 501.00
- 751.00
- 60.98
- COMPOSITION >0.98
- 41%1.00
- vr1.00
- WOW FACTOR >0.99
- 22%1.00
- 601.00
- COMMON1.00
- 12%1.00
- < PRESENCE >0.99
- 56%1.00
- air1.00
- TRACK 9· JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.99
- World's Fair0.96
- Example:Brand adherence1.00
- @reducto.ai0.94
- Document work1.00
- starts here.0.98
- Colors1.00
- → → Decompose → →0.95
- Typography1.00
- Aa1.00
- Aa1.00
- Buttons & icons0.96
- air1.00
- TRACK 9· JUNE 30, 20260.96
- Data Quality1.00
-
- AlEngineer0.99
- World'sFair1.00
- Example: Environment0.99
- Task1.00
- Agent1.00
- Output1.00
- Grading1.00
- Trainer1.00
- Context + Prompt0.99
- Against ground truth1.00
- PRESENTED BY1.00
- reducto.ai0.98
- Microsoft1.00
- Document work1.00
- Preditable pricing0.94
- starts here.1.00
- for every volume0.99
- 4991.00
- Custom1.00
- Aa1.00
- Aa0.86
- air1.00
- TRACK 9• JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.99
- World's Fair0.97
- Why does it feel like slop?0.99
- is the0.89
- Sof Tct ffor a0.79
- air1.00
- TRACK 9 • JUNE 30, 20260.96
- Data Quality1.00
Transcript
181 cues· 3,240 words· 17,807 chars
- 0:13 Hello, everyone.
- 0:15 It's great to meet you all.
- 0:17 I'm Thais.
- 0:17 I'm the founder of Taste Labs.
- 0:20 For those of you who don't know us, we came out of stealth a few weeks ago, and our whole mission is basically how do we end AI slop?
- 0:27 And we believe that to really solve this problem, we have to first
- 0:32 decompose and understand subjective domains.
- 0:34 I think, as probably all of you know, AI has gotten quite good at things like coding and math.
- 0:39 But it's still super behind on things like design, creative writing, personality, emotional intelligence.
- 0:44 And to understand these domains, I think we have to take a little bit of a different approach than we do with objective ones.
- 0:50 So our idea is like, how do we become this data and infrastructure layer to really understand these problems and to become the solution for them across the stack?
- 0:59 So from the foundation model layer all the way to how do we build solutions for agents as well?
- 1:04 So we work primarily in two ways.
- 1:07 We work with the top frontier labs on how do we evaluate, benchmark their models, understand where they're breaking, understand how we can fix them, and how do we determine also which problem is better fixed through each method?
- 1:18 So what things should be turned into RL environments?
- 1:21 Which things should be turned into post-training data problems?
- 1:24 But then we also go and work with a lot of agent and application layer companies on what are things that we actually don't believe should be solved at the foundation model layer.
- 1:32 And that might be better solved through methods like context or understanding user intent, right?
- 1:36 We're basically betting on a world where suddenly you're gonna have billions of people creating that are not necessarily experts.
- 1:43 So this understanding of user intent and context is equally as important as how do we get these models to improve.
- 1:49 For today, I'm going to focus on the model training part.
- 1:51 For those that are here tomorrow, I'll also be giving a chat on the design track where I'll cover more on what we're doing on the agent side of the house.
- 1:58 But there we go.
- 2:01 Okay.
- 2:01 So most of the world is subjective, as I was mentioning.
- 2:04 A lot of the world is subjective, right?
- 2:05 If we talk about these domains of writing, even workflows within companies, right, of sales, marketing, a lot of the times there's this multitude of answers.
- 2:14 There's not one clear right answer, and it's very hard to define what great even means.
- 2:18 I think at the end of the day, those are why these domains are so difficult.
- 2:23 And we oftentimes, I would say, forget to mention why.
- 2:29 We treat, for example, the fact that code is verifiable and measurable as something that is a property about models.
- 2:37 And models are great at coding
- 2:40 because we've made them great at coding.
- 2:42 But realistically, it's actually a fact about code.
- 2:44 Code is something that decomposes, it verifies, it executes, and so it makes it a lot easier for us to be able to train on these domains.
- 2:51 For something like design or writing, how do you decompose it?
- 2:54 How do you verify it?
- 2:55 How do you judge if it's actually good?
- 2:57 So that's why they become so difficult.
- 2:59 So there's two characteristics that I want to touch on today on why fundamentally subjective domains are harder.
- 3:05 One is that capability follows measurability.
- 3:08 So if we can solve the measurability problem, or at least part of it, then we can solve a big portion of these domains.
- 3:14 The second, which I'll touch on later, is basically this collapse to the mean and why the mean is not necessarily optimal in subjective domains.
- 3:25 OK.
- 3:32 So to really start solving this problem, we have to turn something that feels fuzzy, like if I ask you what is great design, into something that is more verifiable.
- 3:40 So there's a few questions here, right?
- 3:42 Because if I ask you this of what is great design, you could ask yourself, OK,
- 3:48 Do you mean great for which type of person, for which type of taste, for which situation?
- 3:52 The same slide could be amazing, for example, if you are a startup and completely inappropriate if you are a finance firm.
- 3:59 So it's contextual, first of all.
loading
Chapters
- 0:00 What ending AI slop means
- 1:06 Working with frontier labs on taste
- 2:33 The verifiable to preference spectrum
- 3:38 Decomposing design to grade it
- 5:30 Verifying whether something is on brand
- 7:40 Collapse to the mean
- 9:18 Shifting preference toward verifiable
- 11:25 Building a preference vector
- 13:05 Human QA tied to specific choices
- 15:00 Training taste in subjective domains