read-only demo

Videos ZyIoTOAbRfs

State of Data — Sean Cai, Independent / State of Data

index_state ready data_status ok

AI Engineer· published 2026-07-26· 0:18:22· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:49, 1 of 1 keyframes kept
  5. Shot 4, 0:49 to 1:07, 1 of 1 keyframes kept
  6. Shot 5, 1:07 to 1:44, 1 of 1 keyframes kept
  7. Shot 6, 1:44 to 2:20, 0 of 1 keyframes kept
  8. Shot 7, 2:20 to 2:46, 1 of 1 keyframes kept
  9. Shot 8, 2:46 to 3:13, 0 of 1 keyframes kept
  10. Shot 9, 3:13 to 3:43, 1 of 1 keyframes kept
  11. Shot 10, 3:43 to 4:14, 0 of 1 keyframes kept
  12. Shot 11, 4:14 to 4:46, 1 of 1 keyframes kept
  13. Shot 12, 4:46 to 5:18, 0 of 1 keyframes kept
  14. Shot 13, 5:18 to 5:51, 0 of 1 keyframes kept
  15. Shot 14, 5:51 to 6:17, 1 of 1 keyframes kept
  16. Shot 15, 6:17 to 6:43, 0 of 1 keyframes kept
  17. Shot 16, 6:43 to 7:10, 0 of 1 keyframes kept
  18. Shot 17, 7:10 to 7:36, 1 of 1 keyframes kept
  19. Shot 18, 7:36 to 8:03, 0 of 1 keyframes kept
  20. Shot 19, 8:03 to 8:32, 1 of 1 keyframes kept
  21. Shot 20, 8:32 to 9:02, 1 of 1 keyframes kept
  22. Shot 21, 9:02 to 9:32, 0 of 1 keyframes kept
  23. Shot 22, 9:32 to 10:02, 0 of 1 keyframes kept
  24. Shot 23, 10:02 to 10:33, 1 of 1 keyframes kept
  25. Shot 24, 10:33 to 11:04, 1 of 1 keyframes kept
  26. Shot 25, 11:04 to 11:36, 0 of 1 keyframes kept
  27. Shot 26, 11:36 to 11:59, 1 of 1 keyframes kept
  28. Shot 27, 11:59 to 12:26, 1 of 1 keyframes kept
  29. Shot 28, 12:26 to 12:54, 0 of 1 keyframes kept
  30. Shot 29, 12:54 to 13:21, 0 of 1 keyframes kept
  31. Shot 30, 13:21 to 13:49, 1 of 1 keyframes kept
  32. Shot 31, 13:49 to 14:16, 0 of 1 keyframes kept
  33. Shot 32, 14:16 to 14:42, 1 of 1 keyframes kept
  34. Shot 33, 14:42 to 15:08, 1 of 1 keyframes kept
  35. Shot 34, 15:08 to 15:34, 0 of 1 keyframes kept
  36. Shot 35, 15:34 to 16:00, 1 of 1 keyframes kept
  37. Shot 36, 16:00 to 16:31, 1 of 1 keyframes kept
  38. Shot 37, 16:31 to 17:03, 0 of 1 keyframes kept
  39. Shot 38, 17:03 to 17:33, 1 of 1 keyframes kept
  40. Shot 39, 17:33 to 17:54, 1 of 1 keyframes kept
  41. Shot 40, 17:54 to 18:05, 1 of 1 keyframes kept
  42. Shot 41, 18:05 to 18:21, 0 of 1 keyframes kept

42 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
196
whisperx 196
chunks
31
from 196 cues
keyframes
25
kept of 42 captured
frames with text
25
581 lines read
chapters
13
from the source metadata
keyframe bytes
4.8 MB
word timings on 196 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:29 0s
stt done 2026-08-09 06:27 20s
chunk done 2026-08-09 06:27 0s
text_embed done 2026-08-10 19:39 0s
keyframe done 2026-08-09 06:27 2m 15s
ocr done 2026-08-09 06:29 14s
frame_embed done 2026-08-10 19:39 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.5

    1. AIEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.0

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2742.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.92
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:41 #3 done2 line(s)

    shot 3·sharpness 307.1

    1. AlEngineer1.00
    2. World's Fair0.98
  • 1:02 #4 done18 line(s)

    shot 4·sharpness 1482.7

    1. AlEngineer0.99
    2. // INTRO0.99
    3. AI ENGINEER0.95
    4. WORLD'S FAIR 20261.00
    5. World'sFair1.00
    6. POST-TRAINING & DATA QUALITY TRACK1.00
    7. STATE OF DATA1.00
    8. PRESENTED BY1.00
    9. Microsoft1.00
    10. A field guide to Al data markets1.00
    11. Every data company is quietly becoming an enterprise Al company.0.99
    12. SEAN CAI1.00
    13. STANDARD DATA1.00
    14. 20 MIN0.99
    15. STATE OF DATA0.98
    16. AI ENGINEERING1.00
    17. Engineering the future of Al0.98
    18. World's Fair0.97
  • 1:32 #5 done25 line(s)

    shot 5·sharpness 2275.3

    1. AlEngineer0.99
    2. // THE COMMODITY1.00
    3. World's Fair0.98
    4. Data is the new-age commodity0.98
    5. TAM = all of human labor. The supply chain is unbundling.0.99
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. RAW DATA1.00
    9. TRACES1.00
    10. ENVS1.00
    11. RUBRICS1.00
    12. EVALS1.00
    13. records of real work0.98
    14. reasoning, decisions1.00
    15. tasks + reward1.00
    16. the finished good1.00
    17. A modern RL dataset isn't labels. It's environments, verifiers, judge-model inference, and the research talent to run them. Selling data is0.99
    18. selling model improvement.0.96
    19. Fragmentation is the equilibrium: two years ago one vertically integrated giant did all of it; now specialists eat every step, because0.99
    20. quality never scales linearly with quantity and labs keep the pool fragmented on purpose.0.99
    21. STATE OF DATA1.00
    22. AI ENGINEERING1.00
    23. TRACK 9· JULY 1, 20260.93
    24. World's Fair0.99
    25. Posttraining&Midtraining1.00
  • 2:02 #6 skipped

    shot 6·duplicate of #5

  • 2:33 #7 done22 line(s)

    shot 7·sharpness 2023.1

    1. AlEngineer0.98
    2. // THE SERIES0.99
    3. World'sFair1.00
    4. Six months of State of Data0.99
    5. One argument, compounding.1.00
    6. JAN1.00
    7. Industrialization & unbundling Antikythera mechanisms; data becomes semi-liquid0.99
    8. FEB1.00
    9. Type 1 vs Type 2 data why the GPQA playbook breaks on long-horizon work0.99
    10. MAR1.00
    11. Verifying the unverifiable subjective self-serve post-training; a real clinical pipeline0.99
    12. APR1.00
    13. The real TAMN-1 datasets; small-model systems get mandated0.99
    14. MAY1.00
    15. Benchmark psychosisthe harness is the product, not the model1.00
    16. JUN1.00
    17. Verticalization & sovereign models the infra to own your own intelligence0.99
    18. STATE OF DATA1.00
    19. AI ENGINEERING1.00
    20. TRACK 9 · JULY 1, 20260.93
    21. World's Fair0.97
    22. Posttraining & Midtraining1.00
  • 3:07 #8 skipped

    shot 8·duplicate of #7

  • 3:37 #9 done54 line(s)

    shot 9·sharpness 1908.5

    1. AlEngineer0.98
    2. // FIGURE0.99
    3. World's Fair0.98
    4. Data is the underfunded leg1.00
    5. STATE OF DATA0.96
    6. APRIL 20261.00
    7. Data Vendors Today will pop the Al Bubble if Left Unchecked0.98
    8. 20241.00
    9. 20251.00
    10. 20260.99
    11. Compute1.00
    12. 250B1.00
    13. 450B1.00
    14. 1T?0.99
    15. Data"0.92
    16. Improvement1.00
    17. Velocity0.96
    18. Model1.00
    19. 40B0.99
    20. 50B1.00
    21. ?0.97
    22. Talent1.00
    23. 0.16x0.99
    24. 0.11x1.00
    25. ?x1.00
    26. L = k(C + D + T)0.92
    27. (3(CDT)1/3)0.93
    28. C+ D+T0.94
    29. β=31.00
    30. R = H·m·q·X0.95
    31. X = SL0.95
    32. L = model improvement velocity T = talent investment0.98
    33. C = compute investment0.99
    34. D = data investment0.96
    35. k = scaling constant0.97
    36. beta = imbalance penalty (3)0.97
    37. R = AI Revenue0.90
    38. S = capex spend0.99
    39. X = model improvement0.99
    40. H = human labor TAM0.99
    41. m € [0, 1] =0.88
    42. q ∈ [0, 1] =0.97
    43. regulation, execution)1.00
    44. market maturity0.98
    45. readiness factor0.97
    46. (competition,0.98
    47. value capture1.00
    48. Model improvement velocity = f(compute, data, talent). Compute races toward a trillion; data spend stays flat. Data is the underfunded leg.1.00
    49. STATE OF DATA0.98
    50. AI ENGINEERING1.00
    51. 04 / 150.95
    52. TRACK 9· JULY 1,20260.96
    53. World's Fair0.99
    54. Posttraining &Midtraining0.98
  • 4:04 #10 skipped

    shot 10·duplicate of #9

  • 4:36 #11 done24 line(s)

    shot 11·sharpness 1860.1

    1. AlEngineer0.98
    2. // FOUNDATION0.99
    3. World'sFair0.97
    4. What “data” actually means0.97
    5. Process-based (how the work got done), not state-based (what got stored).0.99
    6. TYPE 11.00
    7. TYPE1.00
    8. Pure capture of real workflows1.00
    9. Contrived data1.00
    10. GitHub commits, session replays1.00
    11. Hired experts manufacturing examples0.99
    12. Minimal reward shaping by non-experts0.99
    13. Scales badly; QA debt compounds1.00
    14. Realism inherited from the work itself1.00
    15. Reward shaping by non-domain experts1.00
    16. Gets models from 20 → 800.99
    17. Gets models from 0 → 200.97
    18. Data is the most depreciable asset there is. It decays as the frontier moves, so the only durable supply is a live business you partner0.99
    19. with, not a dead startup's codebase.0.98
    20. STATE OF DATA1.00
    21. AI ENGINEERING1.00
    22. TRACK 9· JULY 1, 20260.92
    23. World's Fair0.93
    24. Posttraining & Midtraining1.00
  • 5:08 #12 skipped

    shot 12·duplicate of #11

  • 5:34 #13 skipped

    shot 13·duplicate of #11

  • 6:06 #14 done23 line(s)

    shot 14·sharpness 1792.3

    1. AlEngineer0.99
    2. // VERIFICATION1.00
    3. World's Fair0.96
    4. The three axes of verification1.00
    5. Verifier's Law: trainability ∝ verifiability.0.97
    6. Coding1.00
    7. Finance1.00
    8. Taste1.00
    9. PRESENTED BY1.00
    10. Microsoft1.00
    11. ASYMMETRY1.00
    12. can you break it into checkable steps?1.00
    13. VERACITY1.00
    14. is there consensus on "correct"?0.99
    15. PROLIFERATION1.00
    16. how often does the world hand you proof?1.00
    17. Coding matured first because GitHub solved all three at once. Bio, cyber, finance, health, legal sit low, and the proof is locked inside0.99
    18. enterprises.1.00
    19. STATE OF DATA1.00
    20. AI ENGINEERING1.00
    21. TRACK 9· JULY 1, 20260.93
    22. World's Fair0.99
    23. Posttraining & Midtraining1.00
  • 6:35 #15 skipped

    shot 15·duplicate of #14

  • 6:46 #16 skipped

    shot 16·duplicate of #5

  • 7:23 #17 done22 line(s)

    shot 17·sharpness 1643.4

    1. AlEngineer0.97
    2. // VERIFICATION1.00
    3. World's Fair0.96
    4. The three axes of verification1.00
    5. Verifier's Law: trainability verifiability.0.98
    6. Coding1.00
    7. Finance1.00
    8. Taste1.00
    9. ASYMMETRY1.00
    10. HIGH1.00
    11. can you break it into checkable steps?0.99
    12. VERACITY1.00
    13. is there consensus on "correct"?0.99
    14. PROLIFERATION1.00
    15. how often does the world hand you proof?1.00
    16. Coding matured first because GitHub solved all three at once. Bio, cyber, finance, health, legal sit low, and the proof is locked inside0.99
    17. enterprises.1.00
    18. STATE OF DATA1.00
    19. AI ENGINEERING1.00
    20. TRACK 9· JULY 1,20260.95
    21. World's Fair0.99
    22. Posttraining & Midtraining1.00
  • 7:39 #18 skipped

    shot 18·duplicate of #17

  • 8:09 #19 done24 line(s)

    shot 19·sharpness 1719.9

    1. AlEngineer0.98
    2. // BENCHMARKS0.99
    3. World'sFair1.00
    4. You're hillclimbing imaginary mountains1.00
    5. How contrived "realistic” data manufactures fake benchmarks.0.98
    6. Hire experts1.00
    7. ChatGPT-gen tasks0.97
    8. Cherry-pick fails1.00
    9. Sell as "north-star"0.99
    10. then sell the data to hillclimb the same benchmark.1.00
    11. GOODHART'S LAW1.00
    12. THE HARNESS IS THE PRODUCT0.99
    13. measures nothing real, and CapEx gets pointed at imaginary1.00
    14. When a measure becomes a target, it stops measuring. Set by0.99
    15. people who aren't true domain experts, the benchmark0.99
    16. Even clean benchmarks measure the scaffold. Format swings of0.98
    17. 76 pts; a GLM 5.1 stop-token shim worth 0.34 reward, bigger0.99
    18. than most published frontier deltas.1.00
    19. mountains.1.00
    20. STATE OF DATA1.00
    21. AI ENGINEERING1.00
    22. TRACK 9· JULY 1, 20260.96
    23. World's Fair0.96
    24. Posttraining & Midtraining1.00
  • 8:39 #20 done24 line(s)

    shot 20·sharpness 1720.5

    1. AlEngineer0.98
    2. // BENCHMARKS0.99
    3. World'sFair1.00
    4. You're hillclimbing imaginary mountains1.00
    5. How contrived "realistic” data manufactures fake benchmarks.0.98
    6. Hire experts1.00
    7. ChatGPT-gen tasks0.97
    8. Cherry-pick fails1.00
    9. Sell as "north-star"0.99
    10. then sell the data to hillclimb the same benchmark.0.99
    11. GOODHART'S LAW1.00
    12. THE HARNESS IS THE PRODUCT0.99
    13. measures nothing real, and CapEx gets pointed at imaginary0.99
    14. When a measure becomes a target, it stops measuring. Set by0.99
    15. people who aren't true domain experts, the benchmark0.99
    16. Even clean benchmarks measure the scaffold. Format swings of0.98
    17. 76 pts; a GLM 5.1 stop-token shim worth 0.34 reward, bigger0.99
    18. than most published frontier deltas.1.00
    19. mountains.1.00
    20. STATE OF DATA1.00
    21. AI ENGINEERING1.00
    22. TRACK 9· JULY 1, 20260.95
    23. World'sFair1.00
    24. Posttraining & Midtraining1.00
  • 9:11 #21 skipped

    shot 21·duplicate of #19

  • 9:38 #22 skipped

    shot 22·duplicate of #19

  • 10:08 #23 done25 line(s)

    shot 23·sharpness 1916.1

    1. AlEngineer0.98
    2. // PROOF0.94
    3. World'sFair1.00
    4. The leaderboard is one noisy sample1.00
    5. Three real finance tasks, every frontier model, re-graded criterion by criterion.0.99
    6. Frontier models can go backwards1.00
    7. REGRESSION1.00
    8. Opus 4.7 → 4.8 got worse on analyst rubrics: over-reflection made it average beginning ARR and conflate logo0.98
    9. with revenue retention, in 35-40% of rollouts. Errors 4.7 never made.1.00
    10. GPT 5.5 and Opus 4.8 tie, then split0.99
    11. OPPOSITE FAILURES1.00
    12. GPT nails the arithmetic and loses the methodology; Opus nails the methodology and loses the arithmetic.1.00
    13. Trained for different deliverables.1.00
    14. A one-line shim beats a model upgrade1.00
    15. HARNESS > MODEL1.00
    16. GLM 5.1 looks unusable until a stop-token fix lifts it 0.34 reward and ties the Anthropic flagship, larger than the1.00
    17. gap between adjacent leaderboard models.1.00
    18. A single-scaffold benchmark number is one sample from a distribution nobody measured.1.00
    19. STATE OF DATA1.00
    20. AI ENGINEERING1.00
    21. 08 /150.93
    22. TRACK 9· JULY 1, 20260.95
    23. AlEngineer1.00
    24. World's Fair0.98
    25. Posttraining & Midtraining1.00

Transcript

196 cues· 3,134 words· 18,006 chars

  1. 0:12 Good to give this talk.
  2. 0:14 And I just want to say right off the bat, this is probably going to be a little different from what you've seen so far at AI Conference.
  3. 0:20 I'm here to not deliver an agenda on any company's behalf, but just to expose a lot of alpha in data markets, but also just tell you what's really going on behind the scenes in a very murky landscape where nobody seems to know how Mercore, Handshake, and a lot of these folks actually produce data.
  4. 0:39 You know, quick reframe before we start.
  5. 0:41 When people hear data markets, they picture scale around 2019.
  6. 0:44 They picture these rooms of annotators and manila labeling images.
  7. 0:48 And that's real, and it's maybe $10 to $15 billion a year per lab, but it's kind of the least interesting part.
  8. 0:55 The models work now.
  9. 0:56 What's scarce and badly priced is the sort of data that takes them from generalist competence into real expertise.
  10. 1:02 We knew this since 2024 when scale got acquired, and we were spamming a lot of GPQA data sets.
  11. 1:09 So, look, let me start with this framing.
  12. 1:13 Data is to the white collar revolution what coal and iron basically is to the Victorian age and the information age right now.
  13. 1:20 I have this piece called the TAM is in a vertical, it's all of labor.
  14. 1:24 And so the supply chain is basically just doing what every sort of industrializing supply chain does.
  15. 1:29 It unbundles.
  16. 1:30 Two years ago, one vertically integrated giant like a Mercor or Surge or Scale AI, if you guys are unfamiliar, they're massive data companies, they had to do all of it.
  17. 1:40 But that was because that's the only way that unit economics worked in an immature market today.
  18. 1:46 That's increasingly not the case.
  19. 1:48 Specialists out-compete the giants at a lot of steps, sourcing the people, building environments, designing rewards, running evals.
  20. 1:57 The fragmentation, I would say, is pretty permanent.
  21. 1:59 It's not transitional.
  22. 2:01 And quality increasingly does not scale linearly with quantity, which leads to a sort of cottage industry in data land right now where you have labs literally mandate vendor diversification on the scale of like 20 to 30 different vendors because they inherently distrust.
  23. 2:16 their ability to scale quality with quantity.
  24. 2:19 So to orient you, here's the sort of argument as it developed this year.
  25. 2:26 January, we saw a lot of industrialization and unbundling.
  26. 2:30 You'll hear me come back to this mechanism I call the Antikythera mechanisms, which is this sort of bespoke systems that translate messy business context into evals, increasingly important in labs hungry quest for real-world seed data to set seed ends.
  27. 2:44 I also explore this concept called type one versus type two data, contrived versus non-contrived for the more researcher types in the audience and why the GPQA cell playbook that works in 2024 for data acquisition kind of falls apart on long horizon non-verifiable work.
  28. 2:59 So, you know, going through all these topics, all of which are online, I'll actually just skip to the more interesting part, but I'm bringing this up just here in case any of the particular topics I talk about are particularly interesting and you want to dive in more.
  29. 3:14 So model improvement is a function of three inputs.
  30. 3:17 We all sort of see this commonly expressed in the computes, data, and talent.
  31. 3:22 I put together this very rudimentary sort of chart online, and one of my pieces in a very rudimentary equation just to express the fact that if there's any sort of imbalance in this compute, data, and talent, then you start seeing a sort of inefficiency in producing what I call generalized AI model performance.
  32. 3:42 But also, we see a sort of inefficiency in capex spend that is a sort of result of this equation right now, exponentially increasing capex spend, but sort of like AI revenues are sort of falling far behind.
  33. 3:57 Data is sort of the underfunded leg here.
  34. 4:00 It's the one that sort of turns a generalist model into a real expert.
  35. 4:02 It's the one that's actually, I believe, quite lacking in this equation, and thus, because of the imbalance,
  36. 4:09 parameter concepts that I'm expressing presents a whole opportunity.
  37. 4:15 But going back to what I said at the start, what is data actually?
  38. 4:20 So most people picture state-based data, the rows in an ERP, which is like the final output or saved file.
  39. 4:26 That's kind of the 2023 like next token prediction model.
  40. 4:30 And it's mostly personal data wrapped in privacy law.
  41. 4:33 What is actually available nowadays is process-based data, which is the trajectory, the reasoning trace, the sequence of decisions.
  42. 4:41 So it's kind of what gets a professional to get from a blank page to a finished work output and sort of delineates how the work gets done.
  43. 4:49 So on top of that, it's a quality access.
  44. 4:51 This is the vocabulary I'll use all talk.
  45. 4:54 Type one data is a sort of pure capture of real workflows like GitHub commits or session replays.
  46. 5:00 with minimal reward shaping by non-experts, and type two is contrived data, where you sort of hire experts, you sit them in an arbitrary setting, you have them manufacture examples.
  47. 5:09 Type two, right place to start.
  48. 5:11 Models, when they were reading at a first grade level, I think anybody in the world could teach them as a third grade teacher.
  49. 5:19 But type 1 is what gets you from 20% to 80%, so to say.
  50. 5:23 Because the realism is inherited from the work itself.

Chapters

  1. 0:00 The data market nobody sees
  2. 1:15 Data as industrial fuel
  3. 2:31 Type one and type two data
  4. 3:23 Compute, data, and talent
  5. 4:24 State data versus process data
  6. 5:54 The three axes of verifiability
  7. 8:14 When benchmarks become snake oil
  8. 10:08 Three finance benchmark tests
  9. 12:04 Predicting the next AI domain
  10. 13:21 The robotics counterexample
  11. 14:22 Where the economic value lives
  12. 16:02 Why data companies move enterprise
  13. 17:09 The durable moat

Open at this second