read-only demo

Videos _PdK6x7PQNM

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:19:05· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:41, 1 of 1 keyframes kept
  5. Shot 4, 0:41 to 1:10, 1 of 1 keyframes kept
  6. Shot 5, 1:10 to 1:39, 1 of 1 keyframes kept
  7. Shot 6, 1:39 to 2:08, 1 of 1 keyframes kept
  8. Shot 7, 2:08 to 2:54, 1 of 1 keyframes kept
  9. Shot 8, 2:54 to 3:25, 1 of 1 keyframes kept
  10. Shot 9, 3:25 to 3:55, 0 of 1 keyframes kept
  11. Shot 10, 3:55 to 4:25, 1 of 1 keyframes kept
  12. Shot 11, 4:25 to 4:55, 0 of 1 keyframes kept
  13. Shot 12, 4:55 to 5:25, 0 of 1 keyframes kept
  14. Shot 13, 5:25 to 5:55, 0 of 1 keyframes kept
  15. Shot 14, 5:55 to 6:03, 1 of 1 keyframes kept
  16. Shot 15, 6:03 to 6:29, 1 of 1 keyframes kept
  17. Shot 16, 6:29 to 6:55, 0 of 1 keyframes kept
  18. Shot 17, 6:55 to 7:03, 0 of 1 keyframes kept
  19. Shot 18, 7:03 to 7:32, 1 of 1 keyframes kept
  20. Shot 19, 7:32 to 8:01, 0 of 1 keyframes kept
  21. Shot 20, 8:01 to 8:30, 1 of 1 keyframes kept
  22. Shot 21, 8:30 to 9:14, 1 of 1 keyframes kept
  23. Shot 22, 9:14 to 9:42, 0 of 1 keyframes kept
  24. Shot 23, 9:42 to 10:14, 0 of 1 keyframes kept
  25. Shot 24, 10:14 to 10:46, 1 of 1 keyframes kept
  26. Shot 25, 10:46 to 11:18, 0 of 1 keyframes kept
  27. Shot 26, 11:18 to 11:46, 1 of 1 keyframes kept
  28. Shot 27, 11:46 to 12:13, 1 of 1 keyframes kept
  29. Shot 28, 12:13 to 12:47, 0 of 1 keyframes kept
  30. Shot 29, 12:47 to 12:54, 0 of 1 keyframes kept
  31. Shot 30, 12:54 to 13:40, 1 of 1 keyframes kept
  32. Shot 31, 13:40 to 13:55, 1 of 1 keyframes kept
  33. Shot 32, 13:55 to 14:04, 1 of 1 keyframes kept
  34. Shot 33, 14:04 to 14:35, 1 of 1 keyframes kept
  35. Shot 34, 14:35 to 15:07, 0 of 1 keyframes kept
  36. Shot 35, 15:07 to 15:38, 1 of 1 keyframes kept
  37. Shot 36, 15:38 to 16:18, 1 of 1 keyframes kept
  38. Shot 37, 16:18 to 16:46, 1 of 1 keyframes kept
  39. Shot 38, 16:46 to 17:14, 0 of 1 keyframes kept
  40. Shot 39, 17:14 to 17:42, 0 of 1 keyframes kept
  41. Shot 40, 17:42 to 18:31, 1 of 1 keyframes kept
  42. Shot 41, 18:31 to 18:43, 1 of 1 keyframes kept
  43. Shot 42, 18:43 to 18:48, 1 of 1 keyframes kept
  44. Shot 43, 18:48 to 18:50, 0 of 1 keyframes kept
  45. Shot 44, 18:50 to 19:03, 0 of 1 keyframes kept
  46. Shot 45, 19:03 to 19:04, 1 of 1 keyframes kept

46 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
199
whisperx 199
chunks
34
from 199 cues
keyframes
29
kept of 46 captured
frames with text
28
778 lines read
chapters
11
from the source metadata
keyframe bytes
6.0 MB
word timings on 199 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:25 0s
stt done 2026-08-09 14:16 23s
chunk done 2026-08-09 14:17 0s
text_embed done 2026-08-10 19:42 1s
keyframe done 2026-08-09 14:17 2m 26s
ocr done 2026-08-09 14:19 14s
frame_embed done 2026-08-10 19:42 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.7

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2743.2

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:18 #3 done3 line(s)

    shot 3·sharpness 323.8

    1. datologyai1.00
    2. AlEngineer1.00
    3. World's Fair0.99
  • 1:01 #4 done3 line(s)

    shot 4·sharpness 304.1

    1. datologyai0.99
    2. AlEngineer1.00
    3. World's Fair0.98
  • 1:35 #5 done41 line(s)

    shot 5·sharpness 2820.4

    1. AlEngineer0.98
    2. Compute availability is becoming increasingly scarce0.99
    3. World's Fair0.98
    4. GPU prices are rising1.00
    5. Models use more and1.00
    6. Access to APl tokens0.98
    7. more tokens1.00
    8. may go away0.97
    9. PRESENTED BY1.00
    10. H100 1-year rental ·$/GPU/hr0.96
    11. +40%1.00
    12. $2.401.00
    13. output tokens per response1.00
    14. reasoning +5×/yr0.97
    15. Artificialintelie + Addeign myFT0.70
    16. FINANCIAL TIMES1.00
    17. Google caps Meta's Gemini use as AI demand strains0.99
    18. Microsoft1.00
    19. capacity1.00
    20. reasoning1.00
    21. rech industr's carcest comodity0.89
    22. Surging appetite for advanced models is turning computing power into the0.98
    23. ~8×0.95
    24. OpenAl0.94
    25. Guaranteed1.00
    26. $1.701.00
    27. non-reasoning1.00
    28. Capacity1.00
    29. Guarantee long-term access to OpenAl0.99
    30. compute for the products, agents, and0.96
    31. Oct 20251.00
    32. Apr 20261.00
    33. 20231.00
    34. 20261.00
    35. customer workflows that mattermost.0.96
    36. Plan for capacity0.92
    37. Contact sales0.96
    38. SemiAnalysis H100 Rental Index0.99
    39. Epoch Al1.00
    40. World'sFair1.00
    41. Engineering the future of Al1.00
  • 1:48 #6 done42 line(s)

    shot 6·sharpness 2761.7

    1. AlEngineer0.99
    2. Compute availability is becoming increasingly scarce0.99
    3. World's Fair0.97
    4. GPU prices are rising1.00
    5. Models use more and1.00
    6. Access to APl tokens0.98
    7. more tokens1.00
    8. may go away0.96
    9. PRESENTED BY1.00
    10. H100 1-year rental·$/GPU/hr0.97
    11. output tokens per response1.00
    12. FINANCIAL TIMES1.00
    13. +40%1.00
    14. $2.401.00
    15. reasoning +5×/yr0.98
    16. Artificialitililic + AdligenmyFT0.75
    17. Google caps Meta's Gemini use as AI demand strains1.00
    18. Microsoft1.00
    19. capacity1.00
    20. reasoning1.00
    21. tech industr's carcest comodity0.92
    22. Surging appetite for advanced models is turning computing power into the0.99
    23. ~8×0.95
    24. OpenAl0.93
    25. Guaranteed1.00
    26. $1.701.00
    27. non-reasoning1.00
    28. Capacity1.00
    29. Guarantee long-term access to OpenAl0.99
    30. Oct 20251.00
    31. Apr 20261.00
    32. 20231.00
    33. 20261.00
    34. customer workflows that mattermost.0.97
    35. compute for the products, agents, and1.00
    36. Plan for capacity0.95
    37. Contact sales0.98
    38. SemiAnalysis H100 Rental Index0.98
    39. Epoch Al1.00
    40. World's Fair0.98
    41. TRACK 9·JUNE 30,20260.98
    42. Data Quality1.00
  • 2:35 #7 done17 line(s)

    shot 7·sharpness 2836.5

    1. Data quality is a compute multiplier1.00
    2. AlEngineer0.97
    3. Better data leads to better performance for the same compute budget1.00
    4. World'sFair0.96
    5. Better1.00
    6. performance1.00
    7. Pernce0.73
    8. Increase in0.96
    9. performance for1.00
    10. the same0.96
    11. budget1.00
    12. Worse1.00
    13. performance1.00
    14. Number of Data Points1.00
    15. World'sFair1.00
    16. TRACK 9· JUNE 30, 20260.97
    17. Data Quality1.00
  • 3:01 #8 done12 line(s)

    shot 8·sharpness 3518.1

    1. AlEngineer0.97
    2. World'sFair1.00
    3. Data Quality for Al Model Training0.99
    4. More signal per token, not more tokens1.00
    5. Relevant - to your use-cases0.99
    6. Diverse - covers the gaps, no blind spots0.99
    7. Information dense - every token earns its place0.99
    8. In the right mix - optimize allocation of general vs.1.00
    9. specific data1.00
    10. World'sFair1.00
    11. TRACK 9· JUNE 30, 20260.92
    12. Data Quality0.98
  • 3:31 #9 skipped

    shot 9·duplicate of #8

  • 4:10 #10 done72 line(s)

    shot 10·sharpness 3030.0

    1. AlEngineer0.98
    2. The Data Curation Pipeline0.99
    3. World's Fair0.99
    4. YOUR DATA1.00
    5. Public Datasets1.00
    6. CLEAN1.00
    7. CURATE1.00
    8. CREATE1.00
    9. COMPOSE1.00
    10. Open-source datasets1.00
    11. at internet scale1.00
    12. Licensed Data1.00
    13. Publishers, annotators,1.00
    14. CommonCrawd · RefinedWeb0.96
    15. Proprietary Data1.00
    16. Customer-owned1.00
    17. enterprise data1.00
    18. S3 - GCS · Azure0.95
    19. based on length, symbol-to-word0.99
    20. unify schema, repair encoding0.99
    21. Remove degenerate samples1.00
    22. Connect customer sources;1.00
    23. Heuristic Filtering1.00
    24. Source Ingestion1.00
    25. ratio, etc.1.00
    26. quality and taxonomy classifiers1.00
    27. n-gram and embedding-based0.99
    28. Redundancy Reduction1.00
    29. Quality and Taxonomy0.98
    30. sample similarity rejection1.00
    31. Heuristic and learned1.00
    32. Annotation1.00
    33. Multiple signals, including quality1.00
    34. and task-relevance, are used to1.00
    35. select documents for rephrasing1.00
    36. Robust Rephrasing1.00
    37. Targeted Selection1.00
    38. Synthetic Data1.00
    39. Generation1.00
    40. Principled, quantitative mixture1.00
    41. Multi-Phase Composition0.97
    42. mid-training and annealing0.97
    43. Stage-aligned datasets for0.99
    44. Algorithmic Mixing1.00
    45. design1.00
    46. Multi-stage1.00
    47. Datasets1.00
    48. Training1.00
    49. partners, premium content0.99
    50. Commercial licenses1.00
    51. across all training sources1.00
    52. N-gram leakage check1.00
    53. Decontamination1.00
    54. Benchmark1.00
    55. Quality-based Resampling0.98
    56. Adjust the sampling distribution1.00
    57. to emphasize high-quality data1.00
    58. Rephrasing approaches avoid0.99
    59. the failure modes of de novo0.97
    60. Diversity Maximization1.00
    61. generation1.00
    62. Carefully-tuned rephrasing1.00
    63. Task Distribution Matching1.00
    64. diversity and the mileage of your1.00
    65. methodology maximizes1.00
    66. Reshape the corpus to your0.99
    67. data1.00
    68. use-cases, using examples you1.00
    69. provide1.00
    70. World's Fair0.99
    71. TRACK 9• JUNE 30, 20260.95
    72. Data Quality1.00
  • 4:46 #11 skipped

    shot 11·duplicate of #10

  • 5:16 #12 skipped

    shot 12·duplicate of #10

  • 5:38 #13 skipped

    shot 13·duplicate of #10

  • 5:58 #14 done9 line(s)

    shot 14·sharpness 1921.4

    1. AlEngineer0.97
    2. World'sFair1.00
    3. RESEARCH RESULTS + CUSTOMER EVIDENCE1.00
    4. Proof that Curation Works1.00
    5. Across research frontiers, model modalities, and real customer0.99
    6. deployments1.00
    7. World'sFair1.00
    8. TRACK 9· JUNE 30, 20260.94
    9. Data Quality0.99
  • 6:26 #15 done49 line(s)

    shot 15·sharpness 5254.3

    1. Frontier Data Research is the Foundation of DatalogyAl0.99
    2. AlEngineer0.99
    3. World'sFair1.00
    4. Published research spanning curation algorithms, synthetic data, multilingual scaling, VLMs, and more.1.00
    5. Beyond neural scaling laws:0.98
    6. BeyondWeb1.00
    7. ÜberWeb0.99
    8. beating power law scaling via data pruning1.00
    9. Lessons from Scaling Synthetic Data0.98
    10. Insights from Multilingual Curation1.00
    11. for Trillion-scale Pretraining1.00
    12. Multilingual curation that provides0.99
    13. PRESENTED BY1.00
    14. Data curation can beat power-law0.97
    15. scaling.1.00
    16. high-quality documents into more1.00
    17. Targeted rephrasing turns1.00
    18. frontier-quality multilingual1.00
    19. cross-lingual transfer and1.00
    20. Microsoft1.00
    21. NeurlPS Outstanding Paper0.98
    22. learnable pretraining data and1.00
    23. avoids model collapse.1.00
    24. capabilities at 1/10th the training0.99
    25. compute1.00
    26. 20/20 Vision Language Models:0.99
    27. Brevity is the Soul of1.00
    28. The Finetuner's Fallacy1.00
    29. When to Pretrain with Your Finetuning Data1.00
    30. A Prescription for Better VLMs through Data Curation Alone1.00
    31. Inference Efficiency:1.00
    32. Inducing Concision in VLMs via Data Curation0.99
    33. Mixing even very small quantities0.99
    34. Near-frontier model quality at up1.00
    35. of specialized domain data into0.99
    36. to ~ 150× less training compute0.99
    37. Data curation as a lever for inference0.99
    38. pretraining prevents overfitting1.00
    39. achieved only through VLM0.99
    40. efficiency by training models to be0.99
    41. and catastrophic forgetting,1.00
    42. pretraining data curation0.98
    43. more concise; up to 35x more1.00
    44. efficient than frontier models1.00
    45. leading to substantially improved1.00
    46. fine-tuning performance1.00
    47. World's Fair0.97
    48. TRACK 9·JUNE 30,20260.98
    49. Data Quality0.99
  • 6:44 #16 skipped

    shot 16·duplicate of #15

  • 7:02 #17 skipped

    shot 17·duplicate of #14

  • 7:06 #18 done37 line(s)

    shot 18·sharpness 2156.0

    1. AlEngineer0.98
    2. VLM Curation Defines a New Training Pareto Frontier0.99
    3. World's Fair0.99
    4. All Public Benchmarks1.00
    5. Model1.00
    6. MAmmoTH-VL 180.98
    7. Erg - cy)0.68
    8. 50%1.00
    9. MAmmoTH-VL 2B0.98
    10. MAmmoTH-VL 4B1.00
    11. PRESENTED BY1.00
    12. Datology 1B0.99
    13. Microsoft1.00
    14. Datology 2B1.00
    15. Datology 4B1.00
    16. 40%1.00
    17. InternVL3-2B1.00
    18. InternVL3.5-2B0.98
    19. InternVL3.5-4B0.99
    20. better1.00
    21. Qwen3-VL-2B1.00
    22. 30%1.00
    23. Qwen3-VL-4B1.00
    24. Qwen3.5-2B0.98
    25. cheaper1.00
    26. Size1.00
    27. 10200.96
    28. 10{210.78
    29. 10220.87
    30. 10230.86
    31. ● 1B0.85
    32. 2B1.00
    33. ◆4B0.87
    34. Training FLOPs1.00
    35. World's Fair0.99
    36. TRACK 9· JUNE 30,20260.96
    37. Data Quality1.00
  • 7:54 #19 skipped

    shot 19·duplicate of #18

  • 8:13 #20 done41 line(s)

    shot 20·sharpness 2348.7

    1. AlEngineer0.98
    2. VLM Curation Defines a New Training Pareto Frontier0.99
    3. World's Fair0.99
    4. All Public Benchmarks1.00
    5. Model1.00
    6. MAmmoTH-VL 1B0.99
    7. Er - cy)0.74
    8. 50%1.00
    9. MAmmoTH-VL 2B0.96
    10. MAmmoTH-VL 4B1.00
    11. Datology 1B1.00
    12. Datology 2B1.00
    13. MammoTH-VL1.00
    14. +14.0pp over1.00
    15. Datology 4B1.00
    16. 40%1.00
    17. input data1.00
    18. InternVL3-2B1.00
    19. InternVL3.5-2B0.99
    20. InternVL3.5-4B1.00
    21. better1.00
    22. Qwen3-VL-2B1.00
    23. 30%1.00
    24. Qwen3-VL-4B1.00
    25. 145x less train1.00
    26. Qwen3.5-2B0.99
    27. cheaper1.00
    28. compute than1.00
    29. Qwen3.5-4B1.00
    30. Size1.00
    31. 10200.97
    32. 10{210.78
    33. 10220.98
    34. 10230.99
    35. ● 1B0.83
    36. 2B1.00
    37. ◆4B0.86
    38. Training FLOPs1.00
    39. World's Fair0.99
    40. TRACK 9· JUNE 30, 20260.95
    41. Data Quality0.99
  • 9:05 #21 done75 line(s)

    shot 21·sharpness 2331.1

    1. AlEngineer0.99
    2. Data curation → Concise Model → Cheaper Inference0.99
    3. World's Fair0.99
    4. Verbosity: mean tokens per response0.98
    5. Cost: FLOPs per correct answer1.00
    6. 50%1.00
    7. IV3-2B1.00
    8. 301.00
    9. D-1B1.00
    10. D-4B1.00
    11. Ert a - a oer)0.61
    12. 45%1.00
    13. M-4B0.99
    14. D-2B1.00
    15. 421.00
    16. M-2B0.96
    17. 561.00
    18. M-1B1.00
    19. 571.00
    20. M-4B1.00
    21. 601.00
    22. 40%1.00
    23. IV3.5-4B1.00
    24. 851.00
    25. P-Isaac-2B1.00
    26. 901.00
    27. D-1B1.00
    28. Q3VL-2B1.00
    29. Q3VL-4B1.00
    30. 1021.00
    31. 1091.00
    32. 35%1.00
    33. better0.93
    34. D-2B0.99
    35. 35x FLOPs per0.97
    36. IV3.5-2B0.99
    37. Q3.5-2B1.00
    38. 1291.00
    39. 1691.00
    40. 30%1.00
    41. D-4B1.00
    42. Q3.5-2B0.95
    43. reduction vs1.00
    44. Qwen3.5-4B1.00
    45. correct1.00
    46. Q3.5-4B1.00
    47. Q3.5-4B1.00
    48. 1,2840.96
    49. cheaper1.00
    50. 251.00
    51. 501.00
    52. 751.00
    53. 1001.00
    54. 1251.00
    55. 1501.00
    56. 1751.00
    57. 1,2840.96
    58. 1010.86
    59. 10120.96
    60. 10130.96
    61. Mean output tokens per response (lower is better)0.99
    62. FLOPs per correct answer (log, lower is better)1.00
    63. Family1.00
    64. Model1.00
    65. MAmmoTH-VL1.00
    66. DatologyAl0.97
    67. InternVL3InternVL3.50.99
    68. Qwen3-VLQwen3.5Perceptron1.00
    69. Size1.00
    70. ● 1B0.85
    71. 2B1.00
    72. 4B0.99
    73. World's Fair0.98
    74. TRACK 9·JUNE 30,20260.97
    75. Data Quality1.00
  • 9:23 #22 skipped

    shot 22·duplicate of #14

  • 10:01 #23 skipped

    shot 23·duplicate of #7

Transcript

199 cues· 3,822 words· 21,171 chars

  1. 0:13 Good morning, everybody.
  2. 0:15 My name's Ari Marcos.
  3. 0:16 I'm the CEO and co-founder of Datology AI and really excited to kick off the data quality track today.
  4. 0:23 Data quality is what we live and breathe at Datology.
  5. 0:26 It's all we think about.
  6. 0:27 In fact, the company's name literally means the science or study of data.
  7. 0:31 So very excited to see the increasing excitement and interest in this area and an amazing lineup of talks today.
  8. 0:38 So today I'm gonna tell you about why data quality is a compute multiplier that we're all overlooking and where we can make massive gains by just working on better data.
  9. 0:48 We've seen that compute availability over the last six months has become extremely scarce and is only getting worse.
  10. 0:55 We saw H100 prices reverse, their several year long drop, which is normal for hardware, and all of a sudden come up where now they're about 40% up from their lows at the end of last year.
  11. 1:06 As test time compute has become a critical part of models, and as we have put more and more thinking tokens in, we're seeing that the number, the token usage is absolutely skyrocketing.
  12. 1:16 Reasoning models use eight times as many tokens as non-reasoning models, and that's projected to 5x again in the next year or so.
  13. 1:23 So the number of tokens we're pushing through goes higher and higher, and that constrains compute even further.
  14. 1:28 And this has led actually to a world where it's not implausible today that we might see access to some of the frontier APIs actually get limited or go away.
  15. 1:36 As one example, Google just capped Meta's Gemini usage because of inference constraints.
  16. 1:41 OpenAI has effectively started selling token futures, or you can guarantee token capacity some amount of time into the future.
  17. 1:47 This is only necessary because they are legitimately wondering there might be a world where access to frontier API tokens is limited, not as a business decision, but because there's just simply not enough inference.
  18. 1:58 first-party products will be prioritized.
  19. 2:00 So in a world where compute is increasingly scarce, and you need to make models better, well, what do you do?
  20. 2:05 Well, we work on data.
  21. 2:06 You're gonna hear me say this over and over again, data quality is a compute multiplier, because what it does is it makes the learning curve steeper.
  22. 2:13 So this is just a very simple schematic of performance on the y-axis as a function of data on the x-axis.
  23. 2:18 Note that this x-axis, you could swap it out for data, for compute, for time, for dollars, they're all the same x-axis fundamentally.
  24. 2:26 And if you can make data quality better, you can turn this gray curve into this blue curve.
  25. 2:32 And that means that now you can get dramatically better performance for the same compute budget as if you had trained with far more compute.
  26. 2:40 And similarly, you can get the same performance for a much smaller compute budget, which is exactly showing you how you can get performance as if you had spent 10 times as much on compute and really shows this compute multiplier point.
  27. 2:55 So how do you actually do this?
  28. 2:56 Well, fundamentally the idea is we wanna make it so that we get the maximum signal per token and per batch.
  29. 3:02 For those that are a little bit more technically, what we wanna do here is maximize the marginal information gain per data point when we show it to the model.
  30. 3:09 What data is gonna teach the model the most?
  31. 3:11 And that's all about finding data that's relevant to the use cases that you want.
  32. 3:16 One thing that is very important is that there's no one golden data set to rule them all that's good for everything no matter what you wanna do.
  33. 3:23 A data set's only gonna be optimal with respect to a particular set of output tasks that you want the model to do.
  34. 3:28 So if you want a great legal model, you're gonna want legal data more than healthcare data and vice versa.
  35. 3:33 It needs to be diverse.
  36. 3:34 A lot of the issues we see with model robustness and brittleness comes from training on data that's not diverse enough.
  37. 3:40 So the model can answer a question correctly if it's presented just so, but if it's presented a little bit differently, now everything breaks.
  38. 3:47 It needs to be information dense, and you have to mix the data correctly.
  39. 3:49 This is a hugely difficult part.
  40. 3:51 You now have many different sources.
  41. 3:52 How do you combine them to actually drive the largest improvement in performance?
  42. 3:56 So this is a high level of what we do at Datology here.
  43. 4:00 You can think of us as the oil refinery for data.
  44. 4:02 We don't source new tokens like many data providers.
  45. 4:04 Rather, we take existing tokens coming from public data sets, proprietary data sets, and licensed data sets, and make them way better.
  46. 4:11 And how do we do that?
  47. 4:11 Well, we do that through these four Cs, clean, curate, create, and compose.
  48. 4:16 So cleaning is fairly straightforward.
  49. 4:17 This is doing things like heuristic filters all in gopher and things like that.
  50. 4:21 Removing documents that have only 10 characters in them or are all wing ding.

Chapters

  1. 0:00 Data is all we think about
  2. 0:52 Why good data became scarce
  3. 2:19 Swapping compute for data on the curve
  4. 3:48 An oil refinery for data
  5. 5:52 Curation work at DatologyAI
  6. 6:54 Proof: small models beating bigger ones
  7. 8:58 Inference efficiency from better data
  8. 9:24 Multilingual gains
  9. 12:16 Synthetic data done right
  10. 14:10 Thomson Reuters and Arcee results
  11. 17:43 Cheaper than buying compute

Open at this second