Videos _PdK6x7PQNM
Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Scene timeline
46 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 199
- whisperx 199
- chunks
- 34
- from 199 cues
- keyframes
- 29
- kept of 46 captured
- frames with text
- 28
- 778 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 6.0 MB
- word timings on 199 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:25 | 0s |
stt |
done | — | 2026-08-09 14:16 | 23s |
chunk |
done | — | 2026-08-09 14:17 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 1s |
keyframe |
done | — | 2026-08-09 14:17 | 2m 26s |
ocr |
done | — | 2026-08-09 14:19 | 14s |
frame_embed |
done | — | 2026-08-10 19:42 | 5s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- datologyai1.00
- AlEngineer1.00
- World's Fair0.99
-
- datologyai0.99
- AlEngineer1.00
- World's Fair0.98
-
- AlEngineer0.98
- Compute availability is becoming increasingly scarce0.99
- World's Fair0.98
- GPU prices are rising1.00
- Models use more and1.00
- Access to APl tokens0.98
- more tokens1.00
- may go away0.97
- PRESENTED BY1.00
- H100 1-year rental ·$/GPU/hr0.96
- +40%1.00
- $2.401.00
- output tokens per response1.00
- reasoning +5×/yr0.97
- Artificialintelie + Addeign myFT0.70
- FINANCIAL TIMES1.00
- Google caps Meta's Gemini use as AI demand strains0.99
- Microsoft1.00
- capacity1.00
- reasoning1.00
- rech industr's carcest comodity0.89
- Surging appetite for advanced models is turning computing power into the0.98
- ~8×0.95
- OpenAl0.94
- Guaranteed1.00
- $1.701.00
- non-reasoning1.00
- Capacity1.00
- Guarantee long-term access to OpenAl0.99
- compute for the products, agents, and0.96
- Oct 20251.00
- Apr 20261.00
- 20231.00
- 20261.00
- customer workflows that mattermost.0.96
- Plan for capacity0.92
- Contact sales0.96
- SemiAnalysis H100 Rental Index0.99
- Epoch Al1.00
- World'sFair1.00
- Engineering the future of Al1.00
-
- AlEngineer0.99
- Compute availability is becoming increasingly scarce0.99
- World's Fair0.97
- GPU prices are rising1.00
- Models use more and1.00
- Access to APl tokens0.98
- more tokens1.00
- may go away0.96
- PRESENTED BY1.00
- H100 1-year rental·$/GPU/hr0.97
- output tokens per response1.00
- FINANCIAL TIMES1.00
- +40%1.00
- $2.401.00
- reasoning +5×/yr0.98
- Artificialitililic + AdligenmyFT0.75
- Google caps Meta's Gemini use as AI demand strains1.00
- Microsoft1.00
- capacity1.00
- reasoning1.00
- tech industr's carcest comodity0.92
- Surging appetite for advanced models is turning computing power into the0.99
- ~8×0.95
- OpenAl0.93
- Guaranteed1.00
- $1.701.00
- non-reasoning1.00
- Capacity1.00
- Guarantee long-term access to OpenAl0.99
- Oct 20251.00
- Apr 20261.00
- 20231.00
- 20261.00
- customer workflows that mattermost.0.97
- compute for the products, agents, and1.00
- Plan for capacity0.95
- Contact sales0.98
- SemiAnalysis H100 Rental Index0.98
- Epoch Al1.00
- World's Fair0.98
- TRACK 9·JUNE 30,20260.98
- Data Quality1.00
-
- Data quality is a compute multiplier1.00
- AlEngineer0.97
- Better data leads to better performance for the same compute budget1.00
- World'sFair0.96
- Better1.00
- performance1.00
- Pernce0.73
- Increase in0.96
- performance for1.00
- the same0.96
- budget1.00
- Worse1.00
- performance1.00
- Number of Data Points1.00
- World'sFair1.00
- TRACK 9· JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.97
- World'sFair1.00
- Data Quality for Al Model Training0.99
- More signal per token, not more tokens1.00
- Relevant - to your use-cases0.99
- Diverse - covers the gaps, no blind spots0.99
- Information dense - every token earns its place0.99
- In the right mix - optimize allocation of general vs.1.00
- specific data1.00
- World'sFair1.00
- TRACK 9· JUNE 30, 20260.92
- Data Quality0.98
-
- AlEngineer0.98
- The Data Curation Pipeline0.99
- World's Fair0.99
- YOUR DATA1.00
- Public Datasets1.00
- CLEAN1.00
- CURATE1.00
- CREATE1.00
- COMPOSE1.00
- Open-source datasets1.00
- at internet scale1.00
- Licensed Data1.00
- Publishers, annotators,1.00
- CommonCrawd · RefinedWeb0.96
- Proprietary Data1.00
- Customer-owned1.00
- enterprise data1.00
- S3 - GCS · Azure0.95
- based on length, symbol-to-word0.99
- unify schema, repair encoding0.99
- Remove degenerate samples1.00
- Connect customer sources;1.00
- Heuristic Filtering1.00
- Source Ingestion1.00
- ratio, etc.1.00
- quality and taxonomy classifiers1.00
- n-gram and embedding-based0.99
- Redundancy Reduction1.00
- Quality and Taxonomy0.98
- sample similarity rejection1.00
- Heuristic and learned1.00
- Annotation1.00
- Multiple signals, including quality1.00
- and task-relevance, are used to1.00
- select documents for rephrasing1.00
- Robust Rephrasing1.00
- Targeted Selection1.00
- Synthetic Data1.00
- Generation1.00
- Principled, quantitative mixture1.00
- Multi-Phase Composition0.97
- mid-training and annealing0.97
- Stage-aligned datasets for0.99
- Algorithmic Mixing1.00
- design1.00
- Multi-stage1.00
- Datasets1.00
- Training1.00
- partners, premium content0.99
- Commercial licenses1.00
- across all training sources1.00
- N-gram leakage check1.00
- Decontamination1.00
- Benchmark1.00
- Quality-based Resampling0.98
- Adjust the sampling distribution1.00
- to emphasize high-quality data1.00
- Rephrasing approaches avoid0.99
- the failure modes of de novo0.97
- Diversity Maximization1.00
- generation1.00
- Carefully-tuned rephrasing1.00
- Task Distribution Matching1.00
- diversity and the mileage of your1.00
- methodology maximizes1.00
- Reshape the corpus to your0.99
- data1.00
- use-cases, using examples you1.00
- provide1.00
- World's Fair0.99
- TRACK 9• JUNE 30, 20260.95
- Data Quality1.00
-
- AlEngineer0.97
- World'sFair1.00
- RESEARCH RESULTS + CUSTOMER EVIDENCE1.00
- Proof that Curation Works1.00
- Across research frontiers, model modalities, and real customer0.99
- deployments1.00
- World'sFair1.00
- TRACK 9· JUNE 30, 20260.94
- Data Quality0.99
-
- Frontier Data Research is the Foundation of DatalogyAl0.99
- AlEngineer0.99
- World'sFair1.00
- Published research spanning curation algorithms, synthetic data, multilingual scaling, VLMs, and more.1.00
- Beyond neural scaling laws:0.98
- BeyondWeb1.00
- ÜberWeb0.99
- beating power law scaling via data pruning1.00
- Lessons from Scaling Synthetic Data0.98
- Insights from Multilingual Curation1.00
- for Trillion-scale Pretraining1.00
- Multilingual curation that provides0.99
- PRESENTED BY1.00
- Data curation can beat power-law0.97
- scaling.1.00
- high-quality documents into more1.00
- Targeted rephrasing turns1.00
- frontier-quality multilingual1.00
- cross-lingual transfer and1.00
- Microsoft1.00
- NeurlPS Outstanding Paper0.98
- learnable pretraining data and1.00
- avoids model collapse.1.00
- capabilities at 1/10th the training0.99
- compute1.00
- 20/20 Vision Language Models:0.99
- Brevity is the Soul of1.00
- The Finetuner's Fallacy1.00
- When to Pretrain with Your Finetuning Data1.00
- A Prescription for Better VLMs through Data Curation Alone1.00
- Inference Efficiency:1.00
- Inducing Concision in VLMs via Data Curation0.99
- Mixing even very small quantities0.99
- Near-frontier model quality at up1.00
- of specialized domain data into0.99
- to ~ 150× less training compute0.99
- Data curation as a lever for inference0.99
- pretraining prevents overfitting1.00
- achieved only through VLM0.99
- efficiency by training models to be0.99
- and catastrophic forgetting,1.00
- pretraining data curation0.98
- more concise; up to 35x more1.00
- efficient than frontier models1.00
- leading to substantially improved1.00
- fine-tuning performance1.00
- World's Fair0.97
- TRACK 9·JUNE 30,20260.98
- Data Quality0.99
-
- AlEngineer0.98
- VLM Curation Defines a New Training Pareto Frontier0.99
- World's Fair0.99
- All Public Benchmarks1.00
- Model1.00
- MAmmoTH-VL 180.98
- Erg - cy)0.68
- 50%1.00
- MAmmoTH-VL 2B0.98
- MAmmoTH-VL 4B1.00
- PRESENTED BY1.00
- Datology 1B0.99
- Microsoft1.00
- Datology 2B1.00
- Datology 4B1.00
- 40%1.00
- InternVL3-2B1.00
- InternVL3.5-2B0.98
- InternVL3.5-4B0.99
- better1.00
- Qwen3-VL-2B1.00
- 30%1.00
- Qwen3-VL-4B1.00
- Qwen3.5-2B0.98
- cheaper1.00
- Size1.00
- 10200.96
- 10{210.78
- 10220.87
- 10230.86
- ● 1B0.85
- 2B1.00
- ◆4B0.87
- Training FLOPs1.00
- World's Fair0.99
- TRACK 9· JUNE 30,20260.96
- Data Quality1.00
-
- AlEngineer0.98
- VLM Curation Defines a New Training Pareto Frontier0.99
- World's Fair0.99
- All Public Benchmarks1.00
- Model1.00
- MAmmoTH-VL 1B0.99
- Er - cy)0.74
- 50%1.00
- MAmmoTH-VL 2B0.96
- MAmmoTH-VL 4B1.00
- Datology 1B1.00
- Datology 2B1.00
- MammoTH-VL1.00
- +14.0pp over1.00
- Datology 4B1.00
- 40%1.00
- input data1.00
- InternVL3-2B1.00
- InternVL3.5-2B0.99
- InternVL3.5-4B1.00
- better1.00
- Qwen3-VL-2B1.00
- 30%1.00
- Qwen3-VL-4B1.00
- 145x less train1.00
- Qwen3.5-2B0.99
- cheaper1.00
- compute than1.00
- Qwen3.5-4B1.00
- Size1.00
- 10200.97
- 10{210.78
- 10220.98
- 10230.99
- ● 1B0.83
- 2B1.00
- ◆4B0.86
- Training FLOPs1.00
- World's Fair0.99
- TRACK 9· JUNE 30, 20260.95
- Data Quality0.99
-
- AlEngineer0.99
- Data curation → Concise Model → Cheaper Inference0.99
- World's Fair0.99
- Verbosity: mean tokens per response0.98
- Cost: FLOPs per correct answer1.00
- 50%1.00
- IV3-2B1.00
- 301.00
- D-1B1.00
- D-4B1.00
- Ert a - a oer)0.61
- 45%1.00
- M-4B0.99
- D-2B1.00
- 421.00
- M-2B0.96
- 561.00
- M-1B1.00
- 571.00
- M-4B1.00
- 601.00
- 40%1.00
- IV3.5-4B1.00
- 851.00
- P-Isaac-2B1.00
- 901.00
- D-1B1.00
- Q3VL-2B1.00
- Q3VL-4B1.00
- 1021.00
- 1091.00
- 35%1.00
- better0.93
- D-2B0.99
- 35x FLOPs per0.97
- IV3.5-2B0.99
- Q3.5-2B1.00
- 1291.00
- 1691.00
- 30%1.00
- D-4B1.00
- Q3.5-2B0.95
- reduction vs1.00
- Qwen3.5-4B1.00
- correct1.00
- Q3.5-4B1.00
- Q3.5-4B1.00
- 1,2840.96
- cheaper1.00
- 251.00
- 501.00
- 751.00
- 1001.00
- 1251.00
- 1501.00
- 1751.00
- 1,2840.96
- 1010.86
- 10120.96
- 10130.96
- Mean output tokens per response (lower is better)0.99
- FLOPs per correct answer (log, lower is better)1.00
- Family1.00
- Model1.00
- MAmmoTH-VL1.00
- DatologyAl0.97
- InternVL3InternVL3.50.99
- Qwen3-VLQwen3.5Perceptron1.00
- Size1.00
- ● 1B0.85
- 2B1.00
- 4B0.99
- World's Fair0.98
- TRACK 9·JUNE 30,20260.97
- Data Quality1.00
Transcript
199 cues· 3,822 words· 21,171 chars
- 0:13 Good morning, everybody.
- 0:15 My name's Ari Marcos.
- 0:16 I'm the CEO and co-founder of Datology AI and really excited to kick off the data quality track today.
- 0:23 Data quality is what we live and breathe at Datology.
- 0:26 It's all we think about.
- 0:27 In fact, the company's name literally means the science or study of data.
- 0:31 So very excited to see the increasing excitement and interest in this area and an amazing lineup of talks today.
- 0:38 So today I'm gonna tell you about why data quality is a compute multiplier that we're all overlooking and where we can make massive gains by just working on better data.
- 0:48 We've seen that compute availability over the last six months has become extremely scarce and is only getting worse.
- 0:55 We saw H100 prices reverse, their several year long drop, which is normal for hardware, and all of a sudden come up where now they're about 40% up from their lows at the end of last year.
- 1:06 As test time compute has become a critical part of models, and as we have put more and more thinking tokens in, we're seeing that the number, the token usage is absolutely skyrocketing.
- 1:16 Reasoning models use eight times as many tokens as non-reasoning models, and that's projected to 5x again in the next year or so.
- 1:23 So the number of tokens we're pushing through goes higher and higher, and that constrains compute even further.
- 1:28 And this has led actually to a world where it's not implausible today that we might see access to some of the frontier APIs actually get limited or go away.
- 1:36 As one example, Google just capped Meta's Gemini usage because of inference constraints.
- 1:41 OpenAI has effectively started selling token futures, or you can guarantee token capacity some amount of time into the future.
- 1:47 This is only necessary because they are legitimately wondering there might be a world where access to frontier API tokens is limited, not as a business decision, but because there's just simply not enough inference.
- 1:58 first-party products will be prioritized.
- 2:00 So in a world where compute is increasingly scarce, and you need to make models better, well, what do you do?
- 2:05 Well, we work on data.
- 2:06 You're gonna hear me say this over and over again, data quality is a compute multiplier, because what it does is it makes the learning curve steeper.
- 2:13 So this is just a very simple schematic of performance on the y-axis as a function of data on the x-axis.
- 2:18 Note that this x-axis, you could swap it out for data, for compute, for time, for dollars, they're all the same x-axis fundamentally.
- 2:26 And if you can make data quality better, you can turn this gray curve into this blue curve.
- 2:32 And that means that now you can get dramatically better performance for the same compute budget as if you had trained with far more compute.
- 2:40 And similarly, you can get the same performance for a much smaller compute budget, which is exactly showing you how you can get performance as if you had spent 10 times as much on compute and really shows this compute multiplier point.
- 2:55 So how do you actually do this?
- 2:56 Well, fundamentally the idea is we wanna make it so that we get the maximum signal per token and per batch.
- 3:02 For those that are a little bit more technically, what we wanna do here is maximize the marginal information gain per data point when we show it to the model.
- 3:09 What data is gonna teach the model the most?
- 3:11 And that's all about finding data that's relevant to the use cases that you want.
- 3:16 One thing that is very important is that there's no one golden data set to rule them all that's good for everything no matter what you wanna do.
- 3:23 A data set's only gonna be optimal with respect to a particular set of output tasks that you want the model to do.
- 3:28 So if you want a great legal model, you're gonna want legal data more than healthcare data and vice versa.
- 3:33 It needs to be diverse.
- 3:34 A lot of the issues we see with model robustness and brittleness comes from training on data that's not diverse enough.
- 3:40 So the model can answer a question correctly if it's presented just so, but if it's presented a little bit differently, now everything breaks.
- 3:47 It needs to be information dense, and you have to mix the data correctly.
- 3:49 This is a hugely difficult part.
- 3:51 You now have many different sources.
- 3:52 How do you combine them to actually drive the largest improvement in performance?
- 3:56 So this is a high level of what we do at Datology here.
- 4:00 You can think of us as the oil refinery for data.
- 4:02 We don't source new tokens like many data providers.
- 4:04 Rather, we take existing tokens coming from public data sets, proprietary data sets, and licensed data sets, and make them way better.
- 4:11 And how do we do that?
- 4:11 Well, we do that through these four Cs, clean, curate, create, and compose.
- 4:16 So cleaning is fairly straightforward.
- 4:17 This is doing things like heuristic filters all in gopher and things like that.
- 4:21 Removing documents that have only 10 characters in them or are all wing ding.
loading
Chapters
- 0:00 Data is all we think about
- 0:52 Why good data became scarce
- 2:19 Swapping compute for data on the curve
- 3:48 An oil refinery for data
- 5:52 Curation work at DatologyAI
- 6:54 Proof: small models beating bigger ones
- 8:58 Inference efficiency from better data
- 9:24 Multilingual gains
- 12:16 Synthetic data done right
- 14:10 Thomson Reuters and Arcee results
- 17:43 Cheaper than buying compute