Videos KhYifX22yhE
The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside
Scene timeline
42 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 190
- whisperx 190
- chunks
- 30
- from 190 cues
- keyframes
- 21
- kept of 42 captured
- frames with text
- 21
- 506 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 5.6 MB
- word timings on 190 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:31 | 0s |
stt |
done | — | 2026-08-09 06:31 | 19s |
chunk |
done | — | 2026-08-09 06:32 | 0s |
text_embed |
done | — | 2026-08-10 19:39 | 0s |
keyframe |
done | — | 2026-08-09 06:32 | 2m 13s |
ocr |
done | — | 2026-08-09 06:34 | 10s |
frame_embed |
done | — | 2026-08-10 19:39 | 4s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.98
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- World's Fair0.98
-
- AlEngineer0.96
- World'sFair1.00
- AI ENGINEER260.98
- PRESENTED BY1.00
- The Messy Reality of Scale:0.99
- Microsoft1.00
- Synthetic Data and Pre-Training at1.00
- Poolside1.00
- MARAH ABDIN & ROBERT MCHARDY1.00
- COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.99
- PRODUCTION1.00
- v1.01.021.00
- WorldsFalr0.83
- Engineering the future of Al1.00
-
- AlEngineer0.96
- World'sFair1.00
- poolside'sLaguna models0.99
- Laguna is poolside's first public (& open-weight!) family of models, trained from1.00
- scratch:1.00
- PRESENTED BY0.98
- Laguna M.1: 225B total/23B active parameters, pre-trained on 30T tokens1.00
- Microsoft1.00
- Laguna XS.2: 33B total/3B active parameters, pre-trained on 30T tokens1.00
- We released both models together, but XS.2 was built based on learnings on1.00
- data, architecture and training from M.1.0.99
- TECH REPORT1.00
- World'sFair0.98
- Engineering the future of Al0.99
-
- AlEngineer0.96
- World'sFair0.98
- Synthetic Data at1.00
- Scale1.00
- © COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.98
- PRODUCTION1.00
- v1.01.021.00
- World'sFai0.98
- TRACK 9·JUNE 30,20260.98
- Data Quality0.97
-
- AlEngineer0.99
- World's Fair0.96
- poolside's answer: three moves1.00
- focus1.00
- 11.00
- 21.00
- 31.00
- High-recall web data1.00
- Synthetic data at scale1.00
- AutoMixer1.00
- preserve diversity,1.00
- not just precision1.00
- today's focus1.00
- data mixtures1.00
- automated1.00
- Three moves fixed the M.1 bottlenecks0.99
- World'sFai1.00
- TRACK 9· JUNE 30, 20260.97
- Data Quality1.00
-
- AlEngineer0.96
- World'sFair1.00
- Whysynthetic data1.00
- It complements organic data; it doesn't replace it.1.00
- Three jobs: regularize how information is presented or how it's taught, fill0.99
- under-represented formats, and expose structure (plans, rationales, Q&A1.00
- surfaces).1.00
- In Laguna XS.2: ~13% of the mix, across pre-training stages, from a ~6T+1.00
- synthetic token corpus across domains.0.99
- ~13%0.98
- ~6T+0.96
- of the XS.2 mix1.00
- generated tokens1.00
- TRACK 9· JUNE 30,20260.96
- Data Quality0.97
-
- AlEngineer0.98
- World's Fair0.97
- Scale breaks the old data playbook0.99
- The old playbook: filter1.00
- Synthetic rewrites vs. seeds-only across benchmarks0.99
- hard, maximize average1.00
- MMLU-Pro1.00
- ARC-Challenge0.97
- TriviaQA1.00
- quality. At scale, it0.98
- 0.181.00
- 0.161.00
- 0.201.00
- Amrvn0.58
- 0.551.00
- 0.50-0.98
- 0.4-0.89
- 0.51.00
- backfires.1.00
- 0.140.91
- 0.121.00
- 0.451.00
- 0.400.98
- 0.31.00
- hit two bottlenecks:1.00
- 0.101.00
- 0.080.99
- 0.060.98
- 50k 100k 150k 200k 250k 300k 350k 400k0.99
- 0.351.00
- 0.301.00
- 0.25-0.92
- 50k 100k 150k 200k 250k 300k 350k 400k0.98
- 0.21.00
- 0.11.00
- 50k 100k 150k 200k 250k 300k 350k 400k0.95
- repetition in high-value1.00
- GSM8K1.00
- MATH1.00
- EvalPlus pass@10.98
- 0.300.97
- solld: verifed0.97
- dashed: EM0.93
- 0.1751.00
- 0.401.00
- MMm0.51
- subsets, and budget0.99
- allocation across source1.00
- 0.200.95
- 0.150.64
- 0.251.00
- 0.10-0.90
- 0.051.00
- AMMAMMMA0.64
- 0.1251.00
- 0.1001.00
- 0.150-0.96
- 0.075-0.94
- 0.050-0.92
- 0.351.00
- 0.301.00
- 0.251.00
- 0.20-0.95
- 0.15-0.92
- 0.101.00
- New: Regularize tokens;0.99
- lexical rephrasing trio0.99
- 0.001.00
- 50k 100k 150k0.99
- 200k1.00
- step1.00
- 250k 300k 350k 400k0.97
- 0.0251.00
- — seeds only0.95
- 50k 100k 150k 200k 250k 300k 350k 400k0.99
- —seeds + rewrites0.98
- step1.00
- 0.051.00
- 50k 100k 150k0.96
- 200k 250k 300k 350k 400k0.99
- step1.00
- World's Fair0.98
- TRACK 9· JUNE 30, 20260.94
- Data Quality1.00
-
- AlEngineer0.99
- World'sFair1.00
- How: a modularpipeline0.98
- Two principles: match pipeline complexity to teacher capability, and hand the0.99
- generator inputs that narrow the task.1.00
- PRESENTED BY0.99
- Each pipeline composes from six reusable components1.00
- Microsoft1.00
- S0.99
- M1.00
- G1.00
- V1.00
- pre·post0.99
- inputs1.00
- metadata1.00
- generator1.00
- filter0.99
- validator1.00
- wrappers1.00
- P = postfnGn..f1G1°pre0.88
- Grounding1.00
- balances these1.00
- every pipeline0.97
- Entropy1.00
- faithful to the seed0.98
- diversity, variation0.98
- TRACK 9·JUNE 30,20260.98
- Data Quality0.97
-
- AlEngineer0.99
- World's Fair0.97
- The strategies that fill the mix0.98
- Four shapes, by order-of-magnitude token contribution:1.00
- Shape1.00
- Composition· what it does / examples0.96
- Tokens1.00
- P= post fqual Grewite o pre0.91
- Form-rewrite1.00
- Single-call pass conditioned on metadata. Multi-mode rephrasing of web0.99
- ~10120.99
- & STEM docs; code rephrased into code with natural language.1.00
- P= post o fleep o Go ...o(fv≥r o VoGk) o...o G1 o pre0.77
- Multi-stage1.00
- cascade1.00
- Each stage transforms prior output or emits metadata; a V-gate prunes0.99
- intermediates. Textbook synthesis, grounded QA, diff-conditioned1.00
- ~10110.98
- coding tasks.1.00
- P= post o fmt Gconvet preseed0.92
- Cross-domain1.00
- ~10100.99
- transducer1.00
- Casts a seed across modalities; pre-step enforces seed suitability.1.00
- math ↔ code, code-language porting.0.99
- P= ftmd loopT(…..) Ginit o frel0.86
- rollout1.00
- Multi-turn1.00
- Closed-loop interaction; orchestrator decides flow and termination.0.99
- ~10100.97
- Stacktrace-grounded chats, multi-turn math, iterative doc evolution.1.00
- Synthetic data pipeline strategies. Token counts are order-of-magnitude contribution to the pre-training corpus.1.00
- World's Fair0.96
- TRACK 9· JUNE 30, 20260.93
- Data Quality1.00
-
- AlEngineer0.98
- World'sFair1.00
- The payoff: data that scales0.98
- We broke the repetition ceiling and grew usable training signal, without losing1.00
- diversity.1.00
- Now: how do you train at scale?1.00
- DATA1.00
- PRE-TRAINING1.00
- repetition broken1.00
- now that the data scales,0.96
- usable signal grown1.00
- how do we train at scale?0.99
- diversity kept1.00
- Two halves of scale0.99
- World'sFair1.00
- TRACK 9• JUNE 30, 20260.93
- Data Quality1.00
-
- AlEngineer0.97
- World'sFair0.97
- Pre-training at scale0.99
- COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.99
- PRODUCTION1.00
- v1.01.021.00
- World'sFair1.00
- TRACK 9·JUNE 30,20260.98
- Data Quality0.97
Transcript
190 cues· 3,171 words· 17,352 chars
- 0:12 Hi, everybody.
- 0:13 Thanks for coming to our talk.
- 0:16 My name is Mara Abdin.
- 0:19 I'm from Pulsar.
- 0:20 I'm the synthetic lead for our data team.
- 0:23 And today, my colleague Robert and I will be talking a little bit about some of the challenges that we've seen as we scale our models over here.
- 0:33 Particularly, if you haven't heard, we've switched recently from releasing our models towards enterprise to also releasing towards
- 0:42 everybody.
- 0:43 We've actually put out two open weight models that are a few weeks ago.
- 0:50 We also put out a tech report which has a ton of detail if you're interested.
- 0:54 As you can see here by the Laguna M.1 and XS.2, this is actually because we have
- 1:03 switched out quite a few things between those two models.
- 1:06 And so big flavor of this talk is going to be about kind of how to transition from one to two.
- 1:12 And in fact, we've continued to do so.
- 1:14 And now we actually have a newer version and Robert will give a sneak peek about soon to be released model.
- 1:21 Okay.
- 1:22 So I will be particular talking about synthetic data part of things.
- 1:25 So there's three things that we did on the data side to kind of resolve some of the issues that we've seen with scale.
- 1:33 One is that we implemented an automixer that basically just gives us the chance to do a cheaper sweep on clusters of our sets before moving on to more expensive experiments.
- 1:43 And then we improved, we just rethought our sampling of web data for high recall.
- 1:50 And then the third one is that we relied a lot more on synthetic data in a few forms, which I will go into.
- 1:56 Okay, so before we're kind of going into what does that mean and what have we done, et cetera, why would we kind of, sometimes it's fair at least to ask why synthetic data, and the thing is that
- 2:09 at least at Pulsar we don't see it as a way to replace organic data.
- 2:12 I don't see it so in the current state of the world at least, but it is a way to kind of compliment it.
- 2:18 And the thing is that organic data has a lot in it that is basically kind of implicitly hidden.
- 2:23 A lot of things that could teach the model are not very presented in the most optimal way sometimes.
- 2:27 And so synthetic data gives us a track to extract
- 2:31 some of these features and then project them on some new planes.
- 2:34 And this is how we get to expose implicit rationale, implicit planning, implicit structure, and a way for us to fill gaps and regularize not only how we present the tokens, but also how we are teaching the model.
- 2:47 For XS.2 in particular, we settled on 13% of the mix.
- 2:50 This is only pre-training stages before post-training.
- 2:53 And since then, we've just been continuously generating more data in a bunch of directions.
- 2:57 Now we have a six trillion token corpus that's continuously growing.
- 3:03 So, yeah, so kind of what I just said is that we saw some, I guess, limitations switching from Logonim to 0.1 or 0.2 models.
- 3:16 And so one of those things is that we basically started on data, this is really not a,
- 3:22 crazy kind of problem.
- 3:23 We intuitively started from a place on a smaller scale where we were basically focusing on quality versus quantity, maybe a little too much, because eventually when we started scaling our models, we had to scale our training budget, and with that
- 3:40 came some limitations because we started hitting repetition, like non-optimal repetition on some of our high-quality data, which saturated the model a little too early.
- 3:49 So one of the ways that we've, particularly for this token uniqueness problem, we relied on, which is a very common form of synthetic data, rephrasing, which you've just heard in Beyond Web, for example.
- 4:05 It's become pretty trendy these days.
- 4:07 And you can see here that, you know, this is an ablation result, so take the numbers with a grain of salt, but what presents for you consistently is the diff between using the orange would be just the seeds with repetition, and then the green would be replacing some of those repeated tokens with.
- 4:26 higher, like with multimode rewrites, or at least, yeah, all of them, or at least reducing the repetition.
- 4:32 And so for rephrasing in particular, we did to the, you know, what everyone's doing with the, you know, generic kind of multimode, very scalable pipeline, where we also took it a step far.
- 4:42 And we did two other specialized pipelines, one to go from raw code to code and text, and one to go specifically for STEM data,
- 4:50 just because this is a very cheap, scalable pipeline.
- 4:53 So it kind of, you kind of have to rely very heavily on the seed and we push a little further on that for the STEM documents.
- 5:00 Okay.
- 5:01 So, okay.
- 5:03 So if you kind of think of everything as a kind of a modular way, you can think of every synthetic data pipeline is composed of the same six components.
- 5:14 And so you have your seeds, your primary inputs, your metadata, your secondary inputs, your generator function,
loading
Chapters
- 0:00 Introduction: synthetic data and pre-training at poolside
- 1:52 Why synthetic data
- 3:11 Limitations and the training budget
- 4:44 Inside the synthetic data pipeline
- 6:37 Multistage pipelines and porting data
- 7:43 Multi turn chats and policing generations
- 9:03 Pre-training: trust nothing, crash on mismatch
- 10:41 Failures at scale: broken GPUs
- 11:56 Numerical precision and corrupted gradients
- 13:15 A 118B model for agentic coding
- 15:07 Early results vs GLM 4.5 Air