read-only demo

Videos KhYifX22yhE

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

index_state ready data_status ok

AI Engineer· published 2026-07-26· 0:17:31· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:15, 1 of 1 keyframes kept
  5. Shot 4, 0:15 to 0:34, 1 of 1 keyframes kept
  6. Shot 5, 0:34 to 1:20, 1 of 1 keyframes kept
  7. Shot 6, 1:20 to 1:26, 1 of 1 keyframes kept
  8. Shot 7, 1:26 to 1:55, 1 of 1 keyframes kept
  9. Shot 8, 1:55 to 2:28, 1 of 1 keyframes kept
  10. Shot 9, 2:28 to 3:01, 0 of 1 keyframes kept
  11. Shot 10, 3:01 to 3:30, 1 of 1 keyframes kept
  12. Shot 11, 3:30 to 4:00, 0 of 1 keyframes kept
  13. Shot 12, 4:00 to 4:30, 0 of 1 keyframes kept
  14. Shot 13, 4:30 to 4:59, 0 of 1 keyframes kept
  15. Shot 14, 4:59 to 5:27, 1 of 1 keyframes kept
  16. Shot 15, 5:27 to 5:55, 0 of 1 keyframes kept
  17. Shot 16, 5:55 to 6:23, 0 of 1 keyframes kept
  18. Shot 17, 6:23 to 6:54, 1 of 1 keyframes kept
  19. Shot 18, 6:54 to 7:24, 0 of 1 keyframes kept
  20. Shot 19, 7:24 to 7:54, 0 of 1 keyframes kept
  21. Shot 20, 7:54 to 8:24, 0 of 1 keyframes kept
  22. Shot 21, 8:24 to 8:54, 0 of 1 keyframes kept
  23. Shot 22, 8:54 to 9:05, 1 of 1 keyframes kept
  24. Shot 23, 9:05 to 9:32, 1 of 1 keyframes kept
  25. Shot 24, 9:32 to 10:04, 1 of 1 keyframes kept
  26. Shot 25, 10:04 to 10:37, 0 of 1 keyframes kept
  27. Shot 26, 10:37 to 11:02, 1 of 1 keyframes kept
  28. Shot 27, 11:02 to 11:27, 0 of 1 keyframes kept
  29. Shot 28, 11:27 to 11:52, 1 of 1 keyframes kept
  30. Shot 29, 11:52 to 12:17, 0 of 1 keyframes kept
  31. Shot 30, 12:17 to 12:43, 0 of 1 keyframes kept
  32. Shot 31, 12:43 to 13:19, 0 of 1 keyframes kept
  33. Shot 32, 13:19 to 13:58, 0 of 1 keyframes kept
  34. Shot 33, 13:58 to 14:31, 1 of 1 keyframes kept
  35. Shot 34, 14:31 to 15:03, 0 of 1 keyframes kept
  36. Shot 35, 15:03 to 15:33, 1 of 1 keyframes kept
  37. Shot 36, 15:33 to 16:02, 0 of 1 keyframes kept
  38. Shot 37, 16:02 to 16:32, 0 of 1 keyframes kept
  39. Shot 38, 16:32 to 17:02, 0 of 1 keyframes kept
  40. Shot 39, 17:02 to 17:05, 1 of 1 keyframes kept
  41. Shot 40, 17:05 to 17:14, 1 of 1 keyframes kept
  42. Shot 41, 17:14 to 17:30, 0 of 1 keyframes kept

42 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
190
whisperx 190
chunks
30
from 190 cues
keyframes
21
kept of 42 captured
frames with text
21
506 lines read
chapters
11
from the source metadata
keyframe bytes
5.6 MB
word timings on 190 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:31 0s
stt done 2026-08-09 06:31 19s
chunk done 2026-08-09 06:32 0s
text_embed done 2026-08-10 19:39 0s
keyframe done 2026-08-09 06:32 2m 13s
ocr done 2026-08-09 06:34 10s
frame_embed done 2026-08-10 19:39 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.7

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2743.2

    1. LAB & PLATINUM SPONSORS0.98
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:14 #3 done1 line(s)

    shot 3·sharpness 465.4

    1. World's Fair0.98
  • 0:28 #4 done14 line(s)

    shot 4·sharpness 2440.0

    1. AlEngineer0.96
    2. World'sFair1.00
    3. AI ENGINEER260.98
    4. PRESENTED BY1.00
    5. The Messy Reality of Scale:0.99
    6. Microsoft1.00
    7. Synthetic Data and Pre-Training at1.00
    8. Poolside1.00
    9. MARAH ABDIN & ROBERT MCHARDY1.00
    10. COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.99
    11. PRODUCTION1.00
    12. v1.01.021.00
    13. WorldsFalr0.83
    14. Engineering the future of Al1.00
  • 0:39 #5 done14 line(s)

    shot 5·sharpness 4076.3

    1. AlEngineer0.96
    2. World'sFair1.00
    3. poolside'sLaguna models0.99
    4. Laguna is poolside's first public (& open-weight!) family of models, trained from1.00
    5. scratch:1.00
    6. PRESENTED BY0.98
    7. Laguna M.1: 225B total/23B active parameters, pre-trained on 30T tokens1.00
    8. Microsoft1.00
    9. Laguna XS.2: 33B total/3B active parameters, pre-trained on 30T tokens1.00
    10. We released both models together, but XS.2 was built based on learnings on1.00
    11. data, architecture and training from M.1.0.99
    12. TECH REPORT1.00
    13. World'sFair0.98
    14. Engineering the future of Al0.99
  • 1:25 #6 done10 line(s)

    shot 6·sharpness 1180.8

    1. AlEngineer0.96
    2. World'sFair0.98
    3. Synthetic Data at1.00
    4. Scale1.00
    5. © COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.98
    6. PRODUCTION1.00
    7. v1.01.021.00
    8. World'sFai0.98
    9. TRACK 9·JUNE 30,20260.98
    10. Data Quality0.97
  • 1:46 #7 done19 line(s)

    shot 7·sharpness 2055.3

    1. AlEngineer0.99
    2. World's Fair0.96
    3. poolside's answer: three moves1.00
    4. focus1.00
    5. 11.00
    6. 21.00
    7. 31.00
    8. High-recall web data1.00
    9. Synthetic data at scale1.00
    10. AutoMixer1.00
    11. preserve diversity,1.00
    12. not just precision1.00
    13. today's focus1.00
    14. data mixtures1.00
    15. automated1.00
    16. Three moves fixed the M.1 bottlenecks0.99
    17. World'sFai1.00
    18. TRACK 9· JUNE 30, 20260.97
    19. Data Quality1.00
  • 2:21 #8 done15 line(s)

    shot 8·sharpness 3244.0

    1. AlEngineer0.96
    2. World'sFair1.00
    3. Whysynthetic data1.00
    4. It complements organic data; it doesn't replace it.1.00
    5. Three jobs: regularize how information is presented or how it's taught, fill0.99
    6. under-represented formats, and expose structure (plans, rationales, Q&A1.00
    7. surfaces).1.00
    8. In Laguna XS.2: ~13% of the mix, across pre-training stages, from a ~6T+1.00
    9. synthetic token corpus across domains.0.99
    10. ~13%0.98
    11. ~6T+0.96
    12. of the XS.2 mix1.00
    13. generated tokens1.00
    14. TRACK 9· JUNE 30,20260.96
    15. Data Quality0.97
  • 2:57 #9 skipped

    shot 9·duplicate of #8

  • 3:10 #10 done84 line(s)

    shot 10·sharpness 3238.9

    1. AlEngineer0.98
    2. World's Fair0.97
    3. Scale breaks the old data playbook0.99
    4. The old playbook: filter1.00
    5. Synthetic rewrites vs. seeds-only across benchmarks0.99
    6. hard, maximize average1.00
    7. MMLU-Pro1.00
    8. ARC-Challenge0.97
    9. TriviaQA1.00
    10. quality. At scale, it0.98
    11. 0.181.00
    12. 0.161.00
    13. 0.201.00
    14. Amrvn0.58
    15. 0.551.00
    16. 0.50-0.98
    17. 0.4-0.89
    18. 0.51.00
    19. backfires.1.00
    20. 0.140.91
    21. 0.121.00
    22. 0.451.00
    23. 0.400.98
    24. 0.31.00
    25. hit two bottlenecks:1.00
    26. 0.101.00
    27. 0.080.99
    28. 0.060.98
    29. 50k 100k 150k 200k 250k 300k 350k 400k0.99
    30. 0.351.00
    31. 0.301.00
    32. 0.25-0.92
    33. 50k 100k 150k 200k 250k 300k 350k 400k0.98
    34. 0.21.00
    35. 0.11.00
    36. 50k 100k 150k 200k 250k 300k 350k 400k0.95
    37. repetition in high-value1.00
    38. GSM8K1.00
    39. MATH1.00
    40. EvalPlus pass@10.98
    41. 0.300.97
    42. solld: verifed0.97
    43. dashed: EM0.93
    44. 0.1751.00
    45. 0.401.00
    46. MMm0.51
    47. subsets, and budget0.99
    48. allocation across source1.00
    49. 0.200.95
    50. 0.150.64
    51. 0.251.00
    52. 0.10-0.90
    53. 0.051.00
    54. AMMAMMMA0.64
    55. 0.1251.00
    56. 0.1001.00
    57. 0.150-0.96
    58. 0.075-0.94
    59. 0.050-0.92
    60. 0.351.00
    61. 0.301.00
    62. 0.251.00
    63. 0.20-0.95
    64. 0.15-0.92
    65. 0.101.00
    66. New: Regularize tokens;0.99
    67. lexical rephrasing trio0.99
    68. 0.001.00
    69. 50k 100k 150k0.99
    70. 200k1.00
    71. step1.00
    72. 250k 300k 350k 400k0.97
    73. 0.0251.00
    74. — seeds only0.95
    75. 50k 100k 150k 200k 250k 300k 350k 400k0.99
    76. —seeds + rewrites0.98
    77. step1.00
    78. 0.051.00
    79. 50k 100k 150k0.96
    80. 200k 250k 300k 350k 400k0.99
    81. step1.00
    82. World's Fair0.98
    83. TRACK 9· JUNE 30, 20260.94
    84. Data Quality1.00
  • 3:34 #11 skipped

    shot 11·duplicate of #10

  • 4:06 #12 skipped

    shot 12·duplicate of #10

  • 4:45 #13 skipped

    shot 13·duplicate of #10

  • 5:19 #14 done28 line(s)

    shot 14·sharpness 2712.9

    1. AlEngineer0.99
    2. World'sFair1.00
    3. How: a modularpipeline0.98
    4. Two principles: match pipeline complexity to teacher capability, and hand the0.99
    5. generator inputs that narrow the task.1.00
    6. PRESENTED BY0.99
    7. Each pipeline composes from six reusable components1.00
    8. Microsoft1.00
    9. S0.99
    10. M1.00
    11. G1.00
    12. V1.00
    13. pre·post0.99
    14. inputs1.00
    15. metadata1.00
    16. generator1.00
    17. filter0.99
    18. validator1.00
    19. wrappers1.00
    20. P = postfnGn..f1G1°pre0.88
    21. Grounding1.00
    22. balances these1.00
    23. every pipeline0.97
    24. Entropy1.00
    25. faithful to the seed0.98
    26. diversity, variation0.98
    27. TRACK 9·JUNE 30,20260.98
    28. Data Quality0.97
  • 5:52 #15 skipped

    shot 15·duplicate of #14

  • 6:07 #16 skipped

    shot 16·duplicate of #14

  • 6:30 #17 done35 line(s)

    shot 17·sharpness 2863.7

    1. AlEngineer0.99
    2. World's Fair0.97
    3. The strategies that fill the mix0.98
    4. Four shapes, by order-of-magnitude token contribution:1.00
    5. Shape1.00
    6. Composition· what it does / examples0.96
    7. Tokens1.00
    8. P= post fqual Grewite o pre0.91
    9. Form-rewrite1.00
    10. Single-call pass conditioned on metadata. Multi-mode rephrasing of web0.99
    11. ~10120.99
    12. & STEM docs; code rephrased into code with natural language.1.00
    13. P= post o fleep o Go ...o(fv≥r o VoGk) o...o G1 o pre0.77
    14. Multi-stage1.00
    15. cascade1.00
    16. Each stage transforms prior output or emits metadata; a V-gate prunes0.99
    17. intermediates. Textbook synthesis, grounded QA, diff-conditioned1.00
    18. ~10110.98
    19. coding tasks.1.00
    20. P= post o fmt Gconvet preseed0.92
    21. Cross-domain1.00
    22. ~10100.99
    23. transducer1.00
    24. Casts a seed across modalities; pre-step enforces seed suitability.1.00
    25. math ↔ code, code-language porting.0.99
    26. P= ftmd loopT(…..) Ginit o frel0.86
    27. rollout1.00
    28. Multi-turn1.00
    29. Closed-loop interaction; orchestrator decides flow and termination.0.99
    30. ~10100.97
    31. Stacktrace-grounded chats, multi-turn math, iterative doc evolution.1.00
    32. Synthetic data pipeline strategies. Token counts are order-of-magnitude contribution to the pre-training corpus.1.00
    33. World's Fair0.96
    34. TRACK 9· JUNE 30, 20260.93
    35. Data Quality1.00
  • 7:06 #18 skipped

    shot 18·duplicate of #17

  • 7:36 #19 skipped

    shot 19·duplicate of #17

  • 8:12 #20 skipped

    shot 20·duplicate of #8

  • 8:48 #21 skipped

    shot 21·duplicate of #8

  • 9:00 #22 done17 line(s)

    shot 22·sharpness 2460.3

    1. AlEngineer0.98
    2. World'sFair1.00
    3. The payoff: data that scales0.98
    4. We broke the repetition ceiling and grew usable training signal, without losing1.00
    5. diversity.1.00
    6. Now: how do you train at scale?1.00
    7. DATA1.00
    8. PRE-TRAINING1.00
    9. repetition broken1.00
    10. now that the data scales,0.96
    11. usable signal grown1.00
    12. how do we train at scale?0.99
    13. diversity kept1.00
    14. Two halves of scale0.99
    15. World'sFair1.00
    16. TRACK 9• JUNE 30, 20260.93
    17. Data Quality1.00
  • 9:23 #23 done9 line(s)

    shot 23·sharpness 1091.3

    1. AlEngineer0.97
    2. World'sFair0.97
    3. Pre-training at scale0.99
    4. COPYRIGHT 2026, POOLSIDE INC. ALL RIGHTS RESERVED. MADE GLOBALLY.0.99
    5. PRODUCTION1.00
    6. v1.01.021.00
    7. World'sFair1.00
    8. TRACK 9·JUNE 30,20260.98
    9. Data Quality0.97

Transcript

190 cues· 3,171 words· 17,352 chars

  1. 0:12 Hi, everybody.
  2. 0:13 Thanks for coming to our talk.
  3. 0:16 My name is Mara Abdin.
  4. 0:19 I'm from Pulsar.
  5. 0:20 I'm the synthetic lead for our data team.
  6. 0:23 And today, my colleague Robert and I will be talking a little bit about some of the challenges that we've seen as we scale our models over here.
  7. 0:33 Particularly, if you haven't heard, we've switched recently from releasing our models towards enterprise to also releasing towards
  8. 0:42 everybody.
  9. 0:43 We've actually put out two open weight models that are a few weeks ago.
  10. 0:50 We also put out a tech report which has a ton of detail if you're interested.
  11. 0:54 As you can see here by the Laguna M.1 and XS.2, this is actually because we have
  12. 1:03 switched out quite a few things between those two models.
  13. 1:06 And so big flavor of this talk is going to be about kind of how to transition from one to two.
  14. 1:12 And in fact, we've continued to do so.
  15. 1:14 And now we actually have a newer version and Robert will give a sneak peek about soon to be released model.
  16. 1:21 Okay.
  17. 1:22 So I will be particular talking about synthetic data part of things.
  18. 1:25 So there's three things that we did on the data side to kind of resolve some of the issues that we've seen with scale.
  19. 1:33 One is that we implemented an automixer that basically just gives us the chance to do a cheaper sweep on clusters of our sets before moving on to more expensive experiments.
  20. 1:43 And then we improved, we just rethought our sampling of web data for high recall.
  21. 1:50 And then the third one is that we relied a lot more on synthetic data in a few forms, which I will go into.
  22. 1:56 Okay, so before we're kind of going into what does that mean and what have we done, et cetera, why would we kind of, sometimes it's fair at least to ask why synthetic data, and the thing is that
  23. 2:09 at least at Pulsar we don't see it as a way to replace organic data.
  24. 2:12 I don't see it so in the current state of the world at least, but it is a way to kind of compliment it.
  25. 2:18 And the thing is that organic data has a lot in it that is basically kind of implicitly hidden.
  26. 2:23 A lot of things that could teach the model are not very presented in the most optimal way sometimes.
  27. 2:27 And so synthetic data gives us a track to extract
  28. 2:31 some of these features and then project them on some new planes.
  29. 2:34 And this is how we get to expose implicit rationale, implicit planning, implicit structure, and a way for us to fill gaps and regularize not only how we present the tokens, but also how we are teaching the model.
  30. 2:47 For XS.2 in particular, we settled on 13% of the mix.
  31. 2:50 This is only pre-training stages before post-training.
  32. 2:53 And since then, we've just been continuously generating more data in a bunch of directions.
  33. 2:57 Now we have a six trillion token corpus that's continuously growing.
  34. 3:03 So, yeah, so kind of what I just said is that we saw some, I guess, limitations switching from Logonim to 0.1 or 0.2 models.
  35. 3:16 And so one of those things is that we basically started on data, this is really not a,
  36. 3:22 crazy kind of problem.
  37. 3:23 We intuitively started from a place on a smaller scale where we were basically focusing on quality versus quantity, maybe a little too much, because eventually when we started scaling our models, we had to scale our training budget, and with that
  38. 3:40 came some limitations because we started hitting repetition, like non-optimal repetition on some of our high-quality data, which saturated the model a little too early.
  39. 3:49 So one of the ways that we've, particularly for this token uniqueness problem, we relied on, which is a very common form of synthetic data, rephrasing, which you've just heard in Beyond Web, for example.
  40. 4:05 It's become pretty trendy these days.
  41. 4:07 And you can see here that, you know, this is an ablation result, so take the numbers with a grain of salt, but what presents for you consistently is the diff between using the orange would be just the seeds with repetition, and then the green would be replacing some of those repeated tokens with.
  42. 4:26 higher, like with multimode rewrites, or at least, yeah, all of them, or at least reducing the repetition.
  43. 4:32 And so for rephrasing in particular, we did to the, you know, what everyone's doing with the, you know, generic kind of multimode, very scalable pipeline, where we also took it a step far.
  44. 4:42 And we did two other specialized pipelines, one to go from raw code to code and text, and one to go specifically for STEM data,
  45. 4:50 just because this is a very cheap, scalable pipeline.
  46. 4:53 So it kind of, you kind of have to rely very heavily on the seed and we push a little further on that for the STEM documents.
  47. 5:00 Okay.
  48. 5:01 So, okay.
  49. 5:03 So if you kind of think of everything as a kind of a modular way, you can think of every synthetic data pipeline is composed of the same six components.
  50. 5:14 And so you have your seeds, your primary inputs, your metadata, your secondary inputs, your generator function,

Chapters

  1. 0:00 Introduction: synthetic data and pre-training at poolside
  2. 1:52 Why synthetic data
  3. 3:11 Limitations and the training budget
  4. 4:44 Inside the synthetic data pipeline
  5. 6:37 Multistage pipelines and porting data
  6. 7:43 Multi turn chats and policing generations
  7. 9:03 Pre-training: trust nothing, crash on mismatch
  8. 10:41 Failures at scale: broken GPUs
  9. 11:56 Numerical precision and corrupted gradients
  10. 13:15 A 118B model for agentic coding
  11. 15:07 Early results vs GLM 4.5 Air

Open at this second