read-only demo

Videos xbPriQWXtWM

The Base Model Is Dead — Varun Singh, Arcee AI

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:17:44· en-US· indexed 2026-08-10 19:37

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:40, 1 of 1 keyframes kept
  5. Shot 4, 0:40 to 1:19, 1 of 1 keyframes kept
  6. Shot 5, 1:19 to 1:55, 1 of 1 keyframes kept
  7. Shot 6, 1:55 to 2:25, 1 of 1 keyframes kept
  8. Shot 7, 2:25 to 2:55, 0 of 1 keyframes kept
  9. Shot 8, 2:55 to 3:28, 1 of 1 keyframes kept
  10. Shot 9, 3:28 to 4:01, 0 of 1 keyframes kept
  11. Shot 10, 4:01 to 4:29, 1 of 1 keyframes kept
  12. Shot 11, 4:29 to 4:58, 0 of 1 keyframes kept
  13. Shot 12, 4:58 to 5:27, 0 of 1 keyframes kept
  14. Shot 13, 5:27 to 5:52, 1 of 1 keyframes kept
  15. Shot 14, 5:52 to 6:17, 0 of 1 keyframes kept
  16. Shot 15, 6:17 to 6:42, 0 of 1 keyframes kept
  17. Shot 16, 6:42 to 7:07, 0 of 1 keyframes kept
  18. Shot 17, 7:07 to 7:32, 0 of 1 keyframes kept
  19. Shot 18, 7:32 to 7:58, 0 of 1 keyframes kept
  20. Shot 19, 7:58 to 8:23, 0 of 1 keyframes kept
  21. Shot 20, 8:23 to 8:48, 0 of 1 keyframes kept
  22. Shot 21, 8:48 to 9:13, 0 of 1 keyframes kept
  23. Shot 22, 9:13 to 9:39, 1 of 1 keyframes kept
  24. Shot 23, 9:39 to 10:05, 1 of 1 keyframes kept
  25. Shot 24, 10:05 to 10:31, 0 of 1 keyframes kept
  26. Shot 25, 10:31 to 10:56, 0 of 1 keyframes kept
  27. Shot 26, 10:56 to 11:22, 0 of 1 keyframes kept
  28. Shot 27, 11:22 to 11:49, 1 of 1 keyframes kept
  29. Shot 28, 11:49 to 12:15, 0 of 1 keyframes kept
  30. Shot 29, 12:15 to 12:42, 0 of 1 keyframes kept
  31. Shot 30, 12:42 to 13:18, 1 of 1 keyframes kept
  32. Shot 31, 13:18 to 13:46, 1 of 1 keyframes kept
  33. Shot 32, 13:46 to 14:14, 0 of 1 keyframes kept
  34. Shot 33, 14:14 to 14:42, 0 of 1 keyframes kept
  35. Shot 34, 14:42 to 15:13, 1 of 1 keyframes kept
  36. Shot 35, 15:13 to 15:44, 0 of 1 keyframes kept
  37. Shot 36, 15:44 to 16:10, 1 of 1 keyframes kept
  38. Shot 37, 16:10 to 16:36, 0 of 1 keyframes kept
  39. Shot 38, 16:36 to 17:01, 0 of 1 keyframes kept
  40. Shot 39, 17:01 to 17:27, 1 of 1 keyframes kept
  41. Shot 40, 17:27 to 17:44, 0 of 1 keyframes kept

41 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
97
whisperx 97
chunks
29
from 97 cues
keyframes
18
kept of 41 captured
frames with text
18
688 lines read
chapters
11
from the source metadata
keyframe bytes
6.4 MB
word timings on 97 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:36 0s
stt done 2026-08-09 02:35 24s
chunk done 2026-08-09 02:35 0s
text_embed done 2026-08-10 19:37 0s
keyframe done 2026-08-09 02:35 2m 08s
ocr done 2026-08-09 02:37 14s
frame_embed done 2026-08-10 19:37 3s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 455.4

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 661.6

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2736.5

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.97
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. neo4j0.99
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:34 #3 done10 line(s)

    shot 3·sharpness 1969.0

    1. AlEngineer0.98
    2. World's Fair0.97
    3. PRESENTED BY1.00
    4. The Base Model is Dead1.00
    5. Microsoft1.00
    6. (long live the base model)1.00
    7. Varun Singh, Pre-training Lead1.00
    8. Arcee1.00
    9. World's Fair0.92
    10. Engineering the future of Al0.99
  • 1:03 #4 done74 line(s)

    shot 4·sharpness 2888.1

    1. AlEngineer0.98
    2. Pre-training1.00
    3. Mid-training1.00
    4. World'sFair1.00
    5. What is a base0.99
    6. 0.72
    7. model?1.00
    8. Pre-training1.00
    9. General1.00
    10. Pre-training0.91
    11. (500B)0.98
    12. Synthetic1.00
    13. Long Context &1.00
    14. 4K1.00
    15. 4K1.00
    16. 32K1.00
    17. 32K1.00
    18. 128K1.00
    19. rre 3: Pre-training and mid-training stages for GLM-4.5. We adapt a multi-stage training recipe0.99
    20. PRESENTED BY1.00
    21. pre-training1.00
    22. pre-training0.95
    23. phase 20.99
    24. extend the sequence length from 4K to 128K.0.99
    25. Base Model0.99
    26. Post-Training1.00
    27. Microsoft1.00
    28. Pre-training1.00
    29. Base Model1.00
    30. extension1.00
    31. context1.00
    32. pre-training1.00
    33. phase 31.00
    34. Pre-training1.00
    35. General1.00
    36. Corpus1.00
    37. (18T)0.98
    38. Reasoning1.00
    39. Corpus1.00
    40. Code &0.99
    41. (9T)0.98
    42. Overall SFT1.00
    43. logits1.00
    44. 4K1.00
    45. 4K1.00
    46. Reasoning RL1.00
    47. logits1.00
    48. ……>0.71
    49. Mid-training1.00
    50. On-Policy1.00
    51. mid-training0.98
    52. SFT1.00
    53. RL1.00
    54. Reasoning Data0.99
    55. Long Code &1.00
    56. (1T)0.97
    57. Long Context &0.99
    58. (500B/50B)0.99
    59. Agent Data1.00
    60. Agentic RL1.00
    61. General RL1.00
    62. logits1.00
    63. ……>0.73
    64. Cross-Stage1.00
    65. Distillation1.00
    66. Training pipeline for Arcee Trinity Large Thinking.1.00
    67. 32K1.00
    68. 128K/200K0.99
    69. Sparse Attention Adaption (20B, 200K)0.98
    70. GLM-51.00
    71. Figure 5: Overall training pipeline of GLM-5.1.00
    72. World's Fair1.00
    73. TRACK 9· JUNE 30,20260.96
    74. Data Quality0.99
  • 1:34 #5 done51 line(s)

    shot 5·sharpness 2581.2

    1. AlEngineer0.99
    2. World'sFair1.00
    3. Next-token prediction for useful priors1.00
    4. Cross-Entropy Loss1.00
    5. a0.70
    6. outer loop0.95
    7. Prodictios0.79
    8. Output1.00
    9. Probabilities0.98
    10. Learning via SGD during unsupervised pre-training0.99
    11. Softmax1.00
    12. Linear0.99
    13. 5+ 8 -130.77
    14. gaot > goat0.84
    15. RMS Norm1.00
    16. 7+2=90.95
    17. inner loop1.00
    18. 1+8.10.83
    19. brid-- bird0.86
    20. SwiGLU Feed0.98
    21. 3+4=70.84
    22. fsth a> fish0.90
    23. Forward1.00
    24. 5 + 9 = 140.93
    25. otter loutre0.94
    26. decoder1.00
    27. 0.99
    28. RMS Norm1.00
    29. 9 + 8 = 170.94
    30. bread s> pain0.92
    31. layer1.00
    32. sequence a#10.91
    33. sequence #20.96
    34. sequence #30.97
    35. self-attention1.00
    36. GQA1.00
    37. Figure 1.1: Language model meta-learning. During unsupervised pre-training, a language model develops a broad0.99
    38. set of skills and pattern recognition abilities. It then uses these abilities at inference time to rapidly adapt to or recognize0.99
    39. the desired task. We use the term "in-context learning" to describe the inner loop of this process, which occurs within0.99
    40. RMS Norm1.00
    41. the forward-pass upon each sequence. The sequences in this diagram are not intended to be representative of the data a1.00
    42. model would see during pre-training, but are intended to show that there are sometimes repeated sub-tasks embedded0.99
    43. within a single sequence.1.00
    44. Embedding1.00
    45. Causal1.00
    46. mask1.00
    47. Input Sequence1.00
    48. RoPE1.00
    49. World's Fair0.95
    50. TRACK 9·JUNE 30,20260.98
    51. Data Quality0.98
  • 2:02 #6 done41 line(s)

    shot 6·sharpness 4398.0

    1. AlEngineer0.98
    2. World'sFair1.00
    3. What were older base models trained on?0.99
    4. Quantity1.00
    5. Weight in1.00
    6. Epochs elapsed when1.00
    7. Dataset1.00
    8. (tokens)1.00
    9. training mix0.96
    10. training for 300B tokens1.00
    11. Common Crawl (filtered)0.99
    12. 410 billion1.00
    13. 60%1.00
    14. 0.441.00
    15. WebText21.00
    16. 19 billion0.99
    17. 22%1.00
    18. 2.91.00
    19. Books11.00
    20. 12 billion1.00
    21. 8%1.00
    22. 1.91.00
    23. Books21.00
    24. 55 billion1.00
    25. 8%1.00
    26. 0.431.00
    27. Wikipedia1.00
    28. 3 billion1.00
    29. 3%1.00
    30. 3.41.00
    31. Table 2.2: Datasets used to train GPT-3. "Weight in training mix" refers to the fraction of examples during training0.99
    32. that are drawn from a given dataset, which we intentionally do not make proportional to the size of the dataset. As a0.99
    33. result, when we train for 300 billion tokens, some datasets are seen up to 3.4 times during training while other datasets0.99
    34. are seen less than once.0.99
    35. that model on several key benchmarks.0.98
    36. (Llama 3)0.99
    37. Data mix summary. Our final data mix contains roughly 50% of tokens corresponding to general knowledge,1.00
    38. 25% of mathematical and reasoning tokens, 17% code tokens, and 8% multilingual tokens.0.99
    39. World's Fair0.99
    40. TRACK 9·JUNE 30,20260.98
    41. Data Quality1.00
  • 2:34 #7 skipped

    shot 7·duplicate of #6

  • 3:05 #8 done34 line(s)

    shot 8·sharpness 4681.1

    1. AlEngineer0.97
    2. World'sFair1.00
    3. What was post-training like back then?1.00
    4. 3 Methods and experimental details1.00
    5. • Language model post-training. The pre-trained language model has a rich understanding of language0.99
    6. 3.1 High-level methodology1.00
    7. but it does not yet follow instructions or behave in the way we would expect an assistant to. We1.00
    8. align the model with human feedback in several rounds, each of which involves supervised finetuning1.00
    9. Our methodology follows that of Ziegler et al. (2019) and1.00
    10. (SFT) on instruction tuning data and Direct Preference Optimization (DPO; Rafailov et al., 2024).0.99
    11. it in the stylistic continuation and summarization domains.0.99
    12. At this post-training2 stage, we also integrate new capabilities, such as tool-use, and observe strong0.99
    13. model (Radford et al., 2019; Brown et al., 2020; Fedus et al., 20.98
    14. improvements in other areas, such as coding and reasoning. See Section 4 for details. Finally, safety0.99
    15. 2022), a distribution of prompts on which we want our model t0.98
    16. mitigations are also incorporated into the model at the post-training stage, the details of which are0.99
    17. (Figure2).0.98
    18. of trained human labelers (see Sections 3.4 for details). We0.99
    19. described in Section 5.4.1.00
    20. Step 1: Collect demonstration data, and train a supervised policy. Our labelers provide demon-1.00
    21. strations of the desired behavior on the input prompt distribution (see Section 3.2 for details on this0.99
    22. distribution). We then fine-tune a pretrained GPT-3 model on this data using supervised learning.0.99
    23. Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons0.99
    24. between model outputs, where labelers indicate which output they prefer for a given input. We then0.99
    25. train a reward model to predict the human-preferred output.1.00
    26. Step 3: Optimize a policy against the reward model using PPO. We use the output of the0.99
    27. RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO0.99
    28. algorithm (Schulman et al., 2017).1.00
    29. Steps 2 and 3 can be iterated continuously; more comparison data is collected on the current best1.00
    30. policy, which is used to train a new RM and then a new policy. In practice, most of our comparison0.98
    31. data comes from our supervised policies, with some coming from our PPO policies.1.00
    32. World'sFair1.00
    33. TRACK 9· JUNE 30, 20260.94
    34. Data Quality1.00
  • 3:35 #9 skipped

    shot 9·duplicate of #8

  • 4:07 #10 done43 line(s)

    shot 10·sharpness 2749.9

    1. deepseek1.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. Everything changed0.98
    5. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via1.00
    6. Reinforcement Learning1.00
    7. Our large-scale reinforcement learning algorithm teaches the model how to think0.99
    8. productively using its chain of thought in a highly data-efficient training process. We0.99
    9. DeepSeek-AI1.00
    10. have found that the performance of o1 consistently improves with more reinforcement0.99
    11. learning (train-time compute) and with more time spent thinking (test-time compute).0.99
    12. [email protected]0.99
    13. The constraints on scaling this approach differ substantially from those of LLM0.98
    14. pretraining, and we are continuing to investigate them.1.00
    15. Abstract1.00
    16. General reasoning represents a long-standing and formidable challenge in artificial intelli-1.00
    17. 1001.00
    18. o1 AIME accuracy0.96
    19. during training1.00
    20. 1001.00
    21. o1 AlME accuracy0.96
    22. at test time1.00
    23. gence. Recent breakthroughs, exemplified by large language models (LLMs) (Brown et al.,0.98
    24. 2020; OpenAI, 2023) and chain-of-thought prompting (Wei et al., 2022b), have achieved con-1.00
    25. Claude 3.7 Sonnet shows particularly strong improvements in coding and front-1.00
    26. 801.00
    27. 801.00
    28. end web development. Along with the model, we're also introducing a command0.99
    29. line tool for agentic coding, Claude Code. Claude Code is available as a limited0.99
    30. pae @sscy0.65
    31. 601.00
    32. research preview, and enables developers to delegate substantial engineering0.99
    33. tasks to Claude directly from their terminal.1.00
    34. 401.00
    35. 201.00
    36. 201.00
    37. train-time compute (log scale)1.00
    38. test-time compute (log scale)1.00
    39. to the Claude Code research preview1.00
    40. of performance smoothly improves with both train-time and test-time compute0.99
    41. World's Fair1.00
    42. TRACK 9· JUNE 30,20260.95
    43. Data Quality1.00
  • 4:38 #11 skipped

    shot 11·duplicate of #10

  • 5:18 #12 skipped

    shot 12·duplicate of #10

  • 5:35 #13 done62 line(s)

    shot 13·sharpness 3010.6

    1. AlEngineer0.99
    2. Source family1.00
    3. Unique tokens (T)1.00
    4. Training tokens (T)1.00
    5. Mix Percentage (%)0.99
    6. Avg. Epochs1.00
    7. World's Fair0.98
    8. STEM1.00
    9. Code1.00
    10. 2.21.00
    11. 7.41.00
    12. 16.41.00
    13. 4.71.00
    14. 54.61.00
    15. 15.81.00
    16. 2.22×1.00
    17. 2.17×1.00
    18. Math1.00
    19. Books and journals0.99
    20. 0.31.00
    21. 0.61.00
    22. 1.61.00
    23. 0.91.00
    24. 5.41.00
    25. 3.10.90
    26. 5.28×1.00
    27. 1.65×1.00
    28. Modern base0.98
    29. PDFs1.00
    30. 2.71.00
    31. 1.41.00
    32. 4.70.87
    33. 0.53×1.00
    34. Web text1.00
    35. Multilingual (other)1.00
    36. 8.11.00
    37. 8.11.00
    38. 0.51.00
    39. 4.51.00
    40. 14.91.00
    41. 1.61.00
    42. 0.06×1.00
    43. 0.55×1.00
    44. model data recipes0.97
    45. Total1.00
    46. 29.21.00
    47. 30.01.00
    48. 100.01.00
    49. 1.03×1.00
    50. Table 5. MAI-Base-1 pre-training data composition. Unique tokens is the deduplicated token count per source0.99
    51. family. Training tokens is the number of tokens consumed from that family over the full run. Aug. Epochs is the0.99
    52. PRESENTED BY1.00
    53. ratio of training tokens to unique tokens; values above 1× indicate repeated sampling.1.00
    54. Microsoft1.00
    55. Nemotron 3 Ultra : Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning0.99
    56. (a) Data mixture of phase 10.99
    57. (b) Data mixture of phase 21.00
    58. Figure 4 | The data mixtures for both pretraining phases. We design the phase 1 data mixture to0.99
    59. have a bias for diversity and the phase 2 data mixture to have a bias for quality.1.00
    60. World's Fair0.99
    61. TRACK 9·JUNE 30,20260.98
    62. Data Quality0.99
  • 6:09 #14 skipped

    shot 14·duplicate of #13

  • 6:39 #15 skipped

    shot 15·duplicate of #13

  • 6:45 #16 skipped

    shot 16·duplicate of #13

  • 7:10 #17 skipped

    shot 17·duplicate of #13

  • 7:40 #18 skipped

    shot 18·duplicate of #13

  • 8:08 #19 skipped

    shot 19·duplicate of #13

  • 8:40 #20 skipped

    shot 20·duplicate of #13

  • 9:08 #21 skipped

    shot 21·duplicate of #13

  • 9:19 #22 done86 line(s)

    shot 22·sharpness 3947.9

    1. AlEngineer1.00
    2. Quantity1.00
    3. Weight in1.00
    4. Epochs elapsed when1.00
    5. World'sFair1.00
    6. Dataset1.00
    7. (tokens)1.00
    8. training mix1.00
    9. training for 300B tokens1.00
    10. Common Crawl (filtered)0.99
    11. 410 billion1.00
    12. 60%1.00
    13. 0.440.95
    14. WebText21.00
    15. 19 billion1.00
    16. 22%1.00
    17. 2.91.00
    18. Books11.00
    19. 12 billion0.97
    20. 8%1.00
    21. 1.91.00
    22. Books21.00
    23. 55 billion1.00
    24. 8%1.00
    25. 0.431.00
    26. GPT-31.00
    27. Wikipedia1.00
    28. 3 billion1.00
    29. 3%1.00
    30. 3.41.00
    31. Table 2.2: Datasets used to train GPT-3. "Weight in training mix" refers to the fraction of examples during training1.00
    32. that are drawn from a given dataset, which we intentionally do not make proportional to the size of the dataset. As a1.00
    33. result, when we train for 300 billion tokens, some datasets are seen up to 3.4 times during training while other datasets0.99
    34. are seen less than once.0.98
    35. Source family1.00
    36. Unique tokens (T)0.99
    37. Training tokens (T)1.00
    38. Mix Percentage (%)1.00
    39. Avg. Epochs0.99
    40. Code1.00
    41. 7.41.00
    42. 16.41.00
    43. 54.61.00
    44. 2.22×1.00
    45. STEM1.00
    46. 2.21.00
    47. 4.71.00
    48. 15.81.00
    49. 2.17×1.00
    50. Math1.00
    51. 0.31.00
    52. 1.60.87
    53. 5.41.00
    54. 5.28×1.00
    55. Books and journals1.00
    56. 0.61.00
    57. 0.91.00
    58. 3.11.00
    59. 1.65×1.00
    60. MAl-Base-10.97
    61. PDFs1.00
    62. Web text1.00
    63. 8.11.00
    64. 2.71.00
    65. 4.51.00
    66. 1.40.93
    67. 14.91.00
    68. 4.71.00
    69. 0.53×1.00
    70. 0.55×1.00
    71. Multilingual (other)1.00
    72. 8.11.00
    73. 0.51.00
    74. 1.61.00
    75. 0.06×1.00
    76. Total1.00
    77. 29.21.00
    78. 30.01.00
    79. 100.01.00
    80. 1.03×1.00
    81. Table 5. MAI-Base-1 pre-training data composition. Unique tokens is the deduplicated token count per source0.99
    82. family. Training tokens is the number of tokens consumed from that family over the full run. Aog. Epochs is the0.99
    83. ratio of training tokens to unique tokens; values above 1× indicate repeated sampling.0.99
    84. World's Fair0.98
    85. TRACK 9• JUNE 30, 20260.97
    86. Data Quality1.00
  • 9:52 #23 done66 line(s)

    shot 23·sharpness 4495.6

    1. AlEngineer1.00
    2. The advent of synthetic data for0.98
    3. REWRITING PRE-TRAINING DATA BOOSTS LLM PER-0.99
    4. FORMANCE IN MATH AND CODE1.00
    5. World's Fair0.98
    6. pretraining1.00
    7. Takumi Oamoto Shigeki shida Kakeru Hattor Youmi Ma0.60
    8. Taihei Shiotanil Koshiro Saito1 Masanari Oi1 Taishi Nakamura1.20.94
    9. Kazuki Fuji120.90
    10. Yukito Tajima1 Sakae Mizuki12 Masaki Kawamura10.94
    11. Hinari Shimada1.00
    12. Hiroya Takamura20.99
    13. Rio Yokota230.95
    14. Jun Sakuma Naoaki Okazaki1.20.94
    15. project.1.00
    16. National Institute of Advancd Industrial Science and Technology0.78
    17. Institute of Science Tokyo, Institute of Integrated Research, Supercomputing Research Center0.93
    18. Institute of Science Tokyo, Departmen of Computer Science0.98
    19. Trinity Large has 400B total parameters, with 13B activated per token, and is trained on a large 10.99
    20. SwallowCode SwallowMath0.95
    21. combining curated web-scale data and synthetic data. In the remainder of this report, we descril0.99
    22. decisions that govern architecture, training, and post-training, and we evaluate Trinity Large Basc as weu as0.99
    23. Trinity Large Preview across a wide range of standard benchmarks.1.00
    24. Knowledge Data Rephrasing Pre-training on natural, knowledge-intensive text presents a trade-off: a single epoch0.99
    25. is insufficient for comprehensive knowledge absorption, while multi-epoch repetition yields diminishing returns and1.00
    26. rephrasing framework composed of the following key components:0.99
    27. increases the risk of overfitting. To improve the token tility of high-quality knowledge tokens, we propose a synthetic0.98
    28. 4096 tokens0.92
    29. • Style- and perspective-diverse prompting: Inspired by WRAP [50], we apply a range of carefully engineered0.99
    30. split0.89
    31. fall input excerge0.86
    32. together as context1.00
    33. full outpot excerge0.93
    34. concat1.00
    35. prompts to enhance linguistic diversity while maintaining factual integrity. These prompts guide a large language0.99
    36. 256 tokens0.98
    37. Chunk-wise autoregressive generation: To preserve global coherence and avoid information loss in long1.00
    38. model to generate faithful rephrasings of the original texts in varied styles and from different perspectives.0.99
    39. documents, we adopt a chunk-based autoregressive rewriting strategy. Texts are divided into segments, rephrased0.98
    40. partial input excerpt 10.96
    41. rewrite model0.96
    42. partial output excerpt 10.95
    43. individually, and then stitched back together to form complete passages. This method mitigates implicit output0.99
    44. length limitations that typically exist with LLMs. An overview of this pipeline is presented in Figure 4.0.99
    45. partial input excerpt 20.99
    46. • Fidelity verification: To ensure consistency between original and rewritten content, we perform fidelity checks0.99
    47. that compare the semantic alignment of each rephrased passage with its source. This serves as an initial quality0.99
    48. control step prior to training.1.00
    49. 10 epochs, (2) rephrasing the data once and repeating it for 10 epochs, and (3) rephrasing the data 10 times with a1.00
    50. experiment with an early checkpoint of K2 and evaluate three training strategies: (1) repeating the original dataset for0.99
    51. We compare data rephrasing with multi-epoch repetition by testing their corresponding accuracy on SimpleQA. We0.99
    52. into a full rewritten passage.0.99
    53. Figure 4: Auto-regressive chunk-wise rephrasing pipeline for long input excerpts. The input is1.00
    54. split into smaller chunks with preserved context, rewritten sequentially, and then concatenated0.99
    55. effcaourrran-sgmtonWetismoer laralweandamDtaRhasing mcalraniterewtquim0.54
    56. single training pass. As shown in Table 1, the accuracy consistently improves across these strategies, demonstrating the1.00
    57. observed similarly encouraging results, and each corpora is rephrased at most twice.1.00
    58. :al documents into a "learning-note" style, following the methodology introduced in SwallowMath [16]. In addition,0.99
    59. we increased data diversity by translating high-quality mathematical materials from other languages into English.1.00
    60. Although initial experiments with rephrased subsets of our datasets show promising results, the use of synthetic data1.00
    61. as a strategy for continued scaling remains an active area of investigation. Key challenges include generalizing the0.99
    62. approach to diverse source domains without compromising factual accuracy, minimizing hallucinations and unintended0.97
    63. toxicity, and ensuring scalability to large-scale datasets.1.00
    64. World's Fair1.00
    65. TRACK 9• JUNE 30, 20260.95
    66. Data Quality1.00

Transcript

97 cues· 2,123 words· 12,002 chars

  1. 0:12 Hi everyone, my name is Varun.
  2. 0:14 I'm the pre-training lead at RCAI.
  3. 0:17 And the talk I'm gonna be giving today is called the base model is dead, but not really.
  4. 0:24 The idea of the base model that we have kind of is like built on this idea of like training on super large scale web text and the base model kind of being a reflection of like the whole knowledge of like the human internet.
  5. 0:43 You can see in like, I've taken these from a bunch of different papers on like the entire LLM training process.
  6. 0:52 Our own model, training large thinking, the process looked kind of like the simplified diagram on the left.
  7. 1:00 I've taken the top one from GLM 4.5, the bottom one from GLM 5.
  8. 1:07 All these have a pre-training phase and
  9. 1:11 Pre-training is like the stage where the model accumulates world knowledge, builds useful representations, all through next token prediction on web text.
  10. 1:24 I've got a simplified transformer diagram, decoder-only transformer.
  11. 1:32 a screenshot from the GPT-3 paper that talks about how language models can learn how to do in-context learning through unsupervised or self-supervised or some would even just call it supervised learning through Next Token Prediction.
  12. 1:58 The way that older base models were trained was, like I said, mostly on things that reflected the entirety of human knowledge.
  13. 2:07 So Common Crawl, which is like a commonly available web scrape, made up most of the training data set for GPD3.
  14. 2:16 Webtex2, another web scrape data set.
  15. 2:19 Some sources from like books as well.
  16. 2:22 And Wikipedia as like a high quality representation of human knowledge.
  17. 2:29 You can see that webtext alone here, including Wikipedia, makes up roughly 85% of the whole training mix.
  18. 2:40 Looking at the bottom with Llama 3, webtext still kind of makes up a majority of the model's training data, with like 50% of the tokens corresponding to general knowledge.
  19. 2:56 Back then, post-training was mostly shaping the model to surface the knowledge that it accumulated in pre-training in a chat interface.
  20. 3:11 So mostly allowing the model to adapt to a chat template, to the question-answer format, and be useful in an interaction that way.
  21. 3:22 RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself.
  22. 3:36 Now, in this realm of how language models used to be,
  23. 3:43 Pre-training and the base model kind of defined how good you were able to get a model.
  24. 3:50 It was like the bulk of the compute budget and it was the core of the training process.
  25. 4:02 However, this kind of changed a lot last year when
  26. 4:09 I guess 2024, actually.
  27. 4:11 OpenAI released O1, pioneering reasoning models, and DeepSeq also released R1 in January 2025, allowing the whole world to know how to build these types of language models.
  28. 4:28 And now we have this new use for reinforcement learning, which is no longer a cherry on top.
  29. 4:39 but it can dramatically improve the performance of the model on various different tasks.
  30. 4:45 The famous graph sum 01 there talking about AMI performance competitive math contest.
  31. 4:54 And then even later in the year, we saw Cloud Code come into being as a way for developers to easily kind of use language models in a terminal to build out applications as models got stronger and stronger on things like function calling.
  32. 5:13 And then people realized you could, you know, RL this end to end.
  33. 5:18 And now models could learn how to interact with software environments and build software and perform really useful work.
  34. 5:28 And so, the question then becomes like, is your standard base model still what the best, what would be the best
  35. 5:41 prior for this large-scale reinforcement learning phase that Reasoners and agentic models now use.
  36. 5:52 And we can kind of see in a few open research papers where the trend is going.
  37. 6:00 And interestingly enough, it seems like not super clear yet.
  38. 6:09 I have my opinions on synthetic data being the way forward, but I've got two contrasting perspectives here, kind of in the slide.
  39. 6:18 The top image is from the MAI Thinking One paper, where they make a really large point to not use any synthetic data or any data from any other language model, and they really try to filter their web scripts for this as well, in order to kind of adhere to the previous
  40. 6:38 of using human knowledge as a way to bootstrap model representations and capabilities.
  41. 6:50 But I would say that this is also, even though they stuck with no synthetic data, the data mix that they've chosen here is still totally different from what you'd expect in a classical language model.
  42. 7:06 And the main reason for that is that web text, which used to make up to 85% of the trained data in GBT-3, is now all the way down at 15%.
  43. 7:17 And that just shows that the value of web text contributing to the downstream performance of the models on RL and stuff is...
  44. 7:31 kind of, it's still important, but taking a backseat to things like code and stem abilities as the models kind of gain more real-world use cases related to those.
  45. 7:44 The other approach is to bring post-training data and large-scale synthetic data back through, pull it back through the process into the pre-training phase.
  46. 7:56 The bottom chart I've taken from Nemotron 3 Ultra, where they reveal their data recipe, and I'm not sure how readable it is, but these top three on the left pie chart, the top three on the kind of right side of it, they're all labeled SFT, with SFT as a prefix.
  47. 8:22 That's the type of question announcer kind of chat data set that you'd expect to see only in post-training.
  48. 8:28 But by pulling it back into the process, they're able to get the model to learn the shape of these conversations and what kind of tasks they might be expected to do downstream from the very beginning of the pre-training process.
  49. 8:46 This follows a similar trend in diminishing the amount of web text used in the model.
  50. 8:58 Yeah.

Chapters

  1. 0:00 The base model as a mirror of the web
  2. 1:26 How knowledge accumulates in training
  3. 2:49 When instruction data moves earlier
  4. 4:11 After o1: RL and reasoning
  5. 5:41 What prior the base model must carry
  6. 6:18 Filtering web text, adding synthetic
  7. 8:01 Reading the open data recipes
  8. 9:41 Lessons from training Trinity
  9. 12:02 Balancing coefficients and early stability
  10. 13:30 Why RL keeps raising the stakes
  11. 15:55 The base model's shifting job

Open at this second