Videos TUnPNY4E2fw
Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI
Scene timeline
46 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 89
- whisperx 89
- chunks
- 25
- from 89 cues
- keyframes
- 42
- kept of 46 captured
- frames with text
- 42
- 1,114 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.7 MB
- word timings on 89 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 11:01 | 1m 05s |
stt |
done | — | 2026-08-11 11:02 | 15s |
chunk |
done | — | 2026-08-11 11:02 | 0s |
text_embed |
done | — | 2026-08-11 11:02 | 0s |
keyframe |
done | — | 2026-08-11 11:02 | 1m 23s |
ocr |
done | — | 2026-08-11 11:04 | 20s |
frame_embed |
done | — | 2026-08-11 11:04 | 7s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.94
-
- together.ai1.00
- The AI Native Cloud0.93
- Road to 5M0.99
- Sequence Length1.00
- Breaking Memory Barriers in1.00
- Context Parallelism1.00
- Max Ryabinin1.00
- VP R&D, Model Shaping0.99
- AlEngineer1.00
- EUROPE1.00
-
- together.ai1.00
- The AI Native Cloud1.00
- Road to 5M1.00
- ***0.77
- AIE1.00
- Sequence Length1.00
- ★1.00
- Breaking Memory Barriers in0.99
- Context Parallelism1.00
- Max Ryabinin1.00
- VP R&D, Model Shaping1.00
- Brea1.00
- Con1.00
- Max Rye0.92
- Google DeepMind1.00
- VP R&D0.90
-
- Al Native Cloud1.00
- Model Builders1.00
- → App Developers0.97
- GPU Clusters1.00
- Model Shaping1.00
- Inference1.00
- AIE1.00
- ★1.00
- GPUs as a service with1.00
- Tailored customization for1.00
- The fastest way to launch1.00
- ★1.00
- ★1.00
- accelerated training stack0.99
- your tasks0.98
- Al models0.98
- Accelerate model training0.99
- Complete ownership1.00
- 200+ leading models1.00
- Observability1.00
- Fine-tuning + distillation1.00
- Serverless and Dedicated0.99
- B200 + GB200 Leader1.00
- √ RL (private beta)0.93
- Advanced optimizations1.00
- Storage services1.00
- Trusted by:1.00
- CURSOR1.00
- Runware1.00
- hedra1.00
- Decagon1.00
- KREA1.00
- Braintrust1.00
- WorkOS OpenAI0.98
-
- Al Native Cloud0.99
- Model Builders1.00
- → App Developers0.99
- ★1.00
- GPU Clusters1.00
- Model Shaping1.00
- Inference1.00
- AIE1.00
- ★1.00
- GPUs as a service with1.00
- Tailored customization for1.00
- The fastest way to launch1.00
- ★1.00
- ★1.00
- ★1.00
- accelerated training stack0.99
- your tasks0.97
- Al models0.98
- Accelerate model training0.99
- Complete ownership1.00
- 200+ leading models0.99
- Observability1.00
- Fine-tuning + distillation1.00
- Serverless and Dedicated1.00
- B200 + GB200 Leader1.00
- RL (private beta)0.96
- Advanced optimizations1.00
- Storage services1.00
- Trusted by:1.00
- CURSOR1.00
- Runware1.00
- hedra1.00
- Decagon1.00
- KREA1.00
- AlEngineer0.97
- EUROPE1.00
-
- Why do we want to long-context training?0.99
- Did the car actually do1.00
- this?1.00
- /context1.00
- AIE1.00
- Context Usage1.00
- claude-sonnet-4-5-20250929 · 163k/200k toker0.98
- ★1.00
- ★1.00
- System tools: 15.2k tokens (7.6%)0.99
- System prompt: 2.6k tokens (1.3%)0.97
- MCP tools: 94.2k tokens (47.1%)0.99
- Custom agents: 758 tokens (0.4%)0.99
- C0.58
- Messages: 5.4k tokens (2.7%)0.97
- 口区0.80
- Free space: 37k (18.4%)0.97
- 区0.54
- Autocompact buffer: 45.0k tokens (22.5%)0.98
- Engineering the future of Al1.00
-
- Why do we want to long-context training?0.99
- context0.98
- Context Usage0.98
- claude-sonnet-4-5-20250929·163k/200k toke0.97
- System prompt: 2.6k tokens (1.3%)0.99
- System tools: 15.2k tokens (7.6%)0.95
- MCP tools: 94.2k tokens (47.1%)0.98
- Custom agents: 758 tokens (0.4%)0.96
- Messages: 5.4k tokens (2.7%)0.97
- Free space: 37k (18.4%)0.96
- Autocompact buffer: 45.8k tokens (22.5%)0.97
- AlEngineer1.00
- EUROPE1.00
-
- Why do we want to long-context training?0.98
- Did the car actually do1.00
- this?1.00
- /context1.00
- AIE1.00
- ★1.00
- Context Usage1.00
- claude-sonnet-4-5-20250929· 163k/200k toker0.99
- ★1.00
- ★1.00
- ★1.00
- System tools: 15.2k tokens (7.6%)0.97
- System prompt: 2.6k tokens (1.3%)0.97
- MCP tools: 94.2k tokens (47.1%)1.00
- Custom agents: 758 tokens (0.4%)1.00
- C0.68
- Messages: 5.4k tokens (2.7%)0.98
- 区区0.62
- Free space: 37k (18.4%)0.98
- 区区0.53
- Autocompact buffer: 45.0k tokens (22.5%)0.99
- Even not at insane scales, it helps to know where the memory goes!0.99
- AlEngineer0.97
- EUROPE1.00
-
- What's stopping us?0.98
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- AlEngineer0.97
- EUROPE1.00
-
- What's stopping us?0.97
- Long Context1.00
- AIE1.00
- O(N^2)1.00
- O(N)0.99
- Computation1.00
- Memory1.00
- ★1.00
- ★1.00
- Engineering the future of Al0.99
-
- What's stopping us?1.00
- Long Context1.00
- AIE1.00
- O(N^2)1.00
- O(N)1.00
- Computation1.00
- Memory1.00
- ★1.00
- ★1.00
- ★1.00
- Memory Usage for 8B Model0.95
- DP=81.00
- DP=8 Zero-11.00
- DP=8 Zero-21.00
- DP=8 Zero-30.99
- 1601.00
- Memg ug)0.71
- 1401.00
- 1201.00
- Model Parameters1.00
- 1001.00
- Gradients1.00
- 601.00
- 801.00
- Optimizer States0.99
- 401.00
- Activations1.00
- 201.00
- 1024 40960.94
- 163841.00
- 1024 40960.99
- 163841.00
- 10241.00
- 40961.00
- 163841.00
- 10241.00
- 40961.00
- 163841.00
- Sequence Length1.00
- Sequence Length1.00
- Sequence Length1.00
- Sequence Length1.00
- Engineering the future of Al1.00
-
- How far can we get?1.00
- 801.00
- Model1.00
- *★*0.62
- AIE1.00
- Peakk k i)0.76
- 601.00
- Attn act.1.00
- Other1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 401.00
- OOM0.92
- (119)1.00
- 201.00
- 01.00
- Default1.00
- Llama 3-8B, 3M tokens, 8xH1000.99
- Engineering the future of Al0.99
-
- With FSDP1.00
- 801.00
- Model1.00
- AIE1.00
- (GiB)0.74
- 601.00
- Attn act.1.00
- Other1.00
- eak orry0.72
- OOM0.89
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- 401.00
- OOM0.93
- (119)1.00
- (7684)1.00
- 201.00
- 15.01.00
- 01.00
- Default1.00
- FSDP1.00
- Llama 3-8B, 3M tokens, 8xH1000.99
- AlEngineer0.97
- EUROPE1.00
-
- DeepSpeed Ulysses1.00
- rank1 rank20.98
- ranko0.94
- rank31.00
- q0,0.97
- ko0.72
- rankO0.94
- AIE1.00
- q1.0.83
- ★1.00
- ★1.00
- ★1.00
- k10.99
- rank11.00
- All2All0.98
- q2,0.91
- k21.00
- rank21.00
- q3,0.99
- k31.00
- rank31.00
- What if we divide by num_heads dim?0.99
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. Jacobs et al., 20230.99
- What1.00
- Engineering the future of Al0.99
-
- DeepSpeed Ulysses1.00
- q0,0.98
- ko0.78
- o00.79
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- 92,0.87
- k21.00
- q1,0.87
- k10.99
- All2All0.95
- Full Flash Attention1.00
- All2All0.94
- o20.95
- o10.95
- q3,1.00
- k31.00
- o30.91
- What if we divide by num_heads dim?1.00
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. Jacobs et al., 20230.99
- Whati0.99
- Engineering the future of Al0.99
-
- With Ulysses context parallelism1.00
- 801.00
- Model1.00
- Attn act.1.00
- *★*0.61
- AIE1.00
- (Gi)0.74
- 601.00
- Other1.00
- Peak ory0.70
- OOM0.85
- OOM0.90
- ★1.00
- ★1.00
- ★1.00
- 401.00
- OOM0.89
- (119)1.00
- (7684)1.00
- (964)1.00
- 201.00
- 15.01.00
- 15.01.00
- 01.00
- Default1.00
- FSDP1.00
- FSDP1.00
- + Ulysses0.97
- Llama 3-8B, 3M tokens, 8xH1000.98
- Engineering the future of Al0.99
-
- What if the activations are too big?0.99
- We need them for the backward pass, but we can recompute0.99
- Drawing inspired by https://github.comy/cybertronai/gradlent-checkpointing0.93
- ★1.00
- Forward pass0.99
- Orange nodes are the ones1.00
- kept in memory to compute1.00
- AIE1.00
- the gradient update for this0.99
- node1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Backward pass1.00
- Orange nodes are the ones1.00
- kept in memory to compute0.98
- the gradient update for this0.99
- node1.00
- img src: https0.99
- Braintrust1.00
- WorkOS OpenAI0.98
-
- With activation checkpointing1.00
- 801.00
- Model1.00
- Attn act.1.00
- *★★0.62
- AIE1.00
- ★1.00
- (Gi)0.77
- 601.00
- Other1.00
- Pek ory0.75
- OOM0.89
- OOM0.89
- OOM0.93
- ★1.00
- ★1.00
- ★1.00
- 401.00
- OOM0.90
- (119)1.00
- (7684)1.00
- (964)1.00
- (130)1.00
- 201.00
- 15.01.00
- 15.01.00
- 15.01.00
- 01.00
- Default1.00
- FSDP1.00
- FSDP1.00
- FSDP1.00
- + Ulysses0.99
- + Ulysses0.99
- + AC0.98
- Llama 3-8B, 3M tokens, 8xH1000.99
- AlEngineer0.97
- EUROPE1.00
-
- What if the activations are still too big?0.98
- 1. Inputs to each Transformer block can't be recomputed quickly1.00
- 2. Still, we can offload to CPU1.00
- ***0.62
- AIE1.00
- ★0.99
- o0.62
- Prefetch before we reach the layer, similar to weights for FSDP1.00
- ★1.00
- ★1.00
- o0.70
- First implemented in:1.00
- Blog0.98
- Unsloth Gradient1.00
- Checkpointing-4x longer0.99
- context windows1.00
- Apr 9, 2024 · By Daniel & Mike0.97
- unsloth.ai/blog/long-context1.00
- AlEngineer0.97
- EUROPE1.00
Transcript
89 cues· 1,942 words· 11,022 chars
- 0:14 Hi everyone, my name is Max.
- 0:16 I am VP of Research and Development at Together.ai and today I'm going to tell you about our research project which is called Roll to 5 million sequence length breaking memory barriers in context parallelism.
- 0:30 So to begin I'll first say a few words about Together.ai and who we are.
- 0:36 Together.ai is an AI-native cloud which provides services and infrastructure for AI developers and builders at all stages of development, starting from creating a model where you just might need a GPU cluster with heavily optimized computes and highly reliable
- 0:58 systems all the way through model shaping or you can take existing models and customize them for your tasks in terms of performance, in terms of speed, in terms of quality through services such as the funtioning or reinforcement learning.
- 1:15 Also, we are an inference provider, so if you have an app which is reliant on open source model inference, you can work with us and we'll provide you with the fastest way to launch and use AI models.
- 1:33 with more than 200 models in our portfolio, options for deployment which include serverless and dedicated inference, and a ton of advanced optimizations which I will not be able to speak about today.
- 1:48 The purpose of this talk is focused on model training, customization, and fine-tuning in particular.
- 1:56 And I'll start by asking a question.
- 2:00 I think in the last few months, or at least a year or so, we're seeing a lot of interest in the community, both on the system side and on the research side, in training long context models.
- 2:14 The primary reasons for that are twofold, I would say.
- 2:17 First of all, with the explosion in popularity of agents, you can see a lot of different applications where you might want to put as many tokens as you want in your context, and you want the model to leverage that context effectively.
- 2:35 Second, with the development of applications such as video generation, you might often need to keep track of multiple frames per second, which can occupy quite a few tokens in your context pretty quickly.
- 2:57 And you also need models that have good sense of temporal consistency, which means that they are able to see what was happening a few seconds or ideally a few minutes ago.
- 3:10 To do that all effectively, you need to make sure that the models are able to process that context and work with it correctly at the training time.
- 3:23 But even if you're not at the scales of millions of tokens in the context length, it's still quite important to understand where the memory goes.
- 3:34 Because who knows, maybe you might be able to reinvest it in some other ways and speed up your training overall.
- 3:42 So the problem here is that if you are taking a standard transformer-based language model and trying to extend its context, you can run into two bottlenecks.
- 3:57 Bottom line number one is that you are faced with quadratic computation because, long story short, for transformer-based models, you have pairwise interactions across all the elements in a sequence.
- 4:11 The second problem is more insidious, one might say.
- 4:16 As you continue scaling your context, your memory keeps growing linearly, which is not as bad, but still pretty difficult to deal with.
- 4:26 unless you apply a range of specific techniques.
- 4:31 And this is an example from Hugging Face's blog post on model training, which shows that the sequence length growth can affect your memory limits pretty considerably.
- 4:47 Here's the slide.
- 4:48 And our goal of that project was to see how far exactly we are able to get with a range of existing techniques that are pretty well known to some in the community as well as some further optimizations that we wanted to leverage to push this a bit further ahead.
- 5:11 So let's say you're taking a model which is a standard Lemma 3B architecture, you're trying to fit 3 million trained tokens into your context, and you're taking all of this on an 8x H100 GPU node.
- 5:30 The first stage you'll see is that even with just model parameters, you're not able to fit it into the GPU.
- 5:40 You run out of memory just by trying to place the model.
- 5:43 Of course, the next stage is to apply fully-sharded data parallelism, where all the parameters are basically chunked across the eight GPUs that you have.
- 5:53 which is great, but still doesn't solve the problem.
- 5:56 You see that the memory usage for the model drops quite significantly, but you still are running out of memory because of all the attention activations.
- 6:08 The next point that we've leveraged and I encourage you to use as well is taking advantage of context parallelism.
- 6:19 In particular, there is a pretty well-known technique called deep-speed Ulysses, first introduced by Microsoft.
- 6:26 The idea is that instead of computing all of your multi-head attention on every GPU separately for the whole sequence, you can do something more clever.
- 6:38 In particular, you can try to compute the attention for different heads at different points in time or on different GPUs through communicating these activations as they are required.
- 6:54 in such a way that one GPU is only responsible for one attention head here, but it's still computing the attention over the whole sequence.
- 7:05 That technique is quite effective at addressing the problem, and it also allows it to utilize the best possible attention implementation, like flash attention 1, 2, 3, 4.
- 7:18 to optimize that part of the computation.
- 7:22 And then you aggregate the results as you would have previously.
- 7:27 So if you apply Ulysses context parallelism, the utilization drops quite significantly, like approximately 8x here, as it should.
- 7:39 But we are still quite far from our goal of being able to fit that onto just a single H100 node.
- 7:48 So what happens next is that we can try to recompute the activations as they are needed to us at the backward pass.
- 7:58 That technique is known as activation checkpointing, and it's available in pretty much all of the deep learning frameworks these days that you could use.
- 8:08 You just need to enable it in a correct way that does not impose too much of a computational burden on you.
- 8:19 With that, with activation checkpointing, you can drop the activation usage by a further factor of 8, but still something else needs to be done.
- 8:31 The next optimization is also connected to the storage of activations.
- 8:38 you can try to store some of the inputs to each transformer block, not on the GPU, but instead offload them to CPU when they are not required.
- 8:52 This is not very impactful for the performance because you can offload it and prefetch when you are trying to backpropagate to the corresponding layer.
- 9:05 This optimization, to the best of our knowledge, was first implemented by Onsloth, and it allows you to drastically expand the context window.
- 9:17 The next point is that you're getting with uploading next to 37 gigabytes of data, but then comes the other part of out-of-memory usage.
loading