read-only demo

Videos TUnPNY4E2fw

Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI

index_state ready data_status ok

AI Engineer· published 2026-06-08· 0:15:50· en-US· indexed 2026-08-11 11:04

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:24, 1 of 1 keyframes kept
  5. Shot 4, 0:24 to 0:32, 1 of 1 keyframes kept
  6. Shot 5, 0:32 to 1:00, 1 of 1 keyframes kept
  7. Shot 6, 1:00 to 1:28, 0 of 1 keyframes kept
  8. Shot 7, 1:28 to 1:55, 1 of 1 keyframes kept
  9. Shot 8, 1:55 to 2:35, 1 of 1 keyframes kept
  10. Shot 9, 2:35 to 2:57, 1 of 1 keyframes kept
  11. Shot 10, 2:57 to 3:42, 1 of 1 keyframes kept
  12. Shot 11, 3:42 to 3:49, 1 of 1 keyframes kept
  13. Shot 12, 3:49 to 4:29, 1 of 1 keyframes kept
  14. Shot 13, 4:29 to 4:47, 1 of 1 keyframes kept
  15. Shot 14, 4:47 to 5:15, 1 of 1 keyframes kept
  16. Shot 15, 5:15 to 5:43, 0 of 1 keyframes kept
  17. Shot 16, 5:43 to 6:14, 1 of 1 keyframes kept
  18. Shot 17, 6:14 to 6:39, 1 of 1 keyframes kept
  19. Shot 18, 6:39 to 7:04, 0 of 1 keyframes kept
  20. Shot 19, 7:04 to 7:26, 1 of 1 keyframes kept
  21. Shot 20, 7:26 to 7:49, 1 of 1 keyframes kept
  22. Shot 21, 7:49 to 8:17, 1 of 1 keyframes kept
  23. Shot 22, 8:17 to 8:29, 1 of 1 keyframes kept
  24. Shot 23, 8:29 to 9:16, 1 of 1 keyframes kept
  25. Shot 24, 9:16 to 9:28, 1 of 1 keyframes kept
  26. Shot 25, 9:28 to 9:53, 1 of 1 keyframes kept
  27. Shot 26, 9:53 to 10:03, 1 of 1 keyframes kept
  28. Shot 27, 10:03 to 10:33, 1 of 1 keyframes kept
  29. Shot 28, 10:33 to 11:02, 1 of 1 keyframes kept
  30. Shot 29, 11:02 to 11:05, 1 of 1 keyframes kept
  31. Shot 30, 11:05 to 11:16, 1 of 1 keyframes kept
  32. Shot 31, 11:16 to 11:20, 0 of 1 keyframes kept
  33. Shot 32, 11:20 to 11:33, 1 of 1 keyframes kept
  34. Shot 33, 11:33 to 11:45, 1 of 1 keyframes kept
  35. Shot 34, 11:45 to 12:10, 1 of 1 keyframes kept
  36. Shot 35, 12:10 to 12:29, 1 of 1 keyframes kept
  37. Shot 36, 12:29 to 12:46, 1 of 1 keyframes kept
  38. Shot 37, 12:46 to 13:00, 1 of 1 keyframes kept
  39. Shot 38, 13:00 to 13:33, 1 of 1 keyframes kept
  40. Shot 39, 13:33 to 13:47, 1 of 1 keyframes kept
  41. Shot 40, 13:47 to 13:49, 1 of 1 keyframes kept
  42. Shot 41, 13:49 to 14:15, 1 of 1 keyframes kept
  43. Shot 42, 14:15 to 14:42, 1 of 1 keyframes kept
  44. Shot 43, 14:42 to 15:08, 1 of 1 keyframes kept
  45. Shot 44, 15:08 to 15:34, 1 of 1 keyframes kept
  46. Shot 45, 15:34 to 15:49, 1 of 1 keyframes kept

46 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
89
whisperx 89
chunks
25
from 89 cues
keyframes
42
kept of 46 captured
frames with text
42
1,114 lines read
chapters
0
from the source metadata
keyframe bytes
5.7 MB
word timings on 89 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 11:01 1m 05s
stt done 2026-08-11 11:02 15s
chunk done 2026-08-11 11:02 0s
text_embed done 2026-08-11 11:02 0s
keyframe done 2026-08-11 11:02 1m 23s
ocr done 2026-08-11 11:04 20s
frame_embed done 2026-08-11 11:04 7s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 662.1

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 814.7

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:12 #2 done3 line(s)

    shot 2·sharpness 905.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.94
  • 0:23 #3 done10 line(s)

    shot 3·sharpness 405.3

    1. together.ai1.00
    2. The AI Native Cloud0.93
    3. Road to 5M0.99
    4. Sequence Length1.00
    5. Breaking Memory Barriers in1.00
    6. Context Parallelism1.00
    7. Max Ryabinin1.00
    8. VP R&D, Model Shaping0.99
    9. AlEngineer1.00
    10. EUROPE1.00
  • 0:31 #4 done16 line(s)

    shot 4·sharpness 2941.7

    1. together.ai1.00
    2. The AI Native Cloud1.00
    3. Road to 5M1.00
    4. ***0.77
    5. AIE1.00
    6. Sequence Length1.00
    7. 1.00
    8. Breaking Memory Barriers in0.99
    9. Context Parallelism1.00
    10. Max Ryabinin1.00
    11. VP R&D, Model Shaping1.00
    12. Brea1.00
    13. Con1.00
    14. Max Rye0.92
    15. Google DeepMind1.00
    16. VP R&D0.90
  • 0:57 #5 done34 line(s)

    shot 5·sharpness 3517.0

    1. Al Native Cloud1.00
    2. Model Builders1.00
    3. → App Developers0.97
    4. GPU Clusters1.00
    5. Model Shaping1.00
    6. Inference1.00
    7. AIE1.00
    8. 1.00
    9. GPUs as a service with1.00
    10. Tailored customization for1.00
    11. The fastest way to launch1.00
    12. 1.00
    13. 1.00
    14. accelerated training stack0.99
    15. your tasks0.98
    16. Al models0.98
    17. Accelerate model training0.99
    18. Complete ownership1.00
    19. 200+ leading models1.00
    20. Observability1.00
    21. Fine-tuning + distillation1.00
    22. Serverless and Dedicated0.99
    23. B200 + GB200 Leader1.00
    24. √ RL (private beta)0.93
    25. Advanced optimizations1.00
    26. Storage services1.00
    27. Trusted by:1.00
    28. CURSOR1.00
    29. Runware1.00
    30. hedra1.00
    31. Decagon1.00
    32. KREA1.00
    33. Braintrust1.00
    34. WorkOS OpenAI0.98
  • 1:03 #6 skipped

    shot 6·duplicate of #5

  • 1:36 #7 done36 line(s)

    shot 7·sharpness 3419.3

    1. Al Native Cloud0.99
    2. Model Builders1.00
    3. → App Developers0.99
    4. 1.00
    5. GPU Clusters1.00
    6. Model Shaping1.00
    7. Inference1.00
    8. AIE1.00
    9. 1.00
    10. GPUs as a service with1.00
    11. Tailored customization for1.00
    12. The fastest way to launch1.00
    13. 1.00
    14. 1.00
    15. 1.00
    16. accelerated training stack0.99
    17. your tasks0.97
    18. Al models0.98
    19. Accelerate model training0.99
    20. Complete ownership1.00
    21. 200+ leading models0.99
    22. Observability1.00
    23. Fine-tuning + distillation1.00
    24. Serverless and Dedicated1.00
    25. B200 + GB200 Leader1.00
    26. RL (private beta)0.96
    27. Advanced optimizations1.00
    28. Storage services1.00
    29. Trusted by:1.00
    30. CURSOR1.00
    31. Runware1.00
    32. hedra1.00
    33. Decagon1.00
    34. KREA1.00
    35. AlEngineer0.97
    36. EUROPE1.00
  • 2:26 #8 done20 line(s)

    shot 8·sharpness 2528.9

    1. Why do we want to long-context training?0.99
    2. Did the car actually do1.00
    3. this?1.00
    4. /context1.00
    5. AIE1.00
    6. Context Usage1.00
    7. claude-sonnet-4-5-20250929 · 163k/200k toker0.98
    8. 1.00
    9. 1.00
    10. System tools: 15.2k tokens (7.6%)0.99
    11. System prompt: 2.6k tokens (1.3%)0.97
    12. MCP tools: 94.2k tokens (47.1%)0.99
    13. Custom agents: 758 tokens (0.4%)0.99
    14. C0.58
    15. Messages: 5.4k tokens (2.7%)0.97
    16. 口区0.80
    17. Free space: 37k (18.4%)0.97
    18. 0.54
    19. Autocompact buffer: 45.0k tokens (22.5%)0.98
    20. Engineering the future of Al1.00
  • 2:52 #9 done13 line(s)

    shot 9·sharpness 361.6

    1. Why do we want to long-context training?0.99
    2. context0.98
    3. Context Usage0.98
    4. claude-sonnet-4-5-20250929·163k/200k toke0.97
    5. System prompt: 2.6k tokens (1.3%)0.99
    6. System tools: 15.2k tokens (7.6%)0.95
    7. MCP tools: 94.2k tokens (47.1%)0.98
    8. Custom agents: 758 tokens (0.4%)0.96
    9. Messages: 5.4k tokens (2.7%)0.97
    10. Free space: 37k (18.4%)0.96
    11. Autocompact buffer: 45.8k tokens (22.5%)0.97
    12. AlEngineer1.00
    13. EUROPE1.00
  • 3:28 #10 done24 line(s)

    shot 10·sharpness 3136.7

    1. Why do we want to long-context training?0.98
    2. Did the car actually do1.00
    3. this?1.00
    4. /context1.00
    5. AIE1.00
    6. 1.00
    7. Context Usage1.00
    8. claude-sonnet-4-5-20250929· 163k/200k toker0.99
    9. 1.00
    10. 1.00
    11. 1.00
    12. System tools: 15.2k tokens (7.6%)0.97
    13. System prompt: 2.6k tokens (1.3%)0.97
    14. MCP tools: 94.2k tokens (47.1%)1.00
    15. Custom agents: 758 tokens (0.4%)1.00
    16. C0.68
    17. Messages: 5.4k tokens (2.7%)0.98
    18. 区区0.62
    19. Free space: 37k (18.4%)0.98
    20. 区区0.53
    21. Autocompact buffer: 45.0k tokens (22.5%)0.99
    22. Even not at insane scales, it helps to know where the memory goes!0.99
    23. AlEngineer0.97
    24. EUROPE1.00
  • 3:47 #11 done7 line(s)

    shot 11·sharpness 1047.6

    1. What's stopping us?0.98
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. AlEngineer0.97
    7. EUROPE1.00
  • 4:25 #12 done10 line(s)

    shot 12·sharpness 1283.9

    1. What's stopping us?0.97
    2. Long Context1.00
    3. AIE1.00
    4. O(N^2)1.00
    5. O(N)0.99
    6. Computation1.00
    7. Memory1.00
    8. 1.00
    9. 1.00
    10. Engineering the future of Al0.99
  • 4:40 #13 done43 line(s)

    shot 13·sharpness 1695.3

    1. What's stopping us?1.00
    2. Long Context1.00
    3. AIE1.00
    4. O(N^2)1.00
    5. O(N)1.00
    6. Computation1.00
    7. Memory1.00
    8. 1.00
    9. 1.00
    10. 1.00
    11. Memory Usage for 8B Model0.95
    12. DP=81.00
    13. DP=8 Zero-11.00
    14. DP=8 Zero-21.00
    15. DP=8 Zero-30.99
    16. 1601.00
    17. Memg ug)0.71
    18. 1401.00
    19. 1201.00
    20. Model Parameters1.00
    21. 1001.00
    22. Gradients1.00
    23. 601.00
    24. 801.00
    25. Optimizer States0.99
    26. 401.00
    27. Activations1.00
    28. 201.00
    29. 1024 40960.94
    30. 163841.00
    31. 1024 40960.99
    32. 163841.00
    33. 10241.00
    34. 40961.00
    35. 163841.00
    36. 10241.00
    37. 40961.00
    38. 163841.00
    39. Sequence Length1.00
    40. Sequence Length1.00
    41. Sequence Length1.00
    42. Sequence Length1.00
    43. Engineering the future of Al1.00
  • 4:55 #14 done21 line(s)

    shot 14·sharpness 2006.6

    1. How far can we get?1.00
    2. 801.00
    3. Model1.00
    4. *★*0.62
    5. AIE1.00
    6. Peakk k i)0.76
    7. 601.00
    8. Attn act.1.00
    9. Other1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. 401.00
    15. OOM0.92
    16. (119)1.00
    17. 201.00
    18. 01.00
    19. Default1.00
    20. Llama 3-8B, 3M tokens, 8xH1000.99
    21. Engineering the future of Al0.99
  • 5:21 #15 skipped

    shot 15·duplicate of #14

  • 5:50 #16 done26 line(s)

    shot 16·sharpness 1995.1

    1. With FSDP1.00
    2. 801.00
    3. Model1.00
    4. AIE1.00
    5. (GiB)0.74
    6. 601.00
    7. Attn act.1.00
    8. Other1.00
    9. eak orry0.72
    10. OOM0.89
    11. 1.00
    12. 1.00
    13. 1.00
    14. 1.00
    15. 401.00
    16. OOM0.93
    17. (119)1.00
    18. (7684)1.00
    19. 201.00
    20. 15.01.00
    21. 01.00
    22. Default1.00
    23. FSDP1.00
    24. Llama 3-8B, 3M tokens, 8xH1000.99
    25. AlEngineer0.97
    26. EUROPE1.00
  • 6:27 #17 done25 line(s)

    shot 17·sharpness 1959.2

    1. DeepSpeed Ulysses1.00
    2. rank1 rank20.98
    3. ranko0.94
    4. rank31.00
    5. q0,0.97
    6. ko0.72
    7. rankO0.94
    8. AIE1.00
    9. q1.0.83
    10. 1.00
    11. 1.00
    12. 1.00
    13. k10.99
    14. rank11.00
    15. All2All0.98
    16. q2,0.91
    17. k21.00
    18. rank21.00
    19. q3,0.99
    20. k31.00
    21. rank31.00
    22. What if we divide by num_heads dim?0.99
    23. DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. Jacobs et al., 20230.99
    24. What1.00
    25. Engineering the future of Al0.99
  • 6:47 #18 skipped

    shot 18·duplicate of #17

  • 7:15 #19 done25 line(s)

    shot 19·sharpness 2256.7

    1. DeepSpeed Ulysses1.00
    2. q0,0.98
    3. ko0.78
    4. o00.79
    5. 1.00
    6. AIE1.00
    7. 1.00
    8. 1.00
    9. 1.00
    10. 92,0.87
    11. k21.00
    12. q1,0.87
    13. k10.99
    14. All2All0.95
    15. Full Flash Attention1.00
    16. All2All0.94
    17. o20.95
    18. o10.95
    19. q3,1.00
    20. k31.00
    21. o30.91
    22. What if we divide by num_heads dim?1.00
    23. DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. Jacobs et al., 20230.99
    24. Whati0.99
    25. Engineering the future of Al0.99
  • 7:44 #20 done30 line(s)

    shot 20·sharpness 2535.9

    1. With Ulysses context parallelism1.00
    2. 801.00
    3. Model1.00
    4. Attn act.1.00
    5. *★*0.61
    6. AIE1.00
    7. (Gi)0.74
    8. 601.00
    9. Other1.00
    10. Peak ory0.70
    11. OOM0.85
    12. OOM0.90
    13. 1.00
    14. 1.00
    15. 1.00
    16. 401.00
    17. OOM0.89
    18. (119)1.00
    19. (7684)1.00
    20. (964)1.00
    21. 201.00
    22. 15.01.00
    23. 15.01.00
    24. 01.00
    25. Default1.00
    26. FSDP1.00
    27. FSDP1.00
    28. + Ulysses0.97
    29. Llama 3-8B, 3M tokens, 8xH1000.98
    30. Engineering the future of Al0.99
  • 8:06 #21 done22 line(s)

    shot 21·sharpness 2791.8

    1. What if the activations are too big?0.99
    2. We need them for the backward pass, but we can recompute0.99
    3. Drawing inspired by https://github.comy/cybertronai/gradlent-checkpointing0.93
    4. 1.00
    5. Forward pass0.99
    6. Orange nodes are the ones1.00
    7. kept in memory to compute1.00
    8. AIE1.00
    9. the gradient update for this0.99
    10. node1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. 1.00
    15. Backward pass1.00
    16. Orange nodes are the ones1.00
    17. kept in memory to compute0.98
    18. the gradient update for this0.99
    19. node1.00
    20. img src: https0.99
    21. Braintrust1.00
    22. WorkOS OpenAI0.98
  • 8:25 #22 done38 line(s)

    shot 22·sharpness 2751.7

    1. With activation checkpointing1.00
    2. 801.00
    3. Model1.00
    4. Attn act.1.00
    5. *★★0.62
    6. AIE1.00
    7. 1.00
    8. (Gi)0.77
    9. 601.00
    10. Other1.00
    11. Pek ory0.75
    12. OOM0.89
    13. OOM0.89
    14. OOM0.93
    15. 1.00
    16. 1.00
    17. 1.00
    18. 401.00
    19. OOM0.90
    20. (119)1.00
    21. (7684)1.00
    22. (964)1.00
    23. (130)1.00
    24. 201.00
    25. 15.01.00
    26. 15.01.00
    27. 15.01.00
    28. 01.00
    29. Default1.00
    30. FSDP1.00
    31. FSDP1.00
    32. FSDP1.00
    33. + Ulysses0.99
    34. + Ulysses0.99
    35. + AC0.98
    36. Llama 3-8B, 3M tokens, 8xH1000.99
    37. AlEngineer0.97
    38. EUROPE1.00
  • 8:35 #23 done20 line(s)

    shot 23·sharpness 3697.2

    1. What if the activations are still too big?0.98
    2. 1. Inputs to each Transformer block can't be recomputed quickly1.00
    3. 2. Still, we can offload to CPU1.00
    4. ***0.62
    5. AIE1.00
    6. 0.99
    7. o0.62
    8. Prefetch before we reach the layer, similar to weights for FSDP1.00
    9. 1.00
    10. 1.00
    11. o0.70
    12. First implemented in:1.00
    13. Blog0.98
    14. Unsloth Gradient1.00
    15. Checkpointing-4x longer0.99
    16. context windows1.00
    17. Apr 9, 2024 · By Daniel & Mike0.97
    18. unsloth.ai/blog/long-context1.00
    19. AlEngineer0.97
    20. EUROPE1.00

Transcript

89 cues· 1,942 words· 11,022 chars

  1. 0:14 Hi everyone, my name is Max.
  2. 0:16 I am VP of Research and Development at Together.ai and today I'm going to tell you about our research project which is called Roll to 5 million sequence length breaking memory barriers in context parallelism.
  3. 0:30 So to begin I'll first say a few words about Together.ai and who we are.
  4. 0:36 Together.ai is an AI-native cloud which provides services and infrastructure for AI developers and builders at all stages of development, starting from creating a model where you just might need a GPU cluster with heavily optimized computes and highly reliable
  5. 0:58 systems all the way through model shaping or you can take existing models and customize them for your tasks in terms of performance, in terms of speed, in terms of quality through services such as the funtioning or reinforcement learning.
  6. 1:15 Also, we are an inference provider, so if you have an app which is reliant on open source model inference, you can work with us and we'll provide you with the fastest way to launch and use AI models.
  7. 1:33 with more than 200 models in our portfolio, options for deployment which include serverless and dedicated inference, and a ton of advanced optimizations which I will not be able to speak about today.
  8. 1:48 The purpose of this talk is focused on model training, customization, and fine-tuning in particular.
  9. 1:56 And I'll start by asking a question.
  10. 2:00 I think in the last few months, or at least a year or so, we're seeing a lot of interest in the community, both on the system side and on the research side, in training long context models.
  11. 2:14 The primary reasons for that are twofold, I would say.
  12. 2:17 First of all, with the explosion in popularity of agents, you can see a lot of different applications where you might want to put as many tokens as you want in your context, and you want the model to leverage that context effectively.
  13. 2:35 Second, with the development of applications such as video generation, you might often need to keep track of multiple frames per second, which can occupy quite a few tokens in your context pretty quickly.
  14. 2:57 And you also need models that have good sense of temporal consistency, which means that they are able to see what was happening a few seconds or ideally a few minutes ago.
  15. 3:10 To do that all effectively, you need to make sure that the models are able to process that context and work with it correctly at the training time.
  16. 3:23 But even if you're not at the scales of millions of tokens in the context length, it's still quite important to understand where the memory goes.
  17. 3:34 Because who knows, maybe you might be able to reinvest it in some other ways and speed up your training overall.
  18. 3:42 So the problem here is that if you are taking a standard transformer-based language model and trying to extend its context, you can run into two bottlenecks.
  19. 3:57 Bottom line number one is that you are faced with quadratic computation because, long story short, for transformer-based models, you have pairwise interactions across all the elements in a sequence.
  20. 4:11 The second problem is more insidious, one might say.
  21. 4:16 As you continue scaling your context, your memory keeps growing linearly, which is not as bad, but still pretty difficult to deal with.
  22. 4:26 unless you apply a range of specific techniques.
  23. 4:31 And this is an example from Hugging Face's blog post on model training, which shows that the sequence length growth can affect your memory limits pretty considerably.
  24. 4:47 Here's the slide.
  25. 4:48 And our goal of that project was to see how far exactly we are able to get with a range of existing techniques that are pretty well known to some in the community as well as some further optimizations that we wanted to leverage to push this a bit further ahead.
  26. 5:11 So let's say you're taking a model which is a standard Lemma 3B architecture, you're trying to fit 3 million trained tokens into your context, and you're taking all of this on an 8x H100 GPU node.
  27. 5:30 The first stage you'll see is that even with just model parameters, you're not able to fit it into the GPU.
  28. 5:40 You run out of memory just by trying to place the model.
  29. 5:43 Of course, the next stage is to apply fully-sharded data parallelism, where all the parameters are basically chunked across the eight GPUs that you have.
  30. 5:53 which is great, but still doesn't solve the problem.
  31. 5:56 You see that the memory usage for the model drops quite significantly, but you still are running out of memory because of all the attention activations.
  32. 6:08 The next point that we've leveraged and I encourage you to use as well is taking advantage of context parallelism.
  33. 6:19 In particular, there is a pretty well-known technique called deep-speed Ulysses, first introduced by Microsoft.
  34. 6:26 The idea is that instead of computing all of your multi-head attention on every GPU separately for the whole sequence, you can do something more clever.
  35. 6:38 In particular, you can try to compute the attention for different heads at different points in time or on different GPUs through communicating these activations as they are required.
  36. 6:54 in such a way that one GPU is only responsible for one attention head here, but it's still computing the attention over the whole sequence.
  37. 7:05 That technique is quite effective at addressing the problem, and it also allows it to utilize the best possible attention implementation, like flash attention 1, 2, 3, 4.
  38. 7:18 to optimize that part of the computation.
  39. 7:22 And then you aggregate the results as you would have previously.
  40. 7:27 So if you apply Ulysses context parallelism, the utilization drops quite significantly, like approximately 8x here, as it should.
  41. 7:39 But we are still quite far from our goal of being able to fit that onto just a single H100 node.
  42. 7:48 So what happens next is that we can try to recompute the activations as they are needed to us at the backward pass.
  43. 7:58 That technique is known as activation checkpointing, and it's available in pretty much all of the deep learning frameworks these days that you could use.
  44. 8:08 You just need to enable it in a correct way that does not impose too much of a computational burden on you.
  45. 8:19 With that, with activation checkpointing, you can drop the activation usage by a further factor of 8, but still something else needs to be done.
  46. 8:31 The next optimization is also connected to the storage of activations.
  47. 8:38 you can try to store some of the inputs to each transformer block, not on the GPU, but instead offload them to CPU when they are not required.
  48. 8:52 This is not very impactful for the performance because you can offload it and prefetch when you are trying to backpropagate to the corresponding layer.
  49. 9:05 This optimization, to the best of our knowledge, was first implemented by Onsloth, and it allows you to drastically expand the context window.
  50. 9:17 The next point is that you're getting with uploading next to 37 gigabytes of data, but then comes the other part of out-of-memory usage.

Open at this second