read-only demo

Videos 2bvtay8wGYI

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning

index_state ready data_status ok

AI Engineer· published 2026-07-31· 0:18:07· en-US· indexed 2026-08-10 19:41

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:32, 1 of 1 keyframes kept
  5. Shot 4, 0:32 to 0:40, 1 of 1 keyframes kept
  6. Shot 5, 0:40 to 1:05, 1 of 1 keyframes kept
  7. Shot 6, 1:05 to 1:08, 1 of 1 keyframes kept
  8. Shot 7, 1:08 to 1:54, 1 of 1 keyframes kept
  9. Shot 8, 1:54 to 2:30, 1 of 1 keyframes kept
  10. Shot 9, 2:30 to 3:06, 0 of 1 keyframes kept
  11. Shot 10, 3:06 to 3:17, 1 of 1 keyframes kept
  12. Shot 11, 3:17 to 3:56, 1 of 1 keyframes kept
  13. Shot 12, 3:56 to 4:27, 1 of 1 keyframes kept
  14. Shot 13, 4:27 to 4:48, 1 of 1 keyframes kept
  15. Shot 14, 4:48 to 4:56, 0 of 1 keyframes kept
  16. Shot 15, 4:56 to 5:11, 0 of 1 keyframes kept
  17. Shot 16, 5:11 to 5:46, 1 of 1 keyframes kept
  18. Shot 17, 5:46 to 6:22, 0 of 1 keyframes kept
  19. Shot 18, 6:22 to 7:06, 1 of 1 keyframes kept
  20. Shot 19, 7:06 to 7:34, 1 of 1 keyframes kept
  21. Shot 20, 7:34 to 8:02, 0 of 1 keyframes kept
  22. Shot 21, 8:02 to 8:51, 1 of 1 keyframes kept
  23. Shot 22, 8:51 to 8:58, 1 of 1 keyframes kept
  24. Shot 23, 8:58 to 9:06, 1 of 1 keyframes kept
  25. Shot 24, 9:06 to 9:40, 1 of 1 keyframes kept
  26. Shot 25, 9:40 to 10:00, 1 of 1 keyframes kept
  27. Shot 26, 10:00 to 10:31, 1 of 1 keyframes kept
  28. Shot 27, 10:31 to 10:55, 1 of 1 keyframes kept
  29. Shot 28, 10:55 to 11:16, 1 of 1 keyframes kept
  30. Shot 29, 11:16 to 11:51, 1 of 1 keyframes kept
  31. Shot 30, 11:51 to 12:23, 1 of 1 keyframes kept
  32. Shot 31, 12:23 to 12:31, 1 of 1 keyframes kept
  33. Shot 32, 12:31 to 12:57, 1 of 1 keyframes kept
  34. Shot 33, 12:57 to 13:14, 1 of 1 keyframes kept
  35. Shot 34, 13:14 to 13:33, 0 of 1 keyframes kept
  36. Shot 35, 13:33 to 13:59, 0 of 1 keyframes kept
  37. Shot 36, 13:59 to 14:26, 0 of 1 keyframes kept
  38. Shot 37, 14:26 to 14:33, 1 of 1 keyframes kept
  39. Shot 38, 14:33 to 15:01, 1 of 1 keyframes kept
  40. Shot 39, 15:01 to 15:29, 0 of 1 keyframes kept
  41. Shot 40, 15:29 to 15:50, 1 of 1 keyframes kept
  42. Shot 41, 15:50 to 16:15, 1 of 1 keyframes kept
  43. Shot 42, 16:15 to 16:46, 1 of 1 keyframes kept
  44. Shot 43, 16:46 to 16:48, 1 of 1 keyframes kept
  45. Shot 44, 16:48 to 17:15, 0 of 1 keyframes kept
  46. Shot 45, 17:15 to 17:43, 0 of 1 keyframes kept
  47. Shot 46, 17:43 to 17:50, 1 of 1 keyframes kept
  48. Shot 47, 17:50 to 18:06, 0 of 1 keyframes kept

48 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
224
whisperx 224
chunks
31
from 224 cues
keyframes
36
kept of 48 captured
frames with text
36
624 lines read
chapters
11
from the source metadata
keyframe bytes
6.7 MB
word timings on 224 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:09 0s
stt done 2026-08-09 13:27 26s
chunk done 2026-08-09 13:27 0s
text_embed done 2026-08-10 19:41 1s
keyframe done 2026-08-09 13:27 2m 22s
ocr done 2026-08-09 13:29 13s
frame_embed done 2026-08-10 19:41 6s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 457.1

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 663.6

    1. AlEngineer0.96
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2741.5

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.93
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:22 #3 done2 line(s)

    shot 3·sharpness 326.9

    1. AlEngineer0.99
    2. World's Fair0.94
  • 0:36 #4 done12 line(s)

    shot 4·sharpness 2779.8

    1. General Reasoning1.00
    2. @GenReasoning1.00
    3. AlEngineer0.98
    4. World'sFair1.00
    5. PRESENTED BY1.00
    6. Scaling to Long Horizons1.00
    7. Microsoft1.00
    8. Algorithms, Environments, Compute1.00
    9. Ross Taylor1.00
    10. Chengxi Taylor1.00
    11. World'sFair1.00
    12. Engineering the future of Al1.00
  • 0:55 #5 done11 line(s)

    shot 5·sharpness 3488.9

    1. What you'll hear today1.00
    2. AlEngineer0.97
    3. World'sFair0.96
    4. The Journey So Far - a personal perspective0.98
    5. PRESENTED BY1.00
    6. Galactica, Llama, early RL efforts for LLMs0.99
    7. Microsoft1.00
    8. The Journey Ahead - a company perspective0.99
    9. What does it take to scale to long-horizon?0.99
    10. World's Fair0.97
    11. Engineering the future of Al0.99
  • 1:06 #6 done8 line(s)

    shot 6·sharpness 1454.3

    1. AlEngineer0.96
    2. World'sFair1.00
    3. PRESENTED BY1.00
    4. Microsoft1.00
    5. The Journey So Far0.98
    6. World'sFair1.00
    7. TRACK 9· JUNE 30, 20260.96
    8. Data Quality1.00
  • 1:13 #7 done16 line(s)

    shot 7·sharpness 2976.5

    1. My AI journey started here1.00
    2. AlEngineer0.97
    3. World'sFair1.00
    4. Papers with Code1.00
    5. Galactica1.00
    6. PRESENTED BY1.00
    7. Llama 21.00
    8. Microsoft1.00
    9. Llama 31.00
    10. ReasoningLlama1.00
    11. More recent things!1.00
    12. The Open Weight Revolution1.00
    13. started in this room!1.00
    14. World's Fair0.94
    15. TRACK 9• JUNE 30, 20260.94
    16. Data Quality1.00
  • 1:58 #8 done17 line(s)

    shot 8·sharpness 2985.1

    1. Things got crazy in 20221.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. OpenAI0.91
    5. GALACTICA1.00
    6. ChatGPT1.00
    7. November 30 20221.00
    8. November 15 20221.00
    9. Galactica LLM paper/demo released0.98
    10. ChatGPT product released.0.99
    11. SoTA base model, but failed demo0.99
    12. SoTA RLHF pipeline makes LLMs1.00
    13. due to base model hallucinations.0.98
    14. useful for the first time.1.00
    15. World's Fair0.96
    16. TRACK 9· JUNE 30, 20260.97
    17. Data Quality1.00
  • 2:58 #9 skipped

    shot 9·duplicate of #8

  • 3:14 #10 done19 line(s)

    shot 10·sharpness 2820.2

    1. A good base model = not enough1.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. OpenAI0.92
    5. ChatGPT1.00
    6. November 30 20221.00
    7. November 151.00
    8. Galactica LLM1.00
    9. demo released1.00
    10. ChatGPT product released.0.99
    11. SoTA b1.00
    12. demo1.00
    13. SoTA RLHF pipeline makes LLMs1.00
    14. due to0.99
    15. ns.1.00
    16. useful for the first time.1.00
    17. World's Fair0.95
    18. TRACK 9• JUNE 30, 20260.95
    19. Data Quality1.00
  • 3:25 #11 done21 line(s)

    shot 11·sharpness 2283.4

    1. RLHF made LLMs products0.99
    2. AlEngineer0.99
    3. World's Fair0.99
    4. Wit tt 5B0.62
    5. 0.61.00
    6. Model0.91
    7. PPO-ptx1.00
    8. PPO1.00
    9. 0.41.00
    10. SFT1.00
    11. GPT (prompted)0.98
    12. GPT1.00
    13. 0.21.00
    14. 1.3B1.00
    15. 6B1.00
    16. 175B1.00
    17. Model size1.00
    18. InstructGPT - https://arxiv.org/pdf/2203.021550.99
    19. World's Fair0.96
    20. TRACK 9· JUNE 30, 20260.97
    21. Data Quality0.98
  • 4:06 #12 done9 line(s)

    shot 12·sharpness 3982.5

    1. AlEngineer0.97
    2. Galactica demo set off a storm0.98
    3. World'sFair1.00
    4. We put out a demo to let people play with the base model.1.00
    5. People got scared and saw it as a “fake science” machine.0.98
    6. Meta association didn't help us either0.98
    7. Much of the novel work in the paper was overshadowed1.00
    8. TRACK 9· JUNE 30,20260.94
    9. Data Quality0.98
  • 4:41 #13 done112 line(s)

    shot 13·sharpness 6529.7

    1. But it was a good model, sir0.98
    2. AlEngineer0.99
    3. World'sFair1.00
    4. Galactica outperformed PaLM/Chinchilla/GPT-3.5 with much0.99
    5. less compute (2-10x) in scientific domains. It was SoTA (!)0.98
    6. Mathematics MMLU1.00
    7. Model1.00
    8. Params (bn)0.99
    9. A.Algebra1.00
    10. Elem1.00
    11. HS1.00
    12. College1.00
    13. F. Logic1.00
    14. Average1.00
    15. BLOOM (5-shot)1.00
    16. 1761.00
    17. 25.0%1.00
    18. 26.7%1.00
    19. 27.0%1.00
    20. 25.0%1.00
    21. 26.2%1.00
    22. 26.4%1.00
    23. OPT (5-shot)1.00
    24. 1751.00
    25. 21.0%1.00
    26. 25.7%0.96
    27. 24.4%1.00
    28. 33.0%1.00
    29. 29.4%1.00
    30. 26.7%1.00
    31. Gopher (5-shot)1.00
    32. 2801.00
    33. 25.0%1.00
    34. 33.6%1.00
    35. 23.7%1.00
    36. 37.0%1.00
    37. 35.7%1.00
    38. 30.6%1.00
    39. Chinchilla (5-shot)0.99
    40. 701.00
    41. 31.0%1.00
    42. 41.5%1.00
    43. 31.9%1.00
    44. 32.0%1.00
    45. 33.3%1.00
    46. 35.7%1.00
    47. GAL 1.3B1.00
    48. 1.31.00
    49. 28.0%1.00
    50. 27.2%1.00
    51. 26.7%1.00
    52. 30.0%0.99
    53. 24.6%1.00
    54. 27.1%1.00
    55. GAL 6.7B1.00
    56. 6.71.00
    57. 28.0%1.00
    58. 28.9%1.00
    59. 26.7%1.00
    60. 36.0%1.00
    61. 31.0%1.00
    62. 29.2%1.00
    63. GAL 30B1.00
    64. 301.00
    65. 30.0%1.00
    66. 30.2%1.00
    67. 26.3%1.00
    68. 36.0%1.00
    69. 31.7%1.00
    70. 29.9%1.00
    71. GAL 120B1.00
    72. 1201.00
    73. 33.0%1.00
    74. 38.1%1.00
    75. 32.6%1.00
    76. 43.0%1.00
    77. 32.5%1.00
    78. 35.8%1.00
    79. GAL 1.3B <work>0.98
    80. 1.31.00
    81. 22.0%1.00
    82. 24.6%1.00
    83. 18.9%1.00
    84. 25.0%1.00
    85. 31.0%1.00
    86. 24.6%1.00
    87. GAL 6.7B <work>0.97
    88. 6.71.00
    89. 33.3%1.00
    90. 30.7%1.00
    91. 25.2%1.00
    92. 26.0%1.00
    93. 33.3%1.00
    94. 28.0%1.00
    95. GAL 30B <work>0.99
    96. 301.00
    97. 33.0%1.00
    98. 41.5%0.98
    99. 33.3%1.00
    100. 39.0%1.00
    101. 37.3%1.00
    102. 37.1%1.00
    103. GAL 120B <work>0.99
    104. 1201.00
    105. 27.0%1.00
    106. 54.2%1.00
    107. 37.0%1.00
    108. 44.0%1.00
    109. 40.5%1.00
    110. 41.3%1.00
    111. TRACK 9· JUNE 30, 20260.93
    112. Data Quality0.96
  • 4:53 #14 skipped

    shot 14·duplicate of #13

  • 4:59 #15 skipped

    shot 15·duplicate of #13

  • 5:35 #16 done56 line(s)

    shot 16·sharpness 3684.9

    1. Galactica introduced key ideas0.99
    2. AlEngineer1.00
    3. World'sFair1.00
    4. 3.501.00
    5. Question: A needle 35 mm long rests on a water surface at 20°C. What force over and above the needle's weight0.99
    6. GAL 125M1.00
    7. is required to lift the needle from contact with the water surface? σ = 0.0728m.0.99
    8. 3.251.00
    9. GAL 1.381.00
    10. GAL 6.780.99
    11. <work>0.99
    12. 3.001.00
    13. GAL 120B0.93
    14. GAL 3081.00
    15. σ = 0.0728 N/m0.97
    16. PRESENTED BY1.00
    17. Vvali Los0.69
    18. 2.751.00
    19. 2.501.00
    20. calculate.py1.00
    21. 0.0728 = F/(2 × 0.035)0.98
    22. σ = F/L0.92
    23. F = 0.0728(2 × 0.035)0.94
    24. Microsoft1.00
    25. 2.251.00
    26. f = 0.0728*(2+0.035)0.97
    27. 2.001.00
    28. with open("output.txt", "w") as file:0.96
    29. file.write(atr(roumd(f, 5)))0.96
    30. 1.751.00
    31. "run: "calculate.py">0.96
    32. 1.500.90
    33. 501.00
    34. 1001.00
    35. 1501.00
    36. 2001.00
    37. 2501.00
    38. 3001.00
    39. 3501.00
    40. 4001.00
    41. 4501.00
    42. «read: "output.txt"»0.95
    43. Tokens (Billions)0.98
    44. 0.00511.00
    45. Figure 6: Repeated Tokens and Validation Loss. With four epochs of training, we continue to see validation0.99
    46. </vork>0.86
    47. loss fall for all model sizes. For the 120B model we see the first signs of overfitting at the beginning of the0.98
    48. fifth epoch, and we early stop at this point.0.98
    49. Answer: F = 0.0051 N0.97
    50. First major LLM to crack data-efficiency0.99
    51. Introduced thinking tokens0.99
    52. and multi-epoch training0.98
    53. <work></work> to spend inference before1.00
    54. giving an answer1.00
    55. TRACK 9· JUNE 30,20260.96
    56. Data Quality1.00
  • 5:54 #17 skipped

    shot 17·duplicate of #16

  • 6:27 #18 done11 line(s)

    shot 18·sharpness 3971.9

    1. AlEngineer0.97
    2. My goal after Galactica: solve reasoning1.00
    3. World'sFair1.00
    4. What if we applied RL pressure to <work></work> ?0.99
    5. PRESENTED BY1.00
    6. Microsoft1.00
    7. Sound familiar? This is the same DeepSeek-R1 approach.0.99
    8. But we had Llama 2 base models in 2023.0.99
    9. And the context window was 4096. Not fun.1.00
    10. TRACK 9• JUNE 30,20260.94
    11. Data Quality0.99
  • 7:20 #19 done15 line(s)

    shot 19·sharpness 5007.9

    1. This was our 2023 RL Recipe0.99
    2. AlEngineer0.97
    3. World'sFair1.00
    4. Continued pre-training on Llama 2 towards mathematics0.98
    5. and science data.1.00
    6. PPO with verifiable rewards and strong ORM to initialise0.99
    7. the value model.1.00
    8. SoTA on MATH; beat GPT-4! But no inference-time scaling1.00
    9. unlike o1/R1. Why!?0.98
    10. Work with Iliyan Zarov, Anthony Hartshorn, Guillem Cucurull,1.00
    11. Lukas Blecher, Rui Hou and others.1.00
    12. Unpublished:(1.00
    13. World'sFair0.98
    14. TRACK 9· JUNE 30,20260.94
    15. Data Quality0.99
  • 7:40 #20 skipped

    shot 20·duplicate of #19

  • 8:17 #21 done35 line(s)

    shot 21·sharpness 4839.0

    1. Better base models got RL cooking1.00
    2. AlEngineer0.99
    3. World'sFair1.00
    4. When DeepSeek-R1 came out, I was shocked!1.00
    5. We'd tried essentially the same recipe before and it didn't scale like that1.00
    6. b1.00
    7. DeepSeek-R1-Zero average length per response during training0.99
    8. 20,0001.00
    9. I'm a little surprised as when we did PPO on0.99
    10. Llama-2 2 years ago for reasoning, we didn't see1.00
    11. 17,5001.00
    12. this kind of emergence. The winning recipe seems1.00
    13. Avers eense0.63
    14. 15,0001.00
    15. to be just a better base model, larger context, and0.98
    16. 12,5001.00
    17. more RL compute with a diverse set of prompts.0.98
    18. 10,0001.00
    19. Staggeringly simple, and another win for the1.00
    20. 7,5001.00
    21. scaling hypothesis.0.99
    22. 5,0001.00
    23. Also shows the corollary of the bitter lesson:1.00
    24. 2,5001.00
    25. standing on the shoulders of a better base model1.00
    26. allows you to see further. I think we would have0.99
    27. 2,0001.00
    28. 4,0001.00
    29. 6,0001.00
    30. 8,0001.00
    31. 10,0001.00
    32. struggled to make this work with pre L3 models!1.00
    33. Steps1.00
    34. TRACK 9· JUNE 30, 20260.94
    35. Data Quality1.00
  • 8:56 #22 done7 line(s)

    shot 22·sharpness 1808.5

    1. Defeated by ChatGPT and o11.00
    2. AlEngineer0.98
    3. World'sFair1.00
    4. Ross in0.98
    5. 20241.00
    6. TRACK 9· JUNE 30, 20260.97
    7. Data Quality1.00
  • 8:59 #23 done7 line(s)

    shot 23·sharpness 2713.6

    1. But I wanted the next wave!0.98
    2. AlEngineer0.97
    3. World'sFair1.00
    4. What does it take to build truly general reasoning systems that0.99
    5. can take on civilisation-scale tasks?1.00
    6. TRACK 9· JUNE 30, 20260.95
    7. Data Quality0.98

Transcript

224 cues· 3,038 words· 16,636 chars

  1. 0:12 So this talk is called Scaling to Long Horizons.
  2. 0:15 My name's Ross.
  3. 0:16 I'm the CEO of GR.
  4. 0:17 We're a London-based reinforcement learning company.
  5. 0:20 Before GR, I was the reasoning lead at Meta AI, working on Llamas, Galactica, lots of other models back in the day.
  6. 0:27 I'm joined by Cheng Shi, co-founder and president of GR.
  7. 0:30 And yeah, today we're gonna talk about algorithms, environments, compute, all the things you need to do to get agents scaling to kind of longer tasks.
  8. 0:41 So we're gonna have two parts to this talk today.
  9. 0:42 I'm gonna first of all start with a personal perspective about the early days, the golden age of language modeling between like maybe 2020 and 2023.
  10. 0:52 I'll talk about, like I said, all those models and some of our early reinforcement learning efforts for LLMs.
  11. 0:56 And then Chengxi is gonna talk about what's ahead, what are the next frontiers.
  12. 1:00 And yeah, that's gonna be a really interesting talk with a lot of alpha, so I'd encourage you to stick around for that.
  13. 1:06 So the journey so far.
  14. 1:07 So my journey started here.
  15. 1:10 So this was the Papers with Code team.
  16. 1:12 I'm sure many of you used Papers with Code back in the day.
  17. 1:15 So we were a London-based startup, 2019.
  18. 1:18 We were acquired by Meta later that year.
  19. 1:20 And then we had a crazy transition within Meta to do research.
  20. 1:24 So we did, like I said, Galactica.
  21. 1:26 Then after ChatDVT came out, we started the post-training for Llama 2, Llama 3, so all the great work you saw there was folks in this room.
  22. 1:34 And lots of other interesting stuff that never got published as well, Reasoning Llama and lots of other things.
  23. 1:39 So yeah, this small team, I like to think like the open weight kind of revolution started in this room, and it really hit home, this idea to me that kind of small, focused teams, even in the age of scaling, can do amazing things if people are aligned.
  24. 1:55 Now for me, things got particularly crazy in 2022.
  25. 1:59 So let me tell you a story.
  26. 2:01 The media perception is that ChatGPT came out of nowhere, shocked the world, and that's how the modern AI wave started.
  27. 2:10 But I have a different personal perspective on this because two weeks before ChatGPT came along, there was another language model called Galactica.
  28. 2:18 So let's talk about Galactica.
  29. 2:20 Galactica and ChatGPT, they were both, in some respects, quite similar.
  30. 2:25 They were both based on pretty good base models.
  31. 2:27 Galactica itself was a base model.
  32. 2:29 And then ChatGPT was based on GPT 3.5.
  33. 2:33 But there was a clear difference in outcomes.
  34. 2:35 So Galactica at the time shipped with this base model demo.
  35. 2:39 And as you guys know now, base models, they come with a lot of quirks.
  36. 2:42 They hallucinate.
  37. 2:44 You prompt them to do silly things, they will do silly things.
  38. 2:47 Whereas Chat-TPT wasn't just a base model, but had this crucial reinforcement learning from human feedback pipeline.
  39. 2:54 And this was the key thing that made LLMs really products for the first time.
  40. 2:58 So I like to think, in a weird kind of way, this is the first natural experiment showing you that RL provides value.
  41. 3:06 And to my misfortune, it was like a very personal kind of natural experiment, and Galactica blew up.
  42. 3:11 But that's like a good lesson there.
  43. 3:13 A good base model is not enough.
  44. 3:15 So I took that lesson quite early on.
  45. 3:18 So like I said, RLHF made LLMs products.
  46. 3:21 They were the thing that kind of made LLMs cross the Rubicon into something that wasn't just a toy, but used by now billions of people.
  47. 3:27 But you didn't have to wait until ChatGPT to see this.
  48. 3:30 Like even at the time, like InstructGPT in 2022 had these pretty stunning results.
  49. 3:35 Like a one billion parameter model with RLHF was outperforming 175 billion models.
  50. 3:41 two orders of magnitude fewer parameters, but getting better results.

Chapters

  1. 0:00 Introduction and a look back
  2. 1:57 The Galactica story
  3. 5:15 Curated data and thinking tokens
  4. 8:09 What got RL cooking
  5. 9:12 Long horizon as a mindset
  6. 10:16 Why value models help
  7. 11:08 Credit assignment and bootstrapping
  8. 12:38 Trading football matches for real money
  9. 13:44 Why models struggled
  10. 14:36 Off policy staleness versus GPU use
  11. 16:18 openreward.ai and what's next

Open at this second