read-only demo

Videos 3hXJI2q0Jz8

Recursive Coding Agents - Raymond Weitekamp, OpenProse

index_state ready data_status ok

AI Engineer· published 2026-06-25· 0:23:48· en-US· indexed 2026-08-10 22:51

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:27, 1 of 1 keyframes kept
  2. Shot 1, 0:27 to 0:57, 1 of 1 keyframes kept
  3. Shot 2, 0:57 to 1:26, 0 of 1 keyframes kept
  4. Shot 3, 1:26 to 1:55, 1 of 1 keyframes kept
  5. Shot 4, 1:55 to 2:41, 1 of 1 keyframes kept
  6. Shot 5, 2:41 to 3:13, 1 of 1 keyframes kept
  7. Shot 6, 3:13 to 3:46, 0 of 1 keyframes kept
  8. Shot 7, 3:46 to 4:18, 0 of 1 keyframes kept
  9. Shot 8, 4:18 to 4:48, 1 of 1 keyframes kept
  10. Shot 9, 4:48 to 5:17, 1 of 1 keyframes kept
  11. Shot 10, 5:17 to 5:46, 0 of 1 keyframes kept
  12. Shot 11, 5:46 to 6:15, 0 of 1 keyframes kept
  13. Shot 12, 6:15 to 6:44, 0 of 1 keyframes kept
  14. Shot 13, 6:44 to 7:13, 0 of 1 keyframes kept
  15. Shot 14, 7:13 to 7:42, 1 of 1 keyframes kept
  16. Shot 15, 7:42 to 8:10, 1 of 1 keyframes kept
  17. Shot 16, 8:10 to 8:38, 0 of 1 keyframes kept
  18. Shot 17, 8:38 to 9:07, 0 of 1 keyframes kept
  19. Shot 18, 9:07 to 9:35, 1 of 1 keyframes kept
  20. Shot 19, 9:35 to 10:03, 1 of 1 keyframes kept
  21. Shot 20, 10:03 to 10:31, 0 of 1 keyframes kept
  22. Shot 21, 10:31 to 10:43, 1 of 1 keyframes kept
  23. Shot 22, 10:43 to 11:20, 1 of 1 keyframes kept
  24. Shot 23, 11:20 to 11:50, 1 of 1 keyframes kept
  25. Shot 24, 11:50 to 12:19, 0 of 1 keyframes kept
  26. Shot 25, 12:19 to 12:48, 0 of 1 keyframes kept
  27. Shot 26, 12:48 to 13:18, 1 of 1 keyframes kept
  28. Shot 27, 13:18 to 13:47, 1 of 1 keyframes kept
  29. Shot 28, 13:47 to 14:12, 1 of 1 keyframes kept
  30. Shot 29, 14:12 to 14:37, 1 of 1 keyframes kept
  31. Shot 30, 14:37 to 15:02, 0 of 1 keyframes kept
  32. Shot 31, 15:02 to 15:27, 0 of 1 keyframes kept
  33. Shot 32, 15:27 to 15:53, 1 of 1 keyframes kept
  34. Shot 33, 15:53 to 16:20, 0 of 1 keyframes kept
  35. Shot 34, 16:20 to 16:47, 1 of 1 keyframes kept
  36. Shot 35, 16:47 to 17:15, 0 of 1 keyframes kept
  37. Shot 36, 17:15 to 17:41, 1 of 1 keyframes kept
  38. Shot 37, 17:41 to 18:07, 0 of 1 keyframes kept
  39. Shot 38, 18:07 to 18:33, 0 of 1 keyframes kept
  40. Shot 39, 18:33 to 19:00, 1 of 1 keyframes kept
  41. Shot 40, 19:00 to 19:27, 1 of 1 keyframes kept
  42. Shot 41, 19:27 to 19:54, 0 of 1 keyframes kept
  43. Shot 42, 19:54 to 20:22, 1 of 1 keyframes kept
  44. Shot 43, 20:22 to 20:51, 0 of 1 keyframes kept
  45. Shot 44, 20:51 to 21:20, 1 of 1 keyframes kept
  46. Shot 45, 21:20 to 21:49, 0 of 1 keyframes kept
  47. Shot 46, 21:49 to 22:16, 1 of 1 keyframes kept
  48. Shot 47, 22:16 to 22:42, 0 of 1 keyframes kept
  49. Shot 48, 22:42 to 23:09, 0 of 1 keyframes kept
  50. Shot 49, 23:09 to 23:36, 0 of 1 keyframes kept
  51. Shot 50, 23:36 to 23:47, 1 of 1 keyframes kept

51 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
216
whisperx 216
chunks
44
from 216 cues
keyframes
27
kept of 51 captured
frames with text
27
528 lines read
chapters
0
from the source metadata
keyframe bytes
4.0 MB
word timings on 216 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 22:47 1m 32s
stt done 2026-08-10 22:49 24s
chunk done 2026-08-10 22:49 0s
text_embed done 2026-08-10 22:49 0s
keyframe done 2026-08-10 22:49 1m 14s
ocr done 2026-08-10 22:51 13s
frame_embed done 2026-08-10 22:51 5s

Frames, and what the machine read

  • 0:05 #0 done6 line(s)

    shot 0·sharpness 621.1

    1. AI ENGINEER WORLD'S FAIR· 20260.99
    2. Recursive Coding Agents1.00
    3. Raymond Weitekamp1.00
    4. RAW.works | OpenProse0.96
    5. @raw_works1.00
    6. recursivecodingagents.com1.00
  • 0:31 #1 done8 line(s)

    shot 1·sharpness 890.8

    1. MOTIVATION1.00
    2. We all want outcomes.1.00
    3. Agents that work on our behalf — reliable co-workers — while we're out on a hike.0.99
    4. The bottleneck is not intelligence. It's reliability. It's trust.0.99
    5. One day — my agents build me a full SaaS app from a single prompt.0.99
    6. The next day — they empty the entire contents of my Solana wallet.0.98
    7. A0.62
    8. recursivecodingagents.com1.00
  • 1:14 #2 skipped

    shot 2·duplicate of #1

  • 1:40 #3 done9 line(s)

    shot 3·sharpness 937.0

    1. MOTIVATION1.00
    2. We all want outcomes.1.00
    3. Agents that work on our behalf — reliable co-workers — while we're out on a hike.0.99
    4. The bottleneck is not intelligence. It's reliability. It's trust.0.99
    5. One day — my agents build me a full SaaS app from a single prompt.0.99
    6. The next day — they empty the entire contents of my Solana wallet.0.98
    7. A0.77
    8. CODE1.00
    9. recursivecodingagents.com1.00
  • 2:31 #4 done14 line(s)

    shot 4·sharpness 1271.2

    1. THESIS1.00
    2. Today's agents are1.00
    3. mismanaged geniuses1.00
    4. The intelligence is there.1.00
    5. The missing layer is how we specify, manage, reuse, and verify the work.1.00
    6. The Mismanaged Geniuses Hypothesis1.00
    7. Stop Babysitting Agents, Start0.98
    8. Authoring Outcomes0.98
    9. 00000000.60
    10. TURING POST - RAYMOND WEITEKAMP0.98
    11. ALEX ZHANG - ZED LI- OMAR KHATTAB0.96
    12. Stop Babysitting Agents, Start Authoring Outcomes0.99
    13. The Mismanaged Geniuses Hypothesis1.00
    14. recursivecodingagents.com1.00
  • 2:57 #5 done16 line(s)

    shot 5·sharpness 1203.4

    1. RECURSIVE LANGUAGE MODELS1.00
    2. Context itself is the0.99
    3. object of computation0.98
    4. Recursive Language Models1.00
    5. Externalize — the full prompt lives in a REPL,1.00
    6. not the context window.0.98
    7. Operate — the model writes code to inspect0.98
    8. slice, and transform it.1.00
    9. Recurse — it sub-queries itself over the slices.0.99
    10. ARXIV:2512.246011.00
    11. Root RLM (depth=0)0.99
    12. Sub-RLM A (depth=1)0.98
    13. Sub-RLM B (depth=1)1.00
    14. LLM B1 (depth=2)0.97
    15. - LLM B2 (depth=2)0.98
    16. recursivecodingagents.com1.00
  • 3:42 #6 skipped

    shot 6·duplicate of #5

  • 3:56 #7 skipped

    shot 7·duplicate of #5

  • 4:30 #8 done22 line(s)

    shot 8·sharpness 1263.6

    1. RECURSIVE LANGUAGE MODELS1.00
    2. Code1.00
    3. Execution As1.00
    4. Reasoning1.00
    5. Can process inputs way beyond the context window.0.97
    6. RLMs are the new reasoning models0.99
    7. (Oolong) RLM paper1.00
    8. can scalo wit tes-üime compute. Recursve language models (RLMs) ask0.86
    9. Reasoning models were the first clear proof that language model capability0.97
    10. RLM is itself a powerful memory system.1.00
    11. 2:50 PM - Apr 20, 2026 - 64.2KViews0.95
    12. LongMemEval results1.00
    13. Q160.75
    14. t0.81
    15. 5110.77
    16. 日860.69
    17. 0.56
    18. Read 16 replies0.98
    19. RLM can achieve SOTA on long reasoning tasks,0.99
    20. even with very small models. LongCoT results1.00
    21. RLMs are the new reasoning models.1.00
    22. recursivecodingagents.com1.00
  • 4:59 #9 done23 line(s)

    shot 9·sharpness 1279.2

    1. RECURSIVE LANGUAGE MODELS1.00
    2. Code1.00
    3. Execution As1.00
    4. nond Weitekamp0.98
    5. Reasoning1.00
    6. Can process inputs way beyond the context window.0.97
    7. RLMs are the new reasoning models1.00
    8. (Oolong) RLM paper1.00
    9. can scalo wit tes-üime compute. Recursve languago models (RLMs) ask0.88
    10. Reasoning models were the first clear proof that language model capability0.96
    11. RLM is itself a powerful memory system.1.00
    12. 2:50 PM - Apr 20, 2026 - 66.2KViews0.95
    13. LongMemEval results1.00
    14. Q160.71
    15. t70.72
    16. 5110.85
    17. 8610.87
    18. Read 16 replies0.97
    19. RLM can achieve SOTA on long reasoning tasks,0.99
    20. even with very small models. LongCoT results1.00
    21. RLMs are the new reasoning models.1.00
    22. AlEnaineer0.98
    23. recursivecodingagents.com1.00
  • 5:40 #10 skipped

    shot 10·duplicate of #8

  • 6:09 #11 skipped

    shot 11·duplicate of #8

  • 6:24 #12 skipped

    shot 12·duplicate of #8

  • 6:56 #13 skipped

    shot 13·duplicate of #8

  • 7:20 #14 done35 line(s)

    shot 14·sharpness 1491.3

    1. RLMs: Too Hot To Benchmark0.98
    2. Raymond Weitekamp1.00
    3. x0.94
    4. Sumeet Motwani1.00
    5. x0.90
    6. @raw_works - Follow0.94
    7. @sumeetrm - Follow0.96
    8. i'm super confused...1.00
    9. LongCoT is adding two new leaderboards! Due to the1.00
    10. interest in agents (particularly RLMs), we're adding a0.99
    11. ...as far as i can tell, this a "consolation tweet" implying0.99
    12. "Restricted Harness" and an "Open Harness" leaderboard.0.99
    13. that ARC Prize will not verify the symbolica agent despite1.00
    14. their insane score...1.00
    15. at 25.12%. We expect tool-use SOTA to exceed this very1.00
    16. GPT 5.2 RLM from our paper is SOTA on "Open Harness"0.99
    17. ...at least that's how i'm interpreting the community0.99
    18. soon!1.00
    19. leaderboard reminder...0.99
    20. On Show more1.00
    21. ...am i totally mis-reading this?1.00
    22. LongCoT Open Harness Leaderboard1.00
    23. ARC Prize@arcprize0.96
    24. Today's @symbolica hamess is a clear example of what human-crafted0.99
    25. targeting can achieve on ARC-AGI-3 public demo set0.98
    26. You can "buy" performance with benchmark-specific prompts/strategies0.99
    27. Their approach coul stll contain useful ideas, excited to see what the0.98
    28. community finds0.99
    29. 10:47 AM · Mar 27, 20260.95
    30. Reply0.99
    31. Copy link0.95
    32. Read more on X1.00
    33. I personally do not care if my Al programs do their reasoning in latent space or code.0.99
    34. I want results.1.00
    35. recursivecodingagents.com1.00
  • 8:07 #15 done35 line(s)

    shot 15·sharpness 1497.8

    1. RLMs: Too Hot To Benchmark0.98
    2. Raymond Weitekamp1.00
    3. x0.95
    4. Sumeet Motwani1.00
    5. x0.91
    6. @raw_works - Follow0.94
    7. @sumeetrm - Follow0.93
    8. i'm super confused...0.99
    9. LongCoT is adding two new leaderboards! Due to the0.99
    10. interest in agents (particularly RLMs), we're adding a0.99
    11. ...as far as i can tell, this a "consolation tweet" implying0.99
    12. "Restricted Harness" and an "Open Harness" leaderboard.0.98
    13. that ARC Prize will not verify the symbolica agent despite1.00
    14. their insane score...1.00
    15. at 25.12%. We expect tool-use SOTA to exceed this very1.00
    16. GPT 5.2 RLM from our paper is SOTA on "Open Harness"0.99
    17. ...at least that's how i'm interpreting the community0.99
    18. soon!1.00
    19. leaderboard reminder...0.99
    20. On Show more1.00
    21. ...am i totally mis-reading this?0.99
    22. LongCoT Open Harness Leaderboard1.00
    23. ARC Prize@arcprize0.98
    24. Today's @symbolica hamness is a clear example of what human-crafted0.99
    25. targeting can achieve on ARC-AGI-3 public demo set0.98
    26. You can "buy" performance with benchmark-specific prompts/strategies0.99
    27. Their approach coul stil contain useful ideas, excited to see what the0.97
    28. community finds1.00
    29. 10:47 AM · Mar 27, 20260.95
    30. Reply1.00
    31. Copy link0.95
    32. Read more on X0.97
    33. I personally do not care if my Al programs do their reasoning in latent space or code.0.99
    34. I want results.1.00
    35. recursivecodingagents.com1.00
  • 8:16 #16 skipped

    shot 16·duplicate of #14

  • 8:55 #17 skipped

    shot 17·duplicate of #14

  • 9:23 #18 done22 line(s)

    shot 18·sharpness 596.7

    1. THE RLM RUBRIC0.97
    2. Lots of things feel close.1.00
    3. Executable1.00
    4. Prompt1.00
    5. Code calls0.97
    6. Model picks0.99
    7. State stays1.00
    8. environment1.00
    9. externalized1.00
    10. the model1.00
    11. decomposition1.00
    12. symbolic1.00
    13. Plain long-context call0.98
    14. RAG / reasoning-only0.98
    15. Coding agents + subagents0.98
    16. including loops1.00
    17. Hardcoded map-reduce1.00
    18. developer-authored pipeline — e.g. λ-RLM0.99
    19. Recursive Language Model1.00
    20. passes every check0.99
    21. Open the RLM rubric1.00
    22. recursivecodingagents.com1.00
  • 9:44 #19 done22 line(s)

    shot 19·sharpness 612.1

    1. THE RLM RUBRIC0.97
    2. Lots of things feel close.1.00
    3. Executable1.00
    4. Prompt1.00
    5. Code calls0.97
    6. Model picks0.99
    7. State stays1.00
    8. environment1.00
    9. externalized1.00
    10. the model1.00
    11. decomposition1.00
    12. symbolic1.00
    13. Plain long-context call1.00
    14. RAG / reasoning-only0.95
    15. Coding agents + subagents1.00
    16. including loops1.00
    17. Hardcoded map-reduce1.00
    18. developer-authored pipeline — e.g. λ-RLM0.98
    19. Recursive Language Model1.00
    20. passes every check1.00
    21. Open the RLM rubric1.00
    22. recursivecodingagents.com1.00
  • 10:27 #20 skipped

    shot 20·duplicate of #19

  • 10:36 #21 done16 line(s)

    shot 21·sharpness 802.3

    1. TOWARDS RECURSIVE CODING AGENTS0.99
    2. RLM/LLM1.00
    3. Agent / Sub-Agent0.96
    4. Root RLM (depth=0)0.99
    5. Root Agent (depth=0)1.00
    6. Sub-RLM A (depth=1)0.99
    7. Sub-Agent A (depth=1)1.00
    8. LLM A1 (depth=2)0.94
    9. Sub-Agent A1 (depth=2)0.99
    10. LLM A2 (depth=2)0.98
    11. — Sub-Agent A2 (depth=2)0.98
    12. Sub-RLM B (depth=1)1.00
    13. Sub-Agent B (depth=1)1.00
    14. — Sub-Agent B1 (depth=2)0.98
    15. Sub-Agent B2 (depth=2)1.00
    16. recursivecodingagents.com1.00
  • 11:09 #22 done6 line(s)

    shot 22·sharpness 1095.4

    1. TOWARDS RECURSIVE CODING AGENTS1.00
    2. Either... Trick question: RLMs1.00
    3. are Recursive Coding Agents.0.98
    4. Or... How can we apply the principles0.99
    5. of RLMs to coding agents?1.00
    6. recursivecodingagents.com1.00
  • 11:46 #23 done16 line(s)

    shot 23·sharpness 740.8

    1. MY EXPERIMENTS1.00
    2. Finding ypi1.00
    3. Built on Pi (minimal, extensible). Previously pi extensions could not support recursion — so I0.99
    4. forked it. Y is for the Y-combinator.1.00
    5. Wrapper CLI – ypi - a fully recursive Pi agent.0.96
    6. Pi Extension — pi-recursive - make any existing Pi config recursive.0.98
    7. rawwerks/rlm-cli1.00
    8. rawwerks/ypi1.00
    9. CLI for Recursive Language Models.1.00
    10. A recursive coding agent inspired by RLMs.1.00
    11. Python804Updated Jun 16, 20260.97
    12. Shell339 29 MIT Updated Jun 15, 20260.95
    13. Open repo0.95
    14. Openrepo0.99
    15. Homepage1.00
    16. recursivecodingagents.com1.00

Transcript

216 cues· 3,505 words· 18,821 chars

  1. 0:00 Hello there!
  2. 0:01 My name is Raymond Weidekamp, and today I'm going to talk about recursive coding agents, which is this idea of applying the lessons of recursive language models, RLMs, to coding agents.
  3. 0:14 This is some work that I have done both in my independent research, Raw Works, and also more recently in my role at OpenProse.
  4. 0:28 to motivate this a little bit.
  5. 0:30 We all want outcomes.
  6. 0:31 We all want agents that are working on our behalf.
  7. 0:35 We want reliable coworkers that are getting things done while we're doing something fun, while we're out on a hike, while we're cold chilling, while we're doing the do.
  8. 0:45 And my argument and my experience is that the bottleneck to this is not intelligence.
  9. 0:54 The models are intelligent enough.
  10. 0:57 They know all kinds of things.
  11. 0:59 They know the entire internet, but they can't reliably deliver outcomes.
  12. 1:05 And so I can't trust them.
  13. 1:07 So as a very simple example, you know, one day I get almost a fully working SAS app from a single prompt, granted a long prompt the next day.
  14. 1:17 And I swear this actually happened.
  15. 1:20 Cloud code empties the entire contents of my Solana wallet.
  16. 1:24 Oops.
  17. 1:24 Okay.
  18. 1:25 So that doesn't really instill trust.
  19. 1:29 So at the bottom here, we've got this progression.
  20. 1:32 Okay.
  21. 1:33 And we all want to move towards the one on the right where we're just sort of sitting there meditating and things are manifesting.
  22. 1:39 And so where does that come from?
  23. 1:41 This is from the AI engineer code.
  24. 1:46 It's actually from the back of the t-shirt engineer code, November, 2025, man.
  25. 1:51 I hope, I hope you were there.
  26. 1:53 If you weren't watching it on YouTube, it was, it was amazing.
  27. 1:57 So here's the thesis.
  28. 1:58 The thesis is today's agents are mismanaged geniuses.
  29. 2:03 The intelligence is there and the missing layer is how do we specify and manage and reuse and verify
  30. 2:09 the work.
  31. 2:10 So this framing, this phrase, the mismanaged genius, comes from Alex Zhang, Zed Li, and Omar Khattab at MIT.
  32. 2:19 And Alex and Omar are part of the authors of the original recursive language models paper.
  33. 2:25 I've also talked a little bit about this recently on Turing Post.
  34. 2:29 I forgot to mention that these slides
  35. 2:32 are actually a website recursivecodingagents.com so you can click on them by going to this website so everything i'm going to show in here is is interactive okay what are recursive language models
  36. 2:47 so i like to say that in an rlm the context itself is the object of computation and this is essentially a marriage of tool calling and reasoning we're going to talk a lot more more about that in the next slide but the idea is that the full prompt is not
  37. 3:07 a simple user query the full prompt is a variable the full prompt could be a file or many files and we have this read evaluate print loop repl that the agent is interacting with in the original paper that's python and the rlm is instructed to operate symbolically on that prompt so don't just read the whole thing into your context window explore it symbolically and
  38. 3:37 even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have other LLMs, and I guess other RLMs, if you allow the recursion depth to be greater than one, have these other recursive subagents, and again, we'll get a little bit into the weeds of the lingo,
  39. 4:05 um sub rlm sub lms do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer so it looks something like like this in this in this tree below so
  40. 4:21 My take here is that our limbs are the new reasoning models.
  41. 4:26 And I see this as the next paradigm of test time, compute inference, time, compute, whatever you want to call it.
  42. 4:33 And why does it seem obvious or maybe like, Hey, why is this even a thing?
  43. 4:39 Um, I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution.
  44. 4:46 So the code execution is reasoning.
  45. 4:50 And so instead of, we had long, we had chain of thought as a prompting strategy that evolved into reasoning models that explicitly express the chain of thought as their reasoning tokens.
  46. 5:01 We already had function calling, tool calling, parallel tool calling, and RLMs really puts that together in a way that gets amazing results.
  47. 5:12 So three very simple examples, one from the original paper, Oolong,
  48. 5:16 the rlms can process information that is many orders of magnitude larger than their context window tens of millions of tokens what i showed in my own independent work was that the default rlm harness is itself a really powerful memory system
  49. 5:35 So RLM with no modifications is essentially like a top 10 memory system and like, you know, up there with all the people custom making memory systems and there's probably billions of dollars going into that.
  50. 5:49 And with a little bit of modification, you can get really amazing results using it as memory.

Open at this second