Videos 3hXJI2q0Jz8
Recursive Coding Agents - Raymond Weitekamp, OpenProse
Scene timeline
51 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 216
- whisperx 216
- chunks
- 44
- from 216 cues
- keyframes
- 27
- kept of 51 captured
- frames with text
- 27
- 528 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 4.0 MB
- word timings on 216 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 22:47 | 1m 32s |
stt |
done | — | 2026-08-10 22:49 | 24s |
chunk |
done | — | 2026-08-10 22:49 | 0s |
text_embed |
done | — | 2026-08-10 22:49 | 0s |
keyframe |
done | — | 2026-08-10 22:49 | 1m 14s |
ocr |
done | — | 2026-08-10 22:51 | 13s |
frame_embed |
done | — | 2026-08-10 22:51 | 5s |
Frames, and what the machine read
-
- AI ENGINEER WORLD'S FAIR· 20260.99
- Recursive Coding Agents1.00
- Raymond Weitekamp1.00
- RAW.works | OpenProse0.96
- @raw_works1.00
- recursivecodingagents.com1.00
-
- MOTIVATION1.00
- We all want outcomes.1.00
- Agents that work on our behalf — reliable co-workers — while we're out on a hike.0.99
- The bottleneck is not intelligence. It's reliability. It's trust.0.99
- One day — my agents build me a full SaaS app from a single prompt.0.99
- The next day — they empty the entire contents of my Solana wallet.0.98
- A0.62
- recursivecodingagents.com1.00
-
- MOTIVATION1.00
- We all want outcomes.1.00
- Agents that work on our behalf — reliable co-workers — while we're out on a hike.0.99
- The bottleneck is not intelligence. It's reliability. It's trust.0.99
- One day — my agents build me a full SaaS app from a single prompt.0.99
- The next day — they empty the entire contents of my Solana wallet.0.98
- A0.77
- CODE1.00
- recursivecodingagents.com1.00
-
- THESIS1.00
- Today's agents are1.00
- mismanaged geniuses1.00
- The intelligence is there.1.00
- The missing layer is how we specify, manage, reuse, and verify the work.1.00
- The Mismanaged Geniuses Hypothesis1.00
- Stop Babysitting Agents, Start0.98
- Authoring Outcomes0.98
- 00000000.60
- TURING POST - RAYMOND WEITEKAMP0.98
- ALEX ZHANG - ZED LI- OMAR KHATTAB0.96
- Stop Babysitting Agents, Start Authoring Outcomes0.99
- The Mismanaged Geniuses Hypothesis1.00
- recursivecodingagents.com1.00
-
- RECURSIVE LANGUAGE MODELS1.00
- Context itself is the0.99
- object of computation0.98
- Recursive Language Models1.00
- Externalize — the full prompt lives in a REPL,1.00
- not the context window.0.98
- Operate — the model writes code to inspect0.98
- slice, and transform it.1.00
- Recurse — it sub-queries itself over the slices.0.99
- ARXIV:2512.246011.00
- Root RLM (depth=0)0.99
- Sub-RLM A (depth=1)0.98
- Sub-RLM B (depth=1)1.00
- LLM B1 (depth=2)0.97
- - LLM B2 (depth=2)0.98
- recursivecodingagents.com1.00
-
- RECURSIVE LANGUAGE MODELS1.00
- Code1.00
- Execution As1.00
- Reasoning1.00
- Can process inputs way beyond the context window.0.97
- RLMs are the new reasoning models0.99
- (Oolong) RLM paper1.00
- can scalo wit tes-üime compute. Recursve language models (RLMs) ask0.86
- Reasoning models were the first clear proof that language model capability0.97
- RLM is itself a powerful memory system.1.00
- 2:50 PM - Apr 20, 2026 - 64.2KViews0.95
- LongMemEval results1.00
- Q160.75
- t0.81
- 5110.77
- 日860.69
- 土0.56
- Read 16 replies0.98
- RLM can achieve SOTA on long reasoning tasks,0.99
- even with very small models. LongCoT results1.00
- RLMs are the new reasoning models.1.00
- recursivecodingagents.com1.00
-
- RECURSIVE LANGUAGE MODELS1.00
- Code1.00
- Execution As1.00
- nond Weitekamp0.98
- Reasoning1.00
- Can process inputs way beyond the context window.0.97
- RLMs are the new reasoning models1.00
- (Oolong) RLM paper1.00
- can scalo wit tes-üime compute. Recursve languago models (RLMs) ask0.88
- Reasoning models were the first clear proof that language model capability0.96
- RLM is itself a powerful memory system.1.00
- 2:50 PM - Apr 20, 2026 - 66.2KViews0.95
- LongMemEval results1.00
- Q160.71
- t70.72
- 5110.85
- 8610.87
- Read 16 replies0.97
- RLM can achieve SOTA on long reasoning tasks,0.99
- even with very small models. LongCoT results1.00
- RLMs are the new reasoning models.1.00
- AlEnaineer0.98
- recursivecodingagents.com1.00
-
- RLMs: Too Hot To Benchmark0.98
- Raymond Weitekamp1.00
- x0.94
- Sumeet Motwani1.00
- x0.90
- @raw_works - Follow0.94
- @sumeetrm - Follow0.96
- i'm super confused...1.00
- LongCoT is adding two new leaderboards! Due to the1.00
- interest in agents (particularly RLMs), we're adding a0.99
- ...as far as i can tell, this a "consolation tweet" implying0.99
- "Restricted Harness" and an "Open Harness" leaderboard.0.99
- that ARC Prize will not verify the symbolica agent despite1.00
- their insane score...1.00
- at 25.12%. We expect tool-use SOTA to exceed this very1.00
- GPT 5.2 RLM from our paper is SOTA on "Open Harness"0.99
- ...at least that's how i'm interpreting the community0.99
- soon!1.00
- leaderboard reminder...0.99
- On Show more1.00
- ...am i totally mis-reading this?1.00
- LongCoT Open Harness Leaderboard1.00
- ARC Prize@arcprize0.96
- Today's @symbolica hamess is a clear example of what human-crafted0.99
- targeting can achieve on ARC-AGI-3 public demo set0.98
- You can "buy" performance with benchmark-specific prompts/strategies0.99
- Their approach coul stll contain useful ideas, excited to see what the0.98
- community finds0.99
- 10:47 AM · Mar 27, 20260.95
- Reply0.99
- Copy link0.95
- Read more on X1.00
- I personally do not care if my Al programs do their reasoning in latent space or code.0.99
- I want results.1.00
- recursivecodingagents.com1.00
-
- RLMs: Too Hot To Benchmark0.98
- Raymond Weitekamp1.00
- x0.95
- Sumeet Motwani1.00
- x0.91
- @raw_works - Follow0.94
- @sumeetrm - Follow0.93
- i'm super confused...0.99
- LongCoT is adding two new leaderboards! Due to the0.99
- interest in agents (particularly RLMs), we're adding a0.99
- ...as far as i can tell, this a "consolation tweet" implying0.99
- "Restricted Harness" and an "Open Harness" leaderboard.0.98
- that ARC Prize will not verify the symbolica agent despite1.00
- their insane score...1.00
- at 25.12%. We expect tool-use SOTA to exceed this very1.00
- GPT 5.2 RLM from our paper is SOTA on "Open Harness"0.99
- ...at least that's how i'm interpreting the community0.99
- soon!1.00
- leaderboard reminder...0.99
- On Show more1.00
- ...am i totally mis-reading this?0.99
- LongCoT Open Harness Leaderboard1.00
- ARC Prize@arcprize0.98
- Today's @symbolica hamness is a clear example of what human-crafted0.99
- targeting can achieve on ARC-AGI-3 public demo set0.98
- You can "buy" performance with benchmark-specific prompts/strategies0.99
- Their approach coul stil contain useful ideas, excited to see what the0.97
- community finds1.00
- 10:47 AM · Mar 27, 20260.95
- Reply1.00
- Copy link0.95
- Read more on X0.97
- I personally do not care if my Al programs do their reasoning in latent space or code.0.99
- I want results.1.00
- recursivecodingagents.com1.00
-
- THE RLM RUBRIC0.97
- Lots of things feel close.1.00
- Executable1.00
- Prompt1.00
- Code calls0.97
- Model picks0.99
- State stays1.00
- environment1.00
- externalized1.00
- the model1.00
- decomposition1.00
- symbolic1.00
- Plain long-context call0.98
- RAG / reasoning-only0.98
- Coding agents + subagents0.98
- including loops1.00
- Hardcoded map-reduce1.00
- developer-authored pipeline — e.g. λ-RLM0.99
- Recursive Language Model1.00
- passes every check0.99
- Open the RLM rubric1.00
- recursivecodingagents.com1.00
-
- THE RLM RUBRIC0.97
- Lots of things feel close.1.00
- Executable1.00
- Prompt1.00
- Code calls0.97
- Model picks0.99
- State stays1.00
- environment1.00
- externalized1.00
- the model1.00
- decomposition1.00
- symbolic1.00
- Plain long-context call1.00
- RAG / reasoning-only0.95
- Coding agents + subagents1.00
- including loops1.00
- Hardcoded map-reduce1.00
- developer-authored pipeline — e.g. λ-RLM0.98
- Recursive Language Model1.00
- passes every check1.00
- Open the RLM rubric1.00
- recursivecodingagents.com1.00
-
- TOWARDS RECURSIVE CODING AGENTS0.99
- RLM/LLM1.00
- Agent / Sub-Agent0.96
- Root RLM (depth=0)0.99
- Root Agent (depth=0)1.00
- Sub-RLM A (depth=1)0.99
- Sub-Agent A (depth=1)1.00
- LLM A1 (depth=2)0.94
- Sub-Agent A1 (depth=2)0.99
- LLM A2 (depth=2)0.98
- — Sub-Agent A2 (depth=2)0.98
- Sub-RLM B (depth=1)1.00
- Sub-Agent B (depth=1)1.00
- — Sub-Agent B1 (depth=2)0.98
- Sub-Agent B2 (depth=2)1.00
- recursivecodingagents.com1.00
-
- TOWARDS RECURSIVE CODING AGENTS1.00
- Either... Trick question: RLMs1.00
- are Recursive Coding Agents.0.98
- Or... How can we apply the principles0.99
- of RLMs to coding agents?1.00
- recursivecodingagents.com1.00
-
- MY EXPERIMENTS1.00
- Finding ypi1.00
- Built on Pi (minimal, extensible). Previously pi extensions could not support recursion — so I0.99
- forked it. Y is for the Y-combinator.1.00
- Wrapper CLI – ypi - a fully recursive Pi agent.0.96
- Pi Extension — pi-recursive - make any existing Pi config recursive.0.98
- rawwerks/rlm-cli1.00
- rawwerks/ypi1.00
- CLI for Recursive Language Models.1.00
- A recursive coding agent inspired by RLMs.1.00
- Python804Updated Jun 16, 20260.97
- Shell339 29 MIT Updated Jun 15, 20260.95
- Open repo0.95
- Openrepo0.99
- Homepage1.00
- recursivecodingagents.com1.00
Transcript
216 cues· 3,505 words· 18,821 chars
- 0:00 Hello there!
- 0:01 My name is Raymond Weidekamp, and today I'm going to talk about recursive coding agents, which is this idea of applying the lessons of recursive language models, RLMs, to coding agents.
- 0:14 This is some work that I have done both in my independent research, Raw Works, and also more recently in my role at OpenProse.
- 0:28 to motivate this a little bit.
- 0:30 We all want outcomes.
- 0:31 We all want agents that are working on our behalf.
- 0:35 We want reliable coworkers that are getting things done while we're doing something fun, while we're out on a hike, while we're cold chilling, while we're doing the do.
- 0:45 And my argument and my experience is that the bottleneck to this is not intelligence.
- 0:54 The models are intelligent enough.
- 0:57 They know all kinds of things.
- 0:59 They know the entire internet, but they can't reliably deliver outcomes.
- 1:05 And so I can't trust them.
- 1:07 So as a very simple example, you know, one day I get almost a fully working SAS app from a single prompt, granted a long prompt the next day.
- 1:17 And I swear this actually happened.
- 1:20 Cloud code empties the entire contents of my Solana wallet.
- 1:24 Oops.
- 1:24 Okay.
- 1:25 So that doesn't really instill trust.
- 1:29 So at the bottom here, we've got this progression.
- 1:32 Okay.
- 1:33 And we all want to move towards the one on the right where we're just sort of sitting there meditating and things are manifesting.
- 1:39 And so where does that come from?
- 1:41 This is from the AI engineer code.
- 1:46 It's actually from the back of the t-shirt engineer code, November, 2025, man.
- 1:51 I hope, I hope you were there.
- 1:53 If you weren't watching it on YouTube, it was, it was amazing.
- 1:57 So here's the thesis.
- 1:58 The thesis is today's agents are mismanaged geniuses.
- 2:03 The intelligence is there and the missing layer is how do we specify and manage and reuse and verify
- 2:09 the work.
- 2:10 So this framing, this phrase, the mismanaged genius, comes from Alex Zhang, Zed Li, and Omar Khattab at MIT.
- 2:19 And Alex and Omar are part of the authors of the original recursive language models paper.
- 2:25 I've also talked a little bit about this recently on Turing Post.
- 2:29 I forgot to mention that these slides
- 2:32 are actually a website recursivecodingagents.com so you can click on them by going to this website so everything i'm going to show in here is is interactive okay what are recursive language models
- 2:47 so i like to say that in an rlm the context itself is the object of computation and this is essentially a marriage of tool calling and reasoning we're going to talk a lot more more about that in the next slide but the idea is that the full prompt is not
- 3:07 a simple user query the full prompt is a variable the full prompt could be a file or many files and we have this read evaluate print loop repl that the agent is interacting with in the original paper that's python and the rlm is instructed to operate symbolically on that prompt so don't just read the whole thing into your context window explore it symbolically and
- 3:37 even more, you don't even directly export symbolically, or maybe you do a little bit of poking around, but have other LLMs, and I guess other RLMs, if you allow the recursion depth to be greater than one, have these other recursive subagents, and again, we'll get a little bit into the weeds of the lingo,
- 4:05 um sub rlm sub lms do this symbolic manipulation to pick apart the answer and then work our way back up to a final answer so it looks something like like this in this in this tree below so
- 4:21 My take here is that our limbs are the new reasoning models.
- 4:26 And I see this as the next paradigm of test time, compute inference, time, compute, whatever you want to call it.
- 4:33 And why does it seem obvious or maybe like, Hey, why is this even a thing?
- 4:39 Um, I think it's very elegant because it's a very elegant marriage of two things, reasoning and code execution.
- 4:46 So the code execution is reasoning.
- 4:50 And so instead of, we had long, we had chain of thought as a prompting strategy that evolved into reasoning models that explicitly express the chain of thought as their reasoning tokens.
- 5:01 We already had function calling, tool calling, parallel tool calling, and RLMs really puts that together in a way that gets amazing results.
- 5:12 So three very simple examples, one from the original paper, Oolong,
- 5:16 the rlms can process information that is many orders of magnitude larger than their context window tens of millions of tokens what i showed in my own independent work was that the default rlm harness is itself a really powerful memory system
- 5:35 So RLM with no modifications is essentially like a top 10 memory system and like, you know, up there with all the people custom making memory systems and there's probably billions of dollars going into that.
- 5:49 And with a little bit of modification, you can get really amazing results using it as memory.
loading