Videos 8oyalrfwgjw
RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
Scene timeline
131 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 135
- whisperx 135
- chunks
- 32
- from 135 cues
- keyframes
- 96
- kept of 131 captured
- frames with text
- 94
- 2,811 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 18.0 MB
- word timings on 135 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 03:11 | 1m 26s |
stt |
done | — | 2026-08-11 03:13 | 23s |
chunk |
done | — | 2026-08-11 03:13 | 0s |
text_embed |
done | — | 2026-08-11 03:13 | 0s |
keyframe |
done | — | 2026-08-11 03:13 | 2m 43s |
ocr |
done | — | 2026-08-11 03:16 | 1m 11s |
frame_embed |
done | — | 2026-08-11 03:17 | 18s |
Frames, and what the machine read
-
- localhost1.00
- +0.90
- AI ENGINEER WORLD'S FAIR 20260.99
- Online Track1.00
- RLM:Recursive1.00
- LanguageModels1.00
- MIT RLM PAPER1.00
- Recursive Language Models1.00
- for Large0.95
- RLMs use a programmable environment for long-1.00
- context inference.1.00
- Codebases I0.91
- arXiv:2512.246010.99
- A coding-agent harness pattern, inspired by the MIT1.00
- RLM paper.0.99
- MIT RESEARCH1.00
- The RLM concept1.00
- MIT RLM paper: arXiv:2512.246010.99
- Use the RLM control loop as a coding-agent0.98
- harness pattern for large repository investigation1.00
- Shashi Jagtap1.00
- Founder, Superagentic Al1.00
- RLM: Coding-Agent Harness Pattern for Large Codebases0.99
- next back F fullscreen0.93
-
- pecoming an independent0.97
- We gratefully acknowledge support from the Simons1.00
- Donate1.00
- onprofit.1.00
- Foundation, member institutions, and all contributors.0.99
- Search...1.00
- All fields1.00
- Search1.00
- Help | Advanced Search0.98
- Access Paper:1.00
- View P1.00
- HTML (e0.99
- nental)0.99
- TeX Source0.99
- (cc) BY0.93
- view license1.00
- Current browse context:0.99
- ns of inference-time scaling. We propose Recursive Language0.99
- cs.Al1.00
- nment and allows the LLM to programmatically examine,0.98
- < prev0.92
- next >0.98
- lly process inputs up to two orders of magnitude beyond model1.00
- new | recent | 2025-120.94
- ier LLMs and common long-context and coding scaffolds (e.g., on1.00
- Change to browse by:0.98
- deAct with sub-calls, and 13%against Claude Code) across four0.98
- CS0.66
- odel around the RLM. Our model, RLM-Qwen3-8B, outperforms0.99
- cs.CL1.00
- PT-5 on three long-context tasks. Code is available at this https0.98
- References & Citat0.99
- NASA ADS1.00
- Google Scholar1.00
- Semantic Scholar0.99
-
- arxiv.org1.00
- C0.63
- Agent Harness Secret for Massive Codebases1.00
- https://arxi1.00
- Recursive Language Models1.00
- Alex L. Zhang0.99
- Tim Kraska1.00
- Omar Kh1.00
- MIT CSAIL0.95
- MIT CSAIL0.96
- MIT CSA0.98
- [email protected]1.00
- [email protected]1.00
- [email protected]1.00
-
- LM:Recursive1.00
- MIT RLM PAPER1.00
- anguage Models0.97
- Recursive L0.99
- or Large0.98
- RLMs use a progra1.00
- context inference.1.00
- odebases lı0.89
- arXiv:2512.246011.00
- oding-agent harness pattern, inspired by the MIT1.00
- A paper.0.89
- MIT RESEARCH1.00
- The RLM co0.95
- T RLM paper: arXiv:2512.246010.99
- Use the RI M contr0.99
- ha1.00
- fo1.00
- ashi Jagtap1.00
- under, Superagentic Al0.98
-
- <0.88
- localhost1.00
- +0.87
- The problem1.00
- Large repositories require stateful, structured investigation.1.00
- SMALL REPO1.00
- LARGE REPO0.99
- MORE CONTEXT0.99
- REAL PROBLEM1.00
- One file, one test, one0.99
- Lost architecture,1.00
- Bigger windows add1.00
- A codebase is a1.00
- function. Agents handle1.00
- ownership, and cross-file0.99
- noise. Tool output floods1.00
- structured system, not a0.99
- this well.1.00
- links.1.00
- the chat.1.00
- long document.0.99
- RLM: Coding-Agent Harness Pattern for Large Codebases1.00
-
- <0.92
- localhost1.00
- +0.89
- How people tackle it today0.98
- Search, retrieval, and longer context all help.1.00
- FILESYSTEM TOOLS1.00
- SEMANTIC RETRIEVAL1.00
- LONG CONTEXT1.00
- AGENT MEMORY0.97
- Search the repo directly0.99
- Find likely-relevant0.99
- Load or compress0.97
- Persist working state0.99
- grep, file readers, virtual FS.0.99
- code1.00
- evidence1.00
- Scratchpads and task state1.00
- Good when paths are known.1.00
- Embeddings narrow the0.98
- Bigger windows help, but detail1.00
- avoid rereading every token.0.99
- haystack, but can miss1.00
- competes with noise.1.00
- structure.1.00
- RLM: Coding-Agent Harness Pattern for Large Codebases0.99
-
- ontext management1.00
- 1. The repository is treated0.99
- to a programmable0.99
- 2. The model writes code to1.00
- xecution1.00
- nvironment.1.00
- 3. Focused subquestions re1.00
- rness Pattern for Large Codebases1.00
-
- localhost1.00
- THE THESIS0.98
- RLMexternalizes1.00
- MIT RLM paper: arXiv:2512.246010.97
- context management0.98
- 1. The repository is treated as data the model can operate on.0.99
- into aprogrammable0.97
- 2. The model writes code to inspect, slice, and compute.0.99
- execution1.00
- environment.1.00
- 3. Focused subquestions return clean values.1.00
- RLM: Coding-Agent Harness Pattern for Large Codebases0.99
-
- <0.87
- localhost1.00
- +0.95
- 口0.60
- The analogy0.98
- A senior engineer with a notebook.0.98
- Q1.00
- Large project1.00
- Files, docs, tests, configs0.99
- Focused helper1.00
- task1.00
- inspect this1.00
- Notebook1.00
- Persistent REPL state1.00
- Engineer doing1.00
- evidence1.00
- the1.00
- Investigation scripts1.00
- Model-written Python0.99
- investigation1.00
- evidence, and open1.00
- keeps notes,1.00
- questions1.00
- Clean note1.00
- Focused helper1.00
- question1.00
- Ilm_query(...)0.94
- returned1.00
- added back to the0.98
- Returned note1.00
- A clean value for synthesis1.00
- notebook1.00
- Recursion = the parent REPL calling 11m_query(. .. )0.98
- antting a clean value back.0.99
- RLM: Coding-Agent Harness Pattern for Large Codebases0.99
-
- localhost1.00
- +0.92
- The loop0.95
- How the harness actuallyruns.1.00
- Repo context1.00
- Root model1.00
- Python REPL1.00
- context1.00
- writes repl code1.00
- bounded observation1.00
- Recursive subcall0.99
- Clean result1.00
- Finish1.00
- llm_query(...)0.97
- value returns1.00
- FINAL / FINAL_VAR0.99
- RLM: Coding-Agent Harness Pattern for Large Codebases0.98
-
- <0.86
- localhost1.00
- +0.94
- 口0.57
- Why codebases0.98
- Repos are structured and executable.0.99
- REPOSITORY AS STRUCTURED DATA1.00
- runner.py1.00
- orchestration1.00
- environment.py1.00
- REPL state0.97
- 1 Files and directories give the model a map.0.98
- 2 Imports and symbols reveal execution paths.1.00
- termination.py1.00
- FINAL /FINAL_VAR0.97
- Tests and fixtures expose intended behavior.0.99
- tests/1.00
- behavior checks1.00
- 4 Configs and docs explain runtime assumptions.0.99
- README.md1.00
- usage contract1.00
- MODEL-WRIEN INSPECTION0.99
- extract ewnce → llm_query → FINAL0.95
- RLM: Coding-Agent Harness Pattern for Large Codebases0.99
-
- localhost1.00
- è0.72
- +0.77
- 020.93
- RLMCode1.00
- An independent experimental reference implementation used to demonstrate the pattern.0.99
- Landing page1.00
- Docs1.00
- GitHub1.00
- RLM: Coding-Agent Harness Pattern for Large Codebases1.00
-
- 020.98
- RLMCode1.00
- An independent experimental reference implementation used to der1.00
- Landing1.00
- Docs1.00
- GitHub1.00
-
- super-agentic.ai1.00
- X0.67
- +0.94
- RLM: The Coding Agent Harness Secret for Massive Codebases0.98
- RLM Code | Research Playground for Recursive Language Models | Superagentic Al1.00
- SuperagenticAI1.00
- <> Products0.99
- Solutions1.00
- OSS0.76
- Research1.00
- Resources0.95
- Company1.00
- RESEARCH-FIRST EVALUATION OS1.00
- RLMCode1.00
- Research Playground for Recursive0.99
- Language Models1.00
- Run real RLM workflows, benchmark them, replay every step, and1.00
- compare results under controlled budgets and secure sandboxes.1.00
- Open source under Apache-2.0. Explore the project on GitHub.0.98
- Start with Docs0.99
- View GitHub1.00
- Apache-2.01.00
- PyPI Package1.00
- Research-first1.00
- What is included1.00
- Preferred stack1.00
- Code Mode support1.00
- Watch demo1.00
- Quick workflow1.00
- Security1.00
- FAQ1.00
-
- super-agentic.ai1.00
- X0.60
- +0.95
- RLM: The Coding Agent Harness Secret for Massive Codebases1.00
- RLM Code | Research Playground for Recursive Language Models | Superagentic Al0.99
- SuperagenticAI1.00
- <> Products0.99
- Solutions1.00
- OSS0.72
- Research1.00
- Resources0.95
- V0.82
- Company0.91
- RESEARCH-FIRST EVALUATION OS0.99
- RLMCode1.00
- Research Playground for Recursive1.00
- Language Models1.00
- Run real RLM workflows, benchmark them, replay every step, and1.00
- compare results under controlled budgets and secure sandboxes.1.00
- Open source under Apache-2.0. Explore the project on GitHub.0.98
- Start with Docs0.99
- View GitHub1.00
- Apache-2.01.00
- PyPI Package1.00
- Research-first1.00
- What is included0.99
- Preferred stack1.00
- Code Mode support1.00
- Watch demo1.00
- Quick workflow1.00
- Security1.00
- FAQ1.00
-
- super-agentic.ai1.00
- ×0.79
- +0.97
- RLM: The Coding Agent Harness Secret for Massive Codebases1.00
- RLM Code | Research Playground for Recursive Language Models | Superagentic Al0.99
- SuperagenticAI1.00
- <> Products0.95
- Solutions1.00
- Oss0.54
- Research0.95
- Resources0.97
- Company0.97
- V0.90
- RESEARCH-FIRST EVALUATION OS1.00
- RLM Code0.97
- RLM1.00
- RLM Code1.00
- Research Playground for Recursive1.00
- CODE-0.98
- Research Playground & Evaluation OS for RLM Agentic Systems1.00
- Language Models1.00
- RLM0.86
- Files0.95
- Details0.97
- Shell1.00
- Research Lab0.97
- One Screen:0.97
- Quit0.80
- Run real RLM workflows, benchmark them, replay every step, and1.00
- RLM Restarch Lib0.83
- compare results under controlled budgets and secure sandboxes.0.99
- Dashboard1.00
- Trajectory1.00
- Benchnarks0.99
- Replay0.99
- Events0.99
- Open source under Apache-2.0. Explore the project on GitHub.0.99
- Rum Dashboard0.98
- Start with Docs1.00
- View GitHub1.00
- aolong_style0.94
- Apache-2.01.00
- PyPI Package1.00
- Research-first1.00
- What is included0.99
- Preferred stack1.00
- Code Mode support1.00
- Watch demo1.00
- Quick workflow1.00
- Security1.00
- FAQ1.00
-
- RLM: The Coding Agent Harness Secret fic0.97
- SuperagenticAI1.00
- <> Products0.97
- 目0.67
- Solutions0.94
- 0SS0.60
- RESEARCH-FIRST EVALUATION OS0.99
- RLMCode1.00
- Research Playground for Recursive1.00
- Language Models0.97
- Pun real Rl M workflowe. henehmark them mwvay every st ay0.83
Transcript
135 cues· 2,394 words· 12,806 chars
- 0:00 Hello, and welcome to this online track talk for the AI Engineer World's Fair 2026.
- 0:08 Today, we're going to explore the concept of RLM, also known as recursive language models, and how we can use those concepts for larger code bases.
- 0:20 My name is Shashi.
- 0:21 I'm a founder of SuperAgenti.KI.
- 0:23 First of all, let's be clear that RLM paper has been published by MIT and Friends.
- 0:30 As you can see, there's a full paper.
- 0:32 You can read about it.
- 0:34 But the purpose of this talk is how you can use the concepts of RLM.
- 0:39 and you can use into your own workflow to implement your own harnesses so first of all what's the problem if you're using the coding agents for smaller repos or monorepos they work exceptionally well but if you have ever tried it with the monorepos with the large context you know there is a context problem as the context grows the performance degrade and if you're working with the monorepos this problem get worst
- 1:05 In this talk, we will see we selected the code base and the concept of RLMs are relevant for the larger code bases.
- 1:15 If you use the coding agents, then you probably saw that there are different approaches that other coding agent harnesses have been taken to solve this problem.
- 1:23 Most common approach is searching using the tools like grep.
- 1:28 So basically there's a file system and the coding agent harnesses search using these tools.
- 1:34 The second approach you probably seen that the semantic search or the local search.
- 1:40 So idea here is basically you can search through the code and curate the context.
- 1:46 Another approach is the long context get compressed, and you can use the summarized version of the context.
- 1:54 And there are some memory solutions available in the market as well that you can use to persist the memory for the coding agent.
- 2:05 First of all, let's explore the RLM idea.
- 2:08 The core thesis of the RLM is you need to externalize the context management into programmable execution environment, meaning you should have a separate dedicated environment so that model can operate on that.
- 2:23 In this case, for example, your whole repository is treated as a data that model can operate on.
- 2:30 Then model can write the code to inspect, slice, and compute the relevant chunks of value you can then feed into the main context window.
- 2:42 So basically rather than putting everything into the models context, create a separate dedicated environment, give them a coding agent or REPL, and then model write the code to curate the context that can be used into the main.
- 2:59 So it's another context management technique proved to be very effective.
- 3:04 Could be also be used as a memory layer for your coding agents.
- 3:10 Let me summarize this, giving you a simple analogy.
- 3:14 Imagine you are a lead software engineer and assigned to the new project with a huge code base.
- 3:22 Imagine that's Mono repo.
- 3:24 How does that lead engineer deal with the code?
- 3:28 So rather than reading
- 3:31 line of code line by line engineer probably inspect the code base make some notes see what are the project's dependencies how it is structured if it's something else is not understood by the repository engineer probably asked to another engineer or expert to get some ideas and the same concepts applied in rlm so large project like the files and docs and texts and configs because repository has a lot of things
- 4:02 And the programmable REPL, it's kind of a notebook that engineer makes a note about the code base that can be used.
- 4:08 He's searching, he may be using other techniques or maybe he's writing some script to search something from the repo.
- 4:16 And then if he stacks, then he asks another engineer or specialist where it comes to the LLM query.
- 4:25 and LLM query is basically asking another model environment to get answer from and once they get answer then the loop continues and at the end it returns the clean node synthesis so the recursion part here is engineer ask another specialist using LLM query that can be one question or that can be number of questions so this is where the recursion comes in picture
- 4:54 loop is basically your repo as your context and then the model writes the repl code to get some relevant context that returns the bounded observation and if we if loop needs more information it passes through the llm query where it asks another language model or another system to get the response return the value and continue the loop
- 5:25 and the loop get terminated until we get to final results.
- 5:31 Why are we talking about the code base?
- 5:33 and not the big context in terms of like other things, for example, books or dictionaries.
- 5:38 Codebase is different.
- 5:40 It has directories, it has tests, it has some imports, it has dependencies, it has tests, it has pictures, it has configuration files.
- 5:50 So the codebase is not only just the text, it is structured data.
- 5:56 And the model need to understand and reason over the text.
- 6:01 That's why I chose this scenario to use a code basis to prove these concepts of RLM.
- 6:09 Now let's switch the gear and talk about our own library that we created at SuperAgent TKI called RLM Code.
- 6:18 You can see RLM Code's landing page here, where this is just a research playground where you can invent the concepts of RLM.
- 6:29 We have documentation that you can take a look and there's the GitHub repository.
- 6:37 It is completely open source project that you can use it and play with it.
- 6:46 RLM itself is a concept and a pattern.
- 6:51 and you can implement that concept and pattern in your own way there are official authors also wrote some implementation in their github repos it's called rlm and rlm minimal you can refer that implementation of rlm in dspy.rlm so omar
- 7:15 is author of RLM and he's also author of another popular framework called DSPy.
loading