Videos 1IdzkRVmWAA
How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI
Scene timeline
31 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 109
- whisperx 109
- chunks
- 26
- from 109 cues
- keyframes
- 15
- kept of 31 captured
- frames with text
- 15
- 346 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.3 MB
- word timings on 109 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 04:49 | 1m 20s |
stt |
done | — | 2026-08-11 04:50 | 14s |
chunk |
done | — | 2026-08-11 04:50 | 0s |
text_embed |
done | — | 2026-08-11 04:50 | 0s |
keyframe |
done | — | 2026-08-11 04:50 | 42s |
ocr |
done | — | 2026-08-11 04:51 | 7s |
frame_embed |
done | — | 2026-08-11 04:51 | 2s |
Frames, and what the machine read
-
- AE0.68
- How we taught1.00
- agents to use good0.99
- retrieval1.00
- Aamir Shakir0.99
- Closing the Oracle Gap with knowledge agents.0.99
- Hanna Lichtenberg & Aamir Shakir · Mixedbread0.99
-
- AMixedbread0.97
- Al Engineer0.98
- Hanna Lichtenberg1.00
- Aamir Shakir1.00
- Aamir Shakir0.97
- Al Engineer0.98
- CEO & Cofounder1.00
- Volkswagen· TU Berlin0.96
- Google·EPFL0.99
- How we taught agents to use good retrieval · Mixedbread0.99
-
- A.Mixedbread0.92
- Al Engineer0.99
- Reasoning1.00
- CAES0.84
- Knowledge0.98
- gap1.00
- Search1.00
- Aamir Shakir0.98
- TIME1.00
- How we taught agents to use good retrieval · Mixedbread0.99
-
- A.Mixedbread0.92
- Al Engineer0.99
- Reasoning1.00
- CAES0.83
- Knowledge1.00
- gap1.00
- Search1.00
- Aamir Shakir0.97
- 41.00
- TIME1.00
- How we taught agents to use good retrieval · Mixedbread0.99
-
- AMixedbread0.97
- Al Engineer0.99
- BrowseComp-Plus1.00
- Office QA Pro0.97
- 961.00
- Oracle retrieval: 93.4%1.00
- 921.00
- 641.00
- A%)C%0.87
- (0)NS%)0.65
- 881.00
- 601.00
- 841.00
- Codex1.00
- -9.4pp0.97
- 561.00
- -8.2 pp0.92
- Codex1.00
- Aamir Shakir0.98
- 801.00
- 101.00
- 121.00
- 141.00
- 161.00
- 181.00
- 201.00
- 261.00
- 321.00
- 361.00
- TOOL CALLS0.99
- TOOL CALLS0.99
- How we taught agents to use good retrieval · Mixedbread0.98
-
- Mixedbread0.95
- Al Engineer0.99
- BrowseComp-Plus1.00
- Office QA Pro0.97
- 961.00
- 40.98
- Oracle retrieval: 93.4%1.00
- Oracle retrieval: 64.6%0.99
- 921.00
- 641.00
- ACA((%)0.67
- CORS%0.68
- 881.00
- 601.00
- 841.00
- Codex1.00
- -9.4 pp0.98
- 561.00
- -8.2 pp0.98
- Codex1.00
- Aamir Shakir1.00
- 801.00
- 101.00
- 121.00
- 141.00
- 161.00
- 181.00
- 201.00
- 261.00
- 321.00
- 361.00
- TOOL CALLS1.00
- TOOL CALLS0.96
- How we taught agents to use good retrieval · Mixedbread0.98
-
- Mixedbread1.00
- Al Engineer0.99
- BrowseComp-Plus1.00
- Office QA Pro0.98
- 961.00
- Oracle retrieval: 93.4%0.99
- Oracle retrieval: 64.6%0.99
- 921.00
- 40.98
- 641.00
- AC(%)0.81
- C0O%)0.67
- 881.00
- 601.00
- 841.00
- Codex1.00
- -9.4 pp0.99
- 561.00
- -8.2 pp0.93
- Codex1.00
- Aamir Shakir1.00
- 801.00
- 81.00
- 101.00
- 121.00
- 141.00
- 161.00
- 181.00
- 201.00
- 261.00
- 321.00
- 361.00
- TOOL CALLS1.00
- TOOL CALLS0.97
- How we taught agents to use good retrieval - Mixedbread0.98
-
- Mixedbread1.00
- Al Engineer0.99
- "Senator woman questions0.98
- billionaires not at company then ok0.99
- thank you staff will check hearing"0.98
- Coding agents are optimized for codebase exploration (grep).0.99
- Retrieval training and evals inherit web-search query distributions.1.00
- Benchmark bias: BEIR / nanoBEIR use "caveman-style", entity-dense queries0.99
- that structurally favor BM25.1.00
- Aamir Shakir0.99
- The agent guesses keywords to increase the overlap between query1.00
- and documents.0.99
- How we taught agents to use good retrieval · Mixedbread0.99
-
- Mixedbread1.00
- Al Engineer0.99
- Goals for our agent1.00
- Learn to use powerful search properly1.00
- Work beyond code for knowledge work1.00
- 3 Precise, fast and cheap1.00
- Hanna Lichtenberg0.97
- How we taught agents to use good retrieval - Mixedbread0.99
-
- Mixedbread1.00
- Al Engineer0.99
- Search agent1.00
- Searcher harness1.00
- speed & exploration:1.00
- harness built on1.00
- 4 rounds max · max 4 parallel tool calls per round0.98
- fewer total iterations1.00
- parallel searches,1.00
- Mixedbread1.00
- search planning1.00
- overview search1.00
- wide recall, top 50 summaries0.97
- Intent + clues0.98
- query1.00
- split intent into max. 4 queries0.99
- pick the best tool for each1.00
- cover separate aspects1.00
- full chunks for aspect A0.98
- semantic search1.00
- Main Search Tools:1.00
- semantic search1.00
- overview search: wide semantic search that returns0.99
- full chunks for aspect B0.99
- summaries of the top 50 retrieved chunks0.98
- filters + grep0.97
- semantic search: main semantic search tool, returns full1.00
- metadata and exact terms1.00
- filter chunks: find and sort chunks based on metadata1.00
- top 10 retrieved chunks0.99
- facets1.00
- submit_ranking1.00
- enough1.00
- evidence0.99
- follow-ups1.00
- target gaps1.00
- exclude seen chunks0.99
- seen chunk filter1.00
- Hanna Lichtenberg1.00
- grep: key-word search tool1.00
- How we taught agents to use good retrieval · Mixedbread0.99
-
- AMixedbread0.96
- AI Engineer0.95
- HARNESS GUIDANCE1.00
- 1.G0AL0.95
- Articulate what evidence is needed before writing the query.0.99
- Our harness1.00
- encourages1.00
- 2. TOOL0.94
- Semantic search for aspects; exact tools for required terms.0.99
- better semantic1.00
- 3. PROMPT0.99
- query."0.99
- Ask for "one concise sentence describing what you want to find," not "write a search0.99
- queries1.00
- 4. EXAMPLES0.99
- Show a few good queries and aspect-exploration moves.0.99
- Hanna Lichtenberg0.99
- 5. $EED0.87
- Use original-query results to reveal corpus language and branches.1.00
- How we taught agents to use good retrieval - Mixedbread0.99
-
- Mixedbread1.00
- AI Engineer0.96
- Targeted rewards1.00
- for better search0.99
- R = retrieval quality + trajectory quality0.99
- behavior1.00
- RETRIEVAL REWARD0.97
- TRAJECTORY REWARD1.00
- Recall / NDCG + LLM retrieval judge1.00
- LLM query and exploration judge1.00
- Fine-tune the agent on the search strategy0.99
- RUBRICS1.00
- RUBRICS1.00
- itself: tool choice, semantic query quality,0.99
- Agent result s relevant for the user query0.98
- Semantic query is a natural sentence1.00
- exploration, ranking, and efficiency.1.00
- All returned chunks support the search intent0.98
- Exploration matches the difficulty of the task0.99
- 1. SFT with teacher LLM1.00
- Chunk ranking is plausible and useful1.00
- Tool calls improve evidence, not just volume1.00
- Hanna Lichtenberg1.00
- 2. On-policy RL with our search reward1.00
- How we taught agents to use good retrieval - Mixedbread1.00
Transcript
109 cues· 1,974 words· 10,988 chars
- 0:00 Hey, today we're going to talk about how we taught agents to use good retrieval, or how we call it internally, closing the oracle gap with knowledge agents.
- 0:10 I'm Hannah, I'm an AI engineer at Mixper, and I'm leading the agentic search at Mixper.
- 0:17 Yeah, I'm Amir.
- 0:18 I'm the co-founder of Mixbread.
- 0:19 And let's jump into it.
- 0:21 So we have seen over the past years with the LLMs just getting better and better that the reasoning capabilities of the models is growing exponential.
- 0:31 So just think back how good GPT-3.5 was and how good GPT-5.5 is, right?
- 0:37 It's like a clearly exponential curve.
- 0:39 But if you look at search over the last 20 years, basically, we see that
- 0:45 it's improving but it's improving very very slowly so there's a huge gap between how basically llms and reasoning is evolving and how retrievers is evolving but this retrieval tools
- 1:01 are the main access pattern for this reasoning layer to get the right knowledge and to be truly useful beyond code, like in legal work, in finance work, and so on and so forth.
- 1:13 Internally, we call this gap happening right now between reasoning and search the knowledge gap.
- 1:19 And it's not just one obscure theory of ours.
- 1:22 We see it actually with real benchmarks and real tasks that this gap exists in real world as well.
- 1:28 Let's take two benchmarks here, BrowseCamp Plus and Office QA Pro.
- 1:33 BrowseCamp Plus is like a browsing benchmark where we have pretty complex queries and a corpus of 100,000 documents and we just try to answer these questions.
- 1:44 It's like real world deep search task basically.
- 1:48 And Office QA Pro, which has the US treasuries of the past 100 years, basically, and we asked a really complex question over this.
- 2:00 Office QA Pro was created by Databricks, and BrowseCamp was created by OpenAI, and BrowseCamp Plus is a version of it with a fixed size corpus, while BrowseCamp was over the open web.
- 2:11 And we see here's an Oracle performance in dotted lines.
- 2:15 And Oracle means what is the maximum theoretical performance of the models if we would put in the right documents with the question.
- 2:24 We see for BrowseComp it's 93% and for Office QA Pro it's 64%.
- 2:34 and now we take something like codex with its default tools and want to see hey what is the performance these tools get right now and you see there's a sharp drop in the quality of the answers codex produces for browse camp it's nine points and for office qa pro it's eight points so we see that the models are extremely capable if they would get the right documents
- 2:58 But if we put them into the noisy corpus, the performance drops sharply, meaning that actually the bottleneck here is not the reasoning.
- 3:11 It's actually the access to the right knowledge it needs to answer this question.
- 3:16 So we see this knowledge gap in real world.
- 3:21 And if you would just drop in way better search using mixed spread, which is a search tool using latent reaction, we can recover most of the performance.
- 3:30 So for BrowseCamp, the difference between the Oracle and GPT-5 with mixed spread is just three point.
- 3:38 And for Office QA Pro, we even almost completely closed the gap.
- 3:41 With giving the model better search tool, we can recover most of the knowledge gap.
- 3:47 But looking at the queries the model asks or writes right now to the search system, it gets super, super interesting.
- 3:55 So here's an example query we found during some benchmarking, which is Senator, woman questions, billionaires, not a company, then okay, thank you, staff, we'll check hearing.
- 4:07 It's basically gibberish, right?
- 4:08 If you have a search system which wants to have neural questions or semantic questions and just gets us keywords over keywords, it gets confused.
- 4:18 And the reason why the models write this type of queries is they're mostly trained for coding tasks, like for coding agents, which are then optimized for code-based exploration using tools like
- 4:30 grab, and grab just tries to find regular expressions, right?
- 4:36 So the model tries to just write as many expressions and things in the document to find the right thing.
- 4:42 The second thing is the models are trained to use the web pretty efficiently.
- 4:47 And to use the web tools properly, which are highly optimized for humans, they try to mimic human-like query patterns.
- 4:55 So just keyword on keyword on keyword.
- 4:57 And number three, obviously, is the benchmark bias.
- 5:00 Most benchmarks we have right now, like beer, nano beer, use caveman-style queries, which are entity-based queries that structurally favor heavily BM-25s.
- 5:11 So right now the agent guesses the keywords to actually increase the overlap between the query and documents and can't really use powerful search tools properly.
- 5:23 This motivated us to make our own search agent, to teach it to use powerful search properly, especially to use the right search tools for different use cases and to work beyond code search for knowledge work.
- 5:40 And most importantly, the agent should be precise, fast and cheap.
- 5:44 The first step for building this agent was to define a very powerful harness.
- 5:50 Our harness is fully built on the Mixbed platform.
- 5:54 Our agent has four main search tools, which are overview search.
- 5:59 This is used as a very wide semantic search where the agent receives up to 50 retrieved chunks.
- 6:07 And it sees only summaries of the chunk contents to really have just like an overview of the corpus.
loading