read-only demo

Videos 1IdzkRVmWAA

How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI

index_state ready data_status ok

AI Engineer· published 2026-07-07· 0:14:28· en-US· indexed 2026-08-11 04:51

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:09, 1 of 1 keyframes kept
  2. Shot 1, 0:09 to 0:21, 1 of 1 keyframes kept
  3. Shot 2, 0:21 to 0:52, 1 of 1 keyframes kept
  4. Shot 3, 0:52 to 1:23, 1 of 1 keyframes kept
  5. Shot 4, 1:23 to 1:52, 1 of 1 keyframes kept
  6. Shot 5, 1:52 to 2:21, 1 of 1 keyframes kept
  7. Shot 6, 2:21 to 2:51, 0 of 1 keyframes kept
  8. Shot 7, 2:51 to 3:20, 1 of 1 keyframes kept
  9. Shot 8, 3:20 to 3:47, 0 of 1 keyframes kept
  10. Shot 9, 3:47 to 4:19, 1 of 1 keyframes kept
  11. Shot 10, 4:19 to 4:51, 0 of 1 keyframes kept
  12. Shot 11, 4:51 to 5:23, 0 of 1 keyframes kept
  13. Shot 12, 5:23 to 5:44, 1 of 1 keyframes kept
  14. Shot 13, 5:44 to 6:11, 1 of 1 keyframes kept
  15. Shot 14, 6:11 to 6:37, 0 of 1 keyframes kept
  16. Shot 15, 6:37 to 7:04, 0 of 1 keyframes kept
  17. Shot 16, 7:04 to 7:31, 0 of 1 keyframes kept
  18. Shot 17, 7:31 to 7:57, 0 of 1 keyframes kept
  19. Shot 18, 7:57 to 8:24, 0 of 1 keyframes kept
  20. Shot 19, 8:24 to 8:56, 1 of 1 keyframes kept
  21. Shot 20, 8:56 to 9:28, 0 of 1 keyframes kept
  22. Shot 21, 9:28 to 10:01, 0 of 1 keyframes kept
  23. Shot 22, 10:01 to 10:30, 1 of 1 keyframes kept
  24. Shot 23, 10:30 to 11:00, 0 of 1 keyframes kept
  25. Shot 24, 11:00 to 11:29, 0 of 1 keyframes kept
  26. Shot 25, 11:29 to 11:59, 0 of 1 keyframes kept
  27. Shot 26, 11:59 to 12:16, 1 of 1 keyframes kept
  28. Shot 27, 12:16 to 12:58, 1 of 1 keyframes kept
  29. Shot 28, 12:58 to 13:27, 1 of 1 keyframes kept
  30. Shot 29, 13:27 to 13:57, 0 of 1 keyframes kept
  31. Shot 30, 13:57 to 14:27, 0 of 1 keyframes kept

31 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
109
whisperx 109
chunks
26
from 109 cues
keyframes
15
kept of 31 captured
frames with text
15
346 lines read
chapters
0
from the source metadata
keyframe bytes
2.3 MB
word timings on 109 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 04:49 1m 20s
stt done 2026-08-11 04:50 14s
chunk done 2026-08-11 04:50 0s
text_embed done 2026-08-11 04:50 0s
keyframe done 2026-08-11 04:50 42s
ocr done 2026-08-11 04:51 7s
frame_embed done 2026-08-11 04:51 2s

Frames, and what the machine read

  • 0:08 #0 done7 line(s)

    shot 0·sharpness 976.9

    1. AE0.68
    2. How we taught1.00
    3. agents to use good0.99
    4. retrieval1.00
    5. Aamir Shakir0.99
    6. Closing the Oracle Gap with knowledge agents.0.99
    7. Hanna Lichtenberg & Aamir Shakir · Mixedbread0.99
  • 0:18 #1 done10 line(s)

    shot 1·sharpness 709.1

    1. AMixedbread0.97
    2. Al Engineer0.98
    3. Hanna Lichtenberg1.00
    4. Aamir Shakir1.00
    5. Aamir Shakir0.97
    6. Al Engineer0.98
    7. CEO & Cofounder1.00
    8. Volkswagen· TU Berlin0.96
    9. Google·EPFL0.99
    10. How we taught agents to use good retrieval · Mixedbread0.99
  • 0:36 #2 done10 line(s)

    shot 2·sharpness 434.8

    1. A.Mixedbread0.92
    2. Al Engineer0.99
    3. Reasoning1.00
    4. CAES0.84
    5. Knowledge0.98
    6. gap1.00
    7. Search1.00
    8. Aamir Shakir0.98
    9. TIME1.00
    10. How we taught agents to use good retrieval · Mixedbread0.99
  • 1:07 #3 done11 line(s)

    shot 3·sharpness 445.8

    1. A.Mixedbread0.92
    2. Al Engineer0.99
    3. Reasoning1.00
    4. CAES0.83
    5. Knowledge1.00
    6. gap1.00
    7. Search1.00
    8. Aamir Shakir0.97
    9. 41.00
    10. TIME1.00
    11. How we taught agents to use good retrieval · Mixedbread0.99
  • 1:37 #4 done32 line(s)

    shot 4·sharpness 438.2

    1. AMixedbread0.97
    2. Al Engineer0.99
    3. BrowseComp-Plus1.00
    4. Office QA Pro0.97
    5. 961.00
    6. Oracle retrieval: 93.4%1.00
    7. 921.00
    8. 641.00
    9. A%)C%0.87
    10. (0)NS%)0.65
    11. 881.00
    12. 601.00
    13. 841.00
    14. Codex1.00
    15. -9.4pp0.97
    16. 561.00
    17. -8.2 pp0.92
    18. Codex1.00
    19. Aamir Shakir0.98
    20. 801.00
    21. 101.00
    22. 121.00
    23. 141.00
    24. 161.00
    25. 181.00
    26. 201.00
    27. 261.00
    28. 321.00
    29. 361.00
    30. TOOL CALLS0.99
    31. TOOL CALLS0.99
    32. How we taught agents to use good retrieval · Mixedbread0.98
  • 2:04 #5 done34 line(s)

    shot 5·sharpness 433.9

    1. Mixedbread0.95
    2. Al Engineer0.99
    3. BrowseComp-Plus1.00
    4. Office QA Pro0.97
    5. 961.00
    6. 40.98
    7. Oracle retrieval: 93.4%1.00
    8. Oracle retrieval: 64.6%0.99
    9. 921.00
    10. 641.00
    11. ACA((%)0.67
    12. CORS%0.68
    13. 881.00
    14. 601.00
    15. 841.00
    16. Codex1.00
    17. -9.4 pp0.98
    18. 561.00
    19. -8.2 pp0.98
    20. Codex1.00
    21. Aamir Shakir1.00
    22. 801.00
    23. 101.00
    24. 121.00
    25. 141.00
    26. 161.00
    27. 181.00
    28. 201.00
    29. 261.00
    30. 321.00
    31. 361.00
    32. TOOL CALLS1.00
    33. TOOL CALLS0.96
    34. How we taught agents to use good retrieval · Mixedbread0.98
  • 2:39 #6 skipped

    shot 6·duplicate of #4

  • 3:11 #7 done35 line(s)

    shot 7·sharpness 429.9

    1. Mixedbread1.00
    2. Al Engineer0.99
    3. BrowseComp-Plus1.00
    4. Office QA Pro0.98
    5. 961.00
    6. Oracle retrieval: 93.4%0.99
    7. Oracle retrieval: 64.6%0.99
    8. 921.00
    9. 40.98
    10. 641.00
    11. AC(%)0.81
    12. C0O%)0.67
    13. 881.00
    14. 601.00
    15. 841.00
    16. Codex1.00
    17. -9.4 pp0.99
    18. 561.00
    19. -8.2 pp0.93
    20. Codex1.00
    21. Aamir Shakir1.00
    22. 801.00
    23. 81.00
    24. 101.00
    25. 121.00
    26. 141.00
    27. 161.00
    28. 181.00
    29. 201.00
    30. 261.00
    31. 321.00
    32. 361.00
    33. TOOL CALLS1.00
    34. TOOL CALLS0.97
    35. How we taught agents to use good retrieval - Mixedbread0.98
  • 3:26 #8 skipped

    shot 8·duplicate of #4

  • 4:09 #9 done13 line(s)

    shot 9·sharpness 1471.2

    1. Mixedbread1.00
    2. Al Engineer0.99
    3. "Senator woman questions0.98
    4. billionaires not at company then ok0.99
    5. thank you staff will check hearing"0.98
    6. Coding agents are optimized for codebase exploration (grep).0.99
    7. Retrieval training and evals inherit web-search query distributions.1.00
    8. Benchmark bias: BEIR / nanoBEIR use "caveman-style", entity-dense queries0.99
    9. that structurally favor BM25.1.00
    10. Aamir Shakir0.99
    11. The agent guesses keywords to increase the overlap between query1.00
    12. and documents.0.99
    13. How we taught agents to use good retrieval · Mixedbread0.99
  • 4:44 #10 skipped

    shot 10·duplicate of #9

  • 5:13 #11 skipped

    shot 11·duplicate of #9

  • 5:31 #12 done8 line(s)

    shot 12·sharpness 1134.1

    1. Mixedbread1.00
    2. Al Engineer0.99
    3. Goals for our agent1.00
    4. Learn to use powerful search properly1.00
    5. Work beyond code for knowledge work1.00
    6. 3 Precise, fast and cheap1.00
    7. Hanna Lichtenberg0.97
    8. How we taught agents to use good retrieval - Mixedbread0.99
  • 5:52 #13 done41 line(s)

    shot 13·sharpness 1157.6

    1. Mixedbread1.00
    2. Al Engineer0.99
    3. Search agent1.00
    4. Searcher harness1.00
    5. speed & exploration:1.00
    6. harness built on1.00
    7. 4 rounds max · max 4 parallel tool calls per round0.98
    8. fewer total iterations1.00
    9. parallel searches,1.00
    10. Mixedbread1.00
    11. search planning1.00
    12. overview search1.00
    13. wide recall, top 50 summaries0.97
    14. Intent + clues0.98
    15. query1.00
    16. split intent into max. 4 queries0.99
    17. pick the best tool for each1.00
    18. cover separate aspects1.00
    19. full chunks for aspect A0.98
    20. semantic search1.00
    21. Main Search Tools:1.00
    22. semantic search1.00
    23. overview search: wide semantic search that returns0.99
    24. full chunks for aspect B0.99
    25. summaries of the top 50 retrieved chunks0.98
    26. filters + grep0.97
    27. semantic search: main semantic search tool, returns full1.00
    28. metadata and exact terms1.00
    29. filter chunks: find and sort chunks based on metadata1.00
    30. top 10 retrieved chunks0.99
    31. facets1.00
    32. submit_ranking1.00
    33. enough1.00
    34. evidence0.99
    35. follow-ups1.00
    36. target gaps1.00
    37. exclude seen chunks0.99
    38. seen chunk filter1.00
    39. Hanna Lichtenberg1.00
    40. grep: key-word search tool1.00
    41. How we taught agents to use good retrieval · Mixedbread0.99
  • 6:24 #14 skipped

    shot 14·duplicate of #13

  • 6:56 #15 skipped

    shot 15·duplicate of #13

  • 7:28 #16 skipped

    shot 16·duplicate of #13

  • 7:34 #17 skipped

    shot 17·duplicate of #13

  • 8:01 #18 skipped

    shot 18·duplicate of #13

  • 8:34 #19 done20 line(s)

    shot 19·sharpness 1187.2

    1. AMixedbread0.96
    2. AI Engineer0.95
    3. HARNESS GUIDANCE1.00
    4. 1.G0AL0.95
    5. Articulate what evidence is needed before writing the query.0.99
    6. Our harness1.00
    7. encourages1.00
    8. 2. TOOL0.94
    9. Semantic search for aspects; exact tools for required terms.0.99
    10. better semantic1.00
    11. 3. PROMPT0.99
    12. query."0.99
    13. Ask for "one concise sentence describing what you want to find," not "write a search0.99
    14. queries1.00
    15. 4. EXAMPLES0.99
    16. Show a few good queries and aspect-exploration moves.0.99
    17. Hanna Lichtenberg0.99
    18. 5. $EED0.87
    19. Use original-query results to reveal corpus language and branches.1.00
    20. How we taught agents to use good retrieval - Mixedbread0.99
  • 9:06 #20 skipped

    shot 20·duplicate of #19

  • 9:38 #21 skipped

    shot 21·duplicate of #19

  • 10:10 #22 done25 line(s)

    shot 22·sharpness 1269.7

    1. Mixedbread1.00
    2. AI Engineer0.96
    3. Targeted rewards1.00
    4. for better search0.99
    5. R = retrieval quality + trajectory quality0.99
    6. behavior1.00
    7. RETRIEVAL REWARD0.97
    8. TRAJECTORY REWARD1.00
    9. Recall / NDCG + LLM retrieval judge1.00
    10. LLM query and exploration judge1.00
    11. Fine-tune the agent on the search strategy0.99
    12. RUBRICS1.00
    13. RUBRICS1.00
    14. itself: tool choice, semantic query quality,0.99
    15. Agent result s relevant for the user query0.98
    16. Semantic query is a natural sentence1.00
    17. exploration, ranking, and efficiency.1.00
    18. All returned chunks support the search intent0.98
    19. Exploration matches the difficulty of the task0.99
    20. 1. SFT with teacher LLM1.00
    21. Chunk ranking is plausible and useful1.00
    22. Tool calls improve evidence, not just volume1.00
    23. Hanna Lichtenberg1.00
    24. 2. On-policy RL with our search reward1.00
    25. How we taught agents to use good retrieval - Mixedbread1.00
  • 10:42 #23 skipped

    shot 23·duplicate of #22

Transcript

109 cues· 1,974 words· 10,988 chars

  1. 0:00 Hey, today we're going to talk about how we taught agents to use good retrieval, or how we call it internally, closing the oracle gap with knowledge agents.
  2. 0:10 I'm Hannah, I'm an AI engineer at Mixper, and I'm leading the agentic search at Mixper.
  3. 0:17 Yeah, I'm Amir.
  4. 0:18 I'm the co-founder of Mixbread.
  5. 0:19 And let's jump into it.
  6. 0:21 So we have seen over the past years with the LLMs just getting better and better that the reasoning capabilities of the models is growing exponential.
  7. 0:31 So just think back how good GPT-3.5 was and how good GPT-5.5 is, right?
  8. 0:37 It's like a clearly exponential curve.
  9. 0:39 But if you look at search over the last 20 years, basically, we see that
  10. 0:45 it's improving but it's improving very very slowly so there's a huge gap between how basically llms and reasoning is evolving and how retrievers is evolving but this retrieval tools
  11. 1:01 are the main access pattern for this reasoning layer to get the right knowledge and to be truly useful beyond code, like in legal work, in finance work, and so on and so forth.
  12. 1:13 Internally, we call this gap happening right now between reasoning and search the knowledge gap.
  13. 1:19 And it's not just one obscure theory of ours.
  14. 1:22 We see it actually with real benchmarks and real tasks that this gap exists in real world as well.
  15. 1:28 Let's take two benchmarks here, BrowseCamp Plus and Office QA Pro.
  16. 1:33 BrowseCamp Plus is like a browsing benchmark where we have pretty complex queries and a corpus of 100,000 documents and we just try to answer these questions.
  17. 1:44 It's like real world deep search task basically.
  18. 1:48 And Office QA Pro, which has the US treasuries of the past 100 years, basically, and we asked a really complex question over this.
  19. 2:00 Office QA Pro was created by Databricks, and BrowseCamp was created by OpenAI, and BrowseCamp Plus is a version of it with a fixed size corpus, while BrowseCamp was over the open web.
  20. 2:11 And we see here's an Oracle performance in dotted lines.
  21. 2:15 And Oracle means what is the maximum theoretical performance of the models if we would put in the right documents with the question.
  22. 2:24 We see for BrowseComp it's 93% and for Office QA Pro it's 64%.
  23. 2:34 and now we take something like codex with its default tools and want to see hey what is the performance these tools get right now and you see there's a sharp drop in the quality of the answers codex produces for browse camp it's nine points and for office qa pro it's eight points so we see that the models are extremely capable if they would get the right documents
  24. 2:58 But if we put them into the noisy corpus, the performance drops sharply, meaning that actually the bottleneck here is not the reasoning.
  25. 3:11 It's actually the access to the right knowledge it needs to answer this question.
  26. 3:16 So we see this knowledge gap in real world.
  27. 3:21 And if you would just drop in way better search using mixed spread, which is a search tool using latent reaction, we can recover most of the performance.
  28. 3:30 So for BrowseCamp, the difference between the Oracle and GPT-5 with mixed spread is just three point.
  29. 3:38 And for Office QA Pro, we even almost completely closed the gap.
  30. 3:41 With giving the model better search tool, we can recover most of the knowledge gap.
  31. 3:47 But looking at the queries the model asks or writes right now to the search system, it gets super, super interesting.
  32. 3:55 So here's an example query we found during some benchmarking, which is Senator, woman questions, billionaires, not a company, then okay, thank you, staff, we'll check hearing.
  33. 4:07 It's basically gibberish, right?
  34. 4:08 If you have a search system which wants to have neural questions or semantic questions and just gets us keywords over keywords, it gets confused.
  35. 4:18 And the reason why the models write this type of queries is they're mostly trained for coding tasks, like for coding agents, which are then optimized for code-based exploration using tools like
  36. 4:30 grab, and grab just tries to find regular expressions, right?
  37. 4:36 So the model tries to just write as many expressions and things in the document to find the right thing.
  38. 4:42 The second thing is the models are trained to use the web pretty efficiently.
  39. 4:47 And to use the web tools properly, which are highly optimized for humans, they try to mimic human-like query patterns.
  40. 4:55 So just keyword on keyword on keyword.
  41. 4:57 And number three, obviously, is the benchmark bias.
  42. 5:00 Most benchmarks we have right now, like beer, nano beer, use caveman-style queries, which are entity-based queries that structurally favor heavily BM-25s.
  43. 5:11 So right now the agent guesses the keywords to actually increase the overlap between the query and documents and can't really use powerful search tools properly.
  44. 5:23 This motivated us to make our own search agent, to teach it to use powerful search properly, especially to use the right search tools for different use cases and to work beyond code search for knowledge work.
  45. 5:40 And most importantly, the agent should be precise, fast and cheap.
  46. 5:44 The first step for building this agent was to define a very powerful harness.
  47. 5:50 Our harness is fully built on the Mixbed platform.
  48. 5:54 Our agent has four main search tools, which are overview search.
  49. 5:59 This is used as a very wide semantic search where the agent receives up to 50 retrieved chunks.
  50. 6:07 And it sees only summaries of the chunk contents to really have just like an overview of the corpus.

Open at this second