read-only demo

Videos UM6sFg_jdlE

RAG is dead, right?? — Kuba Rogut, Turbopuffer

index_state ready data_status ok

AI Engineer· published 2026-06-09· 0:11:13· en-US· indexed 2026-08-10 19:51

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:19, 1 of 1 keyframes kept
  5. Shot 4, 0:19 to 0:48, 1 of 1 keyframes kept
  6. Shot 5, 0:48 to 1:11, 1 of 1 keyframes kept
  7. Shot 6, 1:11 to 1:37, 1 of 1 keyframes kept
  8. Shot 7, 1:37 to 2:09, 1 of 1 keyframes kept
  9. Shot 8, 2:09 to 2:42, 0 of 1 keyframes kept
  10. Shot 9, 2:42 to 3:14, 0 of 1 keyframes kept
  11. Shot 10, 3:14 to 3:45, 1 of 1 keyframes kept
  12. Shot 11, 3:45 to 4:15, 0 of 1 keyframes kept
  13. Shot 12, 4:15 to 4:45, 1 of 1 keyframes kept
  14. Shot 13, 4:45 to 5:13, 1 of 1 keyframes kept
  15. Shot 14, 5:13 to 5:40, 0 of 1 keyframes kept
  16. Shot 15, 5:40 to 6:07, 0 of 1 keyframes kept
  17. Shot 16, 6:07 to 6:29, 0 of 1 keyframes kept
  18. Shot 17, 6:29 to 6:54, 1 of 1 keyframes kept
  19. Shot 18, 6:54 to 7:20, 0 of 1 keyframes kept
  20. Shot 19, 7:20 to 7:46, 0 of 1 keyframes kept
  21. Shot 20, 7:46 to 8:11, 0 of 1 keyframes kept
  22. Shot 21, 8:11 to 8:37, 1 of 1 keyframes kept
  23. Shot 22, 8:37 to 9:02, 1 of 1 keyframes kept
  24. Shot 23, 9:02 to 9:27, 0 of 1 keyframes kept
  25. Shot 24, 9:27 to 9:52, 0 of 1 keyframes kept
  26. Shot 25, 9:52 to 10:17, 0 of 1 keyframes kept
  27. Shot 26, 10:17 to 10:42, 0 of 1 keyframes kept
  28. Shot 27, 10:42 to 10:52, 1 of 1 keyframes kept
  29. Shot 28, 10:52 to 10:58, 1 of 1 keyframes kept
  30. Shot 29, 10:58 to 11:12, 1 of 1 keyframes kept

30 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
114
whisperx 114
chunks
19
from 114 cues
keyframes
17
kept of 30 captured
frames with text
17
471 lines read
chapters
8
from the source metadata
keyframe bytes
3.6 MB
word timings on 114 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 14:35 1m 51s
stt done 2026-08-10 14:37 13s
chunk done 2026-08-10 14:38 0s
text_embed done 2026-08-10 19:51 1s
keyframe done 2026-08-10 14:38 55s
ocr done 2026-08-10 14:38 9s
frame_embed done 2026-08-10 19:51 3s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 662.1

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 816.0

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:12 #2 done3 line(s)

    shot 2·sharpness 905.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.94
  • 0:18 #3 done4 line(s)

    shot 3·sharpness 271.3

    1. turbopuffer1.00
    2. Rag is dead, right??0.98
    3. Kuba Rogut1.00
    4. AIF0.86
  • 0:23 #4 done12 line(s)

    shot 4·sharpness 1303.3

    1. turbopuffer1.00
    2. Rag is dead, right??0.99
    3. AIE1.00
    4. How hybrid, tool-rich retrieval is becoming the default1.00
    5. 1.00
    6. 1.00
    7. 0.99
    8. for serious agentic search1.00
    9. Kuba Rogut0.99
    10. Deployed Engineer1.00
    11. April 9, 20260.97
    12. Google DeepMind1.00
  • 1:02 #5 done108 line(s)

    shot 5·sharpness 3541.0

    1. "RAG is dead"0.99
    2. turbopuffer1.00
    3. RAG is dead that's why we actually made sure the model could simply use1.00
    4. Replying to @dinOs1.00
    5. Antoine Chaffin@antoine_chaffin · Feb 120.97
    6. Pratik Desai@chheplo - Apr 29, 20240.97
    7. RAG is a makeshift solution.1.00
    8. RAG is bad.1.00
    9. σ ...0.52
    10. swyx@swyx·Oct 13, 20250.95
    11. relatively consensus now that "traditional embeddings RAG" is dead.0.99
    12. @jerryjliu0 called it first at the 2024 AIE WF - Agentic RAG was clearly0.96
    13. grep instead of leveraging our models0.98
    14. RAG will die soon.1.00
    15. better in almost every way except speed.1.00
    16. Praveen Naik @p_naix · Dec 29, 20250.95
    17. O10.68
    18. That way we made sure RAG wouldn't try to revive1.00
    19. O70.65
    20. d 2220.86
    21. 0.54
    22. 651.00
    23. Manthan Gupta@manthanguptaa · Feb 250.95
    24. After "RAG is dead* we are now in the "Agents only need filesystems* Al0.98
    25. t2 250.83
    26. 3060.99
    27. th1 64K0.81
    28. 0.57
    29. Q ….0.52
    30. will be sharing more about the journey to Agentic Search and what's next at1.00
    31. @elastic's sf closing0.99
    32. Show more1.00
    33. It's funny how we jokingly said RAG is dead and it actuall died0.98
    34. hype cycle1.00
    35. 0.76
    36. 1.00
    37. 1.00
    38. AIE1.00
    39. 1.00
    40. 1.00
    41. 1.00
    42. MANTU KUMAR0.97
    43. @Mantu_kumar911.00
    44. O10.67
    45. t20.91
    46. h 1020.83
    47. 0.72
    48. Eli Mernit@mernit · Feb 240.91
    49. JUST USE FILES.1.00
    50. SAN FRANCISCO . OCT 210.95
    51. Keynote: The0.99
    52. elastic0.99
    53. Year Agents Ate0.99
    54. 1.00
    55. 1.00
    56. 1.00
    57. Skill sikhte hi dead ho jaata hai0.98
    58. 0.03-0.93
    59. Search1.00
    60. RAGisDead1.00
    61. SHAWN WANG1.00
    62. X Article0.98
    63. Today, nearly every agent framework is using some mix of Postgres,0.99
    64. Redis, and a vector DB to manage context. This is a mistake.1.00
    65. Agents Don't Need Databases0.99
    66. Rag is dead. Spinning up 20 Codex spark subagents to search your0.99
    67. jason liu@jxnlco - Mar 50.96
    68. filesystem is the solution.0.98
    69. 601.00
    70. t7 140.88
    71. 4081.00
    72. da 44K0.84
    73. Q0.53
    74. Boris Cherny1.00
    75. 0 .0.50
    76. @bcherny1.00
    77. Doug Turnbull@softwaredoug · Mar 30.95
    78. 0 .0.66
    79. Early versions of Claude Code used RAG + a local vector db, but we0.99
    80. RAG is dead again1.00
    81. Piyush Garg1.00
    82. found pretty quickly that agentic search generally works better. It is also1.00
    83. simpler and doesn't have the same issues around security, privacy,1.00
    84. staleness, and relliability.0.99
    85. GPT-5.4 leak: 2M token context + persistent state = KV cache explosion0.99
    86. Ben Pouladian@benitoz ·Mar 10.97
    87. hesam@Hesamation · Oct 15, 20250.95
    88. people are posting "RAG is dead" and half are constantly talking about it.0.99
    89. I don't understand why everyone is so obsessed with RAG. Half of the0.98
    90. It's 2025. Shouldn't RAG be done-and-dusted by now?0.99
    91. 9:56 PM - Jan 31, 2026 - 1.1M Views0.97
    92. Q 1510.79
    93. t23770.91
    94. 5x0.65
    95. 口2.6K0.86
    96. This is the Memory Wars in real time0.97
    97. HBM for weights. SRAM for latency-critical inference. Optical..0.98
    98. Q50.85
    99. t0.86
    100. ○90.66
    101. dt 1.8K0.92
    102. 0.64
    103. Relevant1.00
    104. View quotes>0.95
    105. RAGis Dead0.96
    106. Braintrust1.00
    107. WorkOS1.00
    108. OpenAI0.97
  • 1:14 #6 done24 line(s)

    shot 6·sharpness 1656.5

    1. "RAG is dead"?0.99
    2. turbopuffer1.00
    3. Retrieval Augmented Generation1.00
    4. Search term1.00
    5. +0.96
    6. United States1.00
    7. Mar 1, 2021 - Feb 27, 20260.99
    8. Web Search1.00
    9. AIE1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. 1.00
    14. Interest over time ①0.95
    15. United States - Mar 1, 2021 - Feb 27, 20260.98
    16. 1001.00
    17. MGoogle Tronds0.94
    18. 20221.00
    19. 20231.00
    20. 20240.99
    21. 20250.99
    22. 20261.00
    23. AlEngineer0.97
    24. EUROPE1.00
  • 1:59 #7 done29 line(s)

    shot 7·sharpness 4241.3

    1. Let's clarify: RAG vs Agentic Search0.99
    2. turbopuffer1.00
    3. RAG1.00
    4. AgenticSearch1.00
    5. What people *think* it is:0.99
    6. What people *think* it is:0.98
    7. 4 Vector search0.96
    8. 4Filesystem grep0.99
    9. 0.90
    10. AIE1.00
    11. 1.00
    12. What it *actually* is:0.97
    13. What it *actually* is:0.99
    14. 1.00
    15. 1.00
    16. Retrieval Augmented Generation1.00
    17. giving agents a set of tools to0.99
    18. progressively find and reason over1.00
    19. vector1.00
    20. LLM1.00
    21. external context0.97
    22. bm251.00
    23. harness1.00
    24. grep1.00
    25. glob1.00
    26. regex1.00
    27. filter...0.99
    28. AlEngineer0.96
    29. EUROPE1.00
  • 2:32 #8 skipped

    shot 8·duplicate of #7

  • 3:04 #9 skipped

    shot 9·duplicate of #7

  • 3:41 #10 done55 line(s)

    shot 10·sharpness 4577.8

    1. cursor.com/blog/secure-codebase-indexing1.00
    2. turbopuffer1.00
    3. CURSOR1.00
    4. Merkle trees enable fast, incremental sync1.00
    5. Securely indexing large codebases0.99
    6. & Client0.92
    7. Repository service1.00
    8. Jan 27, 2026 by Jeremy Stribling in Research1.00
    9. /docs #g7n2d20.97
    10. compare hashes1.00
    11. of entries1.00
    12. /docs #w5g3u20.97
    13. 1.00
    14. guide.md #b8c9d00.96
    15. sync0.93
    16. guide.md#x9y8z70.93
    17. AIE1.00
    18. 0.99
    19. diagram.py#z8s7130.97
    20. …no updates needed…0.93
    21. diagram,py#z8s7130.97
    22. 1.00
    23. Semantic search is one of the biggest drivers of agent performance. In our recent1.00
    24. 1.00
    25. 1.00
    26. 1.00
    27. evaluation, it improved response accuracy by 12.5% on average, produced code changes0.99
    28. that were more likely to be retained in codebases, and raised overall request satisfaction.1.00
    29. No matching on client1.00
    30. × delete0.94
    31. config.json #z1x2y30.91
    32. To power semantic search, Cursor builds a searchable index of your codebase when you0.99
    33. open a project. For small projects, this happens almost instantly. But large repositories with1.00
    34. Secure index sharing enables fast, safe onboarding0.98
    35. tens of thousands of files can take hours to process if indexed naively, and semantic0.99
    36. if hashes differ0.98
    37. search isn't available until at least 80% of that work is finished.0.99
    38. & Client0.96
    39. Indexing worker1.00
    40. 目 Vector DB0.90
    41. We looked for ways to speed up indexing based on the simple observation that most teams1.00
    42. work from near-identical copies of the same codebase. In fact, clones of the same0.99
    43. Merkle trees0.98
    44. Embedding model1.00
    45. → DynamoDB cache0.94
    46. codebase average 92% similarity across users within an organization.0.99
    47. Repository service1.00
    48. localCodebase DB0.98
    49. This means that rather than rebuilding every index from scratch when someone joins or0.99
    50. Secure index copies0.98
    51. switches machines, we can securely reuse a teammate's existing index. This cuts time-to-0.99
    52. if hashes match0.95
    53. first-query from hours to seconds on the largest repos.1.00
    54. AlEngineer0.96
    55. EUROPE1.00
  • 4:08 #11 skipped

    shot 11·duplicate of #10

  • 4:27 #12 done55 line(s)

    shot 12·sharpness 4602.9

    1. cursor.com/blog/secure-codebase-indexing1.00
    2. turbopuffer1.00
    3. CURSOR1.00
    4. Merkle trees enable fast, incremental sync1.00
    5. Securely indexing large codebases0.99
    6. & Client0.92
    7. Repository service1.00
    8. 1.00
    9. Jan 27, 2026 by Jeremy Stribling in Research1.00
    10. /docs #g7n2d20.97
    11. compare hashes1.00
    12. of entries0.96
    13. /docs #w5g3u20.97
    14. 1.00
    15. guide.md #b8c9d00.96
    16. sync0.94
    17. guide.md #x9y8z70.93
    18. AIE1.00
    19. 1.00
    20. diagram.py#z8s7130.99
    21. …no updates needed0.95
    22. diagram.py #z8s7130.94
    23. Semantic search is one of the biggest drivers of agent performance. In our recent0.99
    24. 1.00
    25. 1.00
    26. 1.00
    27. 1.00
    28. evaluation, it improved response accuracy by 12.5% on average, produced code changes0.99
    29. that were more likely to be retained in codebases, and raised overall request satisfaction.1.00
    30. No matching on client1.00
    31. × delete0.95
    32. config.json #z1x2y30.94
    33. To power semantic search, Cursor builds a searchable index of your codebase when you0.99
    34. open a project. For small projects, this happens almost instantly. But large repositories with1.00
    35. Secure index sharing enables fast, safe onboarding0.98
    36. tens of thousands of files can take hours to process if indexed naively, and semantic0.99
    37. if hashes differ0.98
    38. search isn't available until at least 80% of that work is finished.0.99
    39. & Client0.95
    40. Indexing worker0.99
    41. 目 Vector DB0.94
    42. We looked for ways to speed up indexing based on the simple observation that most teams1.00
    43. work from near-identical copies of the same codebase. In fact, clones of the same0.99
    44. Merkle troes0.96
    45. Embedding model1.00
    46. → 目 DynamoDB cache0.93
    47. codebase average 92% similarity across users within an organization.0.99
    48. Repository service1.00
    49. localCodebase DB0.98
    50. This means that rather than rebuilding every index from scratch when someone joins or0.99
    51. Secure index copies0.99
    52. switches machines, we can securely reuse a teammate's existing index. This cuts time-to-0.99
    53. it hashes match0.95
    54. first-query from hours to seconds on the largest repos.1.00
    55. Engineering the future of Al1.00
  • 5:07 #13 done52 line(s)

    shot 13·sharpness 3859.1

    1. cursor.com/blog/semsearch1.00
    2. turbopuffer1.00
    3. CURSOR1.00
    4. All models improve with semantic search0.99
    5. Improving agent with semantic search1.00
    6. Model (alphabetical)1.00
    7. Relative improvement (Cursor Context Bench)0.99
    8. Nov 6, 2025 by Stefan Heule, Emily Jia & Naman Jain in Research0.99
    9. Composer1.00
    10. 23.5%1.00
    11. 1.00
    12. Table of Contents1.00
    13. Gemini 2.5 Pro1.00
    14. 8.7%1.00
    15. AIE1.00
    16. 1.00
    17. Offline evals1.00
    18. GPT-51.00
    19. 6.5%1.00
    20. 1.00
    21. Online A/B tests0.94
    22. Custom retrieval models0.99
    23. Grok Code1.00
    24. 11.9%1.00
    25. 1.00
    26. 1.00
    27. 1.00
    28. Conclusion1.00
    29. Sonnet 4.51.00
    30. 14.7%1.00
    31. When coding agents receive a prompt, returning the right answer requires building an0.99
    32. understanding of the codebase by reading files and searching for relevant information.0.99
    33. Semantic search improves code retention1.00
    34. matching natural language queries, such as "where do we handle authentication?", in0.99
    35. One tool Cursor's agent uses is semantic search, which retrieves segments of code0.99
    36. and reduces dissatisfied user requests0.99
    37. addition to the regex-based searching provided by a tool like grep.1.00
    38. Code Retention1.00
    39. +0.3%1.00
    40. pipelines for fast retrieval. While you could rely exclusively on grep and similar command-1.00
    41. To support semantic search, we've trained our own embedding model and built indexing0.99
    42. Code Retention (large codebases)1.00
    43. +2.6%1.00
    44. line tools for search, we've found that semantic search significantly improves agent1.00
    45. Dissatisfied User Requests1.00
    46. -2.2%1.00
    47. performance, especially over large codebases:0.99
    48. Achieving on average 12.5% higher accuracy in answering questions (6.5%-23.5%0.99
    49. Improvement of having semantic search (vs same model without semantic search).0.99
    50. For "Code Retention" higher is better, and for "Dissatisfied User Requests" lower is better.0.99
    51. depending on the model).1.00
    52. Engineering the future of Al1.00
  • 5:16 #14 skipped

    shot 14·duplicate of #13

  • 5:46 #15 skipped

    shot 15·duplicate of #13

  • 6:12 #16 skipped

    shot 16·duplicate of #6

  • 6:51 #17 done41 line(s)

    shot 17·sharpness 4348.6

    1. embeddings are cached compute1.00
    2. per-sessiondiscovery1.00
    3. amortized understanding1.00
    4. Agent greps, reads, assesses, repeats0.99
    5. index once, retrieve at runtime1.00
    6. grep "metadata filter" src/1.00
    7. +820 tokens1.00
    8. Index time:1.00
    9. 1.00
    10. read indexing.py (wrong section)1.00
    11. +2,106 tokens1.00
    12. - Codebase chunked, embedded, indexed0.98
    13. 1.00
    14. AIE1.00
    15. grep "ingest pipeline" src/1.00
    16. +640 tokens1.00
    17. - Semantic meaning encoded upfront1.00
    18. 1.00
    19. 1.00
    20. read api_client.ts:280-4001.00
    21. +1,450 tokens1.00
    22. - Model has already "read" every file0.99
    23. 1.00
    24. 1.00
    25. read types.ts (for context)1.00
    26. +1,298 tokens1.00
    27. - One-time cost, amortized across all sessions0.99
    28. Runtime:1.00
    29. Repeated every session by every0.99
    30. Agent query: "How is metadata filtered?"0.98
    31. agent, across every task1.00
    32. Results(ranked chunks):1.00
    33. → 6,314 tokens0.96
    34. - indexing.py1.00
    35. +142 tokens1.00
    36. - api_client.ts0.99
    37. +188 tokens0.96
    38. types.ts1.00
    39. +94 tokens1.00
    40. → 424 tokens0.97
    41. Engineering the future of Al1.00
  • 7:00 #18 skipped

    shot 18·duplicate of #17

  • 7:40 #19 skipped

    shot 19·duplicate of #17

  • 7:59 #20 skipped

    shot 20·duplicate of #17

  • 8:27 #21 done43 line(s)

    shot 21·sharpness 4325.7

    1. embeddings are cached compute1.00
    2. per-session discovery1.00
    3. amortized understanding1.00
    4. Agent greps, reads, assesses, repeats0.99
    5. index once, retrieve at runtime0.99
    6. grep "metadata filter" src/1.00
    7. +820 tokens1.00
    8. Index time:1.00
    9. 1.00
    10. 1.00
    11. read indexing.py (wrong section)0.99
    12. +2,106 tokens1.00
    13. - Codebase chunked, embedded, indexed0.98
    14. 1.00
    15. AIE1.00
    16. grep "ingest pipeline" src/0.99
    17. +640 tokens1.00
    18. - Semantic meaning encoded upfront0.99
    19. 1.00
    20. 1.00
    21. read api_client.ts:280-4000.98
    22. +1,450 tokens0.98
    23. - Model has already "read" every file0.99
    24. 1.00
    25. 1.00
    26. read types.ts (for context)1.00
    27. +1,298 tokens1.00
    28. - One-time cost, amortized across all sessions0.99
    29. Runtime:1.00
    30. Repeated every session by every0.99
    31. Agent query: "How is metadata filtered?"1.00
    32. agent, across every task1.00
    33. Results(ranked chunks):1.00
    34. → 6,314 tokens0.97
    35. - indexing.py1.00
    36. +142 tokens0.99
    37. - api_client.ts0.99
    38. +188 tokens0.96
    39. types.ts1.00
    40. +94 tokens1.00
    41. → 424 tokens0.98
    42. AlEngineer0.95
    43. EUROPE1.00
  • 8:57 #22 done28 line(s)

    shot 22·sharpness 3853.2

    1. from RAG to agentic retrieval0.99
    2. turbopuffer1.00
    3. Haider.1.00
    4. Subscribe1.00
    5. @slow_developer1.00
    6. pattern1.00
    7. Google Jeff Dean says bigger context windows alone are not enough0.99
    8. From RAG → Agentic Retrieval0.98
    9. What matters is staged retrieval: lightweight mechanisms that narrow a0.99
    10. trillion tokens down to 10 million, then to the million you actually need1.00
    11. AIE1.00
    12. old1.00
    13. new1.00
    14. "you don't need a trillion at once, you need the right million"0.99
    15. 1.00
    16. 1.00
    17. 1.00
    18. 1.00
    19. @slow_developer1.00
    20. Retrieve once1.00
    21. Reason in steps1.00
    22. Stuff the prompt1.00
    23. Search as needed1.00
    24. Cross fingers1.00
    25. - Fetch what is useful0.99
    26. Retrieval is now iterative with tools0.99
    27. AlEngineer0.97
    28. EUROPE1.00
  • 9:19 #23 skipped

    shot 23·duplicate of #22

Transcript

114 cues· 2,014 words· 10,912 chars

  1. 0:14 Hi, welcome, everyone.
  2. 0:16 Thanks for coming out.
  3. 0:17 I see it's a full room, so I appreciate everyone coming out.
  4. 0:20 So welcome to the talk about RAG is dead, right?
  5. 0:23 So my name is Kuba.
  6. 0:24 I'm a deployed engineer at TurboPuffer.
  7. 0:26 So for those that don't know what TurboPuffer is, we are a full-text search and vector search database built from first principles on top of object storage.
  8. 0:34 If you would love to learn more, just come find me after the talk if you have any questions.
  9. 0:38 So let's get started.
  10. 0:39 So this talk is about how RAG is dead, how hybrid tool-rich retrieval is becoming a default for serious agentic search.
  11. 0:49 So if you guys have been on Twitter or other social media platforms, or I guess X as they call it now, you might have seen a lot of tweets like this about how RAG is dead.
  12. 0:57 You can see there's lots of tweets, especially in the end of 2025 and in the early of this year, about how RAG is dead, agentic file search is all we need, and there's kind of a lot of tweets and a lot of content about this now.
  13. 1:12 Interestingly, if you were to look at something like the Google search volume over the last two years, or the last couple of years, you can see that in 2023, kind of as AI starts, we have kind of this increase, kind of caps out a little bit in 2024, settles down for about a year.
  14. 1:28 And then about midway through 2025, we hit this new inflection point where search volume just goes through the roof.
  15. 1:35 So take that, Twitter.
  16. 1:38 So let's clarify first, what is RAG and what is agentic search?
  17. 1:41 These are the two terms a lot of people are throwing out these days.
  18. 1:44 So RAG, what a lot of people think RAG is is just simple vector search.
  19. 1:48 They just think that this is just simply embedding a bunch of corpus of contents, passing an embedding vector, and getting it back, passing it through your LLM.
  20. 1:57 And at Turbopuffer, what we think this actually means, if you break down RAG into retrieval, augmented generation, retrieval is not just vector search.
  21. 2:06 It's a lot of different things.
  22. 2:07 It could be vector search, full text search using stuff like bm25, grepping, globbing, using regex, using other just basic filters.
  23. 2:15 And the augmented generation is obviously just passing it into your LLM of choice.
  24. 2:20 And then agentic search.
  25. 2:21 This is kind of the terms people are throwing around a lot these days.
  26. 2:25 And generally, when people start talking about agentic search, what they usually talk about is essentially just file system grep.
  27. 2:31 So if you guys are familiar with something like Cloud Code and Cloud Code Codex, a lot of people call this agentic search.
  28. 2:39 And this just essentially is grepping through your file system.
  29. 2:42 And this is kind of why these terms are so correlated.
  30. 2:45 And what we actually believe it is and, you know, kind of the definition we want to give it is it's really giving the agents a set of tools to kind of progressively and iteratively find and reason over context.
  31. 2:55 So with Cloud Code, you can, you know, if you guys are familiar with it, it can read your file, start repping through your file system, read a file, decide that it hasn't found what it needed to actually complete the task, and it will, you know, find something again and then keep doing this until it's happy, you know, it's reached a happy state where it can continue on with the task.
  32. 3:15 So we're going to take a step back and talk about one of the companies that use TurboPuffer that we believe is doing an excellent job with agentic search.
  33. 3:23 This is a company called Cursor.
  34. 3:24 You might have heard of them.
  35. 3:26 Fun fact, they're actually one of TurboPuffer's very first customers.
  36. 3:30 And they have this excellent blog post that came out in the beginning of 2026 about how they index codebases.
  37. 3:35 So for those unaware, when you open up a new code base or a new branch in Cursor, what happens is that Cursor will start embedding your code base.
  38. 3:42 So what they'll do is chunk out your parse, chunk and embed your code base, and make it available for semantic search.
  39. 3:47 And this blog post goes into an excellent technical detail of how they do this.
  40. 3:53 Just to give you the gist, essentially the cool thing they do
  41. 3:56 is that they found that most people working on a team, let's say there's 100 engineers, when they open up codebases, they're normally the same codebase 99% of the time because you can have a team of 100 people most of the time working on one, two, maybe a few codebases, right?
  42. 4:11 And it's really expensive to have to re-chunk, re-embed, and re-upload these codebases every single time.
  43. 4:18 So they essentially use Merkle trees, which essentially is this crypto hash tree, to calculate some layers between code bases people open on the same team.
  44. 4:26 And if they're similar enough, they will essentially copy over the data and then only update every chunk and re-embed the files that have changed and use Turbo Puffer in order to make sure this is done securely.
  45. 4:38 And yeah, just excellent blog posts.
  46. 4:41 They do some really cool stuff.
  47. 4:43 And you may think, this is a lot of work.
  48. 4:45 Why do they do this?
  49. 4:46 Well, the reason they do this is also covered in a different blog post about how they use semantic search.
  50. 4:51 Again, they use Turbo Puffer for this.

Chapters

  1. 0:00 Introduction to the "RAG is dead" discourse
  2. 1:12 Google search volume trends for RAG
  3. 1:39 Defining RAG vs. Agentic Search
  4. 3:15 Cursor's indexing and semantic search approach
  5. 6:10 Contrasting Claude Code (grep) vs. Cursor (indexed)
  6. 6:40 The concept of embeddings as cached compute
  7. 8:38 The shift from simple RAG to Agentic Retrieval
  8. 9:44 Jeff Dean on context windows and stage retrieval

Open at this second