read-only demo

Videos Jx4ZFEAq6bY

User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:15:37· en-US· indexed 2026-08-11 08:07

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:30, 1 of 1 keyframes kept
  2. Shot 1, 0:30 to 0:56, 1 of 1 keyframes kept
  3. Shot 2, 0:56 to 1:22, 1 of 1 keyframes kept
  4. Shot 3, 1:22 to 1:32, 1 of 1 keyframes kept
  5. Shot 4, 1:32 to 1:48, 1 of 1 keyframes kept
  6. Shot 5, 1:48 to 2:15, 1 of 1 keyframes kept
  7. Shot 6, 2:15 to 2:20, 1 of 1 keyframes kept
  8. Shot 7, 2:20 to 2:57, 1 of 1 keyframes kept
  9. Shot 8, 2:57 to 3:15, 1 of 1 keyframes kept
  10. Shot 9, 3:15 to 3:19, 1 of 1 keyframes kept
  11. Shot 10, 3:19 to 3:34, 1 of 1 keyframes kept
  12. Shot 11, 3:34 to 3:45, 1 of 1 keyframes kept
  13. Shot 12, 3:45 to 4:15, 1 of 1 keyframes kept
  14. Shot 13, 4:15 to 4:45, 0 of 1 keyframes kept
  15. Shot 14, 4:45 to 5:13, 1 of 1 keyframes kept
  16. Shot 15, 5:13 to 5:39, 1 of 1 keyframes kept
  17. Shot 16, 5:39 to 6:09, 1 of 1 keyframes kept
  18. Shot 17, 6:09 to 6:40, 0 of 1 keyframes kept
  19. Shot 18, 6:40 to 6:43, 1 of 1 keyframes kept
  20. Shot 19, 6:43 to 7:09, 1 of 1 keyframes kept
  21. Shot 20, 7:09 to 7:35, 0 of 1 keyframes kept
  22. Shot 21, 7:35 to 8:02, 0 of 1 keyframes kept
  23. Shot 22, 8:02 to 8:05, 1 of 1 keyframes kept
  24. Shot 23, 8:05 to 8:18, 1 of 1 keyframes kept
  25. Shot 24, 8:18 to 8:50, 1 of 1 keyframes kept
  26. Shot 25, 8:50 to 9:19, 1 of 1 keyframes kept
  27. Shot 26, 9:19 to 9:47, 0 of 1 keyframes kept
  28. Shot 27, 9:47 to 10:13, 1 of 1 keyframes kept
  29. Shot 28, 10:13 to 10:36, 1 of 1 keyframes kept
  30. Shot 29, 10:36 to 10:40, 1 of 1 keyframes kept
  31. Shot 30, 10:40 to 10:44, 1 of 1 keyframes kept
  32. Shot 31, 10:44 to 10:48, 1 of 1 keyframes kept
  33. Shot 32, 10:48 to 10:50, 1 of 1 keyframes kept
  34. Shot 33, 10:50 to 11:06, 0 of 1 keyframes kept
  35. Shot 34, 11:06 to 11:21, 1 of 1 keyframes kept
  36. Shot 35, 11:21 to 11:55, 0 of 1 keyframes kept
  37. Shot 36, 11:55 to 12:03, 0 of 1 keyframes kept
  38. Shot 37, 12:03 to 12:27, 1 of 1 keyframes kept
  39. Shot 38, 12:27 to 12:38, 1 of 1 keyframes kept
  40. Shot 39, 12:38 to 13:02, 1 of 1 keyframes kept
  41. Shot 40, 13:02 to 13:11, 1 of 1 keyframes kept
  42. Shot 41, 13:11 to 13:13, 0 of 1 keyframes kept
  43. Shot 42, 13:13 to 13:24, 1 of 1 keyframes kept
  44. Shot 43, 13:24 to 13:25, 1 of 1 keyframes kept
  45. Shot 44, 13:25 to 13:27, 1 of 1 keyframes kept
  46. Shot 45, 13:27 to 13:29, 1 of 1 keyframes kept
  47. Shot 46, 13:29 to 13:30, 1 of 1 keyframes kept
  48. Shot 47, 13:30 to 13:32, 0 of 1 keyframes kept
  49. Shot 48, 13:32 to 13:33, 0 of 1 keyframes kept
  50. Shot 49, 13:33 to 13:35, 1 of 1 keyframes kept
  51. Shot 50, 13:35 to 13:37, 0 of 1 keyframes kept
  52. Shot 51, 13:37 to 13:41, 0 of 1 keyframes kept
  53. Shot 52, 13:41 to 13:44, 1 of 1 keyframes kept
  54. Shot 53, 13:44 to 13:45, 0 of 1 keyframes kept
  55. Shot 54, 13:45 to 13:48, 0 of 1 keyframes kept
  56. Shot 55, 13:48 to 13:49, 0 of 1 keyframes kept
  57. Shot 56, 13:49 to 13:56, 1 of 1 keyframes kept
  58. Shot 57, 13:56 to 14:02, 0 of 1 keyframes kept
  59. Shot 58, 14:02 to 14:05, 0 of 1 keyframes kept
  60. Shot 59, 14:05 to 14:07, 1 of 1 keyframes kept
  61. Shot 60, 14:07 to 14:39, 1 of 1 keyframes kept
  62. Shot 61, 14:39 to 15:25, 0 of 1 keyframes kept
  63. Shot 62, 15:25 to 15:36, 1 of 1 keyframes kept
  64. Shot 63, 15:36 to 15:36, 1 of 1 keyframes kept

64 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
107
whisperx 107
chunks
27
from 107 cues
keyframes
45
kept of 64 captured
frames with text
44
970 lines read
chapters
0
from the source metadata
keyframe bytes
2.2 MB
word timings on 107 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 08:05 1m 11s
stt done 2026-08-11 08:06 14s
chunk done 2026-08-11 08:06 0s
text_embed done 2026-08-11 08:06 0s
keyframe done 2026-08-11 08:06 19s
ocr done 2026-08-11 08:07 14s
frame_embed done 2026-08-11 08:07 1s

Frames, and what the machine read

  • 0:23 #0 done4 line(s)

    shot 0·sharpness 716.7

    1. User Signal dies at the1.00
    2. Retrieval Boundary.1.00
    3. Sonam Pankaj1.00
    4. CEO & Co-Founder StarlightSearch0.99
  • 0:53 #1 done6 line(s)

    shot 1·sharpness 1938.0

    1. What is an agent?1.00
    2. 0.63
    3. An agent is an LLM with agency: it reasons, invokes0.99
    4. TOOLS to interact with the world, RETRIEVES from0.98
    5. memory, and loops until the TASK is complete.0.99
    6. agentRTX0.97
  • 1:02 #2 done13 line(s)

    shot 2·sharpness 2926.7

    1. What is an agent?1.00
    2. 0.51
    3. An agent is an LLM with agency: it reasons, invokes0.99
    4. TOOLS to interact with the world, RETRIEVES from1.00
    5. memory, and loops until the TASK is complete.0.99
    6. React Agent0.96
    7. Tools/1.00
    8. 80.99
    9. LLM1.00
    10. Execute1.00
    11. →Retrieval/0.99
    12. ⅡI0.52
    13. websearch1.00
  • 1:29 #3 done5 line(s)

    shot 3·sharpness 1450.8

    1. 0.80
    2. Agents keep failing at the same tasks.1.00
    3. Gartner's 2025 Al deployment survey found that 85% of Al projects fail in production. McKinsey's 20250.99
    4. State of Al report found that fewer than 20% of Al pilots scale to production within 18 months.0.99
    5. agentRTX0.99
  • 1:40 #4 done2 line(s)

    shot 4·sharpness 487.7

    1. Retrieval is Static1.00
    2. Context Stuffing1.00
  • 1:51 #5 done13 line(s)

    shot 5·sharpness 880.3

    1. Ram Sriharsha - 1st0.91
    2. 0.95
    3. Ex-CTO @ Pinecone | Researching what happens when you stop scaling and s..0.99
    4. 2mo·0.97
    5. If you're building Al agents today, I have a confession: your agent's memory is0.99
    6. probably broken and you are probably paying too much.0.99
    7. Why is the world's most advanced coding agent (Anthropic's Claude Code) using1.00
    8. grep. a tool from 1973 instead of vector search?0.99
    9. Because we've been optimizing for the wrong things. We made the wrong answers0.99
    10. faster and cheaper, but we forgot to make retrieval learn.0.99
    11. I wrote about why the industry overcorrected, why "context stuffing is a money pit.0.99
    12. and what actually comes next for agentic memory.1.00
    13. agentRTX0.91
  • 2:17 #6 done3 line(s)

    shot 6·sharpness 851.2

    1. Retrieval is Static1.00
    2. Context Stuffing1.00
    3. Agents are not outcome-informed.1.00
  • 2:35 #7 done17 line(s)

    shot 7·sharpness 1510.2

    1. The missing layer between evals and action1.00
    2. Observability1.00
    3. Evals1.00
    4. Agent1.00
    5. The GAP1.00
    6. stack captures every1.00
    7. Your observability1.00
    8. Your eval suite judges1.00
    9. ←→0.95
    10. Context, Skills1.00
    11. tool call, every LLM0.99
    12. whether the final output1.00
    13. was correct.1.00
    14. and .md file1.00
    15. ompletion, and1.00
    16. very exception.1.00
    17. agentRTX0.98
  • 3:09 #8 done6 line(s)

    shot 8·sharpness 2417.1

    1. The GAP1.00
    2. It has no access to why yesterday's runs passed or1.00
    3. failed. The eval signal dies in a dashboard.0.99
    4. This is the missing layer: a system that consumes0.99
    5. traces, absorbs eval outcomes, and converts both0.99
    6. into retrievable guidance for future runs.1.00
  • 3:17 #9 done5 line(s)

    shot 9·sharpness 1515.8

    1. The Manual Improvement TAX1.00
    2. Rewrite Prompts and redeploy1.00
    3. Upgrade to expensive models0.98
    4. Restructure tool-call harness1.00
    5. tune custom models0.97
  • 3:31 #10 done5 line(s)

    shot 10·sharpness 1828.8

    1. The Manual Improvement TAX1.00
    2. Rewrite Prompts and redeploy0.99
    3. Upgrade to expensive models0.97
    4. Restructure tool-call harness1.00
    5. Fine-tune custom models1.00
  • 3:40 #11 done2 line(s)

    shot 11·sharpness 800.7

    1. Why are current memories1.00
    2. failing?1.00
  • 4:06 #12 done25 line(s)

    shot 12·sharpness 2636.0

    1. In practice, most agent memory frameworks have focused on user continuity: preferences,0.99
    2. profile facts, conversation history, and long-lived personalization.0.99
    3. chat experiences is not self-improving learning system for production agents.0.99
    4. 0.63
    5. Approach1.00
    6. What it stores0.97
    7. Retrieval signal1.00
    8. Learns from outcomes?1.00
    9. Raw chat history1.00
    10. Recency1.00
    11. No0.98
    12. Extracted facts, preferences Embedding similarity0.99
    13. No1.00
    14. Entity relationships over time Graph traversal • recency0.98
    15. No0.99
    16. Verbal self-reflections0.96
    17. Similarity to current task0.99
    18. Partially: reflections capture0.99
    19. lessons, but retrieval is not0.98
    20. ranked by outcome0.99
    21. agentRTX1.00
    22. Task-tinked reflections with0.99
    23. utility scores1.00
    24. Similarity weighted by0.99
    25. outcome-derived utility1.00
  • 4:39 #13 skipped

    shot 13·duplicate of #12

  • 4:54 #14 done4 line(s)

    shot 14·sharpness 1801.9

    1. agentRTX1.00
    2. agents with runtime experience1.00
    3. It's a runtime learning layer that lets production agents improve from0.99
    4. rience without retraining, fine-tuning, or manual prompt engineering.1.00
  • 5:35 #15 done6 line(s)

    shot 15·sharpness 2080.0

    1. 0.96
    2. Introduces UtilityScore.1.00
    3. You do not retrieve by keyword.0.98
    4. You retrieve by semantic similarity to the current task, weighted by whether those1.00
    5. emories have historically helped or hurt. The eval outcome becomes a first-class signal0.99
    6. in the retrieval ranking, not just a post-mortem footnote.1.00
  • 5:57 #16 done12 line(s)

    shot 16·sharpness 1618.0

    1. Memory as Reasoning1.00
    2. facts1.00
    3. Reasoning1.00
    4. User preferences1.00
    5. "Check settlement before0.98
    6. Refund"0.94
    7. Static,1.00
    8. Reranked based on usefulness1.00
    9. No Context1.00
    10. Context is updated based on task1.00
    11. No history0.97
    12. learned from history0.99
  • 6:19 #17 skipped

    shot 17·duplicate of #16

  • 6:43 #18 done2 line(s)

    shot 18·sharpness 319.1

    1. 0.56
    2. Benchmarks1.00
  • 6:56 #19 done8 line(s)

    shot 19·sharpness 713.3

    1. ref/ect1.00
    2. 0.97
    3. τ2-bench0.93
    4. from baseline to Reflect-enabled runs and then skit-0.98
    5. enhanced performance.0.98
    6. Wie- compare two models on the airlines task in 'tl-bench'.0.92
    7. OPTSA0.97
    8. fect0.69
  • 7:25 #20 skipped

    shot 20·duplicate of #19

  • 7:44 #21 skipped

    shot 21·duplicate of #19

  • 8:04 #22 done1 line(s)

    shot 22·sharpness 616.4

    1. ref/ect1.00
  • 8:13 #23 done10 line(s)

    shot 23·sharpness 1014.7

    1. ref/ect1.00
    2. Festures0.92
    3. Howit works0.92
    4. Integrations0.95
    5. Det Sterted0.94
    6. OLM non-thnking0.95
    7. OPTS.40.93
    8. Reftect0.85
    9. HUI0.54
    10. 080.59

Transcript

107 cues· 1,792 words· 9,857 chars

  1. 0:01 Hey, everyone.
  2. 0:01 I'm Sohim.
  3. 0:03 I'm the CEO and co-founder of Starlight Search.
  4. 0:05 And today, my topic is User Signals Die at Retrieval Boundaries.
  5. 0:10 So look into what are agents, essentially why agent fails, what is the cause of fails in retrieval particularly, and how to make actually signals cross the retrieval boundary and how to make your agent basically up compare.
  6. 0:29 So let's get started.
  7. 0:31 What is an agent?
  8. 0:32 An agent is an LLM that has agency to reason, involve tools, interact with the real world, retrieve the memory to complete the task.
  9. 0:43 One major loop here is missing is learning.
  10. 0:47 It should also learn from what worked and what didn't work.
  11. 0:50 Suppose if I have to explain what is agent, I can explain with react-agent.
  12. 0:58 So if I have to explain, uh, what agent is, I have explained with react agent.
  13. 1:04 So basically use it from the agent, uh, execute it in a loop, uh, client tool retrieval search, and then pause when the task is complete.
  14. 1:13 This is very basic react architecture.
  15. 1:18 One thing that is missing is how to make agent learn from the outcome.
  16. 1:24 So agent keeps failing at the same task.
  17. 1:26 Garten reports that 85% of AI projects fail in production.
  18. 1:31 So it's in McKenzie's 2025 report.
  19. 1:34 The problem came out to be most of the time is that retriever is static.
  20. 1:41 73% of RR pipeline fails because of retrieval non-generation and context stuffing.
  21. 1:49 So a recent
  22. 1:51 post from Ram Sriharsha, the ex-CTO of Pinecon said, we have been optimizing for the wrong thing.
  23. 1:59 You are paying a lot for your agent's memory.
  24. 2:02 This is probably broken and we have been optimizing for the wrong things.
  25. 2:05 We made wrong answers appear faster and cheaper that we forgot to make retrieval learn.
  26. 2:12 So why does this matter?
  27. 2:16 Again, the third problem is agents are not out coming for.
  28. 2:21 So there's a missing layer between evals and action.
  29. 2:25 Your observability has all the traces, all the stack that capture, observability is the stack that capture every tool call, every element completion, every exceptions.
  30. 2:37 Your eval suite judges whether the final output was correct or wrong, basically pass or fail.
  31. 2:44 But these evals are not reflected in agent.
  32. 2:51 context, skills, MD files, or agent action in any ways.
  33. 2:59 So the agent doesn't have any access to why yesterday's runs passed or failed.
  34. 3:05 The eva signal dies in the dashboard.
  35. 3:08 This is a missing layer, a system that consume traces, absorb eva, and convert both into retrieval guidance for future runs.
  36. 3:17 So there's a manual improvement task and engineer actually has to sit and see if the email and also will perform well.
  37. 3:25 Relight the prompt, redeploy it, either upgrade to expensive model, restructure to our harness or find you the custom models.
  38. 3:36 Why are current memories failing?
  39. 3:38 Why memories was designed to actually address this, but it's not.
  40. 3:47 So let's see what we have as a current system and current memory is that they basically store user preferences, profile, conversational history, or long-lived personalization.
  41. 4:04 So chat experience is not self-improving learning systems for production.
  42. 4:09 If you see the already existing approach in the market, there's a launching
  43. 4:17 Does mem0, which does extracted fact references, uses retrieval signal, is an embedding similarity?
  44. 4:26 Does it learn from our code?
  45. 4:27 No.
  46. 4:29 So we have come up with something called utility score, which is a similarity weighted by how useful it is for the agent to execute the task.
  47. 4:41 It has actually the history of past cases and past outcomes.
  48. 4:47 So we came up with agent RKX, and that is agents with runtime experience.
  49. 4:51 It's a runtime layer that lets a bunch of agents improve from experiences without retraining, fine-tuning, or manual prompt training.
  50. 5:01 It's a bit different from compile time like DSPy because you bake in all the lessons in the prompt.

Open at this second