read-only demo

Videos dRmWYHuIJxM

We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:10:42· en-US· indexed 2026-08-10 19:54

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:32, 1 of 1 keyframes kept
  2. Shot 1, 0:32 to 1:04, 0 of 1 keyframes kept
  3. Shot 2, 1:04 to 1:48, 1 of 1 keyframes kept
  4. Shot 3, 1:48 to 2:16, 1 of 1 keyframes kept
  5. Shot 4, 2:16 to 2:45, 0 of 1 keyframes kept
  6. Shot 5, 2:45 to 3:19, 1 of 1 keyframes kept
  7. Shot 6, 3:19 to 3:46, 1 of 1 keyframes kept
  8. Shot 7, 3:46 to 4:12, 0 of 1 keyframes kept
  9. Shot 8, 4:12 to 4:39, 0 of 1 keyframes kept
  10. Shot 9, 4:39 to 5:04, 1 of 1 keyframes kept
  11. Shot 10, 5:04 to 5:30, 0 of 1 keyframes kept
  12. Shot 11, 5:30 to 6:01, 1 of 1 keyframes kept
  13. Shot 12, 6:01 to 6:32, 0 of 1 keyframes kept
  14. Shot 13, 6:32 to 7:21, 1 of 1 keyframes kept
  15. Shot 14, 7:21 to 7:58, 1 of 1 keyframes kept
  16. Shot 15, 7:58 to 8:34, 0 of 1 keyframes kept
  17. Shot 16, 8:34 to 9:01, 1 of 1 keyframes kept
  18. Shot 17, 9:01 to 9:28, 0 of 1 keyframes kept
  19. Shot 18, 9:28 to 10:04, 1 of 1 keyframes kept
  20. Shot 19, 10:04 to 10:42, 1 of 1 keyframes kept

20 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
111
whisperx 111
chunks
20
from 111 cues
keyframes
12
kept of 20 captured
frames with text
12
221 lines read
chapters
0
from the source metadata
keyframe bytes
2.0 MB
word timings on 111 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 17:50 1m 42s
stt done 2026-08-10 17:52 10s
chunk done 2026-08-10 17:52 0s
text_embed done 2026-08-10 19:54 0s
keyframe done 2026-08-10 17:52 19s
ocr done 2026-08-10 17:52 7s
frame_embed done 2026-08-10 19:54 2s

Frames, and what the machine read

  • 0:09 #0 done16 line(s)

    shot 0·sharpness 1176.4

    1. AI ENGINEER WORLD'S FAIR0.98
    2. June 30-July 2,2026·San Francisco0.97
    3. SEARCH & RETRIEVAL TRACK1.00
    4. We Cut 94% of Our Al Coding Tokens0.99
    5. With a Local Code Index1.00
    6. Here's the architecture.1.00
    7. 94%1.00
    8. 0.4ms1.00
    9. 0.901.00
    10. Token Reduction0.99
    11. Search Latency1.00
    12. Recall@101.00
    13. Rajkumar Sakthivel1.00
    14. RS1.00
    15. github.com/elara-labs/code-1.00
    16. context-engine1.00
  • 0:36 #1 skipped

    shot 1·duplicate of #0

  • 1:09 #2 done11 line(s)

    shot 2·sharpness 1351.9

    1. THE ASSUMPTION1.00
    2. Every Al coding tool we tried had the same assumption:1.00
    3. send as much context as possible.1.00
    4. WHAT AGENTS SEND1.00
    5. WHAT'S ACTUALLY USEFUL0.99
    6. 45,0001.00
    7. ~5,0001.00
    8. tokens per query1.00
    9. tokens per query0.98
    10. We didn't notice until we saw the cost and latency impact.0.99
    11. elara-labs/code-context-engine1.00
  • 2:10 #3 done12 line(s)

    shot 3·sharpness 980.9

    1. WHAT WE GOT WRONG0.99
    2. We optimized the model. We should have optimized the0.99
    3. context.1.00
    4. Betterprompts1.00
    5. "Be concise." "Only return relevant code." The model still received 45k tokens of input.0.99
    6. Model settings0.98
    7. Temperature, top-p, max_tokens control output shape. The 45k input was already sent and billed.1.00
    8. Outputcompression1.00
    9. "Talk like a caveman." Saves 75% of output (10% of bill). Net impact: ~8%. Wrong 10%.0.99
    10. A retrieval layer between codebase and agent1.00
    11. Search an index, return only relevant chunks. 94% fewer tokens.0.99
    12. elara-labs/code-context-engine1.00
  • 2:30 #4 skipped

    shot 4·duplicate of #3

  • 3:05 #5 done15 line(s)

    shot 5·sharpness 1109.4

    1. WHY INPUT MATTERS0.99
    2. Where your tokens actually go0.98
    3. Output compression1.00
    4. Saves 75% of output tokens1.00
    5. = ~8% off total bill0.99
    6. 90%0.97
    7. Input retrieval1.00
    8. is input0.98
    9. Saves 94% of input tokens1.00
    10. ~61% off total bill0.99
    11. Both help. But if you're only doing one, do the one that tar0.98
    12. Input tokens (file reads, search, context)1.00
    13. spend.1.00
    14. Output tokens (agent replies, code)1.00
    15. elara-labs/code-context-engine1.00
  • 3:32 #6 done24 line(s)

    shot 6·sharpness 1505.3

    1. ARCHITECTURE1.00
    2. A local retrieval layer between codebase and agent1.00
    3. Tree-sitter1.00
    4. Hybrid1.00
    5. Chunk1.00
    6. Code1.00
    7. Confidence1.00
    8. Chunking1.00
    9. Retrieval1.00
    10. Compression1.00
    11. Graph1.00
    12. Scoring1.00
    13. AST-aware splits1.00
    14. Vector+BM25+RRF1.00
    15. Signatures + docs0.97
    16. CALLS· IMPORTS0.96
    17. Threshold gate1.00
    18. 10 langs1.00
    19. 94%1.00
    20. 89%1.00
    21. related1.00
    22. filter0.99
    23. Everything runs locally. No cloud, no API calls. sqlite-vec + FTS5 + graph in three SQLite files.1.00
    24. elara-labs/code-context-engine1.00
  • 4:07 #7 skipped

    shot 7·duplicate of #6

  • 4:18 #8 skipped

    shot 8·duplicate of #6

  • 4:59 #9 done21 line(s)

    shot 9·sharpness 1707.2

    1. DEEP DIVE1.00
    2. Why not just vector search?1.00
    3. 40.99
    4. abc1.00
    5. Vector Search0.98
    6. FTS5 (BM25)1.00
    7. RRF Fusion1.00
    8. bge-small-en-v1.5 (384d). Finds1.00
    9. Exact keyword matching. Catches1.00
    10. Reciprocal Rank Fusion (k=60) merges1.00
    11. conceptually related code even with0.98
    12. function names and identifiers vector1.00
    13. both. Confidence blends similarity,1.00
    14. different naming.1.00
    15. search fuzzes over.0.97
    16. keywords, recency.1.00
    17. Recall:0.780.99
    18. Recall:0.720.99
    19. Recall:0.900.97
    20. Neither retriever is good enough alone. Together they cover each other's blind spots.0.99
    21. elara-labs/code-context-engine1.00
  • 5:27 #10 skipped

    shot 10·duplicate of #9

  • 5:39 #11 done21 line(s)

    shot 11·sharpness 1149.2

    1. THE HARD PART1.00
    2. The hardest problem wasn't retrieval. It was knowing when1.00
    3. retrieval was wrong.1.00
    4. LLM-based scoring1.00
    5. Confidence scoring blend1.00
    6. Asked the model to rate relevance. Accurate but +2-3s latency1.00
    7. and cost per query.1.00
    8. Similarity1.00
    9. 50%1.00
    10. Fixed thresholds1.00
    11. Keywords1.00
    12. 30%1.00
    13. cosine > 0.7 = relevant. Broke on short queries and long queries1.00
    14. alike.1.00
    15. Recency0.93
    16. 20%1.00
    17. Simple heuristic won1.00
    18. 50% similarity + 30% keyword + 20% recency. Adaptive. 0.4ms,0.99
    19. Lesson: don't reach for an LLM when a weighted averag0.98
    20. no API calls.0.99
    21. elara-labs/code-context-engine1.00
  • 6:20 #12 skipped

    shot 12·duplicate of #11

  • 6:43 #13 done17 line(s)

    shot 13·sharpness 1074.5

    1. BENCHMARK1.00
    2. FastAPI FastAPI: 53 files, 20 real questions, reproducible0.97
    3. Full file baseline1.00
    4. 83,681 tok/q1.00
    5. 94%0.99
    6. After retrieval1.00
    7. 4,927 tok/q0.99
    8. retrieval savings0.97
    9. After compression1.00
    10. 523 tok/q1.00
    11. No cherry-picking. No synthetic queries.1.00
    12. 20 questions a developer would actually ask.1.00
    13. Recall@101.00
    14. 0.901.00
    15. $ python benchmarks/run_benchmark.py0.98
    16. --repo fastapi/fastapi --source-dir fa0.99
    17. elara-labs/code-context-engine1.00
  • 7:29 #14 done17 line(s)

    shot 14·sharpness 1711.9

    1. TRADE-OFFS1.00
    2. What we're honest about1.00
    3. 94% is against full-file reads0.99
    4. Monorepos dilute recall1.00
    5. Claude Code already uses grep and partial reads. Real-world1.00
    6. On Go's fiber (396 files), recall dropped to 0.07@10. One-1.00
    7. savings vs normal behavior are lower. Full-file is our reproducible0.99
    8. feature-per-file repos hit R=1.00. Focused files retrieve best.0.99
    9. baseline.1.00
    10. Embedding model matters1.00
    11. What actually worked1.00
    12. bge-small-en-v1.5 (384d) is fast, not SOTA. Bigger models lift1.00
    13. Simple heuristics over ML. SQLite over specialized DB0.99
    14. recall but add latency. We chose speed; <1s re-index at 96%0.99
    15. over pure vector. Local-first. The boring choices com0.99
    16. cache.1.00
    17. elara-labs/code-context-engine1.00
  • 8:16 #15 skipped

    shot 15·duplicate of #14

  • 8:50 #16 done18 line(s)

    shot 16·sharpness 962.1

    1. MULTI-AGENT1.00
    2. One index. Every agent. Shared memory.1.00
    3. Works with every major Al coding tool via MCP protocol0.99
    4. Claude Code1.00
    5. Cursor1.00
    6. VS Code / Copilot1.00
    7. Codex CLl0.95
    8. Gemini CLl0.97
    9. Tabnine1.00
    10. OpenCode1.00
    11. MCP Protocol0.99
    12. Shared Index1.00
    13. Cross-session1.00
    14. 5 tools + session memory0.98
    15. Per-project, not per-agent0.99
    16. Decisions persist across tools1.00
    17. Decisions made in Claude Code surface in Codex. Memory is per-project, not per-agent.1.00
    18. elara-labs/code-context-engine1.00
  • 9:07 #17 skipped

    shot 17·duplicate of #16

  • 9:36 #18 done34 line(s)

    shot 18·sharpness 751.3

    1. KEY TAKEAWAY1.00
    2. MEASURABLE1.00
    3. Every token tracked. Every dollar counted.1.00
    4. my-project · 247 queries · last query 5m ago0.95
    5. 00000000000.84
    6. 88% tokens saved1.00
    7. Input savings0.98
    8. 12.4M1.00
    9. tokens1.00
    10. $186.001.00
    11. Output savings1.00
    12. 48.2k1.00
    13. tokens1.00
    14. $3.621.00
    15. Total saved1.00
    16. 12.4M1.00
    17. tokens1.00
    18. $189.621.00
    19. Breakdown:1.00
    20. retrieval1.00
    21. 84%1.00
    22. 10.4M1.00
    23. $156.001.00
    24. chunk compression0.98
    25. 3%1.00
    26. 421.5k1.00
    27. $6.321.00
    28. output compress*1.00
    29. <1%1.00
    30. 20000000000.93
    31. 48.2k1.00
    32. $3.621.00
    33. Not estimates. Actual tokens served vs full-file baseline, per bucket. Dollar costs from live model pricing.0.98
    34. Thank you . Rajkumar Sakthivel0.96
  • 10:27 #19 done15 line(s)

    shot 19·sharpness 1650.1

    1. KEY TAKEAWAY0.97
    2. The biggest optimization in Al coding0.99
    3. isn't the model. It's the context.1.00
    4. $ uvx --from "code-context-engine[local]" cce init0.99
    5. 94%0.99
    6. 1oca10.95
    7. MIT1.00
    8. fewer input tokens0.99
    9. no data leaves your machine0.98
    10. free, open source1.00
    11. Try it now0.99
    12. Scan to open the repo. Star it, fork it, run the benchmark0.99
    13. yourself.1.00
    14. github.com/elara-labs/code-context-engine1.00
    15. Thank you. Rajkumar Sakthivel0.97

Transcript

111 cues· 1,315 words· 6,958 chars

  1. 0:01 Hey, I'm Raj.
  2. 0:02 I want to tell a story.
  3. 0:03 Me and my friend Voss, we are building project together.
  4. 0:08 We are using AI coding tools every day.
  5. 0:11 Cloud Code, Cursor, Copilot, Codex, normal stuff.
  6. 0:16 One month our AI bill was fine.
  7. 0:18 next month huge we did nothing different same project same tools just more of it we panicked we looked what was happening and we found something surprising most of the money was not the ai thinking most of it was
  8. 0:39 sending too much context files they don't need context is important code that was not relevant sent anyway every time so me and my friend fos we started to building something to fix it in this talk is about what we built and what we learned
  9. 1:05 Every AI coding tools does same thing.
  10. 1:11 It send you a code to the model as a context.
  11. 1:14 And tools thinks more context is better.
  12. 1:18 We measured typical query on our project.
  13. 1:22 It was sending 45,000 tokens of context, but the part of actually mattered is about 5,000 only.
  14. 1:30 Other 40,000 tokens are not useful, but we paid for them.
  15. 1:35 every single curry it's um that's like ordering a pizza and paying for extra nine pizzas you don't eat every time
  16. 1:49 we tried three things before we found what works first we changed our prompt be short only show relevant code sounds good but it does not work the model already got 45 000 tokens before it's read the prompt cost already happened second we change the model setting like a max token temperature same problem this changes the output not the input
  17. 2:18 money is in the input third output compression this one actually works we told the model to write short answers it cut the output 75 percentage but output only about 10 percentage of the cost so 75 percentage of a small number still small number not enough we need to fix the input
  18. 2:46 This is the most important slide.
  19. 2:48 90% of your AIA cost is input.
  20. 2:51 Files, search results, contacts you send in.
  21. 2:55 Only 10% is output.
  22. 2:57 The code the AIA writes back.
  23. 3:00 So if you cut the output by 75%, you can save about 8% total.
  24. 3:07 But if you cut input by 94%, you can save about 61% total.
  25. 3:13 Same math but different result.
  26. 3:15 Fix the input.
  27. 3:17 That's where your money goes.
  28. 3:20 We built a local search layer.
  29. 3:22 It sits between your code base and the AI.
  30. 3:25 Instead of sending whole files, the AI search an index.
  31. 3:29 It gets back only small piece of code actually it needs.
  32. 3:34 Here how it works, five steps.
  33. 3:36 Step one, we read the code and break into small pieces, functions, classes, methods.
  34. 3:44 not a random chunks proper piece that makes sense step 2 we run two searches at the same time one search find the code by meaning one search find the code by exact words then combine the results this is the big saving comes from
  35. 4:03 Step 3, we can shrink their results even more.
  36. 4:08 Keep only the function name and the description.
  37. 4:10 Cut 50 line function down to 5 lines.
  38. 4:14 Step 4, we track the connection with the which function call which.
  39. 4:20 So if you find one piece of code, you can find everything connected to it.
  40. 4:26 Step 5.
  41. 4:28 Every result gets score.
  42. 4:29 If the score is too low, we don't send it.
  43. 4:33 No bad context.
  44. 4:34 Everything runs on your machine.
  45. 4:36 Nothing goes to the cloud.
  46. 4:38 This is the beauty.
  47. 4:40 Why do we run two searches instead of one?
  48. 4:44 Because each one has a weakness.
  49. 4:48 Meaning based search is good at finding related ideas, but it misses exact names.
  50. 4:55 you search for authenticate user function and it might show you different auth function instead because they are similar meaning the word base search is good at exact names but it misses related ideas you search for login flow and it misses everything that says sign in by themselves both searches miss about 1 in 4 results

Open at this second