read-only demo

Videos vh2VGuQ3zhY

The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:28:27· en· indexed 2026-08-11 08:25

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:28, 1 of 1 keyframes kept
  2. Shot 1, 0:28 to 0:56, 0 of 1 keyframes kept
  3. Shot 2, 0:56 to 1:22, 1 of 1 keyframes kept
  4. Shot 3, 1:22 to 1:47, 0 of 1 keyframes kept
  5. Shot 4, 1:47 to 2:13, 0 of 1 keyframes kept
  6. Shot 5, 2:13 to 2:43, 1 of 1 keyframes kept
  7. Shot 6, 2:43 to 3:13, 0 of 1 keyframes kept
  8. Shot 7, 3:13 to 3:43, 0 of 1 keyframes kept
  9. Shot 8, 3:43 to 4:13, 0 of 1 keyframes kept
  10. Shot 9, 4:13 to 4:43, 1 of 1 keyframes kept
  11. Shot 10, 4:43 to 5:13, 0 of 1 keyframes kept
  12. Shot 11, 5:13 to 5:43, 0 of 1 keyframes kept
  13. Shot 12, 5:43 to 6:10, 1 of 1 keyframes kept
  14. Shot 13, 6:10 to 6:37, 0 of 1 keyframes kept
  15. Shot 14, 6:37 to 7:05, 0 of 1 keyframes kept
  16. Shot 15, 7:05 to 7:32, 0 of 1 keyframes kept
  17. Shot 16, 7:32 to 7:58, 1 of 1 keyframes kept
  18. Shot 17, 7:58 to 8:25, 0 of 1 keyframes kept
  19. Shot 18, 8:25 to 8:51, 0 of 1 keyframes kept
  20. Shot 19, 8:51 to 9:18, 0 of 1 keyframes kept
  21. Shot 20, 9:18 to 9:46, 1 of 1 keyframes kept
  22. Shot 21, 9:46 to 10:13, 0 of 1 keyframes kept
  23. Shot 22, 10:13 to 10:41, 0 of 1 keyframes kept
  24. Shot 23, 10:41 to 11:09, 0 of 1 keyframes kept
  25. Shot 24, 11:09 to 11:34, 1 of 1 keyframes kept
  26. Shot 25, 11:34 to 11:59, 0 of 1 keyframes kept
  27. Shot 26, 11:59 to 12:24, 0 of 1 keyframes kept
  28. Shot 27, 12:24 to 12:49, 0 of 1 keyframes kept
  29. Shot 28, 12:49 to 13:21, 1 of 1 keyframes kept
  30. Shot 29, 13:21 to 13:54, 0 of 1 keyframes kept
  31. Shot 30, 13:54 to 14:26, 0 of 1 keyframes kept
  32. Shot 31, 14:26 to 14:55, 1 of 1 keyframes kept
  33. Shot 32, 14:55 to 15:25, 0 of 1 keyframes kept
  34. Shot 33, 15:25 to 15:55, 0 of 1 keyframes kept
  35. Shot 34, 15:55 to 16:27, 1 of 1 keyframes kept
  36. Shot 35, 16:27 to 16:59, 1 of 1 keyframes kept
  37. Shot 36, 16:59 to 17:31, 0 of 1 keyframes kept
  38. Shot 37, 17:31 to 17:59, 1 of 1 keyframes kept
  39. Shot 38, 17:59 to 18:28, 0 of 1 keyframes kept
  40. Shot 39, 18:28 to 18:56, 0 of 1 keyframes kept
  41. Shot 40, 18:56 to 19:25, 1 of 1 keyframes kept
  42. Shot 41, 19:25 to 19:55, 0 of 1 keyframes kept
  43. Shot 42, 19:55 to 20:24, 0 of 1 keyframes kept
  44. Shot 43, 20:24 to 20:57, 1 of 1 keyframes kept
  45. Shot 44, 20:57 to 21:30, 0 of 1 keyframes kept
  46. Shot 45, 21:30 to 22:03, 0 of 1 keyframes kept
  47. Shot 46, 22:03 to 22:31, 1 of 1 keyframes kept
  48. Shot 47, 22:31 to 22:59, 0 of 1 keyframes kept
  49. Shot 48, 22:59 to 23:27, 0 of 1 keyframes kept
  50. Shot 49, 23:27 to 23:53, 1 of 1 keyframes kept
  51. Shot 50, 23:53 to 24:20, 0 of 1 keyframes kept
  52. Shot 51, 24:20 to 24:24, 0 of 1 keyframes kept
  53. Shot 52, 24:24 to 24:52, 1 of 1 keyframes kept
  54. Shot 53, 24:52 to 25:21, 0 of 1 keyframes kept
  55. Shot 54, 25:21 to 25:50, 0 of 1 keyframes kept
  56. Shot 55, 25:50 to 26:19, 1 of 1 keyframes kept
  57. Shot 56, 26:19 to 26:49, 0 of 1 keyframes kept
  58. Shot 57, 26:49 to 27:18, 0 of 1 keyframes kept
  59. Shot 58, 27:18 to 27:58, 1 of 1 keyframes kept
  60. Shot 59, 27:58 to 28:26, 1 of 1 keyframes kept

60 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
226
whisperx 226
chunks
51
from 226 cues
keyframes
21
kept of 60 captured
frames with text
21
564 lines read
chapters
0
from the source metadata
keyframe bytes
6.5 MB
word timings on 226 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 08:22 1m 20s
stt done 2026-08-11 08:24 25s
chunk done 2026-08-11 08:24 0s
text_embed done 2026-08-11 08:24 0s
keyframe done 2026-08-11 08:24 1m 12s
ocr done 2026-08-11 08:25 11s
frame_embed done 2026-08-11 08:25 3s

Frames, and what the machine read

  • 0:08 #0 done8 line(s)

    shot 0·sharpness 676.7

    1. Ankush Rastogi1.00
    2. The 100-Tool Agent1.00
    3. Is a Trap0.99
    4. Sohail Shaikh1.00
    5. Scaling with Semantic Routers and Just-In-Time Context0.99
    6. Al Engineer World's Fair 20260.99
    7. For Engineers Building LLM Agents0.99
    8. Ankush Rastogi1.00
  • 0:31 #1 skipped

    shot 1·duplicate of #0

  • 1:01 #2 done15 line(s)

    shot 2·sharpness 1069.5

    1. THE PRESENTERS1.00
    2. Ankush Rastogi1.00
    3. Ankush Rastogi1.00
    4. Sohail Shaikh1.00
    5. Sohail Shaikh1.00
    6. Senior Data Solutions Engineer · Prosodica LLC0.99
    7. Data Scientist · Prosodica LLC0.98
    8. IEEE Senior Member0.99
    9. Building Real-World Al Systems0.99
    10. 10+ years across data engineering, Al systems, production1.00
    11. 9+ years in Al, NLP, conversational intelligence, RAG pipelines,1.00
    12. analytics, and enterprise LLM implementation.1.00
    13. semantic search, and production LLM workflows.1.00
    14. 021.00
    15. Ankush Rastogi1.00
  • 1:30 #3 skipped

    shot 3·duplicate of #2

  • 2:00 #4 skipped

    shot 4·duplicate of #2

  • 2:31 #5 done32 line(s)

    shot 5·sharpness 1406.9

    1. THE PROBLEM1.00
    2. The Fat Agent Trap0.99
    3. Ankush Rastogi1.00
    4. The Naive Architecture0.99
    5. EVERY SINGLE REQUEST1.00
    6. The easiest approach is to dump every tool's schema into the0.99
    7. prompt on every request. It works perfectly in demos and falls1.00
    8. User Query0.99
    9. apart in production.0.99
    10. Token Bloat 127,000 tokens for 741 tools0.98
    11. LLM + ALL 100+ Tool Schemas0.98
    12. Sohail Shaikh1.00
    13. Accuracy Crash 78% → 13% as the tool pool grows0.99
    14. T11.00
    15. T21.00
    16. T31.00
    17. T41.00
    18. T50.98
    19. T61.00
    20. T71.00
    21. T81.00
    22. T91.00
    23. T101.00
    24. Cost Explosion Up to 99× more tokens billed1.00
    25. T111.00
    26. T121.00
    27. T1001.00
    28. 85 more1.00
    29. Context Crowding No room left for actual reasoning0.99
    30. Every token, every request is billed and processed in full.1.00
    31. 031.00
    32. Ankush Rastogi1.00
  • 3:04 #6 skipped

    shot 6·duplicate of #5

  • 3:31 #7 skipped

    shot 7·duplicate of #5

  • 4:01 #8 skipped

    shot 8·duplicate of #5

  • 4:34 #9 done25 line(s)

    shot 9·sharpness 814.8

    1. WHY IT FAILS1.00
    2. Accuracy Collapses With Scale1.00
    3. Ankush Rastogi1.00
    4. 78%1.00
    5. 40%1.00
    6. 13%1.00
    7. Accuracy at 10 tools1.00
    8. Accuracy at 100 tools1.00
    9. Accuracy at 741 tools0.98
    10. Tool Selection Accuracy (%) vs. Tool-Pool Size1.00
    11. Sohail Shaikh1.00
    12. 1001.00
    13. 800.99
    14. 601.00
    15. 401.00
    16. 201.00
    17. 101.00
    18. 300.75
    19. 1001.00
    20. 2001.00
    21. 7411.00
    22. Fat Agent (all tools)0.97
    23. With Semantic Router0.99
    24. 041.00
    25. Ankush Rastogi1.00
  • 5:04 #10 skipped

    shot 10·duplicate of #9

  • 5:39 #11 skipped

    shot 11·duplicate of #9

  • 6:04 #12 done26 line(s)

    shot 12·sharpness 853.0

    1. THE HIDDEN TAX1.00
    2. Latency & Cost Scale Against You1.00
    3. Ankush Rastogi1.00
    4. 127K1.00
    5. ~1K0.92
    6. 99%1.00
    7. Tokens: 741 tools loaded0.99
    8. Tokens: with JIT routing1.00
    9. token reduction achieved1.00
    10. Time-to-First-Token (TTFT, ms) vs. Tool Count @ GPT-4o0.97
    11. Sohail Shaikh1.00
    12. 60001.00
    13. 50001.00
    14. 40001.00
    15. 30001.00
    16. 20001.00
    17. 10001.00
    18. 101.00
    19. 500.87
    20. 1001.00
    21. 2001.00
    22. 5001.00
    23. Fat Agent (all tools)0.99
    24. With Semantic Router1.00
    25. 051.00
    26. Ankush Rastogi1.00
  • 6:31 #13 skipped

    shot 13·duplicate of #12

  • 7:01 #14 skipped

    shot 14·duplicate of #12

  • 7:16 #15 skipped

    shot 15·duplicate of #12

  • 7:48 #16 done30 line(s)

    shot 16·sharpness 1336.5

    1. HEAD TO HEAD1.00
    2. Architecture Comparison1.00
    3. Ankush Rastogi1.00
    4. Aspect1.00
    5. Fat Agent (All Tools)1.00
    6. Semantic Router + JIT1.00
    7. Context Tokens1.00
    8. Very high: all schemas, every call1.00
    9. Very low: only 3-5 relevant tools0.98
    10. Latency (TTFT)1.00
    11. Grows linearly with tool count1.00
    12. Near-flat: embedding search is ms-fast1.00
    13. Tool Accuracy1.00
    14. Crashes from 78% → 13% at scale1.00
    15. Stays above 83% even at 700+ tools1.00
    16. Sohail Shaikh1.00
    17. Token Cost1.00
    18. Linear: 127K tokens at 741 tools0.99
    19. ~99% savings: ~1K tokens / request1.00
    20. Scalability0.99
    21. Breaks around 100+ tools1.00
    22. Proven stable at 740+ tools0.98
    23. Modularity1.00
    24. Monolithic, hard to debug1.00
    25. Decoupled: test each component0.99
    26. When to use1.00
    27. < 20 tools, small demos0.98
    28. > 50 tools, production systems0.99
    29. 061.00
    30. Ankush Rastogi1.00
  • 8:04 #17 skipped

    shot 17·duplicate of #16

  • 8:35 #18 skipped

    shot 18·duplicate of #16

  • 9:04 #19 skipped

    shot 19·duplicate of #16

  • 9:32 #20 done30 line(s)

    shot 20·sharpness 1225.1

    1. THE MECHANISM0.99
    2. How Semantic Routing Works1.00
    3. Ankush Rastogi1.00
    4. User1.00
    5. Embed1.00
    6. Vector1.00
    7. Top-K1.00
    8. LLM1.00
    9. Query1.00
    10. Query1.00
    11. Search1.00
    12. Tools1.00
    13. Call1.00
    14. Response1.00
    15. Tool Vector Database0.99
    16. Pre-indexed offline, one-time setup0.99
    17. Sohail Shaikh1.00
    18. get_weather0.98
    19. search_flights1.00
    20. book_hotel1.00
    21. send_email1.00
    22. calendar_event1.00
    23. stock_price1.00
    24. run_sql_query0.99
    25. translate_text1.00
    26. pdf_extract1.00
    27. + 700 more0.95
    28. Think of this as RAG but for tools instead of documents. Same retrieval logic, different artifact type.1.00
    29. 071.00
    30. Ankush Rastogi1.00
  • 10:05 #21 skipped

    shot 21·duplicate of #20

  • 10:38 #22 skipped

    shot 22·duplicate of #20

  • 11:03 #23 skipped

    shot 23·duplicate of #20

Transcript

226 cues· 3,434 words· 18,912 chars

  1. 0:01 Hi everyone, thanks for being here.
  2. 0:04 I'm Sohail here and along with me is Ankush.
  3. 0:07 So today we'll be talking about a mistake that looks harmless at first, which is basically giving an AI agent every tool access it might ever need all at once.
  4. 0:20 So basically that approach works well in a demo.
  5. 0:24 It might even work with a small number of tools, like say, for example, 10 tools.
  6. 0:30 But once the catalog grows, the agent gets slower.
  7. 0:36 It might become more expensive and less accurate as well.
  8. 0:40 That is why we are calling it the 100-tool agent wrap.
  9. 0:43 In the next half an hour or so, we'll show why it breaks, what the numbers look like, and how semantic routing with just-in-time context can help us fix this problem.
  10. 0:57 So a quick introduction about myself.
  11. 1:00 I am Suhail Shaikh.
  12. 1:02 I'm currently working as a data scientist with Prosodica.
  13. 1:06 My background spans across AI, NLP, marketing, analytics, and even engineering.
  14. 1:13 My current focus is on applied AI, NLP, and conversational intelligence, along with RAC systems.
  15. 1:20 I'm especially interested in making AI systems more reliable, measurable, and even scalable beyond the demos in production.
  16. 1:31 And I'm Ankush Astogi.
  17. 1:33 I work as a senior data solutions engineer at Prasodica.
  18. 1:38 I have spent more than a decade in AI, data engineering, and production systems.
  19. 1:43 my focus is the engine side so it's not about what's going to work in notebook but whether it's going to survive in with real load real user and real failures so that is the angle we are taking today
  20. 2:04 Sohail will focus more on the model and routing behavior, and I will focus more on system design, implementation, and production trade-offs.
  21. 2:15 Awesome.
  22. 2:15 Thank you, Ankush.
  23. 2:17 So let's get into it.
  24. 2:20 So let's imagine a common design.
  25. 2:23 You tend to build a system, and it can do many things.
  26. 2:28 Say, for example, querying a database, sending an email,
  27. 2:33 even checking a calendar or looking up an order, calling an API, and so on and so forth.
  28. 2:38 The simplest approach over here would be to give a model every tool definition on every request.
  29. 2:46 Every function name, every description, and even every JSON schema will go into the prompt whether the user might need it or not.
  30. 2:55 So that we are calling that as a fat agent at small scale.
  31. 2:59 It feels fine with 10 tools as well.
  32. 3:02 The model might usually pick the right one.
  33. 3:05 The demo looks good.
  34. 3:07 Then the product grows 10 tools might become 30 or it will keep on increasing and eventually the model starts calling the wrong function.
  35. 3:20 Can starts confusing similar tools may invent
  36. 3:24 tool names and even take longer to respond the important point is basically the design does not fail because one tool is badly written it fails because every request is forced to carry the entire catalog so let's look here there are say for example uh 741 tools in your in your entire schema but
  37. 3:51 and it will basically take up to 127,000 tokens just to have all those tool descriptions in it.
  38. 4:00 And this is even before the user's actual question is even considered.
  39. 4:06 So basically this will lead to context overload and we need to manage that properly.
  40. 4:15 So on this slide, we see why it is,
  41. 4:18 Failing and by the accuracy collapses beyond a point.
  42. 4:22 So when you look at the accuracy curve with the 10 tools, fat agent will get the tools right.
  43. 4:29 Almost 78% of the times that is not perfect, but it's usable at almost a hundred tools.
  44. 4:37 The accuracy drops to around 40%, less than half of the tools that are called are the correct tools.
  45. 4:45 And if it grows beyond that, like say for example, in over here at 741 tools, the accuracy will be a mere 13.6%.
  46. 4:55 So in short, it's roughly one correct tool out of eight tools.
  47. 5:00 So when we compare it with the semantic router, semantic router behaves very differently.
  48. 5:06 It stays about 83% across the same catalog sizes.
  49. 5:11 That is because the model is not choosing from hundreds of tools.
  50. 5:15 It's choosing from a small and relevant set.

Open at this second