Videos vh2VGuQ3zhY
The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica
Scene timeline
60 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 226
- whisperx 226
- chunks
- 51
- from 226 cues
- keyframes
- 21
- kept of 60 captured
- frames with text
- 21
- 564 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.5 MB
- word timings on 226 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 08:22 | 1m 20s |
stt |
done | — | 2026-08-11 08:24 | 25s |
chunk |
done | — | 2026-08-11 08:24 | 0s |
text_embed |
done | — | 2026-08-11 08:24 | 0s |
keyframe |
done | — | 2026-08-11 08:24 | 1m 12s |
ocr |
done | — | 2026-08-11 08:25 | 11s |
frame_embed |
done | — | 2026-08-11 08:25 | 3s |
Frames, and what the machine read
-
- Ankush Rastogi1.00
- The 100-Tool Agent1.00
- Is a Trap0.99
- Sohail Shaikh1.00
- Scaling with Semantic Routers and Just-In-Time Context0.99
- Al Engineer World's Fair 20260.99
- For Engineers Building LLM Agents0.99
- Ankush Rastogi1.00
-
- THE PRESENTERS1.00
- Ankush Rastogi1.00
- Ankush Rastogi1.00
- Sohail Shaikh1.00
- Sohail Shaikh1.00
- Senior Data Solutions Engineer · Prosodica LLC0.99
- Data Scientist · Prosodica LLC0.98
- IEEE Senior Member0.99
- Building Real-World Al Systems0.99
- 10+ years across data engineering, Al systems, production1.00
- 9+ years in Al, NLP, conversational intelligence, RAG pipelines,1.00
- analytics, and enterprise LLM implementation.1.00
- semantic search, and production LLM workflows.1.00
- 021.00
- Ankush Rastogi1.00
-
- THE PROBLEM1.00
- The Fat Agent Trap0.99
- Ankush Rastogi1.00
- The Naive Architecture0.99
- EVERY SINGLE REQUEST1.00
- The easiest approach is to dump every tool's schema into the0.99
- prompt on every request. It works perfectly in demos and falls1.00
- User Query0.99
- apart in production.0.99
- Token Bloat 127,000 tokens for 741 tools0.98
- LLM + ALL 100+ Tool Schemas0.98
- Sohail Shaikh1.00
- Accuracy Crash 78% → 13% as the tool pool grows0.99
- T11.00
- T21.00
- T31.00
- T41.00
- T50.98
- T61.00
- T71.00
- T81.00
- T91.00
- T101.00
- Cost Explosion Up to 99× more tokens billed1.00
- T111.00
- T121.00
- T1001.00
- 85 more1.00
- Context Crowding No room left for actual reasoning0.99
- Every token, every request is billed and processed in full.1.00
- 031.00
- Ankush Rastogi1.00
-
- WHY IT FAILS1.00
- Accuracy Collapses With Scale1.00
- Ankush Rastogi1.00
- 78%1.00
- 40%1.00
- 13%1.00
- Accuracy at 10 tools1.00
- Accuracy at 100 tools1.00
- Accuracy at 741 tools0.98
- Tool Selection Accuracy (%) vs. Tool-Pool Size1.00
- Sohail Shaikh1.00
- 1001.00
- 800.99
- 601.00
- 401.00
- 201.00
- 101.00
- 300.75
- 1001.00
- 2001.00
- 7411.00
- Fat Agent (all tools)0.97
- With Semantic Router0.99
- 041.00
- Ankush Rastogi1.00
-
- THE HIDDEN TAX1.00
- Latency & Cost Scale Against You1.00
- Ankush Rastogi1.00
- 127K1.00
- ~1K0.92
- 99%1.00
- Tokens: 741 tools loaded0.99
- Tokens: with JIT routing1.00
- token reduction achieved1.00
- Time-to-First-Token (TTFT, ms) vs. Tool Count @ GPT-4o0.97
- Sohail Shaikh1.00
- 60001.00
- 50001.00
- 40001.00
- 30001.00
- 20001.00
- 10001.00
- 101.00
- 500.87
- 1001.00
- 2001.00
- 5001.00
- Fat Agent (all tools)0.99
- With Semantic Router1.00
- 051.00
- Ankush Rastogi1.00
-
- HEAD TO HEAD1.00
- Architecture Comparison1.00
- Ankush Rastogi1.00
- Aspect1.00
- Fat Agent (All Tools)1.00
- Semantic Router + JIT1.00
- Context Tokens1.00
- Very high: all schemas, every call1.00
- Very low: only 3-5 relevant tools0.98
- Latency (TTFT)1.00
- Grows linearly with tool count1.00
- Near-flat: embedding search is ms-fast1.00
- Tool Accuracy1.00
- Crashes from 78% → 13% at scale1.00
- Stays above 83% even at 700+ tools1.00
- Sohail Shaikh1.00
- Token Cost1.00
- Linear: 127K tokens at 741 tools0.99
- ~99% savings: ~1K tokens / request1.00
- Scalability0.99
- Breaks around 100+ tools1.00
- Proven stable at 740+ tools0.98
- Modularity1.00
- Monolithic, hard to debug1.00
- Decoupled: test each component0.99
- When to use1.00
- < 20 tools, small demos0.98
- > 50 tools, production systems0.99
- 061.00
- Ankush Rastogi1.00
-
- THE MECHANISM0.99
- How Semantic Routing Works1.00
- Ankush Rastogi1.00
- User1.00
- Embed1.00
- Vector1.00
- Top-K1.00
- LLM1.00
- Query1.00
- Query1.00
- Search1.00
- Tools1.00
- Call1.00
- Response1.00
- Tool Vector Database0.99
- Pre-indexed offline, one-time setup0.99
- Sohail Shaikh1.00
- get_weather0.98
- search_flights1.00
- book_hotel1.00
- send_email1.00
- calendar_event1.00
- stock_price1.00
- run_sql_query0.99
- translate_text1.00
- pdf_extract1.00
- + 700 more0.95
- Think of this as RAG but for tools instead of documents. Same retrieval logic, different artifact type.1.00
- 071.00
- Ankush Rastogi1.00
Transcript
226 cues· 3,434 words· 18,912 chars
- 0:01 Hi everyone, thanks for being here.
- 0:04 I'm Sohail here and along with me is Ankush.
- 0:07 So today we'll be talking about a mistake that looks harmless at first, which is basically giving an AI agent every tool access it might ever need all at once.
- 0:20 So basically that approach works well in a demo.
- 0:24 It might even work with a small number of tools, like say, for example, 10 tools.
- 0:30 But once the catalog grows, the agent gets slower.
- 0:36 It might become more expensive and less accurate as well.
- 0:40 That is why we are calling it the 100-tool agent wrap.
- 0:43 In the next half an hour or so, we'll show why it breaks, what the numbers look like, and how semantic routing with just-in-time context can help us fix this problem.
- 0:57 So a quick introduction about myself.
- 1:00 I am Suhail Shaikh.
- 1:02 I'm currently working as a data scientist with Prosodica.
- 1:06 My background spans across AI, NLP, marketing, analytics, and even engineering.
- 1:13 My current focus is on applied AI, NLP, and conversational intelligence, along with RAC systems.
- 1:20 I'm especially interested in making AI systems more reliable, measurable, and even scalable beyond the demos in production.
- 1:31 And I'm Ankush Astogi.
- 1:33 I work as a senior data solutions engineer at Prasodica.
- 1:38 I have spent more than a decade in AI, data engineering, and production systems.
- 1:43 my focus is the engine side so it's not about what's going to work in notebook but whether it's going to survive in with real load real user and real failures so that is the angle we are taking today
- 2:04 Sohail will focus more on the model and routing behavior, and I will focus more on system design, implementation, and production trade-offs.
- 2:15 Awesome.
- 2:15 Thank you, Ankush.
- 2:17 So let's get into it.
- 2:20 So let's imagine a common design.
- 2:23 You tend to build a system, and it can do many things.
- 2:28 Say, for example, querying a database, sending an email,
- 2:33 even checking a calendar or looking up an order, calling an API, and so on and so forth.
- 2:38 The simplest approach over here would be to give a model every tool definition on every request.
- 2:46 Every function name, every description, and even every JSON schema will go into the prompt whether the user might need it or not.
- 2:55 So that we are calling that as a fat agent at small scale.
- 2:59 It feels fine with 10 tools as well.
- 3:02 The model might usually pick the right one.
- 3:05 The demo looks good.
- 3:07 Then the product grows 10 tools might become 30 or it will keep on increasing and eventually the model starts calling the wrong function.
- 3:20 Can starts confusing similar tools may invent
- 3:24 tool names and even take longer to respond the important point is basically the design does not fail because one tool is badly written it fails because every request is forced to carry the entire catalog so let's look here there are say for example uh 741 tools in your in your entire schema but
- 3:51 and it will basically take up to 127,000 tokens just to have all those tool descriptions in it.
- 4:00 And this is even before the user's actual question is even considered.
- 4:06 So basically this will lead to context overload and we need to manage that properly.
- 4:15 So on this slide, we see why it is,
- 4:18 Failing and by the accuracy collapses beyond a point.
- 4:22 So when you look at the accuracy curve with the 10 tools, fat agent will get the tools right.
- 4:29 Almost 78% of the times that is not perfect, but it's usable at almost a hundred tools.
- 4:37 The accuracy drops to around 40%, less than half of the tools that are called are the correct tools.
- 4:45 And if it grows beyond that, like say for example, in over here at 741 tools, the accuracy will be a mere 13.6%.
- 4:55 So in short, it's roughly one correct tool out of eight tools.
- 5:00 So when we compare it with the semantic router, semantic router behaves very differently.
- 5:06 It stays about 83% across the same catalog sizes.
- 5:11 That is because the model is not choosing from hundreds of tools.
- 5:15 It's choosing from a small and relevant set.
loading