Videos EcqMYoIV57A
Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo
Scene timeline
63 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 244
- whisperx 244
- chunks
- 46
- from 244 cues
- keyframes
- 29
- kept of 63 captured
- frames with text
- 29
- 648 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.8 MB
- word timings on 244 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 11:04 | 1m 24s |
stt |
done | — | 2026-08-11 11:06 | 27s |
chunk |
done | — | 2026-08-11 11:06 | 0s |
text_embed |
done | — | 2026-08-11 11:06 | 0s |
keyframe |
done | — | 2026-08-11 11:06 | 2m 11s |
ocr |
done | — | 2026-08-11 11:08 | 19s |
frame_embed |
done | — | 2026-08-11 11:09 | 5s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.94
-
- qodo1.00
- Hidden Failure Modes1.00
- for Al Agents0.99
- Lessons from Real-World AI Engineering1.00
- qodo0.95
- AlEngin0.95
- EUROPZ0.94
-
- Help0.91
- Hidden Falure Modes for Al0.97
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpKth7ZpPGAE/edit?slide=id.g3d342bdbadb_0_260#slide=id.g3d342bdbadb_0_2600.99
- ☆0.97
- New Chrome avallble0.95
- qodo1.00
- 小70.60
- Hidden Failure Modes1.00
- AIE1.00
- ★1.00
- ★0.99
- for Al Agents0.98
- Lessons from Real-World AI Engineering0.98
- Engineering the future of Al0.97
- AIEr0.92
-
- Tab1.00
- Help1.00
- Hidden Falure Modes for Al0.95
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_0_778#slide=id.g3d342bdbadb_0_7780.99
- ☆0.99
- New Chrome avallable0.99
- AIE1.00
- Nupur Sharma1.00
- ★1.00
- ★1.00
- ★1.00
- Solutions Architect,Qodo0.98
- Scan to1.00
- Connect1.00
- qodol0.86
- Braintrust1.00
- WorkOS OpenAI0.97
-
- Window1.00
- Help1.00
- Hidden Falure Modes for Al0.99
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_0_1308#slide=id.g3d342bdbadb_0_13080.99
- ☆0.98
- New Chrome avallble0.95
- The Evolution of the Agentic Stack1.00
- ***0.55
- PHASE 030.94
- AIE1.00
- PHASE 021.00
- Multi-Agent Systems1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- PHASE 010.94
- Agentic Workflows1.00
- Static Prompts0.98
- pruning of data chunks.1.00
- Short context & manual0.99
- Dynamic tool use and1.00
- iterative reasoning loops.1.00
- Specialized orchestration and1.00
- cooperative problem solving.0.99
- qodo1.00
- Braintrust1.00
- WorkOS OpenAI0.93
- AIE1.00
-
- Tab0.97
- Window1.00
- Help1.00
- Hidden Falure Modes for Al0.99
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_0_1308#slide=id.g3d342bdbadb_0_13080.99
- ☆0.98
- New Chrome avallable0.99
- The Evolution of the Agentic Stack1.00
- ★★*0.50
- PHASE 030.99
- AIE1.00
- PHASE 021.00
- Multi-Agent Systems1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- PHASE 010.94
- Agentic Workflows1.00
- Static Prompts1.00
- pruning of data chunks.1.00
- Short context & manual1.00
- Dynamic tool use and1.00
- iterative reasoning loops.1.00
- Specialized orchestration and1.00
- cooperative problem solving.0.99
- qodo1.00
- AlEngineer0.91
- AlEn0.90
- EUROPE1.00
-
- Tab1.00
- Window1.00
- Help1.00
- Hidden Falure Modes for Al0.98
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_0_1308#slide=id.g3d342bdbadb_0_13080.99
- ☆0.96
- New Chrome avallable0.97
- The Evolution of the Agentic Stack0.99
- ***0.66
- PHASE 030.94
- AIE1.00
- PHASE 021.00
- Multi-Agent Systems1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- PHASE 010.94
- Agentic Workflows1.00
- Static Prompts0.99
- pruning of data chunks.0.98
- Short context & manual0.99
- Dynamic tool use and1.00
- iterative reasoning loops.0.99
- Specialized orchestration and1.00
- cooperative problem solving.0.99
- qodo1.00
- Engineering the future of Al0.99
- AIE0.93
-
- Tab0.99
- Window1.00
- Help1.00
- Hidden Falure Modes for Al0.98
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_22_254#slide=id.g3d342bdbadb_22_2540.98
- ☆0.99
- 00.56
- New Chrome avallable0.96
- Failure Mode 1: The Context Trap0.98
- 10 Total Retrieved Documents1.00
- *★★0.54
- AIE1.00
- ★1.00
- ★1.00
- PROBLEM ANALYSIS1.00
- 751.00
- 751.00
- ★1.00
- The "Lost in the Middle" Phenomenon0.99
- 701.00
- 701.00
- ★1.00
- ★1.00
- ★1.00
- LLMs exhibit a U-shaped performance curve when1.00
- ACaay0.60
- 651.00
- ACccy0.66
- 651.00
- processing long context windows:1.00
- 601.00
- 601.00
- High recall for information at the beginning of0.99
- a prompt.0.97
- 551.00
- 551.00
- High recall for information at the end of a0.99
- 501.00
- 501.00
- prompt.1.00
- 1st1.00
- 5th1.00
- 10th1.00
- Failure: Critical data or instructions buried in1.00
- Position of Document with the Answer1.00
- the middle are frequently ignored or "lost".1.00
- claude-1.31.00
- claude-1.3-100k1.00
- gpt-3.50.98
- qodo1.00
- Engineering the future of Al1.00
- AlEn0.94
-
- Tab1.00
- Window Help0.96
- 16:131.00
- Hidden Fallure Modes for Al0.98
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_22_810#slide=id.g3d342bdbadb_22_8100.99
- ☆0.99
- New Chrome avallable0.98
- Strategic Solutions: Context Optimization0.99
- Solution1.00
- Best For...1.00
- Developer Effort1.00
- Cost Impact1.00
- ***0.60
- ★1.00
- AIE1.00
- ★1.00
- Context Engine1.00
- Large, messy codebases.1.00
- High: Needs Search/Ranking0.99
- logic.1.00
- Moderate (Indexing1.00
- overhead).1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Hierarchical1.00
- Understanding project1.00
- Medium: Needs background1.00
- High upfront1.00
- Summarization1.00
- structure.1.00
- LLM processing.0.99
- (Pre-summarization).1.00
- Knowledge Graph1.00
- dependencies.1.00
- Complex logic and deep1.00
- code-parsing engine.1.00
- Very High: Needs1.00
- Variable (Graph DB hosting).1.00
- Iterative Retrieval1.00
- Proactive agents using tools.0.99
- Low: Needs a 'Read' tool.0.98
- High (Recursive API calls).1.00
- Self-Correction1.00
- accuracy).1.00
- High-stakes tasks (100%1.00
- Medium:Adds "Critic"node.0.98
- Lower (Adds latency/token0.99
- cost).1.00
- Engineering the future of Al0.99
-
- Bookmarks0.99
- Profiles1.00
- Tab1.00
- Window Help0.99
- Thu 9 Apr 16:130.98
- Hidden Falure Modes for Al/0.91
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_22_810#slide=id.g3d342bdbadb_22_8100.98
- ☆0.99
- New Chrome avallable0.95
- Strategic Solutions: Context Optimization0.99
- Solution1.00
- Best For...1.00
- Developer Effort1.00
- Cost Impact1.00
- ***0.62
- ★1.00
- AIE1.00
- ★1.00
- ★0.99
- Context Engine1.00
- Large, messy codebases.1.00
- High: Needs Search/Ranking1.00
- logic.1.00
- Moderate (Indexing1.00
- overhead).1.00
- ★1.00
- Hierarchical1.00
- Understanding project1.00
- Medium: Needs background0.98
- High upfront1.00
- Summarization1.00
- structure.1.00
- LLM processing.0.99
- (Pre-summarization).1.00
- Knowledge Graph1.00
- dependencies.1.00
- Complex logic and deep1.00
- code-parsing engine.1.00
- Very High: Needs1.00
- Variable (Graph DB hosting).1.00
- Iterative Retrieval1.00
- Proactive agents using tools.1.00
- Low: Needs a 'Read' tool.0.98
- High (Recursive API calls).1.00
- Self-Correction1.00
- accuracy).1.00
- High-stakes tasks (100%0.97
- Medium:Adds "Critic"node.0.98
- Lower (Adds latency/token0.99
- cost).1.00
- AlEngineer0.96
- AIE0.96
- EUROPE1.00
-
- Tab0.99
- Window Help0.99
- Hidden Fallure Modes for Al0.99
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_22_810#slide=id.g3d342bdbadb_22_8100.98
- ☆0.97
- New Chrome avallable0.98
- Strategic Solutions: Context Optimization0.99
- Solution1.00
- Best For...1.00
- Developer Effort1.00
- Cost Impact1.00
- ***0.58
- ★1.00
- AIE1.00
- ★1.00
- Context Engine1.00
- Large, messy codebases.1.00
- High: Needs Search/Ranking1.00
- logic.1.00
- Moderate (Indexing1.00
- overhead).1.00
- ★1.00
- ★1.00
- Hierarchical1.00
- Understanding project1.00
- Medium: Needs background1.00
- High upfront1.00
- Summarization1.00
- structure.1.00
- LLM processing.0.99
- (Pre-summarization).1.00
- Knowledge Graph1.00
- dependencies.1.00
- Complex logic and deep1.00
- code-parsing engine.1.00
- Very High: Needs1.00
- Variable (Graph DB hosting).1.00
- Iterative Retrieval1.00
- Proactive agents using tools.0.99
- Low: Needs a 'Read" tool.0.98
- High (Recursive API calls).1.00
- Self-Correction1.00
- High-stakes tasks (100%1.00
- accuracy).1.00
- Medium: Adds "Critic" node.0.99
- Lower (Adds latency/token0.98
- cost).1.00
- Engineering the future of Al1.00
- AIE1.00
-
- Tab0.98
- Window1.00
- Help1.00
- 品回0.66
- Hidden Falure Modes for Al0.97
- docs.google.com/presentation/d/11zJUkX3Upxyszm8AZYgSKpK1h7ZpPGAE/edit?slide=id.g3d342bdbadb_22_2413#slide=id.g3d342bdbadb_22_24130.99
- ☆0.99
- New Chrome avalable0.96
- Failure Mode 2: The Orchestration Paradox0.98
- *★*0.57
- REASONING DRIFT: DYNAMIC LOOP DEGRADATION1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- Infinite Thinking Loops0.99
- B0.97
- Purely dynamic "ReAct" loops can trap agents in0.98
- recursive reasoning, repeating the same logic1.00
- without tool execution.0.98
- Resource Exhaustion1.00
- Agents consume excessive API credits "thinking"0.99
- about tool selection rather than solving the task.0.99
- qodo1.00
- AlEngineer0.97
- AIE1.00
- EUROPE1.00
Transcript
244 cues· 4,076 words· 21,531 chars
- 0:14 I'm Nupur.
- 0:15 I work with Kodo.
- 0:17 At Kodo, we do agentic reviews.
- 0:20 I have a background in DevSecOps, so I'm coming from an industry where everything was deterministic.
- 0:27 The pipelines, they run, they crash.
- 0:29 If they crash, we fix them, to a place where we are doing agents where nothing is deterministic.
- 0:38 So in my last few years, I have learned where and how agents fail, what are the learnings, and today I will be sharing some of my learnings with you.
- 0:52 So if you see the evolution of agents, it started with static prompts where it was a 4K context window and we tried to put whatever was important or whatever we deemed important and the AI models will process it and provide you with the results, right?
- 1:12 When we started with that, that means that it was on us to tell LLMs what they should look into.
- 1:19 That means if we provide wrong inputs, we might not get proper results.
- 1:24 And then we thought maybe if the context window grows, if the context size grows, we can do better.
- 1:29 We can have more inputs.
- 1:32 and we started with agentic workflows so we created an agent we get that get them tools like search tool to go into search into documents and do something uh as a command then again look into the search and do something which again created kind of a loop where the tool does not know where to stop it thinks like i need more inputs again going back and back it's a loop
- 1:57 To improvise on that, nowadays multi-agents is becoming more popular.
- 2:03 Create multi-agents, do a lot of stuff together.
- 2:07 When we see it like that, we have a lot of agents working for you.
- 2:11 So a security agent trying to figure security concerns.
- 2:15 A review agent trying to review the tool, a coding agent trying to fix things.
- 2:20 Now again the more the tools the more issues you have.
- 2:24 Not every agent understands and they have clash in their understandings where you don't get into the results.
- 2:34 So what do we learn from here?
- 2:37 What we see is context is not a problem.
- 2:41 Day by day, the models are coming where you can dump a lot of context, a lot of data.
- 2:47 But does that make sure that the results you are getting is smart enough to give you everything or smart enough to decide what's important?
- 2:56 If you see the current LLM models, we see a pattern where it takes the initial inputs you provide, it takes the last inputs, but the in-between context is basically removed.
- 3:11 So, they don't focus on the in-between context, agents look at the starting point, end point and try to provide you the results.
- 3:18 This is like a U-curve where
- 3:22 Some of the things from the start, some of the things from the end make sense, but whatever you are providing in between, that is not taken up.
- 3:29 Yeah?
- 3:30 How do you know this?
- 3:31 This is something which we are working on and we are actually benchmarking things.
- 3:35 So when we create agents, we try to see this is the context we provide to the agents.
- 3:40 Does it take this context into effect and also give us the results?
- 3:45 So we are working with multi-agent architecture where for each of the tasks we do for code reviews, we give the task to an agent and say, okay,
- 3:54 give us the result now every time we for example code reviews we try to see can we give all the context can we give the whole code base for example and see if we can get the results but we see that that whenever we start working with that the initial prompt or the initial goal which we start with that is in focus if we give something at the end as an input that is in focus but all between context like i have jira i have mcps can you look into that
- 4:22 the LLMs try to get rid of those things and push them to make sense by themselves.
- 4:31 So to have this or to make a way out of this, how we deal with is creating strategic solution for context optimization.
- 4:42 Rather than dumping everything to the models and asking them to be smart enough to find out what is more important,
- 4:50 we usually start to see, okay, what we can do to make it better context for the model.
- 4:55 There are lots of solutions in place, if you see currently, and context engine is a buzzword, like everybody wants to create context engine and everybody wants to provide that, but context engine is like a bouncer, right?
- 5:11 So your high speed car is going and it acts as a bouncer and tells you this is more important.
- 5:17 Now, if you have a large, messy code base,
- 5:20 It makes sense to create a context engine because it creates a search pattern, it creates a ranking logic so that whenever you ask for a task, it looks for those rankings and say this is more important for you, take it and work with it.
- 5:35 The problem is the indexing part takes moderate effort, but the scaling is a challenge.
- 5:41 Like if you start talking about 600 repositories or 700 repositories, the mapping and the indexing starts to slow down and it becomes, again, unpredictable to find or create a context engine if you are not actually into making context engine only.
- 5:59 There are lots of areas where agent can get more context instead of investing highly on context engine.
- 6:07 Hierarchical summarization where instead of creating or going through everything, a summary is created for each file and folder so that when the agents try to find, they can try to read the summary and see if that is
- 6:20 more important to us or not can be a good one.
- 6:24 The only thing is that you need a lot of LLM processing.
- 6:27 So every time a file is created or changed, some of the agents need to go and create a mapping for that.
loading