Videos esY99nYXxR4
How we solved Context Management in Agents — Sally-Ann Delucia
Scene timeline
47 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 225
- whisperx 225
- chunks
- 28
- from 225 cues
- keyframes
- 35
- kept of 47 captured
- frames with text
- 35
- 427 lines read
- chapters
- 13
- from the source metadata
- keyframe bytes
- 4.4 MB
- word timings on 225 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 14:15 | 1m 43s |
stt |
done | — | 2026-08-10 14:17 | 19s |
chunk |
done | — | 2026-08-10 14:17 | 0s |
text_embed |
done | — | 2026-08-10 19:50 | 1s |
keyframe |
done | — | 2026-08-10 14:17 | 1m 22s |
ocr |
done | — | 2026-08-10 14:18 | 10s |
frame_embed |
done | — | 2026-08-10 19:51 | 6s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- AlEngineer0.95
- Escaping the Co0.99
- EUROPE1.00
- Lessons from Buildin1.00
- EU/ACO0.98
- AlEngineer0.97
- EUROPE1.00
-
- arize1.00
- Escaping the Context Window1.00
- AIE1.00
- Lessons from Building Alyx1.00
- ★1.00
- ★1.00
- AlEngineer0.97
- AlEngineer0.97
- 20260.93
- EUROPE1.00
-
- Hi, I'm SallyAnn!1.00
- Head of Product at Arize0.98
- Technical background in data1.00
- science → now building products1.00
- AIE1.00
- for teams1.00
- ★1.00
- ★1.00
- ★1.00
- Hands-on: core contributor to our1.00
- own agent Alyx, so I know the pain0.99
- of building firsthand0.98
- My job: turn that pain into tools1.00
- that actually help1.00
- AlEngineer0.96
- AlEngineer0.98
- 20260.92
- EUROPE1.00
-
- What Is Alyx?0.99
- Ask Alyx1.00
- Identifying Common Person...0.99
- Evaluator Recommendation1.00
- Al Harness to help you build your Al1.00
- Applications1.00
- Advanced Planning1.00
- 40+ Skills0.99
- Core Workflows: Prompt Opt, Data1.00
- Make the most out of Alyx1.00
- AIE1.00
- ★1.00
- ★1.00
- Annotations, etc1.00
- Generation and Augmentation,1.00
- Find critical issues, categorize errors, apply annotations, and build evals0.99
- Summarize my evaluations0.99
- How can I use Arize?0.96
- ★1.00
- Categorize errors and apply annotations1.00
- ★1.00
- ★1.00
- ★1.00
- Find critical issues and create an eval0.99
- Ask Alyx a question, type @ for context0.99
- AlClaude Sonnet 4.50.99
- arize We Make Al Work0.94
- © All Rights Reserved0.97
- AlEngineer0.96
- AlEngineer1.00
- 20260.95
- EUROPE1.00
-
- Agenda1.00
- /011.00
- /021.00
- /030.99
- The Problem1.00
- The Vicious Loop0.99
- Escaping the Loop0.99
- AIE1.00
- ★1.00
- ★1.00
- /041.00
- /051.00
- /060.99
- Long Conversations Break1.00
- Sub-Agents: Moving1.00
- What Still Doesn't Work1.00
- Agents1.00
- Heavy Context Out0.99
- Engineering the future of Al1.00
- AlEngineer0.99
- 20260.99
-
- Andrej Karpathy0.99
- @karpathy1.00
- The stack is0.97
- +1 for "context engineering" over "prompt1.00
- changing.1.00
- AIE1.00
- ★1.00
- engineering".1.00
- .. Too little or of the wrong form and the LLM0.99
- Context is the0.98
- ★1.00
- ★1.00
- doesn't have the right context for optimal1.00
- performance. Too much or too irrelevant and the1.00
- new engineering1.00
- LLM costs might go up and performance might1.00
- problem.1.00
- come down. Doing this well is highly non-trivial.1.00
- And art because of the guiding intuition around0.98
- LLM psychology of people spirits...0.99
- Tweet link1.00
- Engineering the future of Al1.00
- AlEngineer0.99
- 20260.99
-
- iru0.99
- 00.52
- Updates1.00
- 4:591.00
- Update Available0.99
- Google Chrome1.00
- Update Now0.99
- PERSPECTIVE1.00
- Delay 1 Hour0.95
- AIE1.00
- The best context strategy is the one that0.99
- ★1.00
- ★1.00
- lets your agent remember what matters1.00
- and forget what doesn't.0.99
- Engineering the future of Al0.99
- AlEngineer0.97
- 20260.97
-
- Why Context Management Matters1.00
- ★0.54
- AIE1.00
- ★0.99
- Context engineering = choosing what the model sees1.00
- ★1.00
- ★1.00
- Not just staying under a token limit.1.00
- arize We Make Al Work0.96
- All Rights Reserved0.95
- AlEngineer0.97
- AlEngineer1.00
- 20260.99
- EUROPE1.00
-
- ality1.00
- AlEngineer1.00
- EUROPE1.00
- EU/ACC0.97
- Being strategic about contex1.00
- AlEngineer0.99
- EUROPE1.00
-
- Alyx Reality0.98
- One Trace0.98
- wser context0.92
- AIE1.00
- ★1.00
- ★1.00
- this multiptes..0.94
- Being strategic about context was non-negotiable.0.99
- Engineering the future of Al0.97
- AlEngineer1.00
- 20260.97
-
- AIE1.00
- Context management is a product + UX problem0.99
- ★1.00
- Not just an engineering one.1.00
- Engineering the future of Al0.98
- AlEngineer1.00
- 20260.90
-
- AIE1.00
- The system analyzing the data was constrained1.00
- ★1.00
- ★1.00
- by the data.1.00
- AlEngineer0.97
- AlEngineer1.00
- 20260.99
- EUROPE1.00
-
- 11.00
- Control Context1.00
- Escaping the1.00
- AIE1.00
- ★1.00
- Loop1.00
- 21.00
- Separate context from memory0.99
- ★1.00
- ★1.00
- ★1.00
- 31.00
- Move heavy work out1.00
- AlEngineer0.96
- AlEngineer0.99
- 20260.99
- EUROPE1.00
-
- AlEngineer0.99
- Contr1.00
- EUROPE1.00
- Escaping the1.00
- Loop1.00
- Separ1.00
- EU/ACC0.98
- Move1.00
- AlEngineer0.99
- EUROPE1.00
-
- runcation0.99
- AlEngineer0.99
- EUROPE1.00
- NAIVE1.00
- First 100 chars0.96
- rest dropped1.00
- EU/ACC0.91
- AlEngineer0.99
- EUROPE1.00
-
- Naive Truncation0.97
- Result:1.00
- NAIVE1.00
- AIE1.00
- The agent forgot everything.1.00
- ★1.00
- First 100 chars0.98
- ... rest dropped0.90
- Follow-ups looked like new1.00
- conversations.1.00
- Engineering the future of Al1.00
- AlEngineer1.00
- 20260.98
-
- LLM summarization as compression1.00
- Sounded like the obvious solution.1.00
- But:1.00
- AIE1.00
- too inconsistent1.00
- ★1.00
- ★1.00
- No control over what was important1.00
- unreliable1.00
- arizeWe Make Al Work1.00
- Google DeepMind0.99
- AlEngineer1.00
- 20260.89
-
- Solution: Smart Truncation + Memory0.98
- SMART1.00
- Head1.00
- ..truncated0.90
- Tail1.00
- ID: x (retrievable)0.98
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Memory1.00
- Better context strategies1.00
- Retrieve by ID1.00
- remove duplicate messages/1.00
- tool calls0.97
- keep latest result1.00
- don't reset system prompt0.98
- truncate middle (keep head +1.00
- tail)1.00
- Engineering the future of Al1.00
- AlEngineer0.97
- 20260.96
-
- AIE1.00
- Context decides what the model sees.1.00
- ★1.00
- Memory decides what survives.1.00
- AlEngineer0.97
- AlEngineer1.00
- 20260.92
- EUROPE1.00
Transcript
225 cues· 3,442 words· 18,400 chars
- 0:15 All right, welcome.
- 0:15 Thanks so much for coming today.
- 0:17 I'm here to talk a little bit about context windows, and I'm really excited because I get to talk about something that my team and I have been building for, honestly, close to a year now, which is RIA Agent Alex.
- 0:27 So I'm gonna talk a little bit about some of the lessons we learned about context management and escaping the context window.
- 0:34 So, who am I?
- 0:35 I'm Sally Ann.
- 0:36 I'm the head of product at Arise.
- 0:38 I have a technical background.
- 0:39 I started out in data science, and now I build products for teams.
- 0:42 I'm hands-on.
- 0:43 I'm a core contributor of Alex.
- 0:45 I'm not only a PM, but I also function a little bit as a part-time AI engineer as well.
- 0:50 So I know the pain of building these products firsthand, and it is not easy to build a successful agent.
- 0:55 My job today really is to turn those pains into tools that may actually help AI engine AI PMs.
- 1:01 I'm gonna talk a little bit about Alex.
- 1:03 I don't wanna spend a lot of time on Alex.
- 1:04 If you wanna know more about what we built, come find me in the booth downstairs.
- 1:07 I'll give you a demo, but basically what Alex is is an AI harness.
- 1:11 It's here to help you build your AI applications.
- 1:13 We have advanced planning, 40 plus skills built into it, core workflows across prompt engineering, like prompt optimization, data gen, data augmentation, annotations, et cetera.
- 1:24 That's just a screenshot from our product, but yeah, come find me if you'd like a demo.
- 1:29 For today's talk, I'm gonna talk a little bit about the problem of context engineering, context management, tell you a little bit about a vicious loop that we got stuck in, how we escaped that loop, and then how long conversations can break agents, a little bit about what we learned about sub-agents, and then I'll tell you a little bit about what we're still working on, because we certainly haven't figured everything out.
- 1:47 So, the problem.
- 1:48 I think like mid last year, this term context engineering started to become more and more popular.
- 1:52 This is an X from Andre Caparthi about plus one in context engineering over prompt engineering.
- 1:58 I think very early on, everybody was really, really focused on the prompts, but we started to realize that the context is what really made an agent fail or succeed.
- 2:06 And so the stack has really changed.
- 2:08 We're no longer really focused just on the prompts, we're focused on the new engineering problem, which is context.
- 2:13 So,
- 2:14 My little perspective is the best context strategy is one that lets your agents remember what it needs to and forget what it doesn't.
- 2:22 And so we're gonna talk a little bit about how you do that, but first let's talk about why context management even matters.
- 2:28 So I think a lot of folks think, um, like context management is just like what fits in the window, but context engineering is really choosing strategically what the model sees.
- 2:37 It's really important that you think about what the data is that is most important and not just think about, Oh, I only have X amount of tokens.
- 2:43 Let's shove as much as I can in there and see how it does.
- 2:45 So it's not just saying under that token limit,
- 2:49 It's being strategic about it.
- 2:50 And that's why it really matters.
- 2:52 All these different applications, a lot of times, it's running on top of your context.
- 2:55 And so what you choose to let the model see really matters.
- 2:58 It can make or break the experience there.
- 3:00 And so our reality with Alex is Alex is built on top of Arise, which is our observability platform.
- 3:05 So we have to deal with all of the traces that come with AI agents.
- 3:09 And so we have one trace we are getting
- 3:11 the input from the user, there's prompts, there's all of this metadata, then the user is interacting with Alex.
- 3:16 And so it becomes really large.
- 3:18 And that's just when we're talking about one trace, but what happens when they wanna see patterns across all of their traces?
- 3:22 Well, this just continues to multiply and multiply and multiply.
- 3:26 So being strategic about context was a non-negotiable for us.
- 3:29 We really had to figure out, okay, what was most important for Alex to see?
- 3:32 And how do we handle when it needs to kind of see everything?
loading
Chapters
- 0:00 Introduction and speaker background
- 1:02 Overview of the AI agent, Alyx
- 1:29 The problem: Context engineering vs. prompt engineering
- 4:06 The vicious loop of data growth in AI agents
- 5:16 Why naive truncation failed
- 6:14 Why summarization proved unreliable
- 6:46 The solution: Smart truncation and memory stores
- 8:02 Handling long session challenges
- 9:23 Offloading tasks to sub-agents
- 11:19 Ongoing challenges and future work
- 12:57 Findings from the Claude Code source release
- 13:44 Final key takeaways on context management
- 14:58 Q&A session