Videos SbcQYbrvAfI
Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize
Scene timeline
164 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 596
- whisperx 596
- chunks
- 92
- from 596 cues
- keyframes
- 130
- kept of 164 captured
- frames with text
- 130
- 6,091 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 18.6 MB
- word timings on 596 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 23:09 | 1m 05s |
stt |
done | — | 2026-08-10 23:10 | 59s |
chunk |
done | — | 2026-08-10 23:11 | 0s |
text_embed |
done | — | 2026-08-10 23:11 | 1s |
keyframe |
done | — | 2026-08-10 23:11 | 1m 38s |
ocr |
done | — | 2026-08-10 23:12 | 2m 21s |
frame_embed |
done | — | 2026-08-10 23:15 | 22s |
Frames, and what the machine read
-
- ANTHROPIC0.99
- Cline0.99
- Google DeepMind0.97
- HumanLayer1.00
- arize1.00
- replit1.00
- Google DeepMind1.00
- CURSOR1.00
-
- PRESENTING SPONSOR0.99
- Google DeepMind0.99
-
- PLATINUMSPONSOR1.00
- ANTHROP\C1.00
-
- arize1.00
- Applied Prompt Learning:0.98
- Building a Eval-Driven1.00
- Optimization Loop1.00
- SallyAnn DeLucia & Fuad Ali0.98
- 2025-11-22 12:33:590.99
- JFK27-B1.3001.00
-
- arize1.00
- Applied Prompt Learning0.99
- Building a Eval-Driven1.00
- Optimization Loop1.00
- SallyAnn DeLucia & Fuad Ali0.99
- a0.70
- 2025-11-22 12:34:300.97
- JFK27-B1.3001.00
-
- Agenda1.00
- /010.98
- /020.98
- /030.98
- Why Agents Fail0.99
- What Is Prompt1.00
- Case Study: Coding0.98
- Today1.00
- Learning?1.00
- Agents1.00
- 1040.89
- /050.98
- Prompt Learning vs0.98
- Workshop1.00
- GEPA1.00
- 2025-11-2212:35:051.00
- JFK27-B1.3001.00
-
- Agenda1.00
- /010.99
- /020.97
- /030.98
- Why Agents Fail1.00
- What Is Prompt1.00
- Case Study: Coding1.00
- Today1.00
- Learning?1.00
- Agents1.00
- /040.98
- /050.96
- Prompt Learning vs0.99
- Workshop1.00
- GEPA1.00
- 2025-11-22 12:35:150.99
-
- Agenda1.00
- /010.97
- /030.92
- Why Agents Fail0.99
- What Is Prompt1.00
- Case Study: Coding0.98
- Today1.00
- Learning?1.00
- Agents1.00
- /040.96
- /050.98
- Prompt Learning vs1.00
- Workshop1.00
- GEPA1.00
- 2025-11-22 12:35:330.97
- JFK27-B1.3001.00
-
- Where Agents are Breaking in 20251.00
- No System Instructions Learned1.00
- No Planning or1.00
- Missing Tools0.98
- Very Static Planning1.00
- From Environment0.98
- Tool Guidance1.00
- Missing Context / State0.99
- Management1.00
- (Pre Pruned Data)1.00
- 2025-11-2212:36:031.00
- arizeWe Make Al Work0.99
-
- Core Issues Distilled1.00
- Adaptability & Self1.00
- Determinism vs1.00
- Context1.00
- Learning1.00
- Non Determinism1.00
- Engineering1.00
- Balance1.00
- No System Instructions1.00
- No Planning or1.00
- Missing Tools1.00
- Very Static Planning1.00
- Tool Guidance0.99
- Learned From Environment1.00
- Missing Context0.99
- (Pre Pruned Data)1.00
- 2025-11-22 12:36:560.97
- arizeWe Make Al Work0.98
-
- 1 Other Issue l'd Like to Mention0.98
- Technical Users1.00
- Domain Experts0.97
- Al Engineer0.96
- Data1.00
- 1010101011.00
- Scientist1.00
- Subject Matter1.00
- Al Product1.00
- Experts1.00
- Manager1.00
- Developer1.00
- Responsibilities1.00
- Responsibilities1.00
- Code/Automation1.00
- Domain Prompt engineering1.00
- Pipelines/Frameworks1.00
- Track and run evals0.99
- Application Performance / Costs0.99
- Ensure product success1.00
- 2025-11-22 12:37:280.98
-
- 1 Other Issue l'd Like to Mention0.98
- Technical Users1.00
- Domain Experts0.99
- Al Engineer1.00
- Data1.00
- Scientist1.00
- Subject Matter1.00
- Al Product0.99
- Experts1.00
- Manager1.00
- Developer0.98
- Responsibilities1.00
- Responsibilities1.00
- Code/Automation0.99
- Domain Prompt engineering0.99
- Pipelines/Frameworks1.00
- Track and run evals0.98
- Application Performance / Costs0.98
- Ensure product success1.00
- 2025-11-2212:37:521.00
- JFK27-B1.3001.00
-
- Reinforcement Learning0.99
- RL Model1.00
- Action1.00
- Reward Function1.00
- (Student's Brain)1.00
- (Takes Exam)1.00
- (Exam Scorer)1.00
- Update Weights1.00
- Scalar Reward1.00
- (Student's Brain)1.00
- (Exam Score)0.98
- Algorithm: Gradient Descent, PPO,1.00
- Q-learning1.00
- 2025-11-22 12:38:480.99
- arizeWe Make Models Work0.99
-
- Meta Prompting - Almost...0.99
- Meta Prompting - Ask an LLM to improve your prompt0.99
- Agent1.00
- Output1.00
- Scorer1.00
- (Student)1.00
- (Takes Exam)1.00
- (Exam Scorer)1.00
- Update Prompt1.00
- Scalar Reward0.99
- (Lessons, HWs)0.99
- (Exam Score)1.00
- Algorithm: Meta-Prompting1.00
- (Teaching)1.00
- 2025-11-22 12:39:320.98
- arizeWe Make Models Work0.98
-
- Prompt Learning1.00
- Agent1.00
- Output1.00
- LLM Evals1.00
- (Student)1.00
- (Takes Exam)1.00
- (Teacher)1.00
- English Feedback1.00
- Update Prompt1.00
- • which answers0.95
- (Lessons, HWs)1.00
- were wrong1.00
- WHY answers1.00
- were wrong1.00
- Algorithm: Meta-Prompting1.00
- WHERE student1.00
- (Teaching)1.00
- needs to study1.00
- 2025-11-22 12:39:550.99
- arize We Make Models Work0.95
-
- Traditional Prompt Optimization1.00
- Formulated Like an ML Problem1.00
- Data1.00
- Prediction1.00
- Prompt1.00
- Labels1.00
- X0.87
- 二0.70
- Optimize This1.00
- Maximize This0.98
- 2025-11-22 12:40:460.98
- arizeWe Make Models Work0.98
-
- System Prompt Learning1.00
- Data1.00
- Human Instrunctions,1.00
- Eval Explanations1.00
- Prompt1.00
- Prediction1.00
- Labels1.00
- Eval Explanations,1.00
- why Failed0.94
- Why it Failed0.97
- AI Why it Failed0.96
- +0.58
- 十0.51
- 十0.69
- +0.58
- Add Instructions or changes to System Prompt here,0.99
- to help it improve1.00
- 2025-11-22 12:41:091.00
- arizeWe Make Models Work0.97
-
- System Prompt Learning1.00
- Data0.99
- Human Instrunetions,0.96
- Why it Faled0.90
- Eval Explanations0.99
- AI Why it Failed0.91
- Prompt0.98
- Prediction1.00
- Labels0.99
- Eval Explanations,1.00
- Why Failed0.97
- Add Instructions or changes to Systen Prompt here,0.98
- to help it improve0.97
- 2025-11-2212:42:001.00
- JFK27-B1.3001.00
-
- Optimizing Coding Agents, just through their Prompts1.00
- Claude1.00
- cline0.95
- Claude Cade swstem prompt0.90
- Cine system prompt0.97
- You are a Cloude agent bulit on0.85
- You are Cline. a highly skilled0.95
- antheopie's Claude Apent S0K0.74
- softwore engineer with extensive0.98
- knowledge in mong pregranming0.89
- helps wsers with software0.89
- You are an interactive CLI tool that0.91
- patterns. and best peactices.0.92
- languages, fromenorks, design0.78
- engineering tasks. Use the0.91
- ovasloble to you to assist the user0.89
- instructions below and the toola0.97
- CLAUDE1.00
- Rules (--append-systen-prompt)0.89
- Rules (./clinerules)0.99
- <Empty>0.93
- <Empty>0.87
- CODE1.00
- 2025-11-2212:42:151.00
- JFK27-B1.3000.96
-
- Optimizing Coding Agents, just through their Prompts1.00
- 米Claude1.00
- Cline0.99
- The collaborative1.00
- coding agent1.00
- for complex work0.98
- Cline system prompt0.99
- Claude Code system prompt0.99
- You are a Claude agent, built on1.00
- You are Cline, a highly skilled1.00
- Anthropic's Claude Agent SDK.1.00
- software engineer with extensive0.99
- knowledge in many programming1.00
- You are an interactive CLI tool that1.00
- languages, frameworks, design0.99
- helps users with software0.99
- patterns, and best practices...1.00
- engineering tasks. Use the1.00
- instructions below and the tools1.00
- available to you to assist the user.0.99
- *Welcome to Claude Code0.99
- Rules (./clinerules)1.00
- Rules (--append-system-prompt)1.00
- <Empty>1.00
- <Empty>1.00
- Press Enter to continue1.00
- 2025-11-22 12:42:350.99
- arize1.00
- We Make Al Work0.97
-
- Coding Agents on SWE-Bench Lite, No Prompt Changes0.99
- CLAUDE1.00
- Cline1.00
- CODE0.99
- Sonnet 4-50.99
- GPT 4.10.99
- Sonnet 4-50.99
- Haiku4.51.00
- Cost: $3/1M tokens0.99
- Cost: $2/1M tokens0.97
- Cost: $3/1M tokens1.00
- Cost: $1/1M tokens1.00
- Latency:1.00
- Latency:1.00
- Latency:1.00
- Latency:1.00
- 30.00%1.00
- 18.67%1.00
- 40.00%1.00
- 18.67%1.00
- Github Issues0.98
- Github Issues0.96
- GithubIssues1.00
- GithubIssues1.00
- resolved1.00
- resolved1.00
- resolved1.00
- resolved1.00
- 2025-11-2212:42:581.00
- arizeWe Make Al Work0.93
-
- Optimizing Coding Agent System Prompt1.00
- Claude Code system prompt1.00
- OLD1.00
- Claude Code system prompt1.00
- NEW1.00
- You are a Claude agent, built on Anthropic's...0.99
- You are a Claude agent, built on Anthropic's...1.00
- Rules Section0.99
- Rules Section1.00
- <Empty>1.00
- 1.0.99
- When dealing with errors or exceptions,1.00
- consider the immediate cause and0.99
- underlying issues that may contribute1.00
- to the problem.1.00
- 2.0.98
- Ensure changes align with the overall1.00
- system design; avoid ad-hoc fixes that1.00
- introduce technical debt.1.00
- 3.0.97
- Any change should be accompanied by0.99
- appropriate tests, covering edge cases1.00
- and ensuring correctness and0.99
- robustness.1.00
- 4.1.00
- Always consider anomalies, None values,0.98
- and unexpected inputs when modifying1.00
- data flows.1.00
- 2025-11-22 12:43:260.97
- 5.0.99
- Ensure changes don't introduce0.99
Transcript
596 cues· 9,599 words· 51,149 chars
- 0:21 Hey, everyone.
- 0:21 Gonna get started here.
- 0:23 Thanks so much for joining us today.
- 0:25 I'm Sally Anne.
- 0:26 I'm the director of PROMPT at Arise.
- 0:28 I'm gonna be walking you through some of PROMPT learning.
- 0:30 We're actually gonna be building an algorithm optimization loop for the part of the workshop.
- 0:35 I have a particular background in data science.
- 0:38 Before I make my way over to product, I do like to still be touching code today.
- 0:43 I think one of my bigger projects that I work on is building our own agent into our platform.
- 0:48 So I'm very familiar with all of the pain points and how important it is to optimize your prompt.
- 0:53 So I'm going to just set the scene, make sure everybody here has context on what we're going to be doing, and then we'll jump into the code.
- 1:01 I love me, so I'll let you do a little bit of an intro.
- 1:03 Yeah, thank you so much.
- 1:04 Great to meet all of you.
- 1:05 Excited to be walking through prompt learning with you all.
- 1:09 I don't know if you got a chance to see our partners talk yesterday, but hopefully that gave you some good background on how powerful that prompting and prompt learning can be.
- 1:18 So my name is Will.
- 1:19 I'm a product manager here at Arise as well.
- 1:21 And like Sally said, we like to stay in code.
- 1:23 We'll be doing a few slides, then we'll walk through the code and we'll be floating around helping you guys debug and things like that.
- 1:29 My background is also
- 1:30 technical.
- 1:31 So I was back in distributed systems engineering for a long time.
- 1:34 So no stranger to how important observability infrastructure really is.
- 1:38 And I think it's an appropriate setting in AWS for that.
- 1:41 So yeah, excited to dive deep into front loading with you all.
- 1:44 Thank you.
- 1:46 So, all right, so we're going to get started.
- 1:48 Just give you a little bit of an agenda of the things I'm going to be covering.
- 1:50 So we're going to talk about why agents fail today, what is evening prompt learning.
- 1:54 I want to go through a case study, kind of show you all why this actually works.
- 1:58 And we'll talk about prompt learning versus GEPA.
- 1:59 I think every week I have a few people come up to me over the conference about, like, what about GEPA?
- 2:04 We have some benchmarking against that, and then we'll hop into our workshop.
- 2:08 But with this, I want to ask a question.
- 2:09 How many people here are building agents today?
- 2:12 Okay, that's what I expected.
- 2:14 And how many people actually feel like the agents they're building are reliable?
- 2:19 Yeah, that's what I also thought.
- 2:20 So let's talk a little bit about why agents fail today.
- 2:23 So why do they fail?
- 2:24 Well, there's a few things that we're seeing with a lot of our folks, and we're seeing even internally as we build with Alex for why agents are breaking.
- 2:31 So I think that a lot of times it's not because the models are weak.
- 2:35 It's a lot of times the environment and the instructions are weak.
- 2:39 So having no instructions from their learned environment
- 2:44 No planning or very static planning.
- 2:46 I feel like a lot of agents right now don't have planning.
- 2:48 We do have some good examples of planning, like we have Cloud Code, Cursor.
- 2:52 Those are really great examples, but I'm not seeing it make its way into every agent that I come across.
loading