read-only demo

Videos 2IxD9OB3XuQ

Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI

index_state ready data_status ok

AI Engineer· published 2026-07-05· 0:22:35· en-US· indexed 2026-08-11 05:18

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:29, 1 of 1 keyframes kept
  2. Shot 1, 0:29 to 0:36, 1 of 1 keyframes kept
  3. Shot 2, 0:36 to 0:49, 1 of 1 keyframes kept
  4. Shot 3, 0:49 to 1:18, 1 of 1 keyframes kept
  5. Shot 4, 1:18 to 1:46, 1 of 1 keyframes kept
  6. Shot 5, 1:46 to 2:30, 1 of 1 keyframes kept
  7. Shot 6, 2:30 to 2:34, 1 of 1 keyframes kept
  8. Shot 7, 2:34 to 3:07, 1 of 1 keyframes kept
  9. Shot 8, 3:07 to 3:44, 1 of 1 keyframes kept
  10. Shot 9, 3:44 to 4:21, 0 of 1 keyframes kept
  11. Shot 10, 4:21 to 4:42, 1 of 1 keyframes kept
  12. Shot 11, 4:42 to 5:12, 1 of 1 keyframes kept
  13. Shot 12, 5:12 to 5:43, 1 of 1 keyframes kept
  14. Shot 13, 5:43 to 6:13, 1 of 1 keyframes kept
  15. Shot 14, 6:13 to 6:24, 1 of 1 keyframes kept
  16. Shot 15, 6:24 to 6:49, 1 of 1 keyframes kept
  17. Shot 16, 6:49 to 7:15, 0 of 1 keyframes kept
  18. Shot 17, 7:15 to 7:40, 1 of 1 keyframes kept
  19. Shot 18, 7:40 to 8:14, 1 of 1 keyframes kept
  20. Shot 19, 8:14 to 8:47, 0 of 1 keyframes kept
  21. Shot 20, 8:47 to 9:12, 1 of 1 keyframes kept
  22. Shot 21, 9:12 to 9:37, 0 of 1 keyframes kept
  23. Shot 22, 9:37 to 10:02, 1 of 1 keyframes kept
  24. Shot 23, 10:02 to 10:32, 1 of 1 keyframes kept
  25. Shot 24, 10:32 to 11:02, 0 of 1 keyframes kept
  26. Shot 25, 11:02 to 11:08, 1 of 1 keyframes kept
  27. Shot 26, 11:08 to 11:45, 1 of 1 keyframes kept
  28. Shot 27, 11:45 to 11:59, 1 of 1 keyframes kept
  29. Shot 28, 11:59 to 12:34, 1 of 1 keyframes kept
  30. Shot 29, 12:34 to 13:00, 1 of 1 keyframes kept
  31. Shot 30, 13:00 to 13:26, 1 of 1 keyframes kept
  32. Shot 31, 13:26 to 14:02, 1 of 1 keyframes kept
  33. Shot 32, 14:02 to 14:37, 0 of 1 keyframes kept
  34. Shot 33, 14:37 to 15:03, 1 of 1 keyframes kept
  35. Shot 34, 15:03 to 15:29, 1 of 1 keyframes kept
  36. Shot 35, 15:29 to 15:43, 1 of 1 keyframes kept
  37. Shot 36, 15:43 to 15:56, 1 of 1 keyframes kept
  38. Shot 37, 15:56 to 16:22, 1 of 1 keyframes kept
  39. Shot 38, 16:22 to 16:49, 1 of 1 keyframes kept
  40. Shot 39, 16:49 to 17:15, 1 of 1 keyframes kept
  41. Shot 40, 17:15 to 17:45, 1 of 1 keyframes kept
  42. Shot 41, 17:45 to 18:15, 0 of 1 keyframes kept
  43. Shot 42, 18:15 to 18:40, 1 of 1 keyframes kept
  44. Shot 43, 18:40 to 19:06, 0 of 1 keyframes kept
  45. Shot 44, 19:06 to 19:29, 1 of 1 keyframes kept
  46. Shot 45, 19:29 to 19:50, 1 of 1 keyframes kept
  47. Shot 46, 19:50 to 20:15, 1 of 1 keyframes kept
  48. Shot 47, 20:15 to 20:37, 1 of 1 keyframes kept
  49. Shot 48, 20:37 to 21:13, 1 of 1 keyframes kept
  50. Shot 49, 21:13 to 21:34, 1 of 1 keyframes kept
  51. Shot 50, 21:34 to 22:04, 1 of 1 keyframes kept
  52. Shot 51, 22:04 to 22:34, 1 of 1 keyframes kept

52 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
207
whisperx 207
chunks
40
from 207 cues
keyframes
44
kept of 52 captured
frames with text
44
602 lines read
chapters
0
from the source metadata
keyframe bytes
4.7 MB
word timings on 206 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:15 1m 36s
stt done 2026-08-11 05:17 22s
chunk done 2026-08-11 05:17 0s
text_embed done 2026-08-11 05:17 0s
keyframe done 2026-08-11 05:17 58s
ocr done 2026-08-11 05:18 20s
frame_embed done 2026-08-11 05:18 7s

Frames, and what the machine read

  • 0:20 #0 done6 line(s)

    shot 0·sharpness 1152.7

    1. Continual Learning for Al Agents:0.99
    2. From Failures to Durable Improvements0.99
    3. Soheil Feizi1.00
    4. Founder & Chief Scientist, RELAl0.97
    5. Associate Prof, CS @ University of Maryland0.98
    6. https://relai.ai1.00
  • 0:30 #1 done5 line(s)

    shot 1·sharpness 567.1

    1. Humans learn from experience.1.00
    2. Human1.00
    3. feedback1.00
    4. act1.00
    5. World1.00
  • 0:47 #2 done7 line(s)

    shot 2·sharpness 970.8

    1. Humans learn from experience. Agents should too.0.98
    2. Agent1.00
    3. feedback1.00
    4. act1.00
    5. World1.00
    6. Continual Learning Loop:0.99
    7. act → get feedback → improve without forgetting1.00
  • 1:09 #3 done20 line(s)

    shot 3·sharpness 2957.6

    1. Continual Learning for an Al Agent0.98
    2. AGENT1.00
    3. MODEL1.00
    4. HARNESS1.00
    5. MEMORY1.00
    6. LLM(s): weights1.00
    7. prompts · skills ·0.94
    8. in-session state1.00
    9. model selection1.00
    10. tools·code·0.95
    11. persistent1.00
    12. workflow1.00
    13. knowledge1.00
    14. Goal: continuously improve0.98
    15. the agent from its experiences1.00
    16. World1.00
    17. without forgetting.1.00
    18. users·tools·data0.99
    19. policies1.00
    20. Agent logs / outputs0.99
  • 1:35 #4 done20 line(s)

    shot 4·sharpness 2952.6

    1. Continual Learning for an Al Agent0.99
    2. AGENT1.00
    3. MODEL1.00
    4. HARNESS1.00
    5. MEMORY1.00
    6. LLM(s): weights1.00
    7. prompts · skills ·0.94
    8. in-session state1.00
    9. model selection1.00
    10. tools·code·0.95
    11. persistent1.00
    12. workflow1.00
    13. knowledge1.00
    14. Goal: continuously improve0.99
    15. the agent from its experiences1.00
    16. World1.00
    17. without forgetting.1.00
    18. users·tools·data1.00
    19. policies1.00
    20. Agent logs / outputs0.99
  • 2:00 #5 done23 line(s)

    shot 5·sharpness 2800.7

    1. Two Challenges in Continual Learning0.99
    2. AGENT1.00
    3. MODEL1.00
    4. HARNESS1.00
    5. MEMORY1.00
    6. LLM(s): weights0.98
    7. prompts · skills ·0.95
    8. in-session state1.00
    9. model selection1.00
    10. tools·code·0.97
    11. persistent1.00
    12. workflow1.00
    13. knowledge1.00
    14. P2: agent optimization0.98
    15. Which layer/ components do0.98
    16. we change, and how?1.00
    17. World1.00
    18. users·tools·data0.99
    19. policies1.00
    20. P1:gettingfeedback1.00
    21. Did the agent do well, or what0.97
    22. Agent logs / outputs0.99
    23. should it have done instead?1.00
  • 2:31 #6 done2 line(s)

    shot 6·sharpness 524.0

    1. Problem 10.99
    2. Where does feedback come from?1.00
  • 2:38 #7 done10 line(s)

    shot 7·sharpness 1506.7

    1. The easy case: benchmark + evaluator0.99
    2. Benchmark1.00
    3. Agent1.00
    4. Evaluator1.00
    5. curated task1.00
    6. runs task0.99
    7. scores output1.00
    8. PASS / FAIL / REWARD0.96
    9. + Feedback0.95
    10. 81.00
  • 3:25 #8 done19 line(s)

    shot 8·sharpness 1880.0

    1. In production, a raw log isn't feedback0.99
    2. Session log1.00
    3. An LLM / code analyzes the log0.99
    4. AUTOMATIC1.00
    5. A model or eval code reads the trace and writes a critique on what to0.99
    6. change.1.00
    7. user: book me a flight to NYC1.00
    8. Scales to every session1.00
    9. agent: searching flights...0.96
    10. agent: called tool get_flights()1.00
    11. agent: returned 3 options1.00
    12. A human gives expert feedback0.99
    13. CRITICAL1.00
    14. user: none of these work - wrong date0.98
    15. Domain experts catch what models miss: subtle correctness, policy,1.00
    16. and taste.0.99
    17. Lower volume0.98
    18. Either way, we now have: session log + feedback0.99
    19. 91.00
  • 4:10 #9 skipped

    shot 9·duplicate of #8

  • 4:37 #10 done12 line(s)

    shot 10·sharpness 1234.8

    1. But it still isn't testable1.00
    2. 0.63
    3. What we have0.97
    4. What we need0.99
    5. log+feedback1.00
    6. a replayable learning environment1.00
    7. "the agent used the wrong date; it should0.99
    8. A simulation that you can re-run with defined grading on0.99
    9. confirm dates first."0.99
    10. thegap1.00
    11. what success looks like0.98
    12. 111.00
  • 4:48 #11 done14 line(s)

    shot 11·sharpness 1899.7

    1. What is a learning environment?0.99
    2. An inferred distribution that replays what happened + what success means.1.00
    3. Observed trace +0.99
    4. Mocked /real tools0.98
    5. Synthetic user1.00
    6. Evaluators1.00
    7. feedback1.00
    8. what the agent can call0.98
    9. whatinteractionrepeats1.00
    10. whatsuccess means0.99
    11. whathappened1.00
    12. The output is executable:0.98
    13. run candidate agents against it, then keep the fix only if it passes.1.00
    14. 121.00
  • 5:16 #12 done14 line(s)

    shot 12·sharpness 1895.1

    1. What is a learning environment?0.98
    2. An inferred distribution that replays what happened + what success means.0.99
    3. Observed trace +0.99
    4. Mocked /real tools0.98
    5. Synthetic user1.00
    6. Evaluators1.00
    7. feedback1.00
    8. what the agent can call0.98
    9. whatinteractionrepeats1.00
    10. whatsuccess means1.00
    11. whathappened1.00
    12. The output is executable:0.98
    13. run candidate agents against it, then keep the fix only if it passes.1.00
    14. 121.00
  • 6:10 #13 done14 line(s)

    shot 13·sharpness 1884.8

    1. What is a learning environment?0.98
    2. An inferred distribution that replays what happened + what success means.1.00
    3. Observed trace +0.99
    4. feedback1.00
    5. Mocked /real tools0.98
    6. Synthetic user1.00
    7. Evaluators1.00
    8. what the agent can call0.98
    9. whatinteractionrepeats1.00
    10. what success means0.97
    11. whathappened1.00
    12. The output is executable:0.99
    13. run candidate agents against it, then keep the fix only if it passes.1.00
    14. 121.00
  • 6:17 #14 done3 line(s)

    shot 14·sharpness 433.7

    1. Problem 21.00
    2. How to optimize the agent?0.97
    3. 131.00
  • 6:34 #15 done17 line(s)

    shot 15·sharpness 1306.2

    1. Three layers to improve the agent1.00
    2. Model1.00
    3. SFT · RL post-training0.98
    4. most expensive0.97
    5. update the weights0.98
    6. Harness1.00
    7. GEPA· trace-to-harness0.98
    8. mostflexible1.00
    9. edit prompts, skills, tools, code0.98
    10. Memory1.00
    11. Letta1.00
    12. mem01.00
    13. consolidation1.00
    14. cheapest1.00
    15. store facts and learned skills1.00
    16. A good learning engine asks for the smallest durable change at the right la0.99
    17. 141.00
  • 7:02 #16 skipped

    shot 16·duplicate of #15

  • 7:32 #17 done17 line(s)

    shot 17·sharpness 1317.2

    1. Three layers to improve the agent1.00
    2. Model1.00
    3. SFT · RL post-training0.98
    4. most expensive0.98
    5. update the weights0.98
    6. Harness1.00
    7. GEPA· trace-to-harness0.98
    8. mostflexible1.00
    9. edit prompts, skills, tools, code0.98
    10. Memory1.00
    11. Letta1.00
    12. mem01.00
    13. consolidation1.00
    14. cheapest1.00
    15. store facts and learned skills1.00
    16. A good learning engine asks for the smallest durable change at the right lay1.00
    17. 141.00
  • 8:00 #18 done13 line(s)

    shot 18·sharpness 3082.4

    1. Updating the model weights1.00
    2. SFT1.00
    3. imitate correct trajectories; needs labeled examples of the right behavior0.99
    4. Supervised fine-tuning0.98
    5. RL post-training1.00
    6. sample, score against a reward or preference signal, reinforce what wins1.00
    7. DPO·GRPO·RLVR1.00
    8. LoRA1.00
    9. limits the set of parameters that can change; cheaper, safer updates0.99
    10. Low-Rank Adaptation0.98
    11. They need: benchmark + evaluator.0.96
    12. Hard to apply to a raw production log (unless we lift it into a replayable envs0.99
    13. 151.00
  • 8:21 #19 skipped

    shot 19·duplicate of #18

  • 9:00 #20 done12 line(s)

    shot 20·sharpness 2014.5

    1. Updating the harness1.00
    2. Rewrite the prompts, skills, and code around the model.0.99
    3. Trace-to-harness1.00
    4. GEPA & prompt search0.99
    5. A coding agent reads the log + feedback and rewrites1.00
    6. Mutate prompts, score each candidate, keep the1.00
    7. a prompt, adds a tool, or patches the workflow.1.00
    8. winners; evolutionary optimization of the harness.1.00
    9. Works on (log + feedback) but mostly vibe-based: no0.99
    10. Testable but needs a benchmark to score against.1.00
    11. test that the change helped.1.00
    12. 161.00
  • 9:17 #21 skipped

    shot 21·duplicate of #20

  • 9:57 #22 done13 line(s)

    shot 22·sharpness 2014.5

    1. Updating the harness1.00
    2. Rewrite the prompts, skills, and code around the model.0.99
    3. 0.87
    4. Trace-to-harness1.00
    5. GEPA & prompt search0.99
    6. A coding agent reads the log + feedback and rewrites1.00
    7. Mutate prompts, score each candidate, keep the1.00
    8. a prompt, adds a tool, or patches the workflow.1.00
    9. winners; evolutionary optimization of the harness.1.00
    10. Works on (log + feedback) but mostly vibe-based: no0.99
    11. Testable but needs a benchmark to score against.1.00
    12. test that the change helped.1.00
    13. 161.00
  • 10:20 #23 done12 line(s)

    shot 23·sharpness 2726.9

    1. Updatingmemory1.00
    2. Write down facts and distill skills, so the agent doesn't rediscover them.1.00
    3. Information memory1.00
    4. store a fact or correction; e.g., "always confirm the date before booking"1.00
    5. Letta1.00
    6. ·mem00.90
    7. Skill distillation0.99
    8. compress a successful trajectory into a reusable how-to packet0.99
    9. skills·SKILL.md1.00
    10. (sometimes viewed as a part of harness)0.99
    11. Cheapest and fastest; works directly on (log + feedback) but usually unverifie0.99
    12. 171.00

Transcript

207 cues· 3,120 words· 17,895 chars

  1. 0:01 Hi, everyone.
  2. 0:02 My name is Sohail Faizi.
  3. 0:03 I'm founder and CSO at Relai.
  4. 0:06 I'm also an associate professor in the computer science department at University of Maryland.
  5. 0:10 Today, I'm going to talk about continual learning for AI agents, how we can go from failures to durable improvements.
  6. 0:19 And if you're interested in any of the tools that I'll be talking in this presentation, you can visit at our website, Relai.ai.
  7. 0:28 Let's get started.
  8. 0:30 Humans learn mainly from experience by interacting with the world and getting feedback.
  9. 0:37 The goal of continual learning is to imitate the same for agents so they can also learn from experience by acting, getting feedback and improving without forgetting.
  10. 0:50 All right, so here's basically a bigger picture of how continual learning for agents look like.
  11. 0:58 So here is an agent that interact with the world, with diverse users, with complex tools, with various data policies.
  12. 1:07 And as I mentioned, the goal is to continuously improve the agent from its experience without forgetting.
  13. 1:12 And this learning can happen in different layers of the agent.
  14. 1:16 It can happen in the model layer where potentially we can change weights of LLMs or other models used in the agent or use different types of models in the agent.
  15. 1:27 It can happen in the harness layer where it brings the context
  16. 1:32 proper context to the LLM with components like prompts, skills, tools, code, workflow.
  17. 1:39 And it can also happen in the memory, either in session memory or persistent memory of the agent.
  18. 1:48 But there are two, I would say, fundamental challenges in continual learning for agents.
  19. 1:53 The first challenge is how to get feedback.
  20. 1:55 How do we know if the agent did well?
  21. 1:58 And if not, what should it have done instead?
  22. 2:01 That's basically the first part, getting the feedback.
  23. 2:04 And the second part is how we can act upon that feedback, how we can optimize and improve agent and learn from that feedback.
  24. 2:13 Which layer, which component do we need to change and also how?
  25. 2:17 I'll be talking about these two challenges, current approaches in order to deal with them and also provide some perspective of how we think about these two problems.
  26. 2:28 So let's get started with the first problem.
  27. 2:31 Where does the feedback come from?
  28. 2:34 So the easy case is when we have a benchmark and some evaluators on that benchmark, so the agent can run tasks from the benchmark.
  29. 2:43 Now we have the evaluators in order to score, and we can get grades like pass, fail, or reward, as well as potentially some feedback on the agent behavior and agent performance.
  30. 2:55 This is usually what is happening during the development time that different teams, they curate benchmarks in order to understand the performance of the agent
  31. 3:03 in certain applications.
  32. 3:08 But in production, we don't have such benchmark.
  33. 3:11 We have logs.
  34. 3:12 Here's an example of a session log where a user is interacting with the agent.
  35. 3:17 Maybe the user is not very happy with the way the agent is behaving, but we don't have any explicit feedback.
  36. 3:23 So there are two ways of getting such feedback.
  37. 3:26 on such session logs.
  38. 3:29 One is automatic using some other models or LLMs or code in order to analyze the log and provide feedback.
  39. 3:37 In some cases, even the agent itself can look at it is log and provide some critics of it.
  40. 3:43 It is automatic and it is scalable.
  41. 3:47 The second approach is where we have human experts to look at some handful of these logs and provide some domain expert feedback on those agent outputs.
  42. 3:59 This is lowering the volume, but it is critical because it provides expert knowledge on the behavior of the agent and it is alignment with the way that we want agent to behave in those applications.
  43. 4:12 Either way, now we have
  44. 4:14 session log plus some feedback on those logs.
  45. 4:19 Is it enough?
  46. 4:21 The answer is no, because it is still not testable.
  47. 4:25 Here we have log and feedback, but what we really need is a replayable learning environment, a simulation that we can rerun with defined grading on what success looks like, not one instance of what happened and the feedback on top of it.
  48. 4:43 So what is a learning environment?
  49. 4:45 Here we are inferring a distribution from one observation that replace what happened and what success means.
  50. 4:53 The input is what we have, some session logs and feedback.

Open at this second