read-only demo

Videos 9HbzAWnKbo4

From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize

index_state ready data_status ok

AI Engineer· published 2026-07-24· 0:20:35· en-US· indexed 2026-08-10 19:40

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:52, 1 of 1 keyframes kept
  5. Shot 4, 0:52 to 1:02, 1 of 1 keyframes kept
  6. Shot 5, 1:02 to 1:34, 1 of 1 keyframes kept
  7. Shot 6, 1:34 to 1:47, 1 of 1 keyframes kept
  8. Shot 7, 1:47 to 1:49, 1 of 1 keyframes kept
  9. Shot 8, 1:49 to 1:53, 0 of 1 keyframes kept
  10. Shot 9, 1:53 to 2:25, 0 of 1 keyframes kept
  11. Shot 10, 2:25 to 2:57, 1 of 1 keyframes kept
  12. Shot 11, 2:57 to 3:29, 0 of 1 keyframes kept
  13. Shot 12, 3:29 to 3:41, 1 of 1 keyframes kept
  14. Shot 13, 3:41 to 4:10, 1 of 1 keyframes kept
  15. Shot 14, 4:10 to 4:40, 1 of 1 keyframes kept
  16. Shot 15, 4:40 to 5:11, 0 of 1 keyframes kept
  17. Shot 16, 5:11 to 5:38, 1 of 1 keyframes kept
  18. Shot 17, 5:38 to 6:05, 0 of 1 keyframes kept
  19. Shot 18, 6:05 to 6:28, 1 of 1 keyframes kept
  20. Shot 19, 6:28 to 6:54, 1 of 1 keyframes kept
  21. Shot 20, 6:54 to 7:19, 1 of 1 keyframes kept
  22. Shot 21, 7:19 to 7:58, 1 of 1 keyframes kept
  23. Shot 22, 7:58 to 8:04, 1 of 1 keyframes kept
  24. Shot 23, 8:04 to 8:10, 1 of 1 keyframes kept
  25. Shot 24, 8:10 to 8:12, 1 of 1 keyframes kept
  26. Shot 25, 8:12 to 8:13, 1 of 1 keyframes kept
  27. Shot 26, 8:13 to 8:24, 0 of 1 keyframes kept
  28. Shot 27, 8:24 to 8:49, 0 of 1 keyframes kept
  29. Shot 28, 8:49 to 9:03, 1 of 1 keyframes kept
  30. Shot 29, 9:03 to 9:29, 1 of 1 keyframes kept
  31. Shot 30, 9:29 to 9:31, 1 of 1 keyframes kept
  32. Shot 31, 9:31 to 10:16, 1 of 1 keyframes kept
  33. Shot 32, 10:16 to 10:17, 1 of 1 keyframes kept
  34. Shot 33, 10:17 to 10:19, 0 of 1 keyframes kept
  35. Shot 34, 10:19 to 11:01, 1 of 1 keyframes kept
  36. Shot 35, 11:01 to 11:03, 1 of 1 keyframes kept
  37. Shot 36, 11:03 to 11:04, 1 of 1 keyframes kept
  38. Shot 37, 11:04 to 11:06, 1 of 1 keyframes kept
  39. Shot 38, 11:06 to 11:08, 1 of 1 keyframes kept
  40. Shot 39, 11:08 to 11:10, 1 of 1 keyframes kept
  41. Shot 40, 11:10 to 11:36, 1 of 1 keyframes kept
  42. Shot 41, 11:36 to 12:05, 1 of 1 keyframes kept
  43. Shot 42, 12:05 to 12:21, 1 of 1 keyframes kept
  44. Shot 43, 12:21 to 12:24, 1 of 1 keyframes kept
  45. Shot 44, 12:24 to 13:10, 0 of 1 keyframes kept
  46. Shot 45, 13:10 to 13:11, 1 of 1 keyframes kept
  47. Shot 46, 13:11 to 13:41, 1 of 1 keyframes kept
  48. Shot 47, 13:41 to 14:11, 0 of 1 keyframes kept
  49. Shot 48, 14:11 to 14:13, 0 of 1 keyframes kept
  50. Shot 49, 14:13 to 14:20, 1 of 1 keyframes kept
  51. Shot 50, 14:20 to 14:21, 1 of 1 keyframes kept
  52. Shot 51, 14:21 to 14:38, 1 of 1 keyframes kept
  53. Shot 52, 14:38 to 15:01, 1 of 1 keyframes kept
  54. Shot 53, 15:01 to 15:02, 1 of 1 keyframes kept
  55. Shot 54, 15:02 to 15:05, 0 of 1 keyframes kept
  56. Shot 55, 15:05 to 15:47, 1 of 1 keyframes kept
  57. Shot 56, 15:47 to 15:49, 0 of 1 keyframes kept
  58. Shot 57, 15:49 to 16:15, 1 of 1 keyframes kept
  59. Shot 58, 16:15 to 16:42, 0 of 1 keyframes kept
  60. Shot 59, 16:42 to 17:08, 0 of 1 keyframes kept
  61. Shot 60, 17:08 to 17:35, 0 of 1 keyframes kept
  62. Shot 61, 17:35 to 18:02, 0 of 1 keyframes kept
  63. Shot 62, 18:02 to 18:33, 1 of 1 keyframes kept
  64. Shot 63, 18:33 to 19:04, 1 of 1 keyframes kept
  65. Shot 64, 19:04 to 19:35, 0 of 1 keyframes kept
  66. Shot 65, 19:35 to 20:06, 0 of 1 keyframes kept
  67. Shot 66, 20:06 to 20:18, 1 of 1 keyframes kept
  68. Shot 67, 20:18 to 20:35, 0 of 1 keyframes kept

68 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
228
whisperx 228
chunks
36
from 228 cues
keyframes
48
kept of 68 captured
frames with text
48
1,766 lines read
chapters
12
from the source metadata
keyframe bytes
7.7 MB
word timings on 228 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:50 0s
stt done 2026-08-09 07:13 22s
chunk done 2026-08-09 07:14 0s
text_embed done 2026-08-10 19:40 1s
keyframe done 2026-08-09 07:14 2m 49s
ocr done 2026-08-09 07:17 31s
frame_embed done 2026-08-10 19:40 8s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 459.9

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.1

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2744.8

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of1.00
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:28 #3 done3 line(s)

    shot 3·sharpness 208.6

    1. arize0.85
    2. AlEngineer0.99
    3. World's Fair1.00
  • 0:57 #4 done2 line(s)

    shot 4·sharpness 216.6

    1. AlEngineer0.99
    2. World's Fair1.00
  • 1:21 #5 done10 line(s)

    shot 5·sharpness 1471.9

    1. AlEngineer0.98
    2. World's Fair0.97
    3. arize1.00
    4. We make agents work1.00
    5. PRESENTED BY1.00
    6. From Signal to PR0.98
    7. Microsoft1.00
    8. Observability & Evaluation1.00
    9. World's Fair0.98
    10. Engineering the future of Al0.99
  • 1:46 #6 done13 line(s)

    shot 6·sharpness 1423.3

    1. AlEngineer0.97
    2. World'sFair1.00
    3. 02:14 - ERROR RATE SPIKING0.99
    4. PRESENTED BY1.00
    5. It's 2am, the platform is down and0.99
    6. Microsoft1.00
    7. the fix is still hours away.1.00
    8. A human still has to wake up, find the trace using1.00
    9. and reconstruct what1.00
    10. happened before a single line changes.0.98
    11. World's Fair0.97
    12. TRACK 5· JULY 1, 20260.95
    13. Evals1.00
  • 1:48 #7 done24 line(s)

    shot 7·sharpness 1166.6

    1. AlEngineer0.98
    2. WHY NOW1.00
    3. World's Fair0.96
    4. Who consumes your telemetry?0.99
    5. Phase 1· Software 1.00.93
    6. Phase 2 · Software 2.00.98
    7. Phase 3·Autonomous1.00
    8. PRESENTED BY1.00
    9. Application1.00
    10. Application1.00
    11. Application1.00
    12. Microsoft1.00
    13. Observability0.93
    14. Observability0.98
    15. Observability1.00
    16. Human dev1.00
    17. + agent0.94
    18. Coding agent1.00
    19. The human is the only reader.0.98
    20. Human reads, prompts an agent to write.1.00
    21. The agent reads it directly.1.00
    22. Worild's Fair0.93
    23. TRACK 5· JULY 1, 20260.95
    24. Evals1.00
  • 1:50 #8 skipped

    shot 8·duplicate of #6

  • 1:59 #9 skipped

    shot 9·duplicate of #7

  • 2:41 #10 done22 line(s)

    shot 10·sharpness 1044.7

    1. AlEngineer0.98
    2. WHY NOW1.00
    3. World's Fair0.99
    4. Who consumes your telemetry?1.00
    5. Phase 1· Software 1.00.94
    6. Phase 2 · Software 2.00.97
    7. Phase 3·Autonomous0.99
    8. Application1.00
    9. Application1.00
    10. Application1.00
    11. Observability1.00
    12. Observability0.96
    13. Observability1.00
    14. Human dev1.00
    15. + agent0.95
    16. Coding agent1.00
    17. The human is the only reader.0.98
    18. Human reads, prompts an agent to write.1.00
    19. The agent reads it directly.1.00
    20. World's F0.94
    21. TRACK 5· JULY 1, 20260.95
    22. Evals1.00
  • 3:04 #11 skipped

    shot 11·duplicate of #10

  • 3:35 #12 done11 line(s)

    shot 12·sharpness 919.0

    1. AlEngineer0.97
    2. World'sFair1.00
    3. THE GAP1.00
    4. You can now build at agent speed0.98
    5. Coding agents· minutes0.97
    6. You still can't improve at agent speed0.99
    7. Humans reading dashboards· days0.97
    8. arize0.96
    9. World'sFair0.99
    10. TRACK 5· JULY 1, 20260.96
    11. Evals1.00
  • 3:58 #13 done14 line(s)

    shot 13·sharpness 1088.4

    1. AlEngineer0.97
    2. World'sFair1.00
    3. The bottleneck is not the fix. It's building1.00
    4. the evidence and confidence in the fix.1.00
    5. Finding the traces, reading the evals, digesting the data and forming a hypothesis.0.99
    6. Find the traces1.00
    7. Read the evals0.99
    8. Digest the data1.00
    9. Form a hypothesis0.97
    10. Localize the bug1.00
    11. Change one line1.00
    12. World'sFa0.96
    13. TRACK 5· JULY 1,20260.94
    14. Evals1.00
  • 4:28 #14 done16 line(s)

    shot 14·sharpness 741.2

    1. AlEngineer0.99
    2. THE MOVE1.00
    3. World'sFair1.00
    4. Invert the loop1.00
    5. before1.00
    6. Evidence1.00
    7. Human debugs1.00
    8. Agent writes fix0.99
    9. Ship1.00
    10. after1.00
    11. Evidence1.00
    12. Agent investigates & writes fix1.00
    13. Human reviews1.00
    14. World'sF1.00
    15. TRACK 5· JULY 1, 20260.95
    16. Evals1.00
  • 4:44 #15 skipped

    shot 15·duplicate of #14

  • 5:24 #16 done20 line(s)

    shot 16·sharpness 1261.8

    1. AlEngineer0.99
    2. World's Fair0.98
    3. THE ANATOMY1.00
    4. What a self-improving agent is made of0.99
    5. 011.00
    6. what happened1.00
    7. 021.00
    8. enough to fix it1.00
    9. 031.00
    10. · when it runs0.97
    11. Event Evidence1.00
    12. Context + Skills1.00
    13. Trigger1.00
    14. Traces + evals0.99
    15. Logs, APM + the repo1.00
    16. Periodic / on error0.97
    17. Evidence + context is everything the agent needs to fix it. A trigger decides when it runs.0.99
    18. World's Fai0.94
    19. TRACK 5· JULY 1,20260.95
    20. Evals1.00
  • 5:57 #17 skipped

    shot 17·duplicate of #16

  • 6:12 #18 done21 line(s)

    shot 18·sharpness 1131.7

    1. AlEngineer0.99
    2. Evidence1.00
    3. Context1.00
    4. Trigger1.00
    5. World'sFair0.99
    6. PRESENTED BY1.00
    7. ANATOMY·010.98
    8. Traces1.00
    9. Microsoft1.00
    10. Evidence1.00
    11. Every span, input, output, latency, and error the full execution path.1.00
    12. Traces are what happened. Evals are whether it0.98
    13. was good. Together they're the ground truth an1.00
    14. Evals1.00
    15. agent reasons from the same evidence you'd1.00
    16. LLM-as-a-judge and code checks that turn raw runs into pass / fail1.00
    17. scores.1.00
    18. open first.1.00
    19. World'sF0.93
    20. TRACK 5· JULY 1, 20260.95
    21. Evals1.00
  • 6:46 #19 done21 line(s)

    shot 19·sharpness 973.6

    1. AlEngineer0.98
    2. Evidence1.00
    3. Context1.00
    4. Trigger1.00
    5. World's Fair0.96
    6. ANATOMY· 020.94
    7. Context1.00
    8. Evidence says something's wrong. The context plus the skills help you fix.1.00
    9. PRESENTED BY1.00
    10. Microsoft1.00
    11. Traces + evals0.99
    12. Logs & APM0.97
    13. The repo1.00
    14. Enough to fix it1.00
    15. the evidence1.00
    16. embedded agents· Datadog0.98
    17. where the fix lands1.00
    18. pinpoints the exact line0.99
    19. World's Fain0.91
    20. TRACK 5· JULY 1,20260.96
    21. Evals1.00
  • 6:57 #20 done21 line(s)

    shot 20·sharpness 966.6

    1. AlEngineer0.97
    2. Evidence1.00
    3. Context1.00
    4. Trigger1.00
    5. World's Fair0.99
    6. ANATOMY· 020.91
    7. Context1.00
    8. Evidence says something's wrong. The context plus the skills help you fix.1.00
    9. PRESENTED BY1.00
    10. Microsoft1.00
    11. Traces + evals0.99
    12. Logs & APM0.97
    13. The repo1.00
    14. Enough to fix it1.00
    15. the evidence1.00
    16. embedded agents· Datadog0.98
    17. where the fix lands0.99
    18. pinpoints the exact line1.00
    19. Wor0.97
    20. TRACK 5· JULY 1,20260.96
    21. Evals0.93
  • 7:46 #21 done22 line(s)

    shot 21·sharpness 1105.6

    1. AlEngineer0.97
    2. FROM LOCAL TO LONG-RUNNING1.00
    3. World's Fair0.99
    4. The same agent you run locally runs on events1.00
    5. RUNNING LOCAL1.00
    6. SANDBOX1.00
    7. ·event-triggered0.96
    8. Human driving1.00
    9. In the cloud1.00
    10. Coding agent1.00
    11. on events1.00
    12. same agent, running unattended0.98
    13. Coding agent1.00
    14. Arize1.00
    15. Pyroscope0.96
    16. Gcloud Logs1.00
    17. Arize1.00
    18. Pyroscope1.00
    19. Gcloud Logs0.99
    20. World's Fair0.94
    21. TRACK 5· JULY 1, 20260.95
    22. Evals1.00
  • 8:00 #22 done17 line(s)

    shot 22·sharpness 895.5

    1. AlEngineer0.98
    2. Evidence1.00
    3. Context1.00
    4. Trigger1.00
    5. World'sFair1.00
    6. ANATOMY· 030.93
    7. Trigger1.00
    8. periodic1.00
    9. event-driven1.00
    10. On a schedule0.98
    11. On every errored trace1.00
    12. Sweep recent traces and evals on any cadence nightly, hourly, or near1.00
    13. The moment a span fails, the loop kicks off.0.99
    14. real-time. Batch up what's worth fixing.0.99
    15. Work's Fair0.92
    16. TRACK 5· JULY 1, 20260.96
    17. Evals1.00
  • 8:08 #23 done13 line(s)

    shot 23·sharpness 785.5

    1. AlEngineer0.97
    2. ASSEMBLED1.00
    3. World'sFair1.00
    4. The loop, assembled0.97
    5. Evidence1.00
    6. Agent investigates0.98
    7. Pull request1.00
    8. You review1.00
    9. &merge1.00
    10. deploy → the next trace feeds the next pass1.00
    11. World's Fair0.96
    12. TRACK 5· JULY 1, 20260.95
    13. Evals1.00

Transcript

228 cues· 3,262 words· 16,875 chars

  1. 0:12 Well, thank you all.
  2. 0:15 Let me just get set up here.
  3. 0:16 So not just the founder of Verizon, but I tend to build an incredible amount of stuff.
  4. 0:25 Let's see if we get this going here.
  5. 0:29 Oh, sorry, one more second.
  6. 0:34 So not just a founder here, but also a builder.
  7. 0:38 And I do my best to,
  8. 0:42 to build agents, assistants.
  9. 0:45 We have an agent in product.
  10. 0:48 We have an agent in product called Alex, and I think a lot of my experience has come from actually trying to make the stuff work and work well.
  11. 0:59 Our first version of our own agent frankly sucked.
  12. 1:03 It was many years ago, probably two years ago, we weren't the first in the space to do it.
  13. 1:09 And a lot of what we have built has come out of our own experience in building this agent.
  14. 1:16 And Signal is kind of our next generation of this, which is trying to automate a bunch of things, which we do every day, and build it into products that people can use.
  15. 1:26 So I'm going to go through this materials here.
  16. 1:29 I'll try to go fast and try to show you a lot of product, too.
  17. 1:32 I'm a product person.
  18. 1:36 if you built a startup before, you've experienced this, your platform's down, it's late at night and you wanna go fix it.
  19. 1:45 And really, it takes a lot of energy to go do that.
  20. 1:49 And we're going to talk about the automation we built a little bit and what the future looks like.
  21. 1:54 And I truly believe that the future of the observability space is actually changing massively right now.
  22. 2:02 Why is that?
  23. 2:03 Well, observability used to be for humans.
  24. 2:06 It used to be a UI you click, a graph you click, something you look at.
  25. 2:11 And today, I would argue it's a lot of 2.0, which is like this combination of coding agent.
  26. 2:17 So those of you who've built skills, skills for PyroScope, Google Cloud, or whatnot, these skills help you with your human debugging these systems.
  27. 2:27 And really, telemetry is like this smoke thrown off of your system that can allow these agents to go make fixes.
  28. 2:38 It tells you what path in the code it took.
  29. 2:41 Without that, you're guessing, and there's a million paths it could have taken.
  30. 2:44 The data thrown off by your system allows you to go use agents to go debug your software.
  31. 2:55 Evals add another layer to this, but really what we're at here is how do I build systems that autonomously fix themselves?
  32. 3:05 Really, that is what we're after, both AI agents.
  33. 3:08 I put AI into my system.
  34. 3:10 How do I have this thing just improve itself?
  35. 3:13 And today, we're kind of in the 2.0, which is a human making fixes and reviewing things, but there's a future we're all driving towards and throwing off traces, throwing off logs, throwing off way more than you normally would, and having agents run at this for a continuous loop is where we're going.
  36. 3:30 You can build at agent speed, but today,
  37. 3:33 you can't improve your systems really at this agent speed.
  38. 3:36 So those of us feel this kind of governor happening within our products.
  39. 3:41 And the bottleneck is actually not the fix anymore.
  40. 3:46 So those of us who've used these systems and used coding agents with skills, the bottlenecks, a lot of the confidence in, do I have it right?
  41. 3:57 A lot of this is about, is this fix the right one to push?
  42. 4:03 And so these are kind of the challenges here.
  43. 4:06 And then how do you build this loop in a way that just moves faster?
  44. 4:10 And a little bit of the way we've kind of come to do it, and we do it in our system, is we've kind of inverted this loop, which is like a human looks at things and an agent,
  45. 4:23 fixes it to a person now can wake up with an idea of the issues based upon the errors occurred in their system.
  46. 4:31 So the agent is actually, you know, maybe it's not a fix itself, but it's putting up an issue.
  47. 4:36 It's looking at the data before a human even looks at it.
  48. 4:40 And what you move from there is kind of humans grabbing tickets to having some amount of evidence
  49. 4:48 some deep evidence relative to whatever you're looking at already sitting in front of you by the time you actually even look at it.
  50. 4:55 And human review is kind of one thing, but a lot of times maybe you're driving this little investigation a bit from where it started.

Chapters

  1. 0:00 Arize, its agent Alyx, and why v1 sucked
  2. 1:36 Observability is changing: from dashboards to telemetry for agents
  3. 2:55 The goal: systems that fix themselves
  4. 4:14 Inverting the loop: the agent investigates first
  5. 6:08 Traces on a filesystem, the key unlock
  6. 7:23 From your laptop to sandboxes
  7. 8:13 A real fix: the Alyx stream canceled bug
  8. 9:39 Why you should trace ten times more
  9. 11:10 Product demo: Signal, AX, and Phoenix
  10. 13:09 Sandboxes, VPC, and why customers won't call out
  11. 16:20 Q&A: why not just point Claude Code at your data?
  12. 18:04 Q&A: where do the evals come in?

Open at this second