Videos 9HbzAWnKbo4
From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize
Scene timeline
68 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 228
- whisperx 228
- chunks
- 36
- from 228 cues
- keyframes
- 48
- kept of 68 captured
- frames with text
- 48
- 1,766 lines read
- chapters
- 12
- from the source metadata
- keyframe bytes
- 7.7 MB
- word timings on 228 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:50 | 0s |
stt |
done | — | 2026-08-09 07:13 | 22s |
chunk |
done | — | 2026-08-09 07:14 | 0s |
text_embed |
done | — | 2026-08-10 19:40 | 1s |
keyframe |
done | — | 2026-08-09 07:14 | 2m 49s |
ocr |
done | — | 2026-08-09 07:17 | 31s |
frame_embed |
done | — | 2026-08-10 19:40 | 8s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of1.00
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- arize0.85
- AlEngineer0.99
- World's Fair1.00
-
- AlEngineer0.99
- World's Fair1.00
-
- AlEngineer0.98
- World's Fair0.97
- arize1.00
- We make agents work1.00
- PRESENTED BY1.00
- From Signal to PR0.98
- Microsoft1.00
- Observability & Evaluation1.00
- World's Fair0.98
- Engineering the future of Al0.99
-
- AlEngineer0.97
- World'sFair1.00
- 02:14 - ERROR RATE SPIKING0.99
- PRESENTED BY1.00
- It's 2am, the platform is down and0.99
- Microsoft1.00
- the fix is still hours away.1.00
- A human still has to wake up, find the trace using1.00
- and reconstruct what1.00
- happened before a single line changes.0.98
- World's Fair0.97
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.98
- WHY NOW1.00
- World's Fair0.96
- Who consumes your telemetry?0.99
- Phase 1· Software 1.00.93
- Phase 2 · Software 2.00.98
- Phase 3·Autonomous1.00
- PRESENTED BY1.00
- Application1.00
- Application1.00
- Application1.00
- Microsoft1.00
- Observability0.93
- Observability0.98
- Observability1.00
- Human dev1.00
- + agent0.94
- Coding agent1.00
- The human is the only reader.0.98
- Human reads, prompts an agent to write.1.00
- The agent reads it directly.1.00
- Worild's Fair0.93
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.98
- WHY NOW1.00
- World's Fair0.99
- Who consumes your telemetry?1.00
- Phase 1· Software 1.00.94
- Phase 2 · Software 2.00.97
- Phase 3·Autonomous0.99
- Application1.00
- Application1.00
- Application1.00
- Observability1.00
- Observability0.96
- Observability1.00
- Human dev1.00
- + agent0.95
- Coding agent1.00
- The human is the only reader.0.98
- Human reads, prompts an agent to write.1.00
- The agent reads it directly.1.00
- World's F0.94
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.97
- World'sFair1.00
- THE GAP1.00
- You can now build at agent speed0.98
- Coding agents· minutes0.97
- You still can't improve at agent speed0.99
- Humans reading dashboards· days0.97
- arize0.96
- World'sFair0.99
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- AlEngineer0.97
- World'sFair1.00
- The bottleneck is not the fix. It's building1.00
- the evidence and confidence in the fix.1.00
- Finding the traces, reading the evals, digesting the data and forming a hypothesis.0.99
- Find the traces1.00
- Read the evals0.99
- Digest the data1.00
- Form a hypothesis0.97
- Localize the bug1.00
- Change one line1.00
- World'sFa0.96
- TRACK 5· JULY 1,20260.94
- Evals1.00
-
- AlEngineer0.99
- THE MOVE1.00
- World'sFair1.00
- Invert the loop1.00
- before1.00
- Evidence1.00
- Human debugs1.00
- Agent writes fix0.99
- Ship1.00
- after1.00
- Evidence1.00
- Agent investigates & writes fix1.00
- Human reviews1.00
- World'sF1.00
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.99
- World's Fair0.98
- THE ANATOMY1.00
- What a self-improving agent is made of0.99
- 011.00
- what happened1.00
- 021.00
- enough to fix it1.00
- 031.00
- · when it runs0.97
- Event Evidence1.00
- Context + Skills1.00
- Trigger1.00
- Traces + evals0.99
- Logs, APM + the repo1.00
- Periodic / on error0.97
- Evidence + context is everything the agent needs to fix it. A trigger decides when it runs.0.99
- World's Fai0.94
- TRACK 5· JULY 1,20260.95
- Evals1.00
-
- AlEngineer0.99
- Evidence1.00
- Context1.00
- Trigger1.00
- World'sFair0.99
- PRESENTED BY1.00
- ANATOMY·010.98
- Traces1.00
- Microsoft1.00
- Evidence1.00
- Every span, input, output, latency, and error the full execution path.1.00
- Traces are what happened. Evals are whether it0.98
- was good. Together they're the ground truth an1.00
- Evals1.00
- agent reasons from the same evidence you'd1.00
- LLM-as-a-judge and code checks that turn raw runs into pass / fail1.00
- scores.1.00
- open first.1.00
- World'sF0.93
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.98
- Evidence1.00
- Context1.00
- Trigger1.00
- World's Fair0.96
- ANATOMY· 020.94
- Context1.00
- Evidence says something's wrong. The context plus the skills help you fix.1.00
- PRESENTED BY1.00
- Microsoft1.00
- Traces + evals0.99
- Logs & APM0.97
- The repo1.00
- Enough to fix it1.00
- the evidence1.00
- embedded agents· Datadog0.98
- where the fix lands1.00
- pinpoints the exact line0.99
- World's Fain0.91
- TRACK 5· JULY 1,20260.96
- Evals1.00
-
- AlEngineer0.97
- Evidence1.00
- Context1.00
- Trigger1.00
- World's Fair0.99
- ANATOMY· 020.91
- Context1.00
- Evidence says something's wrong. The context plus the skills help you fix.1.00
- PRESENTED BY1.00
- Microsoft1.00
- Traces + evals0.99
- Logs & APM0.97
- The repo1.00
- Enough to fix it1.00
- the evidence1.00
- embedded agents· Datadog0.98
- where the fix lands0.99
- pinpoints the exact line1.00
- Wor0.97
- TRACK 5· JULY 1,20260.96
- Evals0.93
-
- AlEngineer0.97
- FROM LOCAL TO LONG-RUNNING1.00
- World's Fair0.99
- The same agent you run locally runs on events1.00
- RUNNING LOCAL1.00
- SANDBOX1.00
- ·event-triggered0.96
- Human driving1.00
- In the cloud1.00
- Coding agent1.00
- on events1.00
- same agent, running unattended0.98
- Coding agent1.00
- Arize1.00
- Pyroscope0.96
- Gcloud Logs1.00
- Arize1.00
- Pyroscope1.00
- Gcloud Logs0.99
- World's Fair0.94
- TRACK 5· JULY 1, 20260.95
- Evals1.00
-
- AlEngineer0.98
- Evidence1.00
- Context1.00
- Trigger1.00
- World'sFair1.00
- ANATOMY· 030.93
- Trigger1.00
- periodic1.00
- event-driven1.00
- On a schedule0.98
- On every errored trace1.00
- Sweep recent traces and evals on any cadence nightly, hourly, or near1.00
- The moment a span fails, the loop kicks off.0.99
- real-time. Batch up what's worth fixing.0.99
- Work's Fair0.92
- TRACK 5· JULY 1, 20260.96
- Evals1.00
-
- AlEngineer0.97
- ASSEMBLED1.00
- World'sFair1.00
- The loop, assembled0.97
- Evidence1.00
- Agent investigates0.98
- Pull request1.00
- You review1.00
- &merge1.00
- deploy → the next trace feeds the next pass1.00
- World's Fair0.96
- TRACK 5· JULY 1, 20260.95
- Evals1.00
Transcript
228 cues· 3,262 words· 16,875 chars
- 0:12 Well, thank you all.
- 0:15 Let me just get set up here.
- 0:16 So not just the founder of Verizon, but I tend to build an incredible amount of stuff.
- 0:25 Let's see if we get this going here.
- 0:29 Oh, sorry, one more second.
- 0:34 So not just a founder here, but also a builder.
- 0:38 And I do my best to,
- 0:42 to build agents, assistants.
- 0:45 We have an agent in product.
- 0:48 We have an agent in product called Alex, and I think a lot of my experience has come from actually trying to make the stuff work and work well.
- 0:59 Our first version of our own agent frankly sucked.
- 1:03 It was many years ago, probably two years ago, we weren't the first in the space to do it.
- 1:09 And a lot of what we have built has come out of our own experience in building this agent.
- 1:16 And Signal is kind of our next generation of this, which is trying to automate a bunch of things, which we do every day, and build it into products that people can use.
- 1:26 So I'm going to go through this materials here.
- 1:29 I'll try to go fast and try to show you a lot of product, too.
- 1:32 I'm a product person.
- 1:36 if you built a startup before, you've experienced this, your platform's down, it's late at night and you wanna go fix it.
- 1:45 And really, it takes a lot of energy to go do that.
- 1:49 And we're going to talk about the automation we built a little bit and what the future looks like.
- 1:54 And I truly believe that the future of the observability space is actually changing massively right now.
- 2:02 Why is that?
- 2:03 Well, observability used to be for humans.
- 2:06 It used to be a UI you click, a graph you click, something you look at.
- 2:11 And today, I would argue it's a lot of 2.0, which is like this combination of coding agent.
- 2:17 So those of you who've built skills, skills for PyroScope, Google Cloud, or whatnot, these skills help you with your human debugging these systems.
- 2:27 And really, telemetry is like this smoke thrown off of your system that can allow these agents to go make fixes.
- 2:38 It tells you what path in the code it took.
- 2:41 Without that, you're guessing, and there's a million paths it could have taken.
- 2:44 The data thrown off by your system allows you to go use agents to go debug your software.
- 2:55 Evals add another layer to this, but really what we're at here is how do I build systems that autonomously fix themselves?
- 3:05 Really, that is what we're after, both AI agents.
- 3:08 I put AI into my system.
- 3:10 How do I have this thing just improve itself?
- 3:13 And today, we're kind of in the 2.0, which is a human making fixes and reviewing things, but there's a future we're all driving towards and throwing off traces, throwing off logs, throwing off way more than you normally would, and having agents run at this for a continuous loop is where we're going.
- 3:30 You can build at agent speed, but today,
- 3:33 you can't improve your systems really at this agent speed.
- 3:36 So those of us feel this kind of governor happening within our products.
- 3:41 And the bottleneck is actually not the fix anymore.
- 3:46 So those of us who've used these systems and used coding agents with skills, the bottlenecks, a lot of the confidence in, do I have it right?
- 3:57 A lot of this is about, is this fix the right one to push?
- 4:03 And so these are kind of the challenges here.
- 4:06 And then how do you build this loop in a way that just moves faster?
- 4:10 And a little bit of the way we've kind of come to do it, and we do it in our system, is we've kind of inverted this loop, which is like a human looks at things and an agent,
- 4:23 fixes it to a person now can wake up with an idea of the issues based upon the errors occurred in their system.
- 4:31 So the agent is actually, you know, maybe it's not a fix itself, but it's putting up an issue.
- 4:36 It's looking at the data before a human even looks at it.
- 4:40 And what you move from there is kind of humans grabbing tickets to having some amount of evidence
- 4:48 some deep evidence relative to whatever you're looking at already sitting in front of you by the time you actually even look at it.
- 4:55 And human review is kind of one thing, but a lot of times maybe you're driving this little investigation a bit from where it started.
loading
Chapters
- 0:00 Arize, its agent Alyx, and why v1 sucked
- 1:36 Observability is changing: from dashboards to telemetry for agents
- 2:55 The goal: systems that fix themselves
- 4:14 Inverting the loop: the agent investigates first
- 6:08 Traces on a filesystem, the key unlock
- 7:23 From your laptop to sandboxes
- 8:13 A real fix: the Alyx stream canceled bug
- 9:39 Why you should trace ten times more
- 11:10 Product demo: Signal, AX, and Phoenix
- 13:09 Sandboxes, VPC, and why customers won't call out
- 16:20 Q&A: why not just point Claude Code at your data?
- 18:04 Q&A: where do the evals come in?