Videos JJGbw4ggaFs
From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
Scene timeline
54 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 185
- whisperx 185
- chunks
- 40
- from 185 cues
- keyframes
- 38
- kept of 54 captured
- frames with text
- 38
- 539 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.9 MB
- word timings on 185 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:10 | 1m 15s |
stt |
done | — | 2026-08-11 00:12 | 23s |
chunk |
done | — | 2026-08-11 00:12 | 0s |
text_embed |
done | — | 2026-08-11 00:12 | 1s |
keyframe |
done | — | 2026-08-11 00:12 | 1m 05s |
ocr |
done | — | 2026-08-11 00:13 | 16s |
frame_embed |
done | — | 2026-08-11 00:13 | 7s |
Frames, and what the machine read
-
- {ud0.99
- From Blind-Spots to Merged PRs1.00
- Agentic Engineering Case Study1.00
- May Walter0.99
- Hud.io | June 20260.98
- 40.70
-
- {ud1.00
- THE INVESTIGATION TRAP0.98
- # perf-backlog 2 members0.98
- Jenny Product · 11:02 AM0.97
- ok it's ACTUALLY too slow now. prioritizing0.99
- it, how long?1.00
- Ben Engineering· 11:03 AM0.99
- B1.00
- somewhere between an hour and a week. won't0.99
- know till l look0.96
- Jenny Product· 11:04 AM0.97
- ...how do I put that on a roadmap0.99
- 40.92
-
- {ud1.00
- THE INVESTIGATION TRAP0.99
- # perf-backlog 2 members0.98
- Jenny Product · 2:41 PM0.95
- and who can even take it?1.00
- Ben Engineering ·2:42 PM0.98
- B1.00
- only Dave. he kinda-knows that code0.99
- Ben Engineering ·2:42 PM0.99
- B1.00
- everyone who wrote it left a decade ago1.00
- 40.50
-
- {ud1.00
- Hi! I'm May.1.00
- Co-Founder, CTO @ Hud.io0.98
- Runtime Intelligence for Coding Agents1.00
- 40.95
- The team behind Hud1.00
- 1.monday0.86
- AXONIUS1.00
- Guardz.1.00
- Z0.77
- zoominfo1.00
- Lemonade1.00
- CYERA1.00
-
- {Hud0.92
- LGTM1.00
- LIGHTYEAR1.00
- A BUNCH OF PULL REQUESTS0.99
- RUBBER STAMPED WITH LGTM0.97
- imgflip.com0.98
-
- {Hud0.89
- CODE1.00
- YÚNOWORK?0.98
- memegenerator.net1.00
-
- {ud1.00
- The landscape of Al's impact0.97
- Estimated effect of Al adoption on key outcomes, with 89% credible intervals0.99
- Individual effectiveness1.00
- Note: An increase here is not a desirable outcome0.99
- Software delivery instability1.00
- Organizational performance1.00
- Valuable work1.00
- Code quality1.00
- Product performance1.00
- Software delivery throughput1.00
- Team performance0.97
- Burnout1.00
- Friction1.00
- -0.051.00
- 0.001.00
- 0.051.00
- 0.101.00
- 0.151.00
- 0.201.00
- Estimated effect (standardized)1.00
- For outcomes in orange (such as burnout), a negative effect is desirable.1.00
- Figure 4: Standardized estimated effects of Al adoption1.00
-
- {Hud0.88
- The landscape of Al's impact0.99
- Estimated effect of Al adoption on key outcomes, with 89% credible intervals1.00
- Individual effectiveness1.00
- Note: An increase here is not a desirable outcome0.99
- Software delivery instability1.00
- Organizational performance0.99
- Valuable work1.00
- Code quality1.00
- Product performance1.00
- Software delivery throughput0.99
- Team performance1.00
- Burnout1.00
- Friction1.00
- -0.051.00
- 0.001.00
- 0.051.00
- 0.101.00
- 0.151.00
- 0.201.00
- Estimated effect (standardized)1.00
- For outcomes in orange (such as burnout), a negative effect is desirable.0.99
- Figure 4: Standardized estimated effects of Al adoption0.99
-
- {ud1.00
- THE INVESTIGATION TRAP1.00
- AT&T 5G0.76
- 9:42 AM0.99
- eb0.86
- Agent1.00
- EP0.62
- TWC0.90
- There's an App1.00
- P1.00
- For That1.00
- B0.92
-
- {Hud0.88
- Infra1.00
- On Our Agenda1.00
- 011.00
- Why?1.00
- 020.99
- How: Tech0.98
- 031.00
- How: Process0.99
- 40.68
- 041.00
- Gotchas and Takeaways1.00
-
- {ud1.00
- Debt leaks faster than we bail.1.00
- Ignore0.99
- Degrade1.00
- Crisis1.00
- Emergency fix0.97
- Then straight back to ignore0.99
- 40.85
-
- {ud1.00
- WHYISITHARD?1.00
- The research phase is a black box.1.00
- Unpredictablecost1.00
- Payjust to find out0.97
- 1hr→3wks1.00
- 40.98
- To find out if there's a fix, you spend the very thing you're trying to protect: seni0.99
- engineeringtime.1.00
-
- {ud1.00
- AGENTIC WORKFLOWSTOTHERESCUE0.98
- Automate the investigation.1.00
- A weekly run with real production context analyzes our0.99
- sweet spots and flags the high-ROl opportunities, scored.0.99
- Weeklyrun1.00
- Production context0.99
- High-ROl,scored0.99
- 40.80
- That “performance sprint” you run every few months: automated, every wee0.98
-
- {ud1.00
- Let's dive under the hood.1.00
- process·challenges·scoring·output1.00
- 40.50
-
- {ud1.00
- Infra1.00
- Agentic Workflows Infra - Where to Run?0.99
- 011.00
- VendorNeutral:1.00
- Compute, Harness,Model0.99
- 020.99
- Secure1.00
- Tool calls, Permissions, Auth0.97
- 031.00
- Trigger System0.98
- Webhooks,Schedules1.00
- 40.88
- 041.00
- Easy to maintain and update1.00
- The scoring and the guardrails are what make it reliable0.99
-
- {ud0.90
- GitHub Agentic Workflows0.99
- Repository automation, running the coding agents you know1.00
- and love, with strong guardrails in GitHub Actions.1.00
- Quick Start with CLI →0.99
- Creating Workflows →0.98
- 40.82
-
- {Hud0.89
- p main0.81
- hud-agentic-workflows-recipes / full-examples / weekly-report-gh-aw / .github / workflows / weekly-report.md0.98
- ↑ Top0.86
- Preview1.00
- Code1.00
- Blame1.00
- 378 Lines (277 loc) · 12 KB0.94
- 80.55
- G0.91
- Raw1.00
- 回0.62
- 1001.00
- 1011.00
- # Weekly Hud Report0.97
- 1021.00
- 1031.00
- > Analyze production regressions, generate fixes, annotate contributors, and deliver a Slack report.0.98
- 1040.98
- 1051.00
- ## Job Description0.99
- 1061.00
- 1070.89
- You are an AI performance and reliability engineer for `$({ github.repository }}`.0.97
- 1080.99
- Your task: generate a weekly deep-insights report analyzing production data, propose fixes for ongoing issues, annotate contributors, apply quality gates, optionally self-heal the most impactf1.00
- 1090.97
- 1101.00
- **Parameters for this run:**1.00
- 1111.00
- -**Investigation mode:**SINVESTIGATION_MODE'(default: weekly)0.99
- 1121.00
- **Slack channel:**$SLACK_CHANNEL1.00
- 1131.00
- **Services filter:** $HUD_SERVICES (empty = all services)0.99
- 1141.00
- **Additional context:**$ADDITIONAL_CONTEXT1.00
- 1151.00
- **Self-heal enabled:** $SELF_HEAL'(default: true)0.98
- 1161.00
- 1171.00
- Execute the following phases **sequentially**. Each phase has a detailed prompt file - read it and follow its instructions precisely.0.99
- 1181.00
- 1191.00
- 1200.99
- 1211.00
- ## Phase 1- Analysis0.97
- 1221.00
- 1231.00
- **Goal:** Analyze Hud production data and write findings to /tmp/analysis_findings.md'.0.99
- 1241.00
- 1251.00
- 1. Determine which prompt file to use:0.99
- 1261.00
- -If investigation mode is `audit': read '$PROMPT_DIR/deep-insights/health-audit.txt`0.95
- 1271.00
- - Otherwise: read $PROMPT_DIR/deep-insights/investigate.txt`0.97
- 1281.00
- 1291.00
- 2. Read the selected prompt file in its entirety.1.00
- 1300.99
- 1311.00
- 3. **Services filter:** If sHUD_SERVICES is non-empty, scope ALL queries to only those service names (comma-separated). Do not include data from any other service.0.99
- 1321.00
- 1331.00
- 4. If additional context was provided above, incorporate it into your analysis.0.99
- 1341.00
- 1351.00
- 5. Follow ALL instructions in the prompt file. Use the **hud-mcp** MCP server for all data queries (metrics, forensics, metadata).0.99
- 1361.00
- 1371.00
- 6. Write your complete findings to /tmp/analysis_findings.md'.0.99
-
- {4ud0.89
- 1011.00
- Weekly Hud Report0.99
- 1021.00
- 1031.00
- > Analyze production regressions, generate fixes, annotate contributors, and deliver a1.00
- 1041.00
- 1051.00
- ## Job Description0.98
- 1061.00
- 1071.00
- You are an AI performance and reliability engineer for `${{ github.repository }}`.0.99
- 1081.00
- Your task: generate a weekly deep-insights report analyzing production data, propose f:0.99
- 1091.00
- 40.80
-
- {Hud0.92
- Github Actions1.00
- Weekly1.00
- Sends1.00
- MCP1.00
- Report1.00
- Slack1.00
- Claude1.00
- Hud1.00
- Weekly Report0.99
- Coding Agent0.98
- Runtime Intelligence0.98
-
- {Hud0.90
- PROCESS1.00
- From production context to a merged PR.1.00
- Production1.00
- Agent1.00
- Score &0.99
- Diff+0.95
- Merged PR0.98
- context1.00
- analysis1.00
- flag1.00
- evidence1.00
- traces, queries,0.99
- high-ROI0.99
- latency1.00
- finds anti-patterns0.99
- opportunities1.00
- fix, with the proof0.93
- humanreview gate0.98
- 40.63
- The loop closes: the same signals that found the problem verify the fix.0.99
-
- {Hud0.92
- CHALLENGES1.00
- Gotchas.1.00
- 011.00
- 021.00
- 031.00
- Plausible,1.00
- Complex queries0.97
- The lazy fix0.97
- unverified1.00
- Agents suggest fixes0.99
- Real query patterns are0.99
- Agents default to the1.00
- that sound right. We1.00
- gnarly. Shallow analysis0.99
- easy change. We push1.00
- ground every one in real1.00
- misses the fix that0.99
- for the one that actually0.96
- runtime evidence1.00
- actually matters1.00
- moves the metric.1.00
- before it ships.1.00
Transcript
185 cues· 3,713 words· 19,762 chars
- 0:01 Hi, everyone.
- 0:02 I don't know if you're familiar with what I'm about to show, but remember when someone from the product is saying that some page is too slow and maybe we can optimize it?
- 0:12 And then someone from engineering would say, yeah, probably, but we'd have to dig in to find out.
- 0:18 we just kind of leave it as is and then a few weeks later jenny would say no no but it's actually way too slow now we have to prioritize it how long is it going to take and we would say something like well somewhere between an hour and a week we'd have to look at it to find out
- 0:38 And then when she asks who can even dig it, the answer is only Dave.
- 0:42 He's the only one who kind of knows that code.
- 0:45 And everyone else who wrote it left a decade ago.
- 0:50 And amazingly, this still happens in teams all the time because we would know how much time it takes to do a specific optimization, but we never know how long it's going to take to investigate it and to find it.
- 1:04 And miraculously, every time we look, we actually find things that can be done around performance and stability, but we never stop to proactively look for them because it doesn't make sense.
- 1:16 So hi, I'm Mai, co-founder and CTO at HUD.
- 1:19 We're building a runtime intelligence layer for coding agent that captures function level context and deep forensic context on things that matter so that coding agents can help you fix what's going on with production.
- 1:31 And as a part of that, we built an agentic workflow that helps
- 1:35 with continuously optimizing performance of production applications.
- 1:39 And I'm going to share a bit more about what we built, what were the challenges along the way, and hopefully you can take something out of it and apply it in your day to day.
- 1:50 So we're surrounded by a lot of agentic PRs that all look kind of good to us.
- 1:57 And
- 1:58 one agent wrote them and the other one says, yeah, I went over it and it looks fine.
- 2:04 And we still feel that urge to verify before we just merge it to production.
- 2:12 And then when it doesn't work, we end up asking ourselves, how the hell did we get to this place?
- 2:18 And if you feel like that, I want you to know that you're not alone.
- 2:21 So Google just published their Dora metrics for 2026.
- 2:25 And we can see that the biggest impact of AI adoption on engineering is individual effectiveness or that feeling of, oh, my God, I'm so fast.
- 2:35 I can do everything in the world.
- 2:37 And for me personally, lack of sleep is another symptom of that.
- 2:42 But the second one is software delivery and stability.
- 2:45 And the throughput is actually not impacted as much as we expected.
- 2:49 So we basically feel more effective.
- 2:51 We're more effective individually.
- 2:53 But as a team, our throughput is kind of the same and our software breaks more often, which is not exactly what we were hoping for with this AI revolution just yet.
- 3:04 But maybe there is an agent for that.
- 3:05 And if we can build faster with AI and we can fix faster with AI,
- 3:10 then we can get those gains that we were talking about.
- 3:14 So I'm going to start with why we even wanted to do this.
- 3:17 Then we're going to go over the tech and the process, which I think is also really important, and then share some gotchas and takeaways along the way.
- 3:26 So, debt leaks kind of faster than we can bail.
- 3:30 We have these issues that we ignore because they're not important enough and then they degrade to the point where they are important enough and then we reach this crisis mode where we are all hands on deck.
- 3:41 We fix it in emergency mode and then we go straight back to ignoring, which is a leaky bucket by definition.
- 3:51 that mostly happens because the research phase is a black box.
- 3:55 It could take an hour or weeks and we have to pay that debt and to make sure that we spend time and engineering time on it in order to even know what can be done about it.
- 4:07 And that's hard.
- 4:07 It's legitimately hard to prioritize something where you're not sure exactly what you're gonna get out of it.
- 4:13 But what if we can automate that investigation?
- 4:16 So we can basically run on a weekly basis with real production context, analyze the sweet spot and flag the high ROI opportunities in a way that's scored.
- 4:27 So it runs automatically without us having to stop and do something about it.
- 4:31 It has the production context in mind and can give a high ROI scored performance opportunities of the things that are easy and impactful.
- 4:41 Kind of like that performance sprint that you run every few months, just automated.
- 4:46 Now let's talk about how we can actually do it because the dream is very nice, but devil's in the details.
- 4:54 So first of all, we wanted an infrastructure for the agentic workflows that would be vendor-neutral in terms of compute and where it runs, in terms of the harness, and also in terms of the model.
- 5:05 Things are changing all the time.
- 5:06 We wouldn't want to constrain ourselves to a specific vendor, specific model, or anything like that.
loading