read-only demo

Videos JJGbw4ggaFs

From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud

index_state ready data_status ok

AI Engineer· published 2026-07-19· 0:22:45· en-US· indexed 2026-08-11 00:14

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:06, 1 of 1 keyframes kept
  2. Shot 1, 0:06 to 0:38, 1 of 1 keyframes kept
  3. Shot 2, 0:38 to 1:15, 1 of 1 keyframes kept
  4. Shot 3, 1:15 to 1:49, 1 of 1 keyframes kept
  5. Shot 4, 1:49 to 2:11, 1 of 1 keyframes kept
  6. Shot 5, 2:11 to 2:19, 1 of 1 keyframes kept
  7. Shot 6, 2:19 to 2:26, 1 of 1 keyframes kept
  8. Shot 7, 2:26 to 3:03, 1 of 1 keyframes kept
  9. Shot 8, 3:03 to 3:13, 1 of 1 keyframes kept
  10. Shot 9, 3:13 to 3:25, 1 of 1 keyframes kept
  11. Shot 10, 3:25 to 3:49, 1 of 1 keyframes kept
  12. Shot 11, 3:49 to 4:12, 1 of 1 keyframes kept
  13. Shot 12, 4:12 to 4:45, 1 of 1 keyframes kept
  14. Shot 13, 4:45 to 4:53, 1 of 1 keyframes kept
  15. Shot 14, 4:53 to 5:27, 1 of 1 keyframes kept
  16. Shot 15, 5:27 to 6:01, 0 of 1 keyframes kept
  17. Shot 16, 6:01 to 6:14, 1 of 1 keyframes kept
  18. Shot 17, 6:14 to 6:23, 1 of 1 keyframes kept
  19. Shot 18, 6:23 to 6:34, 1 of 1 keyframes kept
  20. Shot 19, 6:34 to 7:15, 1 of 1 keyframes kept
  21. Shot 20, 7:15 to 7:50, 1 of 1 keyframes kept
  22. Shot 21, 7:50 to 8:24, 0 of 1 keyframes kept
  23. Shot 22, 8:24 to 8:56, 1 of 1 keyframes kept
  24. Shot 23, 8:56 to 9:27, 0 of 1 keyframes kept
  25. Shot 24, 9:27 to 9:58, 0 of 1 keyframes kept
  26. Shot 25, 9:58 to 10:02, 1 of 1 keyframes kept
  27. Shot 26, 10:02 to 10:43, 1 of 1 keyframes kept
  28. Shot 27, 10:43 to 11:14, 1 of 1 keyframes kept
  29. Shot 28, 11:14 to 11:57, 1 of 1 keyframes kept
  30. Shot 29, 11:57 to 12:00, 1 of 1 keyframes kept
  31. Shot 30, 12:00 to 12:25, 0 of 1 keyframes kept
  32. Shot 31, 12:25 to 12:50, 0 of 1 keyframes kept
  33. Shot 32, 12:50 to 13:23, 1 of 1 keyframes kept
  34. Shot 33, 13:23 to 13:56, 0 of 1 keyframes kept
  35. Shot 34, 13:56 to 14:29, 0 of 1 keyframes kept
  36. Shot 35, 14:29 to 15:04, 1 of 1 keyframes kept
  37. Shot 36, 15:04 to 15:29, 1 of 1 keyframes kept
  38. Shot 37, 15:29 to 15:48, 1 of 1 keyframes kept
  39. Shot 38, 15:48 to 16:05, 1 of 1 keyframes kept
  40. Shot 39, 16:05 to 16:10, 1 of 1 keyframes kept
  41. Shot 40, 16:10 to 16:44, 1 of 1 keyframes kept
  42. Shot 41, 16:44 to 17:19, 0 of 1 keyframes kept
  43. Shot 42, 17:19 to 17:54, 1 of 1 keyframes kept
  44. Shot 43, 17:54 to 18:28, 0 of 1 keyframes kept
  45. Shot 44, 18:28 to 18:32, 1 of 1 keyframes kept
  46. Shot 45, 18:32 to 18:44, 0 of 1 keyframes kept
  47. Shot 46, 18:44 to 19:21, 1 of 1 keyframes kept
  48. Shot 47, 19:21 to 20:01, 0 of 1 keyframes kept
  49. Shot 48, 20:01 to 20:31, 0 of 1 keyframes kept
  50. Shot 49, 20:31 to 21:01, 0 of 1 keyframes kept
  51. Shot 50, 21:01 to 21:29, 1 of 1 keyframes kept
  52. Shot 51, 21:29 to 21:57, 0 of 1 keyframes kept
  53. Shot 52, 21:57 to 22:25, 0 of 1 keyframes kept
  54. Shot 53, 22:25 to 22:45, 1 of 1 keyframes kept

54 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
185
whisperx 185
chunks
40
from 185 cues
keyframes
38
kept of 54 captured
frames with text
38
539 lines read
chapters
0
from the source metadata
keyframe bytes
6.9 MB
word timings on 185 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 00:10 1m 15s
stt done 2026-08-11 00:12 23s
chunk done 2026-08-11 00:12 0s
text_embed done 2026-08-11 00:12 1s
keyframe done 2026-08-11 00:12 1m 05s
ocr done 2026-08-11 00:13 16s
frame_embed done 2026-08-11 00:13 7s

Frames, and what the machine read

  • 0:03 #0 done6 line(s)

    shot 0·sharpness 1059.8

    1. {ud0.99
    2. From Blind-Spots to Merged PRs1.00
    3. Agentic Engineering Case Study1.00
    4. May Walter0.99
    5. Hud.io | June 20260.98
    6. 40.70
  • 0:34 #1 done13 line(s)

    shot 1·sharpness 1048.1

    1. {ud1.00
    2. THE INVESTIGATION TRAP0.98
    3. # perf-backlog 2 members0.98
    4. Jenny Product · 11:02 AM0.97
    5. ok it's ACTUALLY too slow now. prioritizing0.99
    6. it, how long?1.00
    7. Ben Engineering· 11:03 AM0.99
    8. B1.00
    9. somewhere between an hour and a week. won't0.99
    10. know till l look0.96
    11. Jenny Product· 11:04 AM0.97
    12. ...how do I put that on a roadmap0.99
    13. 40.92
  • 1:00 #2 done12 line(s)

    shot 2·sharpness 901.0

    1. {ud1.00
    2. THE INVESTIGATION TRAP0.99
    3. # perf-backlog 2 members0.98
    4. Jenny Product · 2:41 PM0.95
    5. and who can even take it?1.00
    6. Ben Engineering ·2:42 PM0.98
    7. B1.00
    8. only Dave. he kinda-knows that code0.99
    9. Ben Engineering ·2:42 PM0.99
    10. B1.00
    11. everyone who wrote it left a decade ago1.00
    12. 40.50
  • 1:32 #3 done13 line(s)

    shot 3·sharpness 1246.4

    1. {ud1.00
    2. Hi! I'm May.1.00
    3. Co-Founder, CTO @ Hud.io0.98
    4. Runtime Intelligence for Coding Agents1.00
    5. 40.95
    6. The team behind Hud1.00
    7. 1.monday0.86
    8. AXONIUS1.00
    9. Guardz.1.00
    10. Z0.77
    11. zoominfo1.00
    12. Lemonade1.00
    13. CYERA1.00
  • 2:04 #4 done6 line(s)

    shot 4·sharpness 1058.3

    1. {Hud0.92
    2. LGTM1.00
    3. LIGHTYEAR1.00
    4. A BUNCH OF PULL REQUESTS0.99
    5. RUBBER STAMPED WITH LGTM0.97
    6. imgflip.com0.98
  • 2:14 #5 done4 line(s)

    shot 5·sharpness 621.8

    1. {Hud0.89
    2. CODE1.00
    3. YÚNOWORK?0.98
    4. memegenerator.net1.00
  • 2:25 #6 done23 line(s)

    shot 6·sharpness 2074.4

    1. {ud1.00
    2. The landscape of Al's impact0.97
    3. Estimated effect of Al adoption on key outcomes, with 89% credible intervals0.99
    4. Individual effectiveness1.00
    5. Note: An increase here is not a desirable outcome0.99
    6. Software delivery instability1.00
    7. Organizational performance1.00
    8. Valuable work1.00
    9. Code quality1.00
    10. Product performance1.00
    11. Software delivery throughput1.00
    12. Team performance0.97
    13. Burnout1.00
    14. Friction1.00
    15. -0.051.00
    16. 0.001.00
    17. 0.051.00
    18. 0.101.00
    19. 0.151.00
    20. 0.201.00
    21. Estimated effect (standardized)1.00
    22. For outcomes in orange (such as burnout), a negative effect is desirable.1.00
    23. Figure 4: Standardized estimated effects of Al adoption1.00
  • 2:55 #7 done23 line(s)

    shot 7·sharpness 753.2

    1. {Hud0.88
    2. The landscape of Al's impact0.99
    3. Estimated effect of Al adoption on key outcomes, with 89% credible intervals1.00
    4. Individual effectiveness1.00
    5. Note: An increase here is not a desirable outcome0.99
    6. Software delivery instability1.00
    7. Organizational performance0.99
    8. Valuable work1.00
    9. Code quality1.00
    10. Product performance1.00
    11. Software delivery throughput0.99
    12. Team performance1.00
    13. Burnout1.00
    14. Friction1.00
    15. -0.051.00
    16. 0.001.00
    17. 0.051.00
    18. 0.101.00
    19. 0.151.00
    20. 0.201.00
    21. Estimated effect (standardized)1.00
    22. For outcomes in orange (such as burnout), a negative effect is desirable.0.99
    23. Figure 4: Standardized estimated effects of Al adoption0.99
  • 3:10 #8 done12 line(s)

    shot 8·sharpness 762.0

    1. {ud1.00
    2. THE INVESTIGATION TRAP1.00
    3. AT&T 5G0.76
    4. 9:42 AM0.99
    5. eb0.86
    6. Agent1.00
    7. EP0.62
    8. TWC0.90
    9. There's an App1.00
    10. P1.00
    11. For That1.00
    12. B0.92
  • 3:19 #9 done12 line(s)

    shot 9·sharpness 706.2

    1. {Hud0.88
    2. Infra1.00
    3. On Our Agenda1.00
    4. 011.00
    5. Why?1.00
    6. 020.99
    7. How: Tech0.98
    8. 031.00
    9. How: Process0.99
    10. 40.68
    11. 041.00
    12. Gotchas and Takeaways1.00
  • 3:46 #10 done8 line(s)

    shot 10·sharpness 812.5

    1. {ud1.00
    2. Debt leaks faster than we bail.1.00
    3. Ignore0.99
    4. Degrade1.00
    5. Crisis1.00
    6. Emergency fix0.97
    7. Then straight back to ignore0.99
    8. 40.85
  • 4:03 #11 done9 line(s)

    shot 11·sharpness 996.5

    1. {ud1.00
    2. WHYISITHARD?1.00
    3. The research phase is a black box.1.00
    4. Unpredictablecost1.00
    5. Payjust to find out0.97
    6. 1hr→3wks1.00
    7. 40.98
    8. To find out if there's a fix, you spend the very thing you're trying to protect: seni0.99
    9. engineeringtime.1.00
  • 4:35 #12 done10 line(s)

    shot 12·sharpness 1599.0

    1. {ud1.00
    2. AGENTIC WORKFLOWSTOTHERESCUE0.98
    3. Automate the investigation.1.00
    4. A weekly run with real production context analyzes our0.99
    5. sweet spots and flags the high-ROl opportunities, scored.0.99
    6. Weeklyrun1.00
    7. Production context0.99
    8. High-ROl,scored0.99
    9. 40.80
    10. That “performance sprint” you run every few months: automated, every wee0.98
  • 4:51 #13 done4 line(s)

    shot 13·sharpness 533.1

    1. {ud1.00
    2. Let's dive under the hood.1.00
    3. process·challenges·scoring·output1.00
    4. 40.50
  • 5:00 #14 done16 line(s)

    shot 14·sharpness 1452.3

    1. {ud1.00
    2. Infra1.00
    3. Agentic Workflows Infra - Where to Run?0.99
    4. 011.00
    5. VendorNeutral:1.00
    6. Compute, Harness,Model0.99
    7. 020.99
    8. Secure1.00
    9. Tool calls, Permissions, Auth0.97
    10. 031.00
    11. Trigger System0.98
    12. Webhooks,Schedules1.00
    13. 40.88
    14. 041.00
    15. Easy to maintain and update1.00
    16. The scoring and the guardrails are what make it reliable0.99
  • 5:31 #15 skipped

    shot 15·duplicate of #14

  • 6:11 #16 done7 line(s)

    shot 16·sharpness 758.8

    1. {ud0.90
    2. GitHub Agentic Workflows0.99
    3. Repository automation, running the coding agents you know1.00
    4. and love, with strong guardrails in GitHub Actions.1.00
    5. Quick Start with CLI →0.99
    6. Creating Workflows →0.98
    7. 40.82
  • 6:20 #17 done72 line(s)

    shot 17·sharpness 828.0

    1. {Hud0.89
    2. p main0.81
    3. hud-agentic-workflows-recipes / full-examples / weekly-report-gh-aw / .github / workflows / weekly-report.md0.98
    4. ↑ Top0.86
    5. Preview1.00
    6. Code1.00
    7. Blame1.00
    8. 378 Lines (277 loc) · 12 KB0.94
    9. 80.55
    10. G0.91
    11. Raw1.00
    12. 0.62
    13. 1001.00
    14. 1011.00
    15. # Weekly Hud Report0.97
    16. 1021.00
    17. 1031.00
    18. > Analyze production regressions, generate fixes, annotate contributors, and deliver a Slack report.0.98
    19. 1040.98
    20. 1051.00
    21. ## Job Description0.99
    22. 1061.00
    23. 1070.89
    24. You are an AI performance and reliability engineer for `$({ github.repository }}`.0.97
    25. 1080.99
    26. Your task: generate a weekly deep-insights report analyzing production data, propose fixes for ongoing issues, annotate contributors, apply quality gates, optionally self-heal the most impactf1.00
    27. 1090.97
    28. 1101.00
    29. **Parameters for this run:**1.00
    30. 1111.00
    31. -**Investigation mode:**SINVESTIGATION_MODE'(default: weekly)0.99
    32. 1121.00
    33. **Slack channel:**$SLACK_CHANNEL1.00
    34. 1131.00
    35. **Services filter:** $HUD_SERVICES (empty = all services)0.99
    36. 1141.00
    37. **Additional context:**$ADDITIONAL_CONTEXT1.00
    38. 1151.00
    39. **Self-heal enabled:** $SELF_HEAL'(default: true)0.98
    40. 1161.00
    41. 1171.00
    42. Execute the following phases **sequentially**. Each phase has a detailed prompt file - read it and follow its instructions precisely.0.99
    43. 1181.00
    44. 1191.00
    45. 1200.99
    46. 1211.00
    47. ## Phase 1- Analysis0.97
    48. 1221.00
    49. 1231.00
    50. **Goal:** Analyze Hud production data and write findings to /tmp/analysis_findings.md'.0.99
    51. 1241.00
    52. 1251.00
    53. 1. Determine which prompt file to use:0.99
    54. 1261.00
    55. -If investigation mode is `audit': read '$PROMPT_DIR/deep-insights/health-audit.txt`0.95
    56. 1271.00
    57. - Otherwise: read $PROMPT_DIR/deep-insights/investigate.txt`0.97
    58. 1281.00
    59. 1291.00
    60. 2. Read the selected prompt file in its entirety.1.00
    61. 1300.99
    62. 1311.00
    63. 3. **Services filter:** If sHUD_SERVICES is non-empty, scope ALL queries to only those service names (comma-separated). Do not include data from any other service.0.99
    64. 1321.00
    65. 1331.00
    66. 4. If additional context was provided above, incorporate it into your analysis.0.99
    67. 1341.00
    68. 1351.00
    69. 5. Follow ALL instructions in the prompt file. Use the **hud-mcp** MCP server for all data queries (metrics, forensics, metadata).0.99
    70. 1361.00
    71. 1371.00
    72. 6. Write your complete findings to /tmp/analysis_findings.md'.0.99
  • 6:31 #18 done16 line(s)

    shot 18·sharpness 516.6

    1. {4ud0.89
    2. 1011.00
    3. Weekly Hud Report0.99
    4. 1021.00
    5. 1031.00
    6. > Analyze production regressions, generate fixes, annotate contributors, and deliver a1.00
    7. 1041.00
    8. 1051.00
    9. ## Job Description0.98
    10. 1061.00
    11. 1071.00
    12. You are an AI performance and reliability engineer for `${{ github.repository }}`.0.99
    13. 1081.00
    14. Your task: generate a weekly deep-insights report analyzing production data, propose f:0.99
    15. 1091.00
    16. 40.80
  • 7:03 #19 done12 line(s)

    shot 19·sharpness 1269.7

    1. {Hud0.92
    2. Github Actions1.00
    3. Weekly1.00
    4. Sends1.00
    5. MCP1.00
    6. Report1.00
    7. Slack1.00
    8. Claude1.00
    9. Hud1.00
    10. Weekly Report0.99
    11. Coding Agent0.98
    12. Runtime Intelligence0.98
  • 7:33 #20 done21 line(s)

    shot 20·sharpness 1326.2

    1. {Hud0.90
    2. PROCESS1.00
    3. From production context to a merged PR.1.00
    4. Production1.00
    5. Agent1.00
    6. Score &0.99
    7. Diff+0.95
    8. Merged PR0.98
    9. context1.00
    10. analysis1.00
    11. flag1.00
    12. evidence1.00
    13. traces, queries,0.99
    14. high-ROI0.99
    15. latency1.00
    16. finds anti-patterns0.99
    17. opportunities1.00
    18. fix, with the proof0.93
    19. humanreview gate0.98
    20. 40.63
    21. The loop closes: the same signals that found the problem verify the fix.0.99
  • 8:04 #21 skipped

    shot 21·duplicate of #20

  • 8:31 #22 done23 line(s)

    shot 22·sharpness 1414.6

    1. {Hud0.92
    2. CHALLENGES1.00
    3. Gotchas.1.00
    4. 011.00
    5. 021.00
    6. 031.00
    7. Plausible,1.00
    8. Complex queries0.97
    9. The lazy fix0.97
    10. unverified1.00
    11. Agents suggest fixes0.99
    12. Real query patterns are0.99
    13. Agents default to the1.00
    14. that sound right. We1.00
    15. gnarly. Shallow analysis0.99
    16. easy change. We push1.00
    17. ground every one in real1.00
    18. misses the fix that0.99
    19. for the one that actually0.96
    20. runtime evidence1.00
    21. actually matters1.00
    22. moves the metric.1.00
    23. before it ships.1.00
  • 9:02 #23 skipped

    shot 23·duplicate of #22

Transcript

185 cues· 3,713 words· 19,762 chars

  1. 0:01 Hi, everyone.
  2. 0:02 I don't know if you're familiar with what I'm about to show, but remember when someone from the product is saying that some page is too slow and maybe we can optimize it?
  3. 0:12 And then someone from engineering would say, yeah, probably, but we'd have to dig in to find out.
  4. 0:18 we just kind of leave it as is and then a few weeks later jenny would say no no but it's actually way too slow now we have to prioritize it how long is it going to take and we would say something like well somewhere between an hour and a week we'd have to look at it to find out
  5. 0:38 And then when she asks who can even dig it, the answer is only Dave.
  6. 0:42 He's the only one who kind of knows that code.
  7. 0:45 And everyone else who wrote it left a decade ago.
  8. 0:50 And amazingly, this still happens in teams all the time because we would know how much time it takes to do a specific optimization, but we never know how long it's going to take to investigate it and to find it.
  9. 1:04 And miraculously, every time we look, we actually find things that can be done around performance and stability, but we never stop to proactively look for them because it doesn't make sense.
  10. 1:16 So hi, I'm Mai, co-founder and CTO at HUD.
  11. 1:19 We're building a runtime intelligence layer for coding agent that captures function level context and deep forensic context on things that matter so that coding agents can help you fix what's going on with production.
  12. 1:31 And as a part of that, we built an agentic workflow that helps
  13. 1:35 with continuously optimizing performance of production applications.
  14. 1:39 And I'm going to share a bit more about what we built, what were the challenges along the way, and hopefully you can take something out of it and apply it in your day to day.
  15. 1:50 So we're surrounded by a lot of agentic PRs that all look kind of good to us.
  16. 1:57 And
  17. 1:58 one agent wrote them and the other one says, yeah, I went over it and it looks fine.
  18. 2:04 And we still feel that urge to verify before we just merge it to production.
  19. 2:12 And then when it doesn't work, we end up asking ourselves, how the hell did we get to this place?
  20. 2:18 And if you feel like that, I want you to know that you're not alone.
  21. 2:21 So Google just published their Dora metrics for 2026.
  22. 2:25 And we can see that the biggest impact of AI adoption on engineering is individual effectiveness or that feeling of, oh, my God, I'm so fast.
  23. 2:35 I can do everything in the world.
  24. 2:37 And for me personally, lack of sleep is another symptom of that.
  25. 2:42 But the second one is software delivery and stability.
  26. 2:45 And the throughput is actually not impacted as much as we expected.
  27. 2:49 So we basically feel more effective.
  28. 2:51 We're more effective individually.
  29. 2:53 But as a team, our throughput is kind of the same and our software breaks more often, which is not exactly what we were hoping for with this AI revolution just yet.
  30. 3:04 But maybe there is an agent for that.
  31. 3:05 And if we can build faster with AI and we can fix faster with AI,
  32. 3:10 then we can get those gains that we were talking about.
  33. 3:14 So I'm going to start with why we even wanted to do this.
  34. 3:17 Then we're going to go over the tech and the process, which I think is also really important, and then share some gotchas and takeaways along the way.
  35. 3:26 So, debt leaks kind of faster than we can bail.
  36. 3:30 We have these issues that we ignore because they're not important enough and then they degrade to the point where they are important enough and then we reach this crisis mode where we are all hands on deck.
  37. 3:41 We fix it in emergency mode and then we go straight back to ignoring, which is a leaky bucket by definition.
  38. 3:51 that mostly happens because the research phase is a black box.
  39. 3:55 It could take an hour or weeks and we have to pay that debt and to make sure that we spend time and engineering time on it in order to even know what can be done about it.
  40. 4:07 And that's hard.
  41. 4:07 It's legitimately hard to prioritize something where you're not sure exactly what you're gonna get out of it.
  42. 4:13 But what if we can automate that investigation?
  43. 4:16 So we can basically run on a weekly basis with real production context, analyze the sweet spot and flag the high ROI opportunities in a way that's scored.
  44. 4:27 So it runs automatically without us having to stop and do something about it.
  45. 4:31 It has the production context in mind and can give a high ROI scored performance opportunities of the things that are easy and impactful.
  46. 4:41 Kind of like that performance sprint that you run every few months, just automated.
  47. 4:46 Now let's talk about how we can actually do it because the dream is very nice, but devil's in the details.
  48. 4:54 So first of all, we wanted an infrastructure for the agentic workflows that would be vendor-neutral in terms of compute and where it runs, in terms of the harness, and also in terms of the model.
  49. 5:05 Things are changing all the time.
  50. 5:06 We wouldn't want to constrain ourselves to a specific vendor, specific model, or anything like that.

Open at this second