read-only demo

Videos CgsWxRUY5Eo

AI Agents for Performance: Ship Faster, Pay Less — Rajat Shah, Netflix

index_state ready data_status ok

AI Engineer· published 2026-07-28· 0:33:38· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:40, 1 of 1 keyframes kept
  2. Shot 1, 0:40 to 0:48, 1 of 1 keyframes kept
  3. Shot 2, 0:48 to 1:08, 1 of 1 keyframes kept
  4. Shot 3, 1:08 to 1:21, 1 of 1 keyframes kept
  5. Shot 4, 1:21 to 1:48, 1 of 1 keyframes kept
  6. Shot 5, 1:48 to 2:15, 0 of 1 keyframes kept
  7. Shot 6, 2:15 to 2:40, 1 of 1 keyframes kept
  8. Shot 7, 2:40 to 3:05, 0 of 1 keyframes kept
  9. Shot 8, 3:05 to 3:44, 1 of 1 keyframes kept
  10. Shot 9, 3:44 to 4:13, 1 of 1 keyframes kept
  11. Shot 10, 4:13 to 4:42, 0 of 1 keyframes kept
  12. Shot 11, 4:42 to 5:07, 1 of 1 keyframes kept
  13. Shot 12, 5:07 to 5:38, 1 of 1 keyframes kept
  14. Shot 13, 5:38 to 6:09, 0 of 1 keyframes kept
  15. Shot 14, 6:09 to 6:36, 1 of 1 keyframes kept
  16. Shot 15, 6:36 to 7:03, 1 of 1 keyframes kept
  17. Shot 16, 7:03 to 7:30, 0 of 1 keyframes kept
  18. Shot 17, 7:30 to 7:53, 1 of 1 keyframes kept
  19. Shot 18, 7:53 to 8:26, 0 of 1 keyframes kept
  20. Shot 19, 8:26 to 9:00, 1 of 1 keyframes kept
  21. Shot 20, 9:00 to 9:27, 1 of 1 keyframes kept
  22. Shot 21, 9:27 to 9:54, 1 of 1 keyframes kept
  23. Shot 22, 9:54 to 10:22, 1 of 1 keyframes kept
  24. Shot 23, 10:22 to 10:34, 0 of 1 keyframes kept
  25. Shot 24, 10:34 to 11:00, 1 of 1 keyframes kept
  26. Shot 25, 11:00 to 11:29, 1 of 1 keyframes kept
  27. Shot 26, 11:29 to 11:58, 0 of 1 keyframes kept
  28. Shot 27, 11:58 to 12:28, 0 of 1 keyframes kept
  29. Shot 28, 12:28 to 12:58, 1 of 1 keyframes kept
  30. Shot 29, 12:58 to 13:28, 0 of 1 keyframes kept
  31. Shot 30, 13:28 to 13:58, 0 of 1 keyframes kept
  32. Shot 31, 13:58 to 14:46, 1 of 1 keyframes kept
  33. Shot 32, 14:46 to 15:27, 1 of 1 keyframes kept
  34. Shot 33, 15:27 to 16:01, 1 of 1 keyframes kept
  35. Shot 34, 16:01 to 16:34, 0 of 1 keyframes kept
  36. Shot 35, 16:34 to 17:07, 1 of 1 keyframes kept
  37. Shot 36, 17:07 to 17:38, 1 of 1 keyframes kept
  38. Shot 37, 17:38 to 18:10, 0 of 1 keyframes kept
  39. Shot 38, 18:10 to 18:36, 1 of 1 keyframes kept
  40. Shot 39, 18:36 to 19:03, 1 of 1 keyframes kept
  41. Shot 40, 19:03 to 19:30, 0 of 1 keyframes kept
  42. Shot 41, 19:30 to 20:13, 1 of 1 keyframes kept
  43. Shot 42, 20:13 to 20:14, 1 of 1 keyframes kept
  44. Shot 43, 20:14 to 20:51, 1 of 1 keyframes kept
  45. Shot 44, 20:51 to 21:19, 1 of 1 keyframes kept
  46. Shot 45, 21:19 to 21:47, 1 of 1 keyframes kept
  47. Shot 46, 21:47 to 22:15, 1 of 1 keyframes kept
  48. Shot 47, 22:15 to 22:44, 0 of 1 keyframes kept
  49. Shot 48, 22:44 to 23:12, 1 of 1 keyframes kept
  50. Shot 49, 23:12 to 23:40, 1 of 1 keyframes kept
  51. Shot 50, 23:40 to 23:59, 1 of 1 keyframes kept
  52. Shot 51, 23:59 to 24:29, 1 of 1 keyframes kept
  53. Shot 52, 24:29 to 24:58, 1 of 1 keyframes kept
  54. Shot 53, 24:58 to 25:28, 0 of 1 keyframes kept
  55. Shot 54, 25:28 to 25:54, 1 of 1 keyframes kept
  56. Shot 55, 25:54 to 26:20, 1 of 1 keyframes kept
  57. Shot 56, 26:20 to 26:45, 0 of 1 keyframes kept
  58. Shot 57, 26:45 to 27:11, 1 of 1 keyframes kept
  59. Shot 58, 27:11 to 27:45, 1 of 1 keyframes kept
  60. Shot 59, 27:45 to 28:18, 0 of 1 keyframes kept
  61. Shot 60, 28:18 to 28:22, 1 of 1 keyframes kept
  62. Shot 61, 28:22 to 28:59, 1 of 1 keyframes kept
  63. Shot 62, 28:59 to 29:36, 0 of 1 keyframes kept
  64. Shot 63, 29:36 to 29:57, 1 of 1 keyframes kept
  65. Shot 64, 29:57 to 30:22, 1 of 1 keyframes kept
  66. Shot 65, 30:22 to 30:47, 1 of 1 keyframes kept
  67. Shot 66, 30:47 to 31:17, 1 of 1 keyframes kept
  68. Shot 67, 31:17 to 31:46, 1 of 1 keyframes kept
  69. Shot 68, 31:46 to 32:16, 1 of 1 keyframes kept
  70. Shot 69, 32:16 to 32:45, 0 of 1 keyframes kept
  71. Shot 70, 32:45 to 33:15, 1 of 1 keyframes kept
  72. Shot 71, 33:15 to 33:38, 1 of 1 keyframes kept

72 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
257
whisperx 257
chunks
57
from 257 cues
keyframes
52
kept of 72 captured
frames with text
52
776 lines read
chapters
15
from the source metadata
keyframe bytes
6.9 MB
word timings on 257 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:24 0s
stt done 2026-08-09 06:16 39s
chunk done 2026-08-09 06:17 0s
text_embed done 2026-08-10 19:39 0s
keyframe done 2026-08-09 06:17 2m 05s
ocr done 2026-08-09 06:19 21s
frame_embed done 2026-08-10 19:39 9s

Frames, and what the machine read

  • 0:24 #0 done9 line(s)

    shot 0·sharpness 1405.4

    1. AlEngineer0.97
    2. World'sFair1.00
    3. SAN FRANCISCO0.98
    4. JUNE 29 - JULY 2, 20260.95
    5. Al Agents for Performance0.99
    6. Ship Faster, Pay Less0.99
    7. A practitioner's guide to catalog-backed performance agents.1.00
    8. Rajat Shah·Staff Software Engineer0.98
    9. Al Platform, Netflix0.96
  • 0:47 #1 done5 line(s)

    shot 1·sharpness 665.1

    1. PART 011.00
    2. The Problem1.00
    3. Why Performance Engineering doesn't scale.1.00
    4. What it costs.1.00
    5. Rajat Shah · Al Platform, Netflix0.97
  • 1:04 #2 done3 line(s)

    shot 2·sharpness 623.5

    1. THE VIBE CODING ERA1.00
    2. We ship 10x faster now.0.99
    3. Rajat Shah · Al Platform, Netflix0.97
  • 1:19 #3 done5 line(s)

    shot 3·sharpness 805.5

    1. THE VIBE CODING ERA0.98
    2. We ship 10x faster now.1.00
    3. Our CPU bills ship 10x faster too.1.00
    4. Vibe coding giveth. The AWS invoice taketh away.1.00
    5. Rajat Shah· Al Platform, Netflix0.98
  • 1:30 #4 done6 line(s)

    shot 4·sharpness 1145.5

    1. THE NEW PROBLEM1.00
    2. Your Al ships code. Not0.99
    3. always fast code.0.97
    4. LLMs don't know your platform's0.99
    5. performance patterns.0.99
    6. Rajat Shah· Al Platform, Netflix0.97
  • 2:04 #5 skipped

    shot 5·duplicate of #4

  • 2:37 #6 done11 line(s)

    shot 6·sharpness 824.2

    1. THE STATUS QUO1.00
    2. What a (human) perf engineer does0.98
    3. today1.00
    4. TRIGGER1.00
    5. Run the profiler on the0.99
    6. production instance.1.00
    7. 21.00
    8. 31.00
    9. 41.00
    10. 51.00
    11. Rajat Shah · Al Platform, Netflix0.97
  • 3:02 #7 skipped

    shot 7·duplicate of #6

  • 3:39 #8 done26 line(s)

    shot 8·sharpness 1872.0

    1. THE STATUS QUO0.99
    2. What a (human) perf engineer does1.00
    3. today1.00
    4. Complete Visibility1.00
    5. Java Mixed-Mode Flame Graph via Linux perf_events1.00
    6. TRIGGER1.00
    7. SQUINT1.00
    8. Run the profiler on the0.99
    9. production instance.1.00
    10. Stare at the colors for 201.00
    11. minutes straight.1.00
    12. (JVM)1.00
    13. C++1.00
    14. Java1.00
    15. Java1.00
    16. (User)1.00
    17. C0.99
    18. (Inlined)1.00
    19. C0.99
    20. (Kernel)1.00
    21. 21.00
    22. 31.00
    23. DOWNLOAD1.00
    24. Export the flamegraph for1.00
    25. local analysis.1.00
    26. Rajat Shah· Al Platform, Netflix0.97
  • 4:04 #9 done24 line(s)

    shot 9·sharpness 1351.5

    1. THE STATUS QUO1.00
    2. What a (human) perf engineer does0.99
    3. today1.00
    4. TRIGGER1.00
    5. SQUINT1.00
    6. FIND (?)0.99
    7. Run the profiler on the0.99
    8. Stare at the colors for 200.99
    9. Maybe identify the root cause1.00
    10. production instance.0.99
    11. minutes straight.0.98
    12. of the bug.0.99
    13. 21.00
    14. 31.00
    15. 41.00
    16. 51.00
    17. 61.00
    18. DOWNLOAD1.00
    19. SEARCH1.00
    20. Export the flamegraph for0.98
    21. Search cryptic function and1.00
    22. local analysis.1.00
    23. frame names.1.00
    24. Rajat Shah· Al Platform, Netflix0.96
  • 4:33 #10 skipped

    shot 10·duplicate of #9

  • 4:59 #11 done6 line(s)

    shot 11·sharpness 580.7

    1. PART 020.99
    2. The1.00
    3. Experiment1.00
    4. Can an LLM read profiling data?1.00
    5. We ran it against a live service.1.00
    6. Rajat Shah· Al Platform, Netflix0.98
  • 5:28 #12 done20 line(s)

    shot 12·sharpness 1465.5

    1. UNDER THE HOOD0.99
    2. Every profiler speaks the same language0.99
    3. PROFILER1.00
    4. RUNTIME1.00
    5. WHAT IT CAPTURES0.98
    6. async-profiler / JFR0.97
    7. JVM (Java·Kotlin·Scala)0.98
    8. Call stacks· inclusive + self CPU·100Hz sampling0.98
    9. py-spy1.00
    10. Python1.00
    11. Call stacks·inclusive + self CPU·zero overhead, async-safe0.98
    12. pprof1.00
    13. Go1.00
    14. Call stacks· inclusive + self CPU·goroutine profiles0.97
    15. PyTorch Profiler0.99
    16. GPU/Python1.00
    17. Call stacks·inclusive + self CPU·GPU kernel ti0.97
    18. Regardless of language or runtime, every profiler yields the same insight: call stacks ranked by inclusive0.99
    19. cost.1.00
    20. Rajat Shah· Al Platform, Netflix0.97
  • 6:02 #13 skipped

    shot 13·duplicate of #12

  • 6:32 #14 done12 line(s)

    shot 14·sharpness 1277.3

    1. UNDER THE HOOD0.96
    2. Al Agent knows the patterns1.00
    3. O(N²)loops0.95
    4. hot method called for every item × every other item0.99
    5. Loopinvariants1.00
    6. value recomputed every pass that could be computed once1.00
    7. Allocation in hot paths0.97
    8. objects created per call (boxing, streams, lambdas)1.00
    9. Lock contention / N+1 patterns0.98
    10. shared state hits on every request; per-item calls that could be batched1.00
    11. Not guessing. Matching known patterns against actual call st0.99
    12. Rajat Shah· Al Platform,Netflix0.97
  • 6:44 #15 done12 line(s)

    shot 15·sharpness 1273.8

    1. UNDER THE HOOD0.96
    2. Al Agent knows the patterns1.00
    3. O(N²)loops0.96
    4. hot method called for every item × every other item0.99
    5. Loopinvariants1.00
    6. value recomputed every pass that could be computed once1.00
    7. Allocation in hot paths0.96
    8. objects created per call (boxing, streams, lambdas)1.00
    9. Lock contention / N+1 patterns0.95
    10. shared state hits on every request; per-item calls that could be batched1.00
    11. Not guessing. Matching known patterns against actual call st0.99
    12. Rajat Shah·Al Platform,Netflix0.99
  • 7:17 #16 skipped

    shot 16·duplicate of #14

  • 7:50 #17 done22 line(s)

    shot 17·sharpness 1966.2

    1. UNDER THE HOOD·THE DATA0.95
    2. What the agent reads: ranked method data, not the flame image.0.99
    3. TensorSet.merge (TensorSet.java:87)1.00
    4. 8.8% Self0.99
    5. 11.00
    6. ↑ EncoderPipeline.process0.98
    7. ↑ RequestHandler.handle0.99
    8. 10.2% Incl1.00
    9. MetricsCounter.resolve (Counter.java:42)1.00
    10. 4.1% Self1.00
    11. 21.00
    12. ↑ RequestRecorder.record0.99
    13. 4.1% Incl0.98
    14. ↑ [framework instrumentation]1.00
    15. ImmutableMap.copyOf1.00
    16. 1.2% Self0.95
    17. 31.00
    18. (ImmutableMap.java:256)0.98
    19. 3.4% Incl0.99
    20. ↑ TensorSet.merge← same call path0.95
    21. self% = method is the leaf doing actual work · inclusive% = full stack cost · call stack traces who called it0.99
    22. Rajat Shah· Al Platform, Netflix0.98
  • 8:13 #18 skipped

    shot 18·duplicate of #17

  • 8:33 #19 done32 line(s)

    shot 19·sharpness 2745.7

    1. UNDER THE HOOD·THE DATA0.95
    2. What the agent reads: ranked method data, not the flame image.0.99
    3. TensorSet.merge (TensorSet.java:87)1.00
    4. 8.8% Self0.98
    5. Agent Analysis1.00
    6. 11.00
    7. ↑ EncoderPipeline.process0.98
    8. ↑ RequestHandler.handle0.99
    9. 10.2% Incl1.00
    10. Detected O(N²) Loop0.98
    11. Rows 1 and 3 share the same call path.0.99
    12. ImmutableMap.copyOf sits nested0.99
    13. MetricsCounter.resolve (Counter.java:42)0.99
    14. 4.1% Self1.00
    15. inside TensorSet.merge.1.00
    16. 21.00
    17. ↑RequestRecorder.record1.00
    18. 4.1% Incl0.99
    19. This creates the quadratic hotspot: the0.99
    20. ↑ [framework instrumentation]0.97
    21. merge copies all N entries on every single0.97
    22. reduce step.1.00
    23. Self% identifies exactly where the CPU is0.98
    24. ImmutableMap.copyOf1.00
    25. 1.2% Self1.00
    26. spent.1.00
    27. 31.00
    28. (ImmutableMap.java:256)1.00
    29. 3.4% Incl1.00
    30. ↑ TensorSet.merge ← same call path0.94
    31. self% = method is the leaf doing actual work · inclusive% = full stack cost · call stack traces who called it0.98
    32. Rajat Shah·Al Platform, Netflix0.96
  • 9:24 #20 done9 line(s)

    shot 20·sharpness 1000.4

    1. UNDER THE HOOD0.99
    2. What the Al Agent does with it0.97
    3. 011.00
    4. 020.99
    5. Parse profiler output1.00
    6. Check exact commit1.00
    7. Methods ranked by inclusive CPU and self CPU cost.1.00
    8. Identifies the precise code version running during the profile.1.00
    9. Rajat Shah· Al Platform, Netflix0.97
  • 9:51 #21 done12 line(s)

    shot 21·sharpness 1240.6

    1. UNDER THE HOOD0.98
    2. What the Al Agent does with it0.97
    3. 011.00
    4. 020.99
    5. Parse profiler output0.99
    6. Check exact commit1.00
    7. Methods ranked by inclusive CPU and self CPU cost.0.99
    8. Identifies the precise code version running during the profile.0.98
    9. 031.00
    10. Filter to repo code1.00
    11. Focuses on hotspots in owned code; skips library internals.1.00
    12. Rajat Shah· Al Platform, Netflix0.97
  • 10:18 #22 done15 line(s)

    shot 22·sharpness 1493.7

    1. UNDER THE HOOD0.99
    2. What the Al Agent does with it0.97
    3. 011.00
    4. 020.99
    5. Parse profiler output0.99
    6. Check exact commit1.00
    7. Methods ranked by inclusive CPU and self CPU cost.1.00
    8. Identifies the precise code version running during the profile.0.99
    9. 031.00
    10. 041.00
    11. Filter to repo code1.00
    12. Trace full call path1.00
    13. Focuses on hotspots in owned code; skips library internals.1.00
    14. Reads source of hot methods from entry point to leaf0.98
    15. Rajat Shah· Al Platform, Netflix0.97
  • 10:32 #23 skipped

    shot 23·duplicate of #22

Transcript

257 cues· 4,862 words· 26,740 chars

  1. 0:00 Hi there.
  2. 0:02 Welcome to AI Engineer World's Fair 2026 event.
  3. 0:05 I'm Rajat Shah.
  4. 0:06 I'm a staff software engineer at Netflix, where I work in the AI platform organization building large-scale distributed systems for machine learning model hosting.
  5. 0:15 In this talk, I'm here to share how we did improve our performance engineering throughput by introducing AI agents into the mix.
  6. 0:25 And this is more of a playbook or a practitioner's guide to help you also replicate similar learnings in your own organizations to improve the infrastructure cost and ship faster.
  7. 0:42 Let's first talk about the problem.
  8. 0:44 Why does performance engineering doesn't scale?
  9. 0:46 And what does it cost to actually do it right?
  10. 0:51 The problem is arising from the fact is that you are authoring code now at a 10x faster speed.
  11. 0:56 The coding agents are getting better and better at solving problems.
  12. 1:02 And as more and more engineers adopt it, it gets very easy to produce code
  13. 1:10 in your system.
  14. 1:11 And this is slight exaggeration, but the compute cost also is increasing at a similar pace because it doesn't always write the fastest code.
  15. 1:24 So this is where the problem arises because of that new wipe coding error.
  16. 1:29 The AI agent ships code.
  17. 1:31 It is pretty much tuned to ship code fast.
  18. 1:35 And of course, you could say that as the coding agents are evolving, newer models are coming into play.
  19. 1:42 They get better and better at simply writing performant code.
  20. 1:47 But that's not always true.
  21. 1:50 Agent doesn't know specific details about your platform and your frameworks and your internal code base patterns.
  22. 1:57 So it tends to just produce code based on what it might have already seen other code bases using or inventing new patterns in your code bases that you did not anticipate an engineer to use as a pattern to use your framework.
  23. 2:16 So let's look at what a performance engineer typically does.
  24. 2:21 I'm calling this as a human performance engineer, which is responsible for identifying bottlenecks in a service and fixing them.
  25. 2:30 Typically, a human would trigger profiling on a single production instance of a fleet of production instances.
  26. 2:39 You would go and download it, potentially open it in a visualizer.
  27. 2:44 The raw data that you download typically
  28. 2:46 great to look at.
  29. 2:48 It could be, for example, a JSON structure data of the call stack and where the CPU is spent.
  30. 2:55 So using a visualizer helps you at least see and visualize the call stack and CPU time of various method in your services better.
  31. 3:07 Once you have that visualizer open, you pretty much end up spending a lot of time in just looking at and finding in this treasure hunt on the potential places where you could improve the code to make it more performant.
  32. 3:27 This takes a lot of time in order to even learn how to look at it.
  33. 3:32 And there's a learning curve to it.
  34. 3:36 And this is where the real bottleneck ends up being.
  35. 3:38 You end up having to spend straight many, many minutes to identify the bottlenecks.
  36. 3:46 Once you have identified some code parts and some packages that are spending significant CPU cycles, you would end up searching it in your code bases, in your code repos, and see if it has a potential to improvement.
  37. 4:01 Hopefully you have luck here and you find a root cause and you produce a code review out.
  38. 4:10 You merge it and you get some performance wins.
  39. 4:15 And then you repeat all of this again.
  40. 4:17 You see the problem, right?
  41. 4:18 This is a very manual effort and very tedious effort to get right.
  42. 4:23 And this ends up being a bottleneck if you were to do it across
  43. 4:27 many of your code bases and code paths.
  44. 4:30 And that's why this is done very rarely.
  45. 4:31 People typically end up looking at profiling data only when something is going wrong at 2am and somebody needs to fix a problem because your CPU is unbearable.
  46. 4:44 So we asked this question internally.
  47. 4:45 Can an LLM read this profiling data?
  48. 4:49 The 20 minutes that I mentioned an engineer spends in identifying hot paths, can an LLM agent, which is fed that data, also do it much faster?
  49. 5:01 And we tried to answer this question through some live services.
  50. 5:05 So the next couple of slides will be about this experiment and how we do it.

Chapters

  1. 0:00 Introduction: performance for ML serving
  2. 2:24 The manual profiling loop today
  3. 4:46 The experiment: can an LLM read a profile?
  4. 7:37 From call stack to the exact method
  5. 9:07 How the agent locates and reads the code
  6. 11:01 First finding: an O(N) fix, canary confirmed
  7. 12:41 The same antipattern across seven services
  8. 15:31 Building a shared pattern catalog
  9. 17:16 Storing and sharing findings across services
  10. 19:40 Feeding the catalog to coding agents
  11. 20:56 Human approval and verification
  12. 22:13 Canary validation on real traffic
  13. 24:18 Reactive vs proactive paths
  14. 26:54 Catching waste before production
  15. 29:32 Autonomy levels and what is next

Open at this second