read-only demo

Videos TJPInBjhE4Q

ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay

index_state ready data_status ok

AI Engineer· published 2026-07-12· 0:25:00· en-US· indexed 2026-08-11 03:33

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:48, 1 of 1 keyframes kept
  2. Shot 1, 0:48 to 1:16, 1 of 1 keyframes kept
  3. Shot 2, 1:16 to 1:43, 0 of 1 keyframes kept
  4. Shot 3, 1:43 to 2:11, 0 of 1 keyframes kept
  5. Shot 4, 2:11 to 2:39, 1 of 1 keyframes kept
  6. Shot 5, 2:39 to 3:08, 0 of 1 keyframes kept
  7. Shot 6, 3:08 to 3:37, 0 of 1 keyframes kept
  8. Shot 7, 3:37 to 4:11, 1 of 1 keyframes kept
  9. Shot 8, 4:11 to 4:45, 0 of 1 keyframes kept
  10. Shot 9, 4:45 to 5:22, 1 of 1 keyframes kept
  11. Shot 10, 5:22 to 6:00, 0 of 1 keyframes kept
  12. Shot 11, 6:00 to 6:35, 1 of 1 keyframes kept
  13. Shot 12, 6:35 to 7:11, 0 of 1 keyframes kept
  14. Shot 13, 7:11 to 8:01, 1 of 1 keyframes kept
  15. Shot 14, 8:01 to 8:31, 1 of 1 keyframes kept
  16. Shot 15, 8:31 to 9:02, 0 of 1 keyframes kept
  17. Shot 16, 9:02 to 9:27, 1 of 1 keyframes kept
  18. Shot 17, 9:27 to 9:52, 0 of 1 keyframes kept
  19. Shot 18, 9:52 to 10:19, 1 of 1 keyframes kept
  20. Shot 19, 10:19 to 10:45, 0 of 1 keyframes kept
  21. Shot 20, 10:45 to 11:11, 0 of 1 keyframes kept
  22. Shot 21, 11:11 to 11:46, 1 of 1 keyframes kept
  23. Shot 22, 11:46 to 12:21, 0 of 1 keyframes kept
  24. Shot 23, 12:21 to 12:50, 1 of 1 keyframes kept
  25. Shot 24, 12:50 to 13:18, 0 of 1 keyframes kept
  26. Shot 25, 13:18 to 13:44, 1 of 1 keyframes kept
  27. Shot 26, 13:44 to 14:10, 0 of 1 keyframes kept
  28. Shot 27, 14:10 to 14:35, 0 of 1 keyframes kept
  29. Shot 28, 14:35 to 15:09, 1 of 1 keyframes kept
  30. Shot 29, 15:09 to 15:42, 0 of 1 keyframes kept
  31. Shot 30, 15:42 to 16:15, 0 of 1 keyframes kept
  32. Shot 31, 16:15 to 16:49, 1 of 1 keyframes kept
  33. Shot 32, 16:49 to 17:23, 0 of 1 keyframes kept
  34. Shot 33, 17:23 to 17:49, 1 of 1 keyframes kept
  35. Shot 34, 17:49 to 18:16, 0 of 1 keyframes kept
  36. Shot 35, 18:16 to 18:42, 0 of 1 keyframes kept
  37. Shot 36, 18:42 to 19:20, 1 of 1 keyframes kept
  38. Shot 37, 19:20 to 19:57, 0 of 1 keyframes kept
  39. Shot 38, 19:57 to 20:22, 1 of 1 keyframes kept
  40. Shot 39, 20:22 to 20:48, 0 of 1 keyframes kept
  41. Shot 40, 20:48 to 21:13, 0 of 1 keyframes kept
  42. Shot 41, 21:13 to 21:47, 1 of 1 keyframes kept
  43. Shot 42, 21:47 to 22:21, 0 of 1 keyframes kept
  44. Shot 43, 22:21 to 22:48, 1 of 1 keyframes kept
  45. Shot 44, 22:48 to 23:15, 0 of 1 keyframes kept
  46. Shot 45, 23:15 to 23:42, 0 of 1 keyframes kept
  47. Shot 46, 23:42 to 24:08, 1 of 1 keyframes kept
  48. Shot 47, 24:08 to 24:33, 0 of 1 keyframes kept
  49. Shot 48, 24:33 to 24:59, 0 of 1 keyframes kept
  50. Shot 49, 24:59 to 25:00, 1 of 1 keyframes kept

50 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
274
whisperx 274
chunks
48
from 274 cues
keyframes
22
kept of 50 captured
frames with text
22
401 lines read
chapters
0
from the source metadata
keyframe bytes
4.7 MB
word timings on 274 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 03:30 1m 08s
stt done 2026-08-11 03:31 24s
chunk done 2026-08-11 03:32 0s
text_embed done 2026-08-11 03:32 1s
keyframe done 2026-08-11 03:32 57s
ocr done 2026-08-11 03:33 16s
frame_embed done 2026-08-11 03:33 3s

Frames, and what the machine read

  • 0:05 #0 done3 line(s)

    shot 0·sharpness 966.6

    1. YOUR CODING AGENT IS1.00
    2. CREATING REVIEW DEBT0.98
    3. Sachin Gupta1.00
  • 1:00 #1 done14 line(s)

    shot 1·sharpness 1624.1

    1. THE GAP NOBODYIS0.97
    2. MEASURING1.00
    3. I0.98
    4. Coding agents ship PRs faster than humans can trust them.1.00
    5. +25%0.90
    6. -27%0.89
    7. Commits, year over year0.98
    8. Comments on commits, same year0.99
    9. GitHub Octoverse 20251.00
    10. GitHub Octoverse 20250.99
    11. And at high-Al-adoption teams, median PR review time is up 441.5% — and 31.3% of PRs are merged with zero review. Faros Al 20260.99
    12. benchmark1.00
    13. The gap between those two numbers is review debt.1.00
    14. 21.00
  • 1:32 #2 skipped

    shot 2·duplicate of #1

  • 1:57 #3 skipped

    shot 3·duplicate of #1

  • 2:14 #4 done16 line(s)

    shot 4·sharpness 1597.3

    1. THE STORY EVERYONE TELLS0.99
    2. I0.97
    3. How most teams measure coding agents today1.00
    4. PRs per developer1.00
    5. +16%1.00
    6. Vanity1.00
    7. Faros Al 2026, "Acceleration Whiplash"0.99
    8. Median PR size0.98
    9. +63%1.00
    10. Vanity1.00
    11. DX 2026 longitudinal study (44 → 72 lines)0.99
    12. Cycle time, open → merge0.99
    13. Modestly down1.00
    14. Misleading1.00
    15. DX 2026 (PR throughput +8% as Al usage rose 65%)0.98
    16. 31.00
  • 2:43 #5 skipped

    shot 5·duplicate of #4

  • 3:19 #6 skipped

    shot 6·duplicate of #4

  • 3:57 #7 done20 line(s)

    shot 7·sharpness 1911.7

    1. THE STORY NOBODY TELLS1.00
    2. I0.99
    3. What those numbers quietly stop measuring1.00
    4. Reviewer fatigue1.00
    5. Late-night merges1.00
    6. Test theater1.00
    7. Senior engineers are carrying dramatically0.99
    8. PRs unreviewed for 4+ days get a thumbs-up0.99
    9. Tests get added — but they assert what the0.99
    10. more review load than a year ago. They are0.99
    11. at 11pm before a Friday deadline.0.99
    12. code did, not what it should do.1.00
    13. not happier about it.0.99
    14. Architectural drift1.00
    15. Incident lag1.00
    16. The same problem solved three different1.00
    17. Bugs traced back to Al PRs land weeks or0.99
    18. ways in three files. Nobody noticed.1.00
    19. months later. Nobody connects the dots.1.00
    20. 41.00
  • 4:15 #8 skipped

    shot 8·duplicate of #7

  • 5:00 #9 done17 line(s)

    shot 9·sharpness 2663.6

    1. DEFINITION1.00
    2. I0.98
    3. Review debt — definition + why it compounds0.99
    4. Review debt (noun): the accumulating gap between the code a coding agent has produced and the code0.99
    5. humans have actually reviewed, trusted, and understood.1.00
    6. Like financial debt — it compounds. The interest is paid in human attention.0.99
    7. Why it compounds — three feedback loops:0.99
    8. 11.00
    9. Agents learn from your codebase1.00
    10. Yesterday's un-reviewed code becomes tomorrow's Al suggestion.1.00
    11. 21.00
    12. Reviewers cede the architectural call0.99
    13. Attention contracts to syntax; big-picture decisions move to never.0.99
    14. 31.00
    15. Velocity expectations reset1.00
    16. Throughput rises; headcount-per-review falls. No slack to pay del0.99
    17. 51.00
  • 5:41 #10 skipped

    shot 10·duplicate of #9

  • 6:28 #11 done21 line(s)

    shot 11·sharpness 2135.0

    1. MEASURING REVIEW1.00
    2. DEBT1.00
    3. Five signal families. Ten deterministic checks. No LLM in the loop.1.00
    4. 0.99
    5. ..0.52
    6. Diff size & coupling0.99
    7. Test evidence gap1.00
    8. Al-authorship indicators1.00
    9. Evidence & rationale gaps1.00
    10. How big, how spread1.00
    11. Code added vs. tests added1.00
    12. Metadata signals of Al0.98
    13. What the PR doesn't explain1.00
    14. authorship1.00
    15. 0.58
    16. Directory & ownership spread0.99
    17. Deterministic = the score is defensible in a real engineer1.00
    18. How many CODEOWNERS touched1.00
    19. review. No model-of-the-week.0.99
    20. Each family rolls up multiple checks: ten checks total, Full mapping in O&A0.97
    21. 61.00
  • 7:00 #12 skipped

    shot 12·duplicate of #11

  • 7:40 #13 done11 line(s)

    shot 13·sharpness 1785.6

    1. SIGNAL 1 / 50.92
    2. I0.99
    3. Diff size & coupling0.97
    4. Why agents struggle here1.00
    5. What it measures0.99
    6. Agents bias toward "just fix the failing call site." Human1.00
    7. Net lines changed, plus how many files are touched and1.00
    8. engineers route fixes to the cause. Agent PRs sprawl across files0.99
    9. whether they cluster (one module) or sprawl (across0.99
    10. searching for the same problem.0.98
    11. modules).1.00
  • 8:07 #14 done11 line(s)

    shot 14·sharpness 2526.6

    1. SIGNAL 2 / 50.97
    2. I0.98
    3. Test evidence gap0.97
    4. Al-authored PRs ship with a far lower test-to-code ratio1.00
    5. than human-authored ones.1.00
    6. Test LOC added ÷ production LOC added — per PR, not aggregate.0.98
    7. And the tests that do show up tend to assert what the code did, not what it should do — locking in0.99
    8. behavior, including bugs.0.99
    9. Why this number is brutal: Agents happily generate tests — but they assert what the code did, not what it should do. The patter0.99
    10. doesn't even capture that quality gap.1.00
    11. 81.00
  • 8:41 #15 skipped

    shot 15·duplicate of #14

  • 9:17 #16 done14 line(s)

    shot 16·sharpness 1510.3

    1. SIGNAL 3 /50.94
    2. Directory & ownership spread0.99
    3. Distinct CODEOWNERS teams or individuals whose files appear in the diff.0.99
    4. FEW1.00
    5. MANY1.00
    6. HUMAN PR0.95
    7. AIP R0.89
    8. Concentrated ownership1.00
    9. Spread ownership0.99
    10. Reviewer holds the whole mental model. One approval, one0.99
    11. Multiple reviewers, multiple contexts. No single human holds the1.00
    12. context.1.00
    13. whole thing.0.99
    14. 91.00
  • 9:37 #17 skipped

    shot 17·duplicate of #16

  • 10:03 #18 done23 line(s)

    shot 18·sharpness 1695.3

    1. SIGNAL 4 / 50.95
    2. I0.96
    3. Al-authorship indicators1.00
    4. What it measures1.00
    5. Real-data check1.00
    6. Detects metadata signals of Al-assisted authorship. Not based on parsing the code0.99
    7. itself — based on what the author published about the code.0.99
    8. Scanned 3 public repos, 524 PRs1.00
    9. Steady-state firing rate: 5-20% of PRs per week1.00
    10. Coauthor footer0.99
    11. likely1.00
    12. Co-authored-by: Copilot <[email protected]>1.00
    13. Repos: A, B, C — public OSS, names withheld0.97
    14. Branch name pattern1.00
    15. possible1.00
    16. Highest signal: likely via coauthor footer (Repo A)0.98
    17. codex/, copilot/, cursor/ prefixes1.00
    18. Lowest signal: 0% on Repo C (no convention)0.98
    19. PR body / commit phrases0.97
    20. possible1.00
    21. "generated by", "assisted by", "powered by Claude"0.99
    22. Amplifier only. Score impact +2 to +5 points (max), and never fires score impact when tests pass and Cl is green.1.00
    23. 101.00
  • 10:32 #19 skipped

    shot 19·duplicate of #18

  • 11:08 #20 skipped

    shot 20·duplicate of #18

  • 11:32 #21 done17 line(s)

    shot 21·sharpness 1146.1

    1. SIGNAL 5 / 50.97
    2. I0.99
    3. Evidence & rationale gaps1.00
    4. Whether the PR explains why — not just what.1.00
    5. HIGH GAP1.00
    6. LOW GAP0.99
    7. "Fix flaky tests"1.00
    8. "Stop using getRawHeader for SNI...0.98
    9. PR body length0.98
    10. PR body length0.98
    11. 18 characters0.97
    12. 420 characters + repro1.00
    13. Commit message1.00
    14. Linked1.00
    15. "updates"1.00
    16. Issue · Design doc· Bench0.96
    17. 111.00
  • 12:03 #22 skipped

    shot 22·duplicate of #21

  • 12:35 #23 done21 line(s)

    shot 23·sharpness 1789.5

    1. THE SCORING RUBRIC1.00
    2. I0.98
    3. How the five signals combine0.99
    4. ReviewDebt(PR)1.00
    5. 0.25·DiffCoupling1.00
    6. 0.25·TestEvidenceGap1.00
    7. 0.20·OwnerSpread0.99
    8. +0.68
    9. 0.15·AI-indicators0.98
    10. +0.83
    11. 0.15·RationaleGap0.99
    12. 0-251.00
    13. 26-501.00
    14. 51-751.00
    15. 76-1001.00
    16. Healthy1.00
    17. Watch1.00
    18. High debt0.99
    19. Reject / refactor0.96
    20. Weights are starting defaults. Calibrate them against your last 200 PRs.1.00
    21. 121.00

Transcript

274 cues· 3,872 words· 20,888 chars

  1. 0:00 Hi everyone.
  2. 0:01 I'm Sachin Gupta and I'm a software engineer.
  3. 0:04 The title is exactly what it sounds like.
  4. 0:06 Your coding agent is creating review debt.
  5. 0:09 Before I start, sit with the title for a second.
  6. 0:12 Notice what it is not saying.
  7. 0:14 I'm not saying that coding agents are bad.
  8. 0:17 I'm not saying they don't make us faster.
  9. 0:20 I'm saying they're creating a kind of debt that nobody is measuring.
  10. 0:25 And that date is going to come to you.
  11. 0:28 Over the next few minutes, here is what I'm going to do.
  12. 0:31 I will define what review date is.
  13. 0:33 I'll walk five signal families that are going to compose it.
  14. 0:37 I'll score three real pull requests side by side.
  15. 0:39 And I will show you across repo scan of over 500 PRs, which is nothing but three public code bases.
  16. 0:45 So let's go.
  17. 0:46 So look here.
  18. 0:50 Here's the gap nobody's measuring.
  19. 0:52 GitHub's 2025 October's report that covers almost every public pull request on the planet shows that the commits climbed 25% year over year.
  20. 1:03 And now over the same year, comments on the commits dropped 27%.
  21. 1:07 Now, these comments are nothing but the proxy for review activity.
  22. 1:11 Code production volume actually went up
  23. 1:14 but the review attention went down.
  24. 1:16 They moved in opposite directions that too in the same year.
  25. 1:19 Now look at the teams for this along the AI adoption curve.
  26. 1:23 Paros AI tracks this cohort in their 2026 benchmark.
  27. 1:28 Median PR review time is up by 441.5%.
  28. 1:33 And if you see, like if you calculate, you'll figure out that the review PRs take 5.4 times longer than what they used to.
  29. 1:41 and then 31% more PR are now merged with no review at all.
  30. 1:46 So AI is producing the code very fast.
  31. 1:50 AI is producing the pull request very fast, but humans cannot responsibly review them at that pace.
  32. 1:56 This gap is called as review debt.
  33. 1:59 It actually accrues quietly, it compounds, and right now, nobody has a number for it.
  34. 2:05 But by the end of this talk, I will make sure that you will have a certain number.
  35. 2:12 Now, if you look at this particular slide, here is the story every team is telling right now.
  36. 2:17 PR per developer are up 16%.
  37. 2:20 That's the Pharos AI acceleration whiplash benchmark, and that too from April 2026.
  38. 2:27 There were like 22,000 developers, 4,000 teams, and the median PR size is up 63%, 44 lines to 72 lines per pull request.
  39. 2:38 That is basically 16 months of long-term data from DX2026 study, which includes 400 organizations.
  40. 2:46 Now, if you see the cycle time open to merge, that is modestly down and it is coming from the same DX study.
  41. 2:54 The framing is very generous.
  42. 2:55 PR throughput grew actually about 8% and the AI usage rose about 65%.
  43. 3:02 The gain that we see today is real, but it is actually smaller than the hype that is present.
  44. 3:08 Every one of these numbers are real.
  45. 3:10 None of them is a lie.
  46. 3:11 But everyone is a vanity metric.
  47. 3:14 PR count goes up when one PR splits into seven PRs.
  48. 3:18 Median PR size going up is not a benefit.
  49. 3:21 It's actually a bloating.
  50. 3:22 Cycle time going down when reviewer stop pushing back.

Open at this second