read-only demo

Videos 7vn4WpqNpck

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

index_state ready data_status ok

AI Engineer· published 2026-08-08· 0:18:07· en-US· indexed 2026-08-10 19:34

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:23, 1 of 1 keyframes kept
  5. Shot 4, 0:23 to 0:54, 1 of 1 keyframes kept
  6. Shot 5, 0:54 to 1:31, 1 of 1 keyframes kept
  7. Shot 6, 1:31 to 1:53, 1 of 1 keyframes kept
  8. Shot 7, 1:53 to 2:19, 0 of 1 keyframes kept
  9. Shot 8, 2:19 to 2:45, 1 of 1 keyframes kept
  10. Shot 9, 2:45 to 3:29, 0 of 1 keyframes kept
  11. Shot 10, 3:29 to 3:45, 0 of 1 keyframes kept
  12. Shot 11, 3:45 to 4:11, 0 of 1 keyframes kept
  13. Shot 12, 4:11 to 4:59, 1 of 1 keyframes kept
  14. Shot 13, 4:59 to 5:19, 0 of 1 keyframes kept
  15. Shot 14, 5:19 to 5:50, 0 of 1 keyframes kept
  16. Shot 15, 5:50 to 6:18, 0 of 1 keyframes kept
  17. Shot 16, 6:18 to 6:47, 1 of 1 keyframes kept
  18. Shot 17, 6:47 to 6:53, 0 of 1 keyframes kept
  19. Shot 18, 6:53 to 7:28, 0 of 1 keyframes kept
  20. Shot 19, 7:28 to 7:49, 0 of 1 keyframes kept
  21. Shot 20, 7:49 to 8:23, 0 of 1 keyframes kept
  22. Shot 21, 8:23 to 8:57, 0 of 1 keyframes kept
  23. Shot 22, 8:57 to 9:41, 1 of 1 keyframes kept
  24. Shot 23, 9:41 to 10:20, 0 of 1 keyframes kept
  25. Shot 24, 10:20 to 11:01, 0 of 1 keyframes kept
  26. Shot 25, 11:01 to 11:33, 1 of 1 keyframes kept
  27. Shot 26, 11:33 to 11:45, 0 of 1 keyframes kept
  28. Shot 27, 11:45 to 12:03, 0 of 1 keyframes kept
  29. Shot 28, 12:03 to 12:19, 0 of 1 keyframes kept
  30. Shot 29, 12:19 to 12:39, 0 of 1 keyframes kept
  31. Shot 30, 12:39 to 13:02, 0 of 1 keyframes kept
  32. Shot 31, 13:02 to 13:32, 0 of 1 keyframes kept
  33. Shot 32, 13:32 to 13:59, 0 of 1 keyframes kept
  34. Shot 33, 13:59 to 14:26, 0 of 1 keyframes kept
  35. Shot 34, 14:26 to 14:58, 0 of 1 keyframes kept
  36. Shot 35, 14:58 to 15:26, 1 of 1 keyframes kept
  37. Shot 36, 15:26 to 15:55, 0 of 1 keyframes kept
  38. Shot 37, 15:55 to 16:24, 0 of 1 keyframes kept
  39. Shot 38, 16:24 to 16:53, 0 of 1 keyframes kept
  40. Shot 39, 16:53 to 17:22, 0 of 1 keyframes kept
  41. Shot 40, 17:22 to 17:51, 0 of 1 keyframes kept
  42. Shot 41, 17:51 to 18:07, 0 of 1 keyframes kept

42 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
202
whisperx 202
chunks
32
from 202 cues
keyframes
13
kept of 42 captured
frames with text
13
228 lines read
chapters
0
from the source metadata
keyframe bytes
5.0 MB
word timings on 202 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 21:03 0s
stt done 2026-08-08 23:23 33s
chunk done 2026-08-08 23:23 0s
text_embed done 2026-08-10 19:34 0s
keyframe done 2026-08-08 23:23 4m 14s
ocr done 2026-08-08 23:28 7s
frame_embed done 2026-08-10 19:34 3s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 457.7

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 661.9

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2740.4

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.97
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:21 #3 done2 line(s)

    shot 3·sharpness 254.8

    1. AlEngineer0.96
    2. World's Fair0.97
  • 0:38 #4 done10 line(s)

    shot 4·sharpness 1838.5

    1. AlEngineer0.99
    2. World'sFair1.00
    3. Three foundational issues1.00
    4. PRESENTED BY1.00
    5. Too slow to meet customer demand1.00
    6. Microsoft1.00
    7. Too complicated to update0.99
    8. No one wants to touch the 10+ repos where the code lives0.99
    9. Engineering the future of Al0.97
    10. sFair1.00
  • 1:12 #5 done19 line(s)

    shot 5·sharpness 2142.3

    1. AlEngineer0.98
    2. World'sFair1.00
    3. DEFINING THE USE CASE0.98
    4. Processing Complex Medical Claims 10k+ Pages0.99
    5. PRESENTED BY0.96
    6. Microsoft1.00
    7. Claims Analysis UI1.00
    8. Pipeline1.00
    9. HITL1.00
    10. Medical Claims0.99
    11. Case Report1.00
    12. USERS1.00
    13. Doctors1.00
    14. Claims Adjusters1.00
    15. Government1.00
    16. How to scale this 100x?1.00
    17. TRACK 8· JULY 2, 20260.93
    18. sFair0.94
    19. Agentic Engineering1.00
  • 1:49 #6 done17 line(s)

    shot 6·sharpness 1664.6

    1. AlEngineer0.98
    2. World'sFair1.00
    3. WhereIstarted0.99
    4. import java.awt.image.BufferedImage;1.00
    5. import java.io.File;0.99
    6. import javax.imageio.ImageIO;0.98
    7. public class Main {1.00
    8. public static void main(String[] args) throws Exception {0.99
    9. BufferedImage slime = ImageIO.read(new File("slime.png"));0.99
    10. // Prints the Image object representation0.99
    11. System.out.println(slime);1.00
    12. 30.50
    13. 30.84
    14. Fifteen plus years ago. I had no idea what I was doing.0.99
    15. TRACK 8· JULY 2,20260.95
    16. sFair1.00
    17. Agentic Engineering1.00
  • 2:04 #7 skipped

    shot 7·duplicate of #6

  • 2:24 #8 done20 line(s)

    shot 8·sharpness 2077.3

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Al Engineering Allows Teams To Go Faster1.00
    4. REFERENCE, SPOTIFY VP ENGINEERING0.97
    5. REFERENCE, STRIPE ENGINEERING0.99
    6. 4,500deploys a day 73% PRs Al-assisted0.98
    7. By hand, a team0.98
    8. Two months1.00
    9. Boris sat down wth Spotify VP of Engineering Niklas Gustavsson.0.94
    10. With a model, alone0.97
    11. Spotify ships 4,500 production dleploys a day, and 73% of Pis are now0.93
    12. 1day1.00
    13. A codebase-wide migration across a 50-million-line Ruby0.99
    14. codebase.1.00
    15. 12:05 PM ·Aon 29. 2026 5.4MVe0.77
    16. A small amount of debt per change compounds fast at that pace, and it clears just as fast when it needs to.0.99
    17. Watch the interviewRead the Anthropic post0.99
    18. TRACK 8· JULY 2, 20260.94
    19. sFair1.00
    20. Agentic Engineering1.00
  • 2:59 #9 skipped

    shot 9·duplicate of #6

  • 3:43 #10 skipped

    shot 10·duplicate of #6

  • 4:08 #11 skipped

    shot 11·duplicate of #6

  • 4:44 #12 done28 line(s)

    shot 12·sharpness 1916.8

    1. AlEngineer0.99
    2. World's Fair0.97
    3. PRODUCT AND ARCHITECTURE RESEARCH0.98
    4. This process could be 90% faster today0.99
    5. 2025,Manual Confluence Doc0.99
    6. Retrying now, with Deep Research0.99
    7. Profect0.80
    8. nye0.90
    9. Ray0.73
    10. med % smv Pe morkt0.62
    11. Yes we can fetch a0.92
    12. asyncio ibrary0.94
    13. Yens, we can use o0.86
    14. vigper o an asy0.88
    15. create I tont of Ou0.62
    16. Architectural Evaluation and Feature Mapping of Distributed Orchestration0.98
    17. Engines1.00
    18. Criteria1.00
    19. Problem Statement0.97
    20. Deepresearch1.00
    21. Subagents1.00
    22. POC building0.98
    23. Aggregate Scoring0.97
    24. CAUTION0.94
    25. Deep research still needs deep scrutiny. The deep research doc had many surface level assumptions and needed follow up prompts to achieve our standard0.99
    26. TRACK 8·JULY 2,20260.98
    27. sFair1.00
    28. Agentic Engineering0.99
  • 5:16 #13 skipped

    shot 13·duplicate of #6

  • 5:38 #14 skipped

    shot 14·duplicate of #6

  • 5:54 #15 skipped

    shot 15·duplicate of #6

  • 6:22 #16 done47 line(s)

    shot 16·sharpness 2244.9

    1. AlEngineer0.94
    2. TASKTWO·MODELCOMPARISON1.00
    3. World's Fair0.96
    4. Three models, one task0.97
    5. 030.99
    6. Sonnet4.6·medium0.99
    7. Opus 4.8 · medium0.98
    8. COST1.00
    9. WALL TIME0.99
    10. DEV TIME1.00
    11. COST1.00
    12. WALL TIME0.99
    13. DEV TIME1.00
    14. COST1.00
    15. WALL TIME0.97
    16. DEV TIME1.00
    17. PRESENTED BY1.00
    18. $0.161.00
    19. N/A1.00
    20. 3h+0.98
    21. $0.701.00
    22. 6min1.00
    23. 25min1.00
    24. $2.061.00
    25. 12min1.00
    26. 20min1.00
    27. Microsoft1.00
    28. MISTAKES MADE0.97
    29. ONE ADDITIONAL TURN NEEDED· 2 MISTAKES0.97
    30. 1 MINOR MISTAKE1.00
    31. 10+, previous slide1.00
    32. Wrote mock tests instead of integration tests1.00
    33. Referenced the wrong base model (minor)1.00
    34. Wrong docstringparameter0.99
    35. Lots more tests, validations, pytest runs0.98
    36. TOOL CALLS1.00
    37. 42 total - no terminal calls · 7 grep/glob0.96
    38. TOOL CALLS1.00
    39. TOOL CALLS1.00
    40. 49 total · 14 terminal · plus a subagent with 30.98
    41. 84 total - 42 shell, plan, and link calls0.98
    42. shell calls1.00
    43. Models and harnesses have both improved fast enough that this class of task is shifting from hours of hand-debugging to a review pass,1.00
    44. changing the mental model for how much of a refactor an agent can own.0.99
    45. TRACK 8· JULY 2,20260.97
    46. sFair1.00
    47. Agentic Engineering0.99
  • 6:47 #17 skipped

    shot 17·duplicate of #6

  • 7:20 #18 skipped

    shot 18·duplicate of #6

  • 7:38 #19 skipped

    shot 19·duplicate of #6

  • 8:16 #20 skipped

    shot 20·duplicate of #6

  • 8:50 #21 skipped

    shot 21·duplicate of #6

  • 9:28 #22 done27 line(s)

    shot 22·sharpness 1980.4

    1. AlEngineer0.97
    2. World'sFair0.99
    3. WHY RELIABILITY OF AGENT CODING ALLOWS A STRONG MENTAL MODEL0.99
    4. Even Mythos Gets Some Short Tasks Wrong1.00
    5. Predicted 50% time horizon: 17 hr for Claude Mythos Preview (early)0.99
    6. Measurements above 16 hrs are unreliable with our current task suite0.98
    7. 100%1.00
    8. Es te0.66
    9. 80%-0.98
    10. 60%-0.99
    11. 40%-0.96
    12. 20%-0.99
    13. 0%-0.99
    14. ::.:0.57
    15. 1s1.00
    16. 4s1.00
    17. 15s1.00
    18. 1m1.00
    19. 4m1.00
    20. 15m1.00
    21. 1h0.99
    22. 4h1.00
    23. 16h1.00
    24. 64h1.00
    25. TRACK 8· JULY 2, 20260.95
    26. sFair1.00
    27. Agentic Engineering1.00
  • 10:01 #23 skipped

    shot 23·duplicate of #6

Transcript

202 cues· 3,402 words· 18,565 chars

  1. 0:12 So it's not just my AI pipeline that's on fire, but also my PowerPoint.
  2. 0:15 So it's 2025, we're scaling as a business and things are going poorly.
  3. 0:20 We're adding too many customers, we're not getting the throughput we need, and we need to improve our underlying technology.
  4. 0:27 And there's three main issues that we're facing.
  5. 0:30 The first one is that we're too slow to meet customer demand.
  6. 0:33 The second one is that this AI pipeline that we've built is too complicated to update.
  7. 0:37 And the third one is because it's a legacy code base or actually more than 10 repos, nobody actually wants to touch the code.
  8. 0:42 It's not a fun experience.
  9. 0:45 So we made this decision to refactor over the course of six months.
  10. 0:49 And the real question for this talk today, was this the right move to do?
  11. 0:54 So I'll spend this time answering this question, but let's start off with the use case.
  12. 0:57 So the company I work at, WiseDocs, processes complex medical claims, which are PDFs that are more than 10,000 pages in size.
  13. 1:04 Some of these files are bigger than video files.
  14. 1:07 So it's a pretty complex application, and because of this, it's actually non-trivial to scale the different parts.
  15. 1:12 So we're gonna talk about the pipeline today, which has a number of ML models.
  16. 1:17 So divide this talk into a number of chapters.
  17. 1:19 We'll start off with the first one, which is the concept of tech debt.
  18. 1:23 So I think we all have this feeling, universally, if we've been developers for a while, that we all write bad code.
  19. 1:28 The question is, do we do this intentionally or not?
  20. 1:31 If I look back to some of the earliest code I used to write, it was bad.
  21. 1:35 This was more than 15 years ago.
  22. 1:37 I tried to print an image of this character from a video game, and I didn't understand that you can't system.out.println in Java to render something on the screen.
  23. 1:47 So hopefully I've come further from that point in time, but there's these moments where we all know that we've written bad code before.
  24. 1:55 If we think about technical debt as financial debt, it compounds in mysterious and sometimes unexpected ways, but you should think about it in a rigorous format as well.
  25. 2:04 For us to achieve some kind of ROI by taking on technical debt, such as building a feature or getting new customers, we want to make sure that the ROI makes sense.
  26. 2:12 If we introduce additional complexity into our code base, we can very quickly outrun the ROI we've generated.
  27. 2:21 Now with AI engineering, you've probably seen a number of different stories that have come out to showcase the progress that's been made.
  28. 2:26 These are two case studies from Anthropic, one from Spotify and the other from Stripe, talking about the immense progress that they've made, both in shipping velocity and also the ability to refactor code.
  29. 2:37 So at this point in time, writing code or making changes is something that teams are doing faster and faster.
  30. 2:43 Now I'll pause here.
  31. 2:46 Who here thinks that products have gotten better in the past 20 years?
  32. 2:49 Technical products, also raise your hand.
  33. 2:52 I hope everybody, right?
  34. 2:54 Phones are pretty cool.
  35. 2:54 How about five years?
  36. 2:58 How about past year?
  37. 3:01 So the challenge is that we're going faster and faster through the technology lifecycle, but we've lost something.
  38. 3:07 The product focused on customers in some way has degraded.
  39. 3:10 The maintainability of the code and the reliability has degraded.
  40. 3:13 You can see some of the up times here from two leading companies.
  41. 3:16 I've blurred out their names for, it doesn't actually matter who they are, but we are below a 3.9 or even a 4.9 reliability.
  42. 3:24 So even though we're shipping faster and faster, the code quality and the product quality has not necessarily gone up.
  43. 3:31 So let's talk about the refactor that we did.
  44. 3:33 So we started this refactor with the actual code implementation in April and did some pre-work earlier.
  45. 3:38 So I'll go through five different tasks that we did and share some of the findings that we had before and after, especially with as new models have come out.
  46. 3:46 So we spent around two months evaluating orchestrators for our AI pipeline.
  47. 3:50 We looked at five open source projects and we wanted to benchmark and see how effective they were for our use case.
  48. 3:56 And we started this off before deep research came out as part of Google and OpenAI so that web search capability to do a comprehensive analysis was still not there.
  49. 4:05 Now, after we actually gathered these requirements, we built out proof of concepts with a team of three to make sure that we actually got the right results.
  50. 4:12 Now, I'm pretty confident we could do this 90% faster now with the tooling that we have.

Open at this second