Videos 7vn4WpqNpck
Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
Scene timeline
42 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 202
- whisperx 202
- chunks
- 32
- from 202 cues
- keyframes
- 13
- kept of 42 captured
- frames with text
- 13
- 228 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 5.0 MB
- word timings on 202 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 21:03 | 0s |
stt |
done | — | 2026-08-08 23:23 | 33s |
chunk |
done | — | 2026-08-08 23:23 | 0s |
text_embed |
done | — | 2026-08-10 19:34 | 0s |
keyframe |
done | — | 2026-08-08 23:23 | 4m 14s |
ocr |
done | — | 2026-08-08 23:28 | 7s |
frame_embed |
done | — | 2026-08-10 19:34 | 3s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.97
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.99
- World'sFair1.00
- Three foundational issues1.00
- PRESENTED BY1.00
- Too slow to meet customer demand1.00
- Microsoft1.00
- Too complicated to update0.99
- No one wants to touch the 10+ repos where the code lives0.99
- Engineering the future of Al0.97
- sFair1.00
-
- AlEngineer0.98
- World'sFair1.00
- DEFINING THE USE CASE0.98
- Processing Complex Medical Claims 10k+ Pages0.99
- PRESENTED BY0.96
- Microsoft1.00
- Claims Analysis UI1.00
- Pipeline1.00
- HITL1.00
- Medical Claims0.99
- Case Report1.00
- USERS1.00
- Doctors1.00
- Claims Adjusters1.00
- Government1.00
- How to scale this 100x?1.00
- TRACK 8· JULY 2, 20260.93
- sFair0.94
- Agentic Engineering1.00
-
- AlEngineer0.98
- World'sFair1.00
- WhereIstarted0.99
- import java.awt.image.BufferedImage;1.00
- import java.io.File;0.99
- import javax.imageio.ImageIO;0.98
- public class Main {1.00
- public static void main(String[] args) throws Exception {0.99
- BufferedImage slime = ImageIO.read(new File("slime.png"));0.99
- // Prints the Image object representation0.99
- System.out.println(slime);1.00
- 30.50
- 30.84
- Fifteen plus years ago. I had no idea what I was doing.0.99
- TRACK 8· JULY 2,20260.95
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- World'sFair1.00
- Al Engineering Allows Teams To Go Faster1.00
- REFERENCE, SPOTIFY VP ENGINEERING0.97
- REFERENCE, STRIPE ENGINEERING0.99
- 4,500deploys a day 73% PRs Al-assisted0.98
- By hand, a team0.98
- Two months1.00
- Boris sat down wth Spotify VP of Engineering Niklas Gustavsson.0.94
- With a model, alone0.97
- Spotify ships 4,500 production dleploys a day, and 73% of Pis are now0.93
- 1day1.00
- A codebase-wide migration across a 50-million-line Ruby0.99
- codebase.1.00
- 12:05 PM ·Aon 29. 2026 5.4MVe0.77
- A small amount of debt per change compounds fast at that pace, and it clears just as fast when it needs to.0.99
- Watch the interviewRead the Anthropic post0.99
- TRACK 8· JULY 2, 20260.94
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.99
- World's Fair0.97
- PRODUCT AND ARCHITECTURE RESEARCH0.98
- This process could be 90% faster today0.99
- 2025,Manual Confluence Doc0.99
- Retrying now, with Deep Research0.99
- Profect0.80
- nye0.90
- Ray0.73
- med % smv Pe morkt0.62
- Yes we can fetch a0.92
- asyncio ibrary0.94
- Yens, we can use o0.86
- vigper o an asy0.88
- create I tont of Ou0.62
- Architectural Evaluation and Feature Mapping of Distributed Orchestration0.98
- Engines1.00
- Criteria1.00
- Problem Statement0.97
- Deepresearch1.00
- Subagents1.00
- POC building0.98
- Aggregate Scoring0.97
- CAUTION0.94
- Deep research still needs deep scrutiny. The deep research doc had many surface level assumptions and needed follow up prompts to achieve our standard0.99
- TRACK 8·JULY 2,20260.98
- sFair1.00
- Agentic Engineering0.99
-
- AlEngineer0.94
- TASKTWO·MODELCOMPARISON1.00
- World's Fair0.96
- Three models, one task0.97
- 030.99
- Sonnet4.6·medium0.99
- Opus 4.8 · medium0.98
- COST1.00
- WALL TIME0.99
- DEV TIME1.00
- COST1.00
- WALL TIME0.99
- DEV TIME1.00
- COST1.00
- WALL TIME0.97
- DEV TIME1.00
- PRESENTED BY1.00
- $0.161.00
- N/A1.00
- 3h+0.98
- $0.701.00
- 6min1.00
- 25min1.00
- $2.061.00
- 12min1.00
- 20min1.00
- Microsoft1.00
- MISTAKES MADE0.97
- ONE ADDITIONAL TURN NEEDED· 2 MISTAKES0.97
- 1 MINOR MISTAKE1.00
- 10+, previous slide1.00
- Wrote mock tests instead of integration tests1.00
- Referenced the wrong base model (minor)1.00
- Wrong docstringparameter0.99
- Lots more tests, validations, pytest runs0.98
- TOOL CALLS1.00
- 42 total - no terminal calls · 7 grep/glob0.96
- TOOL CALLS1.00
- TOOL CALLS1.00
- 49 total · 14 terminal · plus a subagent with 30.98
- 84 total - 42 shell, plan, and link calls0.98
- shell calls1.00
- Models and harnesses have both improved fast enough that this class of task is shifting from hours of hand-debugging to a review pass,1.00
- changing the mental model for how much of a refactor an agent can own.0.99
- TRACK 8· JULY 2,20260.97
- sFair1.00
- Agentic Engineering0.99
-
- AlEngineer0.97
- World'sFair0.99
- WHY RELIABILITY OF AGENT CODING ALLOWS A STRONG MENTAL MODEL0.99
- Even Mythos Gets Some Short Tasks Wrong1.00
- Predicted 50% time horizon: 17 hr for Claude Mythos Preview (early)0.99
- Measurements above 16 hrs are unreliable with our current task suite0.98
- 100%1.00
- Es te0.66
- 80%-0.98
- 60%-0.99
- 40%-0.96
- 20%-0.99
- 0%-0.99
- ::.:0.57
- 1s1.00
- 4s1.00
- 15s1.00
- 1m1.00
- 4m1.00
- 15m1.00
- 1h0.99
- 4h1.00
- 16h1.00
- 64h1.00
- TRACK 8· JULY 2, 20260.95
- sFair1.00
- Agentic Engineering1.00
Transcript
202 cues· 3,402 words· 18,565 chars
- 0:12 So it's not just my AI pipeline that's on fire, but also my PowerPoint.
- 0:15 So it's 2025, we're scaling as a business and things are going poorly.
- 0:20 We're adding too many customers, we're not getting the throughput we need, and we need to improve our underlying technology.
- 0:27 And there's three main issues that we're facing.
- 0:30 The first one is that we're too slow to meet customer demand.
- 0:33 The second one is that this AI pipeline that we've built is too complicated to update.
- 0:37 And the third one is because it's a legacy code base or actually more than 10 repos, nobody actually wants to touch the code.
- 0:42 It's not a fun experience.
- 0:45 So we made this decision to refactor over the course of six months.
- 0:49 And the real question for this talk today, was this the right move to do?
- 0:54 So I'll spend this time answering this question, but let's start off with the use case.
- 0:57 So the company I work at, WiseDocs, processes complex medical claims, which are PDFs that are more than 10,000 pages in size.
- 1:04 Some of these files are bigger than video files.
- 1:07 So it's a pretty complex application, and because of this, it's actually non-trivial to scale the different parts.
- 1:12 So we're gonna talk about the pipeline today, which has a number of ML models.
- 1:17 So divide this talk into a number of chapters.
- 1:19 We'll start off with the first one, which is the concept of tech debt.
- 1:23 So I think we all have this feeling, universally, if we've been developers for a while, that we all write bad code.
- 1:28 The question is, do we do this intentionally or not?
- 1:31 If I look back to some of the earliest code I used to write, it was bad.
- 1:35 This was more than 15 years ago.
- 1:37 I tried to print an image of this character from a video game, and I didn't understand that you can't system.out.println in Java to render something on the screen.
- 1:47 So hopefully I've come further from that point in time, but there's these moments where we all know that we've written bad code before.
- 1:55 If we think about technical debt as financial debt, it compounds in mysterious and sometimes unexpected ways, but you should think about it in a rigorous format as well.
- 2:04 For us to achieve some kind of ROI by taking on technical debt, such as building a feature or getting new customers, we want to make sure that the ROI makes sense.
- 2:12 If we introduce additional complexity into our code base, we can very quickly outrun the ROI we've generated.
- 2:21 Now with AI engineering, you've probably seen a number of different stories that have come out to showcase the progress that's been made.
- 2:26 These are two case studies from Anthropic, one from Spotify and the other from Stripe, talking about the immense progress that they've made, both in shipping velocity and also the ability to refactor code.
- 2:37 So at this point in time, writing code or making changes is something that teams are doing faster and faster.
- 2:43 Now I'll pause here.
- 2:46 Who here thinks that products have gotten better in the past 20 years?
- 2:49 Technical products, also raise your hand.
- 2:52 I hope everybody, right?
- 2:54 Phones are pretty cool.
- 2:54 How about five years?
- 2:58 How about past year?
- 3:01 So the challenge is that we're going faster and faster through the technology lifecycle, but we've lost something.
- 3:07 The product focused on customers in some way has degraded.
- 3:10 The maintainability of the code and the reliability has degraded.
- 3:13 You can see some of the up times here from two leading companies.
- 3:16 I've blurred out their names for, it doesn't actually matter who they are, but we are below a 3.9 or even a 4.9 reliability.
- 3:24 So even though we're shipping faster and faster, the code quality and the product quality has not necessarily gone up.
- 3:31 So let's talk about the refactor that we did.
- 3:33 So we started this refactor with the actual code implementation in April and did some pre-work earlier.
- 3:38 So I'll go through five different tasks that we did and share some of the findings that we had before and after, especially with as new models have come out.
- 3:46 So we spent around two months evaluating orchestrators for our AI pipeline.
- 3:50 We looked at five open source projects and we wanted to benchmark and see how effective they were for our use case.
- 3:56 And we started this off before deep research came out as part of Google and OpenAI so that web search capability to do a comprehensive analysis was still not there.
- 4:05 Now, after we actually gathered these requirements, we built out proof of concepts with a team of three to make sure that we actually got the right results.
- 4:12 Now, I'm pretty confident we could do this 90% faster now with the tooling that we have.
loading