Videos 03l29gJXpCE
Guide, Verify, Solve — Anirban Chatterjee, Sonar
Scene timeline
55 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 229
- whisperx 229
- chunks
- 39
- from 229 cues
- keyframes
- 26
- kept of 55 captured
- frames with text
- 25
- 521 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 7.2 MB
- word timings on 229 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:27 | 1m 00s |
stt |
done | — | 2026-08-11 00:28 | 26s |
chunk |
done | — | 2026-08-11 00:29 | 0s |
text_embed |
done | — | 2026-08-11 00:29 | 1s |
keyframe |
done | — | 2026-08-11 00:29 | 2m 12s |
ocr |
done | — | 2026-08-11 00:31 | 11s |
frame_embed |
done | — | 2026-08-11 00:31 | 5s |
Frames, and what the machine read
-
- AIEngineer0.95
- World's Fair0.96
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.94
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.99
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- Sonar1.00
- AlEngineer0.98
- World's Fair1.00
-
- AlEngineer0.98
- AlEngineer0.99
- World'sFair1.00
- World'sFair1.00
- Guide, Verify, Solve1.00
- PRESENTED BY1.00
- The engineering discipline1.00
- Microsoft1.00
- agentic development demands1.00
- Anirban Chatterjee | Director of Product Marketing1.00
- Sonar1.00
- TRACK81.00
- sFair1.00
-
- AlEngineer0.98
- AlEngineer1.00
- World'sFair1.00
- World'sFair1.00
- Guide, Verify, Solve0.99
- PRESENTED BY1.00
- The engineering discipline1.00
- Microsoft1.00
- agentic development demands1.00
- Anirban Chatterjee | Director of Product Marketing1.00
- Sonar1.00
- TRACK 8• JULY 2, 20260.94
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.97
- Al tools are great ... but they can introduce risk0.99
- World'sFair1.00
- Carnegie Mellon researchers studied projects that adopted Cursor and measured1.00
- the impact on code quality using SonarQube. They saw:1.00
- A temporary 3-5x velocity spike that disappears by the third month of usage.0.99
- Commits1.00
- Lines Added1.00
- Static Analysis Wamings0.99
- Duplicated Lines Density1.00
- Code Complexity1.00
- 1.01.00
- Treat tet0.70
- 1.01.00
- 0.51.00
- 1.51.00
- 21.00
- 0.51.00
- 0.501.00
- 0.251.00
- 0.51.00
- 1.01.00
- 0.01.00
- 0.001.00
- 0.01.00
- -0.51.00
- -0.251.00
- -0.51.00
- -6-5-4-3-2-1 0 1 2 3 4 5 60.97
- -6-5-4-3-2-1 0 1 2 3 4 5 60.94
- Months Relative to Cursor Adoption0.99
- -6-5-4-3-2-1 0 1 2 3 4 5 60.95
- -6-5-4-3-2-1 0 1 2 3 4 5 60.94
- -6-5-4-3-2-1 0 1 2 3 4 5 60.97
- Filled dots: p < 0.05 (significant), Hollow dots: p ≥ 0.05 (non-significant)0.98
- ©2026, SonarSource Sàrl0.98
- Source: Does Al-Assisted Coding Deliver? A Difference-in-Differences0.99
- Study of Cursor's Impact on Software Projects0.99
- TRACK 8· JULY 2, 20260.95
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- The growing quality gap1.00
- World'sFair1.00
- Needed quality level1.00
- Reats ti uiard0.72
- Agent default quality level1.00
- Software criticality0.98
- Internal / non-critical0.99
- Enterprise / mission-critical0.98
- Simple < 50K LOCs1.00
- Complex < 500K LOCs1.00
- Short lived, will run for months not years0.99
- Long lived, will run for years1.00
- Limited number of users, all friendly0.98
- Many, potentially adversarial users1.00
- ©2026, SonarSource Sàrl0.99
- TRACK 8• JULY 2, 20260.96
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- Why is this happening?0.94
- World's Fair0.98
- 氏0.93
- The models are0.97
- but they make1.00
- and they are1.00
- smart1.00
- mistakes1.00
- missing context1.00
- The latest coding models are1.00
- They are error-prone, and any1.00
- They do not understand your1.00
- extremely intelligent and can1.00
- of their mistakes could prove0.97
- codebase, your context, or0.98
- do incredible work.1.00
- catastrophic for your1.00
- your objectives.1.00
- organization.0.99
- ©2026, SonarSource Sàrl0.97
- Source: Measuring and mitigating overreliance is necessary for building human-compatible Al0.99
- TRACK 8· JULY 2,20260.96
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- Models have diverse quality issues0.99
- World's Fair0.98
- To understand them, we benchmark the code that models produce on 4000+ tasks.0.98
- 82.5 | 81.8%0.97
- Correctness1.00
- 132.1124.8/KLOC1.00
- 0.23 | 1.17%0.94
- Low Complexity1.00
- Unsolved1.00
- Tasks1.00
- Maintainability1.00
- Security1.00
- 1678719312/MLOC1.00
- 154 | 305/MLOC0.98
- Reliability1.00
- 680|698/MLOC0.97
- Claude Opus 4.61.00
- Claude Sonnet 4.61.00
- Best per0.97
- Thinking1.00
- Thinking1.00
- metric1.00
- sonar.com/leaderboard1.00
- ©2026, SonarSource Sàrl0.99
- TRACK 8· JULY 2, 20260.95
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.99
- Human review is compromised1.00
- World'sFair1.00
- Shaw and Nave (Wharton, 2026) ran trials measuring1.00
- "Whereas cognitive offloading is0.99
- what people actually do when they consult Al.1.00
- PRESENTED BY1.00
- a strategic delegation of0.98
- deliberation, using a tool to aid0.99
- Microsoft1.00
- one's own reasoning, cognitive1.00
- Participants followed Al advice:1.00
- surrender is an uncritical0.99
- 92.7%1.00
- abdication of reasoning itself."1.00
- of the time when the Al was correct0.98
- Source: Shaw & Nave (2026), Wharton1.00
- — "Thinking Fast, Slow, and Artificial"0.99
- 79.8%1.00
- of the time when the Al was wrong0.99
- ©2026, SonarSource Sàrl0.98
- TRACK 8• JULY 2, 20260.97
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.98
- World'sFair1.00
- Code is provable.1.00
- Software is not.1.00
- Verification is the key to Al success.0.99
- ©2026, SonarSource Sàrl0.98
- TRACK 8• JULY 2, 20260.96
- sFair1.00
- Agentic Engineering1.00
-
- AlEngineer0.97
- Verification should be zero trust and multilayered0.99
- World'sFair1.00
- Zero trust1.00
- Works the same no matter how the1.00
- code was written1.00
- Uses a different methodology than1.00
- what was used to generate the code1.00
- Has a clear segregation of duties1.00
- Completely auditable1.00
- Perfectly explainable0.99
- Algorithmic and repeatable1.00
- ©2026, SonarSource Sàrl0.98
- 121.00
- TRACK 8• JULY 2, 20260.97
- sFair1.00
- Agentic Engineering1.00
Transcript
229 cues· 4,555 words· 25,038 chars
- 0:13 All right.
- 0:17 Thank you, that's very helpful.
- 0:18 My name's Anirban Chatterjee.
- 0:19 I do product marketing at Sonar.
- 0:21 I'm really excited to be talking to this group today.
- 0:23 It's actually my first time here at this conference, and so I've been having a blast, along with the rest of my team here, meeting a whole bunch of AI engineers, as well as leaders,
- 0:35 you know, influencers and founders.
- 0:39 There's a lot going on in this space.
- 0:40 I think this year there's really been a turning point from experimentation to engineering, and that warms my heart very deeply because I started my career many, many, many, many, many years ago as a software engineer writing code for servers, if you can believe it.
- 0:59 And I think there's a turning point that's happening right now where we're starting to add the capabilities that we need to add to these systems in order to make them repeatable, make them scalable, make them consistent, much in the way we were doing with cloud computing not too long ago in order to expand the access that IT technology gave to small businesses and other innovators.
- 1:21 I think AI is going to do the same thing for software development going forward.
- 1:24 But in order to do that, in order to get there,
- 1:26 We need to start adding safety and trust to these systems so that they can be used more widely across a wide variety of use cases so we can build new things and solve bigger problems.
- 1:36 And how we get there is what we're going to talk about today.
- 1:39 And for those of you who were in Tarek's keynote yesterday, he presented some of this data, and I'm going to talk about it a little bit deeper today.
- 1:46 So there was a study that Carnegie Mellon did.
- 1:49 where they actually looked at projects that were posted on GitHub, and they were able to use the metadata to sort them into projects where they were just traditional tools that were being used, and projects where an AI tool was used to write the code.
- 2:03 In this case, it was Cursor, although it could have been any AI tool.
- 2:05 And what they found was interesting.
- 2:07 They found that there was, in fact, a temporary spike in productivity, but it lasted about three months, and then it went back down.
- 2:16 And the reason for that, we think, is because there was also a persistent increase in static analysis warnings and code complexity.
- 2:23 They're actually using SonarQube to actually collect the data on this.
- 2:27 And they saw that there was a persistent increase in these types of issues that went beyond the three-month mark and persisted well into the future.
- 2:34 And so it's these types of issues that end up actually slowing developers down even more.
- 2:39 And this is what makes it a challenge to deliver high-quality code using AI tools.
- 2:46 The reason for this is that there is a differing need for quality depending on the criticality of the application, right?
- 2:51 If you're experimenting, if you're playing around, if you're just one person building things to see what's possible, it's an internal non-critical application with just a few users, maybe it's just you, maybe it's a small team, maybe it's just a short-lived project that's not gonna last very long.
- 3:06 The gap between the quality that you're getting from the AI tool and the quality you need from the application is quite small.
- 3:12 And you can live with that gap.
- 3:14 But as you move to higher levels of criticality, as you run into situations where you're supporting many, many users, it's a larger code base with many lines of code and many changes happening across that code base all the time.
- 3:27 You have many, many users.
- 3:28 Some of them could be adversarial users actively trying to break your software.
- 3:32 And so in those cases, the quality level you need is quite a bit higher
- 3:37 and the quality level you're getting by default from these AI tools.
- 3:40 And that's where this verification debt comes in.
- 3:42 That's where you have to bring the humans in, bring your software engineers in to try to close that gap and make sure that the quality level is brought up to an acceptable level before you ship that code
- 3:54 into production.
- 3:55 So why is this happening?
- 3:56 Why is this gap actually occurring?
- 3:57 We know these models are excellent.
- 3:58 They're getting better and better all the time.
- 4:00 I'm really excited to start playing with Fable now that that's out to see what levels of code we can get out of Fable going forward.
- 4:08 But we do know that because of the technology, because of the way that these models are built, they will still make mistakes.
- 4:13 They will still have quality issues.
- 4:15 They are still somewhat error prone.
- 4:17 And if you let these errors go into production code, you could have a catastrophic effect
- 4:22 to your organization.
- 4:23 They're also missing context, right?
- 4:24 They only know what you tell it.
- 4:26 They don't know the broader things that are happening elsewhere in the code base.
loading