read-only demo

Videos 03l29gJXpCE

Guide, Verify, Solve — Anirban Chatterjee, Sonar

index_state ready data_status ok

AI Engineer· published 2026-08-09· 0:22:31· en-US· indexed 2026-08-11 00:31

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:29, 1 of 1 keyframes kept
  5. Shot 4, 0:29 to 1:03, 1 of 1 keyframes kept
  6. Shot 5, 1:03 to 1:37, 1 of 1 keyframes kept
  7. Shot 6, 1:37 to 2:11, 1 of 1 keyframes kept
  8. Shot 7, 2:11 to 2:45, 0 of 1 keyframes kept
  9. Shot 8, 2:45 to 3:11, 1 of 1 keyframes kept
  10. Shot 9, 3:11 to 3:37, 0 of 1 keyframes kept
  11. Shot 10, 3:37 to 4:03, 1 of 1 keyframes kept
  12. Shot 11, 4:03 to 4:29, 0 of 1 keyframes kept
  13. Shot 12, 4:29 to 4:56, 0 of 1 keyframes kept
  14. Shot 13, 4:56 to 5:22, 1 of 1 keyframes kept
  15. Shot 14, 5:22 to 5:48, 0 of 1 keyframes kept
  16. Shot 15, 5:48 to 6:14, 0 of 1 keyframes kept
  17. Shot 16, 6:14 to 6:40, 1 of 1 keyframes kept
  18. Shot 17, 6:40 to 7:07, 0 of 1 keyframes kept
  19. Shot 18, 7:07 to 7:33, 0 of 1 keyframes kept
  20. Shot 19, 7:33 to 7:59, 0 of 1 keyframes kept
  21. Shot 20, 7:59 to 8:25, 1 of 1 keyframes kept
  22. Shot 21, 8:25 to 8:51, 0 of 1 keyframes kept
  23. Shot 22, 8:51 to 9:17, 0 of 1 keyframes kept
  24. Shot 23, 9:17 to 9:44, 1 of 1 keyframes kept
  25. Shot 24, 9:44 to 10:10, 0 of 1 keyframes kept
  26. Shot 25, 10:10 to 10:36, 1 of 1 keyframes kept
  27. Shot 26, 10:36 to 11:04, 1 of 1 keyframes kept
  28. Shot 27, 11:04 to 11:32, 0 of 1 keyframes kept
  29. Shot 28, 11:32 to 12:00, 0 of 1 keyframes kept
  30. Shot 29, 12:00 to 12:28, 1 of 1 keyframes kept
  31. Shot 30, 12:28 to 12:56, 0 of 1 keyframes kept
  32. Shot 31, 12:56 to 13:24, 0 of 1 keyframes kept
  33. Shot 32, 13:24 to 13:52, 0 of 1 keyframes kept
  34. Shot 33, 13:52 to 14:20, 0 of 1 keyframes kept
  35. Shot 34, 14:20 to 14:48, 0 of 1 keyframes kept
  36. Shot 35, 14:48 to 15:16, 0 of 1 keyframes kept
  37. Shot 36, 15:16 to 15:44, 1 of 1 keyframes kept
  38. Shot 37, 15:44 to 16:13, 1 of 1 keyframes kept
  39. Shot 38, 16:13 to 16:41, 0 of 1 keyframes kept
  40. Shot 39, 16:41 to 17:10, 1 of 1 keyframes kept
  41. Shot 40, 17:10 to 17:38, 0 of 1 keyframes kept
  42. Shot 41, 17:38 to 18:07, 0 of 1 keyframes kept
  43. Shot 42, 18:07 to 18:35, 0 of 1 keyframes kept
  44. Shot 43, 18:35 to 19:04, 0 of 1 keyframes kept
  45. Shot 44, 19:04 to 19:06, 1 of 1 keyframes kept
  46. Shot 45, 19:06 to 19:08, 0 of 1 keyframes kept
  47. Shot 46, 19:08 to 19:41, 1 of 1 keyframes kept
  48. Shot 47, 19:41 to 20:15, 0 of 1 keyframes kept
  49. Shot 48, 20:15 to 20:34, 1 of 1 keyframes kept
  50. Shot 49, 20:34 to 20:50, 1 of 1 keyframes kept
  51. Shot 50, 20:50 to 21:12, 1 of 1 keyframes kept
  52. Shot 51, 21:12 to 21:43, 1 of 1 keyframes kept
  53. Shot 52, 21:43 to 22:14, 0 of 1 keyframes kept
  54. Shot 53, 22:14 to 22:29, 0 of 1 keyframes kept
  55. Shot 54, 22:29 to 22:30, 1 of 1 keyframes kept

55 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
229
whisperx 229
chunks
39
from 229 cues
keyframes
26
kept of 55 captured
frames with text
25
521 lines read
chapters
0
from the source metadata
keyframe bytes
7.2 MB
word timings on 229 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 00:27 1m 00s
stt done 2026-08-11 00:28 26s
chunk done 2026-08-11 00:29 0s
text_embed done 2026-08-11 00:29 1s
keyframe done 2026-08-11 00:29 2m 12s
ocr done 2026-08-11 00:31 11s
frame_embed done 2026-08-11 00:31 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 453.7

    1. AIEngineer0.95
    2. World's Fair0.96
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 659.6

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2733.7

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:24 #3 done3 line(s)

    shot 3·sharpness 206.7

    1. Sonar1.00
    2. AlEngineer0.98
    3. World's Fair1.00
  • 0:59 #4 done13 line(s)

    shot 4·sharpness 1852.8

    1. AlEngineer0.98
    2. AlEngineer0.99
    3. World'sFair1.00
    4. World'sFair1.00
    5. Guide, Verify, Solve1.00
    6. PRESENTED BY1.00
    7. The engineering discipline1.00
    8. Microsoft1.00
    9. agentic development demands1.00
    10. Anirban Chatterjee | Director of Product Marketing1.00
    11. Sonar1.00
    12. TRACK81.00
    13. sFair1.00
  • 1:07 #5 done14 line(s)

    shot 5·sharpness 1889.7

    1. AlEngineer0.98
    2. AlEngineer1.00
    3. World'sFair1.00
    4. World'sFair1.00
    5. Guide, Verify, Solve0.99
    6. PRESENTED BY1.00
    7. The engineering discipline1.00
    8. Microsoft1.00
    9. agentic development demands1.00
    10. Anirban Chatterjee | Director of Product Marketing1.00
    11. Sonar1.00
    12. TRACK 8• JULY 2, 20260.94
    13. sFair1.00
    14. Agentic Engineering1.00
  • 2:07 #6 done41 line(s)

    shot 6·sharpness 3371.2

    1. AlEngineer0.97
    2. Al tools are great ... but they can introduce risk0.99
    3. World'sFair1.00
    4. Carnegie Mellon researchers studied projects that adopted Cursor and measured1.00
    5. the impact on code quality using SonarQube. They saw:1.00
    6. A temporary 3-5x velocity spike that disappears by the third month of usage.0.99
    7. Commits1.00
    8. Lines Added1.00
    9. Static Analysis Wamings0.99
    10. Duplicated Lines Density1.00
    11. Code Complexity1.00
    12. 1.01.00
    13. Treat tet0.70
    14. 1.01.00
    15. 0.51.00
    16. 1.51.00
    17. 21.00
    18. 0.51.00
    19. 0.501.00
    20. 0.251.00
    21. 0.51.00
    22. 1.01.00
    23. 0.01.00
    24. 0.001.00
    25. 0.01.00
    26. -0.51.00
    27. -0.251.00
    28. -0.51.00
    29. -6-5-4-3-2-1 0 1 2 3 4 5 60.97
    30. -6-5-4-3-2-1 0 1 2 3 4 5 60.94
    31. Months Relative to Cursor Adoption0.99
    32. -6-5-4-3-2-1 0 1 2 3 4 5 60.95
    33. -6-5-4-3-2-1 0 1 2 3 4 5 60.94
    34. -6-5-4-3-2-1 0 1 2 3 4 5 60.97
    35. Filled dots: p < 0.05 (significant), Hollow dots: p ≥ 0.05 (non-significant)0.98
    36. ©2026, SonarSource Sàrl0.98
    37. Source: Does Al-Assisted Coding Deliver? A Difference-in-Differences0.99
    38. Study of Cursor's Impact on Software Projects0.99
    39. TRACK 8· JULY 2, 20260.95
    40. sFair1.00
    41. Agentic Engineering1.00
  • 2:37 #7 skipped

    shot 7·duplicate of #6

  • 2:53 #8 done19 line(s)

    shot 8·sharpness 2148.7

    1. AlEngineer0.98
    2. The growing quality gap1.00
    3. World'sFair1.00
    4. Needed quality level1.00
    5. Reats ti uiard0.72
    6. Agent default quality level1.00
    7. Software criticality0.98
    8. Internal / non-critical0.99
    9. Enterprise / mission-critical0.98
    10. Simple < 50K LOCs1.00
    11. Complex < 500K LOCs1.00
    12. Short lived, will run for months not years0.99
    13. Long lived, will run for years1.00
    14. Limited number of users, all friendly0.98
    15. Many, potentially adversarial users1.00
    16. ©2026, SonarSource Sàrl0.99
    17. TRACK 8• JULY 2, 20260.96
    18. sFair1.00
    19. Agentic Engineering1.00
  • 3:21 #9 skipped

    shot 9·duplicate of #8

  • 4:00 #10 done25 line(s)

    shot 10·sharpness 3055.5

    1. AlEngineer0.98
    2. Why is this happening?0.94
    3. World's Fair0.98
    4. 0.93
    5. The models are0.97
    6. but they make1.00
    7. and they are1.00
    8. smart1.00
    9. mistakes1.00
    10. missing context1.00
    11. The latest coding models are1.00
    12. They are error-prone, and any1.00
    13. They do not understand your1.00
    14. extremely intelligent and can1.00
    15. of their mistakes could prove0.97
    16. codebase, your context, or0.98
    17. do incredible work.1.00
    18. catastrophic for your1.00
    19. your objectives.1.00
    20. organization.0.99
    21. ©2026, SonarSource Sàrl0.97
    22. Source: Measuring and mitigating overreliance is necessary for building human-compatible Al0.99
    23. TRACK 8· JULY 2,20260.96
    24. sFair1.00
    25. Agentic Engineering1.00
  • 4:06 #11 skipped

    shot 11·duplicate of #10

  • 4:40 #12 skipped

    shot 12·duplicate of #10

  • 5:16 #13 done28 line(s)

    shot 13·sharpness 2353.1

    1. AlEngineer0.98
    2. Models have diverse quality issues0.99
    3. World's Fair0.98
    4. To understand them, we benchmark the code that models produce on 4000+ tasks.0.98
    5. 82.5 | 81.8%0.97
    6. Correctness1.00
    7. 132.1124.8/KLOC1.00
    8. 0.23 | 1.17%0.94
    9. Low Complexity1.00
    10. Unsolved1.00
    11. Tasks1.00
    12. Maintainability1.00
    13. Security1.00
    14. 1678719312/MLOC1.00
    15. 154 | 305/MLOC0.98
    16. Reliability1.00
    17. 680|698/MLOC0.97
    18. Claude Opus 4.61.00
    19. Claude Sonnet 4.61.00
    20. Best per0.97
    21. Thinking1.00
    22. Thinking1.00
    23. metric1.00
    24. sonar.com/leaderboard1.00
    25. ©2026, SonarSource Sàrl0.99
    26. TRACK 8· JULY 2, 20260.95
    27. sFair1.00
    28. Agentic Engineering1.00
  • 5:37 #14 skipped

    shot 14·duplicate of #13

  • 6:09 #15 skipped

    shot 15·duplicate of #13

  • 6:27 #16 done24 line(s)

    shot 16·sharpness 3553.9

    1. AlEngineer0.99
    2. Human review is compromised1.00
    3. World'sFair1.00
    4. Shaw and Nave (Wharton, 2026) ran trials measuring1.00
    5. "Whereas cognitive offloading is0.99
    6. what people actually do when they consult Al.1.00
    7. PRESENTED BY1.00
    8. a strategic delegation of0.98
    9. deliberation, using a tool to aid0.99
    10. Microsoft1.00
    11. one's own reasoning, cognitive1.00
    12. Participants followed Al advice:1.00
    13. surrender is an uncritical0.99
    14. 92.7%1.00
    15. abdication of reasoning itself."1.00
    16. of the time when the Al was correct0.98
    17. Source: Shaw & Nave (2026), Wharton1.00
    18. — "Thinking Fast, Slow, and Artificial"0.99
    19. 79.8%1.00
    20. of the time when the Al was wrong0.99
    21. ©2026, SonarSource Sàrl0.98
    22. TRACK 8• JULY 2, 20260.97
    23. sFair1.00
    24. Agentic Engineering1.00
  • 6:58 #17 skipped

    shot 17·duplicate of #16

  • 7:10 #18 skipped

    shot 18·duplicate of #16

  • 7:36 #19 skipped

    shot 19·duplicate of #16

  • 8:05 #20 done9 line(s)

    shot 20·sharpness 1255.8

    1. AlEngineer0.98
    2. World'sFair1.00
    3. Code is provable.1.00
    4. Software is not.1.00
    5. Verification is the key to Al success.0.99
    6. ©2026, SonarSource Sàrl0.98
    7. TRACK 8• JULY 2, 20260.96
    8. sFair1.00
    9. Agentic Engineering1.00
  • 8:33 #21 skipped

    shot 21·duplicate of #20

  • 9:12 #22 skipped

    shot 22·duplicate of #20

  • 9:28 #23 done17 line(s)

    shot 23·sharpness 2448.9

    1. AlEngineer0.97
    2. Verification should be zero trust and multilayered0.99
    3. World'sFair1.00
    4. Zero trust1.00
    5. Works the same no matter how the1.00
    6. code was written1.00
    7. Uses a different methodology than1.00
    8. what was used to generate the code1.00
    9. Has a clear segregation of duties1.00
    10. Completely auditable1.00
    11. Perfectly explainable0.99
    12. Algorithmic and repeatable1.00
    13. ©2026, SonarSource Sàrl0.98
    14. 121.00
    15. TRACK 8• JULY 2, 20260.97
    16. sFair1.00
    17. Agentic Engineering1.00

Transcript

229 cues· 4,555 words· 25,038 chars

  1. 0:13 All right.
  2. 0:17 Thank you, that's very helpful.
  3. 0:18 My name's Anirban Chatterjee.
  4. 0:19 I do product marketing at Sonar.
  5. 0:21 I'm really excited to be talking to this group today.
  6. 0:23 It's actually my first time here at this conference, and so I've been having a blast, along with the rest of my team here, meeting a whole bunch of AI engineers, as well as leaders,
  7. 0:35 you know, influencers and founders.
  8. 0:39 There's a lot going on in this space.
  9. 0:40 I think this year there's really been a turning point from experimentation to engineering, and that warms my heart very deeply because I started my career many, many, many, many, many years ago as a software engineer writing code for servers, if you can believe it.
  10. 0:59 And I think there's a turning point that's happening right now where we're starting to add the capabilities that we need to add to these systems in order to make them repeatable, make them scalable, make them consistent, much in the way we were doing with cloud computing not too long ago in order to expand the access that IT technology gave to small businesses and other innovators.
  11. 1:21 I think AI is going to do the same thing for software development going forward.
  12. 1:24 But in order to do that, in order to get there,
  13. 1:26 We need to start adding safety and trust to these systems so that they can be used more widely across a wide variety of use cases so we can build new things and solve bigger problems.
  14. 1:36 And how we get there is what we're going to talk about today.
  15. 1:39 And for those of you who were in Tarek's keynote yesterday, he presented some of this data, and I'm going to talk about it a little bit deeper today.
  16. 1:46 So there was a study that Carnegie Mellon did.
  17. 1:49 where they actually looked at projects that were posted on GitHub, and they were able to use the metadata to sort them into projects where they were just traditional tools that were being used, and projects where an AI tool was used to write the code.
  18. 2:03 In this case, it was Cursor, although it could have been any AI tool.
  19. 2:05 And what they found was interesting.
  20. 2:07 They found that there was, in fact, a temporary spike in productivity, but it lasted about three months, and then it went back down.
  21. 2:16 And the reason for that, we think, is because there was also a persistent increase in static analysis warnings and code complexity.
  22. 2:23 They're actually using SonarQube to actually collect the data on this.
  23. 2:27 And they saw that there was a persistent increase in these types of issues that went beyond the three-month mark and persisted well into the future.
  24. 2:34 And so it's these types of issues that end up actually slowing developers down even more.
  25. 2:39 And this is what makes it a challenge to deliver high-quality code using AI tools.
  26. 2:46 The reason for this is that there is a differing need for quality depending on the criticality of the application, right?
  27. 2:51 If you're experimenting, if you're playing around, if you're just one person building things to see what's possible, it's an internal non-critical application with just a few users, maybe it's just you, maybe it's a small team, maybe it's just a short-lived project that's not gonna last very long.
  28. 3:06 The gap between the quality that you're getting from the AI tool and the quality you need from the application is quite small.
  29. 3:12 And you can live with that gap.
  30. 3:14 But as you move to higher levels of criticality, as you run into situations where you're supporting many, many users, it's a larger code base with many lines of code and many changes happening across that code base all the time.
  31. 3:27 You have many, many users.
  32. 3:28 Some of them could be adversarial users actively trying to break your software.
  33. 3:32 And so in those cases, the quality level you need is quite a bit higher
  34. 3:37 and the quality level you're getting by default from these AI tools.
  35. 3:40 And that's where this verification debt comes in.
  36. 3:42 That's where you have to bring the humans in, bring your software engineers in to try to close that gap and make sure that the quality level is brought up to an acceptable level before you ship that code
  37. 3:54 into production.
  38. 3:55 So why is this happening?
  39. 3:56 Why is this gap actually occurring?
  40. 3:57 We know these models are excellent.
  41. 3:58 They're getting better and better all the time.
  42. 4:00 I'm really excited to start playing with Fable now that that's out to see what levels of code we can get out of Fable going forward.
  43. 4:08 But we do know that because of the technology, because of the way that these models are built, they will still make mistakes.
  44. 4:13 They will still have quality issues.
  45. 4:15 They are still somewhat error prone.
  46. 4:17 And if you let these errors go into production code, you could have a catastrophic effect
  47. 4:22 to your organization.
  48. 4:23 They're also missing context, right?
  49. 4:24 They only know what you tell it.
  50. 4:26 They don't know the broader things that are happening elsewhere in the code base.

Open at this second