Videos VrpEyglYgeU
In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar
Scene timeline
50 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 197
- whisperx 197
- chunks
- 32
- from 197 cues
- keyframes
- 47
- kept of 50 captured
- frames with text
- 47
- 1,699 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 7.8 MB
- word timings on 197 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 00:24 | 1m 48s |
stt |
done | — | 2026-08-10 00:25 | 23s |
chunk |
done | — | 2026-08-10 00:26 | 0s |
text_embed |
done | — | 2026-08-10 19:43 | 1s |
keyframe |
done | — | 2026-08-10 00:26 | 2m 11s |
ocr |
done | — | 2026-08-10 00:28 | 22s |
frame_embed |
done | — | 2026-08-10 19:43 | 8s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.97
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.98
- OpenAl0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust1.00
- bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- neo4j0.98
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto0.99
- Sonar1.00
- Makers of0.98
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- Sonar1.00
- Makers of1.00
- SonarQube1.00
- AIE0.97
-
- TARIQ SHAUKAT0.97
- CHIEF EXECUTIVE OFFICER1.00
- Sonar1.00
- Makers of1.00
- SonarQube1.00
- AIE1.00
-
- TARIQSHAUKAT0.99
- ot0.58
- Make0.83
- promptql0.85
- ZAY0.73
- ineer1.00
- 'sFair0.99
- CHIEF EXECUTIVE OFFICER1.00
- Sonar1.00
- Ckeric0.85
- Makers of0.99
- SonarQube1.00
- AIE1.00
- ZAI0.98
- Ravenina0.94
- brld's Fai0.75
- BeNCORD0.88
-
- World's Fair0.97
- Engineer1.00
- World's Fa0.94
- air0.96
- THE VELOCITY ROOM0.97
- beigha data0.68
- OpenAl1.00
- INNGEST0.95
- World'sFair1.00
- World's Fair0.97
- © BAND0.89
- AUTOMATTIC1.00
- PRIOR0.98
- Z.AI0.98
- World's Fair0.96
- Keycard0.98
- :neo4j0.99
- World's Fair0.98
- Cleric1.00
- Meticdous0.97
- World's Fair0.97
- qodo0.96
- GGRAMITEE0.89
- World's Fair0.97
- Google DeepMind1.00
- World's Fai0.98
- A ATLASSIAN0.92
- World's Fair0.91
- Amazon AGI Lab1.00
- World's Fai0.93
- World's Fai0.94
- arize0.98
- PayPal0.99
- Gradium0.95
- Braintrust1.00
- World'sFai1.00
- Microsoft1.00
- World'sFa0.95
- arlucto0.83
- ANTHROPC0.96
- OpenAl0.96
- ORAC1.00
- Microsoft0.99
- World's Fair0.97
- World'sFai0.99
- MINIMAX1.00
- brighe data0.81
- World's Fai0.93
- Akamai1.00
- World'sFai0.95
- World's Fair0.87
- ANTHROPIC0.96
- World's Fair0.97
- DATADOG1.00
- World's Fai0.96
- Resolve.ai1.00
- @Arbyte0.92
- DATAD1.00
- Optiver0.99
- World's Fair0.98
- builder.io0.97
- Worid'sFai0.93
- Ravenna1.00
- World'sFa1.00
- ilotKit1.00
- World'sFair1.00
- cognee1.00
- World'sFa0.96
- BeNCORD0.91
- tgrs0.76
-
- orld's Far0.91
- ORACLE1.00
- World's Fair0.95
- arize1.00
- World's Fair0.98
- Google DeepMind1.00
- World's Fair0.97
- :neo4j0.91
- World's Fair0.95
- Z.AI0.98
- World's Fair0.95
- bright data1.00
- World's Fair0.98
- Browserbo0.99
- per compute co.0.97
- World's Fair0.96
- extend1.00
- World's Fair0.95
- vast.ai1.00
- World's Fair0.97
- Ref.1.00
- World's Fair0.91
- Red Hat0.93
- World's Far0.98
- mezmo+0.91
- World's Fair0.98
- stigg1.00
- World's Fal0.92
- orld'sFair0.80
- RELAI0.98
- World's Fair0.97
- VAPI1.00
- World's Fair0.99
- Modal1.00
- World's Far0.98
- promptql1.00
- World's Fair0.94
- THE VELOCITY ROOM0.95
- World's Fair0.97
- fiddler1.00
- World's Fair0.99
- Supercondu1.00
- Surreal0.99
- World's Fair0.98
- ZERO0.99
- World's Fair0.96
- AIEngineer0.96
- comet1.00
- World's Fair0.98
- authe0.99
- World's Fal0.94
- orld's Fair0.91
- SOIOIO0.88
- World's Fair0.98
- granica1.00
- World's Far0.92
- dash00.98
- World's Fair0.97
- PRIOR1.00
- World's Fair1.00
- igitalOcean1.00
- World's Fair0.97
- Vence0.98
- World's Fair0.96
- POSTMAN1.00
- World's Fair0.98
- Composio1.00
- World's Fal0.92
- orld's Fair0.94
- Modular1.00
- World's Fair0.99
- MERGE1.00
- World's Fair0.94
- AUTOMATTIC1.00
- World's Fair0.95
- Buildki1.00
- ugabyteDB1.00
- World's Fair0.97
- Zed1.00
- World's Fair0.96
- Keycard1.00
- World's Fair0.94
- Meticulous1.00
- World's Fai0.97
- lorld's Fair0.91
- Daytona1.00
- World's Fair0.95
- twilio0.92
- World's Fair0.92
- GRAVITEE1.00
- World's Fair0.96
- PlanetSca0.95
- Temporal1.00
- World's Fair0.98
- Llamalndex0.99
- World's Fair0.99
- INNGEST0.99
- World's Fair0.99
- BAND1.00
- World's Fair0.96
- Cleric1.00
- World's Far0.92
- A ATLASSIAN0.95
- World's Fair0.97
- FACTORY1.00
- World's Fai0.94
- orld's Fair0.94
- baseten1.00
- World's Fair0.99
- Snorkel0.91
- World's Fair0.97
- Z.AI0.99
- World's Fair0.97
- qodo0.97
- World's Fair0.96
- PayPal1.00
- World's Fair0.98
- Gradium1.00
- World's Fair0.91
- ANTHROP1.00
- zon AGI Lab0.99
- World's Fair0.98
- Browserbase1.00
- Word's Fair0.88
- :neo4j0.92
- World's Fair0.93
- Google DeepMind1.00
- World's Fair0.95
- arize1.00
- World's Fair0.93
- reducto1.00
- World's Fair0.98
- Microsoft0.98
- World's Fal0.92
- orld's Fair0.88
- OpenAl0.98
- World's Fair0.96
- WorkOS0.88
- World's Fair0.96
- Amazon AGI Lab0.97
- World's Fair0.96
- Microsoft1.00
- World's Fair0.98
- ORACLE1.00
- World's Fair0.94
- bright data1.00
- World's Fair0.95
- Google Deeph0.98
- Microsoft0.96
- World's Fair0.95
- docker1.00
- World's Fair0.99
- Braintrust1.00
- World's Far0.98
- Oper1.00
- World's Fair0.99
- MINIMAX0.99
- World's Fair0.97
- Z.AI0.98
- World's Fair0.96
- ANTHROPIC1.00
- World's Fa0.99
- orld's Fair0.90
- snyk1.00
- World's Fair0.98
- aws0.99
- World's Fair0.96
- togetherai0.96
- Wor0.97
- Akamai1.00
- Word's Fair0.90
- Unblocked0.97
- World0.97
- Lightrun1.00
- ain0.98
- World's Fair0.97
- World's Fair0.96
- DA0.79
- World's Fair0.97
- Resolve.ai0.99
- World's Far0.92
- World's Fa0.94
- Op0.96
- WorldsFa0.99
- orld's Fair0.89
- World's Far0.97
- LanceDB1.00
- World's Fair0.95
- builderio0.96
- Worl0.97
- Ravenna1.00
- World's Fair0.98
- CopilotKit1.00
- VERIS1.00
- Wor0.99
- escupe0.93
- Microsoft1.00
- t1.00
- Word's Far0.84
- TOPK1.00
- World's Fair0.98
- cr0.64
- e0.73
- World's Fair0.92
- ENCORD0.92
- World's Fair0.98
- rid's Fa0.83
- Worid's Fa0.92
-
- AlEngineer0.96
- AI0.94
- World's1.00
- AlEngir0.96
- Fair1.00
- World'1.00
- Akan0.99
- AlFnai0.72
- Reso0.98
- DOG1.00
- Worlu1.00
-
- AlEngineer0.99
- OpenAl0.92
- World's Fair0.96
- AlEngineer1.00
- World's Fair0.97
- Akamai1.00
- AlEngineer0.99
- DATADOG1.00
- Jorld's Fair0.99
-
- World's Fair0.96
- World's Fal0.97
- World's Fair0.99
- ted Hat0.88
- MERGE1.00
- World's Fair0.99
- THIE VELOCITY ROOM0.93
- mezmo{0.93
- ©fiddler0.91
- World's Fair0.97
- INNGEST1.00
- PRIOR1.00
- SROK0.63
- World's Fair0.95
- AUTOMATTIC1.00
- World's Fair0.98
- ©BAND0.91
- Keycard1.00
- Bterupx0.59
- Worid's Fair0.92
- World's Fai0.96
- World's Fair0.98
- Cleric1.00
- GRAVITEE0.91
- Meticulous0.97
- Kexa0.84
- World's Fair0.98
- Amazon AGI Lab0.99
- World'sFai1.00
- World's Fa0.96
- qodo0.96
- arize1.00
- PayPal1.00
- À ATLASSIAN0.98
- Gradium0.99
- FACTORY1.00
- PareSode0.59
- dnid0.58
- diy0.57
- orld's Fair0.99
- World's Fair0.98
- Braintrust1.00
- World's Fair0.96
- World's Fa0.96
- World'sFai1.00
- OpenAl0.96
- World's Fal0.91
- Microsoft0.99
- World's Fal0.89
- ACLE1.00
- reducto1.00
- Z.AI0.98
- bright data0.99
- ANTHROPIC1.00
- Microsoft0.99
- ANTHROPIC0.99
- Google DepMind0.88
- Crusce0.85
- OpenAl0.91
- rld'sFair0.99
- World's Fair0.97
- @Artbyte0.87
- gruptie0.65
- Res0.96
- CODERI0.97
- Optiver0.99
- builder.io1.00
- Worid'sFa0.94
- rlid'sFair0.78
- d's Fair0.93
- World'sFair1.00
- Worid'sFai0.93
- World'sFa0.97
- Workd's Fai0.91
- tigus0.70
-
- AlEngineer0.96
- OpenAI0.92
- World's Fair1.00
- AlEngineer0.99
- together.ai0.99
- orld's Fair0.98
- AlEngineer0.97
- World's Fair0.97
- DATADOG1.00
-
- AlEngineer0.98
- OpenAI0.92
- World's Fair0.97
- AlEngineer0.99
- Wor'sF:0.92
- Akamai1.00
- AlEngineer0.98
- DATADO1.00
- World's Fair0.97
-
- AlEngineer0.98
- World'sFair1.00
- CNN Business0.93
- Markets0.98
- Tech Media0.99
- Calculators1.00
- Videos1.00
- • Watch0.89
- Listen0.90
- Another 'hallucinated' court filing highlights the0.99
- difference between Silicon Valley and the rest of0.99
- the world0.97
- APR 23, 20260.85
- SULLIVAN1.00
- CROMWELL1.00
- 80.89
- d's Fair0.91
- ngineer1.00
- Ope1.00
- Sonar e20260.96
- - AlEngi0.96
- gether0.99
- rld1.00
- In the Land of Al Agents, the Verifiers Are King0.99
- ngineer1.00
- d'sFair1.00
- DAT1.00
- Tariq Shaukat / Chief Executive Officer0.97
- Sonar1.00
- Makers of0.93
-
- AlEngineer0.99
- World's Fair0.98
- Every industry is1.00
- struggling with AI slop.1.00
- AIEn0.93
- enAl0.99
- Worlc1.00
- Sonar e20260.97
- Engineer1.00
- Id's0.95
- Ak1.00
- In the Land of Al Agents, the Verifiers Are King1.00
- AIEn0.87
- ATADOr0.97
- Worlc0.97
- Tariq Shaukat / Chief Executive Officer0.99
- Sonar1.00
- Makers of0.99
-
- AlEngineer0.99
- World's Fair0.96
- But, isn't0.99
- software development1.00
- different?1.00
- AIE1.00
- penAl0.91
- Worl1.00
- Sonar e20260.97
- AlEngineer0.95
- ld'sFai0.93
- In the Land of Al Agents, the Verifiers Are King1.00
- AIE0.99
- ATADOG1.00
- Worl1.00
- Tariq Shaukat / Chief Executive Officer0.99
- Sonar1.00
- Makers of0.99
-
- Coding agents are quickly getting a lot better0.99
- AlEngineer0.99
- World'sFair1.00
- Time horizon of software tasks Al can complete0.97
- at 50% accuracy1.00
- 50% success rate1.00
- Claude Mythos0.98
- Preview (early)0.99
- 16 hours1.00
- Tas u (ns)0.76
- Claude Opus 4.61.00
- 12 hours1.00
- GPT-5.2 (high)1.00
- 8 hours1.00
- Gemini 3.1 Pro1.00
- 6 hours0.99
- Claude Opus 4.51.00
- GPT-51.00
- 4 hours1.00
- 030.97
- 2 hours1.00
- 1 hour1.00
- GPT-21.00
- GPT-31.00
- GPT-3.51.00
- GPT-41.00
- OpenAI0.93
- W0.99
- 01.00
- 20191.00
- 20201.00
- 20211.00
- 20221.00
- 20231.00
- 20241.00
- 20250.98
- 20261.00
- 20271.00
- )Sonar e20280.88
- LLM release date1.00
- Source: METR: Task-Compietion Time Horizons of Frontier AI Models0.98
- AlEngineer1.00
- Vorld'sl0.92
- In the Land of Al Agents, the Verifiers Are King1.00
- DATAD1.00
- W0.99
- Tariq Shaukat / Chief Executive Officer0.98
- Sonar1.00
- Makers of0.99
-
- Coding agents are quickly getting a lot better1.00
- AlEngineer1.00
- World'sFair1.00
- at 50% success rate0.99
- Time horizon of software tasks Al can complete0.98
- at 50% accuracy0.98
- 50% success rate1.00
- PRESENTED BY0.99
- Preview (early)1.00
- Claude Mythos0.99
- 16 hours1.00
- Microsoft1.00
- Tas u (tns)0.76
- Claude Opus 4.60.97
- 12 hours1.00
- GPT-5.2 (high)1.00
- 8 hours1.00
- Gemini 3.1 Pro1.00
- 6 hours0.99
- Claude Opus 4.51.00
- GPT-51.00
- 4 hours1.00
- 030.98
- 2 hours1.00
- 1 hour1.00
- GPT-21.00
- GPT-31.00
- GPT-3.51.00
- GPT-41.00
- :air0.90
- UMINIM0.97
- 01.00
- 20191.00
- 20201.00
- 20211.00
- 20221.00
- 20231.00
- 20241.00
The page's on-screen-text budget of 600
lines is spent, so the last cards in this grid list fewer lines than they
hold. Narrow the page with ?frames= to read them.
Transcript
197 cues· 2,917 words· 16,153 chars
- 0:12 Please join me in welcoming the Chief Executive Officer at Sonar, Tariq Shawkat.
- 0:34 Morning, everyone.
- 0:36 Did you enjoy that last talk?
- 0:37 That was amazing.
- 0:39 You particularly loved the end, the being unreasonable part.
- 0:42 I thought that was awesome.
- 0:44 I also wanna just, I'm trying to calculate the odds of Tarek following Tarek as the first two sessions in the morning.
- 0:51 I think the odds are pretty low on this one.
- 0:54 But thrilled to be here today.
- 0:56 As we just mentioned, I am with Sonar.
- 0:59 We are in the code verification space, and I'm here today to talk about verification.
- 1:04 And I think we're all here in large part because we believe to some extent that AGI is here, it's coming.
- 1:12 The models we just heard about, Fable, it's really incredible what is going on in the world today.
- 1:18 And yet we work almost exclusively with enterprises around the world and the conversation that we have more is the question mark version.
- 1:27 Is AGI here and why are they asking these questions?
- 1:31 It's because you can,
- 1:33 read the news every day, and I'm not trying to name and shame here, but if you look at KPMG putting out reports that they have to retract because of hallucinations, EY doing the same thing, law firms getting into lots and lots of trouble because of made up citations, made up case law, things like this.
- 1:56 I think we can really start to question, how do we get value out of AI?
- 2:01 The models are amazing, as we just heard, but the hard part, as the other Tarek just said, is getting value out of it.
- 2:09 The struggle is that AI slop is everywhere.
- 2:14 I'm sure you all see this inside of your organizations.
- 2:16 I'm sure you see this in your everyday life, that AI is amazing.
- 2:21 The models are incredible at generating very plausible output.
- 2:25 They're incredible at generating things that sound correct, but are they correct?
- 2:30 And how do you know that they're correct is a big problem.
- 2:33 And it's a big problem in professional services, as we saw.
- 2:36 It's a big problem in legal.
- 2:38 But really, I think if we're honest, it's a big problem in every sector, in every field, whether it's marketing or finance or you name it.
- 2:46 You have this question of how do you actually know if it's true?
- 2:50 How do you know if it's good or if it is slop?
- 2:53 And the question that we deal in the coding space in particular, we deal with software development.
- 3:00 And the question we get as we talk to I'm sure many of the people here in the room and a lot of our customers is isn't software development different?
- 3:09 And we can look at the data on this and the mythos models.
- 3:15 This is data from meter.
- 3:18 You may have seen this METR.
- 3:20 The coding agents are getting better very quickly.
- 3:23 They're getting a lot better very quickly.
- 3:25 And you can see the progression, the exponential curve here.
- 3:28 What this shows on this chart is how capable are the models at completing tasks
- 3:34 that humans would take.
- 3:36 So can they complete a task that takes one hour, two hours, whatever it is?
- 3:40 The latest mythos model, at least per the benchmarking, which was done a month or so ago in the preview mode, was you're getting to 16 to 18 hours.
- 3:49 So the agents are able to complete long-running tasks.
- 3:54 And it really is starting to transform how work is happening.
- 3:59 But the critical caveat when you read the data is this is at a 50% success rate.
- 4:05 Okay, so it is again able to complete tasks, but is it able to complete tasks correctly is the question.
- 4:13 So if you start looking at, all right, let's dial up.
- 4:16 the accuracy rate, you dial it up to 80%, and there's still progress, but it is much slower progress.
- 4:22 Instead of 18 hours, you're at about three and a half hours or something along these lines.
- 4:27 And by the way, this is still at 80% accuracy.
loading
Chapters
- 0:00 Introduction and the current state of AI adoption
- 1:30 The challenge: Distinguishing AI utility from "AI slop"
- 3:09 Analyzing the performance data of AI coding agents
- 6:17 The productivity paradox: Why gains dissipate after three months
- 8:28 Introducing the AC/DC (Agent-Centric Development Cycle) framework
- 9:31 Stage 1: Guide (Providing context and constraints)
- 11:22 Stage 2: Verify (Zero-trust, multi-layered verification)
- 13:00 Stage 3: Solve (Maintenance loops and technical debt control)
- 14:32 The necessity of systems-level thinking for AI agents
- 16:56 Real-world impact: 92% reduction in issues with disciplined verification
- 17:42 Conclusion and final thoughts on enterprise AI