Videos MpZzWMdmQCE
Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com
Scene timeline
31 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 134
- whisperx 134
- chunks
- 17
- from 134 cues
- keyframes
- 31
- kept of 31 captured
- frames with text
- 31
- 336 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.5 MB
- word timings on 134 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 04:33 | 1m 06s |
stt |
done | — | 2026-08-11 04:34 | 11s |
chunk |
done | — | 2026-08-11 04:34 | 0s |
text_embed |
done | — | 2026-08-11 04:34 | 0s |
keyframe |
done | — | 2026-08-11 04:34 | 52s |
ocr |
done | — | 2026-08-11 04:35 | 10s |
frame_embed |
done | — | 2026-08-11 04:35 | 5s |
Frames, and what the machine read
-
- Al Engineer0.94
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.95
- WorkOS OpenAI0.95
-
- Vector1.00
-
- ktop/dev/vector-harmess/docs/presentation.html0.97
- AIE1.00
- Vector1.00
- ★1.00
- ★1.00
- ★1.00
- Intensity is vanity; Direction is sanity1.00
- Talha Sheikh1.00
- Google DeepMind1.00
-
- ***0.80
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Engineering the future of Al0.99
-
- The Lie Detector - Talha Sh0.95
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- @0.59
- AIE1.00
- You.The Enforcement.1.00
- ★1.00
- ★1.00
- The agent says 'done.' You check anyway.0.98
- Not because you want to — because nothing else does.1.00
- # Braintrust0.94
- WorkOS OpenAI0.95
-
- op/dev/vector-hamess/docs/presentation.html1.00
- Deterministic?1.00
- AIE1.00
- ITriedIt0.98
- ★0.99
- ★1.00
- AlEngineer0.96
- EUROPE1.00
-
- The Lie Detector - Talha Sh0.97
- ① File0.90
- /Users/talhashelkh/Desktop/dev/vector-haress/docs/presentation.html0.99
- ☆0.98
- @0.59
- THE IDEA1.00
- I Tried It0.97
- #.vector/config.yaml0.98
- checks:1.00
- test-pass:1.00
- run: "npm test"0.99
- AIE1.00
- retries: 30.97
- ★1.00
- type-check:1.00
- ★1.00
- ★1.00
- ★1.00
- run: "npx tsc --noEmit"0.99
- coverage:1.00
- run: "npx vitest --coverage --threshold 80"0.98
- retries: 20.95
- lint:0.98
- run: "npm run lint"1.00
- vectors:1.00
- v1:0.99
- checks: [test-pass, type-check, coverage, lint]1.00
- Every check is a shell command. Exit 0 = pass.0.98
- AlEngineer0.97
- EUROPE1.00
-
- The Lie Detector -Talha Sh0.96
- h/Desktop/dev/vector-harness/docs/presentation.html0.98
- AIE1.00
- Trust1.00
- ★0.92
- ★1.00
- ★1.00
- Not because the Al got smarter → because something verified it.0.99
- Then I made the mistake of telling people.0.99
- Engineering the future of Al0.99
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- 00.57
- Trust@0.92
- AIE1.00
- The PushbackX0.96
- 'We're making models better. You won't need this.'1.00
- That shook me.0.97
- Engineering the future of Al1.00
-
- esktop/dev/vector-hamess/docs/presentation.html0.98
- @0.63
- Crisis mo0.95
- AIE1.00
- Project1.00
- ★0.97
- ★1.00
- ★0.99
- ★1.00
- Glasswing1.00
- Securing critical software1.00
- for the AI era0.99
- Continue reading0.99
- Engineering the future of Al0.98
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.99
- Project1.00
- Glasswing1.00
- AIE1.00
- Capability ≠ Reliability0.98
- Engineering the future of Al0.98
- ineer1.00
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- Capability≠Reliability0.99
- AIE1.00
- Instructions ≠ Verification1.00
- AlEngineer0.97
- EUROPE1.00
- er1.00
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- Instructions ≠ Verification0.99
- AIE1.00
- ★0.95
- Guardrails > Model Size0.97
- ★1.00
- ★1.00
- AlEngineer0.96
- EUROPE1.00
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- ThenIHeardIt0.99
- AIE1.00
- Everywhere1.00
- Engineering the future of Al0.99
-
- ktop/dev/vector-haress/docs/presentation.html0.98
- ThenIHeard It0.96
- Everywhere1.00
- AIE1.00
- ★1.00
- ★1.00
- Engineering the future of Al0.99
-
- The Lie Detector -Talha Sh0.96
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.98
- AIE1.00
- The Real Question0.99
- Q0.51
- Engineering the future of Al0.99
-
- The Lie Detector - Talha Sh0.96
- ① File0.87
- /Users/talhashelkh/Desktop/dev/vector-haress/docs/presentation.html0.99
- ☆0.98
- The Pattern0.99
- THE PATTERN1.00
- Verification at Every Layer1.00
- AIE1.00
- Layer1.00
- Check Point1.00
- Checks1.00
- ★1.00
- Conversation1.00
- on-chat1.00
- tests, types, coverage0.99
- ★1.00
- ★1.00
- ★1.00
- Step0.97
- Conmit0.92
- on-commit1.00
- on-step1.00
- build, e2e, security0.99
- lint, tests, docs0.99
- Async1.00
- on-async1.00
- smoke, health, perf1.00
- Adversarial1.00
- on-adversarial1.00
- full suite, coverage 90%+1.00
- Language:0.98
- JS · Python · Go · Rust · anything with a shell0.92
- Agent:1.00
- Claude · GPT · Cursor · Copilot · any0.93
- Checks:1.00
- exit 0 = pass . your commands, your rules0.98
- AlEngineer0.96
- EUROPE1.00
- er1.00
-
- The Lie Detector - Taiha Sh0.94
- File /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.97
- The Contract0.98
- TheIndustryIs0.99
- AIE1.00
- Converging1.00
- Engineering the future of Al0.98
-
- The Lie Detector -Taiha Sh0.92
- ① File0.86
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.99
- ☆0.98
- TheIndustryls0.96
- ANTHROPIC1.00
- Converg1.00
- The advisor strategy0.99
- AIE1.00
- Executor1.00
- Advisor1.00
- ★1.00
- Main loop0.99
- Sonnet1.00
- Tool call—0.97
- Opus1.00
- ★1.00
- ★1.00
- ★1.00
- Runseverytun0.99
- Read / write0.95
- Shared context1.00
- Reviews context0.98
- Advisor reads the same context as Executor0.97
- Advisor Strategy → optimize which model thinks. Doesn't verify the output.0.98
- Engineering the future of Al1.00
-
- /Users/talhashelkh/Desktop/dev/vector-harness/docs/presentation.html0.99
- ☆0.72
- OPENAI1.00
- *★★0.60
- AIE1.00
- leveraging Codex in an0.99
- Harness engineering:0.99
- ★1.00
- agent-first world0.97
- ★1.00
- ★1.00
- ★1.00
- Harness Engineering → 'Humans steer. Agents execute.' The harness is the product.1.00
- Engineering the future of Al1.00
-
- The Lie Detector0.99
- h/Desktop/dev/vector-harness/docs/presentation.html0.99
- ☆0.95
- QODO0.91
- qodo0.79
- AIE1.00
- ★1.00
- ★1.00
- Code Review by Qodo1.00
- Review-first, not copilot-first. Verification as the entire business.0.99
- Engineering the future of Al0.97
- er1.00
-
- The Lie Detector - Talha Sh0.94
- ①File0.75
- /Users/talhashelkh/Desktop/dev/vector-haress/docs/presentation.html0.99
- ☆0.86
- WhatI Actually Learned0.96
- AIE1.00
- Enforce, don't instruct.0.98
- ★1.00
- A pipeline gate fires every time. A prompt gets forgotten by token 50,000.0.98
- ★1.00
- ★1.00
- ★1.00
- Guide, don't prescribe.0.98
- Tell the model where the landmines are. Let it do what it's aiready good at.0.99
- Measure, don't assume.0.99
- Trust is a pass rate, a hash, a delta score. Not a feeling.0.96
- Engineering the future of Al0.99
Transcript
134 cues· 1,754 words· 9,259 chars
- 0:15 Hello, hello.
- 0:17 Have you ever given a task to Cloud Code?
- 0:19 And you give it a feature, and you're like, OK, cool.
- 0:23 Can you build this for me?
- 0:24 And Cloud Code starts putting it out into subtasks.
- 0:27 And you see it, OK, this is pretty cool.
- 0:29 You see it running multiple subagents.
- 0:32 All right, this is really cool.
- 0:33 And you can see it ripping through all of your tasks, subagents being completed.
- 0:36 And it gives you a final output, task completed.
- 0:39 Amazing, great.
- 0:41 But when you actually try to run it, it should be like, oh, well, it's not.
- 0:45 Something has failed.
- 0:45 Hey, Claude, can you fix this little bit thing?
- 0:47 Oh, well, let's try it again.
- 0:48 OK, it's fixed.
- 0:49 Everything should be working.
- 0:50 Oh, no.
- 0:51 Actually, just this tiny little thing is just missing.
- 0:55 And that's what my talk is about.
- 0:58 All I want to do is play Cyberpunk on my Xbox while I have Cloud Code do some work for me.
- 1:04 And what I realized was the problem is that I kept on telling Cloud, like, hey, fix this, fix that, fix this, fix that, even though if I give it a spec, if I give it some instructions, if I give it little to no instructions, every time there is something that I need to tell it.
- 1:19 So what that means is like, I am the enforcement.
- 1:23 I am the enforcement there.
- 1:24 I have to tell Claude on what exactly you need to do and how exactly this needs to be enforced.
- 1:30 So the agent says it's done, but you have to check it anyway because there is nothing else that can check it for you.
- 1:38 So what I wanted was something to be very deterministic.
- 1:41 So when an agent says it's completed, have this enforcement layer deterministically check something like whether it's actually been done or not.
- 1:47 Actually, the way I wanted it to be done, because it says it is done, but is it the way that I want it?
- 1:53 So I needed some deterministic way to do that.
- 1:57 So I tried it, I built my own vector, I call it my own product called Vector V1, and it deterministically checks Cloud's output.
- 2:05 And the way it did that is through using Cloud Hooks, so that way whenever Cloud finishes its session, it automatically, the hook calls my vector product or program, and it checks it for me.
- 2:16 Cool, and this is how it essentially looks.
- 2:17 So I basically give it a config file, define all of my tests over here, like what I needed to be checked, and if it fails, it can actually keep on telling Cloud, like, hey, look, this is failing, try again, try again, try again.
- 2:28 Sorry if it's a little bit, little.
- 2:31 So you can see one of the test outputs over here.
- 2:34 So it's like, okay, first the test pass has failed, then it retries again, then all of the things pass.
- 2:39 Okay, cool.
- 2:40 So what that means is it's not about whether Cloud can actually do the task, it's about trust.
- 2:45 Can I trust Claude to actually do everything for me?
- 2:48 And by the way, when I say Claude, I'm just talking in general about LLM agents in general, is when I give a task to a coding agent, does it actually complete it?
- 2:58 Yeah.
- 2:59 So, and that's something that, and what I started doing was started telling this about to people about and going to different events.
- 3:05 And it was, it was really cool.
- 3:06 Like, Hey, look, what about this verification feature that I built?
- 3:08 It was so good.
- 3:09 It was so amazing.
- 3:10 And then I met one of the anthropic engineers and, um, he, and they just told me that we're not going to need this anymore.
- 3:18 Like we'll have like another agent or another model that it will be so smart that you won't need enforcement.
- 3:25 Okay.
loading