vidtheque.
313 talks watchedbrowse the corpus

the proof · ai engineer 2026

The knowledge of AI Engineer 2026, on tap. Ask it something.

Your agent watched it — every sentence spoken, every line that crossed the screen, every frame. Every answer comes back with the sentence, the slide, and the second it happened.

10 results of ~35

Evaling Video Slop — Maor Bril, Character.ai · AI Engineer · 4:42

youtu.be/b_PmGocP4rc?t=285 ↗

Evaling Video Slop — Maor Bril, Character.ai

AI Engineer2 momentsb_PmGocP4rc

  1. 4:42spokenbenchmark on how we test video that we can rerun over and over and over again. So that combines both metrics, as I said earlier, that knows how to view individual frames, but also consistent LLM as a judge, right, where we also use human annotation to calibrate the LLM as a judge. So for every report that we generate with that harness, we're able to have humans annotated and basically feed that fe…youtu.be/b_PmGocP4rc?t=285 ↗
  2. 18:02spoken…ideos and games and books. Fair. So how would you construct and align sort of like any human judges? Yeah. So this is actually solved at first at the Judge Judy part. where every report it'll generate, a human can go and annotate it. And we actually, we do that. We periodically have sessions where everyone spends 10 to 15 minutes just annotating videos. And that usually happens on, multiple axes. …youtu.be/b_PmGocP4rc?t=1113 ↗

How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads · AI Engineer · 18:16

youtu.be/xyL2Ltkh-SA?t=1094 ↗
  1. 18:16spokenDo we have time for questions, staff? One? Do we have time for questions? Just one. All right. You went up first, sir. Go ahead. Are all your eval judgments being performed by humans, or are you also using LLM as a judge? And if so, what's your calibration process look like for calibrating that judge to provide good evaluations? Yeah, I think that's a good question. I think I wouldn't say all. I t…youtu.be/xyL2Ltkh-SA?t=1094 ↗

The maturity phases of running evals — Phil Hetzel, Braintrust · AI Engineer · 8:04

youtu.be/FB-MLPhL9Ms?t=482 ↗
  1. 8:04spokenReason being is that you need to extract a lot of this domain-specific knowledge out of that human annotator's head so that eventually you can scale that type of knowledge through a technique like LMS Judge. But this is a great first step, performing human annotation. Who in this room is at this step? Okay, this is way more advanced group. That's okay. That's a good place to start. Are you using h…youtu.be/FB-MLPhL9Ms?t=482 ↗
  2. 7:55spokenBut more importantly, you should make that human annotator perform a justification for why they chose that thumbs up or thumbs down.youtu.be/FB-MLPhL9Ms?t=473 ↗
  3. 7:49spokenyou should give a thumbs up or thumbs down. Was this response good, was it bad?youtu.be/FB-MLPhL9Ms?t=467 ↗

The Agentic AI Engineer - Benedikt Sanftl, Mutagent · AI Engineer · 18:05

youtu.be/pSto5YaNGUo?t=1083 ↗
  1. 18:05spokenthen another point is your lm as a judge solution should be calibrated so that you don't have the scoring noise between judges because since lm are lms are undeterministic what you will mostly encounter is the same judge can evaluate a problem different ways on each run And then here you have to make sure your LM as a judge solution deals with this variance problem.youtu.be/pSto5YaNGUo?t=1083 ↗

Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft · AI Engineer · 21:03

youtu.be/3z2uT5aDx_Y?t=1261 ↗
  1. 21:03spokenBut if it's unexpected behavior, we mark it as pass. Okay, so with this, how do we validate if our LLM judge is working as expected? First thing is we need to treat it as a classifier. So how we train our classification models in traditional machine learning, we can also treat our evaluation judges as those traditional ML classifiers with binary outputs. Once we have binary outputs for every metri…youtu.be/3z2uT5aDx_Y?t=1261 ↗
  2. 19:54spokenAnd a binary outcome is very easy to calibrate and train LLM judge that can consistently score your agent trajectory.youtu.be/3z2uT5aDx_Y?t=1192 ↗
  3. 19:18spokenSo we can use these pre-built eval metrics as a baseline, but we shouldn't use them as our core eval metrics because we want eval metrics to be actionable and tied to the business outcome or the product which we are focusing on. Okay, so what LLM as a judge should be. We collaborate very closely with domain experts and utilize their insights. So eval should be framed around a task success or failu…youtu.be/3z2uT5aDx_Y?t=1156 ↗

Add this corpus to your own agent

mcp endpoint

loading…

claude code

loading…

codex

loading…

mistral vibe

loading…

Claude, ChatGPT and Vibe: add a custom connector and paste the endpoint. No sign-in.