Videos 4VhbYlfC7Gs
Malleable Evals: Why Are We Evaluating Adaptive Systems with Static Tests? — Vincent Koc, OpenClaw
Scene timeline
35 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 156
- whisperx 156
- chunks
- 26
- from 156 cues
- keyframes
- 28
- kept of 35 captured
- frames with text
- 28
- 344 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 3.0 MB
- word timings on 156 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 22:27 | 1m 10s |
stt |
done | — | 2026-08-10 22:29 | 17s |
chunk |
done | — | 2026-08-10 22:29 | 0s |
text_embed |
done | — | 2026-08-10 22:29 | 1s |
keyframe |
done | — | 2026-08-10 22:29 | 1m 13s |
ocr |
done | — | 2026-08-10 22:30 | 8s |
frame_embed |
done | — | 2026-08-10 22:30 | 4s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- comet1.00
- Malleable Evals: From Static AI to0.99
- asuring Adaptive Systems0.99
- Al0.91
- Improv0.89
-
- comet1.00
- ***0.62
- Malleable Evals: From Static AI to1.00
- AIE1.00
- ★1.00
- ★1.00
- Measuring Adaptive Systems1.00
- Vincent Koc (@vincent_koc)0.99
- Comet1.00
- Google DeepMind1.00
-
- Hi I'm Vincent Koc1.00
- Your Friendly Clanker1.00
- AIE1.00
- ★1.00
- ★1.00
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- Google DeepMind1.00
- AlEngineei0.94
-
- Hi I'm Vincent Koc1.00
- Your Friendly Clanker1.00
- My love for testing new0.97
- technology (2013)0.99
- AIE1.00
- ★1.00
- ★1.00
- AIE Europe 20260.99
- Vincent Koc1.00
- comet1.00
- Braintrust1.00
- WorkOS OpenAI0.94
- AIE0.99
-
- Hamel Husain1.00
- @HamelHusain1.00
- Recap of tech ragebait I've seen in the last 24 hours0.99
- *★*0.66
- ★1.00
- - Serious work doesn't get done remotely0.98
- AIE1.00
- - Evals are dead0.99
- ★1.00
- - A/B testing is dead0.97
- ★1.00
- ★1.00
- ★1.00
- - If you aren't using an agent framework in favor of while loops it's just1.00
- hubris1.00
- - You can replace the PyData ecosystem only if there was pandas for1.00
- typescript1.00
- Did I miss anything0.98
- 7:14 AM · Sep 7, 2025·26KViews0.93
- AIE Europe 20260.99
- Vincent Koc0.98
- comet1.00
- AlEngineer0.97
- EUROPE1.00
- AIE0.99
-
- Example Focused Unit Tests1.00
- AIE1.00
- Manual Regression Suites1.00
- ★1.00
- ★1.00
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- AlEngineer0.97
- EUROPE1.00
- AIEr.0.84
-
- Example Focused Unit Tests0.98
- AIE1.00
- Manual Regression Suites1.00
- ★1.00
- CI/CD Pipelines0.98
- Chaos Engineering and Observability1.00
- Vincent Koc1.00
- comet1.00
- AIE Europe 20261.00
- Engineering the future of Al0.99
- AIEn0.79
-
- StaticBenchmarks1.00
- AIE1.00
- Hand Curated Evaluations1.00
- ★1.00
- ★1.00
- Pre-Deployment Offline Evaluations0.99
- ??(gap)0.98
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- AlEngineer0.98
- EUROPE1.00
- AIE1.00
-
- Static Benchmarks0.99
- Hand Curated Evaluations1.00
- Pre-Deployment Offline Evaluations0.99
- ??(gap)1.00
- AIE0.94
- improy0.91
-
- StaticBenchmarks1.00
- AIE1.00
- Hand Curated Evaluations1.00
- ★1.00
- ★1.00
- Pre-Deployment Offline Evaluations1.00
- ??(gap)1.00
- AIE Europe 20260.99
- Vincent Koc1.00
- comet1.00
- Engineering the future of Al0.99
- AlEng.0.89
-
- ADAPTIVE TESTING FOR LLM EVALUATION:1.00
- A PSYCHOMETRIC ALTERNATIVE TO STATIC BENCHMARKS0.99
- A PREPRINT0.99
- ★1.00
- Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, Nitesh V. Chawla0.99
- Peiyu Li*, Xiuxiu Tang',0.93
- AIE1.00
- ★1.00
- Notre Dame, IN 46556 USA0.99
- University of Notre Dame0.98
- ★1.00
- {pli9, xtang8, schen34, ycheng4, rmetoyer, thua, nchawla}@nd.edu0.99
- ★1.00
- ★1.00
- ★1.00
- ABSTRACT1.00
- Evaluating large language models (LLMs) typically requires thousands of benchmark items, making0.99
- the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely0.99
- on average accuracy over fixed item sets, treating all items as equally informative despite substantial0.99
- variation in difficulty and discrimination. We introduce ATLAS, an adaptive testing framework based1.00
- on Item Response Theory (IRT) that estimates model ability using Fisher information-guided item0.99
- selection. ATLAS reduces the number of required items by up to 90% while maintaining measurement0.99
- precision. For instance, it matches whole-bank ability estimates using only 41 items (0.157 MAE)1.00
- on HellaSwag (5,600 items). We further reconstruct accuracy from ATLAS's ability estimates1.00
- and find that reconstructed accuracies closely match raw accuracies across all five benchmarks,0.99
- discrimination within accuracy-equivalent models: among more than 3,000 evaluated models, 23-1.00
- indicating that ability θ preserves the global performance structure. At the same time, θ provides finer0.99
- 31% shift by more than 10 rank positions, and models with identical accuracies receive meaningfully0.99
- different ability estimates. Code and calibrated item banks available at https://github. com/0.99
- AIE Europe 20261.00
- Vincent Koc1.00
- Peiyu-Georgia-Li/ATLAS.git.1.00
- comet1.00
- Engineering the future of Al0.98
- AIEr0.88
-
- The Emerging Dawn of Intentful AI Agents0.98
- An Evolution That's Leading to Self-Optimizing AI1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- The sunsetting of0.97
- Prompt Engineering1.00
- (2022-23)1.00
- Door closes on the troubled1.00
- process of writing prompts1.00
- using trial and error "gut feel.0.95
- Killing the cycle of doom0.98
- wordsmithing instructions.1.00
- Vincent Koc1.00
- comet1.00
- AIE Europe 20261.00
- AlEngineer0.96
- EUROPE1.00
- AIEI0.89
-
- The Emerging Dawn of Intentful AI Agents0.99
- An Evolution That's Leading to Self-Optimizing AI1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- The sunsetting of0.97
- The evolution of0.98
- Prompt Engineering1.00
- Context Engineering1.00
- (2022-23)1.00
- (2024-25)1.00
- Door closes on the troubled1.00
- We can now somewhat steer agents,1.00
- process of writing prompts1.00
- a with RAG, reasoning models,0.99
- using trial and error "gut feel.0.95
- long context windows, MCP servers1.00
- Killing the cycle of doom0.98
- and smart long running models. It's1.00
- wordsmithing instructions.1.00
- about managing memory efficently.1.00
- Vincent Koc1.00
- comet1.00
- AIE Europe 20261.00
- AlEngineer0.96
- EUROPE1.00
- AIE1.00
-
- The Emerging Dawn of Intentful AI Agents0.99
- An Evolution That's Leading to Self-Optimizing AI1.00
- **0.89
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- The sunsetting of0.98
- The evolution of0.98
- Prompt Engineering1.00
- Context Engineering1.00
- (2022-23)1.00
- (2024-25)1.00
- Door closes on the troubled1.00
- We can now somewhat steer agents,1.00
- process of writing prompts0.98
- a with RAG, reasoning models,1.00
- using trial and error "gut feel.0.96
- long context windows, MCP servers1.00
- Killing the cycle of doom0.99
- and smart long running models. It's0.99
- wordsmithing instructions.1.00
- about managing memory efficently.1.00
- Vincent Koc1.00
- comet1.00
- AIE Europe 20260.96
- Engineering the future of Al0.99
- AlEngin0.91
-
- The Emerging Dawn of Intentful AI Agents0.99
- AIEr0.78
- improy0.97
-
- AIE1.00
- Models Became Really Good0.98
- ★1.00
- ★1.00
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- Engineering the future of Al1.00
- Al0.78
-
- AIE1.00
- Models Became Really Good0.99
- ★1.00
- ★1.00
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- # Braintrust0.97
- WorkOS OpenAI0.96
- er1.00
-
- The Emerging Dawn of Intentful AI Agents0.99
- An Evolution That's Leading to Self-Optimizing AI1.00
- ★1.00
- ★1.00
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- The sunsetting of0.98
- The evolution of0.97
- The dawn of1.00
- Prompt Engineering1.00
- Context Engineering1.00
- Intent Engineering1.00
- (2022-23)1.00
- (2024-25)1.00
- (2026)1.00
- Door closes on the troubled1.00
- We can now somewhat steer agents,1.00
- Machines can self-optimize based1.00
- process of writing prompts1.00
- a with RAG, reasoning models,1.00
- on the intent and needs of a user or0.99
- using trial and error “gut feel".0.97
- long context windows, MCP servers1.00
- system through approaches like0.98
- Killing the cycle of doom0.99
- and smart long running models. It's0.99
- online prompt optimization and1.00
- wordsmithing instructions.1.00
- about managing memory efficently.0.99
- self-optimizing models.1.00
- Vincent Koc1.00
- comet1.00
- AIE Europe 20261.00
- AlEngineer0.97
- EUROPE1.00
-
- StaticBenchmarks1.00
- AIE1.00
- Hand Curated Evaluations1.00
- ★0.99
- ★0.99
- Pre-Deployment Offline Evaluations1.00
- ??(gap)1.00
- AIE Europe 20261.00
- Vincent Koc1.00
- comet1.00
- Engineering the future of Al0.99
Transcript
156 cues· 2,876 words· 15,393 chars
- 0:15 cool hey everyone uh thanks for joining this session sorry if my sound's a little croaky i've done three talks back to back so one on wednesday one yesterday keynote and then uh workshop style session today so i'm vincent i'm going to be talking about malleable evals from static ai measuring to adaptive systems now
- 0:39 Let's jump into who I am, what I do.
- 0:41 I call myself the friendly canker, I use AI, I use technology, I'm always on the edge.
- 0:46 For those of you that haven't seen my keynote, I do, yeah, I just live on the edge and just do some fun stuff.
- 0:53 So this is me using VR goggles back in 2013 when people hadn't even heard of VR.
- 0:59 It came with a warning label, said only use it for five minutes, I used it for three hours, then I vomited for three hours after that.
- 1:05 So measurement, anything we do in technology, anything on the edge is going to be janky.
- 1:10 It's going to be weird.
- 1:12 And that's kind of fun, in my opinion.
- 1:14 Now, whenever we talk about evals to people and a little bit of pretext, like my role at Comet
- 1:21 I work in evals, I do in eval research, I work with universities, we benchmark and run evals for large set of companies and organizations, everything from like Uber to Netflix to banks even in the UK.
- 1:34 But the thing that's been going on right now is that, hey, like this kind of joke that like evals are a little bit dead.
- 1:40 And it's a little bit of a joke, but there's a little bit of truth to it as well.
- 1:44 And I'm going to hopefully like kind of walk you through the mindset shift and hopefully explain a little bit less about evals, but like what's actually happening in the sort of agentic AI space.
- 1:53 And then how do we then translate that back to evals?
- 1:56 So when we think about software engineering as a practice, when we're thinking about how do we measure things, we kind of look at it from the sense that we're going to start with this thing is meant to do something.
- 2:08 So we would start with a set of examples and write some unit tests.
- 2:12 We might do a manual regression suite, which is like, hey, when we do A and B, sometimes C happens, and C is unfavorable.
- 2:19 Let's not do that.
- 2:20 Let's not make people vomit when they put their VR goggles on.
- 2:24 We could do things like CI-CD pipelines to make sure that the thing ships out and works the way it's meant to and intended to.
- 2:32 But mostly we do things in engineering known as chaos engineering and observability.
- 2:37 For those of you that are unfamiliar with the term chaos engineering, it's basically where you're doing all kinds of random stuff and just breaking it and just having fun with the technology and just seeing where you can stretch it and where you can go.
- 2:49 Now, when we apply this to AI and data science space, as we traditionally know it in the last little while, 2025 included, we do things like static benchmarks.
- 3:00 We have these evaluations.
- 3:01 It's like, oh, I'll give an example.
- 3:04 It's like, how compliant is my AI in risk?
- 3:09 I'm gonna ask it a bunch of questions and make sure it doesn't talk about selling me some financial services, because that's a big no-no.
- 3:17 We will then handcraft a set of questions and examples and sit there and tune this thing up, make sure it's absolutely perfect.
- 3:26 Before we deploy the AI system or the models, we'll do some sort of offline evaluation where we're just cycling through those tests.
- 3:35 but we're missing that sort of chaos engineering space.
- 3:37 We're missing that, you know, like what comes next and how do we mess up with it?
- 3:42 And how do we know where we can stretch this thing?
- 3:43 And I think that's like an honest gap that we see in this space.
- 3:47 And that's why we're just so hyper fixated on benchmarks and evaluations.
- 3:51 If you go to any AI conference in the academic space, all people talk about is like benchmarks.
- 3:56 I created a benchmark for like,
- 3:59 adding numbers and what LLMs think about it.
- 4:01 It's like, well, great, but like, how is this actually helping me?
- 4:04 So then you end up with like this huge, humongous set of like data sets to try and somewhat explain what is happening with your agent until something goes wrong.
- 4:12 And it's a matter of time before something goes wrong and it will, and you're kind of back to the drawing board and trying to figure out what's going on.
- 4:20 And the reason for that is that our AI applications are not static, but we're treating them like they're static software.
- 4:27 Yes, when we ship software, we might change unit tests, they're a little bit quicker to do, but realistically speaking, even software is becoming malleable.
- 4:35 So,
- 4:36 Flip to my keynote I gave yesterday, where I'm one of the core contributors of something called OpenClaw, the harness changes itself, like the harness will shift, like you wanna create skills, you wanna do other things, like it will adapt, right?
- 4:48 So that adaption that we're seeing inside of things where software is being shipped at lightning speed, how does your benchmarks keep up with that?
- 4:55 Like how does your benchmarks adapt to that space?
- 4:58 This is one of many papers that are out there.
- 5:02 I don't remember when this one was published, but this concept of adaptive testing for LLM evals,
- 5:07 this concept is like somewhat revolutionary maybe, but like what happens if our benchmarks would change with our applications?
loading