Videos wEc9aG7cRQc
Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog
Scene timeline
54 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 185
- whisperx 185
- chunks
- 44
- from 185 cues
- keyframes
- 22
- kept of 54 captured
- frames with text
- 22
- 402 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 6.3 MB
- word timings on 185 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 03:43 | 1m 41s |
stt |
done | — | 2026-08-10 03:45 | 25s |
chunk |
done | — | 2026-08-10 03:45 | 0s |
text_embed |
done | — | 2026-08-10 19:46 | 0s |
keyframe |
done | — | 2026-08-10 03:45 | 38s |
ocr |
done | — | 2026-08-10 03:46 | 9s |
frame_embed |
done | — | 2026-08-10 19:46 | 4s |
Frames, and what the machine read
-
- Why Your Al Agent0.97
- Diane Lin0.99
- Disagrees With Itself0.98
- and What to Do About It0.99
- Dianhuan (Diane) Lin1.00
- DATADOG1.00
-
- Background1.00
- Education & Impact1.00
- PhD in Machine Learning, Imperial College London0.99
- Diane Lin1.00
- Research affiliations at MIT's CSAIL1.00
- Amazon (Alexa Team)0.99
- a0.99
- Founding applied scientist1.00
- Vicarious1.00
- (acquired by Gaogle DeepMind)0.98
- Senior Researcher1.00
- Zscaler1.00
- Director of1.00
- zscaler1.00
- Machine Learning1.00
- Culminate1.00
- (acquired by Datadog)1.00
- CTO & Co-founder0.97
- CULMINATE1.00
- Datadog1.00
- Tech Lead0.99
- Self-Evolving Al Agents0.98
- for the Modern SOC1.00
-
- Agenda1.00
- Diane Lin0.99
- The Problem: LLM inconsistency in practice0.99
- Where Does Inconsistency Come From?0.99
- Solutions1.00
- Experimental Results1.00
- DATADOG0.95
-
- Diane Lin0.96
- The Problem1.00
- Same model. Same input. Different output.1.00
- DATADOG41.00
-
- LLM Inconsistency in Practice0.99
- Diane Lin0.95
- Sentiment analysis: classify hotel reviews as positive or negative0.99
- Review A0.98
- "Fantastic hotel in the perfect location. It's just a five-minute walk to all the major museums and restaurants, yet0.99
- tucked away on a quiet street so you don't hear any city noise at night. The concierge, David, gave us the best0.99
- recommendations for local dining that weren't tourist traps. Everything about this place is top-notch. Highly1.00
- recommend to anyone visiting the city."0.99
- Review B0.95
- "It's an okay hotel. It's clean and safe, which are the most important things, but it entirely lacks any charm or0.99
- personality. The check-in process was automated via a kiosk, so we never actually spoke to a human being during our0.99
- entire stay. The room had all the basics but felt very sterile, almost like a hospital room. If you just need a cheap,0.99
- clean place to crash before an early flight, it works fine. Just don't expect a memorable hospitality experience."1.00
- DATADOG 50.90
-
- Where Does the1.00
- Diane Lin0.99
- Inconsistency Come From?1.00
- The gray zone: data that sit near the decision boundary1.00
- DATADOG70.99
-
- verdict_iter0 verdict_iter1 verdict_iter20.99
- Gray zone: data that sit near the de1.00
- suspicious1.00
- benign1.00
- benign1.00
- ry1.00
- Diane Lin1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- benign1.00
- suspicious1.00
- suspicious1.00
- Security Analysts triage alerts1.00
- benign1.00
- benign1.00
- benign1.00
- “Failed authentication attempts detected from suspicious IP"0.99
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- suspicious1.00
- benign1.00
- benign1.00
- suspicious0.99
- benign1.00
- suspicious1.00
- benign1.00
- benign1.00
- benign1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- benign1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- suspicious1.00
- benign1.00
- benign1.00
- suspicious1.00
- benign1.00
- suspicious1.00
- benign1.00
- suspicious1.00
- suspicious1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- suspicious1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- suspicious1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- benign1.00
- DATADOG 90.96
- benign1.00
- benign1.00
- benign1.00
-
- Gray zone: data that sit near the decision boundary1.00
- Diane Lin0.99
- Gray Zone0.99
- DATADOG0.99
-
- Identifying the Gray Zone:1.00
- Diane Lin1.00
- Active learning1.00
- Active learning tells us exactly where to look1.00
- DATADOG 110.95
Transcript
185 cues· 2,838 words· 16,372 chars
- 0:02 Hello.
- 0:04 This session is about why do your AI agents disagree with itself and what to do about it.
- 0:13 And this is a Dear Huang Lin, and you can call me Diane.
- 0:19 A little bit about my background.
- 0:22 My journey with AGI start from my PhD at Imperial College London, where I was working on continual learning.
- 0:31 Later, I worked with Professor Josh Tenenbaum at MIT on one-shot learning, meaning learning from one single example.
- 0:42 Afterwards, I was lucky that Alexa was launching.
- 0:46 I was the first three applying scientists on the question answering team at Alexa.
- 0:53 later attracted by a startup working on AGI at the Silicon Valley called Vicarious.
- 1:01 So today Vicarious is part of the Google demind.
- 1:05 There I got to work on the exciting frontier research on zero-shot transfer learning.
- 1:13 After working on AGI for a few years, I decided to pivot to the applied world at GSCADR, which is cybersecurity company.
- 1:25 There, I got to apply a variety of different machine learning model to address the cybersecurity challenge.
- 1:34 There for five years, I decided to co-found my own company because the pinpoint I have seen at GSCADR
- 1:44 So CloudMini is building AI agent to auto triage your security alerts and which make your SOG more efficient.
- 1:55 We are super proud that CloudMini has been acquired by Datadog earlier this year.
- 2:02 So now I'm part of the Datadog and leading the development of self-evolved agent
- 2:14 so today i'm going to share with you the problem you're probably seeing your day-to-day building ai agent the inconsistency and then we are discussed about where this inconsistency come from in order to figure out the solutions and we will discuss a few trade-offs among the different solutions and then eventually we'll show you the experimental results to show how feasible they are
- 2:46 You probably see this same model, same input, but different output.
- 2:52 Here, I don't mean the wording different.
- 2:55 I mean, semantically different output.
- 3:01 You might be thinking that's the stochastic nature of LLM.
- 3:07 Hold on to that thought.
- 3:09 Let's look at a few concrete examples.
- 3:13 So here, imagine you're giving a task for sentiment analysis, where we're supposed to label each review of the hotel either positive or negative.
- 3:26 This is a typical NLP task, and today's LLM is really good at this.
- 3:33 However, there's one problem.
- 3:36 When you pipe this input to the same LLM model,
- 3:42 Occasionally you will see different verdict coming out of it when you run multiple times.
- 3:49 That means a single evaluation run didn't tell you the whole story.
- 3:55 You need to repeat your evaluation multiple times in order to get a holistic picture.
- 4:05 By the average over the different runs results.
- 4:10 And this sounds like a technical inconvenience, but the reality is it's way more problematic than just inconvenience.
- 4:21 Let me show you an example in a different domain in cybersecurity.
- 4:26 Imagine you are a security analyst where your day-to-day job is triage the security alerts and decide whether it's malicious that you need to take certain actions to stop the attack or it's false alarm, meaning it's actually benign, which you could ignore.
- 4:47 So here, example alert you got is there's a failed logging attempt to Gmail account detected from suspicious IP.
- 4:58 And if we give such similar task to this one, to AI agent, and run it different times, you will see sometimes it stays consistently benign.
- 5:10 Sometimes it stays consistently suspicious across different lines.
- 5:15 However, there are cases where it flip-flops between benign and suspicious.
- 5:24 Now your customer will have a difficulty using such a product because they will wonder which one should I trust?
- 5:34 So the inconsistency is causing a trust issue in your product.
- 5:42 Imagine if you were in a POC bake-off, one vendor gave the consistent verdict all the time while the other
- 5:55 vendor have this flip-flop verdict you can imagine which one will win the deal now i'm pretty sure you want to solve this problem first we have to figure out where the inconsistency come from
- 6:21 The good news is the data points tend to flip-flop, actually concentrate around the decision boundary, the so-called gray zone.
- 6:34 Let me illustrate with the earlier example.
- 6:37 Here in the sentiment analysis, we have two reviews.
- 6:42 The first one
- 6:44 Actually, no matter how many times you run through the LL model, it will always give you the positive label.
- 6:54 In this case, it's pretty obvious the review is super positive.
- 7:00 However, the second case is the one tend to flip.
loading