Videos LC3-P7v3yoI
Skills are the New SDKs - Elvin Aghammadzada, DataRobot
Scene timeline
68 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 221
- whisperx 221
- chunks
- 47
- from 221 cues
- keyframes
- 44
- kept of 68 captured
- frames with text
- 44
- 597 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 7.0 MB
- word timings on 221 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:20 | 1m 29s |
stt |
done | — | 2026-08-11 00:22 | 28s |
chunk |
done | — | 2026-08-11 00:22 | 0s |
text_embed |
done | — | 2026-08-11 00:22 | 1s |
keyframe |
done | — | 2026-08-11 00:22 | 1m 10s |
ocr |
done | — | 2026-08-11 00:23 | 23s |
frame_embed |
done | — | 2026-08-11 00:24 | 8s |
Frames, and what the machine read
-
- = DataRobot0.97
- Elvin1.00
- Aghammadzada1.00
- Data Science Engineer0.98
- AlEngineer0.98
- World's Fair0.96
- Skills Are the New SDKs1.00
-
- DataRobot0.95
- Elvin Aghammadzada1.00
-
- DataRobot1.00
- Promise1.00
- By the end of this session, we WILL cover:0.99
- DataRobot0.99
- Context (engineering) challenges0.98
- —Why Skills matter?0.99
- —MCP vs. Skills0.98
- Elvin Aghammadzada1.00
- —Skills Ecosystem0.97
- 21.00
-
- DataRobot1.00
- Every Al app is actually THREE applications:1.00
- DataRobot0.98
- The one the user sees1.00
- The one the model sees1.00
- The one the data sees1.00
- Elvin Aghammadzada1.00
- Bugs live in the two we don't render.1.00
-
- DataRobot0.97
- Problem 1: Context Window Paradox1.00
- We've been promised 1M, 5M, even infinite context windows. This1.00
- explosion in capacity has reshaped our thinking on:0.98
- DataRobot0.96
- RAG seemed obsolete (why search for the right document when you1.00
- can include them all?)1.00
- MCP appeared limitless (connect every tool and let models handle1.00
- everything)1.00
- Elvin Aghammadzada0.97
- In reality: longer contexts don't create better responses. In fact,1.00
- overloading context causes agents and applications to fail in1.00
- unexpected ways. Contexts become poisoned, distracted by irrelevant0.99
- information, confused, and degraded by conflicting signals.1.00
-
- DataRobot0.96
- Problem 1: Context Window Paradox1.00
- We've been promised 1M, 5M, even infinite context windows. This0.99
- explosion in capacity has reshaped our thinking on:0.99
- DataRobot0.96
- RAG seemed obsolete (why search for the right document when you0.99
- can include them all?)1.00
- MCP appeared limitless (connect every tool and let models handle1.00
- everything)1.00
- Elvin Aghammadzada0.97
- In reality: longer contexts don't create better responses. In fact,1.00
- overloading context causes agents and applications to fail in0.99
- unexpected ways. Contexts become poisoned, distracted by irrelevant0.99
- information, confused, and degraded by conflicting signals.1.00
-
- DataRobot0.97
- THIS IS FINE.0.97
- Day in the life1.00
- DataRobot0.99
- Day 1: "Wow, we can fit everything!"1.00
- Day 7: "Why is it slower?"0.99
- Day 14: "Why is it giving wrong answers?"1.00
- Elvin Aghammadzada1.00
- Day 21: "The context is eating itself"1.00
-
- DataRobot0.97
- THIS IS FINE.0.98
- Day in the life0.99
- DataRobot0.99
- Day 1: "Wow, we can fit everything!"1.00
- Day 7: "Why is it slower?"0.99
- Day 14: "Why is it giving wrong answers?"1.00
- Elvin Aghammadzada1.00
- Day 21: "The context is eating itself"1.00
- absolutely righ0.97
-
- DataRobot0.97
- Problem 2: Docs have a new audience1.00
- DataRobot0.94
- Traffic to documentation sites is now 50% coding agents,0.98
- up from 10% a year ago. Traditional docs assume a human0.99
- reader – someone who can intuit, ask a follow-up, or0.99
- Elvin Aghammadzada1.00
- Google the confusing bit. Models can't do any of that.1.00
- That reasonable parameter name that every developer0.99
- understood? The model will hallucinate its value with1.00
- complete confidence.0.98
-
- DataRobot0.97
- [Context Engineering] Skills0.99
- DataRobot0.98
- Elvin Aghammadzada1.00
- Friction vs. Fluency1.00
- Moats0.98
-
- DataRobot0.97
- Hardware, data, integrations0.99
- DataRobot0.97
- They all have something in common. They are friction moats and they work by0.99
- making it hard to leave.0.99
- Every year, we get better at reducing switching costs. Better tools, better APls,1.00
- better open standards. The customer who was trapped by switching cost0.99
- Elvin Aghammadzada0.98
- eventually finds the friction low enough to leave. And when they do, they leave0.99
- with resentment.1.00
-
- DataRobot1.00
- Skills are entirely different1.00
- Skills create fluency moat.0.99
- DataRobot0.94
- Unlike friction, fluency compounds. The tenth skill you ship1.00
- makes the first nine more valuable, because now an agent can1.00
- Elvin Aghammadzada1.00
- chain them training to deployment to predictions to1.00
- monitoring. That accumulates in a way that is genuinely hard to1.00
- catch up to.1.00
- 91.00
-
- DataRobot0.97
- Experience1.00
- “The degree to which using the platform produces the result0.99
- DataRobot0.96
- you expected, with the minimum friction between intention1.00
- and outcome".0.96
- Friction moats are defensive. You are making it hard to leave.0.99
- Elvin Aghammadzada1.00
- Fluency moats are offensive. You are making it genuinely0.98
- better to stay.0.99
-
- DataRobot1.00
- Skills are the experience layer for agents.1.00
- The feel of the thing and the reliability of the outcome.0.99
- DataRobot0.98
- The SaaS moat eventually commoditizes. What doesn't0.99
- commoditize is the feeling of fluency and the trust when a1.00
- Elvin Aghammadzada0.98
- platform reliably does what you expect. And that's what skills1.00
- encode for agents.0.97
- 111.00
-
- DataRobot1.00
- Teachability is new the enterprise item1.00
- When an enterprise evaluates your platform today: they have a list. Security,0.99
- DataRobot0.97
- Compliance, Governance, SLA guarantees, Audit logging & tracing,0.99
- Integrations.1.00
- In the agentic era, that checklist gains a new item.0.99
- Elvin Aghammadzada1.00
- Teachability.1.00
- Can your platform encode its operational knowledge in a form agents can act0.99
- on? It is the ability to encode platform's knowledge in a form agents can1.00
- govern, version, and audit.1.00
- 121.00
-
- DataRobot0.99
- DataRobot1.00
- Standardizing1.00
- Elvin Aghammad1.00
- Context1.00
- Skills1.00
-
- DataRobot1.00
- Why think about standards?1.00
- HTML launched in 1993.0.99
- DataRobot0.95
- 20 years later, React cameto web-not just as a framework, but as a philosophy.0.99
- React introduced patterns of reactivity and modularity that became the0.99
- standard for building applications.1.00
- Elvin Aghammadzada1.00
- LLM and agent ecosystem feels like 1993 all over again. We're stitching together1.00
- tools, and discovering patterns through trial and error—much like developers0.99
- once wrestled with raw HTML and CSS. We haven't found our "React moment"1.00
- yet – the foundational patterns that make Al dev. as systematic as modern web.1.00
-
- DataRobot1.00
- Context Engineering1.00
- A few years ago, top Al researchers predicted prompt engineering would be0.99
- DataRobot1.00
- obsolete by now. They couldn't have been more wrong.0.99
- What we once called "prompt engineering" (crafting the perfect instruction for a0.99
- chatbot) has transformed into "context engineering"- the art of dynamically0.99
- orchestrating what information fills the context window at exactly the right1.00
- moment.1.00
- Elvin Aghammadzada1.00
- [Context engineering is the] "...delicate art and science of filling the context0.98
- window with just the right information for the next step."1.00
- 91.00
-
- DataRobot1.00
- Context Engineering1.00
- Systen Instructions0.95
- Models these days are remarkably1.00
- Cloude Builtin Tools0.96
- CLAUDE.md0.95
- MCP Tools0.91
- intelligent, but intelligence without1.00
- go bulild me a dope feature0.90
- User messoge0.97
- DataRobot1.00
- context is paralyzed. Context1.00
- 40.84
- engineering isn't about writing better0.99
- prompts - it's about building systems0.98
- "the smart zone"1.00
- 40% context used0.99
- that automatically curate, filter, and0.99
- "the dumb zone"1.00
- 168k tokens0.98
- inject the precise context needed for0.98
- Elvin Aghammadzada1.00
- each step of an agent's journey.1.00
- 23k tokens0.98
- reserved for outo-compact0.99
- 32k tokens0.97
- reserved for output1.00
- 101.00
Transcript
221 cues· 4,215 words· 22,887 chars
- 0:03 We'll talk about skills today, the current challenges about context, and then lastly about the ecosystem that's being built around skills lately, especially with OpenCLO and all of those stuff getting a lot of attention in the industry.
- 0:22 So let's get into it.
- 0:24 Every AI or agentic app typically has three layers.
- 0:28 The one that user sees, which is typically the UI, the user interface.
- 0:33 The one that model sees, which is the system prompt, as well as the tool descriptions.
- 0:38 And the one that data sees, which is the schema of data or the input and output of the tool calls.
- 0:45 And these little bugs typically live on these second and third one, which is slightly hidden from the user.
- 0:53 And so let's talk about these problems.
- 0:57 The first one is that beautiful lie that we've been told that the latest frontier models have infinite context windows.
- 1:05 They're typically promised that each one has 1 million, 5 million, even latest models, they're promised as infinite context window, which incorrectly shapes our thinking perspectives about RAG and NCP in a sense.
- 1:21 Because if we consider that there's infinite context window, we just think that we can dump the whole list of documents to the context and expect it to magically work.
- 1:32 This applies to MCP as well.
- 1:34 Let's say we can write 100 tools and expect the LLM to use or pick the right tool at the right time, which is typically not true.
- 1:42 Because in reality, the longer context typically doesn't mean better performance.
- 1:49 Because each time you put more context into LLM, that's one more place where the LLM might be misled or its context might be poisoned.
- 1:58 In fact, there was a paper called Context Rot and it proves that after 25% usage of the context window, so for example, for 1 million token, if you've used 256K of it, the performance starts to degrade.
- 2:16 So I'm pretty sure it starts really well as day in the life, but as time goes and as the conversation grows, let's say there is a hundred prompt input and output responses in your one chat and add it to a hundred chats.
- 2:31 And then each one of them kind of calls MCPs or tools each time out of them in the middle fails.
- 2:39 And at the end of the day, the context is almost eating itself in a sense that the latest frontier model can perform.
- 2:47 like a really poor model.
- 2:48 I'm pretty sure we all have seen this, you are absolutely right from Cloud Code.
- 2:53 Even though it feels assuring that you're absolutely right, at one point it makes you feel bad that Cloud Code is making the same mistake it made five minutes ago.
- 3:05 The second problem is mostly about documentation websites.
- 3:10 This is an interesting statistic that just from last year to this year, the traffic to documentation websites increased from 10% to 50%.
- 3:21 That's coming from coding agents.
- 3:22 The challenge here is the fact that docs are written for humans, but models are expected in a different format.
- 3:31 Because docs typically require some kind of intuition, asking for a follow-up, or if you don't understand it well, you can just Google it.
- 3:38 quickly and then find the answer for it.
- 3:41 But models typically doesn't have those capabilities until you actually build a context engineering around it.
- 3:48 I want to add one more strategic point about the future of enterprise AI and why we think that skills might be a huge contributor to where the enterprise agents or enterprise AI is heading towards.
- 4:01 Let's talk a little bit about modes these days.
- 4:04 Now, if you want to understand modes today or tomorrow, we can go back in time to the history of software to see how has it been in the past.
- 4:14 Typically, the things that have been the hardest to replicate have been the modes for the software companies or even hardware ones.
- 4:22 So for example, whoever controls the hardware, the data, or especially the integrations in the SaaS era would be the ones that would control the market most of the time.
- 4:32 so the real mode here would be friction where you would make it genuinely hard to you to leave to to make someone leave your platform so the core idea here is the switching cost was too much that it was just hard to switch now these days we're hearing news where cloud code would rewrite
- 4:54 hundred thousand or a million lines of code from Python to Rust.
- 4:58 So basically we're getting really, really good at switching codes and reducing it.
- 5:03 Now we can have better tools, better APIs, just better languages, and then switching it in few days.
- 5:11 Now skills are creating an entirely different ecosystem where instead of creating a friction mode, they're actually creating fluency mode.
- 5:20 So the core idea is switching the delay a little bit so that each time your skill makes the experience better for whoever is the one that's using your platform.
- 5:32 And that fluency compounds by experience.
- 5:36 The real result that the users of your platform is getting from here is the experience.
- 5:42 If you would look at the definition of experience, you can realize that it's satisfaction or minimum amount of friction between intent and then outcome.
- 5:52 So basically whatever the user is trying to achieve through your platform, achieving it in maybe minimum amount of time or trying to optimize the whole satisfaction layer or full experience layer of it.
- 6:04 If you look at the equation, you will see that overall friction modes and fluency modes are the opposite.
- 6:10 Friction modes are typically defensive.
- 6:12 So you're trying to make your experience in a sense that someone wouldn't switch.
- 6:18 Now, fluency modes are offensive in a sense that you're making the whole experience really nice in a sense that you're making it genuinely hard for someone to switch anyways, because they like the platform.
- 6:33 So we think that the feel of the thing or the reliability of the outcome from the start, which is your intent to the result is the whole experience layer.
- 6:44 And the skills are one of the best tools that can actually quote, quote, commoditize that experience layer for agents.
loading