read-only demo

Videos 0vphxNt4wyk

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind

index_state ready data_status ok

AI Engineer· published 2026-07-14· 0:21:45· en-US· indexed 2026-08-10 19:49

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:23, 1 of 1 keyframes kept
  5. Shot 4, 0:23 to 0:37, 1 of 1 keyframes kept
  6. Shot 5, 0:37 to 1:21, 1 of 1 keyframes kept
  7. Shot 6, 1:21 to 1:47, 1 of 1 keyframes kept
  8. Shot 7, 1:47 to 2:13, 0 of 1 keyframes kept
  9. Shot 8, 2:13 to 2:39, 0 of 1 keyframes kept
  10. Shot 9, 2:39 to 3:04, 1 of 1 keyframes kept
  11. Shot 10, 3:04 to 3:30, 1 of 1 keyframes kept
  12. Shot 11, 3:30 to 3:56, 0 of 1 keyframes kept
  13. Shot 12, 3:56 to 4:22, 0 of 1 keyframes kept
  14. Shot 13, 4:22 to 4:48, 1 of 1 keyframes kept
  15. Shot 14, 4:48 to 5:35, 1 of 1 keyframes kept
  16. Shot 15, 5:35 to 6:02, 1 of 1 keyframes kept
  17. Shot 16, 6:02 to 6:29, 0 of 1 keyframes kept
  18. Shot 17, 6:29 to 6:56, 0 of 1 keyframes kept
  19. Shot 18, 6:56 to 7:23, 1 of 1 keyframes kept
  20. Shot 19, 7:23 to 7:50, 1 of 1 keyframes kept
  21. Shot 20, 7:50 to 8:21, 1 of 1 keyframes kept
  22. Shot 21, 8:21 to 8:52, 0 of 1 keyframes kept
  23. Shot 22, 8:52 to 9:23, 0 of 1 keyframes kept
  24. Shot 23, 9:23 to 9:54, 1 of 1 keyframes kept
  25. Shot 24, 9:54 to 10:20, 1 of 1 keyframes kept
  26. Shot 25, 10:20 to 10:45, 1 of 1 keyframes kept
  27. Shot 26, 10:45 to 11:11, 0 of 1 keyframes kept
  28. Shot 27, 11:11 to 11:36, 1 of 1 keyframes kept
  29. Shot 28, 11:36 to 12:02, 1 of 1 keyframes kept
  30. Shot 29, 12:02 to 12:27, 0 of 1 keyframes kept
  31. Shot 30, 12:27 to 12:52, 1 of 1 keyframes kept
  32. Shot 31, 12:52 to 13:18, 0 of 1 keyframes kept
  33. Shot 32, 13:18 to 13:43, 0 of 1 keyframes kept
  34. Shot 33, 13:43 to 14:09, 0 of 1 keyframes kept
  35. Shot 34, 14:09 to 14:34, 0 of 1 keyframes kept
  36. Shot 35, 14:34 to 15:00, 1 of 1 keyframes kept
  37. Shot 36, 15:00 to 15:25, 1 of 1 keyframes kept
  38. Shot 37, 15:25 to 15:51, 1 of 1 keyframes kept
  39. Shot 38, 15:51 to 16:16, 1 of 1 keyframes kept
  40. Shot 39, 16:16 to 16:41, 0 of 1 keyframes kept
  41. Shot 40, 16:41 to 17:07, 1 of 1 keyframes kept
  42. Shot 41, 17:07 to 17:32, 0 of 1 keyframes kept
  43. Shot 42, 17:32 to 17:58, 1 of 1 keyframes kept
  44. Shot 43, 17:58 to 18:23, 0 of 1 keyframes kept
  45. Shot 44, 18:23 to 18:49, 0 of 1 keyframes kept
  46. Shot 45, 18:49 to 19:14, 1 of 1 keyframes kept
  47. Shot 46, 19:14 to 19:40, 0 of 1 keyframes kept
  48. Shot 47, 19:40 to 20:05, 0 of 1 keyframes kept
  49. Shot 48, 20:05 to 20:30, 0 of 1 keyframes kept
  50. Shot 49, 20:30 to 20:56, 1 of 1 keyframes kept
  51. Shot 50, 20:56 to 21:28, 1 of 1 keyframes kept
  52. Shot 51, 21:28 to 21:45, 0 of 1 keyframes kept

52 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
275
whisperx 275
chunks
39
from 275 cues
keyframes
30
kept of 52 captured
frames with text
30
644 lines read
chapters
19
from the source metadata
keyframe bytes
6.0 MB
word timings on 275 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 11:51 1m 48s
stt done 2026-08-10 11:53 23s
chunk done 2026-08-10 11:53 0s
text_embed done 2026-08-10 19:49 1s
keyframe done 2026-08-10 11:53 2m 09s
ocr done 2026-08-10 11:55 15s
frame_embed done 2026-08-10 19:49 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 448.8

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 663.0

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2734.1

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.94
    6. OpenAI0.92
    7. Akamai0.99
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.99
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:16 #3 done2 line(s)

    shot 3·sharpness 263.9

    1. World's Fair0.97
    2. ORI0.85
  • 0:26 #4 done2 line(s)

    shot 4·sharpness 188.4

    1. AlEngineer0.99
    2. World's Fair1.00
  • 0:59 #5 done15 line(s)

    shot 5·sharpness 1949.7

    1. THE PROBLEM0.90
    2. AlEngineer0.98
    3. Vibe Checks Fail in Production1.00
    4. World'sFair1.00
    5. Why shipping untested skills breaks production0.99
    6. PRESENTED BY1.00
    7. SkillsBench indexed 47k+ unique skills across 6,300 repos.1.00
    8. Microsoft1.00
    9. Almost none have tests.1.00
    10. Most skills are tested with two manual runs and shipped.0.99
    11. Bad skills don't crash; they quietly corrupt outputs.1.00
    12. The Rule: You wouldn't merge code without tests. Why ship skills without evals?1.00
    13. 2 / 250.88
    14. World'sFair1.00
    15. Engineering the future of Al1.00
  • 1:34 #6 done24 line(s)

    shot 6·sharpness 2722.6

    1. • CONTEXT0.89
    2. AlEngineer0.98
    3. Agents We Use vs. Agents We Build0.99
    4. World'sFair1.00
    5. The hidden reliability gap1.00
    6. PRESENTED BY1.00
    7. Agents We Use1.00
    8. Agents We Build0.98
    9. Microsoft1.00
    10. - Antigravity, Cursor, Claude Code, Codex0.99
    11. - Customer support bot, internal workflow agent0.99
    12. You (engineer who can recover from errors)1.00
    13. End users who leave on the first failure0.99
    14. You say "No, use /commit-skill" and it recovers0.99
    15. Customer leaves (won't debug your agent)1.00
    16. Model-invoked or user-invoked (slash command)0.99
    17. Model-invoked only (no human fallback)1.00
    18. Human in the loop compensates for errors1.00
    19. Non-negotiable (saves your customers)1.00
    20. The further the user is from the skill system, the higher the reliability bar. Automated evals become essential.1.00
    21. 3 / 250.93
    22. World's Fair0.94
    23. TRACK 5• JULY 1, 20260.96
    24. Evals1.00
  • 2:10 #7 skipped

    shot 7·duplicate of #6

  • 2:26 #8 skipped

    shot 8·duplicate of #6

  • 2:47 #9 done30 line(s)

    shot 9·sharpness 1741.8

    1. ANATOMY0.94
    2. AlEngineer0.99
    3. What Is a Skill?0.96
    4. World's Fair0.98
    5. Customizing agent behavior without retraining1.00
    6. SKILL FILE LAYOUT1.00
    7. my-skill/0.92
    8. SKILL.md1.00
    9. - The only required file0.97
    10. scripts/1.00
    11. - Reusable code agent can run0.98
    12. A skill is a versionable folder containing a SKILL. md file0.99
    13. references/1.00
    14. - Docs agent reads when needed0.98
    15. and supporting assets.0.99
    16. -assets/0.94
    17. - Templates or files used in output0.99
    18. Three Layers of Progressive Disclosure:1.00
    19. SKILL.MD EXAMPLE1.00
    20. • Layer 1: Frontmatter (name + description)1.00
    21. description: Use this skill for generative video editing.0.99
    22. name: gemini-omni-flash-api0.99
    23. • Layer 2: SKILL.md Body0.99
    24. • Layer 3: References & Scripts0.97
    25. # Gemini Omni Flosh Skill0.99
    26. This skill uses the Gemini Omni Flash model ('g0.98
    27. 4 / 250.84
    28. World's Fair0.95
    29. TRACK 5· JULY 1, 20260.95
    30. Evals1.00
  • 3:25 #10 done21 line(s)

    shot 10·sharpness 1976.9

    1. •CATEGORIES0.95
    2. AlEngineer0.98
    3. Capability vs. Preference Skills0.98
    4. World'sFair1.00
    5. Knowing when skills expire1.00
    6. Capability Skills1.00
    7. Preference Skills1.00
    8. Teach models what they can't do consistently yet1.00
    9. Encode team workflows and conventions1.00
    10. Temporary (retire as models improve)1.00
    11. Durable (must match team process)1.00
    12. Evals tell you when to retire a capability skill!0.99
    13. Evals protect against workflow regressions1.00
    14. EXAMPLES1.00
    15. EXAMPLES1.00
    16. PDF parsing, custom internal APIs, database schemas, framework migration rules.0.99
    17. Code review checklists, PR formatting rules, git workflow patterns, deployment procedures.1.00
    18. 5 / 250.85
    19. World's Fair0.98
    20. TRACK 5· JULY 1, 20260.95
    21. Evals1.00
  • 3:36 #11 skipped

    shot 11·duplicate of #10

  • 4:04 #12 skipped

    shot 12·duplicate of #10

  • 4:25 #13 done21 line(s)

    shot 13·sharpness 1598.8

    1. EFFICACY DATA0.94
    2. AlEngineer0.98
    3. Do Skills Work? (SkillsBench 1.1)0.98
    4. World's Fair0.99
    5. Curated skills jump resolution rate from 33.9% → 50.5% (+16.6 pts)0.97
    6. 0.71.00
    7. Skill Lifl0.88
    8. 0.61.00
    9. 0.51.00
    10. 0.41.00
    11. 0.31.00
    12. 0.20.99
    13. 0.11.00
    14. 0.01.00
    15. GCL0.61
    16. PT.550.52
    17. Curated human-written skills boost task resolution by +16.6 points overall (33.9% → 50.5%).0.99
    18. 6 / 250.84
    19. World's Fair0.99
    20. TRACK 5· JULY 1, 20260.94
    21. Evals1.00
  • 5:02 #14 done41 line(s)

    shot 14·sharpness 1859.5

    1. EFFICACY WARNINGS0.98
    2. AlEngineer0.99
    3. Self-Generated & Bloated Skills Fail1.00
    4. World'sFair1.00
    5. Why auto-generated skills and bloat destroy accuracy1.00
    6. 1001.00
    7. No Skills0.98
    8. Claade0.98
    9. SKILL LENGTH VS. PERFORMANCE LIFT1.00
    10. 801.00
    11. Self-Gienerated0.97
    12. Curated Skills0.95
    13. Codex1.00
    14. Gemini1.00
    15. Compact (< 200 lines): +19.0% Lift0.99
    16. Pas l ()0.52
    17. 601.00
    18. Fast, low token overhead, high precision.1.00
    19. 401.00
    20. 43.00.91
    21. 34.90.95
    22. 46.80.99
    23. 35.50.98
    24. Standard (200-500 lines): +21.5% Lift (Sweet Spot)0.99
    25. Optimal balance of guidance and reasoning headroom.0.99
    26. 201.00
    27. Detailed (500-1000 lines): +14.5% Lift0.99
    28. (Claude Code)0.98
    29. Opus 4.70.99
    30. GPT-5.51.00
    31. (Codex)1.00
    32. Gemini 3.1 Pro0.99
    33. (Gemini CLI)1.00
    34. Instruction drift begins to degrade reasoning.1.00
    35. Comprehensive (> 1000 lines): +0.7% Lift (No-Op)0.99
    36. Self-generated skills hurt accuracy: -8.1 to -11.5 point loss.1.00
    37. Bloat burns reasoning tokens with zero gain.0.99
    38. 7 / 250.89
    39. World's Fair0.96
    40. TRACK 5· JULY 1, 20260.95
    41. Evals1.00
  • 5:51 #15 done27 line(s)

    shot 15·sharpness 2878.4

    1. TRIGGERS1.00
    2. AlEngineer0.97
    3. Triggering Skills: Model vs. User0.99
    4. World'sFair1.00
    5. Deciding how and when the skill is injected into context0.99
    6. MODEL-INVOKED (AUTONOMOUS)1.00
    7. USER-INVOKED (SLASH COMMAND)0.98
    8. name: commit-formatter0.98
    9. name: cleanup-worktree1.00
    10. description: Format git commit messages using team rules1.00
    11. disable-model-invocation: true1.00
    12. • Agent fires it autonomously based on the frontmatter0.99
    13. Only triggered when explicitly called by name (e.g. /commit).0.99
    14. PRESENTED BY1.00
    15. description.1.00
    16. Zero ambient context cost until invoked.0.99
    17. Microsoft1.00
    18. Frontmatter sits in context window on every single turn.0.99
    19. Use for workflows you execute intentionally.0.99
    20. Use when the agent must decide on its own when to trigger.0.99
    21. Set disable-model-invocation: true in frontmatter.1.00
    22. Model-invocation is the only option available for production1.00
    23. agents.1.00
    24. 8 / 250.82
    25. World's Fair0.98
    26. TRACK 5· JULY 1, 20260.96
    27. Evals1.00
  • 6:25 #16 skipped

    shot 16·duplicate of #15

  • 6:35 #17 skipped

    shot 17·duplicate of #15

  • 7:07 #18 done24 line(s)

    shot 18·sharpness 2321.9

    1. • TIP 10.93
    2. AlEngineer0.99
    3. Nail the Description (The Trigger)1.00
    4. World'sFair1.00
    5. The trigger mechanism causes 50%+ of all skill failures0.99
    6. • The frontmatter description is the primary trigger mechanism1.00
    7. DESCRIPTION TRIGGER COMPARISON1.00
    8. for model-invoked skills.1.00
    9. Vague descriptions cause the skill to miss triggers or hijack1.00
    10. X Too Vague0.93
    11. unrelated prompts.1.00
    12. "Helps with documents"0.99
    13. "API helper"1.00
    14. Include both the "what" (capability) and the "when" (trigger0.99
    15. context) in the description.1.00
    16. Specific & Actionable1.00
    17. "Create, edit, and analyze .docx files. Use for tracked changes, comments,0.99
    18. Real Result: Rewriting the description alone fixed 5 of 7 failures0.99
    19. formatting, or text extraction."0.99
    20. in our evaluation suite.1.00
    21. 9 / 250.86
    22. World's Fair0.97
    23. TRACK 5· JULY 1,20260.96
    24. Evals1.00
  • 7:31 #19 done24 line(s)

    shot 19·sharpness 2308.9

    1. • TIP 20.92
    2. AlEngineer0.97
    3. Write Directives Instead of Essays1.00
    4. World'sFair1.00
    5. Models follow clear directives better than inferring passive trivia0.99
    6. Directives drive action; passive explanations become ignored1.00
    7. trivia.1.00
    8. GOOD VS. BAD DIRECTIVES1.00
    9. A 5-line code snippet beats a 5-paragraph explanation every0.99
    10. XPassive Essay0.96
    11. time.1.00
    12. handles session state automatically."0.99
    13. "The Interactions API is recommended for multi-turn chat because it1.00
    14. Explain the reasoning behind rules to help the model1.00
    15. generalize across edge cases.1.00
    16. Active Directive1.00
    17. "Always use client.interactions.create() for chat. Never use the0.99
    18. Avoid overfitting to specific prompts by writing directives that1.00
    19. legacy generate_content API."1.00
    20. scale.1.00
    21. 10 / 250.95
    22. World's Fair0.95
    23. TRACK 5· JULY 1, 20260.95
    24. Evals1.00
  • 8:08 #20 done23 line(s)

    shot 20·sharpness 2688.6

    1. • TIP 30.84
    2. AlEngineer0.97
    3. Keep It Lean (Layer Information)1.00
    4. World'sFair1.00
    5. Progressive information loading saves context for the actual task0.99
    6. Frontmatter (name + description) sits in context on every turn1.00
    7. – keep it minimal.0.95
    8. Layer 1: Frontmatter (Always Loaded)0.99
    9. name + description sit in context on every turn.0.98
    10. Keep the main SKILL. md body under 500 lines to preserve0.99
    11. reasoning token headroom.0.99
    12. Layer 2: SKILL.md Body (Loaded on Trigger)0.99
    13. Move detailed docs, scripts, and multi-page guides into1.00
    14. Core instructions injected into context when skill activates.1.00
    15. external reference files.1.00
    16. Layer 3: References & Scripts (Loaded on Demand)0.99
    17. External references incur zero context cost until the agent0.99
    18. External files read or executed only when explicitly needed.1.00
    19. explicitly reads them.0.99
    20. 11/ 250.96
    21. World's Fair0.94
    22. TRACK 5· JULY 1, 20260.96
    23. Evals0.91
  • 8:45 #21 skipped

    shot 21·duplicate of #18

  • 8:59 #22 skipped

    shot 22·duplicate of #18

  • 9:27 #23 done27 line(s)

    shot 23·sharpness 2472.5

    1. • TIP 40.93
    2. AlEngineer0.99
    3. Set the Right Level of Freedom0.97
    4. World's Fair0.96
    5. Describe what you want, not the step-by-step path to get there1.00
    6. Dictating every step strips an agent's ability to adapt, recover0.99
    7. FREEDOM LEVEL COMPARISON0.99
    8. from errors, or find better approaches.1.00
    9. XRigid Step-by-Step0.98
    10. Describe the desired outcome rather than enforcing a rigid1.00
    11. "Step 1: Read config.json0.99
    12. Step 2: Extract port1.00
    13. procedural path.0.97
    14. Step 3: Edit line 40.99
    15. Step 4: Save file"1.00
    16. Provide constraints, not procedures ("Always run tests before1.00
    17. opening PR", not "Step 1, Step 2...").1.00
    18. Goal & Constraints0.99
    19. If exact step-by-step execution is required, write a script0.99
    20. Ensure file parses correctly.1.00
    21. "Update database port in config.json to 5432.0.99
    22. instead of a skill.1.00
    23. Always run tests before opening PR."1.00
    24. 12 / 250.87
    25. World's Fair0.89
    26. TRACK 5· JULY 1,20260.96
    27. Evals1.00

Transcript

275 cues· 3,742 words· 19,875 chars

  1. 0:12 Yeah, so hi, everyone.
  2. 0:13 My name is Philipp.
  3. 0:14 I'm based out of Germany.
  4. 0:16 I am part of the Google DeepMind team, mostly working on Gemini API and agents.
  5. 0:20 And we are going to talk about why you should not ship skills without evals.
  6. 0:25 And maybe before we start, I need a little bit of your help.
  7. 0:27 So if you could raise your hands if you use coding agents to write code.
  8. 0:32 So yeah, hopefully every hand goes up, right?
  9. 0:35 And do you use skills with it?
  10. 0:38 OK.
  11. 0:39 Do you have evals for those skills?
  12. 0:41 OK, that's not a lot of hands.
  13. 0:44 Everyone uses skills.
  14. 0:45 No one has evals.
  15. 0:46 Hopefully, we can fix that today.
  16. 0:48 And very important is wipe checks fail in productions.
  17. 0:51 And SkillBench is a very popular and nice eval or benchmark, which indexed over 50,000 skills from GitHub and tried to look into them.
  18. 1:01 And almost none of those skills had evals.
  19. 1:04 Most of them were AI written.
  20. 1:07 not really tested, and it's very hard to know if your skill is good or bad because agents are really non-deterministic.
  21. 1:14 So you might not know if your task fails because your skill is bad or if your task fails because it's way too challenging for the model.
  22. 1:23 Very important, before we go into it, I want to really make sure that we know the difference between the agents we use and the agents we build.
  23. 1:31 Most of us use agents for writing code, doing productivity work, that's the agents we use.
  24. 1:36 It's like anti-gravity, cursor, cloud code, and there you are the engineer, and you have context about skills.
  25. 1:43 If you write some prompt to, I don't know,
  26. 1:46 help me build a new Gemini API feature.
  27. 1:49 And if your agent does not invoke the skill on the first time, you will notice it very quickly.
  28. 1:54 You stop your task and reprompt it or use slash commands for triggering those skills.
  29. 2:00 When you build an agent inside your application for consumer or customers, they have no idea about what a skill is.
  30. 2:07 They don't start their prompt with use customer support skill to help me refund or use refund skill to
  31. 2:15 help me solve my problems.
  32. 2:16 So there's a big difference between the agents we use and how we use skills and the agents we build and how our customers might want to use skills in like the context there.
  33. 2:27 And what is a skill?
  34. 2:28 I mean, every one of us knows, hopefully by now what a skill is.
  35. 2:31 It's like basically really a folder with a skills MD file in it and then some additional assets to make that skill really work.
  36. 2:38 And the big difference with skills is that they work on progressive disclosure.
  37. 2:43 So most of the skills start very small.
  38. 2:45 So you have the title and the description.
  39. 2:48 The description is normally part of the model's context.
  40. 2:52 So the model knows when to use the skill.
  41. 2:54 Second layer is we have the skills body with more instructions, more details, and hopefully more references to external files.
  42. 3:01 And then you can really go deep in those reference files where there's all of the context the model needs to discover to solve the task.
  43. 3:08 And I like to differentiate between two kinds of skills.
  44. 3:11 So there are capability skills and preference skills.
  45. 3:13 Capability skills teach models something they cannot do consistently at the moment.
  46. 3:19 Maybe it's like, I don't know, like tracing some logs, creating a new React app,
  47. 3:24 And those capability skills are temporary.
  48. 3:27 So the better our model gets, the more likely it is that we can remove those skills.
  49. 3:32 And evalts will tell us when we can retire skill and when not.
  50. 3:36 And then we have preference skills.

Chapters

  1. 0:00 Introduction: Why skills need evals
  2. 0:25 The problem with current agent workflows
  3. 1:25 Agents we use vs. agents we build
  4. 2:28 Defining a 'Skill' and progressive disclosure
  5. 3:08 Capability skills vs. preference skills
  6. 4:17 Do skills actually work? (Skillsbench data)
  7. 5:39 Model-triggered vs. user-invoked skills
  8. 6:39 Best practices for writing skill descriptions
  9. 8:30 Structuring complex, multi-layered skills
  10. 9:04 Defining goals and constraints (avoiding rigid steps)
  11. 9:56 Don't skip negative cases
  12. 10:36 Testing strategy: Evals and regressions
  13. 11:05 Removing 'no-ops' for cost efficiency
  14. 11:47 Knowing when to retire a skill
  15. 12:22 Practical example: Gemini interactions API
  16. 13:36 Building a lightweight eval harness
  17. 14:35 Using regex and LLMs as judges
  18. 17:14 Top 10 best practices summary
  19. 20:20 Homework: How to start testing your skills

Open at this second