Videos 3_gYbhABcAE
Why (Senior) Engineers Struggle to Build AI Agents — Philipp Schmid, Google DeepMind
Scene timeline
26 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 105
- whisperx 105
- chunks
- 18
- from 105 cues
- keyframes
- 19
- kept of 26 captured
- frames with text
- 19
- 342 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.4 MB
- word timings on 105 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 14:02 | 1m 40s |
stt |
done | — | 2026-08-10 14:04 | 11s |
chunk |
done | — | 2026-08-10 14:04 | 0s |
text_embed |
done | — | 2026-08-10 19:50 | 0s |
keyframe |
done | — | 2026-08-10 14:04 | 52s |
ocr |
done | — | 2026-08-10 14:05 | 7s |
frame_embed |
done | — | 2026-08-10 19:50 | 3s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- Why (Senior) Engineers0.98
- Struggle to Build Al0.98
- Agents1.00
- 5 mental model collisions0.99
- engineering. and how to fix them0.99
- AlEngineer0.99
- EUROPE1.00
-
- TOP0.85
- Engineering Instincts1.00
- PHILIPP SCHMID· GOOGLE DEEPMIND0.98
- Strict Types, Determinism1.00
- Why (Senior) Engineers1.00
- ★1.00
- AIE1.00
- Struggle to Build Al0.99
- BOTECK0.75
- ★1.00
- ★1.00
- ★1.00
- Agents1.00
- Mental Model Collision1.00
- Fighting the probability of LLMs1.00
- 5 mental model collisions from traditional0.99
- engineering, and how to fix them.1.00
- BT0.88
- Agent Reality1.00
- Loops, Recovery, Contextual Meaning0.99
- 1/90.99
- AlEngineer0.97
- AlEngineer0.98
- EUROPE1.00
- EUROPE1.00
-
- THE PARADOX1.00
- Traffic Controller vs. Dispatcher0.99
- Traffic Controller Mindset: Traditional software engineering is Deterministic.1.00
- You own the roads, signals, and laws.1.00
- Dispatcher Mindset: Agent engineering is Probabilistic. You give instructions1.00
- to a driver tracking ambiguity.1.00
- ★1.00
- ★1.00
- AIE1.00
- The Trap: The better the engineer, the harder they fight the model. They try to0.99
- "code away" non-determinism with if-statements and rigid schemas.1.00
- ★1.00
- ★1.00
- 3/90.97
- AlEngineer0.98
- AlEngineer0.98
- EUROPE1.00
- EUROPE1.00
-
- Traffic1.00
- vs. Dispatcher1.00
- AlEngineer0.99
- EUROPE1.00
-
- PART 10.98
- Text is the New State0.97
- The Trap: Treating the real world as enums1.00
- and booleans because it feels safe.1.00
- The Collision: Forcing natural language0.99
- intent into discrete booleans lobotomizes1.00
- AIE1.00
- ★1.00
- the context.1.00
- #Software Eng: Lobotomizing intent0.99
- ★1.00
- The Fix: Preserve semantic meaning1.00
- {"plan_id": "123", "status": "APPROVED"}0.99
- ★1.00
- ★1.00
- through raw strings so the agent can adapt0.98
- #Agent Eng: Preserving semantic meaning0.99
- intelligently downstream.1.00
- {"plan_id": "123", "text": "Approved, but focus on the US market and ignore CA."}1.00
- # Another example: boolean vs. semantic preference0.99
- is_celsius: true1.00
- #One-size-fits-all1.00
- "I prefer Celsius for weather, Fahrenheit for cooking." # Task-aware0.99
- 4/90.92
- Engineering the future of Al1.00
- AlEngineer0.98
- EUROPE1.00
-
- PART 21.00
- Hand Over Control0.99
- The Trap: In microservices, user intent maps to a route. We instinctively hard-0.99
- Customer Support0.97
- code the flow.1.00
- The Collision: Real interactions loop, backtrack, and pivot. Hard-coded paths0.98
- I want to cancel my subscription.0.99
- lose customers.1.00
- ★1.00
- INTENT: CHURN0.98
- AIE1.00
- ★1.00
- The Fix: Trust the agent to navigate. You're a dispatcher, not a traffic0.99
- ★1.00
- controller.0.98
- for the next 3 months.1.00
- I understand. Before you go, I can offer you 50% off0.98
- ★1.00
- ★1.00
- Key insight: Describe what you want, not the path to get there. Provide constraints, not0.99
- procedures.1.00
- Actually, yeah, that works.1.00
- INTENT: RETENTION0.99
- If you hard-coded the churn flow, you just lost a customer.1.00
- 5/90.83
- Engineering the future of Al0.98
- AlEngineer1.00
- EUROPE1.00
-
- Hand1.00
- AlEngineer0.99
- EUROPE1.00
-
- PART 31.00
- Errors Are Just Inputs1.00
- The Trap: Failing fast and crashing the entire execution on a minor schema0.99
- fault.1.00
- LOOP0.93
- Agent Loop Running0.99
- The Collision: An agent run takes 5 minutes and costs $0.50. Crashing at0.99
- step 4 of 5 is unacceptable.1.00
- ★1.00
- AIE1.00
- The Fix: Catch the error and feed it back so the agent self-corrects.0.99
- ★1.00
- ★1.00
- ★1.00
- FAULT0.78
- Tool Execution Error Occurs1.00
- #Traditional: Fail Fast0.99
- raise RuntimeError(f"Failed: {e}") # Crash0.98
- #Agent: Feedback Loop0.98
- Traditional Way0.99
- Agent Way0.99
- return f"Error: {e}. Try a different approach." # Recover0.98
- Crash run1.00
- Feed back as String1.00
- $0.50 wasted0.96
- Agent Retry & Succeed0.99
- 6190.99
- AlEngineer0.97
- AlEngineer1.00
- EUROPE1.00
- EUROPE1.00
-
- PART 31.00
- Errors Are Just Inputs1.00
- The Trap: Failing fast and crashing the entire execution on a minor schema0.99
- fault.1.00
- LOOP0.96
- Agent Loop Running0.99
- The Collision: An agent run takes 5 minutes and costs $0.50. Crashing at0.99
- step 4 of 5 is unacceptable.0.98
- ★1.00
- AIE1.00
- The Fix: Catch the error and feed it back so the agent self-corrects.0.99
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- FAULT0.87
- Tool Execution Error Occurs1.00
- #Traditional: Fail Fast1.00
- raise RuntimeError(f"Failed: {e}") # Crash0.99
- #Agent: Feedback Loop0.98
- Traditional Way1.00
- Agent Way1.00
- return f"Error: {e}. Try a different approach." # Recover0.98
- Crash run1.00
- Feed back as String1.00
- $0.50 wasted0.97
- Agent Retry & Succeed0.99
- 6190.96
- Engineering the future of Al0.99
- AlEngineer1.00
- EUROPE1.00
-
- From1.00
- Evals1.00
- AlEngineer0.99
- EUROPE1.00
-
- PART 41.00
- From Unit Tests to Evals0.99
- Evaluate, Don't Assert: Output is nondeterministic. Run 3-5 trials per0.99
- prompt and measure the distribution.0.99
- EVAL FRAMEWORK1.00
- TESTS1.00
- EVALS1.00
- Negative Cases Matter: Test that your agent ignores requests outside its0.99
- scope. Triggering blindly degrades everything.1.00
- Question1.00
- "Did it work?"1.00
- "How often?"1.00
- ★1.00
- ★1.00
- ★0.94
- AIE1.00
- ★1.00
- Grade Outcomes, Not Paths: Did the agent yield the correct result? Don't1.00
- Method1.00
- Binary assertion0.98
- Pass^k (50x)1.00
- ★1.00
- ★1.00
- grade which tools it chose to get there.1.00
- Quality1.00
- N/A1.00
- LLM-as-Judge1.00
- ★1.00
- ★1.00
- ★1.00
- Debug1.00
- Stack trace1.00
- Tracing1.00
- Ready: 45/50 passes with quality score 4.5/5. You're managing risk, not eliminating variance.0.99
- 7190.99
- Braintrust1.00
- WorkOS OpenAI0.97
- AlEngineer1.00
- EUROPE1.00
-
- PART 41.00
- From Unit Tests to Evals0.98
- Evaluate, Don't Assert: Output is nondeterministic. Run 3-5 trials per0.99
- prompt and measure the distribution.0.99
- EVAL FRAMEWORK0.98
- TESTS1.00
- EVALS1.00
- Negative Cases Matter: Test that your agent ignores requests outside its1.00
- scope. Triggering blindly degrades everything.0.98
- Question1.00
- "Did it work?"0.97
- "How often?"0.96
- ★1.00
- AIE1.00
- Grade Outcomes, Not Paths: Did the agent yield the correct result? Don't0.99
- Method1.00
- Binary assertion0.99
- Pass^k (50x)0.99
- ★1.00
- grade which tools it chose to get there.1.00
- Quality1.00
- N/A1.00
- LLM-as-Judge1.00
- ★1.00
- ★1.00
- ★1.00
- Debug1.00
- Stack trace0.99
- Tracing1.00
- Ready: 45/50 passes with quality score 4.5/5. You're managing risk, not eliminating variance.0.98
- 7190.98
- AlEngineer0.98
- AlEngineer1.00
- EUROPE1.00
- EUROPE1.00
-
- PART 51.00
- Agents Evolve, APls Don't0.98
- A0.51
- Human-Grade API0.98
- ✓ Agent-Ready API0.96
- Agents are literalists and will hallucinate ambiguous parameters.1.00
- Explicit, self-documenting semantic interfaces.1.00
- ★1.00
- AIE1.00
- ★1.00
- delete_item(id)1.00
- delete_item_by_uuid(uuid: str)1.00
- No type, no description. Agent guesses id format.1.00
- "Deletes an item by UUID. Returns error if not found."1.00
- ★1.00
- ★1.00
- ★1.00
- ★1.00
- →0.99
- # Vague implementation1.00
- # Explicit, verbose, typed1.00
- def delete_item(id):1.00
- def delete_item_by_uuid(uuid: str):1.00
- pass1.00
- """Deletes an item by its UUID.1.00
- Returns error if not found.""0.99
- Semantic Naming0.98
- Verbose Docstrings1.00
- Just-In-Time Adaptation0.99
- 8/90.94
- AlEngineer0.97
- AlEngineer0.99
- EUROPE1.00
- EUROPE1.00
-
- Agents1.00
- PIs Don't0.96
- Human-Grade API0.98
- Vague implementation0.96
- delete_item(id):0.75
- poss0.93
- AlEngineer0.99
- EUROPE1.00
-
- SUMMARY1.00
- Trust, but Verify1.00
- ★1.00
- AIE1.00
- ★1.00
- Stop Fighting1.00
- Preserve1.00
- Design for0.97
- Evaluate Don't1.00
- Build to1.00
- ★1.00
- the Model0.98
- Meaning1.00
- Recovery1.00
- Assert1.00
- Delete1.00
- ★1.00
- ★1.00
- ★1.00
- You're a dispatcher now. Give1.00
- instructions, let it pathfind.1.00
- Context is everything.1.00
- Text is greater than booleans.1.00
- Errors are inputs in a long-0.99
- running execution loop.1.00
- Manage risk through Pass^k0.97
- and LLM-as-judge.1.00
- harness gets rewritten. Make1.00
- The Bitter Lesson: every0.99
- yours modular.0.99
- 40.79
- "You cannot code away the probability. You must manage it through evals and self-correction."1.00
- Blog philschmid.de/why-engineers-struggle-building-agents1.00
- Philipp Schmid@_philschmid0.98
- 9190.97
- Engineering the future of Al0.99
- AlEngineer1.00
- EUROPE1.00
Transcript
105 cues· 1,748 words· 9,265 chars
- 0:15 OK, cool.
- 0:15 Awesome.
- 0:16 Hi, everyone.
- 0:17 My name is Filip.
- 0:18 I work at DeepMind, everything related to agents on Gemini or Gemini API.
- 0:22 So if you have some questions afterwards, some concerns, some bugs, some issue, please let me know.
- 0:28 We are going to talk today 10 minutes about why engineers struggle to build agents.
- 0:33 And I see this every day internally at Google, but also externally at Google.
- 0:37 And I brought five examples on what's really different to how we built traditional software a few years ago and to now how we build agents.
- 0:45 And if we, on a high level, compare them, when we wrote software, we created a spec, a PRD, wrote code,
- 0:53 sometimes create a test to make sure our code works.
- 0:56 We deployed it and then our user use it.
- 0:59 When building agents, things are a little bit different.
- 1:02 We define instructions on what we want our agent to do.
- 1:05 We run it, we observe what it does, we maybe adjust our prompts, maybe we adjust our tools.
- 1:11 We run it again and we have this iterative loop of how can we improve and make our agent way more reliable, which is very different to how we build software.
- 1:21 Something I like to compare it to is like traditional software is more like we acted as a traffic controller, right?
- 1:27 We had control over the street lights or how fast you can go, which road you can use, basically how the car drives.
- 1:34 And now with agents, we are more of a dispatcher.
- 1:36 We tell the agent, hey, I want to go to London and I'm from like Germany.
- 1:41 I could use the train, I could fly, I could use my car and go like under the water.
- 1:47 And it's more about, okay, we define the goal on what we want the agent to do, but we don't define the exact step the agent needs to take to achieve that goal.
- 1:55 And I mean, every one of you has probably seen in their coding agent that sometimes it does something very weird, but at the end it achieves the outcome.
- 2:03 And that's what we want to do.
- 2:05 So starting with the first example, text is our new state.
- 2:09 I mean, traditionally we had data structures and everything was kind of mapped to Boolean or to like flags we could check.
- 2:17 Initially, when we created, for example, a deep research agent, deep research agent returns a plan to you, okay, I'm going to research this and that.
- 2:27 In traditional software, we might have had an accept plan or deny plan, but we couldn't catch semantic meaning.
- 2:35 And now what we have with LLMs is they can understand the semantic meaning.
- 2:38 So, for example, if I have a deep research agent
- 2:42 request on like doing some market research I can approve the initial plan but I can also on the same time provide additional information so maybe I want to focus on like the US market and ignore California maybe I want to provide something additional and not have like this multiple steps right traditionally I would probably set decline and then it has a follow-up I might need it to provide more input create a new plan and continue and
- 3:07 And another good example is everything related to memory and personalization we do cannot really be mapped to data structures, right?
- 3:16 The example I have here is like, I'm from Europe, so I mostly use Celsius, but what if I would like to use Fahrenheit for cooking, right?
- 3:25 Previously, we might had some flex on like the user profile is Celsius or is Europe or use Fahrenheit, but I couldn't like dynamically adjust based on the user preference, based on what they provide.
- 3:39 So really it's all about text and context.
- 3:42 I mean, it could be images, video or audio as well, but we no longer are really operating in those clear structured data concepts.
- 3:51 The other thing is we should start handing over control and the trap or the example which we might have from like previous customer support is like when a user reached out, hey, I want to cancel my subscription, I might have had a classification model which kind of classified the attend, okay, the user wants to churn and then I had a predefined workflow of okay,
- 4:17 Do you try to sell it?
- 4:18 Do you cancel the subscription?
- 4:20 But there was no dynamic kind of option to react to it dynamically.
- 4:26 And maybe instead of like,
- 4:29 We're going through the subscription cancel flow.
- 4:31 What if your agent tries to understand the meaning and offers something except to the subscription and the user changes their mind and now you have a whole different intent?
- 4:43 And it's very hard to model all of those differences and uniqueness and to do all of those
- 4:51 stateful workflows we had before.
- 4:53 So we need to trust into the LLM or hand over control that we are no longer working in those purely deterministic environments.
- 5:04 The third one is errors are just inputs.
- 5:08 So if something in your agent flow fails, we need to treat it as a normal input, as very similar to a user input.
- 5:16 In Go, we already do this, right?
- 5:18 A function call can be an error or can be a value, and we treat them kind of equally.
loading