Videos JomVvNDjGb8
Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face
Scene timeline
93 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 213
- whisperx 213
- chunks
- 33
- from 213 cues
- keyframes
- 85
- kept of 93 captured
- frames with text
- 81
- 2,276 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 10.2 MB
- word timings on 213 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 18:27 | 1m 46s |
stt |
done | — | 2026-08-10 18:29 | 21s |
chunk |
done | — | 2026-08-10 18:29 | 0s |
text_embed |
done | — | 2026-08-10 19:55 | 0s |
keyframe |
done | — | 2026-08-10 18:29 | 2m 16s |
ocr |
done | — | 2026-08-10 18:31 | 54s |
frame_embed |
done | — | 2026-08-10 19:55 | 15s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- ai sys0.97
- AlEngineer0.99
- OpenAl0.99
- AlEngineer1.00
- EUROPE0.98
- EUROPE0.99
- engin1.00
- AlEngineer0.99
- WorkOS1.00
- AlEngineer0.99
- IIElev0.89
- EUROPE1.00
- EUROPE0.99
- △Trigger.dev0.94
- Cnnala DeepMind0.79
- AlEngineer0.99
- EUROPE1.00
- AlEngineer0.99
- arize1.00
- AlEngineer0.98
- EUROPE0.99
- EUROPE1.00
- PRESENTED BY1.00
- Google DeepMind1.00
- Unblocked1.00
- AIE1.00
- AlEngineer0.98
- EUROPE1.00
-
- Araintru your coding0.96
- AIEngineer0.95
- agent should do0.96
- AIEngineer0.95
- EURC0.91
- EUROPE1.00
- eer1.00
- ai systems1.00
- OpenA1.00
- Google DeepMir0.96
- eer1.00
- engineering1.00
- croso0.89
- RY1.00
- ATEngineer0.86
- crosoft1.00
- AIEngineer0.95
- EUROPE1.00
- EUROPE1.00
- Google DeepMind1.00
- AlEngineer0.96
- EUROPE1.00
- Modal1.00
- ineer1.00
- BeNCORD0.95
-
- your coding0.96
- Braintrus1.00
- agent should do1.00
- eer1.00
- ai systems0.97
- engineering1.00
- croso1.00
- neer1.00
- Google DeepMind1.00
- AlEngineer0.94
- EUROPE1.00
-
- Takeaways1.00
- Praintru0.90
- Fun part: We can use code agents to do Al1.00
- Sys Engineering / ML Engineering0.99
- eer1.00
- Boring part: We need standard repos and1.00
- primitives to know what's going on0.99
- cros1.00
- 10.76
- ineer1.00
- Google DeepMind1.00
- AlEngineer0.98
- EUROPE1.00
-
- eepMind1.00
- AlEngineer0.97
- Takeaways1.00
- EUROPE0.93
- ineer0.96
- ATrige0.79
- Zed0.94
- Fun part: We can use code agents to do Al0.99
- Sys Engineering / ML Engineering0.99
- gineer0.91
- o4j0.95
- AIE1.00
- Boring part: We need standard repos and0.99
- primitives to know what's going on1.00
- AlEngineer0.99
- # Braintrust0.96
- WorkOS1.00
- OpenAI0.93
- EUROPE1.00
-
- eepMind0.99
- AlEngineer1.00
- EUROPE0.99
- Coding Agents0.97
- jineer0.91
- ΔTrigger.dev0.89
- Zed0.99
- AlEngine0.97
- EUROPE0.93
- I was inspired by this so I wanted to see if Claude Code can get into my0.99
- Andrej Karpathy@karpathy - Dec 28, 20250.93
- Q0.90
- "I also get some programmers are eager to tune it out. The hype drones on,0.98
- DHH@dhh Jan 70.92
- Lutron home automation system.1.00
- the fantastical claims are still far off, and there's uncertainty where this0.99
- leaves the profession. But that's not reason to miss out on this incredible0.99
- gineer0.99
- come0.99
- - it found my Lutron controllers on the local wifi network0.98
- - checked for open ports, connected, got some metadata and identified the1.00
- moment in human history!"0.98
- o4j0.93
- AlEngineer1.00
- shades, HVAC temperature control, motion sensors etc.)0.99
- - searched the internet, found the pdf for my system1.00
- - instructed me on what button to press to pair and get the certificates1.00
- - it connected to the system and found all the home devices (lights,0.98
- devices and their firmware0.98
- Promoting Al agents1.00
- me. Partly because the models got better, but mor...0.99
- At the end of last year, Al agents really came alive for0.99
- world.hey.com1.00
- - it turned on and off my kitchen lights to check that things are working0.99
- (lol!)0.96
- 601.00
- t7 1220.86
- 1K0.69
- I0.98
- l 150K0.82
- 口0.66
- have been accepted.1.00
- AlEngineer0.97
- Braintrust1.00
- WorkOS1.00
- OpenAl0.92
- EUROPE1.00
-
- Coding Agents1.00
- Andrej Karpathy@karpathy·Dec 28, 20250.97
- DHH@dhh Jan70.96
- I was inspired by this so I wanted to see if Claude Code can get into my0.98
- "I also get some programmers are eager to tune it out. The hype drones on,0.99
- Lutron home automation system.1.00
- the fantastical claims are still far off, and there's uncertainty where this0.99
- leaves the profession. But that's not reason to miss out on this incredible0.99
- - it found my Lutron controllers on the local wifi network0.99
- moment in human history!"0.99
- - checked for open ports, connected, got some metadata and identified the0.99
- - searched the internet, found the pdf for my system0.98
- - instructed me on what button to press to pair and get the certificates0.99
- devices and their firmware0.99
- world.hey.com1.00
- Promoting Al agents1.00
- - it connected to the system and found all the home devices (lights,0.99
- At the end of last year, Al agents really came alive for1.00
- shades, HVAC temperature control, motion sensors etc.)1.00
- me. Partly because the models got better, but mor...0.98
- - it turned on and off my kitchen lights to check that things are working0.99
- (lol!)0.92
- 601.00
- t1220.99
- 1K0.98
- I0.94
- 150K1.00
- have been accepted.1.00
-
- the bosses1.00
- 1. write custom kernels1.00
- interactively with an agent1.00
- 2. yolo an agent to finetune1.00
- an LLM0.96
- 3. setup a multi-agent0.98
- autoreasearch lab1.00
-
- Mind1.00
- AlEngineer0.92
- BUROPE0.99
- the bosses1.00
- A Triggerd0.87
- 1. write custom kernels0.97
- AIEr0.88
- interactively with an agent1.00
- 2. yolo an agent to finetune1.00
- COI0.87
- an LLM1.00
- 3. setup a multi-agent0.98
- AlEngin1.00
- EUROPE0.98
- autoreasearch lab1.00
- Google DeepMind1.00
- AlEngineer0.98
- EUROPE1.00
-
- the bosses1.00
- 1. write custom kernels0.99
- interactively with an agent1.00
- 2. yolo an agent to finetune1.00
- an LLM1.00
- 3. setup a multi-agent0.97
- autoreasearch lab1.00
- # Braintrust0.97
- WorkOS OpenAI0.92
- AlEngineer0.98
- EUROPE1.00
-
- Microsoft1.00
- BOSS 10.99
- OpenA1.00
- AlEngine0.92
- Son0.99
- stripe0.92
- Snork1.00
- human + coding agent1.00
- Ägent Skills for Kernels0.99
- AlEngineer1.00
- # Braintrust0.96
- WorkOS OpenAI0.96
- EUROPE1.00
-
- Microsoft0.97
- BOSS 10.98
- OpenAl0.94
- NEng0.95
- NEng0.89
- Sonar0.99
- stripe1.00
- Snorkel0.99
- human + coding agent1.00
- Ägent Skills for Kernels0.99
- Engineering the future of Al1.00
- AlEngineer0.99
- EUROPE1.00
-
- onAl0.92
- CLAUDE CODE0.99
- DOES NOT WRITE CUDA KERNELS0.99
- ust0.99
- imgflip.com0.94
- Google DeepMind0.97
- AlEngineer0.99
- EUROPE1.00
-
- Mind0.85
- AlEngineer0.99
- BUROPE0.95
- ATrigger.de0.92
- CLAUDE CODE0.99
- AlEngine1.00
- BUROPE0.91
- con1.00
- AlEngineer1.00
- EUROPE0.96
- DOES NOT WRITE CUDA KERNELS0.99
- imgflip.com0.94
- Google DeepMind0.99
- AlEngineer0.99
- EUROPE1.00
-
- ai1.00
- AlEngineer0.98
- EUROPE0.96
- WTF is a kernel?1.00
- AIE1.00
- Your Python code1.00
- GPU hardware1.00
- Runs on CPU1.00
- Thousands of threads1.00
- Mic1.00
- A function compiled to run0.99
- torch.matmul(A, B)1.00
- Kernel dispatch1.00
- on a GPU typically1.00
- executed from Python.0.98
- F.softmax(x)1.00
- AlEngin0.99
- matmul_kernel0.98
- EUROPE0.99
- Compiled for GPU1.00
- F.gelu(x)1.00
- AlEngineer0.99
- Braintrust1.00
- WorkOS1.00
- OpenAl0.93
- EUROPE1.00
-
- WTF is a kernel?1.00
- Your Python code1.00
- GPU hardware1.00
- Runs on CPU1.00
- Thousands of threads1.00
- A function compiled to run0.98
- torch.matmul(A, B)1.00
- Kernel dispatch1.00
- on a GPU typically1.00
- executed from Python.1.00
- F.softmax(x)1.00
- matmul_kernel1.00
- Compiled for GPU1.00
- F.gelu(x)1.00
-
- Efficiency in Deep Learning1.00
- Gor0.88
- DeepMin0.95
- Compute1.00
- Memory1.00
- Overhead1.00
- Time spent in computing1.00
- Time spent in communicating1.00
- Everything else1.00
- actual FLOPs1.00
- tensors1.00
- :OS0.61
- Hugging Face1.00
- AlEngineer0.99
- # Braintrust0.98
- WorkOS1.00
- OpenAl0.96
- EUROPE1.00
-
- Zed0.97
- Efficiency in Deep Learning1.00
- WorkOS0.91
- AlEngine0.90
- arize1.00
- neo4j0.83
- Compute1.00
- Memory1.00
- Overhead1.00
- Modal1.00
- Time spent in computing0.99
- Time spent in communicating1.00
- Everything else1.00
- actual FLOPs1.00
- tensors1.00
- Hugging Face0.96
- Google DeepMind1.00
- AlEngineer0.98
- EUROPE1.00
-
- Efficiency in Deep Learning1.00
- rkOS0.87
- Compute1.00
- Memory1.00
- Overhead1.00
- Time spent in computing1.00
- Time spent in communicating1.00
- Everything else1.00
- actual FLOPs1.00
- tensors1.00
- Jst0.96
- Hugging Face1.00
- Google DeepMind0.99
- AlEngineer0.99
- EUROPE1.00
-
- flash-attention1.00
- Flash Attention custom kernel1.00
- Increase arithmetic density0.99
- 21.9k1.00
- 2.3k1.00
- 1451.00
- Spend less time in communicating tensors0.99
- Keep the GPUs warm0.97
- Br.0.82
- AIE1.00
- AlEngineer0.99
- Braintrust0.97
- WorkOS1.00
- OpenAI0.93
- EUROPE1.00
Transcript
213 cues· 3,452 words· 18,282 chars
- 0:15 Hi, everyone.
- 0:16 As you heard, I'm Ben from Hugging Face.
- 0:18 And the talk that I'm going to present to you today is called Your Coding Agent Should Do AI Systems Engineering.
- 0:24 So there are two main takeaways that I want you to get from this talk.
- 0:29 One, and probably the fun part, is that we can use coding agents to tackle the hardest engineering problems in AI, so systems engineering and machine learning engineering.
- 0:37 And maybe the boring part is that in order to do this, we're gonna need standard repos, and we're gonna need those on the hub.
- 0:44 And in many cases, we already have them.
- 0:47 So I think in this case, I'm preaching to the choir here, but in case you haven't noticed, coding agents have been accepted.
- 0:55 Many of us have been using them for a few years, but in the last few months, they seem to have crossed a sort of acceptance gradient where a broader group of people are using them.
- 1:05 So with this in mind, how do we keep our careers, our engineering kind of contemporary?
- 1:11 And how do we keep challenging ourselves in new areas?
- 1:14 And my proposal is that we need to go kind of closer to the silicon and tackle harder problems.
- 1:20 And that's where AI systems engineering comes in.
- 1:23 I've broken this talk down into three progressively more complex steps and more autonomous steps as well.
- 1:30 And I've defined those like three bosses from games.
- 1:34 The first one is a hybrid approach where you interactively use an agent to write a CUDA kernel.
- 1:43 The second is a zero shot task where an agent takes a prompt and trains an LLM.
- 1:50 on Hugging Face.
- 1:51 The third is a multi-agent auto research setup, like a kind of automated AI lab.
- 1:57 So let's get started on the first boss, right?
- 2:00 This is writing CUDA kernels.
- 2:02 So for a while writing custom kernels was seen as this unattainable goal for the humble agent.
- 2:09 They required complex DSLs, they required integration with relevant hardware to be benchmarked and to be tested and it was seen as something that couldn't be achieved by agents.
- 2:21 However, that in most cases was wrong.
- 2:23 If you look at kernel hackathons like those on GPU mode, the recent AMD hackathon, if you look at papers like kernel bench, you'll see that agents are able to write valid and optimized CUDA kernels.
- 2:37 And that's really cool and something that totally inspires me.
- 2:40 I'm a part of GPU mode, I contribute to that, and something that I think everyone should be doing.
- 2:45 However, what do we do with them?
- 2:47 How do we distribute them and how do we get them into our inference engine so that we actually are using these optimized kernels that we're generating?
- 2:54 And that's part of the question of this part of the talk.
- 2:58 Let's take a step back now and just say what a kernel is, right?
- 3:02 So when you run an AI model on a GPU, the actual work is executed through a kernel.
- 3:09 This will be defined in a relevant language for that hardware, and it will use relevant features to that hardware that may not be available on other hardware.
- 3:18 We can write custom kernels that will take advantage of that hardware for a specific math operation, kind of squeeze everything we can out of it so that the model will infer faster.
- 3:29 In general, this requires a lot of expertise about writing CUDA kernels, about the hardware.
- 3:34 And it's also a bit of an insulation hell as you deal with a pretty large install matrix from hardware to software to generations and versions of, say, CUDA and these kind of issues.
- 3:44 So in short, it's hard.
- 3:48 Efficiency in deep learning, so efficiency in kernels, is split into three main sections.
- 3:53 One, compute.
- 3:54 Two, memory.
- 3:54 And three, overhead.
- 3:56 Compute is the flops.
- 3:58 These are the matrix multiplications and the real math of the process.
- 4:02 Memory is the time spent moving data or tensors around memory, typically from slow to fast memory.
- 4:09 And overhead is basically everything else, the Python environment, PyTorch dispatch of those kernels, these kinds of things.
- 4:15 In general, most people might assume that the compute is the bottleneck here because it's doing most of the math, right?
- 4:24 That's not correct.
- 4:25 In most cases, memory is usually the bottleneck.
- 4:29 And that's because a modern GPU, let's take a H100 for example, can do a petaflop a second of computation, but its memory bandwidth is three terabytes.
- 4:38 So in short, the GPU is often waiting idle for these tensors to come back for them to be computed.
loading
Chapters
- 0:00 Introduction to AI Systems Engineering
- 1:59 Boss 1: Writing and Distributing CUDA Kernels
- 3:48 Efficiency in Deep Learning
- 6:08 Using Skills for Agentic Workflows
- 8:37 Benchmarking and Evaluating Skills with Upskill
- 9:26 Boss 2: End-to-End Fine-tuning of LLMs
- 10:16 Boss 3: Multi-Agent Auto Research Labs
- 12:09 Architecture of the Multi-Agent Research System
- 13:40 Implementing the Research Agent in OpenCode
- 15:28 Monitoring Experiments with Trackio
- 16:45 Final Takeaways and Conclusion