Videos TeGsFFNqRLA
Fast Models Need Slow Developers — Sarah Chieng, Cerebras
Scene timeline
57 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 166
- whisperx 166
- chunks
- 30
- from 166 cues
- keyframes
- 53
- kept of 57 captured
- frames with text
- 53
- 1,249 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 8.1 MB
- word timings on 166 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 23:15 | 1m 18s |
stt |
done | — | 2026-08-10 23:16 | 26s |
chunk |
done | — | 2026-08-10 23:17 | 0s |
text_embed |
done | — | 2026-08-10 23:17 | 0s |
keyframe |
done | — | 2026-08-10 23:17 | 1m 58s |
ocr |
done | — | 2026-08-10 23:19 | 26s |
frame_embed |
done | — | 2026-08-10 23:19 | 10s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- WorkOs0.91
- IIElevenLabs0.96
- AlEngineer0.94
- EUROPE1.00
- A0.89
- AlEngineer0.98
- DeepMind0.94
- AlEngineer0.99
- EUROPE0.94
- EUROPE1.00
- IEngineer0.99
- :neo4j0.93
- Λarize0.91
- EUROPE1.00
- RESENTED BY1.00
- gle DeepMind1.00
- AlEngineer0.99
- AlEngineer0.98
- EUROPE1.00
- EUROPE0.97
-
- cerebras1.00
- I Sarah Chieng0.96
- neer1.00
- Braintru0.99
- Fast Models1.00
- Engineer1.00
- Need1.00
- Micros1.00
- Slow1.00
- Mc0.70
- AlEngineer0.98
- EUROPE1.00
- Developers1.00
- Google DeepMind1.00
- AlEngineer0.99
- EUROPE1.00
-
- cerebras0.99
- Sarah Chieng0.97
- Engineer1.00
- Brai1.00
- EUROPE1.00
- Fast Models1.00
- pe1.00
- AIEng0.93
- EURC0.91
- Need1.00
- Engir1.00
- Mic0.99
- EUROPE1.00
- Slow1.00
- M1.00
- AIEng0.90
- EURC0.95
- Developers1.00
- AlEngineer0.99
- # Braintrust0.98
- WorkOS1.00
- OpenAl0.93
- EUROPE1.00
-
- AlEngine0.99
- Braintrust1.00
- cerebras1.00
- Sarah Chieng0.94
- Tessl1.00
- AlEngineer0.97
- Fast Models1.00
- Snorkel0.99
- Need1.00
- AlEngineer1.00
- Sonar0.99
- Slow1.00
- Developers1.00
- Google DeepMind1.00
- AlEngineer0.99
- EUROPE1.00
-
- cerebras0.99
- Sarah Chieng0.96
- AlEngineer0.98
- Br1.00
- EUROPE1.00
- Fast Models1.00
- Opel0.85
- AIE1.00
- Need1.00
- AIEng0.96
- EUPE0.96
- Slow1.00
- AIE0.98
- Developers1.00
- Google DeepMind1.00
- AlEngineer1.00
- EUROPE1.00
-
- Sarah1.00
- gineer1.00
- Braint1.00
- @milksandmatcha - 111K subscribers - 296 videos0.91
- JROPE0.97
- instagram.com/milksandmatcha and 1 more link0.96
- [email protected]0.88
- Customize channel0.98
- Manage videosy Commrunity0.92
- Home1.00
- Videos1.00
- Shorts Playlists Posts0.99
- enAl0.94
- Engine1.00
- Papular0.94
- Oldest0.99
- EUROPE1.00
- burned out1.00
- engineer1.00
- @startup0.95
- gineer1.00
- Micro0.92
- MIT ech, hawail0.87
- WEEK IN OUR IFE cupless vlog.0.86
- nemployed at 24 | life 2 years after0.97
- my 9-5 as a startup engineer | ife 20.97
- years after MIT0.96
- JROPE0.97
- Modo0.93
- IEngine0.94
- cerebr0.95
- EUROPE1.00
- Follow1.00
- Sarah Chieng0.99
- Sarah Chieng0.99
- @SarahChieng1.00
- @Cerebras1.00
- Head of DevX @ Cerebras1.00
- prev.@ExaAiLabs, @shopthrifthouse, @MIT0.98
- sarahchieng.com Joined March 20220.94
- 1,251 Following 17.3K Followers0.97
- AlEngineer0.97
- Braintrust1.00
- WorkOS1.00
- OpenAI0.92
- EUROPE1.00
-
- Sarah1.00
- gef0.90
- AlEngineer0.99
- @milksandmatcha - 111K subscribers · 296 videos0.92
- [email protected]0.85
- EUROPE0.95
- instagram.com/milksandmatcha and 1 more link0.98
- Customize channel0.94
- Manage videos Community0.96
- Home1.00
- Videos1.00
- Shorts Playlists Posts0.98
- Papular0.95
- Oldest0.99
- Zed0.88
- burned out0.99
- engineer1.00
- @startup0.94
- WEXK IN OUR IFE couples vlog0.91
- MIT, tech, hawall0.94
- nemmployed at 24 | lfe 2 years after0.93
- my 9-5 as a startup engineer | ife 20.96
- years after MIT1.00
- AlEngineer0.94
- cerebr0.99
- EUROPE1.00
- Follow1.00
- Sarah Chieng0.98
- Sarah Chieng0.98
- neo40.90
- @SarahChieng1.00
- Head of DevX @ Cerebras0.98
- prev.@ExaAiLabs, @shopthrifthouse, @MIT0.98
- @Cerebras1.00
- sarahchieng.com Joined March 20220.97
- 1,251 Following 17.3K Followers0.98
- Google DeepMind0.98
- AlEngineer0.99
- EUROPE1.00
-
- Speed of Top Coding Models over Time1.00
- togethe"0.91
- AlEn0.87
- in Output Speeds0.99
- 2001.00
- Grok Code0.99
- Fast 10.93
- AlEngine0.99
- 1501.00
- EUROPE1.00
- O(utst eed s)0.64
- Gemini 2.50.96
- Pro1.00
- GPT 4.10.99
- 1001.00
- Claude 4.5 Sonnet0.99
- Λari0.82
- Claude 3.50.98
- (AA eval)0.98
- 501.00
- Sonnet1.00
- Claude 3.51.00
- Sonnet1.00
- Claude1.00
- 4.5 Sonnet0.99
- DeepSeek1.00
- (reasoning)1.00
- V3.21.00
- Claude Sonnet1.00
- 4.61.00
- Claude1.00
- (reasoning)1.00
- Sonnet 4.60.96
- 20241.00
- June1.00
- 20240.77
- 20241.00
- Oct1.00
- 20241.00
- Dec1.00
- 20251.00
- Feb1.00
- 202150.55
- 20251.00
- June1.00
- 20250.72
- 20251.00
- Oct1.00
- 20251.00
- Dec1.00
- 20261.00
- Feb1.00
- 2020.54
- AlEngine0.97
- Date of Release1.00
- EUROPE1.00
- SOURCE: ARTIFICIAL ANALYSIS, OPENROUTER0.99
- Google DeepMind1.00
- AlEngineer0.99
- EUROPE1.00
-
- Speed of Top Coding Models over Time1.00
- in Output Speeds0.99
- 14001.00
- 12001.00
- Codex1.00
- Spark1.00
- 10001.00
- Oouet d d s)0.65
- 8001.00
- 6001.00
- 4001.00
- DeepSeek1.00
- V3.2 (reasoning)0.97
- 2001.00
- Gemini 2.51.00
- Grok Code0.99
- Claude 4.5 Sonnet (AA eval)0.99
- Pro1.00
- Fast 10.97
- Claude 3.50.99
- Claude 3.51.00
- GPT 4.10.95
- Claude 4.51.00
- Claude1.00
- Claude Sonnet1.00
- Sonnet1.00
- Sonnet1.00
- Sonnet1.00
- 4.6 (reasoning)1.00
- June1.00
- Aug1.00
- Oct1.00
- Dec1.00
- Feb1.00
- April1.00
- June1.00
- Aug1.00
- Oct1.00
- Dec1.00
- Feb1.00
- April0.99
- 20241.00
- 20241.00
- 20241.00
- 20241.00
- 20251.00
- 20251.00
- 20251.00
- 20251.00
- 20251.00
- 20251.00
- 20261.00
- 20261.00
- Date of Release0.99
- SOURCE: ARTIFICIAL ANALYSIS, OPENROUTER0.98
-
- Al models are getting faster0.98
- because the entire stack is0.98
- being optimized at once0.99
- Hardware1.00
- The physical reason speed is possible0.99
-
- AIEngin0.93
- EUROP1.00
- Al models are getting faster1.00
- because the entire stack is0.98
- bog0.77
- being optimized at once1.00
- Hardware1.00
- The physical reason speed is possible1.00
- AlEngineer0.96
- # Braintrust0.94
- WorkOS1.00
- OpenAI0.94
- EUROPE1.00
-
- Engineer1.00
- EUROPE1.00
- Al models are getting faster1.00
- @0.76
- H1001.00
- NVIDIA0.93
- because the entire stack is0.99
- DeenMin0.93
- being optimized at once1.00
- Memory is located0.99
- OFFCHIP1.00
- Hardware1.00
- ta1.00
- Engineer1.00
- The physical reason speed is possible1.00
- EUROPE1.00
- The Memory-Wall1.00
- CORD1.00
- On-chip memory: Keep data on-chip to0.98
- minimize memory movement and maximize1.00
- effective bandwidth per token.1.00
- cerebras1.00
- AMD0.99
- aws1.00
- NVIDIA0.97
- AlEngineer0.98
- #Braintrust0.98
- WorkOS1.00
- OpenAl0.94
- EUROPE1.00
-
- Ope0.94
- Al models are getting faster0.99
- 900,000 Cores on WSE-31.00
- because the entire stack is0.98
- Fabric Router1.00
- being optimized at once1.00
- Hardware1.00
- The physical reason speed is possible0.99
- Each Core has1.00
- DIRECT ACCESS1.00
- TO MEMORY1.00
- Al0.76
- The Memory-Wall0.97
- On-chip memory: Keep data on-chip to0.98
- minimize memory movement and maximize1.00
- effective bandwidth per token.1.00
- Be0.71
- cerebras1.00
- AMD1.00
- aws1.00
- Al0.82
- NVIDIA0.99
- Google DeepMind1.00
- AlEngineer0.99
- EUROPE1.00
-
- nAI0.84
- WEngineer0.89
- Al models are getting faster0.99
- EUROPE1.00
- because the entire stack is0.99
- traditional inference1.00
- being optimized at once1.00
- ONE PIECE OF HARDWARE1.00
- ner1.00
- SENTPY0.97
- user inpurt0.96
- prefill0.99
- decode1.00
- output0.92
- Hardware1.00
- COMPUTE BOUND1.00
- MEMORY BOUND0.99
- PE0.86
- The physical reason speed is possible0.99
- Engineer0.97
- Disaggregated Inference1.00
- EUROPE1.00
- Separate compute & memory across1.00
- specialized systems so each stage of1.00
- inference runs on hardware optimized for1.00
- its bottleneck.1.00
- NCORD1.00
- cerebras1.00
- AMD1.00
- aws1.00
- NVIDIA0.99
- Google DeepMind1.00
- AlEngineer0.98
- EUROPE1.00
-
- ineer1.00
- Al models are getting faster1.00
- IPE0.83
- because the entire stack is0.98
- traditional inference1.00
- ft1.00
- being optimized at once1.00
- ONE PIECE OF HARDWARE0.99
- user inpurt0.98
- prefill0.91
- decode1.00
- output1.00
- Hardware1.00
- COMPUTE BOUND0.99
- MEMORY BOUND0.99
- The physical reason speed is possible0.99
- er1.00
- Disaggregated Inference1.00
- ale1.00
- Separate compute & memory across0.99
- specialized systems so each stage of1.00
- inference runs on hardware optimized for1.00
- its bottleneck.1.00
- cerebras1.00
- AMD1.00
- aws1.00
- NVIDIA0.97
- AlEngineer0.98
- Braintrust1.00
- WorkOS1.00
- OpenAl0.94
- EUROPE1.00
-
- Al models are getting faster0.98
- because the entire stack is0.99
- traditional inference1.00
- being optimized at once1.00
- ONE PIECE OF HARDWARE0.99
- user input0.96
- prefill1.00
- decode1.00
- output1.00
- Hardware1.00
- COMPUTE BOUND0.97
- MEMORY BOUND1.00
- The physical reason speed is possible1.00
- Disaggregated Inference1.00
- Separate compute & memory across1.00
- specialized systems so each stage of1.00
- inference runs on hardware optimized for1.00
- its bottleneck.0.99
- cerebras0.98
- AMD1.00
- aws1.00
- NVIDIA.0.95
-
- MoE Model1.00
- Al models are getting faster1.00
- because the entire stack is0.99
- Input Data0.97
- being optimized at once1.00
- Gating Network1.00
- Hardware1.00
- The physical reason speed is possible1.00
- Expert 21.00
- Expert 31.00
- Model Architecture1.00
- How we design models to take advantage of0.99
- REAP1.00
- that hardware1.00
- (Router-weighted Expert Activation Pruning)1.00
- Prune low-importance experts using router1.00
- signals to compress MoE models while1.00
- preserving generative performance and0.99
- reducing memory overhead.0.97
- AI_0.99
- MISTRAL1.00
- ANTHROP\C1.00
- DeepMind1.00
- Carnegie1.00
- Mellon1.00
- University1.00
-
- Al models are getting faster1.00
- because the entire stack is0.98
- being optimized at once1.00
- Hardware1.00
- The physical reason speed is possible0.99
- Model Architecture1.00
- How we design models to take advantage of1.00
- thathardware1.00
- Inference Optimizations1.00
- How we squeeze even more performance at1.00
- runtime1.00
-
- Al models are getting faster0.99
- because the entire stack is0.99
- Yes1.00
- Retrieve from1.00
- cache1.00
- being optimized at once1.00
- Sequence1.00
- Input1.00
- Tokenization1.00
- Available?1.00
- Cache1.00
- Generate Token1.00
- No1.00
- Hardware1.00
- Compute KV1.00
- Pairs1.00
- Store in Cache1.00
- The physical reason speed is possible1.00
- Model Architecture1.00
- How we design models to take advantage of0.99
- KV Cache Reuse0.96
- thathardware1.00
- Store and reuse previously computed token0.99
- Inference Optimizations1.00
- representations to avoid recomputing1.00
- How we squeeze even more performance at1.00
- attention over the entire sequence0.98
- runtime1.00
- together.ai1.00
- baseten0.95
- Modal Fireworks Al0.97
-
- Allie K. Miller· 2nd0.97
- + Follow1.00
- Vincent Van Code0.99
- ø ...0.67
- #1 Most Followed Voice in Al Business (2M) | Former Amaz...0.98
- @vincent_vancode1.00
- View my newsletter1.00
- It's 9pm Saturday night.0.99
- 4mo·0.97
- Have you witnessed the level of Al multitasking insanity happening right now?0.99
- 2 projects, 8 agents, 5 screens, chugging through features, completing1.00
- MVP on both projects in 30 hours.0.99
- This is 6 Claude Code terminals running in parallel on my laptop.0.99
- If you keep calling this vibe coding and "it's not programming", then you0.98
- Multi-agent chaos is my new normal0.98
- are lost in technology.0.99
- Mon Dct 27 145AM0.94
- I am 49, I know 8 programming languages, been coding for 35 years, and0.99
- > create à now filo on my desktop of a legal tenplate0.90
- Clauda Code v2.0.270.90
- my final pivot is: Agentic programming.0.99
- • I'd he hapuy to help you create a legal template file on your0.88
- fusers/allisenaillar0.72
- create nen filos on my dosktop and make it synthetic sales data for a b2b0.93
- software company that is focused an cybersecur0.92
- I used to use GPT1 when we downloaded it ourselves. To see in the last1.00
- File fornat Sutmit0.93
- etrl-g to edit pronot in ví0.92
- 12 years how fast things move it gives me goose bumps.0.99
- Reuven Cohen · 2nd0.93
- + Follow0.99
- ∞ Agentic Engineer / CAiO @ Cognitum One0.97
- Book an appointment1.00
- Drop me all ine if your "vibbin" too0.98
- 1yr·Edited·0.97
- Roo Code now runs in multiple windows concurrently! Here's my current multi-0.99
- monitor setup 21k res, 5 concurrent VS codespaces, 500+ agent coding swarm,1.00
- interchange via MCPs deployed serverless and realtime supabase channel1.00
- comand and control. TS/Deno.0.98
- The monitors are organized based on my desk. 55inch Samsung 4k right display.0.99
- 1:55 AM·Feb 7,2026· 10.9K Views0.97
Transcript
166 cues· 3,248 words· 17,584 chars
- 0:16 Hi, everyone.
- 0:17 So we'll just get right into it.
- 0:20 So over the past few years, we have developers have developed a series of bad habits when it comes to developing as a result of slow AI code generation.
- 0:30 And so we're all familiar with it.
- 0:32 We do things like write massive prompts and try to one-shot.
- 0:36 We'll make huge commits.
- 0:38 Or we'll have our 10 agents all on the screen at the same time, combobulating, cogitating, thinking.
- 0:46 And so about a month ago, we at Cerebrus and OpenAI released a new model, state-of-the-art model, called Codex Spark.
- 0:53 Codex Spark can generate code at 1200 tokens per second.
- 0:58 And to put that into perspective, if you look at the Sonnet family or the Opus family, those can generate code at about 40 to 60 tokens per second.
- 1:08 So in this new era, as we're starting to see much faster coding models, this is 20 times faster, not only does it unlock new capabilities and use cases, but it also requires us to rethink how we as developers interact with the coding model.
- 1:23 And a lot of these bad habits that we had before that were generating maybe 50 tokens per second of bad code, unless we fix them, they're going to start generating 1,200 tokens per second of bad code.
- 1:36 And so that is the topic of today's talk.
- 1:41 So to get started, my name is Sarah Cheng.
- 1:43 I'm the head of developer experience at Cerebrus, where we are building the world's largest and fastest AI processor.
- 1:50 A large part of my job is that I get to introduce fast inference and fast coding models to developers for the very first time.
- 1:58 And for most people, it's a very exciting moment.
- 2:00 There's no thinking and waiting and starting up that you might be really annoyed about.
- 2:05 But at the same time, as I said, unless we change our habits,
- 2:10 we are not going to have good code in the future.
- 2:14 And so this talk really is a practical playbook for how we as developers can think about how we interact with the models in this new regime, especially in a future where the models are generating code faster than we the human can keep up.
- 2:29 So I wanna look back at history a little bit.
- 2:31 We've had a very exciting past few years.
- 2:33 The models have gotten bigger, they're getting smarter.
- 2:36 We have bigger context windows.
- 2:38 But the thing that has remained relatively constant over the past few years is coding speeds, is model speed.
- 2:44 So if we look at a lot of the popular families, we have Gemini, Claude, GPT, Sonnet.
- 2:50 Over the past few years, they've always been within 50 to 150 tokens per second.
- 2:57 And this is Codec Spark.
- 2:59 Again, Codec Spark is just the first of many models that we as developers can expect to be much faster than what we were previously used to.
- 3:07 And we even had to change the y-axis because it's so much faster.
- 3:10 And so before we get into the actual playbook and tips, I want to talk about why this is happening.
- 3:15 Why are we suddenly seeing such faster models?
- 3:18 And it's actually a very exciting development.
- 3:20 It's what many of you probably work on on a day-to-day, but there's so many companies that are working on this problem all at the same time.
- 3:28 And as a result, the entire AI inference stack is getting optimized all at once.
- 3:33 And so breaking it down, let's go through it really quickly.
- 3:35 We have hardware.
- 3:36 This is a physical device that inference, training, all of our compute is happening on.
- 3:42 One of the biggest things that we have to think about with hardware is the memory wall.
- 3:46 And this is exactly why hardware and memory movement takes up 50% to 80% of that latency time for inference.
- 3:52 This is where a lot of the frustration comes from.
- 3:55 And so when we are running inference, we have to constantly move our weights and KV cache values between memory and our actual chip.
- 4:03 On the NVIDIA GPU, this is the most traditional type of hardware, all of this memory is stored off-chip on off-chip HBM.
- 4:10 And we now have a memory bandwidth bottleneck.
- 4:13 What a lot of newer companies are doing are thinking about companies like Cerebrus or Grok, they're thinking about how do we move this memory to be as close to the chip as possible?
- 4:21 And so here's an example of the Cerebrus wafer where all of the chip is, all the memory is distributed across the chip in SRAM, so every core has direct access to the values it needs.
- 4:32 Even more exciting, we have disaggregated inference.
- 4:35 And disaggregated inference really has become commercialized in the last few months.
- 4:40 This is why NVIDIA bought Grok for $20 billion a few months ago.
loading
Chapters
- 0:00 Introduction to the impact of fast AI code generation
- 2:29 Historical context of model speeds
- 3:10 Why AI inference speeds are increasing (Hardware/Stack optimization)
- 7:05 The current developer landscape and risks of "slob"
- 8:27 Playbook: Orchestrating models and sub-agents
- 9:56 Playbook: Validation and automated testing
- 10:47 Playbook: Cherrypicking and variety in output
- 12:07 Playbook: Adopting a real-time collaborative mental model
- 12:53 Playbook: Avoiding "slob" and active steering
- 13:54 Playbook: Continuous refactoring
- 14:30 Playbook: Context management and external memory systems