read-only demo

Videos TeGsFFNqRLA

Fast Models Need Slow Developers — Sarah Chieng, Cerebras

index_state ready data_status ok

AI Engineer· published 2026-05-22· 0:18:01· en-US· indexed 2026-08-10 23:19

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:16, 1 of 1 keyframes kept
  4. Shot 3, 0:16 to 0:18, 1 of 1 keyframes kept
  5. Shot 4, 0:18 to 0:50, 1 of 1 keyframes kept
  6. Shot 5, 0:50 to 1:22, 1 of 1 keyframes kept
  7. Shot 6, 1:22 to 1:34, 1 of 1 keyframes kept
  8. Shot 7, 1:34 to 1:40, 1 of 1 keyframes kept
  9. Shot 8, 1:40 to 2:23, 1 of 1 keyframes kept
  10. Shot 9, 2:23 to 2:28, 1 of 1 keyframes kept
  11. Shot 10, 2:28 to 2:57, 1 of 1 keyframes kept
  12. Shot 11, 2:57 to 3:18, 1 of 1 keyframes kept
  13. Shot 12, 3:18 to 3:41, 1 of 1 keyframes kept
  14. Shot 13, 3:41 to 3:42, 1 of 1 keyframes kept
  15. Shot 14, 3:42 to 3:52, 1 of 1 keyframes kept
  16. Shot 15, 3:52 to 4:32, 1 of 1 keyframes kept
  17. Shot 16, 4:32 to 4:52, 1 of 1 keyframes kept
  18. Shot 17, 4:52 to 5:06, 1 of 1 keyframes kept
  19. Shot 18, 5:06 to 5:38, 1 of 1 keyframes kept
  20. Shot 19, 5:38 to 5:52, 0 of 1 keyframes kept
  21. Shot 20, 5:52 to 6:34, 1 of 1 keyframes kept
  22. Shot 21, 6:34 to 6:40, 1 of 1 keyframes kept
  23. Shot 22, 6:40 to 7:03, 1 of 1 keyframes kept
  24. Shot 23, 7:03 to 7:37, 1 of 1 keyframes kept
  25. Shot 24, 7:37 to 8:12, 0 of 1 keyframes kept
  26. Shot 25, 8:12 to 8:25, 1 of 1 keyframes kept
  27. Shot 26, 8:25 to 8:32, 1 of 1 keyframes kept
  28. Shot 27, 8:32 to 8:59, 1 of 1 keyframes kept
  29. Shot 28, 8:59 to 9:12, 1 of 1 keyframes kept
  30. Shot 29, 9:12 to 9:27, 1 of 1 keyframes kept
  31. Shot 30, 9:27 to 9:33, 1 of 1 keyframes kept
  32. Shot 31, 9:33 to 9:56, 1 of 1 keyframes kept
  33. Shot 32, 9:56 to 10:14, 1 of 1 keyframes kept
  34. Shot 33, 10:14 to 10:45, 1 of 1 keyframes kept
  35. Shot 34, 10:45 to 11:11, 1 of 1 keyframes kept
  36. Shot 35, 11:11 to 11:39, 1 of 1 keyframes kept
  37. Shot 36, 11:39 to 12:06, 0 of 1 keyframes kept
  38. Shot 37, 12:06 to 12:52, 1 of 1 keyframes kept
  39. Shot 38, 12:52 to 12:59, 1 of 1 keyframes kept
  40. Shot 39, 12:59 to 13:18, 1 of 1 keyframes kept
  41. Shot 40, 13:18 to 13:34, 1 of 1 keyframes kept
  42. Shot 41, 13:34 to 13:51, 1 of 1 keyframes kept
  43. Shot 42, 13:51 to 13:56, 1 of 1 keyframes kept
  44. Shot 43, 13:56 to 14:29, 1 of 1 keyframes kept
  45. Shot 44, 14:29 to 15:09, 1 of 1 keyframes kept
  46. Shot 45, 15:09 to 15:39, 1 of 1 keyframes kept
  47. Shot 46, 15:39 to 16:23, 1 of 1 keyframes kept
  48. Shot 47, 16:23 to 16:38, 1 of 1 keyframes kept
  49. Shot 48, 16:38 to 16:46, 1 of 1 keyframes kept
  50. Shot 49, 16:46 to 16:53, 1 of 1 keyframes kept
  51. Shot 50, 16:53 to 16:58, 1 of 1 keyframes kept
  52. Shot 51, 16:58 to 17:18, 1 of 1 keyframes kept
  53. Shot 52, 17:18 to 17:29, 1 of 1 keyframes kept
  54. Shot 53, 17:29 to 17:44, 1 of 1 keyframes kept
  55. Shot 54, 17:44 to 17:47, 1 of 1 keyframes kept
  56. Shot 55, 17:47 to 18:00, 1 of 1 keyframes kept
  57. Shot 56, 18:00 to 18:01, 0 of 1 keyframes kept

57 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
166
whisperx 166
chunks
30
from 166 cues
keyframes
53
kept of 57 captured
frames with text
53
1,249 lines read
chapters
11
from the source metadata
keyframe bytes
8.1 MB
word timings on 166 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 23:15 1m 18s
stt done 2026-08-10 23:16 26s
chunk done 2026-08-10 23:17 0s
text_embed done 2026-08-10 23:17 0s
keyframe done 2026-08-10 23:17 1m 58s
ocr done 2026-08-10 23:19 26s
frame_embed done 2026-08-10 23:19 10s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 827.9

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:12 #2 done3 line(s)

    shot 2·sharpness 907.7

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:18 #3 done20 line(s)

    shot 3·sharpness 747.7

    1. WorkOs0.91
    2. IIElevenLabs0.96
    3. AlEngineer0.94
    4. EUROPE1.00
    5. A0.89
    6. AlEngineer0.98
    7. DeepMind0.94
    8. AlEngineer0.99
    9. EUROPE0.94
    10. EUROPE1.00
    11. IEngineer0.99
    12. :neo4j0.93
    13. Λarize0.91
    14. EUROPE1.00
    15. RESENTED BY1.00
    16. gle DeepMind1.00
    17. AlEngineer0.99
    18. AlEngineer0.98
    19. EUROPE1.00
    20. EUROPE0.97
  • 0:34 #4 done16 line(s)

    shot 4·sharpness 1832.0

    1. cerebras1.00
    2. I Sarah Chieng0.96
    3. neer1.00
    4. Braintru0.99
    5. Fast Models1.00
    6. Engineer1.00
    7. Need1.00
    8. Micros1.00
    9. Slow1.00
    10. Mc0.70
    11. AlEngineer0.98
    12. EUROPE1.00
    13. Developers1.00
    14. Google DeepMind1.00
    15. AlEngineer0.99
    16. EUROPE1.00
  • 0:57 #5 done23 line(s)

    shot 5·sharpness 1797.8

    1. cerebras0.99
    2. Sarah Chieng0.97
    3. Engineer1.00
    4. Brai1.00
    5. EUROPE1.00
    6. Fast Models1.00
    7. pe1.00
    8. AIEng0.93
    9. EURC0.91
    10. Need1.00
    11. Engir1.00
    12. Mic0.99
    13. EUROPE1.00
    14. Slow1.00
    15. M1.00
    16. AIEng0.90
    17. EURC0.95
    18. Developers1.00
    19. AlEngineer0.99
    20. # Braintrust0.98
    21. WorkOS1.00
    22. OpenAl0.93
    23. EUROPE1.00
  • 1:23 #6 done16 line(s)

    shot 6·sharpness 1956.8

    1. AlEngine0.99
    2. Braintrust1.00
    3. cerebras1.00
    4. Sarah Chieng0.94
    5. Tessl1.00
    6. AlEngineer0.97
    7. Fast Models1.00
    8. Snorkel0.99
    9. Need1.00
    10. AlEngineer1.00
    11. Sonar0.99
    12. Slow1.00
    13. Developers1.00
    14. Google DeepMind1.00
    15. AlEngineer0.99
    16. EUROPE1.00
  • 1:37 #7 done17 line(s)

    shot 7·sharpness 1817.6

    1. cerebras0.99
    2. Sarah Chieng0.96
    3. AlEngineer0.98
    4. Br1.00
    5. EUROPE1.00
    6. Fast Models1.00
    7. Opel0.85
    8. AIE1.00
    9. Need1.00
    10. AIEng0.96
    11. EUPE0.96
    12. Slow1.00
    13. AIE0.98
    14. Developers1.00
    15. Google DeepMind1.00
    16. AlEngineer1.00
    17. EUROPE1.00
  • 1:45 #8 done46 line(s)

    shot 8·sharpness 1974.3

    1. Sarah1.00
    2. gineer1.00
    3. Braint1.00
    4. @milksandmatcha - 111K subscribers - 296 videos0.91
    5. JROPE0.97
    6. instagram.com/milksandmatcha and 1 more link0.96
    7. [email protected]0.88
    8. Customize channel0.98
    9. Manage videosy Commrunity0.92
    10. Home1.00
    11. Videos1.00
    12. Shorts Playlists Posts0.99
    13. enAl0.94
    14. Engine1.00
    15. Papular0.94
    16. Oldest0.99
    17. EUROPE1.00
    18. burned out1.00
    19. engineer1.00
    20. @startup0.95
    21. gineer1.00
    22. Micro0.92
    23. MIT ech, hawail0.87
    24. WEEK IN OUR IFE cupless vlog.0.86
    25. nemployed at 24 | life 2 years after0.97
    26. my 9-5 as a startup engineer | ife 20.97
    27. years after MIT0.96
    28. JROPE0.97
    29. Modo0.93
    30. IEngine0.94
    31. cerebr0.95
    32. EUROPE1.00
    33. Follow1.00
    34. Sarah Chieng0.99
    35. Sarah Chieng0.99
    36. @SarahChieng1.00
    37. @Cerebras1.00
    38. Head of DevX @ Cerebras1.00
    39. prev.@ExaAiLabs, @shopthrifthouse, @MIT0.98
    40. sarahchieng.com Joined March 20220.94
    41. 1,251 Following 17.3K Followers0.97
    42. AlEngineer0.97
    43. Braintrust1.00
    44. WorkOS1.00
    45. OpenAI0.92
    46. EUROPE1.00
  • 2:24 #9 done39 line(s)

    shot 9·sharpness 1984.1

    1. Sarah1.00
    2. gef0.90
    3. AlEngineer0.99
    4. @milksandmatcha - 111K subscribers · 296 videos0.92
    5. [email protected]0.85
    6. EUROPE0.95
    7. instagram.com/milksandmatcha and 1 more link0.98
    8. Customize channel0.94
    9. Manage videos Community0.96
    10. Home1.00
    11. Videos1.00
    12. Shorts Playlists Posts0.98
    13. Papular0.95
    14. Oldest0.99
    15. Zed0.88
    16. burned out0.99
    17. engineer1.00
    18. @startup0.94
    19. WEXK IN OUR IFE couples vlog0.91
    20. MIT, tech, hawall0.94
    21. nemmployed at 24 | lfe 2 years after0.93
    22. my 9-5 as a startup engineer | ife 20.96
    23. years after MIT1.00
    24. AlEngineer0.94
    25. cerebr0.99
    26. EUROPE1.00
    27. Follow1.00
    28. Sarah Chieng0.98
    29. Sarah Chieng0.98
    30. neo40.90
    31. @SarahChieng1.00
    32. Head of DevX @ Cerebras0.98
    33. prev.@ExaAiLabs, @shopthrifthouse, @MIT0.98
    34. @Cerebras1.00
    35. sarahchieng.com Joined March 20220.97
    36. 1,251 Following 17.3K Followers0.98
    37. Google DeepMind0.98
    38. AlEngineer0.99
    39. EUROPE1.00
  • 2:32 #10 done60 line(s)

    shot 10·sharpness 1341.2

    1. Speed of Top Coding Models over Time1.00
    2. togethe"0.91
    3. AlEn0.87
    4. in Output Speeds0.99
    5. 2001.00
    6. Grok Code0.99
    7. Fast 10.93
    8. AlEngine0.99
    9. 1501.00
    10. EUROPE1.00
    11. O(utst eed s)0.64
    12. Gemini 2.50.96
    13. Pro1.00
    14. GPT 4.10.99
    15. 1001.00
    16. Claude 4.5 Sonnet0.99
    17. Λari0.82
    18. Claude 3.50.98
    19. (AA eval)0.98
    20. 501.00
    21. Sonnet1.00
    22. Claude 3.51.00
    23. Sonnet1.00
    24. Claude1.00
    25. 4.5 Sonnet0.99
    26. DeepSeek1.00
    27. (reasoning)1.00
    28. V3.21.00
    29. Claude Sonnet1.00
    30. 4.61.00
    31. Claude1.00
    32. (reasoning)1.00
    33. Sonnet 4.60.96
    34. 20241.00
    35. June1.00
    36. 20240.77
    37. 20241.00
    38. Oct1.00
    39. 20241.00
    40. Dec1.00
    41. 20251.00
    42. Feb1.00
    43. 202150.55
    44. 20251.00
    45. June1.00
    46. 20250.72
    47. 20251.00
    48. Oct1.00
    49. 20251.00
    50. Dec1.00
    51. 20261.00
    52. Feb1.00
    53. 2020.54
    54. AlEngine0.97
    55. Date of Release1.00
    56. EUROPE1.00
    57. SOURCE: ARTIFICIAL ANALYSIS, OPENROUTER0.99
    58. Google DeepMind1.00
    59. AlEngineer0.99
    60. EUROPE1.00
  • 2:59 #11 done55 line(s)

    shot 11·sharpness 1245.2

    1. Speed of Top Coding Models over Time1.00
    2. in Output Speeds0.99
    3. 14001.00
    4. 12001.00
    5. Codex1.00
    6. Spark1.00
    7. 10001.00
    8. Oouet d d s)0.65
    9. 8001.00
    10. 6001.00
    11. 4001.00
    12. DeepSeek1.00
    13. V3.2 (reasoning)0.97
    14. 2001.00
    15. Gemini 2.51.00
    16. Grok Code0.99
    17. Claude 4.5 Sonnet (AA eval)0.99
    18. Pro1.00
    19. Fast 10.97
    20. Claude 3.50.99
    21. Claude 3.51.00
    22. GPT 4.10.95
    23. Claude 4.51.00
    24. Claude1.00
    25. Claude Sonnet1.00
    26. Sonnet1.00
    27. Sonnet1.00
    28. Sonnet1.00
    29. 4.6 (reasoning)1.00
    30. June1.00
    31. Aug1.00
    32. Oct1.00
    33. Dec1.00
    34. Feb1.00
    35. April1.00
    36. June1.00
    37. Aug1.00
    38. Oct1.00
    39. Dec1.00
    40. Feb1.00
    41. April0.99
    42. 20241.00
    43. 20241.00
    44. 20241.00
    45. 20241.00
    46. 20251.00
    47. 20251.00
    48. 20251.00
    49. 20251.00
    50. 20251.00
    51. 20251.00
    52. 20261.00
    53. 20261.00
    54. Date of Release0.99
    55. SOURCE: ARTIFICIAL ANALYSIS, OPENROUTER0.98
  • 3:36 #12 done5 line(s)

    shot 12·sharpness 2329.9

    1. Al models are getting faster0.98
    2. because the entire stack is0.98
    3. being optimized at once0.99
    4. Hardware1.00
    5. The physical reason speed is possible0.99
  • 3:42 #13 done13 line(s)

    shot 13·sharpness 2361.1

    1. AIEngin0.93
    2. EUROP1.00
    3. Al models are getting faster1.00
    4. because the entire stack is0.98
    5. bog0.77
    6. being optimized at once1.00
    7. Hardware1.00
    8. The physical reason speed is possible1.00
    9. AlEngineer0.96
    10. # Braintrust0.94
    11. WorkOS1.00
    12. OpenAI0.94
    13. EUROPE1.00
  • 3:47 #14 done30 line(s)

    shot 14·sharpness 2370.4

    1. Engineer1.00
    2. EUROPE1.00
    3. Al models are getting faster1.00
    4. @0.76
    5. H1001.00
    6. NVIDIA0.93
    7. because the entire stack is0.99
    8. DeenMin0.93
    9. being optimized at once1.00
    10. Memory is located0.99
    11. OFFCHIP1.00
    12. Hardware1.00
    13. ta1.00
    14. Engineer1.00
    15. The physical reason speed is possible1.00
    16. EUROPE1.00
    17. The Memory-Wall1.00
    18. CORD1.00
    19. On-chip memory: Keep data on-chip to0.98
    20. minimize memory movement and maximize1.00
    21. effective bandwidth per token.1.00
    22. cerebras1.00
    23. AMD0.99
    24. aws1.00
    25. NVIDIA0.97
    26. AlEngineer0.98
    27. #Braintrust0.98
    28. WorkOS1.00
    29. OpenAl0.94
    30. EUROPE1.00
  • 4:20 #15 done25 line(s)

    shot 15·sharpness 2529.9

    1. Ope0.94
    2. Al models are getting faster0.99
    3. 900,000 Cores on WSE-31.00
    4. because the entire stack is0.98
    5. Fabric Router1.00
    6. being optimized at once1.00
    7. Hardware1.00
    8. The physical reason speed is possible0.99
    9. Each Core has1.00
    10. DIRECT ACCESS1.00
    11. TO MEMORY1.00
    12. Al0.76
    13. The Memory-Wall0.97
    14. On-chip memory: Keep data on-chip to0.98
    15. minimize memory movement and maximize1.00
    16. effective bandwidth per token.1.00
    17. Be0.71
    18. cerebras1.00
    19. AMD1.00
    20. aws1.00
    21. Al0.82
    22. NVIDIA0.99
    23. Google DeepMind1.00
    24. AlEngineer0.99
    25. EUROPE1.00
  • 4:35 #16 done34 line(s)

    shot 16·sharpness 2442.9

    1. nAI0.84
    2. WEngineer0.89
    3. Al models are getting faster0.99
    4. EUROPE1.00
    5. because the entire stack is0.99
    6. traditional inference1.00
    7. being optimized at once1.00
    8. ONE PIECE OF HARDWARE1.00
    9. ner1.00
    10. SENTPY0.97
    11. user inpurt0.96
    12. prefill0.99
    13. decode1.00
    14. output0.92
    15. Hardware1.00
    16. COMPUTE BOUND1.00
    17. MEMORY BOUND0.99
    18. PE0.86
    19. The physical reason speed is possible0.99
    20. Engineer0.97
    21. Disaggregated Inference1.00
    22. EUROPE1.00
    23. Separate compute & memory across1.00
    24. specialized systems so each stage of1.00
    25. inference runs on hardware optimized for1.00
    26. its bottleneck.1.00
    27. NCORD1.00
    28. cerebras1.00
    29. AMD1.00
    30. aws1.00
    31. NVIDIA0.99
    32. Google DeepMind1.00
    33. AlEngineer0.98
    34. EUROPE1.00
  • 4:56 #17 done32 line(s)

    shot 17·sharpness 2323.0

    1. ineer1.00
    2. Al models are getting faster1.00
    3. IPE0.83
    4. because the entire stack is0.98
    5. traditional inference1.00
    6. ft1.00
    7. being optimized at once1.00
    8. ONE PIECE OF HARDWARE0.99
    9. user inpurt0.98
    10. prefill0.91
    11. decode1.00
    12. output1.00
    13. Hardware1.00
    14. COMPUTE BOUND0.99
    15. MEMORY BOUND0.99
    16. The physical reason speed is possible0.99
    17. er1.00
    18. Disaggregated Inference1.00
    19. ale1.00
    20. Separate compute & memory across0.99
    21. specialized systems so each stage of1.00
    22. inference runs on hardware optimized for1.00
    23. its bottleneck.1.00
    24. cerebras1.00
    25. AMD1.00
    26. aws1.00
    27. NVIDIA0.97
    28. AlEngineer0.98
    29. Braintrust1.00
    30. WorkOS1.00
    31. OpenAl0.94
    32. EUROPE1.00
  • 5:25 #18 done22 line(s)

    shot 18·sharpness 2285.6

    1. Al models are getting faster0.98
    2. because the entire stack is0.99
    3. traditional inference1.00
    4. being optimized at once1.00
    5. ONE PIECE OF HARDWARE0.99
    6. user input0.96
    7. prefill1.00
    8. decode1.00
    9. output1.00
    10. Hardware1.00
    11. COMPUTE BOUND0.97
    12. MEMORY BOUND1.00
    13. The physical reason speed is possible1.00
    14. Disaggregated Inference1.00
    15. Separate compute & memory across1.00
    16. specialized systems so each stage of1.00
    17. inference runs on hardware optimized for1.00
    18. its bottleneck.0.99
    19. cerebras0.98
    20. AMD1.00
    21. aws1.00
    22. NVIDIA.0.95
  • 5:45 #19 skipped

    shot 19·duplicate of #12

  • 6:29 #20 done26 line(s)

    shot 20·sharpness 2360.0

    1. MoE Model1.00
    2. Al models are getting faster1.00
    3. because the entire stack is0.99
    4. Input Data0.97
    5. being optimized at once1.00
    6. Gating Network1.00
    7. Hardware1.00
    8. The physical reason speed is possible1.00
    9. Expert 21.00
    10. Expert 31.00
    11. Model Architecture1.00
    12. How we design models to take advantage of0.99
    13. REAP1.00
    14. that hardware1.00
    15. (Router-weighted Expert Activation Pruning)1.00
    16. Prune low-importance experts using router1.00
    17. signals to compress MoE models while1.00
    18. preserving generative performance and0.99
    19. reducing memory overhead.0.97
    20. AI_0.99
    21. MISTRAL1.00
    22. ANTHROP\C1.00
    23. DeepMind1.00
    24. Carnegie1.00
    25. Mellon1.00
    26. University1.00
  • 6:39 #21 done11 line(s)

    shot 21·sharpness 2793.1

    1. Al models are getting faster1.00
    2. because the entire stack is0.98
    3. being optimized at once1.00
    4. Hardware1.00
    5. The physical reason speed is possible0.99
    6. Model Architecture1.00
    7. How we design models to take advantage of1.00
    8. thathardware1.00
    9. Inference Optimizations1.00
    10. How we squeeze even more performance at1.00
    11. runtime1.00
  • 6:43 #22 done31 line(s)

    shot 22·sharpness 2538.3

    1. Al models are getting faster0.99
    2. because the entire stack is0.99
    3. Yes1.00
    4. Retrieve from1.00
    5. cache1.00
    6. being optimized at once1.00
    7. Sequence1.00
    8. Input1.00
    9. Tokenization1.00
    10. Available?1.00
    11. Cache1.00
    12. Generate Token1.00
    13. No1.00
    14. Hardware1.00
    15. Compute KV1.00
    16. Pairs1.00
    17. Store in Cache1.00
    18. The physical reason speed is possible1.00
    19. Model Architecture1.00
    20. How we design models to take advantage of0.99
    21. KV Cache Reuse0.96
    22. thathardware1.00
    23. Store and reuse previously computed token0.99
    24. Inference Optimizations1.00
    25. representations to avoid recomputing1.00
    26. How we squeeze even more performance at1.00
    27. attention over the entire sequence0.98
    28. runtime1.00
    29. together.ai1.00
    30. baseten0.95
    31. Modal Fireworks Al0.97
  • 7:33 #23 done41 line(s)

    shot 23·sharpness 3871.0

    1. Allie K. Miller· 2nd0.97
    2. + Follow1.00
    3. Vincent Van Code0.99
    4. ø ...0.67
    5. #1 Most Followed Voice in Al Business (2M) | Former Amaz...0.98
    6. @vincent_vancode1.00
    7. View my newsletter1.00
    8. It's 9pm Saturday night.0.99
    9. 4mo·0.97
    10. Have you witnessed the level of Al multitasking insanity happening right now?0.99
    11. 2 projects, 8 agents, 5 screens, chugging through features, completing1.00
    12. MVP on both projects in 30 hours.0.99
    13. This is 6 Claude Code terminals running in parallel on my laptop.0.99
    14. If you keep calling this vibe coding and "it's not programming", then you0.98
    15. Multi-agent chaos is my new normal0.98
    16. are lost in technology.0.99
    17. Mon Dct 27 145AM0.94
    18. I am 49, I know 8 programming languages, been coding for 35 years, and0.99
    19. > create à now filo on my desktop of a legal tenplate0.90
    20. Clauda Code v2.0.270.90
    21. my final pivot is: Agentic programming.0.99
    22. • I'd he hapuy to help you create a legal template file on your0.88
    23. fusers/allisenaillar0.72
    24. create nen filos on my dosktop and make it synthetic sales data for a b2b0.93
    25. software company that is focused an cybersecur0.92
    26. I used to use GPT1 when we downloaded it ourselves. To see in the last1.00
    27. File fornat Sutmit0.93
    28. etrl-g to edit pronot in ví0.92
    29. 12 years how fast things move it gives me goose bumps.0.99
    30. Reuven Cohen · 2nd0.93
    31. + Follow0.99
    32. ∞ Agentic Engineer / CAiO @ Cognitum One0.97
    33. Book an appointment1.00
    34. Drop me all ine if your "vibbin" too0.98
    35. 1yr·Edited·0.97
    36. Roo Code now runs in multiple windows concurrently! Here's my current multi-0.99
    37. monitor setup 21k res, 5 concurrent VS codespaces, 500+ agent coding swarm,1.00
    38. interchange via MCPs deployed serverless and realtime supabase channel1.00
    39. comand and control. TS/Deno.0.98
    40. The monitors are organized based on my desk. 55inch Samsung 4k right display.0.99
    41. 1:55 AM·Feb 7,2026· 10.9K Views0.97

Transcript

166 cues· 3,248 words· 17,584 chars

  1. 0:16 Hi, everyone.
  2. 0:17 So we'll just get right into it.
  3. 0:20 So over the past few years, we have developers have developed a series of bad habits when it comes to developing as a result of slow AI code generation.
  4. 0:30 And so we're all familiar with it.
  5. 0:32 We do things like write massive prompts and try to one-shot.
  6. 0:36 We'll make huge commits.
  7. 0:38 Or we'll have our 10 agents all on the screen at the same time, combobulating, cogitating, thinking.
  8. 0:46 And so about a month ago, we at Cerebrus and OpenAI released a new model, state-of-the-art model, called Codex Spark.
  9. 0:53 Codex Spark can generate code at 1200 tokens per second.
  10. 0:58 And to put that into perspective, if you look at the Sonnet family or the Opus family, those can generate code at about 40 to 60 tokens per second.
  11. 1:08 So in this new era, as we're starting to see much faster coding models, this is 20 times faster, not only does it unlock new capabilities and use cases, but it also requires us to rethink how we as developers interact with the coding model.
  12. 1:23 And a lot of these bad habits that we had before that were generating maybe 50 tokens per second of bad code, unless we fix them, they're going to start generating 1,200 tokens per second of bad code.
  13. 1:36 And so that is the topic of today's talk.
  14. 1:41 So to get started, my name is Sarah Cheng.
  15. 1:43 I'm the head of developer experience at Cerebrus, where we are building the world's largest and fastest AI processor.
  16. 1:50 A large part of my job is that I get to introduce fast inference and fast coding models to developers for the very first time.
  17. 1:58 And for most people, it's a very exciting moment.
  18. 2:00 There's no thinking and waiting and starting up that you might be really annoyed about.
  19. 2:05 But at the same time, as I said, unless we change our habits,
  20. 2:10 we are not going to have good code in the future.
  21. 2:14 And so this talk really is a practical playbook for how we as developers can think about how we interact with the models in this new regime, especially in a future where the models are generating code faster than we the human can keep up.
  22. 2:29 So I wanna look back at history a little bit.
  23. 2:31 We've had a very exciting past few years.
  24. 2:33 The models have gotten bigger, they're getting smarter.
  25. 2:36 We have bigger context windows.
  26. 2:38 But the thing that has remained relatively constant over the past few years is coding speeds, is model speed.
  27. 2:44 So if we look at a lot of the popular families, we have Gemini, Claude, GPT, Sonnet.
  28. 2:50 Over the past few years, they've always been within 50 to 150 tokens per second.
  29. 2:57 And this is Codec Spark.
  30. 2:59 Again, Codec Spark is just the first of many models that we as developers can expect to be much faster than what we were previously used to.
  31. 3:07 And we even had to change the y-axis because it's so much faster.
  32. 3:10 And so before we get into the actual playbook and tips, I want to talk about why this is happening.
  33. 3:15 Why are we suddenly seeing such faster models?
  34. 3:18 And it's actually a very exciting development.
  35. 3:20 It's what many of you probably work on on a day-to-day, but there's so many companies that are working on this problem all at the same time.
  36. 3:28 And as a result, the entire AI inference stack is getting optimized all at once.
  37. 3:33 And so breaking it down, let's go through it really quickly.
  38. 3:35 We have hardware.
  39. 3:36 This is a physical device that inference, training, all of our compute is happening on.
  40. 3:42 One of the biggest things that we have to think about with hardware is the memory wall.
  41. 3:46 And this is exactly why hardware and memory movement takes up 50% to 80% of that latency time for inference.
  42. 3:52 This is where a lot of the frustration comes from.
  43. 3:55 And so when we are running inference, we have to constantly move our weights and KV cache values between memory and our actual chip.
  44. 4:03 On the NVIDIA GPU, this is the most traditional type of hardware, all of this memory is stored off-chip on off-chip HBM.
  45. 4:10 And we now have a memory bandwidth bottleneck.
  46. 4:13 What a lot of newer companies are doing are thinking about companies like Cerebrus or Grok, they're thinking about how do we move this memory to be as close to the chip as possible?
  47. 4:21 And so here's an example of the Cerebrus wafer where all of the chip is, all the memory is distributed across the chip in SRAM, so every core has direct access to the values it needs.
  48. 4:32 Even more exciting, we have disaggregated inference.
  49. 4:35 And disaggregated inference really has become commercialized in the last few months.
  50. 4:40 This is why NVIDIA bought Grok for $20 billion a few months ago.

Chapters

  1. 0:00 Introduction to the impact of fast AI code generation
  2. 2:29 Historical context of model speeds
  3. 3:10 Why AI inference speeds are increasing (Hardware/Stack optimization)
  4. 7:05 The current developer landscape and risks of "slob"
  5. 8:27 Playbook: Orchestrating models and sub-agents
  6. 9:56 Playbook: Validation and automated testing
  7. 10:47 Playbook: Cherrypicking and variety in output
  8. 12:07 Playbook: Adopting a real-time collaborative mental model
  9. 12:53 Playbook: Avoiding "slob" and active steering
  10. 13:54 Playbook: Continuous refactoring
  11. 14:30 Playbook: Context management and external memory systems

Open at this second