read-only demo

Videos lyL5QhgIOxc

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

index_state ready data_status ok

AI Engineer· published 2026-07-28· 0:21:39· en-US· indexed 2026-08-10 19:39

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:15, 1 of 1 keyframes kept
  5. Shot 4, 0:15 to 0:27, 1 of 1 keyframes kept
  6. Shot 5, 0:27 to 1:01, 1 of 1 keyframes kept
  7. Shot 6, 1:01 to 1:34, 1 of 1 keyframes kept
  8. Shot 7, 1:34 to 2:03, 1 of 1 keyframes kept
  9. Shot 8, 2:03 to 2:32, 0 of 1 keyframes kept
  10. Shot 9, 2:32 to 3:00, 0 of 1 keyframes kept
  11. Shot 10, 3:00 to 3:28, 1 of 1 keyframes kept
  12. Shot 11, 3:28 to 4:05, 1 of 1 keyframes kept
  13. Shot 12, 4:05 to 4:35, 1 of 1 keyframes kept
  14. Shot 13, 4:35 to 5:04, 0 of 1 keyframes kept
  15. Shot 14, 5:04 to 5:33, 0 of 1 keyframes kept
  16. Shot 15, 5:33 to 5:38, 1 of 1 keyframes kept
  17. Shot 16, 5:38 to 6:05, 1 of 1 keyframes kept
  18. Shot 17, 6:05 to 6:32, 0 of 1 keyframes kept
  19. Shot 18, 6:32 to 7:00, 0 of 1 keyframes kept
  20. Shot 19, 7:00 to 7:27, 0 of 1 keyframes kept
  21. Shot 20, 7:27 to 7:54, 0 of 1 keyframes kept
  22. Shot 21, 7:54 to 8:24, 1 of 1 keyframes kept
  23. Shot 22, 8:24 to 8:55, 0 of 1 keyframes kept
  24. Shot 23, 8:55 to 9:25, 0 of 1 keyframes kept
  25. Shot 24, 9:25 to 10:07, 1 of 1 keyframes kept
  26. Shot 25, 10:07 to 10:41, 0 of 1 keyframes kept
  27. Shot 26, 10:41 to 11:16, 0 of 1 keyframes kept
  28. Shot 27, 11:16 to 11:46, 1 of 1 keyframes kept
  29. Shot 28, 11:46 to 12:12, 1 of 1 keyframes kept
  30. Shot 29, 12:12 to 12:38, 0 of 1 keyframes kept
  31. Shot 30, 12:38 to 12:54, 0 of 1 keyframes kept
  32. Shot 31, 12:54 to 13:22, 1 of 1 keyframes kept
  33. Shot 32, 13:22 to 13:50, 0 of 1 keyframes kept
  34. Shot 33, 13:50 to 14:17, 0 of 1 keyframes kept
  35. Shot 34, 14:17 to 14:45, 0 of 1 keyframes kept
  36. Shot 35, 14:45 to 15:13, 1 of 1 keyframes kept
  37. Shot 36, 15:13 to 15:41, 0 of 1 keyframes kept
  38. Shot 37, 15:41 to 16:09, 0 of 1 keyframes kept
  39. Shot 38, 16:09 to 16:36, 0 of 1 keyframes kept
  40. Shot 39, 16:36 to 17:09, 1 of 1 keyframes kept
  41. Shot 40, 17:09 to 17:41, 0 of 1 keyframes kept
  42. Shot 41, 17:41 to 18:13, 0 of 1 keyframes kept
  43. Shot 42, 18:13 to 18:38, 1 of 1 keyframes kept
  44. Shot 43, 18:38 to 19:04, 0 of 1 keyframes kept
  45. Shot 44, 19:04 to 19:53, 1 of 1 keyframes kept
  46. Shot 45, 19:53 to 20:42, 1 of 1 keyframes kept
  47. Shot 46, 20:42 to 21:10, 0 of 1 keyframes kept
  48. Shot 47, 21:10 to 21:15, 1 of 1 keyframes kept
  49. Shot 48, 21:15 to 21:22, 1 of 1 keyframes kept
  50. Shot 49, 21:22 to 21:38, 0 of 1 keyframes kept

50 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
191
whisperx 191
chunks
38
from 191 cues
keyframes
25
kept of 50 captured
frames with text
25
497 lines read
chapters
11
from the source metadata
keyframe bytes
9.0 MB
word timings on 191 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 22:22 0s
stt done 2026-08-09 06:11 17s
chunk done 2026-08-09 06:11 0s
text_embed done 2026-08-10 19:39 0s
keyframe done 2026-08-09 06:11 2m 44s
ocr done 2026-08-09 06:14 8s
frame_embed done 2026-08-10 19:39 5s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 456.5

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 665.3

    1. AIEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2736.3

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.92
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:12 #3 done5 line(s)

    shot 3·sharpness 319.4

    1. Hugging Faca0.94
    2. HUGGING FACE1.00
    3. Hugging Face0.98
    4. AlEngineer0.95
    5. World's Fair0.99
  • 0:16 #4 done3 line(s)

    shot 4·sharpness 155.2

    1. HuggingFue0.77
    2. HUGGING FACE1.00
    3. Fair1.00
  • 0:40 #5 done20 line(s)

    shot 5·sharpness 3587.3

    1. AlEngineer0.99
    2. Serving 3 Million Models0.98
    3. World's Fair0.99
    4. How Hugging Face Scaled the World's Largest Model Hub0.99
    5. TA1.00
    6. PRESENTED BY1.00
    7. 0.52
    8. Microsoft1.00
    9. )0.95
    10. 27K1.00
    11. 50K1.00
    12. 35K1.00
    13. 40K1.00
    14. 41K1.00
    15. 50K1.00
    16. Arek Borucki1.00
    17. ML Platform & Database Engineer @ Hugging Face0.99
    18. MongoDB Champion | Author of MongoDB 8.0 in Action0.99
    19. HUGGING FACE1.00
    20. Engineering the future of Al1.00
  • 1:14 #6 done20 line(s)

    shot 6·sharpness 3807.1

    1. AlEngineer0.99
    2. Serving 3 Million Models0.99
    3. World's Fair0.97
    4. How Hugging Face Scaled the World's Largest Model Hub0.99
    5. TA1.00
    6. PRESENTED BY0.97
    7. Microsoft1.00
    8. )0.91
    9. 27K1.00
    10. 50K1.00
    11. 35K1.00
    12. 40K1.00
    13. 41K1.00
    14. 50K1.00
    15. Arek Borucki1.00
    16. ML Platform & Database Engineer @ Hugging Face0.99
    17. MongoDB Champion | Author of MongoDB 8.0 in Action1.00
    18. HUGGING FACE1.00
    19. LEADERSHIP 2· JUNE 30, 20260.94
    20. Al Architects: Show my Workflow0.99
  • 1:46 #7 done22 line(s)

    shot 7·sharpness 3299.3

    1. AlEngineer0.99
    2. Hugging Face: The Scale0.99
    3. World's Fair0.97
    4. 14M+1.00
    5. 2.9M+1.00
    6. 1M+1.00
    7. 50K+1.00
    8. USERS1.00
    9. PUBLIC MODELS1.00
    10. DATASETS1.00
    11. ORGANIZATIONS1.00
    12. DATA1.00
    13. DATA1.00
    14. DATA1.00
    15. USERS.0.98
    16. PUBLIC MODELS1.00
    17. DATASETS1.00
    18. ORGANIZATIONS1.00
    19. 30%+ of the Fortune 500 use Hugging Face1.00
    20. HUGGING FACE0.96
    21. LEADERSHIP 2· JUNE 30, 20260.95
    22. Al Architects: Show my Workflow0.99
  • 2:23 #8 skipped

    shot 8·duplicate of #7

  • 2:57 #9 skipped

    shot 9·duplicate of #7

  • 3:24 #10 done36 line(s)

    shot 10·sharpness 2629.3

    1. Halfway through 20261.00
    2. AlEngineer0.99
    3. 3M1.00
    4. World's Fair0.96
    5. 3M1.00
    6. Models1.00
    7. Cumsg M i o ace0.61
    8. 2M1.00
    9. 2.5M1.00
    10. Models1.00
    11. GPT-OSS1.00
    12. 2M1.00
    13. DeepSeek-R11.00
    14. 1M1.00
    15. 1.5M1.00
    16. Models1.00
    17. Flux 1.01.00
    18. 1M1.00
    19. Qwen21.00
    20. 500K1.00
    21. Models1.00
    22. 500K1.00
    23. 100K1.00
    24. LLaMA 20.98
    25. Models0.99
    26. BLOOM1.00
    27. o0.63
    28. 20221.00
    29. 20231.00
    30. 20241.00
    31. 20251.00
    32. 20261.00
    33. (H1)1.00
    34. HUGWING FACE0.93
    35. LEADERSHIP 2• JUNE 30, 20260.97
    36. Al Architects: Show my Workflow0.97
  • 3:47 #11 done54 line(s)

    shot 11·sharpness 2135.3

    1. Hugging Face just crossed 1M datasets0.99
    2. AlEngineer1.00
    3. Cumulative datasets published, by task category1.00
    4. World's Fair0.99
    5. 1.0M1.00
    6. 800k1.00
    7. Sep 20251.00
    8. datasets1.00
    9. Other1.00
    10. Ccumst dtets0.72
    11. Tabular0.98
    12. 600k1.00
    13. Oct 20221.00
    14. 10K1.00
    15. Feb 20241.00
    16. datasets1.00
    17. 100K0.99
    18. Reinforcement1.00
    19. Learning1.00
    20. datasets1.00
    21. Audio1.00
    22. Multimodal1.00
    23. 400k1.00
    24. Computer1.00
    25. Vision1.00
    26. Natural1.00
    27. Language1.00
    28. Processing1.00
    29. 200k1.00
    30. Apr1.00
    31. Jul0.96
    32. Oct1.00
    33. Jan1.00
    34. Apr1.00
    35. Jul1.00
    36. Oct1.00
    37. Jan1.00
    38. Apr1.00
    39. Jul1.00
    40. Oct1.00
    41. Jan1.00
    42. Apr0.98
    43. Jul1.00
    44. Oct1.00
    45. Jan1.00
    46. Apr1.00
    47. 20221.00
    48. 20231.00
    49. 20241.00
    50. 20251.00
    51. 20261.00
    52. HUGGING FACE1.00
    53. LEADERSHIP 2· JUNE 30, 20260.97
    54. Al Architects: Show my Workflow0.97
  • 4:09 #12 done12 line(s)

    shot 12·sharpness 4133.4

    1. AlEngineer0.98
    2. When 20k become 3 Million. Latency Matters0.99
    3. World'sFair1.00
    4. Users expect instant results.0.99
    5. Slow search = they leave.0.98
    6. At 20K models any query was fast.1.00
    7. At 3M the same approach breaks.1.00
    8. 14+M users, p50 lies.1.00
    9. What hurts UX is p99.0.99
    10. HUGGING FNE0.92
    11. LEADERSHIP 2· JUNE 30, 20260.95
    12. Al Architects: Show my Workflow0.99
  • 4:46 #13 skipped

    shot 13·duplicate of #12

  • 5:29 #14 skipped

    shot 14·duplicate of #12

  • 5:34 #15 done7 line(s)

    shot 15·sharpness 1229.5

    1. AlEngineer0.99
    2. World'sFair1.00
    3. PRESENTED BY1.00
    4. Microsoft1.00
    5. HUGGING FACE0.99
    6. LEADERSHIP 2· JUNE 30, 20260.95
    7. Al Architects: Show my Workflow0.97
  • 5:56 #16 done21 line(s)

    shot 16·sharpness 3046.7

    1. AlEngineer0.98
    2. High-level Architecture1.00
    3. World's Fair0.97
    4. Hub1.00
    5. Hub1.00
    6. MongoDB Atlas0.97
    7. Users1.00
    8. Frontend1.00
    9. API1.00
    10. (Source of Truth for Metadata)1.00
    11. (Holds metadata, card text,1.00
    12. PRESENTED BY0.97
    13. configuration data, user info)0.99
    14. Microsoft1.00
    15. Cloud Storage0.99
    16. (S3 / GCS - Binary Data & Models)0.98
    17. Holds all model weights, tokenizer files,0.99
    18. card assets, and configuration files1.00
    19. AGING FACE0.93
    20. LEADERSHIP 2· JUNE 30, 20260.95
    21. Al Architects: Show my Workflow0.99
  • 6:19 #17 skipped

    shot 17·duplicate of #16

  • 6:54 #18 skipped

    shot 18·duplicate of #16

  • 7:03 #19 skipped

    shot 19·duplicate of #16

  • 7:41 #20 skipped

    shot 20·duplicate of #16

  • 8:12 #21 done32 line(s)

    shot 21·sharpness 4393.3

    1. How we search 3M Models under 15ms0.99
    2. AlEngineer0.99
    3. World'sFair1.00
    4. Users1.00
    5. “Llama”0.97
    6. Hub1.00
    7. Optimized1.00
    8. results1.00
    9. Ranked1.00
    10. Read Collection1.00
    11. Results1.00
    12. Write time:1.00
    13. meta-llama/Llama-3.1-8B0.98
    14. meta1.00
    15. llama1.00
    16. 3.11.00
    17. 8b1.00
    18. Pre-computed1.00
    19. search tokens1.00
    20. Tokens at Write Time1.00
    21. p99 Under 15 ms0.97
    22. No External Cache1.00
    23. Tokenized once at publish,1.00
    24. One query, one collection,1.00
    25. Right index,1.00
    26. searched millions of times1.00
    27. zero additional lookups1.00
    28. right data model.1.00
    29. One collection. One query. Pre-computed tokens in, ranked results out.0.99
    30. HUBGING FACE0.92
    31. LEADERSHIP 2· JUNE 30, 20260.97
    32. Al Architects: Show my Workflow0.99
  • 8:51 #22 skipped

    shot 22·duplicate of #21

  • 9:01 #23 skipped

    shot 23·duplicate of #21

Transcript

191 cues· 1,944 words· 11,197 chars

  1. 0:13 Good afternoon, everyone.
  2. 0:17 I have a question.
  3. 0:19 How many of you knows Hugging Face?
  4. 0:24 Nice.
  5. 0:27 How many of you already use Hugging Face?
  6. 0:33 Amazing, almost everyone.
  7. 0:36 but I think we still have opportunity to grow our usage.
  8. 0:41 My name is Arek Borucki.
  9. 0:43 I work as machine learning platform and database engineer at Hugging Face.
  10. 0:49 Today, I would like to walk you through how Hugging Face scaled infrastructure and how we ended up serving 3 million models to developers around the world.
  11. 1:09 I would like to share architectural decisions we made, challenges we faced, and lessons we learned while scaling one of the fastest growing open source AI communities in the world.
  12. 1:29 I hope you will enjoy it and let's get started.
  13. 1:37 Before I dive into technical details, let's talk about scale.
  14. 1:44 Today, Hugging Face serves more than 14 million users.
  15. 1:50 And this number is growing very fast, especially in the last couple of months.
  16. 1:57 We host 3 million public models, 1 million data sets,
  17. 2:07 50,000 organizations, and not only hobbyists or scientists.
  18. 2:15 More than 30% of Fortune 500 use Hugging Face as a part of AI workflows.
  19. 2:27 Just to give you some perspective, few years ago, we had 20,000 models
  20. 2:36 Today, three million.
  21. 2:39 It is around 150x increase in just last couple of years.
  22. 2:48 And this grow is exactly why I'm here today talking about infrastructure decisions that keep the hub healthy at scale.
  23. 3:03 This is how fast the number of public models is growing on the hub.
  24. 3:09 Every big release like Lama or DeepSeq generated thousands of new models on top.
  25. 3:20 And our infrastructure needs to handle that.
  26. 3:25 And it is not only models, also data sets.
  27. 3:31 In 2022, we had 10K.
  28. 3:35 In 2024, 100K.
  29. 3:39 Less than a year ago, we had 500K.
  30. 3:45 Today, 1 million.
  31. 3:49 All this data must be stored, indexed, and also must be searchable.
  32. 3:57 And that's the hardest part.
  33. 4:02 And this is also the reason why we had to rethink our search.
  34. 4:09 At 20,000 models, any query is fast, even without an index.
  35. 4:16 Trust me, no one would notice.
  36. 4:18 At three million, same approach breaks.
  37. 4:23 Imagine what would you do if the hub search would be slow.
  38. 4:30 you would just leave and go somewhere else.
  39. 4:32 And this is also what users are doing.
  40. 4:35 They expect fast, instant results.
  41. 4:39 With 14 million users, even 1% is a not small number.
  42. 4:47 It is 140,000 of people hitting slow search.
  43. 4:50 At scale,
  44. 4:58 P99 is much more important than P50.
  45. 5:04 And we are paying lots of attention to P99.
  46. 5:09 And that's the reason why we invest in pre-compute tokens, denormalize, optimize for read, collection in MongoDB, full text search based on Apache Lucene,
  47. 5:27 Kubernetes autoscaling, and soon in database sharding.
  48. 5:34 The next slides will show you how.
  49. 5:41 High-level architecture.
  50. 5:43 When user interact with the Hugging Face Hub, his request flows from the front end to the Hub API.

Chapters

  1. 0:00 Introduction: scaling the Hugging Face Hub
  2. 1:44 The numbers: 14 million users, millions of models
  3. 3:57 Why search at scale is the hard part
  4. 5:09 Full text search on Apache Lucene
  5. 5:46 Request flow: autoscaling, MongoDB Atlas, and S3
  6. 7:55 How a search for "llama" works
  7. 10:11 Ranking and Atlas Search with the $search operator
  8. 13:00 The seven node cluster and a hidden analytics node
  9. 16:42 Sharding the database
  10. 18:14 Kubernetes autoscaling: 10 to 500 pods and CastAI
  11. 20:07 Scaling on event loop utilization with KEDA

Open at this second