Videos lyL5QhgIOxc
Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face
Scene timeline
50 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 191
- whisperx 191
- chunks
- 38
- from 191 cues
- keyframes
- 25
- kept of 50 captured
- frames with text
- 25
- 497 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 9.0 MB
- word timings on 191 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 22:22 | 0s |
stt |
done | — | 2026-08-09 06:11 | 17s |
chunk |
done | — | 2026-08-09 06:11 | 0s |
text_embed |
done | — | 2026-08-10 19:39 | 0s |
keyframe |
done | — | 2026-08-09 06:11 | 2m 44s |
ocr |
done | — | 2026-08-09 06:14 | 8s |
frame_embed |
done | — | 2026-08-10 19:39 | 5s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AIEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.92
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- Hugging Faca0.94
- HUGGING FACE1.00
- Hugging Face0.98
- AlEngineer0.95
- World's Fair0.99
-
- HuggingFue0.77
- HUGGING FACE1.00
- Fair1.00
-
- AlEngineer0.99
- Serving 3 Million Models0.98
- World's Fair0.99
- How Hugging Face Scaled the World's Largest Model Hub0.99
- TA1.00
- PRESENTED BY1.00
- Ⅲ0.52
- Microsoft1.00
- )0.95
- 27K1.00
- 50K1.00
- 35K1.00
- 40K1.00
- 41K1.00
- 50K1.00
- Arek Borucki1.00
- ML Platform & Database Engineer @ Hugging Face0.99
- MongoDB Champion | Author of MongoDB 8.0 in Action0.99
- HUGGING FACE1.00
- Engineering the future of Al1.00
-
- AlEngineer0.99
- Serving 3 Million Models0.99
- World's Fair0.97
- How Hugging Face Scaled the World's Largest Model Hub0.99
- TA1.00
- PRESENTED BY0.97
- Microsoft1.00
- )0.91
- 27K1.00
- 50K1.00
- 35K1.00
- 40K1.00
- 41K1.00
- 50K1.00
- Arek Borucki1.00
- ML Platform & Database Engineer @ Hugging Face0.99
- MongoDB Champion | Author of MongoDB 8.0 in Action1.00
- HUGGING FACE1.00
- LEADERSHIP 2· JUNE 30, 20260.94
- Al Architects: Show my Workflow0.99
-
- AlEngineer0.99
- Hugging Face: The Scale0.99
- World's Fair0.97
- 14M+1.00
- 2.9M+1.00
- 1M+1.00
- 50K+1.00
- USERS1.00
- PUBLIC MODELS1.00
- DATASETS1.00
- ORGANIZATIONS1.00
- DATA1.00
- DATA1.00
- DATA1.00
- USERS.0.98
- PUBLIC MODELS1.00
- DATASETS1.00
- ORGANIZATIONS1.00
- 30%+ of the Fortune 500 use Hugging Face1.00
- HUGGING FACE0.96
- LEADERSHIP 2· JUNE 30, 20260.95
- Al Architects: Show my Workflow0.99
-
- Halfway through 20261.00
- AlEngineer0.99
- 3M1.00
- World's Fair0.96
- 3M1.00
- Models1.00
- Cumsg M i o ace0.61
- 2M1.00
- 2.5M1.00
- Models1.00
- GPT-OSS1.00
- 2M1.00
- DeepSeek-R11.00
- 1M1.00
- 1.5M1.00
- Models1.00
- Flux 1.01.00
- 1M1.00
- Qwen21.00
- 500K1.00
- Models1.00
- 500K1.00
- 100K1.00
- LLaMA 20.98
- Models0.99
- BLOOM1.00
- o0.63
- 20221.00
- 20231.00
- 20241.00
- 20251.00
- 20261.00
- (H1)1.00
- HUGWING FACE0.93
- LEADERSHIP 2• JUNE 30, 20260.97
- Al Architects: Show my Workflow0.97
-
- Hugging Face just crossed 1M datasets0.99
- AlEngineer1.00
- Cumulative datasets published, by task category1.00
- World's Fair0.99
- 1.0M1.00
- 800k1.00
- Sep 20251.00
- datasets1.00
- Other1.00
- Ccumst dtets0.72
- Tabular0.98
- 600k1.00
- Oct 20221.00
- 10K1.00
- Feb 20241.00
- datasets1.00
- 100K0.99
- Reinforcement1.00
- Learning1.00
- datasets1.00
- Audio1.00
- Multimodal1.00
- 400k1.00
- Computer1.00
- Vision1.00
- Natural1.00
- Language1.00
- Processing1.00
- 200k1.00
- Apr1.00
- Jul0.96
- Oct1.00
- Jan1.00
- Apr1.00
- Jul1.00
- Oct1.00
- Jan1.00
- Apr1.00
- Jul1.00
- Oct1.00
- Jan1.00
- Apr0.98
- Jul1.00
- Oct1.00
- Jan1.00
- Apr1.00
- 20221.00
- 20231.00
- 20241.00
- 20251.00
- 20261.00
- HUGGING FACE1.00
- LEADERSHIP 2· JUNE 30, 20260.97
- Al Architects: Show my Workflow0.97
-
- AlEngineer0.98
- When 20k become 3 Million. Latency Matters0.99
- World'sFair1.00
- Users expect instant results.0.99
- Slow search = they leave.0.98
- At 20K models any query was fast.1.00
- At 3M the same approach breaks.1.00
- 14+M users, p50 lies.1.00
- What hurts UX is p99.0.99
- HUGGING FNE0.92
- LEADERSHIP 2· JUNE 30, 20260.95
- Al Architects: Show my Workflow0.99
-
- AlEngineer0.99
- World'sFair1.00
- PRESENTED BY1.00
- Microsoft1.00
- HUGGING FACE0.99
- LEADERSHIP 2· JUNE 30, 20260.95
- Al Architects: Show my Workflow0.97
-
- AlEngineer0.98
- High-level Architecture1.00
- World's Fair0.97
- Hub1.00
- Hub1.00
- MongoDB Atlas0.97
- Users1.00
- Frontend1.00
- API1.00
- (Source of Truth for Metadata)1.00
- (Holds metadata, card text,1.00
- PRESENTED BY0.97
- configuration data, user info)0.99
- Microsoft1.00
- Cloud Storage0.99
- (S3 / GCS - Binary Data & Models)0.98
- Holds all model weights, tokenizer files,0.99
- card assets, and configuration files1.00
- AGING FACE0.93
- LEADERSHIP 2· JUNE 30, 20260.95
- Al Architects: Show my Workflow0.99
-
- How we search 3M Models under 15ms0.99
- AlEngineer0.99
- World'sFair1.00
- Users1.00
- “Llama”0.97
- Hub1.00
- Optimized1.00
- results1.00
- Ranked1.00
- Read Collection1.00
- Results1.00
- Write time:1.00
- meta-llama/Llama-3.1-8B0.98
- meta1.00
- llama1.00
- 3.11.00
- 8b1.00
- Pre-computed1.00
- search tokens1.00
- Tokens at Write Time1.00
- p99 Under 15 ms0.97
- No External Cache1.00
- Tokenized once at publish,1.00
- One query, one collection,1.00
- Right index,1.00
- searched millions of times1.00
- zero additional lookups1.00
- right data model.1.00
- One collection. One query. Pre-computed tokens in, ranked results out.0.99
- HUBGING FACE0.92
- LEADERSHIP 2· JUNE 30, 20260.97
- Al Architects: Show my Workflow0.99
Transcript
191 cues· 1,944 words· 11,197 chars
- 0:13 Good afternoon, everyone.
- 0:17 I have a question.
- 0:19 How many of you knows Hugging Face?
- 0:24 Nice.
- 0:27 How many of you already use Hugging Face?
- 0:33 Amazing, almost everyone.
- 0:36 but I think we still have opportunity to grow our usage.
- 0:41 My name is Arek Borucki.
- 0:43 I work as machine learning platform and database engineer at Hugging Face.
- 0:49 Today, I would like to walk you through how Hugging Face scaled infrastructure and how we ended up serving 3 million models to developers around the world.
- 1:09 I would like to share architectural decisions we made, challenges we faced, and lessons we learned while scaling one of the fastest growing open source AI communities in the world.
- 1:29 I hope you will enjoy it and let's get started.
- 1:37 Before I dive into technical details, let's talk about scale.
- 1:44 Today, Hugging Face serves more than 14 million users.
- 1:50 And this number is growing very fast, especially in the last couple of months.
- 1:57 We host 3 million public models, 1 million data sets,
- 2:07 50,000 organizations, and not only hobbyists or scientists.
- 2:15 More than 30% of Fortune 500 use Hugging Face as a part of AI workflows.
- 2:27 Just to give you some perspective, few years ago, we had 20,000 models
- 2:36 Today, three million.
- 2:39 It is around 150x increase in just last couple of years.
- 2:48 And this grow is exactly why I'm here today talking about infrastructure decisions that keep the hub healthy at scale.
- 3:03 This is how fast the number of public models is growing on the hub.
- 3:09 Every big release like Lama or DeepSeq generated thousands of new models on top.
- 3:20 And our infrastructure needs to handle that.
- 3:25 And it is not only models, also data sets.
- 3:31 In 2022, we had 10K.
- 3:35 In 2024, 100K.
- 3:39 Less than a year ago, we had 500K.
- 3:45 Today, 1 million.
- 3:49 All this data must be stored, indexed, and also must be searchable.
- 3:57 And that's the hardest part.
- 4:02 And this is also the reason why we had to rethink our search.
- 4:09 At 20,000 models, any query is fast, even without an index.
- 4:16 Trust me, no one would notice.
- 4:18 At three million, same approach breaks.
- 4:23 Imagine what would you do if the hub search would be slow.
- 4:30 you would just leave and go somewhere else.
- 4:32 And this is also what users are doing.
- 4:35 They expect fast, instant results.
- 4:39 With 14 million users, even 1% is a not small number.
- 4:47 It is 140,000 of people hitting slow search.
- 4:50 At scale,
- 4:58 P99 is much more important than P50.
- 5:04 And we are paying lots of attention to P99.
- 5:09 And that's the reason why we invest in pre-compute tokens, denormalize, optimize for read, collection in MongoDB, full text search based on Apache Lucene,
- 5:27 Kubernetes autoscaling, and soon in database sharding.
- 5:34 The next slides will show you how.
- 5:41 High-level architecture.
- 5:43 When user interact with the Hugging Face Hub, his request flows from the front end to the Hub API.
loading
Chapters
- 0:00 Introduction: scaling the Hugging Face Hub
- 1:44 The numbers: 14 million users, millions of models
- 3:57 Why search at scale is the hard part
- 5:09 Full text search on Apache Lucene
- 5:46 Request flow: autoscaling, MongoDB Atlas, and S3
- 7:55 How a search for "llama" works
- 10:11 Ranking and Atlas Search with the $search operator
- 13:00 The seven node cluster and a hidden analytics node
- 16:42 Sharding the database
- 18:14 Kubernetes autoscaling: 10 to 500 pods and CastAI
- 20:07 Scaling on event loop utilization with KEDA