Videos qdh_x-uRs9g
The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked
Scene timeline
44 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 175
- whisperx 175
- chunks
- 32
- from 175 cues
- keyframes
- 34
- kept of 44 captured
- frames with text
- 34
- 762 lines read
- chapters
- 9
- from the source metadata
- keyframe bytes
- 4.7 MB
- word timings on 175 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 14:32 | 1m 30s |
stt |
done | — | 2026-08-10 14:33 | 18s |
chunk |
done | — | 2026-08-10 14:34 | 0s |
text_embed |
done | — | 2026-08-10 19:51 | 0s |
keyframe |
done | — | 2026-08-10 14:34 | 1m 26s |
ocr |
done | — | 2026-08-10 14:35 | 13s |
frame_embed |
done | — | 2026-08-10 19:51 | 6s |
Frames, and what the machine read
-
- AlEngineer0.98
- EUROPE1.00
-
- PRESENTINGSPONSOR1.00
- Google DeepMind1.00
-
- PLATINUM SPONSORS0.98
- # Braintrust0.96
- WorkOS OpenAI0.95
-
- The Small-Model1.00
- Infrastructure1.00
- Nobody Built1.00
- (so we did)1.00
- Filip Makraduli·Superlinked1.00
- AlEngineer0.99
- EUROPE1.00
- 20261.00
-
- *★*0.64
- A WRITER'S DIARY ON AI0.97
- AIE1.00
- What Actually Makes Embedding1.00
- Not quite.1.00
- ★1.00
- Model Inference Fast?0.99
- ★1.00
- ★1.00
- ★1.00
- From Flash Attention to Quantization, where is the inference bottleneck? Is it the0.99
- architecture, the maths, or will writing everything in Rust solve all my problems? (Hint:0.99
- It's Not Rust)1.00
- FILIP MAKRADULI0.97
- JAN 22, 20261.00
- Braintrust1.00
- WorkOS OpenAI0.97
- AlEngineer1.00
- EUROPE0.88
- 20260.99
-
- The aspect I overlooked was...0.99
- AIE1.00
- Inference1.00
- ★1.00
- ★1.00
- # Braintrust0.96
- WorkOS OpenAI0.95
- AlEngineer1.00
- EUROPE0.99
- 20260.97
-
- I realised .. I had to learn0.98
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- AlEngineer0.97
- AlEngineer1.00
- EUROPE1.00
- EUROPE0.99
- 20261.00
-
- LEARNING FROM FIRST PRINCIPLES0.98
- I knew models well0.99
- I had overlooked infrastructure0.99
- Training & fine-tuning0.99
- How models run in production0.99
- Evaluation & benchmarking0.99
- GPU utilization & scheduling1.00
- *★★0.66
- ★1.00
- Applied Al research1.00
- Routing & autoscaling0.99
- AIE1.00
- vLLM, Flash Attention1.00
- Deployment & operations0.98
- ★1.00
- 40.59
- ★1.00
- ★1.00
- ★1.00
- years of experience1.00
- blind spot0.99
- Best way to learn?0.99
- Join the team building it.0.99
- AlEngineer0.96
- AlEngineer0.99
- EUROPE1.00
- EUROPE0.99
- 20260.98
-
- So Ijoined the world class infrastructure team at:0.99
- Superlinked1.00
- github.com/superlinked/sie1.00
- ★1.00
- AIE1.00
- ★1.00
- VC - backed Al infrastructure company building the model stack for Al search and document0.99
- ★1.00
- processing1.00
- ★1.00
- ★1.00
- ★1.00
- Works with your favorite tools1.00
- BROWSE INTEGRATIONS1.00
- Chroma1.00
- DOCS →0.95
- LanceDB0.97
- DOcs →0.64
- Qdrant0.93
- DOCS →0.93
- Weaviate1.00
- DOCS →0.93
- "Chroma makes context1.00
- "LanceDB centralizes multi-0.97
- "Modern search systems1.00
- "Weaviate's Query Agent1.00
- engineering simple. SIE adds0.99
- modal training datasets and1.00
- compose the best indexing,1.00
- unlocks natural language1.00
- instruction-following rerankers1.00
- with SIE you can self-host0.99
- scoring, filtering and ranking1.00
- search and with SIE you can0.99
- and relationship extractors for1.00
- inference for all the required1.00
- models. With SIE you can self-0.99
- pre-process your query and0.99
- even more precise retrieval."1.00
- data transformations."1.00
- host them all in one cluster."0.99
- data for better latency."0.99
- CEO & Founder0.99
- Jeff Huber1.00
- CEO & Co-founder0.98
- Chang She0.99
- Andre Zayarni1.00
- CEO & Co-founder1.00
- Bob Van Luijt0.99
- CEO & Co-founder1.00
- Engineering the future of Al0.98
- AlEngineer0.99
- EUROPE0.99
- 20260.99
-
- KEYPOINTS1.00
- 1. Why this matters for your agents, even if you're1.00
- *★★0.58
- AIE1.00
- not doing Al search0.99
- ★1.00
- ★1.00
- ★★0.98
- 2. What small-model inference is not about1.00
- 3. The yin and yang of model inference1.00
- Engineering the future of Al1.00
- AlEngineer1.00
- EUROPE0.98
- 20260.97
-
- KEYPOINTS1.00
- 1. Why this matters for your agents, even if you're1.00
- AIE1.00
- not doing Al search0.99
- ★1.00
- ★★0.98
- 2. What small-model inference is not about1.00
- 3. The yin and yang of model inference1.00
- AlEngineer0.96
- AlEngineer1.00
- EUROPE1.00
- EUROPE0.99
- 20260.99
-
- THE ANTITHESIS1.00
- Why this matters for your1.00
- AIE1.00
- Agents and workflow ?0.99
- AlEngineer0.97
- AlEngineer1.00
- EUROPE1.00
- BUROPE0.89
- 20261.00
-
- WHY IT MATTERS FOR YOUR AGENTS0.98
- Contextrotis real0.98
- Context management1.00
- Performance degrades as context fills up.0.99
- More tokens, worse answers.0.98
- Manages what goes INTO context,1.00
- Repeated Words - Performance by Input Length (Tokens)0.98
- Claude Sonnet 41.00
- ***0.64
- AIE1.00
- Qwen3-3281.00
- Gemini 2.5 Flash0.97
- ★1.00
- Tool calling1.00
- ★1.00
- ★1.00
- ★1.00
- Embeddings, reranking, extraction1.00
- "Why not just Opus 4.6?"1.00
- Works until context rot kills quality1.00
- Input Length (Tokens)0.97
- 1040.79
- Source: Chroma Research1.00
- Engineering the future of Al1.00
- AlEngineer1.00
- EUROPE1.00
-
- WHYIT MATTERS FOR YOUR AGENTS0.97
- Context management requires a level of pre-processing1.00
- Andrej Karpathy1.00
- @karpathy1.00
- Accuracy drop when relevant info lands in the middle of context,0.99
- LLM Knowledge Bases0.99
- Something I'm finding very useful recently: using LLMs to build personal0.99
- knowledge bases for various topics of research interest. In this way, a0.99
- High1.00
- High1.00
- large fraction of my recent token throughput is going less into1.00
- *★*0.52
- AIE1.00
- ★1.00
- ★1.00
- Chroma Context-1: Training1.00
- ★1.00
- ★1.00
- ★1.00
- a Self-Editing Search Agent0.97
- Low1.00
- r/ClaudeCode1.00
- Join1.00
- u/captainkink07· 18h · github.com0.96
- 71.5x token reduction by1.00
- Position in context1.00
- compiling your raw folder into a1.00
- shamsi/graphify1.00
- Start1.00
- Middle1.00
- End1.00
- knowledge graph instead of1.00
- reading files. Built from0.99
- Liu et al. 2023·18 models tested by Chroma · 194K LLM calls0.98
- Karpathy's workflow0.98
- Showcase1.00
- Engineering the future of Al1.00
- AlEngineer1.00
-
- WHY IT MATTERS FOR YOUR AGENTS0.95
- A production story for tool calling0.99
- Context management1.00
- Taxonomy classification1.00
- 10K categories Shopify product catalogue text + images1.00
- Manages what goes INTO context,1.00
- *★★0.62
- ★1.00
- sie.extract()0.98
- NLI classification1.00
- AIE1.00
- ★0.99
- ★1.00
- ★1.00
- Product1.00
- title + image1.00
- sie.encode()1.00
- embed text +1.00
- images1.00
- Electronics1.00
- ,Computers0.95
- Laptops0.95
- Tool calling0.97
- Embeddings, reranking, extraction0.98
- sie.score()1.00
- rerank candidates0.99
- "Why not just Opus 4.6?"0.98
- Switching model type = changing one parameter, not0.99
- rebuilding infra0.97
- Works until context rot kills quality0.99
- Engineering the future of Al1.00
- AlEngineer1.00
- EUROPE1.00
-
- THE ANTITHESIS1.00
- What Inference1.00
- AIE1.00
- ★0.99
- ★1.00
- Doesn't Look Like0.97
- Engineering the future of Al0.98
- AlEngineer1.00
-
- IT'S NOT ABOUT FASTER MODELS OR BIGGER GPUS0.98
- One model per GPU — most capacity wasted0.98
- GPU1·24GB0.98
- GPU2·24GB1.00
- GPU3·24GB0.97
- GPU4·24GB1.00
- GPU5·24GB1.00
- stella1.00
- bge-m31.00
- reranker1.00
- gliner1.00
- clip1.00
- 20GB idle0.96
- 18GB idle1.00
- 21GB idle0.98
- 22GBidle1.00
- 21GB idle0.98
- *★★0.52
- AIE1.00
- ★1.00
- ★1.00
- ★1.00
- Multi-model GPU sharing — hot-swap with LRU eviction0.99
- 1GPU·24GB·80%+utilized1.00
- stella1.00
- bge-m31.00
- reranker1.00
- gliner1.00
- clip1.00
- LRU eviction: coldest model swaps out when memory is tight0.99
- Engineering the future of Al0.99
- AlEngineer0.99
- EUROPE1.00
-
- IT'S NOT ABOUT THE SERVER0.97
- Production1.00
- Model Server1.00
- Routing1.00
- per-model, pool isolation1.00
- Solved1.00
- Autoscaling1.00
- KEDA + spot instances0.99
- Scale-to-zero1.00
- $0 when idle0.97
- AIE1.00
- TEI·vLLM·FastAPI wrapper0.98
- ★1.00
- Monitoring1.00
- Prometheus + Grafana0.99
- ★1.00
- ★1.00
- one model, one container0.99
- MIND THE GAP0.99
- Terraform·Helm1.00
- AWS·GCP1.00
- Every team builds this layer0.99
- alone. It barely exists in0.99
- open source.1.00
- Braintrust1.00
- WorkOS OpenAI0.97
- AlEngineer1.00
- EUROPE1.00
- 20260.94
-
- IT'S NOT ABOUT THE SERVER0.97
- Production1.00
- Model Server1.00
- Routing1.00
- per-model, pool isolation1.00
- Solved1.00
- Autoscaling1.00
- KEDA + spot instances0.99
- Scale-to-zero1.00
- $0 when idle0.98
- AIE1.00
- ★1.00
- TEI·vLLM·FastAPI wrapper0.97
- ★1.00
- Monitoring1.00
- Prometheus + Grafana0.97
- ★1.00
- ★1.00
- one model, one container1.00
- MIND THE GAP0.98
- Terraform·Helm1.00
- AWS·GCP1.00
- Every team builds this layer1.00
- alone. It barely exists in1.00
- open source.1.00
- AlEngineer0.96
- AlEngineer0.99
- EUROPE1.00
- EUROPE0.99
- 20260.99
-
- THE ANTITHESIS1.00
- So what is0.97
- AIE1.00
- ★1.00
- ★1.00
- inference about?0.98
- AlEngineer0.97
- AlEngineer0.95
- EUROPE1.00
- EUROPE1.00
- 20261.00
-
- The Yin and Yang0.99
- AIE1.00
- of1.00
- ★1.00
- ★1.00
- Inference1.00
- AlEngineer0.96
- AlEngineer1.00
- EUROPE1.00
- EUROPE0.99
- 20260.99
Transcript
175 cues· 2,577 words· 13,996 chars
- 0:15 Hello, everyone.
- 0:17 Welcome to this talk.
- 0:19 I'll be speaking about small model inference and a gap that we've recognized in the market.
- 0:27 and what we did about it and why we kind of made this approach.
- 0:32 And as you can see, this background slide here, this is no accident.
- 0:36 So if you can guess what this is, I'll prompt you at the end of the slides, you win a little reward.
- 0:44 So you can catch me at the break afterwards.
- 0:47 So think about this, but also listen to me, so don't think too hard.
- 0:53 So the story starts with me posting an article a few months ago on Substack that got a bit of traction, got a few people interested.
- 1:03 And I explained flash attention, I explained how models worked, how processes can be memory bound, compute bound.
- 1:10 And I felt really good because I kind of went deep into this and as a person who's been in AI for a few years, I felt very confident.
- 1:21 And that was true, but then some people pointed out that actually I had overlooked one key aspect around what makes this model fast in the real world.
- 1:36 And that aspect that I overlooked was inference.
- 1:40 So as someone who kind of wants to understand things in first principles and work to understand the problems and the solutions deep, I realized, okay, I need to figure this out.
- 1:52 I need to know kind of where I've made my mistake.
- 1:57 And as a AI researcher and engineer, I need to find more about
- 2:02 inference.
- 2:03 So I've done a lot of work with VLLM, training models, fine-tuning, kind of doing applied ML and AI, did a bit of also research in academia as well.
- 2:16 But this part around kind of how models run in production, scheduling GPUs, routing and automation, I guess was a bit of a blind spot for me.
- 2:26 So I realized, okay, this is the time I have to now figure this out, learn
- 2:32 and make it work.
- 2:33 And what better way to do this than to actually build stuff?
- 2:38 So I decided to join a team, a team at Superlinked, comprised of very good infrastructure engineers, and actually work with them and build
- 2:52 something around inference.
- 2:54 And that something is this repo, the super linked inference engine that we have open sourced.
- 3:00 And this is kind of like the soft launch that I'm doing today.
- 3:04 So you can have a look at that later.
- 3:06 So basically it's inference for small models around AI search and document processing.
- 3:11 And we've tested this out, as you can see, with some of our partners.
- 3:14 So we've tested out with Chroma, Quadrant, VV8, so a lot of the vector DBs, as well as LensDB.
- 3:23 And as you can see, they've tried it out a bit.
- 3:26 Sounds fun.
- 3:26 Sounds interesting.
- 3:27 So this is working.
- 3:29 And it was the right step for me to kind of figure this out and learn
- 3:36 where inference kind of truly is and how that combines with my ML experience.
- 3:42 So the three key points I want you to go away with from this talk
- 3:48 are these.
- 3:49 So first, I want to tell you why this matters, so why doing inference for iSearch and document processing actually matters when you're building agents or when you're building workflows that involve agents.
- 4:04 So this is very important, the why.
- 4:07 Then I want to talk to you about the second thing, which is what inference is not about.
- 4:12 So there are some misconceptions or ideas of how inference looks like, but it's not all about those things.
- 4:20 And the third thing is how we see inference.
- 4:24 And I call this as like the yin and yang of model inference.
- 4:28 And it's a way of combining a few things around model support and infrastructure.
- 4:34 So why this matters for your agentic workflow?
- 4:39 Well, what you have encountered for sure is context rot.
- 4:44 And as we probably all know, this research paper from Chroma from some time ago showcases this effect that no matter what you do, there is this effect of context rot.
- 4:55 So quality degrades as context increases.
- 4:59 So being able to manage this context and do some context management is very important and a useful way to solve this.
loading
Chapters
- 0:00 Introduction and the gap in small model inference
- 0:53 Moving from research to building inference infrastructure
- 2:54 Introduction of the Superlinked inference engine
- 4:34 The importance of context management for agents
- 7:03 Misconceptions: Why more GPUs isn't the only answer
- 9:33 The "Yin and Yang" of inference: Model support and infrastructure
- 10:43 The challenge of supporting diverse model architectures
- 14:33 Deep dive into infrastructure and scalability
- 16:10 Conclusion and the open-source launch of SAI