read-only demo

Videos qdh_x-uRs9g

The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked

index_state ready data_status ok

AI Engineer· published 2026-05-05· 0:18:29· en-US· indexed 2026-08-10 19:51

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:05, 1 of 1 keyframes kept
  2. Shot 1, 0:05 to 0:09, 1 of 1 keyframes kept
  3. Shot 2, 0:09 to 0:14, 1 of 1 keyframes kept
  4. Shot 3, 0:14 to 0:52, 1 of 1 keyframes kept
  5. Shot 4, 0:52 to 1:36, 1 of 1 keyframes kept
  6. Shot 5, 1:36 to 1:41, 1 of 1 keyframes kept
  7. Shot 6, 1:41 to 2:01, 1 of 1 keyframes kept
  8. Shot 7, 2:01 to 2:41, 1 of 1 keyframes kept
  9. Shot 8, 2:41 to 3:11, 1 of 1 keyframes kept
  10. Shot 9, 3:11 to 3:41, 0 of 1 keyframes kept
  11. Shot 10, 3:41 to 4:07, 1 of 1 keyframes kept
  12. Shot 11, 4:07 to 4:34, 1 of 1 keyframes kept
  13. Shot 12, 4:34 to 4:38, 1 of 1 keyframes kept
  14. Shot 13, 4:38 to 5:10, 1 of 1 keyframes kept
  15. Shot 14, 5:10 to 5:43, 0 of 1 keyframes kept
  16. Shot 15, 5:43 to 6:30, 1 of 1 keyframes kept
  17. Shot 16, 6:30 to 7:02, 1 of 1 keyframes kept
  18. Shot 17, 7:02 to 7:10, 1 of 1 keyframes kept
  19. Shot 18, 7:10 to 7:46, 1 of 1 keyframes kept
  20. Shot 19, 7:46 to 8:22, 0 of 1 keyframes kept
  21. Shot 20, 8:22 to 8:52, 1 of 1 keyframes kept
  22. Shot 21, 8:52 to 9:22, 1 of 1 keyframes kept
  23. Shot 22, 9:22 to 9:35, 1 of 1 keyframes kept
  24. Shot 23, 9:35 to 9:47, 1 of 1 keyframes kept
  25. Shot 24, 9:47 to 10:14, 0 of 1 keyframes kept
  26. Shot 25, 10:14 to 10:42, 0 of 1 keyframes kept
  27. Shot 26, 10:42 to 11:09, 1 of 1 keyframes kept
  28. Shot 27, 11:09 to 11:39, 1 of 1 keyframes kept
  29. Shot 28, 11:39 to 12:09, 1 of 1 keyframes kept
  30. Shot 29, 12:09 to 12:38, 0 of 1 keyframes kept
  31. Shot 30, 12:38 to 13:06, 1 of 1 keyframes kept
  32. Shot 31, 13:06 to 13:34, 1 of 1 keyframes kept
  33. Shot 32, 13:34 to 14:01, 0 of 1 keyframes kept
  34. Shot 33, 14:01 to 14:29, 0 of 1 keyframes kept
  35. Shot 34, 14:29 to 14:34, 1 of 1 keyframes kept
  36. Shot 35, 14:34 to 14:55, 1 of 1 keyframes kept
  37. Shot 36, 14:55 to 15:32, 1 of 1 keyframes kept
  38. Shot 37, 15:32 to 16:09, 1 of 1 keyframes kept
  39. Shot 38, 16:09 to 16:44, 1 of 1 keyframes kept
  40. Shot 39, 16:44 to 17:19, 0 of 1 keyframes kept
  41. Shot 40, 17:19 to 17:25, 1 of 1 keyframes kept
  42. Shot 41, 17:25 to 17:43, 0 of 1 keyframes kept
  43. Shot 42, 17:43 to 18:14, 1 of 1 keyframes kept
  44. Shot 43, 18:14 to 18:29, 1 of 1 keyframes kept

44 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
175
whisperx 175
chunks
32
from 175 cues
keyframes
34
kept of 44 captured
frames with text
34
762 lines read
chapters
9
from the source metadata
keyframe bytes
4.7 MB
word timings on 175 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 14:32 1m 30s
stt done 2026-08-10 14:33 18s
chunk done 2026-08-10 14:34 0s
text_embed done 2026-08-10 19:51 0s
keyframe done 2026-08-10 14:34 1m 26s
ocr done 2026-08-10 14:35 13s
frame_embed done 2026-08-10 19:51 6s

Frames, and what the machine read

  • 0:03 #0 done2 line(s)

    shot 0·sharpness 658.4

    1. AlEngineer0.98
    2. EUROPE1.00
  • 0:08 #1 done2 line(s)

    shot 1·sharpness 829.1

    1. PRESENTINGSPONSOR1.00
    2. Google DeepMind1.00
  • 0:13 #2 done3 line(s)

    shot 2·sharpness 907.6

    1. PLATINUM SPONSORS0.98
    2. # Braintrust0.96
    3. WorkOS OpenAI0.95
  • 0:44 #3 done8 line(s)

    shot 3·sharpness 513.0

    1. The Small-Model1.00
    2. Infrastructure1.00
    3. Nobody Built1.00
    4. (so we did)1.00
    5. Filip Makraduli·Superlinked1.00
    6. AlEngineer0.99
    7. EUROPE1.00
    8. 20261.00
  • 1:26 #4 done20 line(s)

    shot 4·sharpness 1596.2

    1. *★*0.64
    2. A WRITER'S DIARY ON AI0.97
    3. AIE1.00
    4. What Actually Makes Embedding1.00
    5. Not quite.1.00
    6. 1.00
    7. Model Inference Fast?0.99
    8. 1.00
    9. 1.00
    10. 1.00
    11. From Flash Attention to Quantization, where is the inference bottleneck? Is it the0.99
    12. architecture, the maths, or will writing everything in Rust solve all my problems? (Hint:0.99
    13. It's Not Rust)1.00
    14. FILIP MAKRADULI0.97
    15. JAN 22, 20261.00
    16. Braintrust1.00
    17. WorkOS OpenAI0.97
    18. AlEngineer1.00
    19. EUROPE0.88
    20. 20260.99
  • 1:36 #5 done10 line(s)

    shot 5·sharpness 1007.4

    1. The aspect I overlooked was...0.99
    2. AIE1.00
    3. Inference1.00
    4. 1.00
    5. 1.00
    6. # Braintrust0.96
    7. WorkOS OpenAI0.95
    8. AlEngineer1.00
    9. EUROPE0.99
    10. 20260.97
  • 1:56 #6 done10 line(s)

    shot 6·sharpness 946.4

    1. I realised .. I had to learn0.98
    2. AIE1.00
    3. 1.00
    4. 1.00
    5. 1.00
    6. AlEngineer0.97
    7. AlEngineer1.00
    8. EUROPE1.00
    9. EUROPE0.99
    10. 20261.00
  • 2:21 #7 done28 line(s)

    shot 7·sharpness 2394.6

    1. LEARNING FROM FIRST PRINCIPLES0.98
    2. I knew models well0.99
    3. I had overlooked infrastructure0.99
    4. Training & fine-tuning0.99
    5. How models run in production0.99
    6. Evaluation & benchmarking0.99
    7. GPU utilization & scheduling1.00
    8. *★★0.66
    9. 1.00
    10. Applied Al research1.00
    11. Routing & autoscaling0.99
    12. AIE1.00
    13. vLLM, Flash Attention1.00
    14. Deployment & operations0.98
    15. 1.00
    16. 40.59
    17. 1.00
    18. 1.00
    19. 1.00
    20. years of experience1.00
    21. blind spot0.99
    22. Best way to learn?0.99
    23. Join the team building it.0.99
    24. AlEngineer0.96
    25. AlEngineer0.99
    26. EUROPE1.00
    27. EUROPE0.99
    28. 20260.98
  • 2:56 #8 done54 line(s)

    shot 8·sharpness 3114.3

    1. So Ijoined the world class infrastructure team at:0.99
    2. Superlinked1.00
    3. github.com/superlinked/sie1.00
    4. 1.00
    5. AIE1.00
    6. 1.00
    7. VC - backed Al infrastructure company building the model stack for Al search and document0.99
    8. 1.00
    9. processing1.00
    10. 1.00
    11. 1.00
    12. 1.00
    13. Works with your favorite tools1.00
    14. BROWSE INTEGRATIONS1.00
    15. Chroma1.00
    16. DOCS →0.95
    17. LanceDB0.97
    18. DOcs →0.64
    19. Qdrant0.93
    20. DOCS →0.93
    21. Weaviate1.00
    22. DOCS →0.93
    23. "Chroma makes context1.00
    24. "LanceDB centralizes multi-0.97
    25. "Modern search systems1.00
    26. "Weaviate's Query Agent1.00
    27. engineering simple. SIE adds0.99
    28. modal training datasets and1.00
    29. compose the best indexing,1.00
    30. unlocks natural language1.00
    31. instruction-following rerankers1.00
    32. with SIE you can self-host0.99
    33. scoring, filtering and ranking1.00
    34. search and with SIE you can0.99
    35. and relationship extractors for1.00
    36. inference for all the required1.00
    37. models. With SIE you can self-0.99
    38. pre-process your query and0.99
    39. even more precise retrieval."1.00
    40. data transformations."1.00
    41. host them all in one cluster."0.99
    42. data for better latency."0.99
    43. CEO & Founder0.99
    44. Jeff Huber1.00
    45. CEO & Co-founder0.98
    46. Chang She0.99
    47. Andre Zayarni1.00
    48. CEO & Co-founder1.00
    49. Bob Van Luijt0.99
    50. CEO & Co-founder1.00
    51. Engineering the future of Al0.98
    52. AlEngineer0.99
    53. EUROPE0.99
    54. 20260.99
  • 3:23 #9 skipped

    shot 9·duplicate of #8

  • 3:45 #10 done14 line(s)

    shot 10·sharpness 2796.1

    1. KEYPOINTS1.00
    2. 1. Why this matters for your agents, even if you're1.00
    3. *★★0.58
    4. AIE1.00
    5. not doing Al search0.99
    6. 1.00
    7. 1.00
    8. ★★0.98
    9. 2. What small-model inference is not about1.00
    10. 3. The yin and yang of model inference1.00
    11. Engineering the future of Al1.00
    12. AlEngineer1.00
    13. EUROPE0.98
    14. 20260.97
  • 4:16 #11 done13 line(s)

    shot 11·sharpness 2788.5

    1. KEYPOINTS1.00
    2. 1. Why this matters for your agents, even if you're1.00
    3. AIE1.00
    4. not doing Al search0.99
    5. 1.00
    6. ★★0.98
    7. 2. What small-model inference is not about1.00
    8. 3. The yin and yang of model inference1.00
    9. AlEngineer0.96
    10. AlEngineer1.00
    11. EUROPE1.00
    12. EUROPE0.99
    13. 20260.99
  • 4:37 #12 done9 line(s)

    shot 12·sharpness 1529.8

    1. THE ANTITHESIS1.00
    2. Why this matters for your1.00
    3. AIE1.00
    4. Agents and workflow ?0.99
    5. AlEngineer0.97
    6. AlEngineer1.00
    7. EUROPE1.00
    8. BUROPE0.89
    9. 20261.00
  • 4:54 #13 done26 line(s)

    shot 13·sharpness 2030.6

    1. WHY IT MATTERS FOR YOUR AGENTS0.98
    2. Contextrotis real0.98
    3. Context management1.00
    4. Performance degrades as context fills up.0.99
    5. More tokens, worse answers.0.98
    6. Manages what goes INTO context,1.00
    7. Repeated Words - Performance by Input Length (Tokens)0.98
    8. Claude Sonnet 41.00
    9. ***0.64
    10. AIE1.00
    11. Qwen3-3281.00
    12. Gemini 2.5 Flash0.97
    13. 1.00
    14. Tool calling1.00
    15. 1.00
    16. 1.00
    17. 1.00
    18. Embeddings, reranking, extraction1.00
    19. "Why not just Opus 4.6?"1.00
    20. Works until context rot kills quality1.00
    21. Input Length (Tokens)0.97
    22. 1040.79
    23. Source: Chroma Research1.00
    24. Engineering the future of Al1.00
    25. AlEngineer1.00
    26. EUROPE1.00
  • 5:30 #14 skipped

    shot 14·duplicate of #13

  • 5:48 #15 done38 line(s)

    shot 15·sharpness 2498.2

    1. WHYIT MATTERS FOR YOUR AGENTS0.97
    2. Context management requires a level of pre-processing1.00
    3. Andrej Karpathy1.00
    4. @karpathy1.00
    5. Accuracy drop when relevant info lands in the middle of context,0.99
    6. LLM Knowledge Bases0.99
    7. Something I'm finding very useful recently: using LLMs to build personal0.99
    8. knowledge bases for various topics of research interest. In this way, a0.99
    9. High1.00
    10. High1.00
    11. large fraction of my recent token throughput is going less into1.00
    12. *★*0.52
    13. AIE1.00
    14. 1.00
    15. 1.00
    16. Chroma Context-1: Training1.00
    17. 1.00
    18. 1.00
    19. 1.00
    20. a Self-Editing Search Agent0.97
    21. Low1.00
    22. r/ClaudeCode1.00
    23. Join1.00
    24. u/captainkink07· 18h · github.com0.96
    25. 71.5x token reduction by1.00
    26. Position in context1.00
    27. compiling your raw folder into a1.00
    28. shamsi/graphify1.00
    29. Start1.00
    30. Middle1.00
    31. End1.00
    32. knowledge graph instead of1.00
    33. reading files. Built from0.99
    34. Liu et al. 2023·18 models tested by Chroma · 194K LLM calls0.98
    35. Karpathy's workflow0.98
    36. Showcase1.00
    37. Engineering the future of Al1.00
    38. AlEngineer1.00
  • 6:55 #16 done33 line(s)

    shot 16·sharpness 2607.4

    1. WHY IT MATTERS FOR YOUR AGENTS0.95
    2. A production story for tool calling0.99
    3. Context management1.00
    4. Taxonomy classification1.00
    5. 10K categories Shopify product catalogue text + images1.00
    6. Manages what goes INTO context,1.00
    7. *★★0.62
    8. 1.00
    9. sie.extract()0.98
    10. NLI classification1.00
    11. AIE1.00
    12. 0.99
    13. 1.00
    14. 1.00
    15. Product1.00
    16. title + image1.00
    17. sie.encode()1.00
    18. embed text +1.00
    19. images1.00
    20. Electronics1.00
    21. ,Computers0.95
    22. Laptops0.95
    23. Tool calling0.97
    24. Embeddings, reranking, extraction0.98
    25. sie.score()1.00
    26. rerank candidates0.99
    27. "Why not just Opus 4.6?"0.98
    28. Switching model type = changing one parameter, not0.99
    29. rebuilding infra0.97
    30. Works until context rot kills quality0.99
    31. Engineering the future of Al1.00
    32. AlEngineer1.00
    33. EUROPE1.00
  • 7:08 #17 done8 line(s)

    shot 17·sharpness 1299.5

    1. THE ANTITHESIS1.00
    2. What Inference1.00
    3. AIE1.00
    4. 0.99
    5. 1.00
    6. Doesn't Look Like0.97
    7. Engineering the future of Al0.98
    8. AlEngineer1.00
  • 7:28 #18 done33 line(s)

    shot 18·sharpness 1753.1

    1. IT'S NOT ABOUT FASTER MODELS OR BIGGER GPUS0.98
    2. One model per GPU — most capacity wasted0.98
    3. GPU1·24GB0.98
    4. GPU2·24GB1.00
    5. GPU3·24GB0.97
    6. GPU4·24GB1.00
    7. GPU5·24GB1.00
    8. stella1.00
    9. bge-m31.00
    10. reranker1.00
    11. gliner1.00
    12. clip1.00
    13. 20GB idle0.96
    14. 18GB idle1.00
    15. 21GB idle0.98
    16. 22GBidle1.00
    17. 21GB idle0.98
    18. *★★0.52
    19. AIE1.00
    20. 1.00
    21. 1.00
    22. 1.00
    23. Multi-model GPU sharing — hot-swap with LRU eviction0.99
    24. 1GPU·24GB·80%+utilized1.00
    25. stella1.00
    26. bge-m31.00
    27. reranker1.00
    28. gliner1.00
    29. clip1.00
    30. LRU eviction: coldest model swaps out when memory is tight0.99
    31. Engineering the future of Al0.99
    32. AlEngineer0.99
    33. EUROPE1.00
  • 8:04 #19 skipped

    shot 19·duplicate of #10

  • 8:34 #20 done29 line(s)

    shot 20·sharpness 1877.8

    1. IT'S NOT ABOUT THE SERVER0.97
    2. Production1.00
    3. Model Server1.00
    4. Routing1.00
    5. per-model, pool isolation1.00
    6. Solved1.00
    7. Autoscaling1.00
    8. KEDA + spot instances0.99
    9. Scale-to-zero1.00
    10. $0 when idle0.97
    11. AIE1.00
    12. TEI·vLLM·FastAPI wrapper0.98
    13. 1.00
    14. Monitoring1.00
    15. Prometheus + Grafana0.99
    16. 1.00
    17. 1.00
    18. one model, one container0.99
    19. MIND THE GAP0.99
    20. Terraform·Helm1.00
    21. AWS·GCP1.00
    22. Every team builds this layer0.99
    23. alone. It barely exists in0.99
    24. open source.1.00
    25. Braintrust1.00
    26. WorkOS OpenAI0.97
    27. AlEngineer1.00
    28. EUROPE1.00
    29. 20260.94
  • 9:10 #21 done30 line(s)

    shot 21·sharpness 1777.8

    1. IT'S NOT ABOUT THE SERVER0.97
    2. Production1.00
    3. Model Server1.00
    4. Routing1.00
    5. per-model, pool isolation1.00
    6. Solved1.00
    7. Autoscaling1.00
    8. KEDA + spot instances0.99
    9. Scale-to-zero1.00
    10. $0 when idle0.98
    11. AIE1.00
    12. 1.00
    13. TEI·vLLM·FastAPI wrapper0.97
    14. 1.00
    15. Monitoring1.00
    16. Prometheus + Grafana0.97
    17. 1.00
    18. 1.00
    19. one model, one container1.00
    20. MIND THE GAP0.98
    21. Terraform·Helm1.00
    22. AWS·GCP1.00
    23. Every team builds this layer1.00
    24. alone. It barely exists in1.00
    25. open source.1.00
    26. AlEngineer0.96
    27. AlEngineer0.99
    28. EUROPE1.00
    29. EUROPE0.99
    30. 20260.99
  • 9:31 #22 done11 line(s)

    shot 22·sharpness 1159.2

    1. THE ANTITHESIS1.00
    2. So what is0.97
    3. AIE1.00
    4. 1.00
    5. 1.00
    6. inference about?0.98
    7. AlEngineer0.97
    8. AlEngineer0.95
    9. EUROPE1.00
    10. EUROPE1.00
    11. 20261.00
  • 9:43 #23 done11 line(s)

    shot 23·sharpness 1211.3

    1. The Yin and Yang0.99
    2. AIE1.00
    3. of1.00
    4. 1.00
    5. 1.00
    6. Inference1.00
    7. AlEngineer0.96
    8. AlEngineer1.00
    9. EUROPE1.00
    10. EUROPE0.99
    11. 20260.99

Transcript

175 cues· 2,577 words· 13,996 chars

  1. 0:15 Hello, everyone.
  2. 0:17 Welcome to this talk.
  3. 0:19 I'll be speaking about small model inference and a gap that we've recognized in the market.
  4. 0:27 and what we did about it and why we kind of made this approach.
  5. 0:32 And as you can see, this background slide here, this is no accident.
  6. 0:36 So if you can guess what this is, I'll prompt you at the end of the slides, you win a little reward.
  7. 0:44 So you can catch me at the break afterwards.
  8. 0:47 So think about this, but also listen to me, so don't think too hard.
  9. 0:53 So the story starts with me posting an article a few months ago on Substack that got a bit of traction, got a few people interested.
  10. 1:03 And I explained flash attention, I explained how models worked, how processes can be memory bound, compute bound.
  11. 1:10 And I felt really good because I kind of went deep into this and as a person who's been in AI for a few years, I felt very confident.
  12. 1:21 And that was true, but then some people pointed out that actually I had overlooked one key aspect around what makes this model fast in the real world.
  13. 1:36 And that aspect that I overlooked was inference.
  14. 1:40 So as someone who kind of wants to understand things in first principles and work to understand the problems and the solutions deep, I realized, okay, I need to figure this out.
  15. 1:52 I need to know kind of where I've made my mistake.
  16. 1:57 And as a AI researcher and engineer, I need to find more about
  17. 2:02 inference.
  18. 2:03 So I've done a lot of work with VLLM, training models, fine-tuning, kind of doing applied ML and AI, did a bit of also research in academia as well.
  19. 2:16 But this part around kind of how models run in production, scheduling GPUs, routing and automation, I guess was a bit of a blind spot for me.
  20. 2:26 So I realized, okay, this is the time I have to now figure this out, learn
  21. 2:32 and make it work.
  22. 2:33 And what better way to do this than to actually build stuff?
  23. 2:38 So I decided to join a team, a team at Superlinked, comprised of very good infrastructure engineers, and actually work with them and build
  24. 2:52 something around inference.
  25. 2:54 And that something is this repo, the super linked inference engine that we have open sourced.
  26. 3:00 And this is kind of like the soft launch that I'm doing today.
  27. 3:04 So you can have a look at that later.
  28. 3:06 So basically it's inference for small models around AI search and document processing.
  29. 3:11 And we've tested this out, as you can see, with some of our partners.
  30. 3:14 So we've tested out with Chroma, Quadrant, VV8, so a lot of the vector DBs, as well as LensDB.
  31. 3:23 And as you can see, they've tried it out a bit.
  32. 3:26 Sounds fun.
  33. 3:26 Sounds interesting.
  34. 3:27 So this is working.
  35. 3:29 And it was the right step for me to kind of figure this out and learn
  36. 3:36 where inference kind of truly is and how that combines with my ML experience.
  37. 3:42 So the three key points I want you to go away with from this talk
  38. 3:48 are these.
  39. 3:49 So first, I want to tell you why this matters, so why doing inference for iSearch and document processing actually matters when you're building agents or when you're building workflows that involve agents.
  40. 4:04 So this is very important, the why.
  41. 4:07 Then I want to talk to you about the second thing, which is what inference is not about.
  42. 4:12 So there are some misconceptions or ideas of how inference looks like, but it's not all about those things.
  43. 4:20 And the third thing is how we see inference.
  44. 4:24 And I call this as like the yin and yang of model inference.
  45. 4:28 And it's a way of combining a few things around model support and infrastructure.
  46. 4:34 So why this matters for your agentic workflow?
  47. 4:39 Well, what you have encountered for sure is context rot.
  48. 4:44 And as we probably all know, this research paper from Chroma from some time ago showcases this effect that no matter what you do, there is this effect of context rot.
  49. 4:55 So quality degrades as context increases.
  50. 4:59 So being able to manage this context and do some context management is very important and a useful way to solve this.

Chapters

  1. 0:00 Introduction and the gap in small model inference
  2. 0:53 Moving from research to building inference infrastructure
  3. 2:54 Introduction of the Superlinked inference engine
  4. 4:34 The importance of context management for agents
  5. 7:03 Misconceptions: Why more GPUs isn't the only answer
  6. 9:33 The "Yin and Yang" of inference: Model support and infrastructure
  7. 10:43 The challenge of supporting diverse model architectures
  8. 14:33 Deep dive into infrastructure and scalability
  9. 16:10 Conclusion and the open-source launch of SAI

Open at this second