Videos Bck7ABCZRZI
Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
Scene timeline
18 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 87
- whisperx 87
- chunks
- 15
- from 87 cues
- keyframes
- 16
- kept of 18 captured
- frames with text
- 15
- 104 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 1.0 MB
- word timings on 87 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:43 | 1m 06s |
stt |
done | — | 2026-08-11 00:44 | 9s |
chunk |
done | — | 2026-08-11 00:44 | 0s |
text_embed |
done | — | 2026-08-11 00:44 | 0s |
keyframe |
done | — | 2026-08-11 00:44 | 19s |
ocr |
done | — | 2026-08-11 00:44 | 5s |
frame_embed |
done | — | 2026-08-11 00:44 | 2s |
Frames, and what the machine read
-
- THE ROOM YOU WEREN'T SHOWN0.99
- $200M1.00
- One enterprise. On the cost of inference. Then they built their own.1.00
- IMAGE - endless warehouse / hyperscale hall / cargo port.0.99
-
- THE DRAIN0.99
- The moment you rent a0.98
- token, it's expensive.0.99
- $25 per million looks like nothing. You load credits, load credits, load credits,0.99
- and never feel the spend until the budget is gone.0.99
- IMAGE - unopened mail / phone face-down / blinds shut.0.99
-
- THREE WEEKS AGO0.99
- I built UltaSuno on0.98
- rented intelligence.1.00
- Hundreds of users. Thousands of dollars in credits. Then my key got stolen, and0.98
- I watched a stranger drain the rest.0.99
- IMAGE - money in a torn bag / burn rate / counter ticking up.0.99
-
- THREE WEEKS AGO0.99
- I built UltaSuno on0.98
- rented intelligence.1.00
- Hundreds of users. Thousands of dollars in credits. Then my key got stolen, and0.98
- I watched a stranger drain the rest.1.00
- IMAGE - money in a torn bag / burn rate / counter ticking up.0.98
-
- THREE WEEKS AGO0.97
- I built UltaSuno on0.98
- rented intelligence.1.00
- Hundreds of users. Thousands of dollars in credits. Then my key got stolen, and0.99
- I watched a stranger drain the rest.0.99
- IMAGE - money in a torn bag / burn rate / counter ticking up.0.98
-
- THE SWITCH1.00
- Token factories just1.00
- switched the landlord.1.00
- Renting became leasing. Instead of paying Anthropic you pay Fireworks. Lost0.99
- keys, dumped tokens, retries, none of it goes away.1.00
- IMAGE - showroom / row of suits / slick sales floor.0.98
-
- ITRIED IT MYSELF0.99
- Local Inference0.97
- Memory was the1.00
- bottleneck1.00
- Raised the price citing a memory shortage. On a memory i0.98
- factory.1.00
- IMAGE - YOUR DGX Spark photo / brick-wall dead end.0.98
-
- THE POSITION0.96
- For an enterprise, renting1.00
- and leasing don't cut it.1.00
- The bill is only the first problem. The moment you rent, a second set of problems1.00
- shows up, and for an enterprise they're the ones that matter.0.99
- IMAGE - server room behind glass / access denied.0.98
-
- THREE REGULATED SHOPS1.00
- A fund. A hospital.0.99
- A tax practice.0.99
- Afund1.00
- Control. The reasoning can't sit behind someone else's rate limit.0.99
- Ahospital1.00
- Audit. You have to know exactly who you depend on, and prove it.0.99
- A tax practice0.94
- Reproducibility. If intelligence gave the answer, you must retrace it.0.99
-
- FIND YOUR POSITION1.00
- Where do you sit?1.00
- Pre-PMF1.00
- Keep renting. You don't know what you'll need yet.0.99
- Post-PMF1.00
- Start building. You have to support what you've found.1.00
- Enterprise1.00
- No choice. You build your own inference infra.0.98
- IMAGE - three doors / fork in the road / sorting line.0.98
-
- THE TRAP, NAMED0.99
- Treat it like Airbnb.0.99
- You have to buy your own house.1.00
- Rent1.00
- Cheap monthly. Never yours.0.99
- Lease1.00
- Locked in for years. Still not yours.1.00
- Own0.95
- Costs upfront. The only one that's yours.1.00
-
- NOT A HYPOTHESIS0.97
- So I built my own.0.97
- I call it JustInfer.0.99
- Rented, hit two problems. Leased, same two problems. DGX boxes wouldn't hold0.99
- reliability, so I went to bare metal and built the whole stack.0.99
- IMAGE - bare-metal GPU rack / hands-on the hardware.0.98
-
- WHAT I MADE ALONG THE WAY0.98
- Open sourced a tool1.00
- github.com/kalmantic.com/JusTokenMax1.00
- Wrote the book0.99
- PeakInference Infra Economics of AI Inference1.00
- JusTokenMax (MIT)1.00
- Open source. If you still rent or lease, it compresses input, cuts waste, adds1.00
- security.1.00
- The book0.95
- If you're an enterprise that has to build its own, the whole how-to is in here.1.00
-
- BEFORE MONDAY0.97
- Everyone sells one story.1.00
- Find your own answer.1.00
- Jensen, Nvidia0.98
- Token factory. He sells the factory.1.00
- Satya, Microsoft0.98
- Unmetered intelligence. Folded into a seat you already pay for.1.00
- Lin Qiao, Fireworks0.99
- Own your intelligence. On her infrastructure.1.00
- For enterprises, the answer is to own it. Find me at Al Engineer / X / LinkedIn.0.97
-
- Thank you.0.96
- JustInfer · JusTokenMax(MIT) · PeakInference0.96
- mtrajan.com x.com/mtrajan1.00
- IMAGE - an open door, light beyond - the exit. or reuse slide 1.0.99
Transcript
87 cues· 1,546 words· 8,402 chars
- 0:00 One of the largest retailers in the country spent close to $200 million on inference with Anthropic and decided that things got way out of hand and built their own infrastructure.
- 0:09 I'm pretty sure most of you have read the news from Uber CTO on how they had planned a budget of their tokens for an entire year and it got over in month four.
- 0:19 I'm also confident that half of you in this room have come to a very similar conclusion that as time goes by, the cost of intelligence really builds.
- 0:27 Using inference feels like, you know, it's one of the most inexpensive thing.
- 0:30 But then this is very different from using a phone where you get a bill once every month and then you have like
- 0:36 a specific set of amount that you can actually anchor your mind to but in case of using these rented intelligence platform they are like prepaid you load credits it's almost as if you're loading credits inside a casino you put some and then you pull it and then you are so addicted to it then you end up doing more and more of it and by some time you realize that you've blown past the threshold that you had mentally kept in mind and i had this
- 1:01 experience myself i built an app called ultra sono and i experienced the influence cost ballooning here suno.com has anybody heard about suno.com yeah suno.com is this application that allows a user to turn a text prompt into music what i was interested in is doing the reverse which is given a particular song what prompt could have actually generated it this is something that i wanted so i built this
- 1:26 And I was having a lot of fun using this application, shared it with a few friends and spread wide around.
- 1:31 I had hundreds of thousands of users, but then the cost ballooned way more than what I had anticipated.
- 1:38 Hundreds of thousands of dollars had to spend on inference.
- 1:41 Now, this happens for many reasons.
- 1:43 There are many talks that are there at AI itself where people talk about how you need to manage your context better.
- 1:50 And many people forget about doing compression of their input token.
- 1:54 And when there are agent loops, then there are many of these calls that are happening, which are very, very wasteful.
- 2:00 The inference endpoint that is consuming this is completely unaware of the shape of the workload and which is why this happens.
- 2:06 And I have this other issue that had happened.
- 2:08 Three weeks ago, my key got stolen.
- 2:11 Someone in China got hold of it and then was sucking my endpoint dry.
- 2:15 I could see the cost rise up from $7,000 to $7,500 to $8,000 and so on and so forth.
- 2:22 Thanks to my co-founder who heads the research and technology, we were able to arrest it at $10,000.
- 2:28 Otherwise, it could have been $100,000.
- 2:30 Now, many people suggest that the alternative to rented intelligence platform is to use Token Factory.
- 2:35 Token Factory is basically saying that why are you paying money to Anthropic and OpenAI?
- 2:39 Instead, go open source, have these open source models that are already deployed somewhere on the cloud, and then they are provisioned as tokens per second.
- 2:48 There are NeoClouds and then there are inference endpoint providers who actually do this.
- 2:52 In fact, there is also an argument saying that you can build this token factory locally.
- 2:56 There are AI Twitter influencers who actually talk about building inference in your garage, in your basement.
- 3:03 Buy GPU cards, rig them up together, and then you could actually run a local token factory.
- 3:08 In fact, I was inspired by that a little bit.
- 3:11 I bought my own DGX Sparks, and then I first moved Ultasuno from Anthropic to DGX Sparks.
- 3:19 It worked well.
- 3:19 I ran into this one issue of memory being the bottleneck, and then it was good enough that I started building my next applications.
- 3:26 I started having agents.
- 3:28 I have some agents that I need for running my research lab.
- 3:30 So these agents started shaping up inside the DGX Sparks, 2, 6, 8, 12, and it worked all right.
- 3:37 The issue though is that, you know, it may not be reliable for enterprise, which is what I exactly faced.
- 3:43 Three enterprises reached out to me to replicate the same setup for them.
- 3:47 But for enterprises, renting and leasing don't cut it.
- 3:50 Bill is a problem, but then there are secondary set of problems that makes it extremely ineffective approach.
- 3:57 The enterprises that I reached out to me, one was a fund, another a hospital, and the third a tax practice.
- 4:03 And each of them had different wall that they had hit.
- 4:06 The fund, it was an investment fund.
- 4:08 They were running an investment analyst on a Nemo architecture, and they didn't want somebody else to dictate as to what the rate limit that they could consume.
- 4:16 So control become a big issue for them to actually go with token factories.
- 4:21 Hospital had a different issue.
- 4:22 They used the use case, it worked well, but later when they went through an audit,
- 4:28 A third party vendor dependency was redlined and then they couldn't go forward.
- 4:32 The tax practice was a completely different issue.
- 4:35 In a tax practice, what is happening is when intelligent generates a recommendation, you want to be able to recreate it.
- 4:42 And when you don't have access to the in-depth of the model, you will not be able to do this.
loading