Videos XV2oYi7kojc
The Desktop Frontier — Ahmad Osman, Osmantic
Scene timeline
47 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 188
- whisperx 188
- chunks
- 32
- from 188 cues
- keyframes
- 20
- kept of 47 captured
- frames with text
- 19
- 202 lines read
- chapters
- 14
- from the source metadata
- keyframe bytes
- 5.1 MB
- word timings on 188 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 00:07 | 1m 17s |
stt |
done | — | 2026-08-11 00:08 | 17s |
chunk |
done | — | 2026-08-11 00:08 | 0s |
text_embed |
done | — | 2026-08-11 00:08 | 0s |
keyframe |
done | — | 2026-08-11 00:08 | 1m 59s |
ocr |
done | — | 2026-08-11 00:10 | 8s |
frame_embed |
done | — | 2026-08-11 00:10 | 4s |
Frames, and what the machine read
-
- AlEngineer0.95
- World's Fair0.97
-
- AlEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.95
- OpenAI0.92
- Akamai1.00
- arize1.00
- aws1.00
- Braintrust bright data0.97
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.91
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- AlEngineer0.96
- World's Fair0.98
-
- ounder & CEO, Osmantic0.97
- op Frontier1.00
- Workd's Fair0.94
-
- AIEngineer0.95
- World's Fair0.99
-
- AlEngineer0.99
- World's Fair0.96
-
- AIEngineer0.97
- World's Fair0.98
-
- AIEngineer0.95
- World'sFair1.00
- PREDICTION1.00
- PRESENTED BY1.00
- Within roughly 18 months, GLM-5.2-class intelligence1.00
- Microsoft1.00
- will fit on a 32GB GPU0.99
- By late 2027 or early 2028, a model that performs at GLM-5.2 level on important benchmarks and practical agent workloads could0.99
- plausibly run on an RTX 5090-class card with 32GB of VRAM.0.98
- Engineering the future of Al0.99
- World'sFair1.00
-
- AlEngineer0.96
- World'sFair1.00
- THETHESIS1.00
- PRESENTED BY1.00
- Frontier compression.1.00
- Microsoft1.00
- For a long time, the story was simple: bigger models, bigger clusters, better results. That story is still partly true.0.99
- But a second curve has become impossible to ignore over the last three years. Similar capability bands keep appearing in smaller0.99
- active footprints and more ownable systems.1.00
- Engineering the future of Al0.99
-
- AlEngineer0.97
- World'sFair1.00
- FRAMEWORK1.00
- Impact per parameter is0.99
- PRESENTED BY1.00
- Microsoft1.00
- What capability are we talking about?0.99
- What footprint did it require last year vs. now?0.99
- What hardware does that use?1.00
- How quickly is that hardware requirement1.00
- moving down?1.00
- Models are becoming more efficient, more portable, and easier to own.1.00
- TRACK 4· JULY 2,20260.96
- d'sFair1.00
- Local Al0.90
-
- AlEngineer0.96
- World'sFair1.00
- FRAMEWORK1.00
- Impact per parameter is1.00
- What capability are we talking about?0.99
- What footprint did it require last year vs. now?0.99
- What hardware does that use?0.99
- How quickly is that hardware requirement1.00
- moving down?1.00
- Models are becoming more efficient, more portable, and easier to own.1.00
- TRACK 4· JULY 2,20260.95
- Local Al0.91
-
- AlEngineer0.97
- World'sFair1.00
- DEFINITION1.00
- Similar capabilities moving into smaller hardware0.99
- What "similar capabilities" means0.99
- It is not that small models are beating big models, it is that1.00
- Similar or better benchmark scores1.00
- newer more efficient models are beating older less efficient0.99
- Same or better performance1.00
- ones.1.00
- World'sFair1.00
- TRACK 4· JULY 2,20260.96
- Local Al0.92
-
- AlEngineer0.97
- BENCHMARK1.00
- World'sFair1.00
- The Speed of Compression: From Llama 2 to the Frontier0.99
- The open-weight ecosystem has seen incredible progress in efficiency and capability. Looking back at leading models from the recent past, we0.99
- can clearly see the "densing law" in action.0.99
- Llama 2 (Q3 2023 Baseline)0.99
- The "Densing Law" in Action0.98
- Up to 70B total parameters1.00
- Similar or better capabilities with significantly fewer parameters1.00
- MMLU: ~68.9 on 70B model0.98
- Architectural innovations (e.g., MoE, hybrid models)1.00
- Pre-trained on 2 trillion tokens0.99
- Enhanced efficiency leading to smaller active footprints0.99
- Targeted general-purpose chat/reasoning0.99
- Reduced hardware requirements, enabling local inference1.00
- Requires significant GPU resources for inference (e.g., A100s)0.99
- Rapid iteration driving down cost-to-capability ratio1.00
- A major milestone for open-source LLMs, demonstrating strong1.00
- This relentless compression allows frontier-level intelligence to move1.00
- performance across a range of tasks.0.98
- onto more accessible hardware at an accelerating pace.1.00
- TRACK 4· JULY 2, 20260.93
- Local Al0.97
Transcript
188 cues· 2,408 words· 12,863 chars
- 0:12 Hey, everyone.
- 0:14 We are about to start this presentation.
- 0:18 It's called the Desktop Frontier.
- 0:20 And it's basically about where we started and how far we've come with local and open source models.
- 0:35 How many, like, just a quick question, how many of you here follow me on X?
- 0:42 I'm amazing.
- 0:43 Love you all.
- 0:44 Love you all.
- 0:46 So, you know, sometimes every now and then I would say a prediction.
- 0:50 Here is a new one.
- 0:52 Within roughly 18 months, we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 1590 with 32 gigabytes of VRAM.
- 1:07 That's basically late 2027.
- 1:10 This is conservative.
- 1:12 We might actually get there faster.
- 1:17 So for a long time, the story has been bigger models, bigger models, bigger models.
- 1:25 How can we get to the next 5 trillion?
- 1:27 How can we get to the 20 trillion?
- 1:29 And I'm not saying that there won't ever be a gap between frontier intelligence and open source models.
- 1:36 There will always be a gap.
- 1:37 But that gap will shrink, and the efficiency of the models will get exponentially better.
- 1:51 So the term that I like to think about is impact per parameter.
- 1:59 What capability are we talking about?
- 2:01 What could the model do?
- 2:04 What footprint, like hardware footprint, did it have last year in comparison to now?
- 2:10 And what hardware does that use?
- 2:13 And what hardware did it need to use a year ago?
- 2:16 And are we moving down for the same kind of quality
- 2:22 on that hardware.
- 2:23 Again, as I was saying earlier, I used to run Lama 2 on RTX 1390.
- 2:28 It's now running Quin 3.5, 3.6, 27 billion parameter.
- 2:32 That's better than Lama 3 405.
- 2:35 That's a 400 billion plus parameters model that you beat with a 27 billion parameter model a year and a half after.
- 2:48 So...
- 2:50 Yeah, as I was saying, similar capabilities are moving into smaller hardware footprint.
- 2:56 Benchmark scores are one thing, but also, you know, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within Cloud Code, right?
- 3:10 It wasn't until GLM 4.5 that came out in late July.
- 3:15 And GLM 4.5 Air required at least four RTX 1390s or an RTX Pro 6000.
- 3:24 Now, that footprint for hardware is not needed anymore.
- 3:27 All that you need is a single RTX 1390, 1590, and you have something much more capable, much more intelligent.
- 3:36 So is this trend just random?
- 3:39 Or is there more to it?
- 3:41 That's a question that everyone should ask.
- 3:45 Is it just by random chance that we've gotten this far from models that weren't able to sustain more than 4,000 tokens in terms of context lengths?
- 3:56 And now we have things that are million tokens locally on your hardware that you own.
- 4:06 It's not by chance.
- 4:07 It's not just a coincidence that we got here.
- 4:10 There is research being done.
- 4:11 There is efficiency gains to be made.
- 4:14 There are architecture hacks that compound, and they will continue to compound.
- 4:21 And I think I like this line.
loading
Chapters
- 0:00 Introduction and the Desktop Frontier concept
- 0:47 Future predictions: GLM 5.2 on an RTX 5090
- 1:17 Efficiency over raw size: The move toward compact intelligence
- 1:51 The concept of impact per parameter
- 2:48 Shifting hardware footprints: From server-grade to consumer-grade
- 3:38 Architecture hacks and the compounding nature of AI research
- 4:33 Explaining the Densing Law: Getting more intelligence from fewer parameters
- 5:09 Running frontier-class models like GLM 5.2 on local hardware
- 7:32 The case for sovereign AI: Owning your own compute stack
- 9:08 A retrospective on open-weight models: Mistral to Qwen
- 11:12 The evolution of reasoning: DeepSeek R1 and beyond
- 12:08 The rise of agentic performance and tool calling
- 15:33 Economic value: Does hardware appreciate as models become more efficient?
- 16:38 Closing thoughts: Why you should own your own GPU