read-only demo

Videos XV2oYi7kojc

The Desktop Frontier — Ahmad Osman, Osmantic

index_state ready data_status ok

AI Engineer· published 2026-07-21· 0:18:01· en-US· indexed 2026-08-11 00:10

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:24, 1 of 1 keyframes kept
  5. Shot 4, 0:24 to 0:42, 1 of 1 keyframes kept
  6. Shot 5, 0:42 to 0:46, 1 of 1 keyframes kept
  7. Shot 6, 0:46 to 0:52, 1 of 1 keyframes kept
  8. Shot 7, 0:52 to 0:54, 1 of 1 keyframes kept
  9. Shot 8, 0:54 to 1:16, 1 of 1 keyframes kept
  10. Shot 9, 1:16 to 1:50, 1 of 1 keyframes kept
  11. Shot 10, 1:50 to 2:18, 1 of 1 keyframes kept
  12. Shot 11, 2:18 to 2:47, 1 of 1 keyframes kept
  13. Shot 12, 2:47 to 3:13, 1 of 1 keyframes kept
  14. Shot 13, 3:13 to 3:39, 0 of 1 keyframes kept
  15. Shot 14, 3:39 to 4:05, 0 of 1 keyframes kept
  16. Shot 15, 4:05 to 4:31, 0 of 1 keyframes kept
  17. Shot 16, 4:31 to 5:08, 0 of 1 keyframes kept
  18. Shot 17, 5:08 to 5:36, 0 of 1 keyframes kept
  19. Shot 18, 5:36 to 6:03, 0 of 1 keyframes kept
  20. Shot 19, 6:03 to 6:31, 0 of 1 keyframes kept
  21. Shot 20, 6:31 to 6:57, 0 of 1 keyframes kept
  22. Shot 21, 6:57 to 7:23, 0 of 1 keyframes kept
  23. Shot 22, 7:23 to 7:49, 0 of 1 keyframes kept
  24. Shot 23, 7:49 to 8:15, 1 of 1 keyframes kept
  25. Shot 24, 8:15 to 8:40, 0 of 1 keyframes kept
  26. Shot 25, 8:40 to 9:06, 0 of 1 keyframes kept
  27. Shot 26, 9:06 to 9:11, 0 of 1 keyframes kept
  28. Shot 27, 9:11 to 9:50, 0 of 1 keyframes kept
  29. Shot 28, 9:50 to 10:17, 0 of 1 keyframes kept
  30. Shot 29, 10:17 to 10:44, 1 of 1 keyframes kept
  31. Shot 30, 10:44 to 11:11, 1 of 1 keyframes kept
  32. Shot 31, 11:11 to 11:51, 0 of 1 keyframes kept
  33. Shot 32, 11:51 to 12:18, 1 of 1 keyframes kept
  34. Shot 33, 12:18 to 12:46, 0 of 1 keyframes kept
  35. Shot 34, 12:46 to 13:13, 0 of 1 keyframes kept
  36. Shot 35, 13:13 to 13:40, 0 of 1 keyframes kept
  37. Shot 36, 13:40 to 14:06, 0 of 1 keyframes kept
  38. Shot 37, 14:06 to 14:51, 0 of 1 keyframes kept
  39. Shot 38, 14:51 to 15:29, 0 of 1 keyframes kept
  40. Shot 39, 15:29 to 15:56, 0 of 1 keyframes kept
  41. Shot 40, 15:56 to 16:23, 1 of 1 keyframes kept
  42. Shot 41, 16:23 to 16:50, 0 of 1 keyframes kept
  43. Shot 42, 16:50 to 17:18, 0 of 1 keyframes kept
  44. Shot 43, 17:18 to 17:39, 0 of 1 keyframes kept
  45. Shot 44, 17:39 to 17:44, 1 of 1 keyframes kept
  46. Shot 45, 17:44 to 18:00, 0 of 1 keyframes kept
  47. Shot 46, 18:00 to 18:01, 1 of 1 keyframes kept

47 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
188
whisperx 188
chunks
32
from 188 cues
keyframes
20
kept of 47 captured
frames with text
19
202 lines read
chapters
14
from the source metadata
keyframe bytes
5.1 MB
word timings on 188 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 00:07 1m 17s
stt done 2026-08-11 00:08 17s
chunk done 2026-08-11 00:08 0s
text_embed done 2026-08-11 00:08 0s
keyframe done 2026-08-11 00:08 1m 59s
ocr done 2026-08-11 00:10 8s
frame_embed done 2026-08-11 00:10 4s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 457.7

    1. AlEngineer0.95
    2. World's Fair0.97
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 661.9

    1. AlEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2740.4

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.95
    6. OpenAI0.92
    7. Akamai1.00
    8. arize1.00
    9. aws1.00
    10. Braintrust bright data0.97
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.91
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:23 #3 done2 line(s)

    shot 3·sharpness 281.0

    1. AlEngineer0.96
    2. World's Fair0.98
  • 0:31 #4 done3 line(s)

    shot 4·sharpness 450.9

    1. ounder & CEO, Osmantic0.97
    2. op Frontier1.00
    3. Workd's Fair0.94
  • 0:42 #5 done2 line(s)

    shot 5·sharpness 275.7

    1. AIEngineer0.95
    2. World's Fair0.99
  • 0:48 #6 done2 line(s)

    shot 6·sharpness 311.6

    1. AlEngineer0.99
    2. World's Fair0.96
  • 0:53 #7 done2 line(s)

    shot 7·sharpness 248.7

    1. AIEngineer0.97
    2. World's Fair0.98
  • 1:13 #8 done11 line(s)

    shot 8·sharpness 2499.0

    1. AIEngineer0.95
    2. World'sFair1.00
    3. PREDICTION1.00
    4. PRESENTED BY1.00
    5. Within roughly 18 months, GLM-5.2-class intelligence1.00
    6. Microsoft1.00
    7. will fit on a 32GB GPU0.99
    8. By late 2027 or early 2028, a model that performs at GLM-5.2 level on important benchmarks and practical agent workloads could0.99
    9. plausibly run on an RTX 5090-class card with 32GB of VRAM.0.98
    10. Engineering the future of Al0.99
    11. World'sFair1.00
  • 1:23 #9 done10 line(s)

    shot 9·sharpness 2092.4

    1. AlEngineer0.96
    2. World'sFair1.00
    3. THETHESIS1.00
    4. PRESENTED BY1.00
    5. Frontier compression.1.00
    6. Microsoft1.00
    7. For a long time, the story was simple: bigger models, bigger clusters, better results. That story is still partly true.0.99
    8. But a second curve has become impossible to ignore over the last three years. Similar capability bands keep appearing in smaller0.99
    9. active footprints and more ownable systems.1.00
    10. Engineering the future of Al0.99
  • 1:53 #10 done15 line(s)

    shot 10·sharpness 2055.6

    1. AlEngineer0.97
    2. World'sFair1.00
    3. FRAMEWORK1.00
    4. Impact per parameter is0.99
    5. PRESENTED BY1.00
    6. Microsoft1.00
    7. What capability are we talking about?0.99
    8. What footprint did it require last year vs. now?0.99
    9. What hardware does that use?1.00
    10. How quickly is that hardware requirement1.00
    11. moving down?1.00
    12. Models are becoming more efficient, more portable, and easier to own.1.00
    13. TRACK 4· JULY 2,20260.96
    14. d'sFair1.00
    15. Local Al0.90
  • 2:24 #11 done12 line(s)

    shot 11·sharpness 1922.1

    1. AlEngineer0.96
    2. World'sFair1.00
    3. FRAMEWORK1.00
    4. Impact per parameter is1.00
    5. What capability are we talking about?0.99
    6. What footprint did it require last year vs. now?0.99
    7. What hardware does that use?0.99
    8. How quickly is that hardware requirement1.00
    9. moving down?1.00
    10. Models are becoming more efficient, more portable, and easier to own.1.00
    11. TRACK 4· JULY 2,20260.95
    12. Local Al0.91
  • 2:50 #12 done13 line(s)

    shot 12·sharpness 2026.0

    1. AlEngineer0.97
    2. World'sFair1.00
    3. DEFINITION1.00
    4. Similar capabilities moving into smaller hardware0.99
    5. What "similar capabilities" means0.99
    6. It is not that small models are beating big models, it is that1.00
    7. Similar or better benchmark scores1.00
    8. newer more efficient models are beating older less efficient0.99
    9. Same or better performance1.00
    10. ones.1.00
    11. World'sFair1.00
    12. TRACK 4· JULY 2,20260.96
    13. Local Al0.92
  • 3:33 #13 skipped

    shot 13·duplicate of #11

  • 3:49 #14 skipped

    shot 14·duplicate of #11

  • 4:10 #15 skipped

    shot 15·duplicate of #12

  • 4:42 #16 skipped

    shot 16·duplicate of #11

  • 5:16 #17 skipped

    shot 17·duplicate of #11

  • 6:00 #18 skipped

    shot 18·duplicate of #9

  • 6:09 #19 skipped

    shot 19·duplicate of #12

  • 6:34 #20 skipped

    shot 20·duplicate of #10

  • 7:03 #21 skipped

    shot 21·duplicate of #10

  • 7:36 #22 skipped

    shot 22·duplicate of #10

  • 8:11 #23 done24 line(s)

    shot 23·sharpness 3356.3

    1. AlEngineer0.97
    2. BENCHMARK1.00
    3. World'sFair1.00
    4. The Speed of Compression: From Llama 2 to the Frontier0.99
    5. The open-weight ecosystem has seen incredible progress in efficiency and capability. Looking back at leading models from the recent past, we0.99
    6. can clearly see the "densing law" in action.0.99
    7. Llama 2 (Q3 2023 Baseline)0.99
    8. The "Densing Law" in Action0.98
    9. Up to 70B total parameters1.00
    10. Similar or better capabilities with significantly fewer parameters1.00
    11. MMLU: ~68.9 on 70B model0.98
    12. Architectural innovations (e.g., MoE, hybrid models)1.00
    13. Pre-trained on 2 trillion tokens0.99
    14. Enhanced efficiency leading to smaller active footprints0.99
    15. Targeted general-purpose chat/reasoning0.99
    16. Reduced hardware requirements, enabling local inference1.00
    17. Requires significant GPU resources for inference (e.g., A100s)0.99
    18. Rapid iteration driving down cost-to-capability ratio1.00
    19. A major milestone for open-source LLMs, demonstrating strong1.00
    20. This relentless compression allows frontier-level intelligence to move1.00
    21. performance across a range of tasks.0.98
    22. onto more accessible hardware at an accelerating pace.1.00
    23. TRACK 4· JULY 2, 20260.93
    24. Local Al0.97

Transcript

188 cues· 2,408 words· 12,863 chars

  1. 0:12 Hey, everyone.
  2. 0:14 We are about to start this presentation.
  3. 0:18 It's called the Desktop Frontier.
  4. 0:20 And it's basically about where we started and how far we've come with local and open source models.
  5. 0:35 How many, like, just a quick question, how many of you here follow me on X?
  6. 0:42 I'm amazing.
  7. 0:43 Love you all.
  8. 0:44 Love you all.
  9. 0:46 So, you know, sometimes every now and then I would say a prediction.
  10. 0:50 Here is a new one.
  11. 0:52 Within roughly 18 months, we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 1590 with 32 gigabytes of VRAM.
  12. 1:07 That's basically late 2027.
  13. 1:10 This is conservative.
  14. 1:12 We might actually get there faster.
  15. 1:17 So for a long time, the story has been bigger models, bigger models, bigger models.
  16. 1:25 How can we get to the next 5 trillion?
  17. 1:27 How can we get to the 20 trillion?
  18. 1:29 And I'm not saying that there won't ever be a gap between frontier intelligence and open source models.
  19. 1:36 There will always be a gap.
  20. 1:37 But that gap will shrink, and the efficiency of the models will get exponentially better.
  21. 1:51 So the term that I like to think about is impact per parameter.
  22. 1:59 What capability are we talking about?
  23. 2:01 What could the model do?
  24. 2:04 What footprint, like hardware footprint, did it have last year in comparison to now?
  25. 2:10 And what hardware does that use?
  26. 2:13 And what hardware did it need to use a year ago?
  27. 2:16 And are we moving down for the same kind of quality
  28. 2:22 on that hardware.
  29. 2:23 Again, as I was saying earlier, I used to run Lama 2 on RTX 1390.
  30. 2:28 It's now running Quin 3.5, 3.6, 27 billion parameter.
  31. 2:32 That's better than Lama 3 405.
  32. 2:35 That's a 400 billion plus parameters model that you beat with a 27 billion parameter model a year and a half after.
  33. 2:48 So...
  34. 2:50 Yeah, as I was saying, similar capabilities are moving into smaller hardware footprint.
  35. 2:56 Benchmark scores are one thing, but also, you know, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within Cloud Code, right?
  36. 3:10 It wasn't until GLM 4.5 that came out in late July.
  37. 3:15 And GLM 4.5 Air required at least four RTX 1390s or an RTX Pro 6000.
  38. 3:24 Now, that footprint for hardware is not needed anymore.
  39. 3:27 All that you need is a single RTX 1390, 1590, and you have something much more capable, much more intelligent.
  40. 3:36 So is this trend just random?
  41. 3:39 Or is there more to it?
  42. 3:41 That's a question that everyone should ask.
  43. 3:45 Is it just by random chance that we've gotten this far from models that weren't able to sustain more than 4,000 tokens in terms of context lengths?
  44. 3:56 And now we have things that are million tokens locally on your hardware that you own.
  45. 4:06 It's not by chance.
  46. 4:07 It's not just a coincidence that we got here.
  47. 4:10 There is research being done.
  48. 4:11 There is efficiency gains to be made.
  49. 4:14 There are architecture hacks that compound, and they will continue to compound.
  50. 4:21 And I think I like this line.

Chapters

  1. 0:00 Introduction and the Desktop Frontier concept
  2. 0:47 Future predictions: GLM 5.2 on an RTX 5090
  3. 1:17 Efficiency over raw size: The move toward compact intelligence
  4. 1:51 The concept of impact per parameter
  5. 2:48 Shifting hardware footprints: From server-grade to consumer-grade
  6. 3:38 Architecture hacks and the compounding nature of AI research
  7. 4:33 Explaining the Densing Law: Getting more intelligence from fewer parameters
  8. 5:09 Running frontier-class models like GLM 5.2 on local hardware
  9. 7:32 The case for sovereign AI: Owning your own compute stack
  10. 9:08 A retrospective on open-weight models: Mistral to Qwen
  11. 11:12 The evolution of reasoning: DeepSeek R1 and beyond
  12. 12:08 The rise of agentic performance and tool calling
  13. 15:33 Economic value: Does hardware appreciate as models become more efficient?
  14. 16:38 Closing thoughts: Why you should own your own GPU

Open at this second