read-only demo

Videos fWXJM-J0ZB8

Frontier results, on device - RL Nabors, Arize

index_state ready data_status ok

AI Engineer· published 2026-06-29· 0:30:51· en-US· indexed 2026-08-11 05:36

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:10, 1 of 1 keyframes kept
  2. Shot 1, 0:10 to 0:23, 1 of 1 keyframes kept
  3. Shot 2, 0:23 to 0:33, 1 of 1 keyframes kept
  4. Shot 3, 0:33 to 0:56, 1 of 1 keyframes kept
  5. Shot 4, 0:56 to 1:08, 1 of 1 keyframes kept
  6. Shot 5, 1:08 to 1:15, 1 of 1 keyframes kept
  7. Shot 6, 1:15 to 1:37, 1 of 1 keyframes kept
  8. Shot 7, 1:37 to 1:56, 1 of 1 keyframes kept
  9. Shot 8, 1:56 to 2:13, 1 of 1 keyframes kept
  10. Shot 9, 2:13 to 2:33, 1 of 1 keyframes kept
  11. Shot 10, 2:33 to 2:46, 1 of 1 keyframes kept
  12. Shot 11, 2:46 to 2:56, 1 of 1 keyframes kept
  13. Shot 12, 2:56 to 3:02, 1 of 1 keyframes kept
  14. Shot 13, 3:02 to 3:34, 1 of 1 keyframes kept
  15. Shot 14, 3:34 to 3:59, 1 of 1 keyframes kept
  16. Shot 15, 3:59 to 4:29, 1 of 1 keyframes kept
  17. Shot 16, 4:29 to 4:59, 0 of 1 keyframes kept
  18. Shot 17, 4:59 to 5:13, 1 of 1 keyframes kept
  19. Shot 18, 5:13 to 5:43, 1 of 1 keyframes kept
  20. Shot 19, 5:43 to 5:59, 1 of 1 keyframes kept
  21. Shot 20, 5:59 to 6:04, 1 of 1 keyframes kept
  22. Shot 21, 6:04 to 6:17, 1 of 1 keyframes kept
  23. Shot 22, 6:17 to 6:49, 1 of 1 keyframes kept
  24. Shot 23, 6:49 to 7:02, 1 of 1 keyframes kept
  25. Shot 24, 7:02 to 7:10, 1 of 1 keyframes kept
  26. Shot 25, 7:10 to 7:13, 1 of 1 keyframes kept
  27. Shot 26, 7:13 to 7:43, 1 of 1 keyframes kept
  28. Shot 27, 7:43 to 8:13, 1 of 1 keyframes kept
  29. Shot 28, 8:13 to 8:15, 1 of 1 keyframes kept
  30. Shot 29, 8:15 to 8:21, 1 of 1 keyframes kept
  31. Shot 30, 8:21 to 8:47, 1 of 1 keyframes kept
  32. Shot 31, 8:47 to 9:10, 1 of 1 keyframes kept
  33. Shot 32, 9:10 to 9:44, 1 of 1 keyframes kept
  34. Shot 33, 9:44 to 9:50, 1 of 1 keyframes kept
  35. Shot 34, 9:50 to 9:53, 0 of 1 keyframes kept
  36. Shot 35, 9:53 to 9:57, 1 of 1 keyframes kept
  37. Shot 36, 9:57 to 10:03, 1 of 1 keyframes kept
  38. Shot 37, 10:03 to 10:08, 1 of 1 keyframes kept
  39. Shot 38, 10:08 to 10:09, 1 of 1 keyframes kept
  40. Shot 39, 10:09 to 10:13, 1 of 1 keyframes kept
  41. Shot 40, 10:13 to 10:28, 1 of 1 keyframes kept
  42. Shot 41, 10:28 to 10:47, 1 of 1 keyframes kept
  43. Shot 42, 10:47 to 11:00, 1 of 1 keyframes kept
  44. Shot 43, 11:00 to 11:14, 1 of 1 keyframes kept
  45. Shot 44, 11:14 to 11:43, 1 of 1 keyframes kept
  46. Shot 45, 11:43 to 12:08, 1 of 1 keyframes kept
  47. Shot 46, 12:08 to 12:33, 0 of 1 keyframes kept
  48. Shot 47, 12:33 to 12:58, 0 of 1 keyframes kept
  49. Shot 48, 12:58 to 13:14, 0 of 1 keyframes kept
  50. Shot 49, 13:14 to 13:33, 1 of 1 keyframes kept
  51. Shot 50, 13:33 to 13:58, 1 of 1 keyframes kept
  52. Shot 51, 13:58 to 14:32, 1 of 1 keyframes kept
  53. Shot 52, 14:32 to 15:06, 1 of 1 keyframes kept
  54. Shot 53, 15:06 to 15:55, 1 of 1 keyframes kept
  55. Shot 54, 15:55 to 16:15, 1 of 1 keyframes kept
  56. Shot 55, 16:15 to 16:43, 1 of 1 keyframes kept
  57. Shot 56, 16:43 to 17:12, 0 of 1 keyframes kept
  58. Shot 57, 17:12 to 17:41, 0 of 1 keyframes kept
  59. Shot 58, 17:41 to 18:09, 1 of 1 keyframes kept
  60. Shot 59, 18:09 to 18:29, 1 of 1 keyframes kept
  61. Shot 60, 18:29 to 18:56, 1 of 1 keyframes kept
  62. Shot 61, 18:56 to 19:22, 0 of 1 keyframes kept
  63. Shot 62, 19:22 to 19:54, 1 of 1 keyframes kept
  64. Shot 63, 19:54 to 20:27, 0 of 1 keyframes kept
  65. Shot 64, 20:27 to 20:36, 0 of 1 keyframes kept
  66. Shot 65, 20:36 to 21:12, 1 of 1 keyframes kept
  67. Shot 66, 21:12 to 21:31, 1 of 1 keyframes kept
  68. Shot 67, 21:31 to 22:05, 1 of 1 keyframes kept
  69. Shot 68, 22:05 to 22:38, 0 of 1 keyframes kept
  70. Shot 69, 22:38 to 23:04, 1 of 1 keyframes kept
  71. Shot 70, 23:04 to 23:08, 1 of 1 keyframes kept
  72. Shot 71, 23:08 to 23:35, 0 of 1 keyframes kept
  73. Shot 72, 23:35 to 24:02, 0 of 1 keyframes kept
  74. Shot 73, 24:02 to 24:29, 0 of 1 keyframes kept
  75. Shot 74, 24:29 to 24:39, 1 of 1 keyframes kept
  76. Shot 75, 24:39 to 25:09, 1 of 1 keyframes kept
  77. Shot 76, 25:09 to 25:40, 0 of 1 keyframes kept
  78. Shot 77, 25:40 to 25:43, 1 of 1 keyframes kept
  79. Shot 78, 25:43 to 25:47, 0 of 1 keyframes kept
  80. Shot 79, 25:47 to 26:24, 1 of 1 keyframes kept
  81. Shot 80, 26:24 to 27:00, 1 of 1 keyframes kept
  82. Shot 81, 27:00 to 27:45, 1 of 1 keyframes kept
  83. Shot 82, 27:45 to 28:21, 1 of 1 keyframes kept
  84. Shot 83, 28:21 to 28:36, 1 of 1 keyframes kept
  85. Shot 84, 28:36 to 28:39, 1 of 1 keyframes kept
  86. Shot 85, 28:39 to 28:54, 1 of 1 keyframes kept
  87. Shot 86, 28:54 to 29:18, 1 of 1 keyframes kept
  88. Shot 87, 29:18 to 29:36, 0 of 1 keyframes kept
  89. Shot 88, 29:36 to 29:49, 1 of 1 keyframes kept
  90. Shot 89, 29:49 to 29:57, 0 of 1 keyframes kept
  91. Shot 90, 29:57 to 30:06, 1 of 1 keyframes kept
  92. Shot 91, 30:06 to 30:20, 1 of 1 keyframes kept
  93. Shot 92, 30:20 to 30:51, 1 of 1 keyframes kept

93 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
322
whisperx 322
chunks
54
from 322 cues
keyframes
75
kept of 93 captured
frames with text
75
1,354 lines read
chapters
0
from the source metadata
keyframe bytes
9.4 MB
word timings on 322 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:32 1m 07s
stt done 2026-08-11 05:33 33s
chunk done 2026-08-11 05:34 0s
text_embed done 2026-08-11 05:34 0s
keyframe done 2026-08-11 05:34 1m 38s
ocr done 2026-08-11 05:35 34s
frame_embed done 2026-08-11 05:36 12s

Frames, and what the machine read

  • 0:01 #0 done6 line(s)

    shot 0·sharpness 1425.7

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. STOPPAYINGFOR1.00
    3. FRONTIER MODELS0.98
    4. USE1.00
    5. EVALSAND1.00
    6. LOCAL MODELS0.97
  • 0:22 #1 done5 line(s)

    shot 1·sharpness 1375.1

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. MDN web docs0.96
    3. W3C0.98
    4. moz://a1.00
    5. O0.66
  • 0:30 #2 done4 line(s)

    shot 2·sharpness 820.7

    1. Rachel-Lee Nabors (they/them) nearestnabors.com0.99
    2. TinyFish1.00
    3. Google1.00
    4. Arcade1.00
  • 0:44 #3 done3 line(s)

    shot 3·sharpness 720.0

    1. Λ arize0.89
    2. You have agents.0.99
    3. We can test them.0.99
  • 1:05 #4 done5 line(s)

    shot 4·sharpness 1762.0

    1. 0.81
    2. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    3. Every time you reach for foundation0.99
    4. model like GPT5 or Claude,1.00
    5. it's costing you and your users.1.00
  • 1:10 #5 done3 line(s)

    shot 5·sharpness 1412.2

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. THE COST OF0.96
    3. ONE-SIZE-FITS-ALL INFERENCE1.00
  • 1:20 #6 done32 line(s)

    shot 6·sharpness 1251.0

    1. informa1.00
    2. TechTarget and Informa Tech's Digital Business Combine.0.99
    3. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    4. Dark Reading Resource Library0.99
    5. Black Hat News0.97
    6. Omdia Cybersecurity1.00
    7. Advertise1.00
    8. DARKREADING1.00
    9. NEWSLETTER SIGN-UP0.99
    10. Cybersecurity Topics0.97
    11. World1.00
    12. The Edge1.00
    13. DR Technology1.00
    14. Events1.00
    15. Resources1.00
    16. CYBER RISK1.00
    17. CYBERSECURITY OPERATIONS0.99
    18. DATA PRIVACY1.00
    19. REMOTE WORKFORCE0.97
    20. NEWS1.00
    21. Shadow Al, Data Exposure Plague Workplace Chatbot Use1.00
    22. Securitycostsusertrust.ke0.95
    23. Productivity has a down0.99
    24. he generational1.00
    25. Al platforms they use, v1.00
    26. Tara Seals, Managing Editor, News, Dark Reading1.00
    27. 6 Min Read0.97
    28. September 30,20240.99
    29. Editor's Choice1.00
    30. CYBERSECURITY OPERATIONS0.99
    31. Tara Seals, Da1.00
    32. Electronic Warfare Puts Commercial1.00
  • 1:50 #7 done6 line(s)

    shot 7·sharpness 1726.3

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. “[l]atency above 4 seconds degrades0.99
    3. quality of experience...1.00
    4. 990.99
    5. Mitigating Response Delays in Free-Form Conversations with LLM-0.99
    6. powered Intelligent Virtual Agents July 7, 20250.99
  • 2:10 #8 done2 line(s)

    shot 8·sharpness 753.8

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. API calls cost your business.1.00
  • 2:31 #9 done4 line(s)

    shot 9·sharpness 794.1

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. Couldn't connect to Claude0.99
    3. ERR_INTERNET_DISCONNECTED1.00
    4. Refresh1.00
  • 2:40 #10 done6 line(s)

    shot 10·sharpness 2379.8

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. Chief Product Officers should not1.00
    3. confuse the deflation of commodity0.99
    4. tokens with the democratization of1.00
    5. frontier reasoning.1.00
    6. – Gartner, March 20260.97
  • 2:55 #11 done5 line(s)

    shot 11·sharpness 1399.7

    1. Rachel-Lee Nabors (they/them) nearestnabors.com0.99
    2. Why spend1.00
    3. $10/day on API costs1.00
    4. when you could get the same1.00
    5. results forfree?1.00
  • 3:00 #12 done2 line(s)

    shot 12·sharpness 1225.2

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. DO YOU REALLY NEED AN LLM?0.99
  • 3:30 #13 done8 line(s)

    shot 13·sharpness 3152.0

    1. Rachel-Lee Nabors (they/them) nearestnabors.com0.99
    2. TASK-SPECIFIC MODELS CHEAT SHEET1.00
    3. · Is a camera pointing at something?0.98
    4. Use vision models like MobileNet, YOLO, MediaPipe1.00
    5. · Is a microphone recording?0.96
    6. Use audio models like Whisper, Wav2Vec20.99
    7. Chat, translation, analysis?1.00
    8. Use small language models like Gemma, Qwen1.00
  • 3:40 #14 done3 line(s)

    shot 14·sharpness 757.1

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. SLMs0.97
    3. LLMS BUT SMALL0.97
  • 4:06 #15 done5 line(s)

    shot 15·sharpness 2330.2

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. Small(er) Language Models (SLMs)1.00
    3. Smaller versions of LLMs containing several million0.99
    4. to several billion parameters (LLMs may have1.00
    5. hundreds of billions or even a trillion parameter0.99
  • 4:55 #16 skipped

    shot 16·duplicate of #15

  • 5:11 #17 done5 line(s)

    shot 17·sharpness 1772.7

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. “SLMs consume the same or less energy1.00
    3. as LLMs” to produce correct responses0.99
    4. Energy-Aware Code Generation with LLMs: Benchmarking Small vs.1.00
    5. Large Language Models for Sustainable AI Programming Aug 20250.99
  • 5:22 #18 done32 line(s)

    shot 18·sharpness 2810.1

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. SMALL LANGUAGE MODELS (SLMS)1.00
    3. Model1.00
    4. Parameters1.00
    5. Physical Size (FP16)1.00
    6. Phi-4 Mini1.00
    7. 3.8B1.00
    8. ~7.6 GB0.98
    9. Llama 3.2 (1B / 3B)0.96
    10. 1B/3B1.00
    11. ~2 GB/~6 GB0.97
    12. Ministral (3B /8B)0.98
    13. 3B/8B0.97
    14. ~6 GB /~16 GB0.97
    15. Gemma 4 E2B/E4B0.99
    16. 5B/9B0.99
    17. ~10 GB /~18 GB0.95
    18. Qwen 3 (0.6B / 1.7B / 4B/8B)0.94
    19. 0.6B/1.7B/4B/8B0.99
    20. ~1.2/ 3.4/ 8 /16 GB0.91
    21. Zephyr1.00
    22. 7B1.00
    23. ~14 GB0.98
    24. TinyLlama1.00
    25. 1.1B1.00
    26. ~2.2 GB1.00
    27. MiniCPM 40.95
    28. 0.5B/4B/8B1.00
    29. ~1 GB / ~8 GB /~16 GB0.92
    30. SmolLM31.00
    31. 3B1.00
    32. ~6 GB0.99
  • 5:51 #19 done1 line(s)

    shot 19·sharpness 567.1

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
  • 6:02 #20 done2 line(s)

    shot 20·sharpness 1250.6

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. SLMS ARE PRODUCTION READY1.00
  • 6:15 #21 done7 line(s)

    shot 21·sharpness 2747.6

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. “SLMs are sufficiently powerful, inherently1.00
    3. more suitable, and necessarily more1.00
    4. economical for many invocations in agentic1.00
    5. systems, and are therefore the future of1.00
    6. agentic AI."0.94
    7. Small Language Models are the Future of Agentic AI—NVIDIA, 20250.98
  • 6:42 #22 done9 line(s)

    shot 22·sharpness 1425.8

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. Energy consumption comparison1.00
    3. LLMs0.99
    4. SLMs1.00
    5. Task-specific Models1.00
    6. 251.00
    7. 501.00
    8. 751.00
    9. proportional energy consumed1.00
  • 7:00 #23 done7 line(s)

    shot 23·sharpness 1511.4

    1. Rachel-Lee Nabors (they/them) nearestnabors.com1.00
    2. BENEFITS OF SMALL AND LOCAL AI0.97
    3. · More secure0.95
    4. · Works offline0.95
    5. ·No fees0.98
    6. · More efficient0.98
    7. · Lowerlatency0.98

Transcript

322 cues· 4,766 words· 25,840 chars

  1. 0:01 Hi there, I'm Rachel Leigh Neighbors, and today I'm here to talk with you about how to use local models to stop paying for frontier models.
  2. 0:09 Let's dig into it.
  3. 0:11 So I've worked on standards that power today's web with Mozilla on Firefox dev tools and the W3C on web standards, and of course on Microsoft's Edge browser.
  4. 0:21 I've even been on the React team.
  5. 0:23 Now, I've spent the past three years consulting with AI startups and some of our favorite LLM and browser companies on all things web, AI, and UI.
  6. 0:33 And recently I've joined Arise.
  7. 0:36 Have you ever had your CTO ruin your agentic workflow with a slight change of prompt or an LLM migration?
  8. 0:42 Have you ever been that CTO?
  9. 0:44 Well, you probably need Arise's observability platform for models and the agents who love them.
  10. 0:50 Anyway, I'll be actually using one of Arise's open source projects today, Phoenix.
  11. 0:55 We'll talk more about that later.
  12. 0:57 But today, specifically, I'm here to talk with you about how AI is costing you.
  13. 1:02 Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.
  14. 1:10 Let's have a look at the costs of one-size-fits-all inference.
  15. 1:17 So first there's security.
  16. 1:19 Security costs trust.
  17. 1:20 When you use a large LLM that's in the cloud, you're sending data to remote servers, and it always carries the risk of exposure, interception, and retention by third parties.
  18. 1:30 We have cases where the use of remote AI chatbots has led to sensitive business data being stored, breached, or leaked to the public.
  19. 1:38 latency costs the user experience.
  20. 1:40 Now there's been research on mitigating response delays and LLM chats in VR and found that four seconds is the limit of believability for users.
  21. 1:48 And many calls that you will make to large models are going to take longer than four seconds, as we will see when we check Phoenix.
  22. 1:57 And it costs your business.
  23. 1:59 Third-party inference costs are uncontrollable compared to API costs.
  24. 2:05 Agency compounding levels of inference, and this means even if the tokens are cheaper, you may be using more of them, or if you're using four of them, they may be more expensive.
  25. 2:14 And of course, if you aren't connected, remote models simply aren't going to work, which means that unless your software is connected to the web, nobody can use it.
  26. 2:25 Big inference going offline costs productivity.
  27. 2:28 If there's an outage, if you're in a place where you cannot reach Wi-Fi or you're in a very secure environment.
  28. 2:35 Now, token costs have been falling as of late, but total inference spend has been rising because agentic and reasoning workloads consume tokens way faster than prices are dropping.
  29. 2:48 but we can completely eliminate most of these costs.
  30. 2:51 And it starts by asking ourselves exactly how much is this costing?
  31. 2:57 Do we really need an LLM to do this job?
  32. 3:03 All right, you can use task-specific models, which are small in size and power consumption compared to an AI foundation.
  33. 3:10 This is my little cheat sheet.
  34. 3:12 If you're looking for an expert model, is a camera pointing at something?
  35. 3:16 Division, you know, is there a vision component?
  36. 3:20 You can use a model like MobileNet, YOLO, MediaPipe.
  37. 3:23 Is a microphone recording something?
  38. 3:25 There are audio models like Whisper and Wave 2 Vectu.
  39. 3:29 Chat translation or analysis.
  40. 3:30 This is where he might use something like a small language model like Gemma or Quinn.
  41. 3:36 These small models are called SLMs or smaller language models.
  42. 3:43 The definition of small is up to debate here.
  43. 3:47 These are great for times when we do need the language power of a generative pre-trained transformer, a GPT, but we probably don't need the sum total of human knowledge in a black box at our disposal, or we don't need multimodal capabilities.
  44. 4:01 We're not going to be analyzing images and audio at the same time.
  45. 4:05 Now, SLMs, smaller language models, contain millions to billions of parameters.
  46. 4:11 An LLM contains billions, well, trillions.
  47. 4:16 So you see the big dot there?
  48. 4:17 That's one of the smaller LLMs that's out there.
  49. 4:20 But the little dot actually represents one of the larger SLMs.
  50. 4:23 So you can see that there is a huge difference in parameter size, and this means that there is a vast difference in the size of machine you're gonna need to run one of these models.

Open at this second