Videos fWXJM-J0ZB8
Frontier results, on device - RL Nabors, Arize
Scene timeline
93 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 322
- whisperx 322
- chunks
- 54
- from 322 cues
- keyframes
- 75
- kept of 93 captured
- frames with text
- 75
- 1,354 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 9.4 MB
- word timings on 322 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:32 | 1m 07s |
stt |
done | — | 2026-08-11 05:33 | 33s |
chunk |
done | — | 2026-08-11 05:34 | 0s |
text_embed |
done | — | 2026-08-11 05:34 | 0s |
keyframe |
done | — | 2026-08-11 05:34 | 1m 38s |
ocr |
done | — | 2026-08-11 05:35 | 34s |
frame_embed |
done | — | 2026-08-11 05:36 | 12s |
Frames, and what the machine read
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- STOPPAYINGFOR1.00
- FRONTIER MODELS0.98
- USE1.00
- EVALSAND1.00
- LOCAL MODELS0.97
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- MDN web docs0.96
- W3C0.98
- moz://a1.00
- O0.66
-
- Rachel-Lee Nabors (they/them) nearestnabors.com0.99
- TinyFish1.00
- Google1.00
- Arcade1.00
-
- Λ arize0.89
- You have agents.0.99
- We can test them.0.99
-
- 木0.81
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Every time you reach for foundation0.99
- model like GPT5 or Claude,1.00
- it's costing you and your users.1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- THE COST OF0.96
- ONE-SIZE-FITS-ALL INFERENCE1.00
-
- informa1.00
- TechTarget and Informa Tech's Digital Business Combine.0.99
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Dark Reading Resource Library0.99
- Black Hat News0.97
- Omdia Cybersecurity1.00
- Advertise1.00
- DARKREADING1.00
- NEWSLETTER SIGN-UP0.99
- Cybersecurity Topics0.97
- World1.00
- The Edge1.00
- DR Technology1.00
- Events1.00
- Resources1.00
- CYBER RISK1.00
- CYBERSECURITY OPERATIONS0.99
- DATA PRIVACY1.00
- REMOTE WORKFORCE0.97
- NEWS1.00
- Shadow Al, Data Exposure Plague Workplace Chatbot Use1.00
- Securitycostsusertrust.ke0.95
- Productivity has a down0.99
- he generational1.00
- Al platforms they use, v1.00
- Tara Seals, Managing Editor, News, Dark Reading1.00
- 6 Min Read0.97
- September 30,20240.99
- Editor's Choice1.00
- CYBERSECURITY OPERATIONS0.99
- Tara Seals, Da1.00
- Electronic Warfare Puts Commercial1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- “[l]atency above 4 seconds degrades0.99
- quality of experience...1.00
- 990.99
- Mitigating Response Delays in Free-Form Conversations with LLM-0.99
- powered Intelligent Virtual Agents July 7, 20250.99
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- API calls cost your business.1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Couldn't connect to Claude0.99
- ERR_INTERNET_DISCONNECTED1.00
- Refresh1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Chief Product Officers should not1.00
- confuse the deflation of commodity0.99
- tokens with the democratization of1.00
- frontier reasoning.1.00
- – Gartner, March 20260.97
-
- Rachel-Lee Nabors (they/them) nearestnabors.com0.99
- Why spend1.00
- $10/day on API costs1.00
- when you could get the same1.00
- results forfree?1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- DO YOU REALLY NEED AN LLM?0.99
-
- Rachel-Lee Nabors (they/them) nearestnabors.com0.99
- TASK-SPECIFIC MODELS CHEAT SHEET1.00
- · Is a camera pointing at something?0.98
- Use vision models like MobileNet, YOLO, MediaPipe1.00
- · Is a microphone recording?0.96
- Use audio models like Whisper, Wav2Vec20.99
- Chat, translation, analysis?1.00
- Use small language models like Gemma, Qwen1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- SLMs0.97
- LLMS BUT SMALL0.97
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Small(er) Language Models (SLMs)1.00
- Smaller versions of LLMs containing several million0.99
- to several billion parameters (LLMs may have1.00
- hundreds of billions or even a trillion parameter0.99
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- “SLMs consume the same or less energy1.00
- as LLMs” to produce correct responses0.99
- Energy-Aware Code Generation with LLMs: Benchmarking Small vs.1.00
- Large Language Models for Sustainable AI Programming Aug 20250.99
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- SMALL LANGUAGE MODELS (SLMS)1.00
- Model1.00
- Parameters1.00
- Physical Size (FP16)1.00
- Phi-4 Mini1.00
- 3.8B1.00
- ~7.6 GB0.98
- Llama 3.2 (1B / 3B)0.96
- 1B/3B1.00
- ~2 GB/~6 GB0.97
- Ministral (3B /8B)0.98
- 3B/8B0.97
- ~6 GB /~16 GB0.97
- Gemma 4 E2B/E4B0.99
- 5B/9B0.99
- ~10 GB /~18 GB0.95
- Qwen 3 (0.6B / 1.7B / 4B/8B)0.94
- 0.6B/1.7B/4B/8B0.99
- ~1.2/ 3.4/ 8 /16 GB0.91
- Zephyr1.00
- 7B1.00
- ~14 GB0.98
- TinyLlama1.00
- 1.1B1.00
- ~2.2 GB1.00
- MiniCPM 40.95
- 0.5B/4B/8B1.00
- ~1 GB / ~8 GB /~16 GB0.92
- SmolLM31.00
- 3B1.00
- ~6 GB0.99
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- SLMS ARE PRODUCTION READY1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- “SLMs are sufficiently powerful, inherently1.00
- more suitable, and necessarily more1.00
- economical for many invocations in agentic1.00
- systems, and are therefore the future of1.00
- agentic AI."0.94
- Small Language Models are the Future of Agentic AI—NVIDIA, 20250.98
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- Energy consumption comparison1.00
- LLMs0.99
- SLMs1.00
- Task-specific Models1.00
- 251.00
- 501.00
- 751.00
- proportional energy consumed1.00
-
- Rachel-Lee Nabors (they/them) nearestnabors.com1.00
- BENEFITS OF SMALL AND LOCAL AI0.97
- · More secure0.95
- · Works offline0.95
- ·No fees0.98
- · More efficient0.98
- · Lowerlatency0.98
Transcript
322 cues· 4,766 words· 25,840 chars
- 0:01 Hi there, I'm Rachel Leigh Neighbors, and today I'm here to talk with you about how to use local models to stop paying for frontier models.
- 0:09 Let's dig into it.
- 0:11 So I've worked on standards that power today's web with Mozilla on Firefox dev tools and the W3C on web standards, and of course on Microsoft's Edge browser.
- 0:21 I've even been on the React team.
- 0:23 Now, I've spent the past three years consulting with AI startups and some of our favorite LLM and browser companies on all things web, AI, and UI.
- 0:33 And recently I've joined Arise.
- 0:36 Have you ever had your CTO ruin your agentic workflow with a slight change of prompt or an LLM migration?
- 0:42 Have you ever been that CTO?
- 0:44 Well, you probably need Arise's observability platform for models and the agents who love them.
- 0:50 Anyway, I'll be actually using one of Arise's open source projects today, Phoenix.
- 0:55 We'll talk more about that later.
- 0:57 But today, specifically, I'm here to talk with you about how AI is costing you.
- 1:02 Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.
- 1:10 Let's have a look at the costs of one-size-fits-all inference.
- 1:17 So first there's security.
- 1:19 Security costs trust.
- 1:20 When you use a large LLM that's in the cloud, you're sending data to remote servers, and it always carries the risk of exposure, interception, and retention by third parties.
- 1:30 We have cases where the use of remote AI chatbots has led to sensitive business data being stored, breached, or leaked to the public.
- 1:38 latency costs the user experience.
- 1:40 Now there's been research on mitigating response delays and LLM chats in VR and found that four seconds is the limit of believability for users.
- 1:48 And many calls that you will make to large models are going to take longer than four seconds, as we will see when we check Phoenix.
- 1:57 And it costs your business.
- 1:59 Third-party inference costs are uncontrollable compared to API costs.
- 2:05 Agency compounding levels of inference, and this means even if the tokens are cheaper, you may be using more of them, or if you're using four of them, they may be more expensive.
- 2:14 And of course, if you aren't connected, remote models simply aren't going to work, which means that unless your software is connected to the web, nobody can use it.
- 2:25 Big inference going offline costs productivity.
- 2:28 If there's an outage, if you're in a place where you cannot reach Wi-Fi or you're in a very secure environment.
- 2:35 Now, token costs have been falling as of late, but total inference spend has been rising because agentic and reasoning workloads consume tokens way faster than prices are dropping.
- 2:48 but we can completely eliminate most of these costs.
- 2:51 And it starts by asking ourselves exactly how much is this costing?
- 2:57 Do we really need an LLM to do this job?
- 3:03 All right, you can use task-specific models, which are small in size and power consumption compared to an AI foundation.
- 3:10 This is my little cheat sheet.
- 3:12 If you're looking for an expert model, is a camera pointing at something?
- 3:16 Division, you know, is there a vision component?
- 3:20 You can use a model like MobileNet, YOLO, MediaPipe.
- 3:23 Is a microphone recording something?
- 3:25 There are audio models like Whisper and Wave 2 Vectu.
- 3:29 Chat translation or analysis.
- 3:30 This is where he might use something like a small language model like Gemma or Quinn.
- 3:36 These small models are called SLMs or smaller language models.
- 3:43 The definition of small is up to debate here.
- 3:47 These are great for times when we do need the language power of a generative pre-trained transformer, a GPT, but we probably don't need the sum total of human knowledge in a black box at our disposal, or we don't need multimodal capabilities.
- 4:01 We're not going to be analyzing images and audio at the same time.
- 4:05 Now, SLMs, smaller language models, contain millions to billions of parameters.
- 4:11 An LLM contains billions, well, trillions.
- 4:16 So you see the big dot there?
- 4:17 That's one of the smaller LLMs that's out there.
- 4:20 But the little dot actually represents one of the larger SLMs.
- 4:23 So you can see that there is a huge difference in parameter size, and this means that there is a vast difference in the size of machine you're gonna need to run one of these models.
loading