Videos 65X0pQ6Lmbg
Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
Scene timeline
33 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 98
- whisperx 98
- chunks
- 22
- from 98 cues
- keyframes
- 27
- kept of 33 captured
- frames with text
- 27
- 231 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.5 MB
- word timings on 98 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:54 | 1m 09s |
stt |
done | — | 2026-08-11 05:55 | 13s |
chunk |
done | — | 2026-08-11 05:55 | 0s |
text_embed |
done | — | 2026-08-11 05:55 | 1s |
keyframe |
done | — | 2026-08-11 05:55 | 39s |
ocr |
done | — | 2026-08-11 05:56 | 5s |
frame_embed |
done | — | 2026-08-11 05:56 | 5s |
Frames, and what the machine read
-
- AIE1.00
- Voice In, Visuals Out1.00
- The Agony and the Ecstasy0.98
- Allen Pike, Forestwalk Labs0.99
- Al Engineering WF, 20260.97
-
- RODE1.00
- “Audio is the human-preferred0.99
- AIE1.00
- input to AIs, but vision is the0.97
- preferred output from them."0.99
- - Andrej Karpathy, May 20260.98
-
- Current date is Tue1.00
- 1-01-19801.00
- Enter new date:1.00
- Current time is 21:35:24.181.00
- AIE1.00
- Enter new time:1.00
- The IBM Personal Computer DOS1.00
- Version 2.00 (C)Copyright IBM Corp 1981, 1982, 19830.99
- A>_0.80
-
- AIE1.00
- Visuals Out1.00
-
- AIE1.00
- Lorem ipsum1.00
- Lorem ipsum0.99
- Lorem ipsum1.00
- Visuals Out1.00
- A B C D E0.92
- Lorem ipsum dolor0.98
- B1.00
- Lorem ipsum0.99
- Loremipsum1.00
- Lorem ipsum0.99
-
- Output1.00
- Probabilities1.00
- Softmax1.00
- Add & Norm0.98
- AIE1.00
- Forward1.00
- Feed1.00
- Add & Norm0.99
- Add & Norm0.99
- Multi-Head0.96
- Forward1.00
- Feed1.00
- Attention1.00
- N×0.95
- N×0.79
- Add & Norm0.99
- Multi-Head1.00
- Attention1.00
- Add & Norm0.99
- Multi-Head1.00
- Attention1.00
- Masked1.00
- Positional1.00
- Encoding1.00
- Positional1.00
- Encoding1.00
- Input1.00
- Output1.00
- Embedding1.00
- Embedding1.00
- Visualization1.00
- Inputs1.00
- Figure 1: The Transformer - model architecture.0.99
-
- Output1.00
- THEME1.00
- Probabilities0.99
- Softmax1.00
- DARK1.00
- LIGHT1.00
- dit1.00
- Add & Norm0.96
- BREAKPOINT1.00
- AIE1.00
- Forward1.00
- Feed1.00
- XPLORE0.99
- DESKTOP1.00
- TABLET1.00
- MOBILE1.00
- Add &Norm0.99
- Add & Norm0.98
- Multi-Head0.97
- Forward1.00
- Feed1.00
- Attention1.00
- N×0.88
- NETWORK1.00
- N×0.77
- Add & Norm0.98
- Multi-Head1.00
- Attention1.00
- Add & Norm0.99
- Multi-Head1.00
- Masked1.00
- Attention1.00
- Arc width1.00
- Arc color1.00
- 0.61.00
- Arc glow1.00
- 131.00
- Encoding1.00
- Positional1.00
- Positional1.00
- Encoding1.00
- Arc density1.00
- 100%1.00
- Embedding1.00
- Input1.00
- Embedding1.00
- Output1.00
- City size1.00
- Visualization1.00
- Inputs1.00
- Outputs0.99
- Interactivity0.97
- Pulse speed1.00
- 3.4s0.91
- Beauty1.00
- Figure 1: The Transformer - model architecture.0.99
- Show cities1.00
-
- AIE1.00
- Voice In0.94
-
- AIE1.00
- COJO0.60
- Voice In1.00
-
- How Al reacts to OpenAl0.98
- CEO reacting to my video0.99
- How Al reacts to my0.97
- ugly face filter0.99
- How Alreacts to me0.97
- getting pulled over0.96
- simple instructions1.00
- How Al reacts to0.98
- Al tells me there's 10.95
- Ein seventeen0.98
- AIE1.00
- What's going on...0.96
- Idk what to type here rn1.00
- Hopefully Al makes a1.00
- It did too well..0.93
- Just learned this1.00
- 6M views0.99
- 5.3M views1.00
- better lawyer /Twitch ...0.98
- 3M views0.99
- 1.9M views0.99
- 4.3M views1.00
-
- AIE1.00
-
- "okay"0.99
- "okay"0.99
- AIE1.00
-
- I saw a weird thing where the Slack integration1.00
- missed a response in a thread, I think.0.98
- AIE1.00
-
- I saw a weird thing where the Slack integration0.98
- missed a response in a thread, I think.1.00
- AIE1.00
- Oh yeah, I've seen a problem with threading too.0.99
- Okay, let's file that in Linear.0.98
-
- I saw a weird thing where the Slack integration0.99
- missed a response in a thread, I think.0.99
- AIE1.00
- Oh yeah, I've seen a problem with threading too.0.99
- Okay, let's file that in Linear.1.00
- Filed: CED-28170.98
- Slack integration drops threaded replies1.00
-
- The Tyranny1.00
- AIE1.00
- of Latency1.00
-
- AIE1.00
- 1000ms·seamless visuals0.99
- 100ms·feelsinstant0.99
-
- 1000ms·seamless visuals0.99
- AIE1.00
- 100ms·feelsinstant0.99
-
- AIE1.00
- 200ms·seamless voice0.99
- STT1.00
- first token0.99
- 100ms·feelsinstant0.99
-
- AIE1.00
- 1000ms·seamless visuals0.99
- STT1.00
- first token1.00
- variabiity1.00
- 100ms·feelsinstant0.99
-
- AIE1.00
- 1000ms·seamless visuals0.99
- STT1.00
- first token0.99
- variability1.00
- 100ms·feelsinstant0.98
-
- Pillars of low latency1.00
- AIE1.00
- 11.00
- Fast models1.00
Transcript
98 cues· 1,853 words· 10,103 chars
- 0:01 All right, I'm Alan Pike, and today I'm going to be sharing some of what we've learned building voice in, visuals out, experiences using AI.
- 0:14 Here we go.
- 0:16 So this is Andrej Karpathy, and he made an argument last month that voice is the human preferred input
- 0:29 for AIs, but that we prefer visuals as the output.
- 0:36 And he knows a thing or two about this stuff, but that's not how we have been, for the most part, building or using AI.
- 0:49 We've been typing to it, it's been typing back maybe with some markdown.
- 0:55 But there's some breakthroughs over the last few months
- 0:58 where both this visuals in and audio, or the audio in and visuals out experience is now really feasible and we can create really delightful experiences with it.
- 1:12 The visuals outside is pretty intuitive, right?
- 1:16 Of course, a third of our brain is dedicated to processing visual information.
- 1:21 We love looking at things.
- 1:24 And models have recently got to the point where they can generate rich HTML, tool calling.
- 1:30 So we can have these experiences where there's visualizations that come back, explain things, help us understand, communicate the responses from these models.
- 1:42 They can give us interactive controls that allow us to explore and understand and modify and change and direct the models.
- 1:54 And they can even respond with beautiful illustrations and images.
- 2:01 And so the ceiling on the visuals out piece has really lifted, and we have a lot more capability of what we can do in terms of responding from these models.
- 2:15 The more controversial half of Karpathy's argument, though, is this idea of voice as the preferred input.
- 2:23 We've long sort of idealized and fantasized about speaking to an AI and having this real-time conversation where it understands what we need and it reacts appropriately in real time.
- 2:39 But the experiences that most people have had so far with voice interfaces have been more like trying to get Siri to turn the lights on and it's not working, or like this guy that is like trying to get ChatGPT voice mode to do things, but it keeps being like awkward and confused, right?
- 3:01 The models that we have so far, the experiences that most people have seen, are both slow and dumb, which is not a great combination.
- 3:14 And so a lot of people are down on voices and input.
- 3:18 But speaking is the ultimate way that as humans we communicate.
- 3:24 When we're speaking, we have more words per minute than when we're typing.
- 3:31 but we also convey more with each word.
- 3:35 There's a huge difference in between if you say to me, okay, versus if you say, okay.
- 3:46 And that's why when we have something that is really important that we need to communicate or get through to somebody, we will jump on a call.
- 3:59 and we will, or we'll speak in person so that we can get through that high bandwidth communication.
- 4:09 And so what we've done at ForestWalk is we have built an agent that is in our calls that can help us out in real time.
- 4:18 And so the other day I mentioned to my co-founder on a call that I'd seen a bug with the Slack integration, or at least I thought I had.
- 4:29 And she responded that she had seen the same bug.
- 4:34 And so I just said, okay, well, let's file that as a linear issue.
- 4:40 And the voice agent within a second responded that it had done so.
- 4:46 And that feels perfectly natural when you get it really dialed in that you are speaking, either incidentally or with a purpose, intentionally speaking to the AI.
- 4:59 And it responds in a way that is not interruptive of you.
- 5:04 It doesn't need to be voice.
- 5:06 It is just taking action on your intent.
- 5:10 And so we're going to see this experience more and more in the coming months and years.
- 5:16 But there's a huge barrier, a huge challenge to making this actually work and feel good.
- 5:22 And that is the tyranny of latency.
- 5:29 It's really difficult to get a response through that whole chain fast enough.
- 5:35 We've known since the 60s that to have a computer react to us fast enough that it feels
- 5:44 Instead, it needs to react in about 100 milliseconds, a tenth of a second.
- 5:51 And so, of course, we are always aspiring to get our products to react that fast.
- 5:57 And sometimes we can achieve that.
- 5:58 But with networking and everything, it can be challenging, depending on what work needs to get done.
- 6:04 And so sometimes we might flex up to 1,000 milliseconds.
- 6:08 You get a full second.
- 6:09 It's about the limit before people start to lose their train of thought.
- 6:12 They ask people to do something.
- 6:14 It takes more than a second.
loading