read-only demo

Videos 65X0pQ6Lmbg

Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:13:05· en-US· indexed 2026-08-11 05:56

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:16, 1 of 1 keyframes kept
  2. Shot 1, 0:16 to 0:48, 1 of 1 keyframes kept
  3. Shot 2, 0:48 to 1:10, 1 of 1 keyframes kept
  4. Shot 3, 1:10 to 1:20, 1 of 1 keyframes kept
  5. Shot 4, 1:20 to 1:32, 1 of 1 keyframes kept
  6. Shot 5, 1:32 to 1:42, 1 of 1 keyframes kept
  7. Shot 6, 1:42 to 2:19, 1 of 1 keyframes kept
  8. Shot 7, 2:19 to 2:28, 1 of 1 keyframes kept
  9. Shot 8, 2:28 to 2:50, 1 of 1 keyframes kept
  10. Shot 9, 2:50 to 3:19, 1 of 1 keyframes kept
  11. Shot 10, 3:19 to 3:38, 1 of 1 keyframes kept
  12. Shot 11, 3:38 to 4:04, 1 of 1 keyframes kept
  13. Shot 12, 4:04 to 4:30, 1 of 1 keyframes kept
  14. Shot 13, 4:30 to 4:56, 1 of 1 keyframes kept
  15. Shot 14, 4:56 to 5:21, 1 of 1 keyframes kept
  16. Shot 15, 5:21 to 5:47, 1 of 1 keyframes kept
  17. Shot 16, 5:47 to 6:13, 1 of 1 keyframes kept
  18. Shot 17, 6:13 to 6:38, 1 of 1 keyframes kept
  19. Shot 18, 6:38 to 7:04, 0 of 1 keyframes kept
  20. Shot 19, 7:04 to 7:30, 1 of 1 keyframes kept
  21. Shot 20, 7:30 to 7:56, 0 of 1 keyframes kept
  22. Shot 21, 7:56 to 8:21, 1 of 1 keyframes kept
  23. Shot 22, 8:21 to 8:47, 1 of 1 keyframes kept
  24. Shot 23, 8:47 to 9:13, 1 of 1 keyframes kept
  25. Shot 24, 9:13 to 9:39, 1 of 1 keyframes kept
  26. Shot 25, 9:39 to 10:04, 1 of 1 keyframes kept
  27. Shot 26, 10:04 to 10:30, 0 of 1 keyframes kept
  28. Shot 27, 10:30 to 10:56, 0 of 1 keyframes kept
  29. Shot 28, 10:56 to 11:21, 0 of 1 keyframes kept
  30. Shot 29, 11:21 to 11:47, 1 of 1 keyframes kept
  31. Shot 30, 11:47 to 12:13, 1 of 1 keyframes kept
  32. Shot 31, 12:13 to 12:39, 1 of 1 keyframes kept
  33. Shot 32, 12:39 to 13:04, 0 of 1 keyframes kept

33 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
98
whisperx 98
chunks
22
from 98 cues
keyframes
27
kept of 33 captured
frames with text
27
231 lines read
chapters
0
from the source metadata
keyframe bytes
2.5 MB
word timings on 98 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:54 1m 09s
stt done 2026-08-11 05:55 13s
chunk done 2026-08-11 05:55 0s
text_embed done 2026-08-11 05:55 1s
keyframe done 2026-08-11 05:55 39s
ocr done 2026-08-11 05:56 5s
frame_embed done 2026-08-11 05:56 5s

Frames, and what the machine read

  • 0:01 #0 done5 line(s)

    shot 0·sharpness 965.6

    1. AIE1.00
    2. Voice In, Visuals Out1.00
    3. The Agony and the Ecstasy0.98
    4. Allen Pike, Forestwalk Labs0.99
    5. Al Engineering WF, 20260.97
  • 0:41 #1 done6 line(s)

    shot 1·sharpness 1747.5

    1. RODE1.00
    2. “Audio is the human-preferred0.99
    3. AIE1.00
    4. input to AIs, but vision is the0.97
    5. preferred output from them."0.99
    6. - Andrej Karpathy, May 20260.98
  • 1:02 #2 done9 line(s)

    shot 2·sharpness 540.3

    1. Current date is Tue1.00
    2. 1-01-19801.00
    3. Enter new date:1.00
    4. Current time is 21:35:24.181.00
    5. AIE1.00
    6. Enter new time:1.00
    7. The IBM Personal Computer DOS1.00
    8. Version 2.00 (C)Copyright IBM Corp 1981, 1982, 19830.99
    9. A>_0.80
  • 1:16 #3 done2 line(s)

    shot 3·sharpness 517.0

    1. AIE1.00
    2. Visuals Out1.00
  • 1:29 #4 done11 line(s)

    shot 4·sharpness 814.2

    1. AIE1.00
    2. Lorem ipsum1.00
    3. Lorem ipsum0.99
    4. Lorem ipsum1.00
    5. Visuals Out1.00
    6. A B C D E0.92
    7. Lorem ipsum dolor0.98
    8. B1.00
    9. Lorem ipsum0.99
    10. Loremipsum1.00
    11. Lorem ipsum0.99
  • 1:33 #5 done33 line(s)

    shot 5·sharpness 1359.5

    1. Output1.00
    2. Probabilities1.00
    3. Softmax1.00
    4. Add & Norm0.98
    5. AIE1.00
    6. Forward1.00
    7. Feed1.00
    8. Add & Norm0.99
    9. Add & Norm0.99
    10. Multi-Head0.96
    11. Forward1.00
    12. Feed1.00
    13. Attention1.00
    14. 0.95
    15. 0.79
    16. Add & Norm0.99
    17. Multi-Head1.00
    18. Attention1.00
    19. Add & Norm0.99
    20. Multi-Head1.00
    21. Attention1.00
    22. Masked1.00
    23. Positional1.00
    24. Encoding1.00
    25. Positional1.00
    26. Encoding1.00
    27. Input1.00
    28. Output1.00
    29. Embedding1.00
    30. Embedding1.00
    31. Visualization1.00
    32. Inputs1.00
    33. Figure 1: The Transformer - model architecture.0.99
  • 2:15 #6 done57 line(s)

    shot 6·sharpness 2272.8

    1. Output1.00
    2. THEME1.00
    3. Probabilities0.99
    4. Softmax1.00
    5. DARK1.00
    6. LIGHT1.00
    7. dit1.00
    8. Add & Norm0.96
    9. BREAKPOINT1.00
    10. AIE1.00
    11. Forward1.00
    12. Feed1.00
    13. XPLORE0.99
    14. DESKTOP1.00
    15. TABLET1.00
    16. MOBILE1.00
    17. Add &Norm0.99
    18. Add & Norm0.98
    19. Multi-Head0.97
    20. Forward1.00
    21. Feed1.00
    22. Attention1.00
    23. 0.88
    24. NETWORK1.00
    25. 0.77
    26. Add & Norm0.98
    27. Multi-Head1.00
    28. Attention1.00
    29. Add & Norm0.99
    30. Multi-Head1.00
    31. Masked1.00
    32. Attention1.00
    33. Arc width1.00
    34. Arc color1.00
    35. 0.61.00
    36. Arc glow1.00
    37. 131.00
    38. Encoding1.00
    39. Positional1.00
    40. Positional1.00
    41. Encoding1.00
    42. Arc density1.00
    43. 100%1.00
    44. Embedding1.00
    45. Input1.00
    46. Embedding1.00
    47. Output1.00
    48. City size1.00
    49. Visualization1.00
    50. Inputs1.00
    51. Outputs0.99
    52. Interactivity0.97
    53. Pulse speed1.00
    54. 3.4s0.91
    55. Beauty1.00
    56. Figure 1: The Transformer - model architecture.0.99
    57. Show cities1.00
  • 2:25 #7 done2 line(s)

    shot 7·sharpness 421.9

    1. AIE1.00
    2. Voice In0.94
  • 2:47 #8 done3 line(s)

    shot 8·sharpness 376.8

    1. AIE1.00
    2. COJO0.60
    3. Voice In1.00
  • 3:10 #9 done22 line(s)

    shot 9·sharpness 1845.2

    1. How Al reacts to OpenAl0.98
    2. CEO reacting to my video0.99
    3. How Al reacts to my0.97
    4. ugly face filter0.99
    5. How Alreacts to me0.97
    6. getting pulled over0.96
    7. simple instructions1.00
    8. How Al reacts to0.98
    9. Al tells me there's 10.95
    10. Ein seventeen0.98
    11. AIE1.00
    12. What's going on...0.96
    13. Idk what to type here rn1.00
    14. Hopefully Al makes a1.00
    15. It did too well..0.93
    16. Just learned this1.00
    17. 6M views0.99
    18. 5.3M views1.00
    19. better lawyer /Twitch ...0.98
    20. 3M views0.99
    21. 1.9M views0.99
    22. 4.3M views1.00
  • 3:23 #10 done1 line(s)

    shot 10·sharpness 909.2

    1. AIE1.00
  • 3:46 #11 done3 line(s)

    shot 11·sharpness 512.4

    1. "okay"0.99
    2. "okay"0.99
    3. AIE1.00
  • 4:27 #12 done3 line(s)

    shot 12·sharpness 919.9

    1. I saw a weird thing where the Slack integration1.00
    2. missed a response in a thread, I think.0.98
    3. AIE1.00
  • 4:40 #13 done5 line(s)

    shot 13·sharpness 1572.1

    1. I saw a weird thing where the Slack integration0.98
    2. missed a response in a thread, I think.1.00
    3. AIE1.00
    4. Oh yeah, I've seen a problem with threading too.0.99
    5. Okay, let's file that in Linear.0.98
  • 5:08 #14 done7 line(s)

    shot 14·sharpness 810.0

    1. I saw a weird thing where the Slack integration0.99
    2. missed a response in a thread, I think.0.99
    3. AIE1.00
    4. Oh yeah, I've seen a problem with threading too.0.99
    5. Okay, let's file that in Linear.1.00
    6. Filed: CED-28170.98
    7. Slack integration drops threaded replies1.00
  • 5:29 #15 done3 line(s)

    shot 15·sharpness 449.7

    1. The Tyranny1.00
    2. AIE1.00
    3. of Latency1.00
  • 6:10 #16 done3 line(s)

    shot 16·sharpness 277.8

    1. AIE1.00
    2. 1000ms·seamless visuals0.99
    3. 100ms·feelsinstant0.99
  • 6:23 #17 done3 line(s)

    shot 17·sharpness 283.6

    1. 1000ms·seamless visuals0.99
    2. AIE1.00
    3. 100ms·feelsinstant0.99
  • 6:46 #18 skipped

    shot 18·duplicate of #17

  • 7:17 #19 done5 line(s)

    shot 19·sharpness 306.0

    1. AIE1.00
    2. 200ms·seamless voice0.99
    3. STT1.00
    4. first token0.99
    5. 100ms·feelsinstant0.99
  • 7:38 #20 skipped

    shot 20·duplicate of #19

  • 8:01 #21 done6 line(s)

    shot 21·sharpness 312.7

    1. AIE1.00
    2. 1000ms·seamless visuals0.99
    3. STT1.00
    4. first token1.00
    5. variabiity1.00
    6. 100ms·feelsinstant0.99
  • 8:34 #22 done6 line(s)

    shot 22·sharpness 313.9

    1. AIE1.00
    2. 1000ms·seamless visuals0.99
    3. STT1.00
    4. first token0.99
    5. variability1.00
    6. 100ms·feelsinstant0.98
  • 8:55 #23 done4 line(s)

    shot 23·sharpness 678.8

    1. Pillars of low latency1.00
    2. AIE1.00
    3. 11.00
    4. Fast models1.00

Transcript

98 cues· 1,853 words· 10,103 chars

  1. 0:01 All right, I'm Alan Pike, and today I'm going to be sharing some of what we've learned building voice in, visuals out, experiences using AI.
  2. 0:14 Here we go.
  3. 0:16 So this is Andrej Karpathy, and he made an argument last month that voice is the human preferred input
  4. 0:29 for AIs, but that we prefer visuals as the output.
  5. 0:36 And he knows a thing or two about this stuff, but that's not how we have been, for the most part, building or using AI.
  6. 0:49 We've been typing to it, it's been typing back maybe with some markdown.
  7. 0:55 But there's some breakthroughs over the last few months
  8. 0:58 where both this visuals in and audio, or the audio in and visuals out experience is now really feasible and we can create really delightful experiences with it.
  9. 1:12 The visuals outside is pretty intuitive, right?
  10. 1:16 Of course, a third of our brain is dedicated to processing visual information.
  11. 1:21 We love looking at things.
  12. 1:24 And models have recently got to the point where they can generate rich HTML, tool calling.
  13. 1:30 So we can have these experiences where there's visualizations that come back, explain things, help us understand, communicate the responses from these models.
  14. 1:42 They can give us interactive controls that allow us to explore and understand and modify and change and direct the models.
  15. 1:54 And they can even respond with beautiful illustrations and images.
  16. 2:01 And so the ceiling on the visuals out piece has really lifted, and we have a lot more capability of what we can do in terms of responding from these models.
  17. 2:15 The more controversial half of Karpathy's argument, though, is this idea of voice as the preferred input.
  18. 2:23 We've long sort of idealized and fantasized about speaking to an AI and having this real-time conversation where it understands what we need and it reacts appropriately in real time.
  19. 2:39 But the experiences that most people have had so far with voice interfaces have been more like trying to get Siri to turn the lights on and it's not working, or like this guy that is like trying to get ChatGPT voice mode to do things, but it keeps being like awkward and confused, right?
  20. 3:01 The models that we have so far, the experiences that most people have seen, are both slow and dumb, which is not a great combination.
  21. 3:14 And so a lot of people are down on voices and input.
  22. 3:18 But speaking is the ultimate way that as humans we communicate.
  23. 3:24 When we're speaking, we have more words per minute than when we're typing.
  24. 3:31 but we also convey more with each word.
  25. 3:35 There's a huge difference in between if you say to me, okay, versus if you say, okay.
  26. 3:46 And that's why when we have something that is really important that we need to communicate or get through to somebody, we will jump on a call.
  27. 3:59 and we will, or we'll speak in person so that we can get through that high bandwidth communication.
  28. 4:09 And so what we've done at ForestWalk is we have built an agent that is in our calls that can help us out in real time.
  29. 4:18 And so the other day I mentioned to my co-founder on a call that I'd seen a bug with the Slack integration, or at least I thought I had.
  30. 4:29 And she responded that she had seen the same bug.
  31. 4:34 And so I just said, okay, well, let's file that as a linear issue.
  32. 4:40 And the voice agent within a second responded that it had done so.
  33. 4:46 And that feels perfectly natural when you get it really dialed in that you are speaking, either incidentally or with a purpose, intentionally speaking to the AI.
  34. 4:59 And it responds in a way that is not interruptive of you.
  35. 5:04 It doesn't need to be voice.
  36. 5:06 It is just taking action on your intent.
  37. 5:10 And so we're going to see this experience more and more in the coming months and years.
  38. 5:16 But there's a huge barrier, a huge challenge to making this actually work and feel good.
  39. 5:22 And that is the tyranny of latency.
  40. 5:29 It's really difficult to get a response through that whole chain fast enough.
  41. 5:35 We've known since the 60s that to have a computer react to us fast enough that it feels
  42. 5:44 Instead, it needs to react in about 100 milliseconds, a tenth of a second.
  43. 5:51 And so, of course, we are always aspiring to get our products to react that fast.
  44. 5:57 And sometimes we can achieve that.
  45. 5:58 But with networking and everything, it can be challenging, depending on what work needs to get done.
  46. 6:04 And so sometimes we might flex up to 1,000 milliseconds.
  47. 6:08 You get a full second.
  48. 6:09 It's about the limit before people start to lose their train of thought.
  49. 6:12 They ask people to do something.
  50. 6:14 It takes more than a second.

Open at this second