read-only demo

Videos mOf-PP4mVjA

Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs

index_state ready data_status ok

AI Engineer· published 2026-07-23· 0:20:27· en-US· indexed 2026-08-10 19:48

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:03, 1 of 1 keyframes kept
  2. Shot 1, 0:03 to 0:05, 1 of 1 keyframes kept
  3. Shot 2, 0:05 to 0:12, 1 of 1 keyframes kept
  4. Shot 3, 0:12 to 0:55, 1 of 1 keyframes kept
  5. Shot 4, 0:55 to 1:23, 1 of 1 keyframes kept
  6. Shot 5, 1:23 to 1:51, 0 of 1 keyframes kept
  7. Shot 6, 1:51 to 2:19, 0 of 1 keyframes kept
  8. Shot 7, 2:19 to 2:46, 1 of 1 keyframes kept
  9. Shot 8, 2:46 to 3:14, 0 of 1 keyframes kept
  10. Shot 9, 3:14 to 3:41, 0 of 1 keyframes kept
  11. Shot 10, 3:41 to 4:08, 1 of 1 keyframes kept
  12. Shot 11, 4:08 to 4:34, 0 of 1 keyframes kept
  13. Shot 12, 4:34 to 5:00, 0 of 1 keyframes kept
  14. Shot 13, 5:00 to 5:19, 1 of 1 keyframes kept
  15. Shot 14, 5:19 to 5:52, 1 of 1 keyframes kept
  16. Shot 15, 5:52 to 5:55, 0 of 1 keyframes kept
  17. Shot 16, 5:55 to 6:23, 0 of 1 keyframes kept
  18. Shot 17, 6:23 to 6:50, 0 of 1 keyframes kept
  19. Shot 18, 6:50 to 7:19, 1 of 1 keyframes kept
  20. Shot 19, 7:19 to 7:49, 0 of 1 keyframes kept
  21. Shot 20, 7:49 to 8:15, 1 of 1 keyframes kept
  22. Shot 21, 8:15 to 8:41, 0 of 1 keyframes kept
  23. Shot 22, 8:41 to 9:08, 0 of 1 keyframes kept
  24. Shot 23, 9:08 to 9:40, 1 of 1 keyframes kept
  25. Shot 24, 9:40 to 10:13, 0 of 1 keyframes kept
  26. Shot 25, 10:13 to 10:45, 0 of 1 keyframes kept
  27. Shot 26, 10:45 to 11:13, 1 of 1 keyframes kept
  28. Shot 27, 11:13 to 11:42, 0 of 1 keyframes kept
  29. Shot 28, 11:42 to 12:15, 0 of 1 keyframes kept
  30. Shot 29, 12:15 to 12:49, 0 of 1 keyframes kept
  31. Shot 30, 12:49 to 12:50, 1 of 1 keyframes kept
  32. Shot 31, 12:50 to 12:55, 1 of 1 keyframes kept
  33. Shot 32, 12:55 to 13:28, 1 of 1 keyframes kept
  34. Shot 33, 13:28 to 13:52, 1 of 1 keyframes kept
  35. Shot 34, 13:52 to 14:00, 1 of 1 keyframes kept
  36. Shot 35, 14:00 to 14:15, 1 of 1 keyframes kept
  37. Shot 36, 14:15 to 14:19, 1 of 1 keyframes kept
  38. Shot 37, 14:19 to 14:36, 1 of 1 keyframes kept
  39. Shot 38, 14:36 to 14:39, 1 of 1 keyframes kept
  40. Shot 39, 14:39 to 14:48, 0 of 1 keyframes kept
  41. Shot 40, 14:48 to 14:58, 1 of 1 keyframes kept
  42. Shot 41, 14:58 to 15:20, 1 of 1 keyframes kept
  43. Shot 42, 15:20 to 15:26, 1 of 1 keyframes kept
  44. Shot 43, 15:26 to 15:28, 1 of 1 keyframes kept
  45. Shot 44, 15:28 to 15:33, 1 of 1 keyframes kept
  46. Shot 45, 15:33 to 15:43, 1 of 1 keyframes kept
  47. Shot 46, 15:43 to 15:48, 1 of 1 keyframes kept
  48. Shot 47, 15:48 to 15:49, 1 of 1 keyframes kept
  49. Shot 48, 15:49 to 15:57, 1 of 1 keyframes kept
  50. Shot 49, 15:57 to 16:03, 1 of 1 keyframes kept
  51. Shot 50, 16:03 to 16:25, 1 of 1 keyframes kept
  52. Shot 51, 16:25 to 16:28, 1 of 1 keyframes kept
  53. Shot 52, 16:28 to 16:32, 1 of 1 keyframes kept
  54. Shot 53, 16:32 to 16:36, 1 of 1 keyframes kept
  55. Shot 54, 16:36 to 16:39, 1 of 1 keyframes kept
  56. Shot 55, 16:39 to 16:52, 1 of 1 keyframes kept
  57. Shot 56, 16:52 to 16:55, 1 of 1 keyframes kept
  58. Shot 57, 16:55 to 17:07, 1 of 1 keyframes kept
  59. Shot 58, 17:07 to 17:10, 1 of 1 keyframes kept
  60. Shot 59, 17:10 to 17:49, 1 of 1 keyframes kept
  61. Shot 60, 17:49 to 17:57, 1 of 1 keyframes kept
  62. Shot 61, 17:57 to 18:00, 1 of 1 keyframes kept
  63. Shot 62, 18:00 to 18:02, 0 of 1 keyframes kept
  64. Shot 63, 18:02 to 18:05, 1 of 1 keyframes kept
  65. Shot 64, 18:05 to 18:08, 0 of 1 keyframes kept
  66. Shot 65, 18:08 to 18:10, 1 of 1 keyframes kept
  67. Shot 66, 18:10 to 18:14, 1 of 1 keyframes kept
  68. Shot 67, 18:14 to 18:18, 1 of 1 keyframes kept
  69. Shot 68, 18:18 to 18:20, 0 of 1 keyframes kept
  70. Shot 69, 18:20 to 18:21, 0 of 1 keyframes kept
  71. Shot 70, 18:21 to 18:48, 0 of 1 keyframes kept
  72. Shot 71, 18:48 to 19:04, 1 of 1 keyframes kept
  73. Shot 72, 19:04 to 19:43, 0 of 1 keyframes kept
  74. Shot 73, 19:43 to 19:47, 1 of 1 keyframes kept
  75. Shot 74, 19:47 to 19:48, 1 of 1 keyframes kept
  76. Shot 75, 19:48 to 20:10, 1 of 1 keyframes kept
  77. Shot 76, 20:10 to 20:26, 0 of 1 keyframes kept

77 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
237
whisperx 237
chunks
36
from 237 cues
keyframes
52
kept of 77 captured
frames with text
52
1,985 lines read
chapters
11
from the source metadata
keyframe bytes
11.1 MB
word timings on 237 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-10 04:39 1m 39s
stt done 2026-08-10 04:41 22s
chunk done 2026-08-10 04:41 0s
text_embed done 2026-08-10 19:48 1s
keyframe done 2026-08-10 04:41 2m 22s
ocr done 2026-08-10 04:44 35s
frame_embed done 2026-08-10 19:48 9s

Frames, and what the machine read

  • 0:02 #0 done2 line(s)

    shot 0·sharpness 449.7

    1. AlEngineer0.96
    2. World's Fair1.00
  • 0:03 #1 done2 line(s)

    shot 1·sharpness 662.0

    1. AIEngineer0.95
    2. World's Fair0.99
  • 0:10 #2 done24 line(s)

    shot 2·sharpness 2739.9

    1. LAB & PLATINUM SPONSORS0.99
    2. Amazon AGI Lab0.98
    3. ANTHROP\C1.00
    4. Google DeepMind1.00
    5. MINIMAX0.97
    6. OpenAI0.92
    7. Akamai1.00
    8. arize0.92
    9. aws1.00
    10. Braintrust bright data0.98
    11. B1.00
    12. Browserbase1.00
    13. docker1.00
    14. :neo4j0.93
    15. ORACLE1.00
    16. PayPal1.00
    17. qodo1.00
    18. reducto1.00
    19. Sonar1.00
    20. Makers of0.99
    21. togetherai1.00
    22. Unblocked1.00
    23. WorkOS1.00
    24. SonarQube1.00
  • 0:38 #3 done13 line(s)

    shot 3·sharpness 2589.8

    1. Video Has No Memory. Here is0.99
    2. AlEngineer0.97
    3. World'sFair1.00
    4. How We Built One.1.00
    5. James Le1.00
    6. PRESENTED BY1.00
    7. Microsoft1.00
    8. Graph Track1.00
    9. Al Engineering World's Fair, San Francisco1.00
    10. July 2,20261.00
    11. TwelveLabs1.00
    12. Engineering the future of Al0.99
    13. World'sFair1.00
  • 1:14 #4 done38 line(s)

    shot 4·sharpness 3750.4

    1. CHALLENGE1.00
    2. TwelveLabs1.00
    3. AlEngineer0.99
    4. Video Is A Spatiotemporal Volume, Not A Bag of Frames1.00
    5. World'sFair0.97
    6. Video data is incredibly complex1.00
    7. The scale is staggering1.00
    8. Holistic understanding is critical1.00
    9. More than just the sum of its parts (visuals,0.99
    10. Petabytes of footage in enterprises of every0.99
    11. At TwelveLabs, we're building foundation models1.00
    12. audio, movement, and time), video captures the1.00
    13. kind—that's millions and millions of hours of1.00
    14. that understand video the way humans do—not as0.98
    15. spatio-temporal context of the world around us.0.98
    16. video. Finding meaning in it, finding the0.99
    17. a sequence of frames or transcripts, but as a unified1.00
    18. PRESENTED BY1.00
    19. Language models and legacy tagging methods0.98
    20. moments that matter, is even harder.1.00
    21. story across sight, sound, and time.0.98
    22. fail to capture this nuance.0.98
    23. Microsoft1.00
    24. CAMERA 010.91
    25. CAMERA 020.89
    26. 1:03-1:450.97
    27. CAMERA 030.94
    28. CAMERA 040.97
    29. 2:43-3:050.99
    30. Time1.00
    31. 1:56-2:080.97
    32. SOURCE/INPUTS0.96
    33. Space1.00
    34. MONTHS0.99
    35. Time1.00
    36. TRACK 5· JULY 2, 20260.96
    37. Graphs1.00
    38. World'sFair1.00
  • 1:47 #5 skipped

    shot 5·duplicate of #4

  • 2:05 #6 skipped

    shot 6·duplicate of #4

  • 2:22 #7 done28 line(s)

    shot 7·sharpness 2375.4

    1. THE BLOCKER0.98
    2. TwelveLabs1.00
    3. AlEngineer0.99
    4. LLMs and the supporting stacks are1.00
    5. World'sFair1.00
    6. fundamentally wrong for video0.99
    7. Wrong Context0.98
    8. Video isn't text. But LLMs treat it as text0.99
    9. by chopping it into tokens, losing the1.00
    10. 1 PETABYTE OF TOKENS1.00
    11. spatio-temporal context that makes video1.00
    12. unique.1.00
    13. Wrong Memory1.00
    14. context windows. Video requires1.00
    15. Memory for video is not RAG or larger1.00
    16. HOURS OF FOOTAGE0.98
    17. long-term memory that links today's scene0.99
    18. to what happened days, or years ago.1.00
    19. Wrong Reasoning1.00
    20. LLM1.00
    21. Text-first models can't reason over0.97
    22. spatiotemporal structure. They miss motion,1.00
    23. causality, and progression — the very0.99
    24. elements that define video understanding.0.99
    25. TRACK 5• JULY 2, 20260.96
    26. World's Fair0.98
    27. AlEngin0.93
    28. Graphs0.99
  • 3:00 #8 skipped

    shot 8·duplicate of #7

  • 3:33 #9 skipped

    shot 9·duplicate of #7

  • 3:59 #10 done29 line(s)

    shot 10·sharpness 4605.1

    1. AlEngineer0.98
    2. World's Fair0.99
    3. Video is Temporal, Multimodal, Dense, Ambiguous, and1.00
    4. Evidence-Sensitive1.00
    5. Challenges1.00
    6. —mm—mmmm0.75
    7. Temporal1.00
    8. Meaning depends on1.00
    9. before and after1.00
    10. Evidence spans visual,1.00
    11. Transcript1.00
    12. Multimodal1.00
    13. speech, audio, OCR, metadata0.99
    14. Useful moments are1.00
    15. Dense1.00
    16. sparse inside noisy footage0.98
    17. Identity and meaning1.00
    18. Ambiguous1.00
    19. emerge across time0.99
    20. Repeated inspection costs1.00
    21. Expensive1.00
    22. latency and compute0.99
    23. 00:001.00
    24. 09:001.00
    25. 41.00
    26. TRACK 5· JULY 2, 20260.96
    27. World's Fair1.00
    28. AlEngin0.94
    29. Graphs1.00
  • 4:31 #11 skipped

    shot 11·duplicate of #10

  • 4:42 #12 skipped

    shot 12·duplicate of #10

  • 5:13 #13 done20 line(s)

    shot 13·sharpness 1382.8

    1. SOLUTION1.00
    2. OVERVIEW1.00
    3. TwelveLabs1.00
    4. AlEngineer0.99
    5. The TwelveLabs Video Intelligence Stack1.00
    6. World's Fair0.98
    7. Pegasus1.00
    8. API0.99
    9. Video Context Aware Language Model1.00
    10. Search0.92
    11. XEmbed0.91
    12. Aealyze0.75
    13. 02000.51
    14. Spatiotemporal Context Store1.00
    15. Marengo1.00
    16. Spatiotemporal Contexts1.00
    17. Semantic Chunks0.99
    18. TRACK 5• JULY 2, 20260.96
    19. Graphs1.00
    20. World's Fair0.99
  • 5:45 #14 done23 line(s)

    shot 14·sharpness 1515.3

    1. SOLUTION1.00
    2. OVERVIEW1.00
    3. TwelveLabs1.00
    4. AlEngineer1.00
    5. The TwelveLabs Video Intelligence Stack1.00
    6. World's Fair0.96
    7. Pegasus1.00
    8. API1.00
    9. Video Context Aware Language Model1.00
    10. Search0.93
    11. ×Embed0.92
    12. Acalyzn0.79
    13. 82000.64
    14. PRESENTED BY1.00
    15. Spatiotemporal Context Store1.00
    16. Microsoft1.00
    17. Marengo1.00
    18. Spatiotemporal Contexts1.00
    19. Semantic Chunks1.00
    20. TRACK 5• JULY 2, 20260.95
    21. AlEnginee0.94
    22. Graphs1.00
    23. World's Fair0.99
  • 5:54 #15 skipped

    shot 15·duplicate of #14

  • 6:14 #16 skipped

    shot 16·duplicate of #10

  • 6:29 #17 skipped

    shot 17·duplicate of #10

  • 7:08 #18 done22 line(s)

    shot 18·sharpness 2733.4

    1. BUILDING VIDEO REASONING AGENT0.99
    2. AlEngineer0.98
    3. From Clip1.00
    4. World's Fair0.97
    5. Retrieval to0.98
    6. Corpus Memory1.00
    7. Time scaling1.00
    8. Reason over years of footage via0.99
    9. memory-first retrieval: multi-hop1.00
    10. timelines and episodic recall at low1.00
    11. latency and cost0.99
    12. Space1.00
    13. Space scaling1.00
    14. Time1.00
    15. SOURCE/INPUTS1.00
    16. YEARS1.00
    17. Fuse perspectives and live streams1.00
    18. from thousands of sources and streams0.99
    19. to construct coherent understanding1.00
    20. TRACK 5• JULY 2, 20260.95
    21. Graphs1.00
    22. World's Fair0.96
  • 7:31 #19 skipped

    shot 19·duplicate of #18

  • 8:12 #20 done20 line(s)

    shot 20·sharpness 3238.1

    1. The Context Graph As A1.00
    2. AlEngineer0.99
    3. System Concept1.00
    4. World's Fair0.96
    5. [ Corpus-level context ] ← themes,0.98
    6. gaps, patterns, coverage1.00
    7. 0.52
    8. [Relationships ] ← co-occurrence,0.98
    9. sequence, cause, timeline1.00
    10. [ Entities ] ← people, brands, places,0.97
    11. objects, concepts1.00
    12. [Appearances ] ← where + when each0.99
    13. entity shows up1.00
    14. TRANSCRIPT1.00
    15. [ Time-bounded Moments ] ← clips,0.99
    16. scenes, shots with start/end times0.99
    17. 0:16-0:211.00
    18. TRACK 5· JULY 2, 20260.96
    19. Graphs0.98
    20. World'sFair1.00
  • 8:28 #21 skipped

    shot 21·duplicate of #20

  • 8:52 #22 skipped

    shot 22·duplicate of #20

  • 9:27 #23 done39 line(s)

    shot 23·sharpness 4299.4

    1. The 5 Principles for Building1.00
    2. AlEngineer0.99
    3. a Video Memory Layer1.00
    4. World'sFair1.00
    5. 1 - Ingest Once, Reason Many Times:0.96
    6. Move expensive understanding into a1.00
    7. preparation step. Same mental model1.00
    8. as databases1.00
    9. Keep it composable1.00
    10. Let intent shape memory0.99
    11. 2 - Store Primitives, Not Just Answers:1.00
    12. Prepare reusable structure1.00
    13. Plug into apps and agents1.00
    14. Moments, entities, appearances,1.00
    15. relationships, compose. Summaries do0.99
    16. not.1.00
    17. Ingest once,1.00
    18. Represent primitives,1.00
    19. 3 - Ground Every Claim: A timestamp0.99
    20. reason many times0.98
    21. not answers1.00
    22. is a product requirement. Evidence0.99
    23. Store moments, entities, events1.00
    24. Extract what the workflow needs1.00
    25. should survive synthesis.1.00
    26. 4 - Let Intent Shape Memory: Brand0.99
    27. Preserve grounding1.00
    28. safety and sports highlights need1.00
    29. different primitives from the same1.00
    30. Trace claims to evidence1.00
    31. footage1.00
    32. 5 - Keep The Memory Layer0.99
    33. Composable: APl-first. It should plug0.99
    34. into agents, dashboards, and review0.99
    35. tools - not become all of them.0.99
    36. TRACK 5• JULY 2, 20260.95
    37. AlEngin0.96
    38. Graphs1.00
    39. World'sFair1.00

Transcript

237 cues· 3,137 words· 18,255 chars

  1. 0:12 Thanks so much for having me and inviting me to be a speaker at the World Fair.
  2. 0:17 I attended last year and was so impressed by the quality of presenters.
  3. 0:20 So glad to have a chance to be here and present.
  4. 0:23 So the title of my talk is Video Has No Memory.
  5. 0:27 And this might sound strange because video is already a preservation of the past.
  6. 0:32 We think about you have footage, you preserve.
  7. 0:35 Recording, training data, incident, creative work, history, et cetera.
  8. 0:40 But actually, most of the video AI systems these days do not have memory in the system sense.
  9. 0:44 So actually, for this talk, I will try to answer the question, what could it take to build a memory layer for video intelligence?
  10. 0:51 To start, I want to be clear about what makes Vue different from other data type, right?
  11. 0:56 So this is the first mental model that I want to highlight, which is that video is not a bag of frames.
  12. 1:01 So in many of my conversations with developers who are using a product, a lot of them still treat video as like a stack of images, maybe a transcript being attached or, you know,
  13. 1:13 but essentially like a frame level, right?
  14. 1:16 And that is useful approximation for some tasks, but it throw away the thing that makes video very unique, which is continuity, right?
  15. 1:24 So meaning in video derives from space, time, modalities, and sequence.
  16. 1:29 So a better mental model,
  17. 1:30 for video is a spatial temporal volume.
  18. 1:33 So what I mean that inside that volume, you have visual information, speech, sound, motion, OCR, camera changes, scene transition, metadata, and time, right?
  19. 1:43 So the hard part here is really, well, how can you preserve in relationship across this volume so that later an application can traverse it?
  20. 1:51 And then especially at the enterprise scale across like,
  21. 1:54 industry like entertainment sport you know short form content then you're sitting on petabytes of footage right so fighting moment is already hard so how can you preserve meaning across millions of moments in the deeper platform so i work at 12 app which is a series b startup and we do foundation models that understand you know video the way that humans do um and the way we talk about our positioning is like the existing uh stack of dealing with video is not equipped to do that right
  22. 2:24 Obviously, language models are very powerful.
  23. 2:27 They are good reasoning interfaces.
  24. 2:29 They are increasingly multimodal as well.
  25. 2:32 But the supporting stack around that is very, I would say, limited.
  26. 2:37 And that creates three problems.
  27. 2:38 Number one is wrong context.
  28. 2:40 So video is not naturally a sequence of text token.
  29. 2:44 If we force it into that sequence by sampling frames, by extracting a transcript, by dumping everything into a prompt, you lose the spatial temporal relationships that actually define the event.
  30. 2:55 Second, wrong memory.
  31. 2:57 So if you think about tech system, memory here is often mean generation, vector search, or probably like larger context window.
  32. 3:04 Those are very useful, but video memory has a different requirement.
  33. 3:07 It needs to link to the scene for something that happened in another file, another episode, another camera angle, another season, another year.
  34. 3:14 So it actually needs durable continuity.
  35. 3:16 And the last part here is strong reasoning.
  36. 3:18 Like I said, you know, text-first system cannot reason over, you know, natively over motion, causality, all of that.
  37. 3:25 So, you know, they do not automatically build like a persistent structure on, you know, who appear, what happen, what changes, et cetera.
  38. 3:34 And so my argument is that video intelligence need a memory layer that decide what to preverse, how to connect it, and how to reshoot later.
  39. 3:41 So I want to kind of ground it into the properties of video, right, to make it even clearer.
  40. 3:48 There's five challenges dealing with video.
  41. 3:50 Number one is temporal, right?
  42. 3:52 So meaning depends on before and after.
  43. 3:55 So a frame by itself can be misleading, right?
  44. 3:58 The same expression, product shot, physical action can mean different things depending on the sequence around it, right?
  45. 4:03 Second is that video is obviously multimodal, I explained already.
  46. 4:07 A transcript alone may miss the logo, a frame alone may miss the spoken claim.
  47. 4:13 Video is also very dense, right?
  48. 4:15 So a few minutes can contain dozens of shots, people, objects, action, location, claims.
  49. 4:21 The useful signal is uneven across the distribution on the frame.
  50. 4:25 Some seconds are decisive, others are noisy.

Chapters

  1. 0:00 Video has no memory
  2. 0:50 Video is a spatial temporal volume, not a bag of frames
  3. 2:06 Three problems: wrong context, wrong memory, weak reasoning
  4. 3:36 Five properties that make video memory hard
  5. 4:53 The TwelveLabs stack: Marengo, the context store, and Pegasus
  6. 5:56 Search versus memory
  7. 7:48 The context graph
  8. 9:04 Five design principles for a video memory layer
  9. 10:45 From a static model to a video worker
  10. 12:51 Demo: sports and tracking Messi across the World Cup
  11. 15:37 Demos: traffic security and ad placement

Open at this second