Videos mOf-PP4mVjA
Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs
Scene timeline
77 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 237
- whisperx 237
- chunks
- 36
- from 237 cues
- keyframes
- 52
- kept of 77 captured
- frames with text
- 52
- 1,985 lines read
- chapters
- 11
- from the source metadata
- keyframe bytes
- 11.1 MB
- word timings on 237 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-10 04:39 | 1m 39s |
stt |
done | — | 2026-08-10 04:41 | 22s |
chunk |
done | — | 2026-08-10 04:41 | 0s |
text_embed |
done | — | 2026-08-10 19:48 | 1s |
keyframe |
done | — | 2026-08-10 04:41 | 2m 22s |
ocr |
done | — | 2026-08-10 04:44 | 35s |
frame_embed |
done | — | 2026-08-10 19:48 | 9s |
Frames, and what the machine read
-
- AlEngineer0.96
- World's Fair1.00
-
- AIEngineer0.95
- World's Fair0.99
-
- LAB & PLATINUM SPONSORS0.99
- Amazon AGI Lab0.98
- ANTHROP\C1.00
- Google DeepMind1.00
- MINIMAX0.97
- OpenAI0.92
- Akamai1.00
- arize0.92
- aws1.00
- Braintrust bright data0.98
- B1.00
- Browserbase1.00
- docker1.00
- :neo4j0.93
- ORACLE1.00
- PayPal1.00
- qodo1.00
- reducto1.00
- Sonar1.00
- Makers of0.99
- togetherai1.00
- Unblocked1.00
- WorkOS1.00
- SonarQube1.00
-
- Video Has No Memory. Here is0.99
- AlEngineer0.97
- World'sFair1.00
- How We Built One.1.00
- James Le1.00
- PRESENTED BY1.00
- Microsoft1.00
- Graph Track1.00
- Al Engineering World's Fair, San Francisco1.00
- July 2,20261.00
- TwelveLabs1.00
- Engineering the future of Al0.99
- World'sFair1.00
-
- CHALLENGE1.00
- TwelveLabs1.00
- AlEngineer0.99
- Video Is A Spatiotemporal Volume, Not A Bag of Frames1.00
- World'sFair0.97
- Video data is incredibly complex1.00
- The scale is staggering1.00
- Holistic understanding is critical1.00
- More than just the sum of its parts (visuals,0.99
- Petabytes of footage in enterprises of every0.99
- At TwelveLabs, we're building foundation models1.00
- audio, movement, and time), video captures the1.00
- kind—that's millions and millions of hours of1.00
- that understand video the way humans do—not as0.98
- spatio-temporal context of the world around us.0.98
- video. Finding meaning in it, finding the0.99
- a sequence of frames or transcripts, but as a unified1.00
- PRESENTED BY1.00
- Language models and legacy tagging methods0.98
- moments that matter, is even harder.1.00
- story across sight, sound, and time.0.98
- fail to capture this nuance.0.98
- Microsoft1.00
- CAMERA 010.91
- CAMERA 020.89
- 1:03-1:450.97
- CAMERA 030.94
- CAMERA 040.97
- 2:43-3:050.99
- Time1.00
- 1:56-2:080.97
- SOURCE/INPUTS0.96
- Space1.00
- MONTHS0.99
- Time1.00
- TRACK 5· JULY 2, 20260.96
- Graphs1.00
- World'sFair1.00
-
- THE BLOCKER0.98
- TwelveLabs1.00
- AlEngineer0.99
- LLMs and the supporting stacks are1.00
- World'sFair1.00
- fundamentally wrong for video0.99
- Wrong Context0.98
- Video isn't text. But LLMs treat it as text0.99
- by chopping it into tokens, losing the1.00
- 1 PETABYTE OF TOKENS1.00
- spatio-temporal context that makes video1.00
- unique.1.00
- Wrong Memory1.00
- context windows. Video requires1.00
- Memory for video is not RAG or larger1.00
- HOURS OF FOOTAGE0.98
- long-term memory that links today's scene0.99
- to what happened days, or years ago.1.00
- Wrong Reasoning1.00
- LLM1.00
- Text-first models can't reason over0.97
- spatiotemporal structure. They miss motion,1.00
- causality, and progression — the very0.99
- elements that define video understanding.0.99
- TRACK 5• JULY 2, 20260.96
- World's Fair0.98
- AlEngin0.93
- Graphs0.99
-
- AlEngineer0.98
- World's Fair0.99
- Video is Temporal, Multimodal, Dense, Ambiguous, and1.00
- Evidence-Sensitive1.00
- Challenges1.00
- —mm—mmmm0.75
- Temporal1.00
- Meaning depends on1.00
- before and after1.00
- Evidence spans visual,1.00
- Transcript1.00
- Multimodal1.00
- speech, audio, OCR, metadata0.99
- Useful moments are1.00
- Dense1.00
- sparse inside noisy footage0.98
- Identity and meaning1.00
- Ambiguous1.00
- emerge across time0.99
- Repeated inspection costs1.00
- Expensive1.00
- latency and compute0.99
- 00:001.00
- 09:001.00
- 41.00
- TRACK 5· JULY 2, 20260.96
- World's Fair1.00
- AlEngin0.94
- Graphs1.00
-
- SOLUTION1.00
- OVERVIEW1.00
- TwelveLabs1.00
- AlEngineer0.99
- The TwelveLabs Video Intelligence Stack1.00
- World's Fair0.98
- Pegasus1.00
- API0.99
- Video Context Aware Language Model1.00
- Search0.92
- XEmbed0.91
- Aealyze0.75
- 02000.51
- Spatiotemporal Context Store1.00
- Marengo1.00
- Spatiotemporal Contexts1.00
- Semantic Chunks0.99
- TRACK 5• JULY 2, 20260.96
- Graphs1.00
- World's Fair0.99
-
- SOLUTION1.00
- OVERVIEW1.00
- TwelveLabs1.00
- AlEngineer1.00
- The TwelveLabs Video Intelligence Stack1.00
- World's Fair0.96
- Pegasus1.00
- API1.00
- Video Context Aware Language Model1.00
- Search0.93
- ×Embed0.92
- Acalyzn0.79
- 82000.64
- PRESENTED BY1.00
- Spatiotemporal Context Store1.00
- Microsoft1.00
- Marengo1.00
- Spatiotemporal Contexts1.00
- Semantic Chunks1.00
- TRACK 5• JULY 2, 20260.95
- AlEnginee0.94
- Graphs1.00
- World's Fair0.99
-
- BUILDING VIDEO REASONING AGENT0.99
- AlEngineer0.98
- From Clip1.00
- World's Fair0.97
- Retrieval to0.98
- Corpus Memory1.00
- Time scaling1.00
- Reason over years of footage via0.99
- memory-first retrieval: multi-hop1.00
- timelines and episodic recall at low1.00
- latency and cost0.99
- Space1.00
- Space scaling1.00
- Time1.00
- SOURCE/INPUTS1.00
- YEARS1.00
- Fuse perspectives and live streams1.00
- from thousands of sources and streams0.99
- to construct coherent understanding1.00
- TRACK 5• JULY 2, 20260.95
- Graphs1.00
- World's Fair0.96
-
- The Context Graph As A1.00
- AlEngineer0.99
- System Concept1.00
- World's Fair0.96
- [ Corpus-level context ] ← themes,0.98
- gaps, patterns, coverage1.00
- ↑0.52
- [Relationships ] ← co-occurrence,0.98
- sequence, cause, timeline1.00
- [ Entities ] ← people, brands, places,0.97
- objects, concepts1.00
- [Appearances ] ← where + when each0.99
- entity shows up1.00
- TRANSCRIPT1.00
- [ Time-bounded Moments ] ← clips,0.99
- scenes, shots with start/end times0.99
- 0:16-0:211.00
- TRACK 5· JULY 2, 20260.96
- Graphs0.98
- World'sFair1.00
-
- The 5 Principles for Building1.00
- AlEngineer0.99
- a Video Memory Layer1.00
- World'sFair1.00
- 1 - Ingest Once, Reason Many Times:0.96
- Move expensive understanding into a1.00
- preparation step. Same mental model1.00
- as databases1.00
- Keep it composable1.00
- Let intent shape memory0.99
- 2 - Store Primitives, Not Just Answers:1.00
- Prepare reusable structure1.00
- Plug into apps and agents1.00
- Moments, entities, appearances,1.00
- relationships, compose. Summaries do0.99
- not.1.00
- Ingest once,1.00
- Represent primitives,1.00
- 3 - Ground Every Claim: A timestamp0.99
- reason many times0.98
- not answers1.00
- is a product requirement. Evidence0.99
- Store moments, entities, events1.00
- Extract what the workflow needs1.00
- should survive synthesis.1.00
- 4 - Let Intent Shape Memory: Brand0.99
- Preserve grounding1.00
- safety and sports highlights need1.00
- different primitives from the same1.00
- Trace claims to evidence1.00
- footage1.00
- 5 - Keep The Memory Layer0.99
- Composable: APl-first. It should plug0.99
- into agents, dashboards, and review0.99
- tools - not become all of them.0.99
- TRACK 5• JULY 2, 20260.95
- AlEngin0.96
- Graphs1.00
- World'sFair1.00
Transcript
237 cues· 3,137 words· 18,255 chars
- 0:12 Thanks so much for having me and inviting me to be a speaker at the World Fair.
- 0:17 I attended last year and was so impressed by the quality of presenters.
- 0:20 So glad to have a chance to be here and present.
- 0:23 So the title of my talk is Video Has No Memory.
- 0:27 And this might sound strange because video is already a preservation of the past.
- 0:32 We think about you have footage, you preserve.
- 0:35 Recording, training data, incident, creative work, history, et cetera.
- 0:40 But actually, most of the video AI systems these days do not have memory in the system sense.
- 0:44 So actually, for this talk, I will try to answer the question, what could it take to build a memory layer for video intelligence?
- 0:51 To start, I want to be clear about what makes Vue different from other data type, right?
- 0:56 So this is the first mental model that I want to highlight, which is that video is not a bag of frames.
- 1:01 So in many of my conversations with developers who are using a product, a lot of them still treat video as like a stack of images, maybe a transcript being attached or, you know,
- 1:13 but essentially like a frame level, right?
- 1:16 And that is useful approximation for some tasks, but it throw away the thing that makes video very unique, which is continuity, right?
- 1:24 So meaning in video derives from space, time, modalities, and sequence.
- 1:29 So a better mental model,
- 1:30 for video is a spatial temporal volume.
- 1:33 So what I mean that inside that volume, you have visual information, speech, sound, motion, OCR, camera changes, scene transition, metadata, and time, right?
- 1:43 So the hard part here is really, well, how can you preserve in relationship across this volume so that later an application can traverse it?
- 1:51 And then especially at the enterprise scale across like,
- 1:54 industry like entertainment sport you know short form content then you're sitting on petabytes of footage right so fighting moment is already hard so how can you preserve meaning across millions of moments in the deeper platform so i work at 12 app which is a series b startup and we do foundation models that understand you know video the way that humans do um and the way we talk about our positioning is like the existing uh stack of dealing with video is not equipped to do that right
- 2:24 Obviously, language models are very powerful.
- 2:27 They are good reasoning interfaces.
- 2:29 They are increasingly multimodal as well.
- 2:32 But the supporting stack around that is very, I would say, limited.
- 2:37 And that creates three problems.
- 2:38 Number one is wrong context.
- 2:40 So video is not naturally a sequence of text token.
- 2:44 If we force it into that sequence by sampling frames, by extracting a transcript, by dumping everything into a prompt, you lose the spatial temporal relationships that actually define the event.
- 2:55 Second, wrong memory.
- 2:57 So if you think about tech system, memory here is often mean generation, vector search, or probably like larger context window.
- 3:04 Those are very useful, but video memory has a different requirement.
- 3:07 It needs to link to the scene for something that happened in another file, another episode, another camera angle, another season, another year.
- 3:14 So it actually needs durable continuity.
- 3:16 And the last part here is strong reasoning.
- 3:18 Like I said, you know, text-first system cannot reason over, you know, natively over motion, causality, all of that.
- 3:25 So, you know, they do not automatically build like a persistent structure on, you know, who appear, what happen, what changes, et cetera.
- 3:34 And so my argument is that video intelligence need a memory layer that decide what to preverse, how to connect it, and how to reshoot later.
- 3:41 So I want to kind of ground it into the properties of video, right, to make it even clearer.
- 3:48 There's five challenges dealing with video.
- 3:50 Number one is temporal, right?
- 3:52 So meaning depends on before and after.
- 3:55 So a frame by itself can be misleading, right?
- 3:58 The same expression, product shot, physical action can mean different things depending on the sequence around it, right?
- 4:03 Second is that video is obviously multimodal, I explained already.
- 4:07 A transcript alone may miss the logo, a frame alone may miss the spoken claim.
- 4:13 Video is also very dense, right?
- 4:15 So a few minutes can contain dozens of shots, people, objects, action, location, claims.
- 4:21 The useful signal is uneven across the distribution on the frame.
- 4:25 Some seconds are decisive, others are noisy.
loading
Chapters
- 0:00 Video has no memory
- 0:50 Video is a spatial temporal volume, not a bag of frames
- 2:06 Three problems: wrong context, wrong memory, weak reasoning
- 3:36 Five properties that make video memory hard
- 4:53 The TwelveLabs stack: Marengo, the context store, and Pegasus
- 5:56 Search versus memory
- 7:48 The context graph
- 9:04 Five design principles for a video memory layer
- 10:45 From a static model to a video worker
- 12:51 Demo: sports and tracking Messi across the World Cup
- 15:37 Demos: traffic security and ad placement