Videos XovaGv4f39A
When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis
Scene timeline
75 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 53
- whisperx 53
- chunks
- 10
- from 53 cues
- keyframes
- 59
- kept of 75 captured
- frames with text
- 38
- 184 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 8.1 MB
- word timings on 53 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 07:52 | 1m 34s |
stt |
done | — | 2026-08-11 07:53 | 6s |
chunk |
done | — | 2026-08-11 07:53 | 0s |
text_embed |
done | — | 2026-08-11 07:53 | 0s |
keyframe |
done | — | 2026-08-11 07:53 | 48s |
ocr |
done | — | 2026-08-11 07:54 | 10s |
frame_embed |
done | — | 2026-08-11 07:54 | 9s |
Frames, and what the machine read
-
- When all context matters:1.00
- Extended Cache1.00
- Augmented Generation (ECAG)0.99
-
- ORBIS1.00
-
- X0.54
-
- S0.94
-
- 11.00
- 21.00
- 91.00
- 10.98
- 31.00
- 450.96
- 51.00
- 71.00
- 90.87
- 21.00
- 31.00
- 61.00
- 51.00
- 91.00
- B20.65
- 11.00
- 31.00
- 20.82
- 21.00
- 70.92
- 40.75
- 70.98
- 50.59
- a0.60
- 50.96
- 7050.57
- 00.62
- 4A0.75
- 890.98
- 21.00
- 590.92
- 46540.85
- 91.00
- 61.00
- 160.65
- 51.00
- 90.74
- 150.65
- 90.96
- 61.00
- 80.98
-
- Apple1.00
- Fruit1.00
- Orange0.93
- Orange1.00
- Fruit1.00
- Train1.00
- Car1.00
- Train1.00
-
- Apple1.00
- Fr1.00
- Orange1.00
- 011101001100010000.97
- Fruit1.00
- Tr1.00
- Car1.00
-
- -3.69360.95
- 1600.99
- 1800.95
- -1.89800.94
- 1601.00
- 1401.00
- 1201.00
- -0.58800.82
- Apple0.91
- 1001.00
- 01.00
- -500.99
- 301.00
- -11.8330.88
- 0.39880.91
- 201.00
- -5.7330.89
- 21.00
- 31.00
- 41.00
- 51.00
- 61.00
- 71.00
- 81.00
- 91.00
- 101.00
- 111.00
- 131.00
- 131.00
- 141.00
- X0.64
-
- Apple1.00
- 6138436491.00
- 296702109555533937435925679320.97
-
- 50001.00
- 57501.00
- 46501.00
- 35001.00
- Orbi1.00
- X0.68
- 33501.00
- 21001.00
- 10001.00
- -1001.00
- 11.00
- 21.00
- 31.00
- 41.00
- 41.00
- 51.00
- 61.00
- 7 8 9 10 11 12 13 16 29 X0.98
-
- ORBIS0.99
-
- Context Window0.98
Transcript
53 cues· 845 words· 5,000 chars
- 0:10 I'm Luis Romero Sevilla and I'm the VP of AI at Orbis Operations.
- 0:15 I'm on a mission to solve knowledge representation when all context matters.
- 0:19 So, let's start with a very specific example.
- 0:24 Let's say we have a large number of documents.
- 0:26 And all documents represent an event.
- 0:28 And all documents in the collection are relevant to answer a set of questions that the user has.
- 0:34 Not only that, there's one more challenge.
- 0:38 the document in the collection becomes obsolete very fast and all documents get replaced with new information.
- 0:44 Let's start with the simplest approach.
- 0:46 We could start with a simple reg.
- 0:48 For that we just need a vector database and an embedding model.
- 0:53 An embedding model takes the documents and turns them into a learned numerical representation, a vector.
- 1:00 Now we take those vectors and we store them in a database optimized for performing operations with vectors.
- 1:07 Perfect.
- 1:08 Now we can take all of our questions, turn them into vectors, and then look for vectors that are similar to the initial query.
- 1:18 Those vectors that are within the similarity threshold are retrieved, and we can pass them to the LLM to answer the question.
- 1:27 Inserting to a vector database, it's relatively fast.
- 1:30 So whenever a collection becomes obsolete, we can just replace it with a new one.
- 1:36 We still have one problem with our very specific scenario.
- 1:41 All the documents in the collection are relevant for us to answer the question.
- 1:45 So we can't just take all the documents in the collection and pass them to LLM.
- 1:51 And that's just one of the many limitations with this approach.
- 1:55 Now let's get a bit more sophisticated.
- 1:58 All documents are relevant to answer a global question.
- 2:01 Therefore, there must be some connections and relationship between the details within a document in the collections.
- 2:08 For us to map out those relationship, we're going to need a knowledge graph.
- 2:12 And one implementation we could try is GraphRag.
- 2:16 GraphRag had many stats, but basically it uses an LLM to read through all the documents and extract key entities and relationship between them.
- 2:27 It constructs a network, a knowledge graph, where all those connections and details are tied together.
- 2:32 Then when a question is asked, it navigates this graph to synthesize a complete answer drawn across the entire collection.
- 2:39 If your collection of documents isn't changing very often, GraphRack is an excellent approach for finding those relationships within details to answer the user's question.
- 2:51 However, our very specific scenario states that our data is not only deeply interconnected, but also the data gets replaced very often.
- 3:01 Recomputing a knowledge graph every time the data gets replaced is computationally very expensive and it takes a relatively long time.
- 3:08 Okay, what if we take an even simpler approach and we continue to build on top of it?
- 3:13 If we were to use something like GraphRack, each document needs to pass through an LLM for the entity and relationship structure anyway.
- 3:21 Why can't we just throw all the documents into context?
- 3:24 This approach will look something like cache amended generation, CAG, where we use a model with a large context window, load the documents into the context, and cache the context by storing the model's KB matrix.
- 3:38 The problem here is that the context window is limited, and if you fill the context window too much, the quality of the answer gets degraded too.
- 3:46 The solution, what if we use four kegs in parallel and distribute the documents across different context buckets?
- 3:54 Now each keg can answer questions regarding its content.
- 3:57 And now we just need something to ask the right questions to the right buckets.
- 4:03 So for this, we can use a smarter model to interrogate each bucket and eventually synthesize an answer.
- 4:11 How do we distribute the documents?
- 4:13 It sounds tempting to organize the documents by domains and tell the supervisor, here are the different categories.
- 4:20 But in practice, with very dense relationship between documents, the supervisor tends to ignore domains that at first glance seem irrelevant.
- 4:30 For this reason, all documents are distributed in no particular order.
- 4:34 The only requirement is to balance the number of documents in a way that the least amount of documents are needed.
- 4:40 Then the supervisor model starts exploring the buckets and progressively builds its internal understanding.
- 4:47 And if it finds something interesting, it can ask a specific bucket follow-up questions.
- 4:52 Because all caches can be loaded in parallel, the knowledge building process is significantly faster than graph rag while providing more accurate answers than a simple rag.
loading