read-only demo

Videos XovaGv4f39A

When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:05:52· en-US· indexed 2026-08-11 07:54

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:01, 1 of 1 keyframes kept
  2. Shot 1, 0:01 to 0:10, 1 of 1 keyframes kept
  3. Shot 2, 0:10 to 0:23, 1 of 1 keyframes kept
  4. Shot 3, 0:23 to 0:33, 1 of 1 keyframes kept
  5. Shot 4, 0:33 to 0:37, 0 of 1 keyframes kept
  6. Shot 5, 0:37 to 0:39, 1 of 1 keyframes kept
  7. Shot 6, 0:39 to 0:42, 1 of 1 keyframes kept
  8. Shot 7, 0:42 to 0:47, 1 of 1 keyframes kept
  9. Shot 8, 0:47 to 0:52, 0 of 1 keyframes kept
  10. Shot 9, 0:52 to 0:54, 1 of 1 keyframes kept
  11. Shot 10, 0:54 to 0:57, 1 of 1 keyframes kept
  12. Shot 11, 0:57 to 1:02, 1 of 1 keyframes kept
  13. Shot 12, 1:02 to 1:04, 0 of 1 keyframes kept
  14. Shot 13, 1:04 to 1:06, 1 of 1 keyframes kept
  15. Shot 14, 1:06 to 1:08, 1 of 1 keyframes kept
  16. Shot 15, 1:08 to 1:10, 1 of 1 keyframes kept
  17. Shot 16, 1:10 to 1:12, 1 of 1 keyframes kept
  18. Shot 17, 1:12 to 1:14, 1 of 1 keyframes kept
  19. Shot 18, 1:14 to 1:20, 1 of 1 keyframes kept
  20. Shot 19, 1:20 to 1:22, 1 of 1 keyframes kept
  21. Shot 20, 1:22 to 1:25, 1 of 1 keyframes kept
  22. Shot 21, 1:25 to 1:26, 1 of 1 keyframes kept
  23. Shot 22, 1:26 to 1:30, 1 of 1 keyframes kept
  24. Shot 23, 1:30 to 1:33, 0 of 1 keyframes kept
  25. Shot 24, 1:33 to 1:34, 1 of 1 keyframes kept
  26. Shot 25, 1:34 to 1:41, 1 of 1 keyframes kept
  27. Shot 26, 1:41 to 1:43, 1 of 1 keyframes kept
  28. Shot 27, 1:43 to 1:46, 0 of 1 keyframes kept
  29. Shot 28, 1:46 to 1:48, 1 of 1 keyframes kept
  30. Shot 29, 1:48 to 1:51, 1 of 1 keyframes kept
  31. Shot 30, 1:51 to 1:53, 1 of 1 keyframes kept
  32. Shot 31, 1:53 to 1:56, 1 of 1 keyframes kept
  33. Shot 32, 1:56 to 2:10, 0 of 1 keyframes kept
  34. Shot 33, 2:10 to 2:15, 1 of 1 keyframes kept
  35. Shot 34, 2:15 to 2:17, 1 of 1 keyframes kept
  36. Shot 35, 2:17 to 2:20, 1 of 1 keyframes kept
  37. Shot 36, 2:20 to 2:22, 0 of 1 keyframes kept
  38. Shot 37, 2:22 to 2:25, 1 of 1 keyframes kept
  39. Shot 38, 2:25 to 2:27, 1 of 1 keyframes kept
  40. Shot 39, 2:27 to 2:31, 1 of 1 keyframes kept
  41. Shot 40, 2:31 to 2:32, 1 of 1 keyframes kept
  42. Shot 41, 2:32 to 2:35, 0 of 1 keyframes kept
  43. Shot 42, 2:35 to 2:39, 1 of 1 keyframes kept
  44. Shot 43, 2:39 to 2:40, 0 of 1 keyframes kept
  45. Shot 44, 2:40 to 2:49, 1 of 1 keyframes kept
  46. Shot 45, 2:49 to 2:50, 1 of 1 keyframes kept
  47. Shot 46, 2:50 to 3:24, 0 of 1 keyframes kept
  48. Shot 47, 3:24 to 3:27, 1 of 1 keyframes kept
  49. Shot 48, 3:27 to 3:32, 1 of 1 keyframes kept
  50. Shot 49, 3:32 to 3:34, 1 of 1 keyframes kept
  51. Shot 50, 3:34 to 3:48, 0 of 1 keyframes kept
  52. Shot 51, 3:48 to 3:51, 1 of 1 keyframes kept
  53. Shot 52, 3:51 to 3:52, 1 of 1 keyframes kept
  54. Shot 53, 3:52 to 3:58, 1 of 1 keyframes kept
  55. Shot 54, 3:58 to 4:02, 0 of 1 keyframes kept
  56. Shot 55, 4:02 to 4:05, 1 of 1 keyframes kept
  57. Shot 56, 4:05 to 4:07, 1 of 1 keyframes kept
  58. Shot 57, 4:07 to 4:12, 1 of 1 keyframes kept
  59. Shot 58, 4:12 to 4:17, 1 of 1 keyframes kept
  60. Shot 59, 4:17 to 4:21, 1 of 1 keyframes kept
  61. Shot 60, 4:21 to 4:24, 1 of 1 keyframes kept
  62. Shot 61, 4:24 to 4:27, 1 of 1 keyframes kept
  63. Shot 62, 4:27 to 4:30, 0 of 1 keyframes kept
  64. Shot 63, 4:30 to 4:33, 1 of 1 keyframes kept
  65. Shot 64, 4:33 to 4:37, 1 of 1 keyframes kept
  66. Shot 65, 4:37 to 4:40, 1 of 1 keyframes kept
  67. Shot 66, 4:40 to 4:51, 0 of 1 keyframes kept
  68. Shot 67, 4:51 to 4:55, 0 of 1 keyframes kept
  69. Shot 68, 4:55 to 5:02, 1 of 1 keyframes kept
  70. Shot 69, 5:02 to 5:05, 1 of 1 keyframes kept
  71. Shot 70, 5:05 to 5:16, 0 of 1 keyframes kept
  72. Shot 71, 5:16 to 5:34, 1 of 1 keyframes kept
  73. Shot 72, 5:34 to 5:41, 1 of 1 keyframes kept
  74. Shot 73, 5:41 to 5:46, 1 of 1 keyframes kept
  75. Shot 74, 5:46 to 5:51, 1 of 1 keyframes kept

75 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
53
whisperx 53
chunks
10
from 53 cues
keyframes
59
kept of 75 captured
frames with text
38
184 lines read
chapters
0
from the source metadata
keyframe bytes
8.1 MB
word timings on 53 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 07:52 1m 34s
stt done 2026-08-11 07:53 6s
chunk done 2026-08-11 07:53 0s
text_embed done 2026-08-11 07:53 0s
keyframe done 2026-08-11 07:53 48s
ocr done 2026-08-11 07:54 10s
frame_embed done 2026-08-11 07:54 9s

Frames, and what the machine read

  • 0:01 #0 empty

    shot 0·sharpness 103.7

  • 0:04 #1 done3 line(s)

    shot 1·sharpness 195.2

    1. When all context matters:1.00
    2. Extended Cache1.00
    3. Augmented Generation (ECAG)0.99
  • 0:16 #2 done1 line(s)

    shot 2·sharpness 176.7

    1. ORBIS1.00
  • 0:29 #3 done1 line(s)

    shot 3·sharpness 240.6

    1. X0.54
  • 0:34 #4 skipped

    shot 4·duplicate of #2

  • 0:37 #5 done1 line(s)

    shot 5·sharpness 444.4

    1. S0.94
  • 0:40 #6 empty

    shot 6·sharpness 290.8

  • 0:43 #7 empty

    shot 7·sharpness 208.9

  • 0:50 #8 skipped

    shot 8·duplicate of #2

  • 0:54 #9 empty

    shot 9·sharpness 33.5

  • 0:55 #10 done41 line(s)

    shot 10·sharpness 125.6

    1. 11.00
    2. 21.00
    3. 91.00
    4. 10.98
    5. 31.00
    6. 450.96
    7. 51.00
    8. 71.00
    9. 90.87
    10. 21.00
    11. 31.00
    12. 61.00
    13. 51.00
    14. 91.00
    15. B20.65
    16. 11.00
    17. 31.00
    18. 20.82
    19. 21.00
    20. 70.92
    21. 40.75
    22. 70.98
    23. 50.59
    24. a0.60
    25. 50.96
    26. 7050.57
    27. 00.62
    28. 4A0.75
    29. 890.98
    30. 21.00
    31. 590.92
    32. 46540.85
    33. 91.00
    34. 61.00
    35. 160.65
    36. 51.00
    37. 90.74
    38. 150.65
    39. 90.96
    40. 61.00
    41. 80.98
  • 0:59 #11 empty

    shot 11·sharpness 178.8

  • 1:02 #12 skipped

    shot 12·duplicate of #2

  • 1:05 #13 done8 line(s)

    shot 13·sharpness 258.0

    1. Apple1.00
    2. Fruit1.00
    3. Orange0.93
    4. Orange1.00
    5. Fruit1.00
    6. Train1.00
    7. Car1.00
    8. Train1.00
  • 1:06 #14 done7 line(s)

    shot 14·sharpness 184.1

    1. Apple1.00
    2. Fr1.00
    3. Orange1.00
    4. 011101001100010000.97
    5. Fruit1.00
    6. Tr1.00
    7. Car1.00
  • 1:10 #15 done31 line(s)

    shot 15·sharpness 246.2

    1. -3.69360.95
    2. 1600.99
    3. 1800.95
    4. -1.89800.94
    5. 1601.00
    6. 1401.00
    7. 1201.00
    8. -0.58800.82
    9. Apple0.91
    10. 1001.00
    11. 01.00
    12. -500.99
    13. 301.00
    14. -11.8330.88
    15. 0.39880.91
    16. 201.00
    17. -5.7330.89
    18. 21.00
    19. 31.00
    20. 41.00
    21. 51.00
    22. 61.00
    23. 71.00
    24. 81.00
    25. 91.00
    26. 101.00
    27. 111.00
    28. 131.00
    29. 131.00
    30. 141.00
    31. X0.64
  • 1:10 #16 done3 line(s)

    shot 16·sharpness 153.4

    1. Apple1.00
    2. 6138436491.00
    3. 296702109555533937435925679320.97
  • 1:13 #17 done18 line(s)

    shot 17·sharpness 362.4

    1. 50001.00
    2. 57501.00
    3. 46501.00
    4. 35001.00
    5. Orbi1.00
    6. X0.68
    7. 33501.00
    8. 21001.00
    9. 10001.00
    10. -1001.00
    11. 11.00
    12. 21.00
    13. 31.00
    14. 41.00
    15. 41.00
    16. 51.00
    17. 61.00
    18. 7 8 9 10 11 12 13 16 29 X0.98
  • 1:18 #18 done1 line(s)

    shot 18·sharpness 190.9

    1. ORBIS0.99
  • 1:21 #19 empty

    shot 19·sharpness 285.3

  • 1:24 #20 empty

    shot 20·sharpness 291.6

  • 1:26 #21 empty

    shot 21·sharpness 205.0

  • 1:29 #22 done1 line(s)

    shot 22·sharpness 210.1

    1. Context Window0.98
  • 1:32 #23 skipped

    shot 23·duplicate of #2

Transcript

53 cues· 845 words· 5,000 chars

  1. 0:10 I'm Luis Romero Sevilla and I'm the VP of AI at Orbis Operations.
  2. 0:15 I'm on a mission to solve knowledge representation when all context matters.
  3. 0:19 So, let's start with a very specific example.
  4. 0:24 Let's say we have a large number of documents.
  5. 0:26 And all documents represent an event.
  6. 0:28 And all documents in the collection are relevant to answer a set of questions that the user has.
  7. 0:34 Not only that, there's one more challenge.
  8. 0:38 the document in the collection becomes obsolete very fast and all documents get replaced with new information.
  9. 0:44 Let's start with the simplest approach.
  10. 0:46 We could start with a simple reg.
  11. 0:48 For that we just need a vector database and an embedding model.
  12. 0:53 An embedding model takes the documents and turns them into a learned numerical representation, a vector.
  13. 1:00 Now we take those vectors and we store them in a database optimized for performing operations with vectors.
  14. 1:07 Perfect.
  15. 1:08 Now we can take all of our questions, turn them into vectors, and then look for vectors that are similar to the initial query.
  16. 1:18 Those vectors that are within the similarity threshold are retrieved, and we can pass them to the LLM to answer the question.
  17. 1:27 Inserting to a vector database, it's relatively fast.
  18. 1:30 So whenever a collection becomes obsolete, we can just replace it with a new one.
  19. 1:36 We still have one problem with our very specific scenario.
  20. 1:41 All the documents in the collection are relevant for us to answer the question.
  21. 1:45 So we can't just take all the documents in the collection and pass them to LLM.
  22. 1:51 And that's just one of the many limitations with this approach.
  23. 1:55 Now let's get a bit more sophisticated.
  24. 1:58 All documents are relevant to answer a global question.
  25. 2:01 Therefore, there must be some connections and relationship between the details within a document in the collections.
  26. 2:08 For us to map out those relationship, we're going to need a knowledge graph.
  27. 2:12 And one implementation we could try is GraphRag.
  28. 2:16 GraphRag had many stats, but basically it uses an LLM to read through all the documents and extract key entities and relationship between them.
  29. 2:27 It constructs a network, a knowledge graph, where all those connections and details are tied together.
  30. 2:32 Then when a question is asked, it navigates this graph to synthesize a complete answer drawn across the entire collection.
  31. 2:39 If your collection of documents isn't changing very often, GraphRack is an excellent approach for finding those relationships within details to answer the user's question.
  32. 2:51 However, our very specific scenario states that our data is not only deeply interconnected, but also the data gets replaced very often.
  33. 3:01 Recomputing a knowledge graph every time the data gets replaced is computationally very expensive and it takes a relatively long time.
  34. 3:08 Okay, what if we take an even simpler approach and we continue to build on top of it?
  35. 3:13 If we were to use something like GraphRack, each document needs to pass through an LLM for the entity and relationship structure anyway.
  36. 3:21 Why can't we just throw all the documents into context?
  37. 3:24 This approach will look something like cache amended generation, CAG, where we use a model with a large context window, load the documents into the context, and cache the context by storing the model's KB matrix.
  38. 3:38 The problem here is that the context window is limited, and if you fill the context window too much, the quality of the answer gets degraded too.
  39. 3:46 The solution, what if we use four kegs in parallel and distribute the documents across different context buckets?
  40. 3:54 Now each keg can answer questions regarding its content.
  41. 3:57 And now we just need something to ask the right questions to the right buckets.
  42. 4:03 So for this, we can use a smarter model to interrogate each bucket and eventually synthesize an answer.
  43. 4:11 How do we distribute the documents?
  44. 4:13 It sounds tempting to organize the documents by domains and tell the supervisor, here are the different categories.
  45. 4:20 But in practice, with very dense relationship between documents, the supervisor tends to ignore domains that at first glance seem irrelevant.
  46. 4:30 For this reason, all documents are distributed in no particular order.
  47. 4:34 The only requirement is to balance the number of documents in a way that the least amount of documents are needed.
  48. 4:40 Then the supervisor model starts exploring the buckets and progressively builds its internal understanding.
  49. 4:47 And if it finds something interesting, it can ask a specific bucket follow-up questions.
  50. 4:52 Because all caches can be loaded in parallel, the knowledge building process is significantly faster than graph rag while providing more accurate answers than a simple rag.

Open at this second