read-only demo

Videos -x5GEVnkuRw

Structuring the Unstructured - Cedric Clyburn, Red Hat

index_state ready data_status ok

AI Engineer· published 2026-06-28· 0:20:40· en-US· indexed 2026-08-11 08:15

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:27, 1 of 1 keyframes kept
  2. Shot 1, 0:27 to 0:55, 0 of 1 keyframes kept
  3. Shot 2, 0:55 to 1:12, 1 of 1 keyframes kept
  4. Shot 3, 1:12 to 1:25, 1 of 1 keyframes kept
  5. Shot 4, 1:25 to 1:55, 1 of 1 keyframes kept
  6. Shot 5, 1:55 to 2:06, 1 of 1 keyframes kept
  7. Shot 6, 2:06 to 2:09, 1 of 1 keyframes kept
  8. Shot 7, 2:09 to 2:26, 1 of 1 keyframes kept
  9. Shot 8, 2:26 to 2:57, 1 of 1 keyframes kept
  10. Shot 9, 2:57 to 3:22, 1 of 1 keyframes kept
  11. Shot 10, 3:22 to 3:42, 1 of 1 keyframes kept
  12. Shot 11, 3:42 to 3:57, 1 of 1 keyframes kept
  13. Shot 12, 3:57 to 4:45, 1 of 1 keyframes kept
  14. Shot 13, 4:45 to 4:53, 1 of 1 keyframes kept
  15. Shot 14, 4:53 to 5:40, 1 of 1 keyframes kept
  16. Shot 15, 5:40 to 6:11, 1 of 1 keyframes kept
  17. Shot 16, 6:11 to 6:42, 0 of 1 keyframes kept
  18. Shot 17, 6:42 to 7:07, 1 of 1 keyframes kept
  19. Shot 18, 7:07 to 7:32, 1 of 1 keyframes kept
  20. Shot 19, 7:32 to 7:59, 1 of 1 keyframes kept
  21. Shot 20, 7:59 to 8:25, 0 of 1 keyframes kept
  22. Shot 21, 8:25 to 9:03, 1 of 1 keyframes kept
  23. Shot 22, 9:03 to 9:39, 1 of 1 keyframes kept
  24. Shot 23, 9:39 to 9:45, 1 of 1 keyframes kept
  25. Shot 24, 9:45 to 10:05, 1 of 1 keyframes kept
  26. Shot 25, 10:05 to 10:07, 1 of 1 keyframes kept
  27. Shot 26, 10:07 to 10:10, 1 of 1 keyframes kept
  28. Shot 27, 10:10 to 10:12, 1 of 1 keyframes kept
  29. Shot 28, 10:12 to 10:18, 1 of 1 keyframes kept
  30. Shot 29, 10:18 to 10:20, 1 of 1 keyframes kept
  31. Shot 30, 10:20 to 10:23, 1 of 1 keyframes kept
  32. Shot 31, 10:23 to 10:31, 1 of 1 keyframes kept
  33. Shot 32, 10:31 to 10:34, 1 of 1 keyframes kept
  34. Shot 33, 10:34 to 10:36, 1 of 1 keyframes kept
  35. Shot 34, 10:36 to 10:45, 1 of 1 keyframes kept
  36. Shot 35, 10:45 to 10:48, 1 of 1 keyframes kept
  37. Shot 36, 10:48 to 10:51, 1 of 1 keyframes kept
  38. Shot 37, 10:51 to 10:54, 1 of 1 keyframes kept
  39. Shot 38, 10:54 to 11:00, 1 of 1 keyframes kept
  40. Shot 39, 11:00 to 11:04, 1 of 1 keyframes kept
  41. Shot 40, 11:04 to 11:08, 1 of 1 keyframes kept
  42. Shot 41, 11:08 to 11:10, 1 of 1 keyframes kept
  43. Shot 42, 11:10 to 11:13, 1 of 1 keyframes kept
  44. Shot 43, 11:13 to 11:15, 1 of 1 keyframes kept
  45. Shot 44, 11:15 to 11:19, 1 of 1 keyframes kept
  46. Shot 45, 11:19 to 11:20, 1 of 1 keyframes kept
  47. Shot 46, 11:20 to 11:30, 1 of 1 keyframes kept
  48. Shot 47, 11:30 to 11:34, 1 of 1 keyframes kept
  49. Shot 48, 11:34 to 11:36, 1 of 1 keyframes kept
  50. Shot 49, 11:36 to 11:43, 1 of 1 keyframes kept
  51. Shot 50, 11:43 to 11:46, 1 of 1 keyframes kept
  52. Shot 51, 11:46 to 11:49, 1 of 1 keyframes kept
  53. Shot 52, 11:49 to 11:55, 1 of 1 keyframes kept
  54. Shot 53, 11:55 to 12:02, 1 of 1 keyframes kept
  55. Shot 54, 12:02 to 12:05, 1 of 1 keyframes kept
  56. Shot 55, 12:05 to 12:10, 1 of 1 keyframes kept
  57. Shot 56, 12:10 to 12:13, 1 of 1 keyframes kept
  58. Shot 57, 12:13 to 12:15, 1 of 1 keyframes kept
  59. Shot 58, 12:15 to 12:18, 1 of 1 keyframes kept
  60. Shot 59, 12:18 to 12:22, 1 of 1 keyframes kept
  61. Shot 60, 12:22 to 12:25, 1 of 1 keyframes kept
  62. Shot 61, 12:25 to 12:29, 1 of 1 keyframes kept
  63. Shot 62, 12:29 to 12:31, 1 of 1 keyframes kept
  64. Shot 63, 12:31 to 12:35, 0 of 1 keyframes kept
  65. Shot 64, 12:35 to 12:46, 1 of 1 keyframes kept
  66. Shot 65, 12:46 to 12:48, 1 of 1 keyframes kept
  67. Shot 66, 12:48 to 12:53, 1 of 1 keyframes kept
  68. Shot 67, 12:53 to 12:57, 1 of 1 keyframes kept
  69. Shot 68, 12:57 to 12:59, 1 of 1 keyframes kept
  70. Shot 69, 12:59 to 13:01, 1 of 1 keyframes kept
  71. Shot 70, 13:01 to 13:09, 1 of 1 keyframes kept
  72. Shot 71, 13:09 to 13:13, 1 of 1 keyframes kept
  73. Shot 72, 13:13 to 13:20, 1 of 1 keyframes kept
  74. Shot 73, 13:20 to 13:23, 1 of 1 keyframes kept
  75. Shot 74, 13:23 to 13:34, 1 of 1 keyframes kept
  76. Shot 75, 13:34 to 13:35, 1 of 1 keyframes kept
  77. Shot 76, 13:35 to 13:38, 1 of 1 keyframes kept
  78. Shot 77, 13:38 to 13:39, 1 of 1 keyframes kept
  79. Shot 78, 13:39 to 13:45, 1 of 1 keyframes kept
  80. Shot 79, 13:45 to 13:56, 1 of 1 keyframes kept
  81. Shot 80, 13:56 to 13:57, 1 of 1 keyframes kept
  82. Shot 81, 13:57 to 14:04, 1 of 1 keyframes kept
  83. Shot 82, 14:04 to 14:33, 1 of 1 keyframes kept
  84. Shot 83, 14:33 to 14:36, 1 of 1 keyframes kept
  85. Shot 84, 14:36 to 14:47, 1 of 1 keyframes kept
  86. Shot 85, 14:47 to 15:07, 1 of 1 keyframes kept
  87. Shot 86, 15:07 to 15:15, 1 of 1 keyframes kept
  88. Shot 87, 15:15 to 15:17, 1 of 1 keyframes kept
  89. Shot 88, 15:17 to 15:23, 1 of 1 keyframes kept
  90. Shot 89, 15:23 to 15:27, 1 of 1 keyframes kept
  91. Shot 90, 15:27 to 15:30, 1 of 1 keyframes kept
  92. Shot 91, 15:30 to 15:33, 1 of 1 keyframes kept
  93. Shot 92, 15:33 to 15:35, 1 of 1 keyframes kept
  94. Shot 93, 15:35 to 15:38, 1 of 1 keyframes kept
  95. Shot 94, 15:38 to 15:43, 1 of 1 keyframes kept
  96. Shot 95, 15:43 to 15:54, 1 of 1 keyframes kept
  97. Shot 96, 15:54 to 15:56, 1 of 1 keyframes kept
  98. Shot 97, 15:56 to 16:02, 1 of 1 keyframes kept
  99. Shot 98, 16:02 to 16:03, 1 of 1 keyframes kept
  100. Shot 99, 16:03 to 16:08, 1 of 1 keyframes kept
  101. Shot 100, 16:08 to 16:10, 1 of 1 keyframes kept
  102. Shot 101, 16:10 to 16:12, 1 of 1 keyframes kept
  103. Shot 102, 16:12 to 16:14, 1 of 1 keyframes kept
  104. Shot 103, 16:14 to 16:22, 1 of 1 keyframes kept
  105. Shot 104, 16:22 to 16:23, 1 of 1 keyframes kept
  106. Shot 105, 16:23 to 16:25, 1 of 1 keyframes kept
  107. Shot 106, 16:25 to 16:27, 1 of 1 keyframes kept
  108. Shot 107, 16:27 to 16:30, 1 of 1 keyframes kept
  109. Shot 108, 16:30 to 16:35, 1 of 1 keyframes kept
  110. Shot 109, 16:35 to 16:42, 1 of 1 keyframes kept
  111. Shot 110, 16:42 to 16:55, 1 of 1 keyframes kept
  112. Shot 111, 16:55 to 16:56, 1 of 1 keyframes kept
  113. Shot 112, 16:56 to 17:00, 1 of 1 keyframes kept
  114. Shot 113, 17:00 to 17:01, 1 of 1 keyframes kept
  115. Shot 114, 17:01 to 17:04, 1 of 1 keyframes kept
  116. Shot 115, 17:04 to 17:07, 1 of 1 keyframes kept
  117. Shot 116, 17:07 to 17:10, 1 of 1 keyframes kept
  118. Shot 117, 17:10 to 17:31, 1 of 1 keyframes kept
  119. Shot 118, 17:31 to 17:33, 1 of 1 keyframes kept
  120. Shot 119, 17:33 to 17:36, 1 of 1 keyframes kept
  121. Shot 120, 17:36 to 17:41, 1 of 1 keyframes kept
  122. Shot 121, 17:41 to 17:43, 1 of 1 keyframes kept
  123. Shot 122, 17:43 to 18:05, 1 of 1 keyframes kept
  124. Shot 123, 18:05 to 18:07, 1 of 1 keyframes kept
  125. Shot 124, 18:07 to 18:10, 1 of 1 keyframes kept
  126. Shot 125, 18:10 to 18:17, 1 of 1 keyframes kept
  127. Shot 126, 18:17 to 18:19, 1 of 1 keyframes kept
  128. Shot 127, 18:19 to 18:24, 1 of 1 keyframes kept
  129. Shot 128, 18:24 to 18:27, 1 of 1 keyframes kept
  130. Shot 129, 18:27 to 18:29, 1 of 1 keyframes kept
  131. Shot 130, 18:29 to 18:31, 1 of 1 keyframes kept
  132. Shot 131, 18:31 to 18:34, 1 of 1 keyframes kept
  133. Shot 132, 18:34 to 18:37, 1 of 1 keyframes kept
  134. Shot 133, 18:37 to 18:42, 1 of 1 keyframes kept
  135. Shot 134, 18:42 to 18:49, 1 of 1 keyframes kept
  136. Shot 135, 18:49 to 18:54, 1 of 1 keyframes kept
  137. Shot 136, 18:54 to 18:57, 1 of 1 keyframes kept
  138. Shot 137, 18:57 to 19:00, 1 of 1 keyframes kept
  139. Shot 138, 19:00 to 19:02, 1 of 1 keyframes kept
  140. Shot 139, 19:02 to 19:09, 1 of 1 keyframes kept
  141. Shot 140, 19:09 to 19:28, 1 of 1 keyframes kept
  142. Shot 141, 19:28 to 19:43, 1 of 1 keyframes kept
  143. Shot 142, 19:43 to 20:03, 1 of 1 keyframes kept
  144. Shot 143, 20:03 to 20:10, 1 of 1 keyframes kept
  145. Shot 144, 20:10 to 20:39, 1 of 1 keyframes kept
  146. Shot 145, 20:39 to 20:40, 0 of 1 keyframes kept

146 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
174
whisperx 174
chunks
36
from 174 cues
keyframes
141
kept of 146 captured
frames with text
141
10,196 lines read
chapters
0
from the source metadata
keyframe bytes
23.8 MB
word timings on 174 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 08:07 1m 15s
stt done 2026-08-11 08:08 23s
chunk done 2026-08-11 08:09 0s
text_embed done 2026-08-11 08:09 1s
keyframe done 2026-08-11 08:09 1m 51s
ocr done 2026-08-11 08:11 3m 50s
frame_embed done 2026-08-11 08:14 28s

Frames, and what the machine read

  • 0:16 #0 done10 line(s)

    shot 0·sharpness 2334.7

    1. AlEngineer0.99
    2. Red Hat0.99
    3. World's Fair0.97
    4. Developer1.00
    5. Structuring the Unstructured: Advanced1.00
    6. Document Parsing for Al Workflows0.99
    7. AI Engineer: 20260.97
    8. Cedric Clyburn0.98
    9. Senior Developer Advocate1.00
    10. @cedricclyburnn0.98
  • 0:36 #1 skipped

    shot 1·duplicate of #0

  • 1:07 #2 done12 line(s)

    shot 2·sharpness 1887.3

    1. Unstructured Data is the Context of Al0.99
    2. 100's of Zettabytes Per Year of Unstructured Data – Growing Exponentially0.99
    3. Enterprises1.00
    4. FSI0.82
    5. RBC0.83
    6. EnterpriseSoftware1.00
    7. Storage Platform0.99
    8. Enterprise Databases1.00
    9. cSP Engines0.98
    10. OSS Engines0.92
    11. cuVS1.00
    12. Source: IDC Global DataSphere, 20250.99
  • 1:22 #3 done5 line(s)

    shot 3·sharpness 1064.8

    1. We've got a lot to cover today!1.00
    2. PDF1.00
    3. Wait, so 85% of the0.99
    4. world's data is...1.00
    5. unstructured?!1.00
  • 1:43 #4 done23 line(s)

    shot 4·sharpness 4256.6

    1. We've got a lot to cover today!0.98
    2. Azure Al Document Intelligence0.97
    3. 0.84
    4. $1.00
    5. Accelerate information extraction from documents.0.98
    6. Amazon Textract1.00
    7. PDF1.00
    8. LlamaParse: Transform unstructured1.00
    9. But current solutions are0.98
    10. data into LLM optimized formats0.99
    11. Automatically extract printed text, handwriting, layout elements, and1.00
    12. data from any document0.99
    13. proprietary, and require1.00
    14. sending your private data!0.98
    15. Wait, so 85% of the0.99
    16. world's data is...1.00
    17. unstructured?!1.00
    18. let's learn about0.97
    19. extraction, parsing,1.00
    20. chunking, and much more!1.00
    21. docling1.00
    22. So, how can we easily parse d0.98
    23. graphs, tables, etc to formats1.00
  • 2:01 #5 done19 line(s)

    shot 5·sharpness 3170.6

    1. Agenda1.00
    2. Today's Schedule1.00
    3. Data Preparation for1.00
    4. Demo #1:0.94
    5. Al: It's not easy!1.00
    6. Extracting/Parsing1.00
    7. Unstructured Data1.00
    8. Parsing PDF's,0.99
    9. tables, images, etc.1.00
    10. Demo #2: Chunking &1.00
    11. Embedding1.00
    12. Building an Al0.98
    13. Session Slides0.99
    14. document pipeline1.00
    15. Demo #3: Building &0.98
    16. red.ht/structuring1.00
    17. Deploying a Q&A app!1.00
    18. Red Hat0.98
    19. Developer1.00
  • 2:08 #6 done5 line(s)

    shot 6·sharpness 1219.0

    1. Why's there a need for0.97
    2. advanced document1.00
    3. processing?1.00
    4. Red Hat0.88
    5. Developer1.00
  • 2:22 #7 done53 line(s)

    shot 7·sharpness 1927.0

    1. Data is the key ingredient behind Al applications!1.00
    2. Table of contets0.90
    3. Board Meeting1.00
    4. 0.58
    5. Chapter 2. Creating an Amazon S3 client using1.00
    6. Jine 24 03 0 AM 1:300 AM0.67
    7. notebook cells1.00
    8. ATTENDEES:0.93
    9. HOST0.91
    10. ane Rotrguer0.69
    11. ▲ POF0.89
    12. that service.0.97
    13. 900 AM- 9:300 AME0.77
    14. 1.1 0paring0.85
    15. Teple0.86
    16. YartaAArata0.70
    17. Presecter0.82
    18. Jose Rodripuat0.84
    19. 0.51
    20. Prerequisites0.90
    21. • Access o a Jupyter notebok server runng on Red Hat penShit Al.0.93
    22. 1.2 Atandance0.88
    23. Dvdn maters0.60
    24. Financial1.00
    25. Defi us or t A an vie varables0.58
    26. when you star your notebook server sing the values from your Amazon Web Services0.95
    27. 1.3 Appreval of Agenda0.90
    28. Jef Yon0.64
    29. Documents1.00
    30. er Portal0.99
    31. Procedure0.98
    32. 1 In new notebook el, iprt th equied irarie y dding te folowing:0.61
    33. account under My Security Credentials.0.95
    34. e0.66
    35. 9.30 10 1 0 . Revew of Perio Monues0.52
    36. 2.1 Aproetlt telete Jof Kaih0.51
    37. Yerta Amata0.76
    38. How to control the access of podman to system0.99
    39. users1.00
    40. inpert botal0.66
    41. fron boto3 lapert session0.81
    42. Jone Rodrigunz0.75
    43. Environment1.00
    44. 2.2 Cloing trsteages0.68
    45. Jone Rodiguit0.59
    46. Issue1.00
    47. i Define your credendials.0.93
    48. Meeting Minutes1.00
    49. Technical1.00
    50. Resolution1.00
    51. Documentation1.00
    52. Knowledge1.00
    53. Articles1.00
  • 2:50 #8 done51 line(s)

    shot 8·sharpness 2861.8

    1. Data is the key ingredient behind Al applications!1.00
    2. Board Meeting1.00
    3. ATTENODEES0.84
    4. Powering:1.00
    5. Table of conteents0.90
    6. 00.830 AM0.68
    7. Jonn Rodigu0.63
    8. RAG (Document Q&A)0.97
    9. Chapter 2. Creating an Amazon S3 client using0.98
    10. notebook cells1.00
    11. Fine-Tuning1.00
    12. 4 PDF0.84
    13. 1.3 Aporoval of Agante0.78
    14. To interactwth datain mazon S3 buckets, youmust reste alocal lient t ande requests to0.76
    15. that service.0.97
    16. rowelldt Prolous Woute0.53
    17. Tewrtoa Amata0.69
    18. Financial1.00
    19. etc.1.00
    20. Prerequisites0.93
    21. Documents1.00
    22. •Access t a Jupyter notebook server rnning on Red Hat OpenShift A Al.0.89
    23. Defin vaue t an T A v vaables0.56
    24. when you start our notebok erver using hevalues from you Amazon Web Sevies0.83
    25. or Portal0.89
    26. account under My Security Credentials.0.76
    27. Meeting Minutes1.00
    28. Procedure1.00
    29. How to control the access of podman to system1.00
    30. I In a new notebook cel, mport the required libraries by adding the folowing:0.90
    31. e0.74
    32. users1.00
    33. Hugging Face1.00
    34. Environment1.00
    35. NVIDIA.0.96
    36. fron botu3 (aport session0.84
    37. databricks1.00
    38. 2. In another new notebok cel, define the folowing to create your session and clent.0.95
    39. Issue1.00
    40. 0Meta0.86
    41. L. Define your credentials.0.96
    42. Technical1.00
    43. Resolution1.00
    44. Mic0.99
    45. Documentation1.00
    46. + much more!1.00
    47. Google1.00
    48. Knowledge Base0.99
    49. Articles1.00
    50. MISTRAL1.00
    51. AI_0.98
  • 3:16 #9 done56 line(s)

    shot 9·sharpness 1388.7

    1. Data processing & prep is quite important!0.98
    2. gurovdigital15 h0.99
    3. ..0.83
    4. lol, over 20 scientific papers now feature the1.00
    5. nonsensical term 'vegetative electron0.99
    6. microscopy'.1.00
    7. "vegetative electron microsct ×0.95
    8. all because an Al misinterpreted a 1959 article,0.99
    9. merging 'vegetative' and 'electron microscopy'0.99
    10. Scholar1.00
    11. YEAR -0.85
    12. from separate columns.1.00
    13. hydrophila and Yersinia ruckeri bacteria isolated0.93
    14. Study of CNT@ Fe304 effects on Aeromonas0.99
    15. from fish.1.00
    16. on of the mmayme fmin 8.0.72
    17. spore moala of 8.0.63
    18. M Alshai Taae A.Minastal. .Joumnai of .20190.55
    19. search.ebscohost.com0.97
    20. ie ensyme did not attack0.86
    21. Norris of Leeds University0.97
    22. carbon nenocubes syrtheszed by spectroscopic and0.93
    23. tion). He treated spores0.98
    24. s in the vegetative cell,0.97
    25. preparation of lytic enzy0.96
    26. à sporangium. It is by no0.98
    27. spores and examined the0.97
    28. pore is released. In Clos-1.00
    29. ears that at lenst part of0.97
    30. happens to the vegetative1.00
    31. hed as an outer membrane0.99
    32. electryu mieroscopy. No ev0.96
    33. exospcsium was obtained.0.95
    34. in spores, or another enzy0.98
    35. for lysin of the sporangial0.90
    36. It was not known whethe0.99
    37. Green synthesis of silver nanoparticles via0.99
    38. Ganoderma lucidum fungus extract and its0.98
    39. antibacteraleffects on Klebsiellapneumonia0.91
    40. isolates from urinary tract ..0.97
    41. M. JamsnidianMoaver, M A-Aborz Untversity. 20210.67
    42. Vegettve electon microscopy was0.75
    43. R analysin was also0.93
    44. a to measure the0.85
    45. 6921.00
    46. Q120.99
    47. 431.00
    48. ☆ Cited by 1 Related aricles0.87
    49. performed to investigate possble organic compounos thet0.90
    50. METALLOGRAPHIC STUDIES OF1.00
    51. [POF| res0.79
    52. BRONZE PIECES FROM JEYRÄN TEPE, OZBAKI0.97
    53. IRAN'S IRON AGE: CASE STUDY0.98
    54. ESODAEL H RA-NEMA - reseorchgate.net0.90
    55. This stuty in a report of the resalts of metalogrphic stady of0.92
    56. Sb ps foud in Jn Tee datig bac e lron0.58
  • 3:28 #10 done80 line(s)

    shot 10·sharpness 3099.9

    1. Data processing & prep is quite important!0.98
    2. gurovdigital15 h0.99
    3. 0.97
    4. Date syrup (as one of the agricultural wastes)1.00
    5. lol, over 20 scientific papers now feature the0.99
    6. was used to produce bacterial cellulose using1.00
    7. nonsensical term 'vegetative electron0.99
    8. microscopy'.1.00
    9. Gluconastobacter xylinus. Fourier transform0.99
    10. all because an Al misinterpreted a 1959 article,1.00
    11. "vegetative electron microsct ×0.95
    12. infrared spectroscopy (FTIR), vegetative electron1.00
    13. merging 'vegetative' and 'electron microscopy'0.98
    14. from separate columns.1.00
    15. hydrophila and Yersinia ruckeri bacteria isolated0.96
    16. from fish.0.95
    17. Study of CNT@ Fe304 effects on Aeromonas0.99
    18. Scholar0.99
    19. YEAR·0.91
    20. microscopy, and X-ray diffraction were used to1.00
    21. determine the structure of bacterial cellulose,0.99
    22. M Alsha KR Taae AMinastf .Joumnal o.2019-0.53
    23. on of the emaytue from 8.0.79
    24. search.ebacohost.com0.99
    25. cellulose fibers, and crystallinity of the samples1.00
    26. ie enayme did not attack0.87
    27. Norris of Leeds University0.97
    28. carbon nanetubes synthesized by spectroscopic and0.97
    29. tion). He treated spores0.97
    30. (Moosavi and Yousefi, 2011). After 14 days of1.00
    31. s in the vegetative cell,1.00
    32. preparation of lytic enzy0.99
    33. à sporangium. It is by no0.96
    34. happens to the vegetative0.98
    35. spores and examined the0.97
    36. electron mieroscopy. No ev0.99
    37. incubation at 28 °C, the highest yield of cellulose0.99
    38. pore is released. In Clos-0.99
    39. ears that at lenst part of0.98
    40. hed as an outer membrane0.98
    41. exosporium was obtained.1.00
    42. in spores, or another enzy0.98
    43. for lysin of the sporangial0.92
    44. It was not known whethe1.00
    45. Green synthesis of silver nanoparticles via1.00
    46. Ganoderma lucidum fungus extract and its1.00
    47. antibacteral ffects on Klebsiella pneumonia0.91
    48. isolates from urinary tract..0.97
    49. M Jamshidian-Mopaver. MAi.- Alborz Untversity. 2021 -10.79
    50. Vegetattve electon micoscopy as0.81
    51. R analysin was also0.94
    52. d to measure the0.87
    53. antimicrobial purposes against1.00
    54. Silver and gold nanoparticles for1.00
    55. [HTML] m0.96
    56. 6921.00
    57. Q121.00
    58. 431.00
    59. ☆ Chted by 1 Related aricies 80.81
    60. multi-drug resistance bacteria1.00
    61. METALLOGRAPHIC STUDIES OF0.99
    62. [POF] res0.84
    63. BRONZE PIECES FROM JEYRÄN TEPE, OZBAKI0.99
    64. IRAN'S IRON AGE: CASE STUDY0.98
    65. N Rabiee, S Ahmadi, O Akhavan, R Lugue - Materials, 2022 -0.98
    66. SODAEL HRA-NEM - reseorchgute.net0.71
    67. mdpi.com1.00
    68. This stul s a reort of the resut of metalographic stady of0.72
    69. S bronze pleces found in Jeyn Tepe dating back to the ron0.83
    70. ... Dead bacteria have been observed by imaging and1.00
    71. elem1.00
    72. ntai analysis using transmission erect0.97
    73. on mic0.94
    74. (TEM1.00
    75. vegetative electron microscopy, and1.00
    76. dEDX (0.90
    77. Micro0.97
    78. 0.99
    79. Cited by 1121.00
    80. telated articles0.99
  • 3:54 #11 done37 line(s)

    shot 11·sharpness 4888.4

    1. Data processing & prep is quite important!0.98
    2. were incuvateu witni an extract iromn spores uis-0.91
    3. acteristic type. 1t was conciuueu tnat at reast0.80
    4. integrated at pH 7.0. Peptide was released0.98
    5. part of the sporangial wall was dissolved away0.98
    6. which established that the coats contained sub-0.99
    7. to allow release of the spore. It appears likely0.98
    8. strate for the lytic enzyme present in spores.1.00
    9. that the exosporium of B. cereus does not have1.00
    10. Peptide was also released from spore coats of B.1.00
    11. a composition similar to that of the vegetative1.00
    12. megaterum by the action of the enzyme from B.1.00
    13. cell wall, from the results obtained by Dr. J. R.1.00
    14. cereus spores. The lytic enzyme did not attack1.00
    15. intact resting spor0.99
    16. The spore develops in the vegetative cell, which thus becomes a sporangium. It is by no means certain what happens to the1.00
    17. The spore develops in the vegetative cell,1.00
    18. vegetative cell wall when the spore is released. In Clostridium species it appears that at least part of this structure is retained as an0.99
    19. which thus becomes a sporangium. It is by no1.00
    20. outer membrane around the spore. It is the opinion of some workers that the wall of the sporulating cell forms the exosporium which0.99
    21. exists as an outer coat around spores of several Bacilus species. Spores of several varieties of B. cereus had exosporia whereas these1.00
    22. cell wall when the spore is released. In Clos-0.99
    23. structures appeared to be absent from spores of B. megaterium and B. subtilis. It seems, however, that in Bacillus species at least,0.99
    24. tridium species it appears that at least part of1.00
    25. the greater part of the vegetative cell wall is dissolved away before the developed spore is released. If this is true, then soluble0.99
    26. around the spore. It is the opinion of some1.00
    27. this structure is retained as an outer membrane1.00
    28. components containing the characteristic constituents should appear in the medium during spore release. Culture filtrates from B.0.99
    29. cereus organisms at various stages of growth and sporulation were hydrolyzed and the hydrolyzates analyzed for amino sugars and0.99
    30. diaminopimelic acid (28). Results showed that a large increase in the concentration of these substances in the culture filtrate0.99
    31. workers that the wall of the sporulating cell1.00
    32. occurred during spore release (table 2); they were found to be present in a nondialyzable peptide of the characteristic type. It was0.99
    33. forms the exosporium which exists as an outer0.98
    34. concluded that at least part of the sporangial wall was dissolved away to allow release of the spore. It appears likely that the1.00
    35. exosporium of B. cereus does not have a composition similar to that of the vegetative cell wall, from the results obtained by Dr. J. R.0.99
    36. Norris of Leeds University (personal communication). He treated spores with a highly active preparation of lytic enzyme from0.99
    37. cereus spores and examined the effect by means of electron microscopy. No evidence of lysis of the exosporium was obta1.00
  • 4:03 #12 done39 line(s)

    shot 12·sharpness 1746.9

    1. So, let's try a simple PDF parser...0.99
    2. KDD '22, August 14-18, 2022, Washington, DC, USA Birgit Pfitzmann, Christoph Auer, Michel0.98
    3. Nassar, and Peter Staar1.00
    4. Table 1: DocLayNet dataset overview. Along with the frequency of each class label, we prese0.98
    5. occurrence (as %0.99
    6. of row "Total") in the train, test and validation sets. The inter-annotator agreement is com0.98
    7. [email protected] metric1.00
    8. 10.55
    9. between pairwise annotations from the triple-annotated pages, from which we obtain accu0.98
    10. Very fast and cheap0.98
    11. % of Total1.00
    12. triple inter-annotator mAP @ 0.5-0.95 (%)0.99
    13. X Incomplete0.93
    14. [.-]0.58
    15. Count1.00
    16. 225241.00
    17. X Loss of structure0.98
    18. 63181.00
    19. 250271.00
    20. 1856601.00
    21. ×Noisy0.97
    22. 708781.00
    23. 580221.00
    24. 1428841.00
    25. 459761.00
    26. Unfit for most use1.00
    27. 5103771.00
    28. 347331.00
    29. 11074701.00
    30. 50711.00
    31. cases1.00
    32. [..]0.77
    33. include publication repositories such as arXiv3, government o"ces,0.98
    34. company websites as well as data directory services for #nancial1.00
    35. reports and patents. Scanned documents were excluded wherever0.99
    36. possible because they can be rotated or skewed. This would not0.99
    37. and therefore complicate the annotation process.1.00
    38. allow us to perform annotation with rectangular bounding-boxes1.00
    39. [.-]0.60
  • 4:48 #13 done50 line(s)

    shot 13·sharpness 2212.3

    1. So, let's try a simple PDF parser... okay that won't cut it!0.98
    2. undesired1.00
    3. page headers1.00
    4. KDD '22, August 14–18, 2022, Washington, DC, USA Birgit Pfitzmann, Christoph Auer0.98
    5. Nassar, and Peter Staar1.00
    6. KOD ZL Angml 14–16. 202, Wolangkm, DC. 40A Baglt Plikm0.64
    7. Table 1: DocLayNet dataset overview. Along with the frequency of each class labol, we pres0.99
    8. of row "Total") in the train, test and validation sets. The inter-annotator agreement is compu0.98
    9. occurrence (as %0.99
    10. quqn0.56
    11. [email protected] metric1.00
    12. between pairwise annotations from the triple-annotated pages, from which we obtain accur0.99
    13. Very fast and cheap1.00
    14. % of Total1.00
    15. triple inter-annotator mAP @ 0.5-0.95 (%)0.99
    16. X Incomplete0.94
    17. [...]0.94
    18. 225241.00
    19. Count1.00
    20. Tables not0.97
    21. X Loss of structure0.96
    22. 63181.00
    23. 1856601.00
    24. 708781.00
    25. 250271.00
    26. understood1.00
    27. X Noisy0.92
    28. 580221.00
    29. 459761.00
    30. 5103771.00
    31. 50711.00
    32. 1428841.00
    33. 347331.00
    34. Image content1.00
    35. missing1.00
    36. Unfit for most use0.98
    37. cases1.00
    38. 11074701.00
    39. [...1]0.74
    40. include publication repositories such as arXiv3, government o"ces.0.99
    41. company websites as well as data directory services for #nancial0.99
    42. Line wraps not1.00
    43. reports and patents. Scanned documents were excluded wherever1.00
    44. allow us to perform annotation with rectangular bounding-boxes0.99
    45. possible because they can be rotated or skewed. This would not0.98
    46. understood1.00
    47. and therefore complicate the annotation process.0.99
    48. [...]0.96
    49. Multi-column1.00
    50. often breaks order1.00
  • 5:12 #14 done145 line(s)

    shot 14·sharpness 2477.9

    1. But powerful frontier models? Not bad!1.00
    2. DocLayNet dataset overview1.00
    3. Table 10.99
    4. Along with the frequency of each las labe( we present the relative occurence (as 1% of row Total') in mhe train, lnst an0.80
    5. agreement is computed as te mAP0.5-C.95 metric between pairwise annotations from the briple-acnotated oages,0.92
    6. class label0.89
    7. Count0.92
    8. % of Total0.77
    9. metaner mAP 0.5-0.05 (%)0.88
    10. Test0.98
    11. w0.77
    12. MI0.65
    13. re0.55
    14. Man0.99
    15. Sui0.59
    16. Lov0.54
    17. Good quality and0.98
    18. Caption1.00
    19. 226240.95
    20. 2.040.99
    21. 1970.68
    22. 2.320.91
    23. 84-890.90
    24. 40-410.93
    25. 86-020.87
    26. 81-090.65
    27. Footrote0.87
    28. 6m180.70
    29. 0.800.82
    30. c.390.73
    31. 0.580.90
    32. 83-910.70
    33. no0.65
    34. 000.84
    35. 62-050.56
    36. robustness1.00
    37. Formuia0.82
    38. 260270.96
    39. 2.260.98
    40. 1.800.79
    41. 2.900.84
    42. 83-850.96
    43. n0.70
    44. 84-870.93
    45. Ust-tem0.76
    46. 1856600.87
    47. 7:90.66
    48. 13.340.99
    49. 15.820.97
    50. 87-880.87
    51. 74-630.97
    52. 80-000.84
    53. 07.470.76
    54. Page-footer0.93
    55. 708780.94
    56. 6.610.93
    57. 5.560.84
    58. 6.000.99
    59. 03-940.86
    60. 88-800.83
    61. 06-060.71
    62. 1000.95
    63. Expensive (for now)1.00
    64. Pago-heeder0.86
    65. 680220.99
    66. 6.100.89
    67. 6.700.94
    68. 6.060.98
    69. 85-890.84
    70. 66-760.95
    71. 90-940.77
    72. 08-1000.87
    73. Peturs0.77
    74. Sectian-header0.90
    75. 1428040.92
    76. 469760.92
    77. 4.210.97
    78. 12.600.89
    79. 16.770.85
    80. 2.780.98
    81. 6.310.65
    82. 12.850.98
    83. 83-840.91
    84. 69-710.88
    85. 56-590.85
    86. 76-010.89
    87. 82-000.64
    88. 80-920.89
    89. 94-000.97
    90. 60-820.87
    91. Hard to achieve consistent1.00
    92. Tabie0.91
    93. Test0.80
    94. 6808770.89
    95. 347330.89
    96. 3.200.88
    97. 46.820.99
    98. 2.270.91
    99. 49.280.86
    100. 1600.74
    101. 46.000.92
    102. 77-a10.79
    103. 84-800.87
    104. 85.800.82
    105. 76-000.92
    106. 82-000.90
    107. 88-030.81
    108. 00-000.85
    109. 80.910.76
    110. structured output1.00
    111. Tite0.88
    112. 60710.89
    113. 0.471.00
    114. 0.301.00
    115. 0.501.00
    116. 60-720.88
    117. 24-430.99
    118. 50-630.85
    119. 94-1000.94
    120. Toral0.93
    121. 9074700.72
    122. 9411230.88
    123. 098160.92
    124. 666310.93
    125. 82-830.81
    126. 71-240.81
    127. 79-810.83
    128. 80-040.87
    129. Our inclusion criteria for documents were descbed in Section 3. A arge effort went into ensuring that al documents are Iro0.88
    130. Phase 1: Data selection and preparation0.99
    131. Possible hallucinations1.00
    132. therefore complicate the annotation process.0.99
    133. pubrication repositories such as arxk, govemment offices, company websites as well as data directory services for finuncial reports0.96
    134. were excluded wherever possible becase they can be rotated or skewed. This would not alow us to perform annotation w/th rec0.93
    135. Very costly at scale1.00
    136. Phase 2: Label selection and guideline0.99
    137. Phose I: flots melar0.70
    138. and lead us to the definition of 11 distinct class labels. These 11 class labels are Caption, Footnore, Formula Lis-itomn Pope-0.86
    139. heade able, Text and Titie Critical factors that were considered for he choice of these clas labels were () the ouvrall occc0.88
    140. We revriewed the collected doouments and identified the most common structural features they exnibit. This was achieved bay0.96
    141. not always faithfu1.00
    142. the label(, () necognisabiity on a single page (.e no ned for cotet from pelous or net pag) and (4) overall couerage of te0.71
    143. choice oflabel is not ambiguous, whle coverage ensures tha llmeeninglulitems on a pege can be annotuted. Wo relraind from0.81
    144. to a decument category, such as Abstract in the Sclentilic Articles category. We also avoided cless labe's that are lightly Inkend lo0.91
    145. such as Autfor and Affilation, as seen in Doclhank, are often only distinguisheble by discriminating on0.94
  • 5:44 #15 done15 line(s)

    shot 15·sharpness 2955.1

    1. Maybe there's a middle ground... Welcome to Docling!0.99
    2. KOD 21. Asapat 14-t8. 2022, Wachnglom, DC. USA Bregt Pitemae0.71
    3. occurrence (as of row Total) in the tran, test and valdation sets. The inter-annotator agrement is computed0.79
    4. as the [email protected] metric between painwise annotations from the triple-annotated pages. from which we0.92
    5. class label0.99
    6. Table 1: DocLayNet dataset overview. Along with the frequency of each class label, we present the relative0.94
    7. Count1.00
    8. Train0.98
    9. Test1.00
    10. % of Total0.99
    11. obtain accuracy ranges.0.99
    12. Val1.00
    13. All1.00
    14. Fin1.00
    15. triple inter-annotator mAP @ 0.5-0.95 (%)0.99
  • 6:15 #16 skipped

    shot 16·duplicate of #15

  • 6:45 #17 done

    shot 17·sharpness 4077.9

  • 7:29 #18 done

    shot 18·sharpness 4189.3

  • 7:45 #19 done

    shot 19·sharpness 2520.6

  • 8:15 #20 skipped

    shot 20·duplicate of #19

  • 8:58 #21 done

    shot 21·sharpness 2755.0

  • 9:31 #22 done

    shot 22·sharpness 1709.5

  • 9:43 #23 done

    shot 23·sharpness 1980.8

The page's on-screen-text budget of 600 lines is spent, so the last cards in this grid list fewer lines than they hold. Narrow the page with ?frames= to read them.

Transcript

174 cues· 3,939 words· 21,100 chars

  1. 0:00 Hey, hey, welcome.
  2. 0:00 My name is Cedric Clyburn.
  3. 0:02 I'm an open source engineer here at Red Hat.
  4. 0:04 And I think we can all agree that context is the most important aspect to building an AI application or an agent, right?
  5. 0:11 It's the reason that harnesses have become so popular in order to manage the LLMs context.
  6. 0:15 But the thing is, no matter what model or agent that you're using, there is so much data that we're not able to use properly because it's in unstructured formats.
  7. 0:25 I'm talking everything from PDFs to presentations to contracts and technical docs, even meeting notes, scanned documents, diagrams, tables, images, and more.
  8. 0:35 And sorry, I know that's a lot, but you understand what I mean, right?
  9. 0:38 All this data needs to be transformed into something that an LLM can actually understand.
  10. 0:43 And that's why by the end of this session, you'll understand
  11. 0:46 how to extract structure between raw enterprise documents and use it to power better downstream AI systems like RAG and Agents.
  12. 0:53 So let's get started.
  13. 0:55 Now, I think Jensen from NVIDIA made this point super clear at his keynote that unstructured data is becoming this new context layer for AI.
  14. 1:03 And the reality though, for many teams, and I know this personally working at Red Hat, that PDFs and data are spread across dozens of different systems.
  15. 1:12 So we've got a lot to cover today.
  16. 1:14 As you might know, a large majority of the world's data is unstructured, and so no matter what model you're using, if you're working with data like PDFs and unstructured types of formats, this is a bit tricky to work with.
  17. 1:26 Because there are solutions out there, but they might be proprietary or require you to send your private data to someone else's server.
  18. 1:33 And for not just text, how do we take documents and their graphs or tables and images to formats that LLMs can understand like Markdown or JSON?
  19. 1:42 And I'm going to show you how in the session today because we're going to be using an open source tool, part of the Linux foundation that is called Dockling, and learn about extraction, parsing, chunking, and much more.
  20. 1:53 And I've got some live demos for you, so we're going to have some fun.
  21. 1:56 And just in case you'd like it, we have the session slides here on the right and a little overview of the specifics I'll be showing you in today's session.
  22. 2:04 But without further ado, let's get started.
  23. 2:07 So why is there a need for advanced document processing?
  24. 2:10 As I briefly mentioned before, you might have a lot of technical documentation or meeting minutes or different types of documents and invoices that you need to use and maybe RAG or different type of applications where the context is provided to an LLM.
  25. 2:26 So whether it's RAG or retrieval augmented generation to answer questions based on this data, or you're using this to fine tune a new specialized model, well, data is this key ingredient behind those applications.
  26. 2:39 And it doesn't matter if you're using NVIDIA acceleration or an open source or proprietary model, that data and the way you process it is the key determining factor in whether your answer is going to be correct or incorrect for the user or customer at the end of the day.
  27. 2:55 And that's what's most important.
  28. 2:58 And how important is it?
  29. 2:59 Well, I had this viral tweet from earlier where 20 scientific papers now feature a new nonsensical term that doesn't exist because AI misinterpreted a very old article that was scanned and taken to a PDF, merging two different words from two different columns in this PDF.
  30. 3:17 And because researchers are using these models in order to help them write, now we have different types of scientific papers that all feature this word and are even being cited by other people.
  31. 3:28 And so that's how important it is to make sure that the data that we're processing is processed in a way that's accurate and not hallucinated and able to be used confidently in our applications that we're delivering to users and customers.
  32. 3:42 So it's quite important.
  33. 3:43 Now, if we were to use a tool like Docling that I can run on my own machine, you could see that these two words are quite far away from each other and shouldn't have been combined in the first place.
  34. 3:54 But that's how we're going to learn about extracting this text here in a second.
  35. 3:58 Now, if we were to try a simple PDF parser for a PDF like this that includes a table here, it has an image, there's captions, and there's regular sections of text,
  36. 4:08 Well, we might get an answer like this here on the right in Markdown.
  37. 4:12 You know, this might be very fast and cheap to run even on CPU.
  38. 4:16 But the issue is, is that a lot of this text has been truncated, has been merged, and isn't decipherable even by me as a human.
  39. 4:25 And if I sent this to a model, I don't think I could trust that the model could extract specifics from, say, for example, this table.
  40. 4:31 because the table has been kind of just spit out linearly, and this information isn't fit for most use cases where I need to ask questions or have an agent do validation and extraction on this source data.
  41. 4:46 So this isn't going to cut it, right?
  42. 4:48 There's undesired page headers, we don't understand the table, and where's the content from the image, right?
  43. 4:52 It's not even there.
  44. 4:54 when we're using frontier models this is um kind of not bad right but quite expensive i'm sending this to a model that's maybe thirty dollars per million output tokens you can see how this can get quite expensive as i scale this up to dozens or hundreds or in a lot of cases thousands of pdfs that organizations have to work through to use an ai application
  45. 5:16 And the differences between maybe a 5.1 of a model that was depreciated and a 5.2 version of a model make it tricky to have structured output that's consistent every single time.
  46. 5:27 And so while it might be good quality, and I can see that most of this looks accurate in the exported markdown, we might be susceptible to hallucinations because models are non-deterministic.
  47. 5:37 And this is really tricky at scale.
  48. 5:40 And so what is the middle ground?
  49. 5:42 Well, that's where Dockland comes in.
  50. 5:44 It's a fast and cheap and most importantly local CLI and library that I can use to take various types of input sources and convert this to Markdown, JSON, and a Pydantic data type that I can use in my applications and that I can scale up if I have thousands of different types of formats that need to be used or translated to something like Markdown.

Open at this second