Videos -x5GEVnkuRw
Structuring the Unstructured - Cedric Clyburn, Red Hat
Scene timeline
146 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 174
- whisperx 174
- chunks
- 36
- from 174 cues
- keyframes
- 141
- kept of 146 captured
- frames with text
- 141
- 10,196 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 23.8 MB
- word timings on 174 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 08:07 | 1m 15s |
stt |
done | — | 2026-08-11 08:08 | 23s |
chunk |
done | — | 2026-08-11 08:09 | 0s |
text_embed |
done | — | 2026-08-11 08:09 | 1s |
keyframe |
done | — | 2026-08-11 08:09 | 1m 51s |
ocr |
done | — | 2026-08-11 08:11 | 3m 50s |
frame_embed |
done | — | 2026-08-11 08:14 | 28s |
Frames, and what the machine read
-
- AlEngineer0.99
- Red Hat0.99
- World's Fair0.97
- Developer1.00
- Structuring the Unstructured: Advanced1.00
- Document Parsing for Al Workflows0.99
- AI Engineer: 20260.97
- Cedric Clyburn0.98
- Senior Developer Advocate1.00
- @cedricclyburnn0.98
-
- Unstructured Data is the Context of Al0.99
- 100's of Zettabytes Per Year of Unstructured Data – Growing Exponentially0.99
- Enterprises1.00
- FSI0.82
- RBC0.83
- EnterpriseSoftware1.00
- Storage Platform0.99
- Enterprise Databases1.00
- cSP Engines0.98
- OSS Engines0.92
- cuVS1.00
- Source: IDC Global DataSphere, 20250.99
-
- We've got a lot to cover today!1.00
- PDF1.00
- Wait, so 85% of the0.99
- world's data is...1.00
- unstructured?!1.00
-
- We've got a lot to cover today!0.98
- Azure Al Document Intelligence0.97
- 人0.84
- $1.00
- Accelerate information extraction from documents.0.98
- Amazon Textract1.00
- PDF1.00
- LlamaParse: Transform unstructured1.00
- But current solutions are0.98
- data into LLM optimized formats0.99
- Automatically extract printed text, handwriting, layout elements, and1.00
- data from any document0.99
- proprietary, and require1.00
- sending your private data!0.98
- Wait, so 85% of the0.99
- world's data is...1.00
- unstructured?!1.00
- let's learn about0.97
- extraction, parsing,1.00
- chunking, and much more!1.00
- docling1.00
- So, how can we easily parse d0.98
- graphs, tables, etc to formats1.00
-
- Agenda1.00
- Today's Schedule1.00
- Data Preparation for1.00
- Demo #1:0.94
- Al: It's not easy!1.00
- Extracting/Parsing1.00
- Unstructured Data1.00
- Parsing PDF's,0.99
- tables, images, etc.1.00
- Demo #2: Chunking &1.00
- Embedding1.00
- Building an Al0.98
- Session Slides0.99
- document pipeline1.00
- Demo #3: Building &0.98
- red.ht/structuring1.00
- Deploying a Q&A app!1.00
- Red Hat0.98
- Developer1.00
-
- Why's there a need for0.97
- advanced document1.00
- processing?1.00
- Red Hat0.88
- Developer1.00
-
- Data is the key ingredient behind Al applications!1.00
- Table of contets0.90
- Board Meeting1.00
- 二0.58
- Chapter 2. Creating an Amazon S3 client using1.00
- Jine 24 03 0 AM 1:300 AM0.67
- notebook cells1.00
- ATTENDEES:0.93
- HOST0.91
- ane Rotrguer0.69
- ▲ POF0.89
- that service.0.97
- 900 AM- 9:300 AME0.77
- 1.1 0paring0.85
- Teple0.86
- YartaAArata0.70
- Presecter0.82
- Jose Rodripuat0.84
- 三0.51
- Prerequisites0.90
- • Access o a Jupyter notebok server runng on Red Hat penShit Al.0.93
- 1.2 Atandance0.88
- Dvdn maters0.60
- Financial1.00
- Defi us or t A an vie varables0.58
- when you star your notebook server sing the values from your Amazon Web Services0.95
- 1.3 Appreval of Agenda0.90
- Jef Yon0.64
- Documents1.00
- er Portal0.99
- Procedure0.98
- 1 In new notebook el, iprt th equied irarie y dding te folowing:0.61
- account under My Security Credentials.0.95
- e0.66
- 9.30 10 1 0 . Revew of Perio Monues0.52
- 2.1 Aproetlt telete Jof Kaih0.51
- Yerta Amata0.76
- How to control the access of podman to system0.99
- users1.00
- inpert botal0.66
- fron boto3 lapert session0.81
- Jone Rodrigunz0.75
- Environment1.00
- 2.2 Cloing trsteages0.68
- Jone Rodiguit0.59
- Issue1.00
- i Define your credendials.0.93
- Meeting Minutes1.00
- Technical1.00
- Resolution1.00
- Documentation1.00
- Knowledge1.00
- Articles1.00
-
- Data is the key ingredient behind Al applications!1.00
- Board Meeting1.00
- ATTENODEES0.84
- Powering:1.00
- Table of conteents0.90
- 00.830 AM0.68
- Jonn Rodigu0.63
- RAG (Document Q&A)0.97
- Chapter 2. Creating an Amazon S3 client using0.98
- notebook cells1.00
- Fine-Tuning1.00
- 4 PDF0.84
- 1.3 Aporoval of Agante0.78
- To interactwth datain mazon S3 buckets, youmust reste alocal lient t ande requests to0.76
- that service.0.97
- rowelldt Prolous Woute0.53
- Tewrtoa Amata0.69
- Financial1.00
- etc.1.00
- Prerequisites0.93
- Documents1.00
- •Access t a Jupyter notebook server rnning on Red Hat OpenShift A Al.0.89
- Defin vaue t an T A v vaables0.56
- when you start our notebok erver using hevalues from you Amazon Web Sevies0.83
- or Portal0.89
- account under My Security Credentials.0.76
- Meeting Minutes1.00
- Procedure1.00
- How to control the access of podman to system1.00
- I In a new notebook cel, mport the required libraries by adding the folowing:0.90
- e0.74
- users1.00
- Hugging Face1.00
- Environment1.00
- NVIDIA.0.96
- fron botu3 (aport session0.84
- databricks1.00
- 2. In another new notebok cel, define the folowing to create your session and clent.0.95
- Issue1.00
- 0Meta0.86
- L. Define your credentials.0.96
- Technical1.00
- Resolution1.00
- Mic0.99
- Documentation1.00
- + much more!1.00
- Google1.00
- Knowledge Base0.99
- Articles1.00
- MISTRAL1.00
- AI_0.98
-
- Data processing & prep is quite important!0.98
- gurovdigital15 h0.99
- ..0.83
- lol, over 20 scientific papers now feature the1.00
- nonsensical term 'vegetative electron0.99
- microscopy'.1.00
- "vegetative electron microsct ×0.95
- all because an Al misinterpreted a 1959 article,0.99
- merging 'vegetative' and 'electron microscopy'0.99
- Scholar1.00
- YEAR -0.85
- from separate columns.1.00
- hydrophila and Yersinia ruckeri bacteria isolated0.93
- Study of CNT@ Fe304 effects on Aeromonas0.99
- from fish.1.00
- on of the mmayme fmin 8.0.72
- spore moala of 8.0.63
- M Alshai Taae A.Minastal. .Joumnai of .20190.55
- search.ebscohost.com0.97
- ie ensyme did not attack0.86
- Norris of Leeds University0.97
- carbon nenocubes syrtheszed by spectroscopic and0.93
- tion). He treated spores0.98
- s in the vegetative cell,0.97
- preparation of lytic enzy0.96
- à sporangium. It is by no0.98
- spores and examined the0.97
- pore is released. In Clos-1.00
- ears that at lenst part of0.97
- happens to the vegetative1.00
- hed as an outer membrane0.99
- electryu mieroscopy. No ev0.96
- exospcsium was obtained.0.95
- in spores, or another enzy0.98
- for lysin of the sporangial0.90
- It was not known whethe0.99
- Green synthesis of silver nanoparticles via0.99
- Ganoderma lucidum fungus extract and its0.98
- antibacteraleffects on Klebsiellapneumonia0.91
- isolates from urinary tract ..0.97
- M. JamsnidianMoaver, M A-Aborz Untversity. 20210.67
- Vegettve electon microscopy was0.75
- R analysin was also0.93
- a to measure the0.85
- 6921.00
- Q120.99
- 431.00
- ☆ Cited by 1 Related aricles0.87
- performed to investigate possble organic compounos thet0.90
- METALLOGRAPHIC STUDIES OF1.00
- [POF| res0.79
- BRONZE PIECES FROM JEYRÄN TEPE, OZBAKI0.97
- IRAN'S IRON AGE: CASE STUDY0.98
- ESODAEL H RA-NEMA - reseorchgate.net0.90
- This stuty in a report of the resalts of metalogrphic stady of0.92
- Sb ps foud in Jn Tee datig bac e lron0.58
-
- Data processing & prep is quite important!0.98
- gurovdigital15 h0.99
- …0.97
- Date syrup (as one of the agricultural wastes)1.00
- lol, over 20 scientific papers now feature the0.99
- was used to produce bacterial cellulose using1.00
- nonsensical term 'vegetative electron0.99
- microscopy'.1.00
- Gluconastobacter xylinus. Fourier transform0.99
- all because an Al misinterpreted a 1959 article,1.00
- "vegetative electron microsct ×0.95
- infrared spectroscopy (FTIR), vegetative electron1.00
- merging 'vegetative' and 'electron microscopy'0.98
- from separate columns.1.00
- hydrophila and Yersinia ruckeri bacteria isolated0.96
- from fish.0.95
- Study of CNT@ Fe304 effects on Aeromonas0.99
- Scholar0.99
- YEAR·0.91
- microscopy, and X-ray diffraction were used to1.00
- determine the structure of bacterial cellulose,0.99
- M Alsha KR Taae AMinastf .Joumnal o.2019-0.53
- on of the emaytue from 8.0.79
- search.ebacohost.com0.99
- cellulose fibers, and crystallinity of the samples1.00
- ie enayme did not attack0.87
- Norris of Leeds University0.97
- carbon nanetubes synthesized by spectroscopic and0.97
- tion). He treated spores0.97
- (Moosavi and Yousefi, 2011). After 14 days of1.00
- s in the vegetative cell,1.00
- preparation of lytic enzy0.99
- à sporangium. It is by no0.96
- happens to the vegetative0.98
- spores and examined the0.97
- electron mieroscopy. No ev0.99
- incubation at 28 °C, the highest yield of cellulose0.99
- pore is released. In Clos-0.99
- ears that at lenst part of0.98
- hed as an outer membrane0.98
- exosporium was obtained.1.00
- in spores, or another enzy0.98
- for lysin of the sporangial0.92
- It was not known whethe1.00
- Green synthesis of silver nanoparticles via1.00
- Ganoderma lucidum fungus extract and its1.00
- antibacteral ffects on Klebsiella pneumonia0.91
- isolates from urinary tract..0.97
- M Jamshidian-Mopaver. MAi.- Alborz Untversity. 2021 -10.79
- Vegetattve electon micoscopy as0.81
- R analysin was also0.94
- d to measure the0.87
- antimicrobial purposes against1.00
- Silver and gold nanoparticles for1.00
- [HTML] m0.96
- 6921.00
- Q121.00
- 431.00
- ☆ Chted by 1 Related aricies 80.81
- multi-drug resistance bacteria1.00
- METALLOGRAPHIC STUDIES OF0.99
- [POF] res0.84
- BRONZE PIECES FROM JEYRÄN TEPE, OZBAKI0.99
- IRAN'S IRON AGE: CASE STUDY0.98
- N Rabiee, S Ahmadi, O Akhavan, R Lugue - Materials, 2022 -0.98
- SODAEL HRA-NEM - reseorchgute.net0.71
- mdpi.com1.00
- This stul s a reort of the resut of metalographic stady of0.72
- S bronze pleces found in Jeyn Tepe dating back to the ron0.83
- ... Dead bacteria have been observed by imaging and1.00
- elem1.00
- ntai analysis using transmission erect0.97
- on mic0.94
- (TEM1.00
- vegetative electron microscopy, and1.00
- dEDX (0.90
- Micro0.97
- ☆0.99
- Cited by 1121.00
- telated articles0.99
-
- Data processing & prep is quite important!0.98
- were incuvateu witni an extract iromn spores uis-0.91
- acteristic type. 1t was conciuueu tnat at reast0.80
- integrated at pH 7.0. Peptide was released0.98
- part of the sporangial wall was dissolved away0.98
- which established that the coats contained sub-0.99
- to allow release of the spore. It appears likely0.98
- strate for the lytic enzyme present in spores.1.00
- that the exosporium of B. cereus does not have1.00
- Peptide was also released from spore coats of B.1.00
- a composition similar to that of the vegetative1.00
- megaterum by the action of the enzyme from B.1.00
- cell wall, from the results obtained by Dr. J. R.1.00
- cereus spores. The lytic enzyme did not attack1.00
- intact resting spor0.99
- The spore develops in the vegetative cell, which thus becomes a sporangium. It is by no means certain what happens to the1.00
- The spore develops in the vegetative cell,1.00
- vegetative cell wall when the spore is released. In Clostridium species it appears that at least part of this structure is retained as an0.99
- which thus becomes a sporangium. It is by no1.00
- outer membrane around the spore. It is the opinion of some workers that the wall of the sporulating cell forms the exosporium which0.99
- exists as an outer coat around spores of several Bacilus species. Spores of several varieties of B. cereus had exosporia whereas these1.00
- cell wall when the spore is released. In Clos-0.99
- structures appeared to be absent from spores of B. megaterium and B. subtilis. It seems, however, that in Bacillus species at least,0.99
- tridium species it appears that at least part of1.00
- the greater part of the vegetative cell wall is dissolved away before the developed spore is released. If this is true, then soluble0.99
- around the spore. It is the opinion of some1.00
- this structure is retained as an outer membrane1.00
- components containing the characteristic constituents should appear in the medium during spore release. Culture filtrates from B.0.99
- cereus organisms at various stages of growth and sporulation were hydrolyzed and the hydrolyzates analyzed for amino sugars and0.99
- diaminopimelic acid (28). Results showed that a large increase in the concentration of these substances in the culture filtrate0.99
- workers that the wall of the sporulating cell1.00
- occurred during spore release (table 2); they were found to be present in a nondialyzable peptide of the characteristic type. It was0.99
- forms the exosporium which exists as an outer0.98
- concluded that at least part of the sporangial wall was dissolved away to allow release of the spore. It appears likely that the1.00
- exosporium of B. cereus does not have a composition similar to that of the vegetative cell wall, from the results obtained by Dr. J. R.0.99
- Norris of Leeds University (personal communication). He treated spores with a highly active preparation of lytic enzyme from0.99
- cereus spores and examined the effect by means of electron microscopy. No evidence of lysis of the exosporium was obta1.00
-
- So, let's try a simple PDF parser...0.99
- KDD '22, August 14-18, 2022, Washington, DC, USA Birgit Pfitzmann, Christoph Auer, Michel0.98
- Nassar, and Peter Staar1.00
- Table 1: DocLayNet dataset overview. Along with the frequency of each class label, we prese0.98
- occurrence (as %0.99
- of row "Total") in the train, test and validation sets. The inter-annotator agreement is com0.98
- [email protected] metric1.00
- 10.55
- between pairwise annotations from the triple-annotated pages, from which we obtain accu0.98
- Very fast and cheap0.98
- % of Total1.00
- triple inter-annotator mAP @ 0.5-0.95 (%)0.99
- X Incomplete0.93
- [.-]0.58
- Count1.00
- 225241.00
- X Loss of structure0.98
- 63181.00
- 250271.00
- 1856601.00
- ×Noisy0.97
- 708781.00
- 580221.00
- 1428841.00
- 459761.00
- Unfit for most use1.00
- 5103771.00
- 347331.00
- 11074701.00
- 50711.00
- cases1.00
- [..]0.77
- include publication repositories such as arXiv3, government o"ces,0.98
- company websites as well as data directory services for #nancial1.00
- reports and patents. Scanned documents were excluded wherever0.99
- possible because they can be rotated or skewed. This would not0.99
- and therefore complicate the annotation process.1.00
- allow us to perform annotation with rectangular bounding-boxes1.00
- [.-]0.60
-
- So, let's try a simple PDF parser... okay that won't cut it!0.98
- undesired1.00
- page headers1.00
- KDD '22, August 14–18, 2022, Washington, DC, USA Birgit Pfitzmann, Christoph Auer0.98
- Nassar, and Peter Staar1.00
- KOD ZL Angml 14–16. 202, Wolangkm, DC. 40A Baglt Plikm0.64
- Table 1: DocLayNet dataset overview. Along with the frequency of each class labol, we pres0.99
- of row "Total") in the train, test and validation sets. The inter-annotator agreement is compu0.98
- occurrence (as %0.99
- quqn0.56
- [email protected] metric1.00
- between pairwise annotations from the triple-annotated pages, from which we obtain accur0.99
- Very fast and cheap1.00
- % of Total1.00
- triple inter-annotator mAP @ 0.5-0.95 (%)0.99
- X Incomplete0.94
- [...]0.94
- 225241.00
- Count1.00
- Tables not0.97
- X Loss of structure0.96
- 63181.00
- 1856601.00
- 708781.00
- 250271.00
- understood1.00
- X Noisy0.92
- 580221.00
- 459761.00
- 5103771.00
- 50711.00
- 1428841.00
- 347331.00
- Image content1.00
- missing1.00
- Unfit for most use0.98
- cases1.00
- 11074701.00
- [...1]0.74
- include publication repositories such as arXiv3, government o"ces.0.99
- company websites as well as data directory services for #nancial0.99
- Line wraps not1.00
- reports and patents. Scanned documents were excluded wherever1.00
- allow us to perform annotation with rectangular bounding-boxes0.99
- possible because they can be rotated or skewed. This would not0.98
- understood1.00
- and therefore complicate the annotation process.0.99
- [...]0.96
- Multi-column1.00
- often breaks order1.00
-
- But powerful frontier models? Not bad!1.00
- DocLayNet dataset overview1.00
- Table 10.99
- Along with the frequency of each las labe( we present the relative occurence (as 1% of row Total') in mhe train, lnst an0.80
- agreement is computed as te mAP0.5-C.95 metric between pairwise annotations from the briple-acnotated oages,0.92
- class label0.89
- Count0.92
- % of Total0.77
- metaner mAP 0.5-0.05 (%)0.88
- Test0.98
- w0.77
- MI0.65
- re0.55
- Man0.99
- Sui0.59
- Lov0.54
- Good quality and0.98
- Caption1.00
- 226240.95
- 2.040.99
- 1970.68
- 2.320.91
- 84-890.90
- 40-410.93
- 86-020.87
- 81-090.65
- Footrote0.87
- 6m180.70
- 0.800.82
- c.390.73
- 0.580.90
- 83-910.70
- no0.65
- 000.84
- 62-050.56
- robustness1.00
- Formuia0.82
- 260270.96
- 2.260.98
- 1.800.79
- 2.900.84
- 83-850.96
- n0.70
- 84-870.93
- Ust-tem0.76
- 1856600.87
- 7:90.66
- 13.340.99
- 15.820.97
- 87-880.87
- 74-630.97
- 80-000.84
- 07.470.76
- Page-footer0.93
- 708780.94
- 6.610.93
- 5.560.84
- 6.000.99
- 03-940.86
- 88-800.83
- 06-060.71
- 1000.95
- Expensive (for now)1.00
- Pago-heeder0.86
- 680220.99
- 6.100.89
- 6.700.94
- 6.060.98
- 85-890.84
- 66-760.95
- 90-940.77
- 08-1000.87
- Peturs0.77
- Sectian-header0.90
- 1428040.92
- 469760.92
- 4.210.97
- 12.600.89
- 16.770.85
- 2.780.98
- 6.310.65
- 12.850.98
- 83-840.91
- 69-710.88
- 56-590.85
- 76-010.89
- 82-000.64
- 80-920.89
- 94-000.97
- 60-820.87
- Hard to achieve consistent1.00
- Tabie0.91
- Test0.80
- 6808770.89
- 347330.89
- 3.200.88
- 46.820.99
- 2.270.91
- 49.280.86
- 1600.74
- 46.000.92
- 77-a10.79
- 84-800.87
- 85.800.82
- 76-000.92
- 82-000.90
- 88-030.81
- 00-000.85
- 80.910.76
- structured output1.00
- Tite0.88
- 60710.89
- 0.471.00
- 0.301.00
- 0.501.00
- 60-720.88
- 24-430.99
- 50-630.85
- 94-1000.94
- Toral0.93
- 9074700.72
- 9411230.88
- 098160.92
- 666310.93
- 82-830.81
- 71-240.81
- 79-810.83
- 80-040.87
- Our inclusion criteria for documents were descbed in Section 3. A arge effort went into ensuring that al documents are Iro0.88
- Phase 1: Data selection and preparation0.99
- Possible hallucinations1.00
- therefore complicate the annotation process.0.99
- pubrication repositories such as arxk, govemment offices, company websites as well as data directory services for finuncial reports0.96
- were excluded wherever possible becase they can be rotated or skewed. This would not alow us to perform annotation w/th rec0.93
- Very costly at scale1.00
- Phase 2: Label selection and guideline0.99
- Phose I: flots melar0.70
- and lead us to the definition of 11 distinct class labels. These 11 class labels are Caption, Footnore, Formula Lis-itomn Pope-0.86
- heade able, Text and Titie Critical factors that were considered for he choice of these clas labels were () the ouvrall occc0.88
- We revriewed the collected doouments and identified the most common structural features they exnibit. This was achieved bay0.96
- not always faithfu1.00
- the label(, () necognisabiity on a single page (.e no ned for cotet from pelous or net pag) and (4) overall couerage of te0.71
- choice oflabel is not ambiguous, whle coverage ensures tha llmeeninglulitems on a pege can be annotuted. Wo relraind from0.81
- to a decument category, such as Abstract in the Sclentilic Articles category. We also avoided cless labe's that are lightly Inkend lo0.91
- such as Autfor and Affilation, as seen in Doclhank, are often only distinguisheble by discriminating on0.94
-
- Maybe there's a middle ground... Welcome to Docling!0.99
- KOD 21. Asapat 14-t8. 2022, Wachnglom, DC. USA Bregt Pitemae0.71
- occurrence (as of row Total) in the tran, test and valdation sets. The inter-annotator agrement is computed0.79
- as the [email protected] metric between painwise annotations from the triple-annotated pages. from which we0.92
- class label0.99
- Table 1: DocLayNet dataset overview. Along with the frequency of each class label, we present the relative0.94
- Count1.00
- Train0.98
- Test1.00
- % of Total0.99
- obtain accuracy ranges.0.99
- Val1.00
- All1.00
- Fin1.00
- triple inter-annotator mAP @ 0.5-0.95 (%)0.99
The page's on-screen-text budget of 600
lines is spent, so the last cards in this grid list fewer lines than they
hold. Narrow the page with ?frames= to read them.
Transcript
174 cues· 3,939 words· 21,100 chars
- 0:00 Hey, hey, welcome.
- 0:00 My name is Cedric Clyburn.
- 0:02 I'm an open source engineer here at Red Hat.
- 0:04 And I think we can all agree that context is the most important aspect to building an AI application or an agent, right?
- 0:11 It's the reason that harnesses have become so popular in order to manage the LLMs context.
- 0:15 But the thing is, no matter what model or agent that you're using, there is so much data that we're not able to use properly because it's in unstructured formats.
- 0:25 I'm talking everything from PDFs to presentations to contracts and technical docs, even meeting notes, scanned documents, diagrams, tables, images, and more.
- 0:35 And sorry, I know that's a lot, but you understand what I mean, right?
- 0:38 All this data needs to be transformed into something that an LLM can actually understand.
- 0:43 And that's why by the end of this session, you'll understand
- 0:46 how to extract structure between raw enterprise documents and use it to power better downstream AI systems like RAG and Agents.
- 0:53 So let's get started.
- 0:55 Now, I think Jensen from NVIDIA made this point super clear at his keynote that unstructured data is becoming this new context layer for AI.
- 1:03 And the reality though, for many teams, and I know this personally working at Red Hat, that PDFs and data are spread across dozens of different systems.
- 1:12 So we've got a lot to cover today.
- 1:14 As you might know, a large majority of the world's data is unstructured, and so no matter what model you're using, if you're working with data like PDFs and unstructured types of formats, this is a bit tricky to work with.
- 1:26 Because there are solutions out there, but they might be proprietary or require you to send your private data to someone else's server.
- 1:33 And for not just text, how do we take documents and their graphs or tables and images to formats that LLMs can understand like Markdown or JSON?
- 1:42 And I'm going to show you how in the session today because we're going to be using an open source tool, part of the Linux foundation that is called Dockling, and learn about extraction, parsing, chunking, and much more.
- 1:53 And I've got some live demos for you, so we're going to have some fun.
- 1:56 And just in case you'd like it, we have the session slides here on the right and a little overview of the specifics I'll be showing you in today's session.
- 2:04 But without further ado, let's get started.
- 2:07 So why is there a need for advanced document processing?
- 2:10 As I briefly mentioned before, you might have a lot of technical documentation or meeting minutes or different types of documents and invoices that you need to use and maybe RAG or different type of applications where the context is provided to an LLM.
- 2:26 So whether it's RAG or retrieval augmented generation to answer questions based on this data, or you're using this to fine tune a new specialized model, well, data is this key ingredient behind those applications.
- 2:39 And it doesn't matter if you're using NVIDIA acceleration or an open source or proprietary model, that data and the way you process it is the key determining factor in whether your answer is going to be correct or incorrect for the user or customer at the end of the day.
- 2:55 And that's what's most important.
- 2:58 And how important is it?
- 2:59 Well, I had this viral tweet from earlier where 20 scientific papers now feature a new nonsensical term that doesn't exist because AI misinterpreted a very old article that was scanned and taken to a PDF, merging two different words from two different columns in this PDF.
- 3:17 And because researchers are using these models in order to help them write, now we have different types of scientific papers that all feature this word and are even being cited by other people.
- 3:28 And so that's how important it is to make sure that the data that we're processing is processed in a way that's accurate and not hallucinated and able to be used confidently in our applications that we're delivering to users and customers.
- 3:42 So it's quite important.
- 3:43 Now, if we were to use a tool like Docling that I can run on my own machine, you could see that these two words are quite far away from each other and shouldn't have been combined in the first place.
- 3:54 But that's how we're going to learn about extracting this text here in a second.
- 3:58 Now, if we were to try a simple PDF parser for a PDF like this that includes a table here, it has an image, there's captions, and there's regular sections of text,
- 4:08 Well, we might get an answer like this here on the right in Markdown.
- 4:12 You know, this might be very fast and cheap to run even on CPU.
- 4:16 But the issue is, is that a lot of this text has been truncated, has been merged, and isn't decipherable even by me as a human.
- 4:25 And if I sent this to a model, I don't think I could trust that the model could extract specifics from, say, for example, this table.
- 4:31 because the table has been kind of just spit out linearly, and this information isn't fit for most use cases where I need to ask questions or have an agent do validation and extraction on this source data.
- 4:46 So this isn't going to cut it, right?
- 4:48 There's undesired page headers, we don't understand the table, and where's the content from the image, right?
- 4:52 It's not even there.
- 4:54 when we're using frontier models this is um kind of not bad right but quite expensive i'm sending this to a model that's maybe thirty dollars per million output tokens you can see how this can get quite expensive as i scale this up to dozens or hundreds or in a lot of cases thousands of pdfs that organizations have to work through to use an ai application
- 5:16 And the differences between maybe a 5.1 of a model that was depreciated and a 5.2 version of a model make it tricky to have structured output that's consistent every single time.
- 5:27 And so while it might be good quality, and I can see that most of this looks accurate in the exported markdown, we might be susceptible to hallucinations because models are non-deterministic.
- 5:37 And this is really tricky at scale.
- 5:40 And so what is the middle ground?
- 5:42 Well, that's where Dockland comes in.
- 5:44 It's a fast and cheap and most importantly local CLI and library that I can use to take various types of input sources and convert this to Markdown, JSON, and a Pydantic data type that I can use in my applications and that I can scale up if I have thousands of different types of formats that need to be used or translated to something like Markdown.
loading