Videos Akm1sqvWG4A
Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy
Scene timeline
208 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 390
- whisperx 390
- chunks
- 81
- from 390 cues
- keyframes
- 86
- kept of 208 captured
- frames with text
- 86
- 8,850 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 27.3 MB
- word timings on 390 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 07:41 | 1m 14s |
stt |
done | — | 2026-08-11 07:42 | 49s |
chunk |
done | — | 2026-08-11 07:43 | 0s |
text_embed |
done | — | 2026-08-11 07:43 | 5s |
keyframe |
done | — | 2026-08-11 07:43 | 2m 54s |
ocr |
done | — | 2026-08-11 07:46 | 2m 03s |
frame_embed |
done | — | 2026-08-11 07:48 | 14s |
Frames, and what the machine read
-
- ○×0.79
- ×0.94
- FAC ×0.94
- Tra×0.97
- Mic1.00
- x0.55
- Scr0.94
- ×0.82
- M0.90
- hea0.87
- x0.70
- M hea0.81
- ×0.78
- Pre1.00
- x0.75
- Go:0.90
- tre1.00
- x0.58
- Byp1.00
- x0.78
- obs rec0.99
- World1.00
- Present1.00
- +0.96
- ā0.69
- ×0.83
- C0.78
- localhost:3030/1?clicks=01.00
- ☆1.00
- C0.98
- M0.89
- WPP Bookmarks0.99
- 品0.98
- Free Pictures0.98
- Ogilvy1.00
- 口0.76
- M&M1.00
- AI0.98
- BTrust0.93
- PGPMS1.00
- Websites1.00
- Website Test0.99
- Decoupled1.00
- RAG Tutorials1.00
- Al Trend0.96
- Follow Up1.00
- temp1.00
- 口0.88
- Google Cloud1.00
- »0.85
- ← Back to Help Centre0.96
- FAQ Assistant1.00
- ×0.51
- hi0.99
- Bypassing the Multimodal Tax0.98
- Hi there!I'm the FAQ assistant.1.00
- Framework-Free Hybrid RAG, Raw SQL RRF, and Live UI Telemetry1.00
- A0.82
- I can help with questions answered in1.00
- our uploaded FAQ documents —0.97
- policies, procedures, product details,1.00
- troubleshooting steps, and more.1.00
- What would you like to know?1.00
- Abed Matini1.00
- Senior Backend Developer· Ogilvy0.99
- AI Engineer World's Fair 2026 - Online Track0.97
- K0.86
- ←1.00
- →1.00
- Q0.53
- 00.61
- p0.77
- 1/261.00
- Type a message...0.98
- x0.51
- localhost:51741.00
- x0.86
-
- BO0.99
- ×0.82
- x0.66
- S0.50
- FAQs—0.90
- x0.51
- Tracing0.97
- ×0.71
- Microso0.93
- x0.71
- Screens1.00
- ×0.73
- heading1.00
- x0.82
- M0.81
- heading1.00
- Present:0.96
- x0.57
- treat illr0.85
- x0.52
- Bypassi0.97
- x0.73
- obs rec0.94
- +0.80
- ā0.80
- ×0.87
- C0.65
- localhost:3030/1?clicks=00.98
- ☆1.00
- C0.63
- obs record screen with best quality -1.00
- Google Search1.00
- Co0.64
- WPP Bookmarks0.99
- 品0.98
- Free Pictures0.99
- Ogilvy0.97
- 口0.80
- M&M1.00
- AI0.98
- BTrust0.99
- PGPMS1.00
- Websites1.00
- Website Test0.99
- Decoupled1.00
- RAG Tutorials1.00
- Al Trend0.96
- Fc0.69
- google.com1.00
- »0.98
- G0.87
- Memory usage: 165 MB1.00
- hi1.00
- Bypassing the Multimodal Tax1.00
- Hi there!I'm the FAQ assistant.0.99
- I can help with questions answered in1.00
- Framework-Free Hybrid RAG, Raw SQL RRF, and Live UI Telemetry1.00
- our uploaded FAQ documents —0.98
- policies, procedures, product details,1.00
- troubleshooting steps, and more.0.99
- What would you like to know?1.00
- Abed Matini1.00
- Senior Backend Developer· Ogilvy0.99
- AI Engineer World's Fair 2026 - Online Track0.97
- K0.85
- ←1.00
- →1.00
- Q0.62
- 00.61
- p0.89
- 1/261.00
- Type a message...1.00
- x0.53
- localhost:51741.00
- x0.77
-
- x0.79
- FA( ×0.93
- FAQs — Help Ce0.90
- ×0.80
- Tracing | Langfu:0.96
- x0.58
- Microsoft Word0.98
- ×0.78
- Screenshot 2021.00
- x0.89
- heading-faq-1.×0.98
- heading-faq-2.1.00
- x0.50
- Presenter - Byp0.99
- x0.65
- +0.98
- ā0.81
- ×0.85
- C0.86
- localhost:3030/20.98
- ☆1.00
- C0.98
- M0.79
- Co0.61
- WPP Bookmarks0.98
- 品0.99
- Free Pictures0.97
- Ogilvy0.95
- M&M1.00
- AI0.99
- BTrust1.00
- PGPMS1.00
- Websites1.00
- Website Test0.98
- Decoupled1.00
- RAG Tutorials1.00
- Al Trend0.96
- Follow Up0.98
- 口0.88
- temp1.00
- Google Cloud1.00
- »0.91
- ← Back to Help Centre0.98
- Two problems every document chatbot hits1.00
- FAQ Assistant1.00
- ×0.50
- hi0.99
- Paying to read documents twice1.00
- Search split across too many tools0.99
- Hi there!I'm the FAQ assistant.0.99
- Cloud vision APIs charge ~500-1,000 tokens per0.98
- Good answers need meaning (vectors) and exact1.00
- I can help with questions answered in1.00
- page just to turn a PDF into text —before a user asks0.98
- words (keywords). Teams often run a vector DB, a1.00
- our uploaded FAQ documents —0.97
- anything. A 200-page manual can cost 100k+ tokens1.00
- search engine, and wrapper code to combine them —0.97
- troubleshooting steps, and more.1.00
- policies, procedures, product details,1.00
- at ingest, and tables still break.0.98
- so when results are wrong, you fix config files, not the0.99
- query.1.00
- What would you like to know?0.99
- This talk: parse locally · one Postgres database · hybrid search in plain Python0.98
- K0.94
- 70.99
- ←1.00
- 00.52
- p0.79
- 2/261.00
- Type a message...0.97
- localhost:5174 ×0.93
-
- x0.78
- FA( ×0.95
- FAQs—Help Ce0.95
- ×0.77
- Tracing | Langfu:0.97
- x0.77
- Microsoft Word0.94
- S0.51
- Screenshot 20260.99
- x0.81
- heading-faq-1.×0.94
- heading-faq-2.0.99
- ×0.61
- Presenter - Byp0.99
- x0.59
- +0.98
- ā0.55
- ×0.77
- C0.86
- localhost:3030/20.94
- ☆1.00
- C0.99
- M0.78
- WPP Bookmarks1.00
- 品0.99
- Free Pictures1.00
- Ogilvy0.95
- 口0.82
- M&M1.00
- AI0.99
- BTrust1.00
- PGPMS1.00
- Websites0.93
- Website Test0.98
- Decoupled0.96
- RAG Tutorials1.00
- Al Trend0.94
- Follow Up0.95
- 口0.91
- temp1.00
- Google Cloud1.00
- »0.96
- ← Back to Help Centre0.99
- Two problems every document chatbot hits1.00
- FAQ Assistant1.00
- ×0.52
- hi1.00
- Paying to read documents twice0.98
- Search split across too many tools1.00
- Hi there!I'm the FAQ assistant.1.00
- Cloud vision APIs charge ~500-1,000 tokens per0.98
- Good answers need meaning (vectors) and exact1.00
- I can help with questions answered in1.00
- page just to turn a PDF into text — before a user asks0.98
- words (keywords). Teams often run a vector DB, a0.99
- our uploaded FAQ documents —0.98
- anything. A 200-page manual can cost 100k+ tokens1.00
- search engine, and wrapper code to combine them —0.98
- policies, procedures, product details,1.00
- troubleshooting steps, and more.0.99
- at ingest, and tables still break.0.99
- so when results are wrong, you fix config files, not the0.99
- query.1.00
- What would you like to know?1.00
- This talk: parse locally · one Postgres database · hybrid search in plain Python1.00
- 40.97
- K0.88
- 70.99
- ←1.00
- 00.63
- p0.85
- 2/261.00
- Type a message...0.98
- localhost:5174 X0.93
-
- Byp0.99
- O0.82
- S0.74
- FAQs1.00
- FAQs-Help C×0.94
- Tracing|Langfu:×0.98
- Microsoft Word0.99
- S0.58
- Screenshot 202×0.99
- heading-faq-1.×heading-faq-2.0.98
- ×0.57
- Presenter- Byp0.93
- x0.73
- +0.94
- ā0.56
- ×0.88
- C0.74
- ① File0.94
- C:/Users/AbedMatini/OneDrive%20-%20Ogilvy/Documents/HCP%20Chat/Sample%20Employee%20Handbook%20-%20National%20C...1.00
- C0.96
- M0.90
- 区0.70
- 白0.59
- 10.55
- WPP Bookmarks0.99
- 品0.97
- Free Pictures0.96
- Ogilvy0.93
- M&M0.99
- □AI0.80
- BTrust0.99
- PGPMS1.00
- Websites1.00
- Website Test1.00
- Decoupled1.00
- RAG Tutorials0.99
- Al Trend0.96
- Follow Up1.00
- temp1.00
- 口0.87
- Google Cloud1.00
- »0.97
- Ⅲ0.53
- Microsoft Word - Sample Employee Handbook for web.doc1.00
- 1/281.00
- 100%1.00
- +0.98
- 日0.99
- 50.53
- 40.98
- SAMPLE EMPLOYEE HANDBOOK0.99
Transcript
390 cues· 5,746 words· 30,153 chars
- 0:01 Hello, everyone.
- 0:03 I'm Abed Mattini.
- 0:05 I'm senior backend developer at Ogilvy.
- 0:09 Thanks to AI Engineer World Fair 2026 for giving me a time slot for online track.
- 0:18 And I'm going to walk you through bypassing the multimodal text today, how we can have a framework-free hybrid drag.
- 0:29 Withdraw SQLRF and live telemetry.
- 0:32 As you can see, I have got my live demo in the right side of the screen, so any slide that I pass and if there is anything related.
- 0:46 Related to my.
- 0:49 presentation I can show you I have few tabs also ready here for you to show and walk you through the slides and the demos okay let's begin so
- 1:11 We have two problems that we try to solve today.
- 1:17 One is when we start to chat with any LLM and if we have to upload a document, we usually going to just drag and drop a PDF or a Word doc or an image and ask the chatbot to get ready for the questions that are going to come up.
- 1:34 after we upload those documents.
- 1:36 The issue with that is as soon as we upload these documents, we're going to basically lose some of our tokens that we're supposed to have only for processing these documents without even asking a question.
- 1:54 So we already spent some tokens there without even asking any questions.
- 2:01 That would be the first issue.
- 2:03 And then if we're building a chatbot and we want to have vector database there, we want to have keyword search and semantic search there, and we want to just add more tools, this is going to be too many tools combined altogether.
- 2:19 So in the production, basically chatbot, we're going to have too many tools and it's going to be complicated for us to manage
- 2:32 our production chatbot properly.
- 2:35 So today I'm going to walk you through in three main sections.
- 2:44 One is how we can upload our documents and get our documents and make it ready.
- 2:52 for our chatbot before users gonna chat.
- 2:56 So on the right side, I do have a FAQ assistant chatbot that's running locally now for me.
- 3:07 And
- 3:09 I'm going to show you the dashboard, the dashboard that I've created for it and how we can chat and how we get the results.
- 3:15 So this is going to be a supposedly FAQ assistant for a sample employee handbook.
- 3:26 So imagine that you have an HR company and this HR has uploaded the handbook
- 3:35 for the employees to ask their question, any question they have about leave, about different kinds of leave, a sickness, or parental leave, or any other questions related to possibly an HR would have as a common question that they get every day.
- 4:00 So that would be one sample that we're going to use to upload.
- 4:06 So our documents can get in the different formats or we can upload in different formats of PDF, like PowerPoint documents and Word documents or even image.
- 4:18 so we're going to talk about that we're going to talk about how we have to chunk these because we're talking about drag here we need to know how we can chunk this information uh properly based on the document so we're going to talk about uh different type of chunking here as well we can show you in the admin how we can choose different uh
- 4:45 different strategies to chunk the documents.
- 4:48 Then we're going to see how the search is going to work, how the embedding is going to work, how our database is going to be, how the search is going to be based on the keyword search or the semantic search.
- 5:00 And we're going to see about how RRF in the Python is going to work.
- 5:06 And then we're going to talk about the observability, how we're going to check the safety, how we're going to observe all the process with blank views, and how
- 5:16 we can prevent problem injection pattern or
- 5:24 any risky question that might come to LLM.
- 5:27 I'm planning this talk for about under an hour.
- 5:30 And this, uh, the whole thing, the plan is the entire code base can be easily be running on a GitHub code space, which you can just pull from the repo and one command run, and it's gonna be ready to test on GitHub code spaces.
- 5:52 Um,
- 5:52 So the stack we're using here are Python, FastAPI for the back end, React for the front end, PostgreSQL for our database, Docker for just easily being able to have the containers and we can just reproduce it anywhere, rerun it anywhere, for example, in Codespaces or any other environment we want.
- 6:19 Then we have Ulama, local language models, some local language models, and embedding models that are running only locally.
- 6:26 That can be run on any server as well.
- 6:30 It doesn't need a GPU.
- 6:32 The CPU is going to be good enough for it.
- 6:34 So any staging server, any code space are going to be fine for it.
- 6:43 and length views for tracking what's going on and how we can see the chat and the latency so we can improve our chats better um
- 6:59 This is the end-to-end blueprint of what we're going to talk about.
- 7:03 So we're going to have a few documents as a source of truth for our chatbot.
- 7:12 So we're going to upload some documents.
- 7:15 So what we can do with Python is we're using Duckling to convert those raw documents to Markdown files.
loading