Videos Iwe_RY-fYgI
AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
Scene timeline
37 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 148
- whisperx 148
- chunks
- 34
- from 148 cues
- keyframes
- 16
- kept of 37 captured
- frames with text
- 16
- 273 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 4.7 MB
- word timings on 148 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:56 | 1m 03s |
stt |
done | — | 2026-08-11 05:57 | 17s |
chunk |
done | — | 2026-08-11 05:57 | 0s |
text_embed |
done | — | 2026-08-11 05:57 | 0s |
keyframe |
done | — | 2026-08-11 05:57 | 54s |
ocr |
done | — | 2026-08-11 05:58 | 9s |
frame_embed |
done | — | 2026-08-11 05:58 | 3s |
Frames, and what the machine read
-
- AI-Driven Multi-Document1.00
- Correlation for Enterprise1.00
- Financial Compliance and Fraud0.99
- Detection1.00
- A framework for cross-document fraud detection through relational intelligence, evaluated0.99
- across 3 million anonymized records and four jurisdictions.1.00
- By Varsha Shah, Enterprise Technical Architect, USA0.99
-
- The Compliance Gap No One Is Closing0.99
- Multi-Jurisdictional1.00
- Growing Data Volumes1.00
- Sophisticated Fraud1.00
- Complexity1.00
- Patterns1.00
- Enterprise financial records span0.99
- Overlapping and divergent1.00
- payroll, tax, procurement, and1.00
- Modern fraud exploits1.00
- regulations across global systems1.00
- transactions, generating data at a1.00
- inconsistencies that only surface1.00
- create blind spots that manual and0.97
- scale that overwhelms0.98
- when comparing across related1.00
- rule-based tools cannot reconcile.1.00
- conventional review.1.00
- document sets, not within any0.99
- single record.1.00
- Rule-based and NLP-augmented systems operating at the document level are structurally incapable of1.00
- detecting these cross-document anomalies.1.00
-
- Why Document-Level1.00
- Analysis Falls Short0.99
- SANE1.00
- BIUNDE0.99
- ED1.00
- MASK1.00
- Traditional compliance tools evaluate records in isolation. Fraud that1.00
- 20.980.94
- 2,681.00
- 3,301.00
- 000.78
- 600.98
- 11,981.00
- arises from discrepancies between payroll registers, vendor invoices,1.00
- 25.501.00
- 2,501.00
- 2,501.00
- and tax filings remains invisible when each document passes its own1.00
- 24,501.00
- 2,001.00
- 3,501.00
- internal validation.1.00
- MKME1.00
- DASKY1.00
- BUNSE1.00
- OORTU1.00
- FEUN0.96
- 9,631.00
- 2,001.00
- 3,771.00
- 8,801.00
- 2,551.00
- 8,881.00
- 14,151.00
- 17.601.00
- 2,501.00
- 11.501.00
- 6,631.00
- 14.101.00
- 16,951.00
- 5,701.00
- 10.801.00
- The most costly fraud patterns are not found within a single0.99
- 12101.00
- 15.501.00
- 10,701.00
- 1,801.00
- 60,501.00
- document. They emerge in the space between documento0.98
-
- Framework Architecture Overview1.00
- Entity Correlation0.98
- Graph engine linking financial records1.00
- Risk Modeling0.98
- Adaptive probabilistic aggregation of1.00
- anomalies1.00
- Normalization Layer0.98
- Harmonize currencies and reporting1.00
- standards1.00
- The three components operate in concert: entities are linked relationally, risk signals are aggregated1.00
- and calibrated, and jurisdictional variance is normalized before scoring.1.00
-
- COMPONENT 10.97
- Graph-Based Entity Correlation1.00
- Engine1.00
- 10.57
- The foundation of the framework is a knowledge graph that establishes1.00
- relational links across payroll, tax, procurement, and transactional records.1.00
- Entities normalized and resolved across disparate data sources0.99
- Relationships surfaced between vendors, employees, accounts, and filings0.99
- Structural anomalies detected through graph topology analysis0.99
-
- Porbability Distribution's Sovon Ryst a0.99
- Risk Scoring Model1.00
- COMPONENT 20.98
- Adaptive Probabilistic Risk Model1.00
- 10001.00
- 2501.00
- Proear Snity0.53
- Anomaly signals from correlated documents are aggregated into calibrated1.00
- 4001.00
- A0.60
- 000.66
- risk scores that prioritize cases requiring human review.0.99
- 3501.00
- 2001.00
- Probabilistic scoring accounts for signal strength and source reliability0.99
- B:150.82
- Reis0.58
- Model adapts continuously from audit feedback loops0.99
- D:200.89
- ●0.71
- Scoridly Divew Risek0.99
- Low Risk0.98
- Risk thresholds configurable per jurisdiction and business unit1.00
- High RiskPolorisk Pow RiskLow Risk0.98
- Rub Risk Righ RiskLow Risk Dow Risk0.95
- Low Risk Low Risk Low Risk Low Risk0.96
-
- COMPONENT 30.97
- Cross-Jurisdictional Normalization1.00
- Layer1.00
- Before correlation can occur, disparate financial reporting standards must be0.99
- reconciled. This layer harmonizes:1.00
- Currency conversions and exchange rate adjustments1.00
- Tax structure differences across jurisdictions1.00
- Reporting period and classification schema alignment1.00
- This normalization enables the risk model to score entities0.99
- consistently regardless of geographic origin.0.99
-
- Evaluation Conditions1.00
- Dataset Scale1.00
- Jurisdictions1.00
- ...0.72
- Approximately 3 million1.00
- Four distinct regulatory1.00
- anonymized financial records1.00
- environments evaluated in1.00
- parallel1.00
- Time Horizon1.00
- Five years of historical data reflecting real-world enterprise1.00
- conditions1.00
-
- Detection Performance Results0.99
- ~91%0.86
- ~87%0.88
- ~0.891.00
- Precision1.00
- Recall1.00
- F1 Score0.99
- Flagged cases confirmed as genuine1.00
- True fraud cases successfully surfaced0.99
- Harmonic mean balancing precision1.00
- anomalies0.98
- by the system0.99
- and recall0.99
- Performance was measured against a labeled ground truth derived from confirmed audit findings across1.00
- all four jurisdictions and the full five-year evaluation window.0.99
-
- Operational Efficiency Gains1.00
- ~9%0.83
- ~40%0.87
- False Positive Rate1.00
- Audit Workload Reduction0.99
- Down from approximately 38% in rule-based baseline systems1.00
- Fewer manual reviews required per compliance cycle1.00
- What This Means in Practice0.99
- A 76% reduction in false positives translates directly to fewer investigator hours spent on non-events. A 40% reduction in1.00
- manual audit workload frees compliance teams to focus analytical capacity on high-confidence risk cases rather than routine1.00
- triage.1.00
Transcript
148 cues· 2,238 words· 15,632 chars
- 0:04 Hello, everyone.
- 0:05 Thank you to the AI Engineering World's Fair team for providing this wonderful opportunity to share my research.
- 0:13 It is truly an honor to be speaking alongside so many talented researchers and practitioners.
- 0:18 My name is Varshasha.
- 0:20 I am an Enterprise Technical Architect working at Tata Consultancy Services, working for Microsoft.
- 0:26 I'm focused on artificial intelligence, enterprise compliance, finance governance, and intelligent automation.
- 0:34 Today, I would like to share my research on AI-driven multi-document correlation for enterprise financial compliance and fraud detection.
- 0:43 As organizations continue to digitalize their operations today, they generate a numerous amount of data for the financial system across payroll, tax, procurement, transaction system.
- 0:56 Ironically, while we have more data than ever before, compliance teams continue to struggle with hidden fraud patterns, regulatory risk.
- 1:05 The reason is the most existing solution analyze the documents independently.
- 1:10 while many of the most critical risks only become visible when the information is connected across the multiple systems.
- 1:18 In this presentation, I'll introduce you to a framework that combines the graph-based entity correlation, probabilistic risk modeling, and cross-judictional normalization to uncover these hidden relationships and transform enterprise compliance from reactive process into proactive intelligence capability.
- 1:38 The framework was evaluated using approximately 3 million financial records across four judicial demonstrating both the strong detection performance and meaningful operations improvement.
- 1:51 With that context, let's begin by looking at the compliance gap that organizations continue to face today.
- 2:01 Let's begin by understanding the compliance gap that many organizations continue to face.
- 2:05 Enterprise compliance has become significantly more complex over the last decade.
- 2:10 Organizations now operate across multiple countries, regulatory frameworks, and financial system, each with own reporting standards and compliance requirements.
- 2:20 At the same time, the volume of the enterprise has grown exponentially.
- 2:25 Payroll records, tax filings, procurement transactions, and financial documents are generated every day, making manual reviews both time-consuming and increasingly impractical.
- 2:38 Adding to this challenge, fraud has been evolved.
- 2:42 Modern fraud rarely appears as an obvious error within a single document.
- 2:48 Instead, it exploits subtle inconsistency across multiple systems, patterns that often remain invisible when the records are reviewed independently.
- 2:59 This creates a fundamental limitation.
- 3:02 Traditional rule-based and document-level NLP system are designed to validate individual records, but they are not built to understand relationship across the documents.
- 3:14 And that's the gap this research aims to address.
- 3:17 Moving beyond this isolated document analysis to uncover the hidden risk through the cross-document correlation.
- 3:24 to better understand this limitation let's look at why traditional document level analysis often fail to detect the most sophisticated fraud patterns to understand why traditional approaches struggle let's consider how most compliance system operates nowadays they evaluate each document independently
- 3:47 A payroll register is validated against the payroll rules.
- 3:51 The vendor invoices are checked against the procurement policies.
- 3:55 A tax filing is reviewed under the tax regulations.
- 3:59 If each document passes its individual validation, the transaction is generally considered compliant.
- 4:05 The challenge is that many sophisticated fraud patterns doesn't really appear within a single document that emerges only when the multiple documents are analyzed together.
- 4:16 For example, a payroll record may appear accurate to us.
- 4:20 A vendor invoice may seem legitimate to us.
- 4:23 A tax filing may be correctly submitted.
- 4:27 But when these records are connected, they are revealing the inconsistencies and indicate the fraud and compliance risk.
- 4:34 The information already exists.
- 4:37 What missing is the ability to understand the relationship between these documents.
- 4:41 That is why the research shifts the focus from document level validation to cross document intelligence, enabling the organizations to detect the risk that would otherwise remain hidden.
- 4:53 So if the problem is understanding relationships rather than individual documents, what kind of architecture can solve this?
- 5:01 Let me introduce you to the framework.
- 5:05 Now that we have established the problem, let's look at the proposed framework.
- 5:09 Rather than relying on a single model or algorithm, the solution is built on three complementary components that work together to transform the raw enterprise data into the actionable compliance intelligence.
- 5:22 The first component is the entity correlation engine.
- 5:27 The purpose of this is to connect the related information across the payroll, tax procurement, financial systems,
- 5:34 creating a unified view of enterprise activity rather than isolated records.
- 5:39 Once these relationships are established, the second component that is adaptive probabilistic risk model, which evaluates the connected data to determine which patterns represent meaningful compliance.
- 5:53 Instead of generating alerts based on single rule, it prioritized the cases using the multiple risk signals here.
- 6:03 The cross-judictional normalization layer provides the regulatory context by standardizing the currency, tax structure, reporting standards, and compliance rules across the different jurisdictions.
- 6:16 This ensures that the risks are evaluated consistently regardless of where the transaction originated.
loading