Videos LrGCT7G_rU8
Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon
Scene timeline
33 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 146
- whisperx 146
- chunks
- 26
- from 146 cues
- keyframes
- 17
- kept of 33 captured
- frames with text
- 17
- 319 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 7.2 MB
- word timings on 146 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-11 05:48 | 1m 26s |
stt |
done | — | 2026-08-11 05:49 | 15s |
chunk |
done | — | 2026-08-11 05:50 | 0s |
text_embed |
done | — | 2026-08-11 05:50 | 0s |
keyframe |
done | — | 2026-08-11 05:50 | 1m 20s |
ocr |
done | — | 2026-08-11 05:51 | 8s |
frame_embed |
done | — | 2026-08-11 05:51 | 3s |
Frames, and what the machine read
-
- Using RL-based Agent to1.00
- Detect and Remediat0.99
- Presenter:1.00
- Anna Marie Benzon1.00
- University of the Philippines, Diliman1.00
- github.com/ambenzon27/rl-etl-remediation-agent1.00
-
- Using RL-based Agent to1.00
- Detect and Remediate0.99
- ETL Pipeline Failures0.98
- Presenter:1.00
- Anna Marie Benzon1.00
- University of the Philippines, Diliman1.00
- github.com/ambenzon27/rl-etl-remediation-agent1.00
-
- Failure1.00
- Inspect Logs1.00
- Diagnose1.00
- Repair1.00
- Rerun1.00
- Validate1.00
- Late or unavailable source data0.99
- The1.00
- Schema drift1.00
- X0.95
- Problem1.00
- Datetime parsing incompatibilities0.99
- Null-rate spikes1.00
- X0.82
- Cloud ETL jobs break because of:0.99
- Type changes0.98
- Unknown runtime errors1.00
- STEN0.90
- Modeled manual MTTR ~ 2.5 working days0.99
-
- Phase 1 Architecture — Al Pipeline Health Agent0.99
- Incident-Response Loop Using 6 AWS Services0.99
- AWS Glue ETL Jobs1.00
- 1. job FALED evarnt0.85
- Amazon EventBridge1.00
- Catches RALED state euents0.95
- AnEnd-to-EndRLPipeline1.00
- Health Agent for Anomaly0.95
- Detection and Autonomous0.96
- CloudWatch Logs0.98
- AWS Lambda1.00
- Glue Data Catalog1.00
- Self-Healing of Cloud Data0.95
- Pipelines1.00
- RL Agent (Decide + Act)0.99
- Action selection engine0.99
- AWS Glue API1.00
- Amazon 531.00
- Starpoblun Inetrigger)0.86
- EventBridge / CloudWatch0.98
- Existing Pipeline0.97
- RL Decision Engine1.00
- Lambda (Agent Runtime)1.00
- Glue / Data Catalog0.98
- 53 (Artifacts)0.95
- NOTE: GENERALIZED PUBLIC REFERENCE ARCHITECTURE.0.99
- The1.00
- Monitor → Diagnose → Score → Decide → Safety → Act → Validate0.98
- 1.Observe logs, schema, and data-quality conditions0.99
- 2.Diagnose the likely failure family1.00
- Solution1.00
- 3.Estimate operational risk0.98
- 4.Select a bounded remediation action1.00
- 5.Validate whether the action restored a healthy state0.99
-
- Deterministic Anomaly Rules1.00
- Schema drift, null spikes, field0.98
- removals, type changes1.00
- Q-Learning Decision Policy0.99
- Retry, coerce schema, rollback,1.00
- quarantine, escalate, log0.98
- The1.00
- Safety Override0.99
- Intelligence1.00
- Critical anomaly + passive action1.00
- →escalate1.00
- Layer1.00
-
- Detectionand Diagnosis0.98
- Establish the Facts Before Choosing an Action0.99
- Component1.00
- Responsibility1.00
- Schema profiler0.99
- Structure, types, nesting, null rates0.98
- Drift detector1.00
- Additions, removals, type changes0.98
- Data-quality analyzer0.99
- Completeness, validity, consistency0.99
- Error classifier0.99
- Converts log patterns into failure categories1.00
- Risk scorer1.00
- Produces an operational risk level0.99
- Deterministic prototype with limited dataset; ML-ready with richer incident history.1.00
-
- RL Selects the Response, Not the Facts1.00
- Action:1.00
- State:1.00
- retry1.00
- coerce1.00
- rollback1.00
- quarantine1.00
- escalate1.00
- log1.00
- 1.Failure category1.00
- 2.Risk level1.00
- 1.Tabular Q-learning1.00
- 3.Retry count1.00
- 2.Small, interpretable state space1.00
- 4.Drift severity1.00
- 3.Low-memory inference0.99
- 5.Data-quality condition1.00
- 4.Inspectable Q-values for every decision1.00
- *TECHNICALLY, THIS IS A SINGLE-STEP CONTEXTUAL DECISION PROBLEM IMPLEMENTED WITH TABULAR Q-LEARNING.0.99
-
- Safe Autonomy0.99
- The Learned Policy Does Not Have Final Authority1.00
- 1.Q-learning policy proposes an action1.00
- 2.Safety layer evaluates critical anomalies1.00
- 3.Unsafe passive actions are overridden0.99
- 4.High-risk and unknown conditions escalate1.00
- 5.Every decision produces an audit record1.00
- Escalation is a correct action, not a failure of autonomy.1.00
-
- Example Failure1.00
- "job_name": " synthetic_etl_job ",0.98
- "triggered_at": "2026-05-21T02:05:20Z",0.99
- [A] Real Glue job failure0.96
- "error_classification": {1.00
- "error_type": "DATETIME FORMAT ERROR",0.97
- [B] Classified by Error Classifier0.99
- "confidence": 0.90,1.00
- "root_cause": "Spark 3 datetime format incompatibility",1.00
- "recommended_action": "APPLY_SCHEMA_COERCION" [C] Rule engine says: coerce0.99
- schema/date format0.99
- },0.90
- "selected_action": "APPLY_SCHEMA_COERCION",1.00
- ← [D] Agent selects schema coercion0.98
- "anomaly_override": false,1.00
- [E] Safety override did not fire0.97
- "remediation": {1.00
- "success": false,0.99
- "note": "Manual Glue script update required"0.99
- [F] Logged for review0.99
- }0.98
- }0.98
- Observed condition1.00
- Decision1.00
- 1.Glue-style job failure event0.99
- 1.Policy selects schema coercion1.00
- received1.00
- 2.Safety override does not fire0.99
- 2.Datetime-format incompatibility1.00
- 3.Automatic coercion is unavailable1.00
- detected1.00
- 4.Incident is logged for manual review0.98
- 3.Error classified with 0.900.98
- confidence1.00
-
- Reproducible Evaluation1.00
- Designed for Independent Reproduction1.00
- ambenzon27/rl-etl-1.00
- 1.Generalized AWS Lambda-style1.00
- remediation-agent1.00
- architecture1.00
- RL-guided ETL remediation agent with schema-drift0.98
- detection, error classification, Q-learning decisions,0.99
- 2.Synthetic schemas, records, logs, and1.00
- synthetic benchmarks, and AWS Lambda1.00
- deployment templates.1.00
- 810.72
- ⊙00.68
- ☆00.87
- y o0.54
- C0.54
- incidents1.00
- Contributor1.00
- Issues0.99
- Stars1.00
- Forks1.00
- 3.No production data or infrastructure0.99
- ambenzon27/rl-etl-remediation-agent: RL-guided ETL remediation agent1.00
- with schema-drift detection, error classification, Q-learning...1.00
- identifiers1.00
- RL-guided ETL remediation agent with schema-drift detection, error classification,0.98
- Q-learning decisions, synthetic benchmarks, and AWS Lambda deployment1.00
- 4.Four controlled experiments: El-E40.99
- templates. - ambenzon27/rl-etl-remediation-a...0.97
- OGitHub0.86
- 5.Robustness check across 30 runs, seeds0.99
- 42-711.00
- 6.Results reported with 95% confidence0.99
- Public benchmark and tests are1.00
- intervals1.00
- available in the GitHub1.00
- repository.1.00
-
- Evaluation Results1.00
- Controlled synthetic benchmark0.99
- MTTR Drops from Days to Minutes in the Benchmark1.00
- - 30 seeds, mean ± 95% Cl0.94
- Mean time to recovery, shown on a log scale1.00
- 10,0001.00
- 2.5 working days0.99
- 1.Rule-based anomaly detector:1.00
- Precision 1.0001.00
- 1,0001.00
- ~99.85% lower MTTR1.00
- MTnn utes0.79
- 2.Recall: 0.8001.00
- 3.F1: 0.8890.96
- 1001.00
- 4.RL successful-case resolution time:1.00
- approximately 5.2 minutes1.00
- 101.00
- 5.24 +1- 0.14 ryn0.93
- 5.RL simulated success rate: 74.63 ± 1.51%1.00
- 6.RL non-escalation rate: 88.63 ± 0.89%0.98
- Manual1.00
- RL health agent1.00
- incident response1.00
- synthetic benchmark1.00
- Controlled synthetic benchmark; not production incident evidence.0.99
-
- Robustness and Ablation0.99
- What produced reliability?1.00
- 1.RL success matched the0.98
- deterministic policy: 0.00 ± 0.190.98
- 30 synthetic seeds; mean ± 95% Cl0.99
- percentage-point difference1.00
- 2.Deterministic rules beat random1.00
- −15.03 ± 0.66 pp0.97
- selection by 15.63 ± 1.86 points0.99
- Safety override1.00
- H0.79
- 3.The safety override reduced0.98
- vs none1.00
- +0.00 ± 0.19 pp0.98
- non-escalation by 15.03 ± 0.661.00
- RL vs rules1.00
- points1.00
- +15.63 ± 1.86 pp1.00
- 4.RL provided an inspectable1.00
- Rules vs random1.00
- learned policy, but did not1.00
- outperform rules in this1.00
- -151.00
- -101.00
- -51.00
- 01.00
- 51.00
- 101.00
- 151.00
- benchmark1.00
- Difference (percentage points)1.00
- 5.Structured decision'logic and0.98
- external guardrails produced1.00
- most of the reliability1.00
Transcript
146 cues· 1,835 words· 12,067 chars
- 0:00 Imagine you are this engineer.
- 0:03 A production data job failed hours ago.
- 0:06 The dashboard went stale.
- 0:08 You have spent all day checking the logs, the schema, and the upstream data.
- 0:12 And now it is past midnight.
- 0:15 The same question keeps coming back.
- 0:18 What changed?
- 0:19 The failure itself may be small, but the expensive part is everything around it.
- 0:26 inspection, diagnosis, choosing a safe response, re-running the job, and confirming that we did not make the data worse.
- 0:35 Hi, I'm Anna Marie Benzon.
- 0:37 In this talk, I will show an RL-guided system that selects bounded remediation action for ETL failures.
- 0:45 The central question is not simply whether an agent can act, but whether it can act usefully, explainably, and within boundaries that an operations team would actually trust.
- 0:59 Cloud ETL failures rarely arrive as one clean, well-labeled exemption.
- 1:04 We see late or unavailable sources, schema drift, daytime incompatibilities, null rate spikes, type changes, and runtime errors that do not match anything in the runbook.
- 1:15 The usual response is a human workflow.
- 1:17 Inspect the logs, form a diagnosis, attempt a repair, rerun the job, and validate the output.
- 1:25 Each step is reasonable.
- 1:27 The latency comes from handoffs in complete context and the need to avoid an unsafe fix.
- 1:34 In the capstone evaluation, the manual recovery baseline was modeled at roughly 2.5 working days.
- 1:41 This represents an incident moving through normal queuing, investigation, and approval.
- 1:47 So the engineering objective is specific.
- 1:50 Compress that loop for routine, recognizable failures while escalating the cases that are uncertain, novel, or high risk.
- 1:59 This diagram shows the end-to-end AWS architecture from my capstone.
- 2:04 An existing AWS Glue ETL job emits a job failed event.
- 2:09 Amazon EventBridge catches that event and triggers the Lambda function that runs the agent.
- 2:15 Lambda gathers evidence from two read-only sources.
- 2:19 CloudWatch provides the error logs, while the Glue Data Catalog provides the current schema metadata.
- 2:26 The system uses those signals to classify the failure, assess the data quality and operational risk.
- 2:33 and construct the state pass to the RL decision engine.
- 2:37 The policy then proposes a bounded response.
- 2:40 The safety layer checks that proposal before the executor can use the Glue API to re-trigger the job or apply an approved remediation.
- 2:50 Amazon S3 stores agent artifacts, audit logs, and quarantined outputs.
- 2:57 Finally, the job is rerun and validated.
- 3:01 So this is close operational look, monitor, diagnose, score, decide, check safety, act and verify recovery.
- 3:12 The capstone implementation use synthetic data provided by the client.
- 3:17 The public repository preserve this pattern through a sanitized generalized deployment template.
- 3:23 The intelligence layer deliberately separates three concerns.
- 3:27 Deterministic anomaly rules establish observable facts.
- 3:31 A field disappeared, a tide changed, or the null rate crossed a threshold.
- 3:35 The Q-learning policy handles contextual action selection.
- 3:38 Given the current incident state, should the system retry, coerce the schema, rollback, quarantine, escalate, or simply log the event?
- 3:48 Then, a safety override sits outside the learned policy.
- 3:52 For example, if the anomaly is critical and the policy proposes a passive action such as logging, the override converts that choice into an escalation.
- 4:02 This separation is the design thesis of the project, rules for facts, learning for bounded choices, and guardrails for authority.
- 4:11 Before selecting an action, the system has to establish what actually happened.
- 4:15 The schema profiler extracts a structure
- 4:18 types, nesting, and null rate statistics.
- 4:21 The drift detector compares the current profiler with the baseline area and identifies additions, removals, and type changes.
- 4:29 The data quality analyzer checks completeness, validity, and consistency.
- 4:34 The error classifier maps log patterns into failure families, and the risk scorer turns those signals into an operational risk level.
loading