read-only demo

Videos LrGCT7G_rU8

Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon

index_state ready data_status ok

AI Engineer· published 2026-06-29· 0:14:40· en-US· indexed 2026-08-11 05:51

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:07, 1 of 1 keyframes kept
  2. Shot 1, 0:07 to 0:33, 1 of 1 keyframes kept
  3. Shot 2, 0:33 to 0:58, 0 of 1 keyframes kept
  4. Shot 3, 0:58 to 1:29, 1 of 1 keyframes kept
  5. Shot 4, 1:29 to 2:00, 0 of 1 keyframes kept
  6. Shot 5, 2:00 to 2:27, 1 of 1 keyframes kept
  7. Shot 6, 2:27 to 2:55, 0 of 1 keyframes kept
  8. Shot 7, 2:55 to 3:23, 0 of 1 keyframes kept
  9. Shot 8, 3:23 to 4:10, 1 of 1 keyframes kept
  10. Shot 9, 4:10 to 4:41, 1 of 1 keyframes kept
  11. Shot 10, 4:41 to 5:11, 0 of 1 keyframes kept
  12. Shot 11, 5:11 to 5:39, 1 of 1 keyframes kept
  13. Shot 12, 5:39 to 6:08, 0 of 1 keyframes kept
  14. Shot 13, 6:08 to 6:37, 1 of 1 keyframes kept
  15. Shot 14, 6:37 to 7:06, 0 of 1 keyframes kept
  16. Shot 15, 7:06 to 7:35, 1 of 1 keyframes kept
  17. Shot 16, 7:35 to 8:05, 0 of 1 keyframes kept
  18. Shot 17, 8:05 to 8:34, 1 of 1 keyframes kept
  19. Shot 18, 8:34 to 9:03, 0 of 1 keyframes kept
  20. Shot 19, 9:03 to 9:29, 1 of 1 keyframes kept
  21. Shot 20, 9:29 to 9:55, 0 of 1 keyframes kept
  22. Shot 21, 9:55 to 10:21, 0 of 1 keyframes kept
  23. Shot 22, 10:21 to 10:48, 1 of 1 keyframes kept
  24. Shot 23, 10:48 to 11:15, 0 of 1 keyframes kept
  25. Shot 24, 11:15 to 11:41, 0 of 1 keyframes kept
  26. Shot 25, 11:41 to 12:07, 1 of 1 keyframes kept
  27. Shot 26, 12:07 to 12:33, 0 of 1 keyframes kept
  28. Shot 27, 12:33 to 12:58, 1 of 1 keyframes kept
  29. Shot 28, 12:58 to 13:24, 0 of 1 keyframes kept
  30. Shot 29, 13:24 to 13:50, 1 of 1 keyframes kept
  31. Shot 30, 13:50 to 14:15, 0 of 1 keyframes kept
  32. Shot 31, 14:15 to 14:16, 1 of 1 keyframes kept
  33. Shot 32, 14:16 to 14:40, 1 of 1 keyframes kept

33 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
146
whisperx 146
chunks
26
from 146 cues
keyframes
17
kept of 33 captured
frames with text
17
319 lines read
chapters
0
from the source metadata
keyframe bytes
7.2 MB
word timings on 146 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-11 05:48 1m 26s
stt done 2026-08-11 05:49 15s
chunk done 2026-08-11 05:50 0s
text_embed done 2026-08-11 05:50 0s
keyframe done 2026-08-11 05:50 1m 20s
ocr done 2026-08-11 05:51 8s
frame_embed done 2026-08-11 05:51 3s

Frames, and what the machine read

  • 0:05 #0 done6 line(s)

    shot 0·sharpness 1352.8

    1. Using RL-based Agent to1.00
    2. Detect and Remediat0.99
    3. Presenter:1.00
    4. Anna Marie Benzon1.00
    5. University of the Philippines, Diliman1.00
    6. github.com/ambenzon27/rl-etl-remediation-agent1.00
  • 0:27 #1 done7 line(s)

    shot 1·sharpness 1366.4

    1. Using RL-based Agent to1.00
    2. Detect and Remediate0.99
    3. ETL Pipeline Failures0.98
    4. Presenter:1.00
    5. Anna Marie Benzon1.00
    6. University of the Philippines, Diliman1.00
    7. github.com/ambenzon27/rl-etl-remediation-agent1.00
  • 0:38 #2 skipped

    shot 2·duplicate of #1

  • 1:20 #3 done19 line(s)

    shot 3·sharpness 2829.6

    1. Failure1.00
    2. Inspect Logs1.00
    3. Diagnose1.00
    4. Repair1.00
    5. Rerun1.00
    6. Validate1.00
    7. Late or unavailable source data0.99
    8. The1.00
    9. Schema drift1.00
    10. X0.95
    11. Problem1.00
    12. Datetime parsing incompatibilities0.99
    13. Null-rate spikes1.00
    14. X0.82
    15. Cloud ETL jobs break because of:0.99
    16. Type changes0.98
    17. Unknown runtime errors1.00
    18. STEN0.90
    19. Modeled manual MTTR ~ 2.5 working days0.99
  • 1:53 #4 skipped

    shot 4·duplicate of #3

  • 2:13 #5 done34 line(s)

    shot 5·sharpness 2933.3

    1. Phase 1 Architecture — Al Pipeline Health Agent0.99
    2. Incident-Response Loop Using 6 AWS Services0.99
    3. AWS Glue ETL Jobs1.00
    4. 1. job FALED evarnt0.85
    5. Amazon EventBridge1.00
    6. Catches RALED state euents0.95
    7. AnEnd-to-EndRLPipeline1.00
    8. Health Agent for Anomaly0.95
    9. Detection and Autonomous0.96
    10. CloudWatch Logs0.98
    11. AWS Lambda1.00
    12. Glue Data Catalog1.00
    13. Self-Healing of Cloud Data0.95
    14. Pipelines1.00
    15. RL Agent (Decide + Act)0.99
    16. Action selection engine0.99
    17. AWS Glue API1.00
    18. Amazon 531.00
    19. Starpoblun Inetrigger)0.86
    20. EventBridge / CloudWatch0.98
    21. Existing Pipeline0.97
    22. RL Decision Engine1.00
    23. Lambda (Agent Runtime)1.00
    24. Glue / Data Catalog0.98
    25. 53 (Artifacts)0.95
    26. NOTE: GENERALIZED PUBLIC REFERENCE ARCHITECTURE.0.99
    27. The1.00
    28. Monitor → Diagnose → Score → Decide → Safety → Act → Validate0.98
    29. 1.Observe logs, schema, and data-quality conditions0.99
    30. 2.Diagnose the likely failure family1.00
    31. Solution1.00
    32. 3.Estimate operational risk0.98
    33. 4.Select a bounded remediation action1.00
    34. 5.Validate whether the action restored a healthy state0.99
  • 2:41 #6 skipped

    shot 6·duplicate of #5

  • 3:17 #7 skipped

    shot 7·duplicate of #5

  • 3:28 #8 done12 line(s)

    shot 8·sharpness 1732.4

    1. Deterministic Anomaly Rules1.00
    2. Schema drift, null spikes, field0.98
    3. removals, type changes1.00
    4. Q-Learning Decision Policy0.99
    5. Retry, coerce schema, rollback,1.00
    6. quarantine, escalate, log0.98
    7. The1.00
    8. Safety Override0.99
    9. Intelligence1.00
    10. Critical anomaly + passive action1.00
    11. →escalate1.00
    12. Layer1.00
  • 4:25 #9 done15 line(s)

    shot 9·sharpness 4961.7

    1. Detectionand Diagnosis0.98
    2. Establish the Facts Before Choosing an Action0.99
    3. Component1.00
    4. Responsibility1.00
    5. Schema profiler0.99
    6. Structure, types, nesting, null rates0.98
    7. Drift detector1.00
    8. Additions, removals, type changes0.98
    9. Data-quality analyzer0.99
    10. Completeness, validity, consistency0.99
    11. Error classifier0.99
    12. Converts log patterns into failure categories1.00
    13. Risk scorer1.00
    14. Produces an operational risk level0.99
    15. Deterministic prototype with limited dataset; ML-ready with richer incident history.1.00
  • 4:50 #10 skipped

    shot 10·duplicate of #9

  • 5:28 #11 done19 line(s)

    shot 11·sharpness 2673.7

    1. RL Selects the Response, Not the Facts1.00
    2. Action:1.00
    3. State:1.00
    4. retry1.00
    5. coerce1.00
    6. rollback1.00
    7. quarantine1.00
    8. escalate1.00
    9. log1.00
    10. 1.Failure category1.00
    11. 2.Risk level1.00
    12. 1.Tabular Q-learning1.00
    13. 3.Retry count1.00
    14. 2.Small, interpretable state space1.00
    15. 4.Drift severity1.00
    16. 3.Low-memory inference0.99
    17. 5.Data-quality condition1.00
    18. 4.Inspectable Q-values for every decision1.00
    19. *TECHNICALLY, THIS IS A SINGLE-STEP CONTEXTUAL DECISION PROBLEM IMPLEMENTED WITH TABULAR Q-LEARNING.0.99
  • 5:59 #12 skipped

    shot 12·duplicate of #11

  • 6:33 #13 done8 line(s)

    shot 13·sharpness 2852.3

    1. Safe Autonomy0.99
    2. The Learned Policy Does Not Have Final Authority1.00
    3. 1.Q-learning policy proposes an action1.00
    4. 2.Safety layer evaluates critical anomalies1.00
    5. 3.Unsafe passive actions are overridden0.99
    6. 4.High-risk and unknown conditions escalate1.00
    7. 5.Every decision produces an audit record1.00
    8. Escalation is a correct action, not a failure of autonomy.1.00
  • 6:54 #14 skipped

    shot 14·duplicate of #13

  • 7:20 #15 done34 line(s)

    shot 15·sharpness 4691.2

    1. Example Failure1.00
    2. "job_name": " synthetic_etl_job ",0.98
    3. "triggered_at": "2026-05-21T02:05:20Z",0.99
    4. [A] Real Glue job failure0.96
    5. "error_classification": {1.00
    6. "error_type": "DATETIME FORMAT ERROR",0.97
    7. [B] Classified by Error Classifier0.99
    8. "confidence": 0.90,1.00
    9. "root_cause": "Spark 3 datetime format incompatibility",1.00
    10. "recommended_action": "APPLY_SCHEMA_COERCION" [C] Rule engine says: coerce0.99
    11. schema/date format0.99
    12. },0.90
    13. "selected_action": "APPLY_SCHEMA_COERCION",1.00
    14. ← [D] Agent selects schema coercion0.98
    15. "anomaly_override": false,1.00
    16. [E] Safety override did not fire0.97
    17. "remediation": {1.00
    18. "success": false,0.99
    19. "note": "Manual Glue script update required"0.99
    20. [F] Logged for review0.99
    21. }0.98
    22. }0.98
    23. Observed condition1.00
    24. Decision1.00
    25. 1.Glue-style job failure event0.99
    26. 1.Policy selects schema coercion1.00
    27. received1.00
    28. 2.Safety override does not fire0.99
    29. 2.Datetime-format incompatibility1.00
    30. 3.Automatic coercion is unavailable1.00
    31. detected1.00
    32. 4.Incident is logged for manual review0.98
    33. 3.Error classified with 0.900.98
    34. confidence1.00
  • 7:39 #16 skipped

    shot 16·duplicate of #15

  • 8:25 #17 done37 line(s)

    shot 17·sharpness 3739.7

    1. Reproducible Evaluation1.00
    2. Designed for Independent Reproduction1.00
    3. ambenzon27/rl-etl-1.00
    4. 1.Generalized AWS Lambda-style1.00
    5. remediation-agent1.00
    6. architecture1.00
    7. RL-guided ETL remediation agent with schema-drift0.98
    8. detection, error classification, Q-learning decisions,0.99
    9. 2.Synthetic schemas, records, logs, and1.00
    10. synthetic benchmarks, and AWS Lambda1.00
    11. deployment templates.1.00
    12. 810.72
    13. ⊙00.68
    14. ☆00.87
    15. y o0.54
    16. C0.54
    17. incidents1.00
    18. Contributor1.00
    19. Issues0.99
    20. Stars1.00
    21. Forks1.00
    22. 3.No production data or infrastructure0.99
    23. ambenzon27/rl-etl-remediation-agent: RL-guided ETL remediation agent1.00
    24. with schema-drift detection, error classification, Q-learning...1.00
    25. identifiers1.00
    26. RL-guided ETL remediation agent with schema-drift detection, error classification,0.98
    27. Q-learning decisions, synthetic benchmarks, and AWS Lambda deployment1.00
    28. 4.Four controlled experiments: El-E40.99
    29. templates. - ambenzon27/rl-etl-remediation-a...0.97
    30. OGitHub0.86
    31. 5.Robustness check across 30 runs, seeds0.99
    32. 42-711.00
    33. 6.Results reported with 95% confidence0.99
    34. Public benchmark and tests are1.00
    35. intervals1.00
    36. available in the GitHub1.00
    37. repository.1.00
  • 8:43 #18 skipped

    shot 18·duplicate of #17

  • 9:08 #19 done26 line(s)

    shot 19·sharpness 2969.4

    1. Evaluation Results1.00
    2. Controlled synthetic benchmark0.99
    3. MTTR Drops from Days to Minutes in the Benchmark1.00
    4. - 30 seeds, mean ± 95% Cl0.94
    5. Mean time to recovery, shown on a log scale1.00
    6. 10,0001.00
    7. 2.5 working days0.99
    8. 1.Rule-based anomaly detector:1.00
    9. Precision 1.0001.00
    10. 1,0001.00
    11. ~99.85% lower MTTR1.00
    12. MTnn utes0.79
    13. 2.Recall: 0.8001.00
    14. 3.F1: 0.8890.96
    15. 1001.00
    16. 4.RL successful-case resolution time:1.00
    17. approximately 5.2 minutes1.00
    18. 101.00
    19. 5.24 +1- 0.14 ryn0.93
    20. 5.RL simulated success rate: 74.63 ± 1.51%1.00
    21. 6.RL non-escalation rate: 88.63 ± 0.89%0.98
    22. Manual1.00
    23. RL health agent1.00
    24. incident response1.00
    25. synthetic benchmark1.00
    26. Controlled synthetic benchmark; not production incident evidence.0.99
  • 9:39 #20 skipped

    shot 20·duplicate of #19

  • 10:13 #21 skipped

    shot 21·duplicate of #19

  • 10:29 #22 done34 line(s)

    shot 22·sharpness 4761.2

    1. Robustness and Ablation0.99
    2. What produced reliability?1.00
    3. 1.RL success matched the0.98
    4. deterministic policy: 0.00 ± 0.190.98
    5. 30 synthetic seeds; mean ± 95% Cl0.99
    6. percentage-point difference1.00
    7. 2.Deterministic rules beat random1.00
    8. −15.03 ± 0.66 pp0.97
    9. selection by 15.63 ± 1.86 points0.99
    10. Safety override1.00
    11. H0.79
    12. 3.The safety override reduced0.98
    13. vs none1.00
    14. +0.00 ± 0.19 pp0.98
    15. non-escalation by 15.03 ± 0.661.00
    16. RL vs rules1.00
    17. points1.00
    18. +15.63 ± 1.86 pp1.00
    19. 4.RL provided an inspectable1.00
    20. Rules vs random1.00
    21. learned policy, but did not1.00
    22. outperform rules in this1.00
    23. -151.00
    24. -101.00
    25. -51.00
    26. 01.00
    27. 51.00
    28. 101.00
    29. 151.00
    30. benchmark1.00
    31. Difference (percentage points)1.00
    32. 5.Structured decision'logic and0.98
    33. external guardrails produced1.00
    34. most of the reliability1.00
  • 11:12 #23 skipped

    shot 23·duplicate of #22

Transcript

146 cues· 1,835 words· 12,067 chars

  1. 0:00 Imagine you are this engineer.
  2. 0:03 A production data job failed hours ago.
  3. 0:06 The dashboard went stale.
  4. 0:08 You have spent all day checking the logs, the schema, and the upstream data.
  5. 0:12 And now it is past midnight.
  6. 0:15 The same question keeps coming back.
  7. 0:18 What changed?
  8. 0:19 The failure itself may be small, but the expensive part is everything around it.
  9. 0:26 inspection, diagnosis, choosing a safe response, re-running the job, and confirming that we did not make the data worse.
  10. 0:35 Hi, I'm Anna Marie Benzon.
  11. 0:37 In this talk, I will show an RL-guided system that selects bounded remediation action for ETL failures.
  12. 0:45 The central question is not simply whether an agent can act, but whether it can act usefully, explainably, and within boundaries that an operations team would actually trust.
  13. 0:59 Cloud ETL failures rarely arrive as one clean, well-labeled exemption.
  14. 1:04 We see late or unavailable sources, schema drift, daytime incompatibilities, null rate spikes, type changes, and runtime errors that do not match anything in the runbook.
  15. 1:15 The usual response is a human workflow.
  16. 1:17 Inspect the logs, form a diagnosis, attempt a repair, rerun the job, and validate the output.
  17. 1:25 Each step is reasonable.
  18. 1:27 The latency comes from handoffs in complete context and the need to avoid an unsafe fix.
  19. 1:34 In the capstone evaluation, the manual recovery baseline was modeled at roughly 2.5 working days.
  20. 1:41 This represents an incident moving through normal queuing, investigation, and approval.
  21. 1:47 So the engineering objective is specific.
  22. 1:50 Compress that loop for routine, recognizable failures while escalating the cases that are uncertain, novel, or high risk.
  23. 1:59 This diagram shows the end-to-end AWS architecture from my capstone.
  24. 2:04 An existing AWS Glue ETL job emits a job failed event.
  25. 2:09 Amazon EventBridge catches that event and triggers the Lambda function that runs the agent.
  26. 2:15 Lambda gathers evidence from two read-only sources.
  27. 2:19 CloudWatch provides the error logs, while the Glue Data Catalog provides the current schema metadata.
  28. 2:26 The system uses those signals to classify the failure, assess the data quality and operational risk.
  29. 2:33 and construct the state pass to the RL decision engine.
  30. 2:37 The policy then proposes a bounded response.
  31. 2:40 The safety layer checks that proposal before the executor can use the Glue API to re-trigger the job or apply an approved remediation.
  32. 2:50 Amazon S3 stores agent artifacts, audit logs, and quarantined outputs.
  33. 2:57 Finally, the job is rerun and validated.
  34. 3:01 So this is close operational look, monitor, diagnose, score, decide, check safety, act and verify recovery.
  35. 3:12 The capstone implementation use synthetic data provided by the client.
  36. 3:17 The public repository preserve this pattern through a sanitized generalized deployment template.
  37. 3:23 The intelligence layer deliberately separates three concerns.
  38. 3:27 Deterministic anomaly rules establish observable facts.
  39. 3:31 A field disappeared, a tide changed, or the null rate crossed a threshold.
  40. 3:35 The Q-learning policy handles contextual action selection.
  41. 3:38 Given the current incident state, should the system retry, coerce the schema, rollback, quarantine, escalate, or simply log the event?
  42. 3:48 Then, a safety override sits outside the learned policy.
  43. 3:52 For example, if the anomaly is critical and the policy proposes a passive action such as logging, the override converts that choice into an escalation.
  44. 4:02 This separation is the design thesis of the project, rules for facts, learning for bounded choices, and guardrails for authority.
  45. 4:11 Before selecting an action, the system has to establish what actually happened.
  46. 4:15 The schema profiler extracts a structure
  47. 4:18 types, nesting, and null rate statistics.
  48. 4:21 The drift detector compares the current profiler with the baseline area and identifies additions, removals, and type changes.
  49. 4:29 The data quality analyzer checks completeness, validity, and consistency.
  50. 4:34 The error classifier maps log patterns into failure families, and the risk scorer turns those signals into an operational risk level.

Open at this second