read-only demo

Videos 0RNNfxpdbQk

Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest

index_state ready data_status ok

AI Engineer· published 2026-07-20· 0:11:20· en-US· indexed 2026-08-10 19:42

Open on YouTube

Scene timeline

  1. Shot 0, 0:00 to 0:16, 1 of 1 keyframes kept
  2. Shot 1, 0:16 to 0:24, 1 of 1 keyframes kept
  3. Shot 2, 0:24 to 1:01, 1 of 1 keyframes kept
  4. Shot 3, 1:01 to 1:38, 0 of 1 keyframes kept
  5. Shot 4, 1:38 to 2:03, 1 of 1 keyframes kept
  6. Shot 5, 2:03 to 2:28, 1 of 1 keyframes kept
  7. Shot 6, 2:28 to 2:58, 1 of 1 keyframes kept
  8. Shot 7, 2:58 to 3:25, 1 of 1 keyframes kept
  9. Shot 8, 3:25 to 3:52, 0 of 1 keyframes kept
  10. Shot 9, 3:52 to 4:34, 1 of 1 keyframes kept
  11. Shot 10, 4:34 to 5:07, 1 of 1 keyframes kept
  12. Shot 11, 5:07 to 5:31, 1 of 1 keyframes kept
  13. Shot 12, 5:31 to 5:57, 1 of 1 keyframes kept
  14. Shot 13, 5:57 to 6:22, 0 of 1 keyframes kept
  15. Shot 14, 6:22 to 6:48, 0 of 1 keyframes kept
  16. Shot 15, 6:48 to 7:19, 1 of 1 keyframes kept
  17. Shot 16, 7:19 to 7:50, 0 of 1 keyframes kept
  18. Shot 17, 7:50 to 8:20, 0 of 1 keyframes kept
  19. Shot 18, 8:20 to 8:54, 1 of 1 keyframes kept
  20. Shot 19, 8:54 to 9:27, 0 of 1 keyframes kept
  21. Shot 20, 9:27 to 9:43, 1 of 1 keyframes kept
  22. Shot 21, 9:43 to 10:25, 1 of 1 keyframes kept
  23. Shot 22, 10:25 to 11:06, 1 of 1 keyframes kept
  24. Shot 23, 11:06 to 11:15, 1 of 1 keyframes kept
  25. Shot 24, 11:15 to 11:20, 1 of 1 keyframes kept

25 shot(s).

keyframes kept every frame deduplicated

What was stored

cues
98
whisperx 98
chunks
20
from 98 cues
keyframes
18
kept of 25 captured
frames with text
17
254 lines read
chapters
0
from the source metadata
keyframe bytes
2.6 MB
word timings on 98 cues

Provenance

Each pipeline stage, its state and the model that produced it
stage state model started took
fetch done 2026-08-09 23:41 0s
stt done 2026-08-09 14:56 11s
chunk done 2026-08-09 14:56 0s
text_embed done 2026-08-10 19:42 0s
keyframe done 2026-08-09 14:56 1m 04s
ocr done 2026-08-09 14:58 12s
frame_embed done 2026-08-10 19:42 3s

Frames, and what the machine read

  • 0:11 #0 done5 line(s)

    shot 0·sharpness 2539.6

    1. Medic for Apache Spark1.00
    2. Save1.00
    3. First Aid for Failing Jobs0.99
    4. Drasko Profirovic1.00
    5. Staff Engineer @ Pinterest0.99
  • 0:18 #1 done5 line(s)

    shot 1·sharpness 1410.0

    1. Agenda1.00
    2. Motivation1.00
    3. Our Journey building Medic for Apache Spark0.99
    4. Lessons learnt1.00
    5. Upcoming opportunities0.99
  • 0:39 #2 done11 line(s)

    shot 2·sharpness 4189.2

    1. Motivation1.00
    2. Initial support for Spark jobs at Pinterest1.00
    3. Pager rotation for critical jobs1.00
    4. Slack channel for less urgent questions1.00
    5. What makes diagnosing Spark hard1.00
    6. Disparate data sources to review0.99
    7. Company specific patches1.00
    8. Job owners have varying degrees of familiarity with Spark1.00
    9. Business impact by improving the status quo1.00
    10. Reduce KTLO burden1.00
    11. Timelier remediation of issues0.99
  • 1:13 #3 skipped

    shot 3·duplicate of #2

  • 1:43 #4 done12 line(s)

    shot 4·sharpness 5423.1

    1. North Star design1.00
    2. High-confidence root cause identification1.00
    3. Focus on identifying one primary root cause1.00
    4. Evidence cited in the report0.98
    5. Excel with supported use cases, perform decently for less frequent scenarios, and1.00
    6. include mechanisms for continuous improvement1.00
    7. Opinionated, actionable remediation guidance0.99
    8. Provide prioritized next steps1.00
    9. Tie suggestions to Pinterest's best practices1.00
    10. Offer ready-to-use code snippets0.98
    11. Minimal friction for Spark users1.00
    12. Meet users where they currently operate1.00
  • 2:22 #5 done17 line(s)

    shot 5·sharpness 3419.7

    1. Building the Spark MCP Foundation0.99
    2. MCP server & deployment1.00
    3. Spark MCP server and K8s production deployment1.00
    4. Connecting to data sources0.99
    5. Job Submission Service for job metadata0.99
    6. Spark History Server1.00
    7. Logs1.00
    8. o0.50
    9. Driver / Executor0.97
    10. O0.71
    11. K8s Pod0.99
    12. Time series metrics1.00
    13. MCP Server0.98
    14. SHS1.00
    15. JSS1.00
    16. Logs1.00
    17. Metrics1.00
  • 2:48 #6 done12 line(s)

    shot 6·sharpness 2991.2

    1. Exposed a ReAct Agent0.98
    2. Manually tuned generic prompt which outlined1.00
    3. Problem solving approach1.00
    4. How to present the report1.00
    5. How to handle certain common failure1.00
    6. patterns, and their suggested fixes0.99
    7. MCP Server1.00
    8. ReAct Agent (Medic)1.00
    9. SHS1.00
    10. JSS1.00
    11. Logs1.00
    12. Metrics1.00
  • 3:04 #7 done10 line(s)

    shot 7·sharpness 5076.9

    1. Agent shortcomings1.00
    2. Prompt tuning presented several issues1.00
    3. Unbounded prompt length as we tried to handle edge cases1.00
    4. Difficulty in steering the agent's reasoning and outcomes1.00
    5. Limitations with context windows1.00
    6. Tool calls (e.g., logs, metrics) to the MCP would exceed the LLM models' token0.99
    7. limits1.00
    8. Manual testing lacked confidence0.99
    9. The absence of a regression framework made it impossible to accurately judge0.99
    10. the agent's quality across supported scenarios1.00
  • 3:28 #8 skipped

    shot 8·duplicate of #7

  • 4:05 #9 done15 line(s)

    shot 9·sharpness 4930.2

    1. Observability and Testability1.00
    2. OTel to track traces1.00
    3. Allows us to see the tool invocations and reasoning steps0.99
    4. Used Langfuse1.00
    5. Purpose built E2E test harness1.00
    6. Needed a way to snapshot production data for future replay0.99
    7. Used offline evals to grade the Agent's response based on key metrics0.99
    8. Expanded our corpus of test cases by1.00
    9. O0.84
    10. Identifying production use cases1.00
    11. O0.61
    12. Recording production state1.00
    13. O0.68
    14. Setting expectations in evals1.00
    15. Tuning prompts until evals pass1.00
  • 4:38 #10 done31 line(s)

    shot 10·sharpness 2036.8

    1. Testability1.00
    2. Playback Mode0.99
    3. Records Mode1.00
    4. Report1.00
    5. SHS1.00
    6. JSS1.00
    7. Logs1.00
    8. Metrics1.00
    9. Evals1.00
    10. Fixtures1.00
    11. Fixtures1.00
    12. Fixtures0.95
    13. Fixtures1.00
    14. 0.86
    15. E2E Test Harness1.00
    16. E2E Test Harness1.00
    17. MCP Server1.00
    18. ReAct Agent (Medic)0.99
    19. ReAct Agent (Medic)0.98
    20. SHS1.00
    21. JSS1.00
    22. Logs1.00
    23. Metrics1.00
    24. SHS1.00
    25. JSS1.00
    26. Logs1.00
    27. Metrics1.00
    28. Fixtures1.00
    29. Fixtures1.00
    30. Fixtures1.00
    31. Fixtures1.00
  • 5:14 #11 done19 line(s)

    shot 11·sharpness 7687.3

    1. Testability1.00
    2. root_cause_check evaluation:1.00
    3. Score:5/50.95
    4. Explanation: The response clearly identifies the expected root cause—schema mismatch or incompatible data types in Spark SQL—by specifying that1.00
    5. the 'volume' column was expected as 'bigint' but found as 'INT32'. It provides detailed supporting evidence from exception messages, stack traces, and0.99
    6. file paths, and explains why the error occurred. The diagnosis is thorough, with alternative hypotheses considered and ruled out, and actionable fixes0.99
    7. are provided.0.99
    8. suggested_fix_check evaluation:0.99
    9. Score: 5/50.93
    10. Explanation: The response provides multiple specific, actionable fixes that directly address schema alignment, including explicit casting, enabling1.00
    11. schema merging, and repairing Parquet files for schema consistency. Each fix includes clear implementation steps and validation guidance, fully1.00
    12. aligning with the expected fix theme.0.99
    13. root_cause_justification_check evaluation:1.00
    14. Score:4/51.00
    15. Explanation: The justification provides a clear and technically accurate explanation of a schema mismatch between expected and actual column types1.00
    16. in a Parquet file, with supporting evidence and actionable fixes. However, it does not explicitly connect the failure to joins or operations on tables with1.00
    17. incompatible schemas.1.00
    18. Overall: 7/7 evaluations passed0.99
    19. All evaluations passed!0.99
  • 5:49 #12 done12 line(s)

    shot 12·sharpness 4981.5

    1. Handling Logs0.98
    2. Problems1.00
    3. Multiple exceptions per job, often benign0.98
    4. Heuristic if-else rules were brittle & hard to maintain0.99
    5. Exception classifier pipeline0.98
    6. Continuous sampling of production jobs1.00
    7. Fingerprinting and clustering of logs0.99
    8. Improve signal to noise ratio by filtering out known red-herrings0.99
    9. Ranking driven by content relevance and time-decay scoring0.99
    10. MCP tools1.00
    11. Top-k truncated logs, ordered by their ranking score1.00
    12. Fully detailed logs by exception id0.98
  • 6:07 #13 skipped

    shot 13·duplicate of #12

  • 6:33 #14 skipped

    shot 14·duplicate of #12

  • 6:55 #15 done7 line(s)

    shot 15·sharpness 3923.3

    1. HandlingMetrics1.00
    2. Quarantined analysis of metrics data in sub-agent0.99
    3. Parent Agent provides instructions what observations the sub-agent needs to1.00
    4. make with the data1.00
    5. Moved from raw text time series data to encoding the data as a graph, attached1.00
    6. as an image to the conversation0.98
    7. Ability to collage multiple time-series datasets into one attached image1.00
  • 7:23 #16 skipped

    shot 16·duplicate of #15

  • 8:05 #17 skipped

    shot 17·duplicate of #15

  • 8:40 #18 done13 line(s)

    shot 18·sharpness 8303.9

    1. From single Agent to Multi Agent architecture1.00
    2. Built using langraph's deepagents library1.00
    3. Agents have a dedicated prompt and subset of MCP tools1.00
    4. Built-in tools to assist with staying on track and for exchange of information0.99
    5. Decomposed previous systems into a series of logical Agents0.99
    6. Supervisor that orchestrates the whole flow and return the response to the caller1.00
    7. Triage agent that determines the lifecycle of a Spark job and generates hypothesis1.00
    8. Research agents that are assigned a hypothesis to review and validate1.00
    9. Healer agent that generates suggested fixes for the identified root cause0.99
    10. Benefits1.00
    11. Reduced hallucinations by controlling the context window in each LLM convo0.99
    12. Improved readability of the system0.98
    13. Paved path for authoring specialized agents1.00
  • 9:14 #19 skipped

    shot 19·duplicate of #18

  • 9:31 #20 done29 line(s)

    shot 20·sharpness 864.4

    1. From single Agent to Multi Agent architecture0.99
    2. User Request1.00
    3. Job Name Parser - URL.0.97
    4. Extraction1.00
    5. Intent Classifier1.00
    6. Diagnosis Request1.00
    7. Supervisor Agent -0.98
    8. Orchestrator0.99
    9. Phase 1: Triage0.99
    10. Generic Query0.99
    11. Triage Agent - Job State0.96
    12. and Hypotheses1.00
    13. job_info0.87
    14. exceptions1.00
    15. analyze_spark_application_metrics1.00
    16. Job Info1.00
    17. Exceptions1.00
    18. Spark Metrics - Memory,0.98
    19. GC, 100.97
    20. Job State?0.99
    21. RUNNING1.00
    22. SUCCEEDED1.00
    23. FALED + Hypotheses0.93
    24. Recommend Wait0.98
    25. Optimization Suggstions0.96
    26. Failed Job Diagnosis - see0.97
    27. Diagram 20.99
    28. Generic React Agent0.99
    29. User Response0.98
  • 10:04 #21 done34 line(s)

    shot 21·sharpness 839.6

    1. From single Agent to Multi Agent architecture0.99
    2. EAILED + Hypotheses0.91
    3. Phase 3: Irvestigation0.99
    4. Research Agent 10.98
    5. Hypothesis 10.99
    6. Research Agent 20.98
    7. Hypothesis 21.00
    8. Research Agent 30.95
    9. Hypothesis 31.00
    10. Confidence Score1.00
    11. Evidence0.99
    12. Confidence Score0.97
    13. Evidence +0.98
    14. Confidence Score0.99
    15. Evidence +0.88
    16. Phase 4: Selectioion0.88
    17. ALL 30- Tools0.86
    18. ALL 30= Tools0.92
    19. ALL. 30+ Tools0.88
    20. Winting Root Cause0.97
    21. Highest Confidence1.00
    22. Phase 5: Healing0.99
    23. Solution Generation1.00
    24. Healer Agent1.00
    25. ALL 30+ Toois0.85
    26. answer_question_using_spark_documentation0.96
    27. Spark Analysis0.95
    28. Tools1.00
    29. Actionable Flers0.95
    30. Up to 31.00
    31. Phase 6: Final Report0.98
    32. Comprehensive Report0.98
    33. final_report.md0.94
    34. User Response0.93
  • 11:01 #22 done9 line(s)

    shot 22·sharpness 4097.4

    1. Learnings and What's next0.98
    2. What worked well1.00
    3. Multi-Agent architecture with built-in tools1.00
    4. Investing in ways to boost signal to noise ratio for logs0.99
    5. Hard-earned lessons0.98
    6. Using langraph workflows to represent the state of the system is brittle0.99
    7. Looking ahead0.99
    8. Moving away from prompt tuning for domain specific knowledge to agentic RAG1.00
    9. Extending the framework beyond Spark1.00
  • 11:11 #23 done13 line(s)

    shot 23·sharpness 1635.3

    1. Thank you1.00
    2. Project contributors1.00
    3. Frida1.00
    4. Yifei1.00
    5. Vanessa1.00
    6. Tucker1.00
    7. Jingyuan1.00
    8. Drasko1.00
    9. Special thanks to1.00
    10. Ang1.00
    11. Zaheen1.00
    12. Chen1.00
    13. Kingsley1.00

Transcript

98 cues· 1,570 words· 9,334 chars

  1. 0:01 Hi, my name is Drasko Profirovich.
  2. 0:04 I'm a staff engineer at Pinterest.
  3. 0:07 Today, I'll cover Medic for Apache Spark, which is our agentech diagnostics tool built to troubleshoot Spark failures.
  4. 0:16 We'll dive into why we built the Medic, the journey from prototype to the current architecture, lessons learned along the way, and what's next.
  5. 0:26 A bit of background about myself.
  6. 0:28 I had the opportunity to work at a few companies under the data platform org.
  7. 0:33 Despite many differences between those companies, there's at least one similarity.
  8. 0:38 The high bar for providing quality support to partner teams who rely on the infrastructure owned by the data platform org.
  9. 0:46 I'm sure I'm not alone when I say that the support rotation feels like a never ending stream of questions or problems to resolve.
  10. 0:55 Moreover, it's easy to forget how difficult it is to troubleshoot Spark or any distributed system for that matter.
  11. 1:02 This is particularly true for anyone just getting started with the framework.
  12. 1:07 The other challenge with supporting a load-bearing system comes down to ambiguous priorities.
  13. 1:12 Do you focus on helping one team with their failing job or do you unblock another team with a looming deadline?
  14. 1:20 It's not always straightforward to rank these asks, but as humans, we often have to decide how we'll spend our time.
  15. 1:28 The same is not true for LLMs.
  16. 1:30 We can easily scale out knowledge and capabilities on demand.
  17. 1:36 Our vision for a diagnostics agent was to ask it simply, why did a job fail?
  18. 1:41 And get back a deep research document, which provides evidence on the root cause of the failure.
  19. 1:47 The agent would also need to provide suggested fixes that are grounded in the context of the job.
  20. 1:54 Needless to say, this agent would need to be available in all the surfaces where our users operate today, like Slack or the Airflow UI, to name a few.
  21. 2:05 We started by exposing our data resources by way of the model context protocol as a way to connect them to the LLMs.
  22. 2:13 At this point, we could start an LLM conversation with the MCP tools enabled and ask the model to reason about our Spark job.
  23. 2:22 This worked in practice, but it required a lot of careful prompting from the human operator.
  24. 2:29 We extended our prototype by creating a single reasoning and acting agent, React for short.
  25. 2:36 The agent was given a single prompt, which embodied the problem solving approach it would take, how to structure the responses as a report, and specific examples to common failure patterns.
  26. 2:49 At this point, we had enough capabilities to start trialing the solution with our beta users.
  27. 2:55 From those early trials, we found a lot of shortcomings with our solution.
  28. 3:00 Prompt tuning became unsustainable.
  29. 3:03 One prompt had to do everything and adding detail in one area degraded the behavior in another.
  30. 3:10 Response quality was inconsistent.
  31. 3:13 Sometimes analysis was shallow or other times too verbose.
  32. 3:18 We lacked controls to keep the agent on track.
  33. 3:22 and we often hit context window issues for production jobs.
  34. 3:25 As an example, large tool outputs from logs would quickly consume tokens and brought a halt to the agent's reasoning.
  35. 3:35 Lastly, our end-to-end testing strategy up to this point relied on manual tests from production.
  36. 3:42 This felt anecdotal since production data would be retention the way.
  37. 3:47 Overall, it was hard to know if changes broke earlier wins.
  38. 3:53 To improve the system, we invested in observability and testability.
  39. 3:57 We used OpenTelemetry to publish traces to LangFuse, and by viewing the agent's execution as a waterfall diagram of steps, we could better understand the cause of lower quality responses.
  40. 4:11 The reliance on manual end-to-end testing highlighted the need for a more reliable and scalable solution.
  41. 4:18 We built an end-to-end test harness to snapshot production state,
  42. 4:23 and we could codify expectations as offline evaluations.
  43. 4:28 Lastly, this allowed us to tune our prompt based on the results.
  44. 4:34 In practice, the end-to-end test harness is simple.
  45. 4:38 In record mode, the agent calls real downstream systems and tool responses are captured as fixtures.
  46. 4:45 These are then saved to the file system and checked in as code.
  47. 4:49 Playback mode, the agent runs against fixtures instead of production data, but this time performs the analysis and generates the report.
  48. 4:59 The test suite then grades the report based on the offline evals we have authored.
  49. 5:05 For example, an offline eval might check for a limit of three suggested fixes.
  50. 5:11 The eval would score lower if the agent provided too many fixes towards managing the verbosity of the final report.

Open at this second