Videos 0RNNfxpdbQk
Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest
Scene timeline
25 shot(s).
keyframes kept every frame deduplicated
What was stored
- cues
- 98
- whisperx 98
- chunks
- 20
- from 98 cues
- keyframes
- 18
- kept of 25 captured
- frames with text
- 17
- 254 lines read
- chapters
- 0
- from the source metadata
- keyframe bytes
- 2.6 MB
- word timings on 98 cues
Provenance
| stage | state | model | started | took |
|---|---|---|---|---|
fetch |
done | — | 2026-08-09 23:41 | 0s |
stt |
done | — | 2026-08-09 14:56 | 11s |
chunk |
done | — | 2026-08-09 14:56 | 0s |
text_embed |
done | — | 2026-08-10 19:42 | 0s |
keyframe |
done | — | 2026-08-09 14:56 | 1m 04s |
ocr |
done | — | 2026-08-09 14:58 | 12s |
frame_embed |
done | — | 2026-08-10 19:42 | 3s |
Frames, and what the machine read
-
- Medic for Apache Spark1.00
- Save1.00
- First Aid for Failing Jobs0.99
- Drasko Profirovic1.00
- Staff Engineer @ Pinterest0.99
-
- Agenda1.00
- Motivation1.00
- Our Journey building Medic for Apache Spark0.99
- Lessons learnt1.00
- Upcoming opportunities0.99
-
- Motivation1.00
- Initial support for Spark jobs at Pinterest1.00
- Pager rotation for critical jobs1.00
- Slack channel for less urgent questions1.00
- What makes diagnosing Spark hard1.00
- Disparate data sources to review0.99
- Company specific patches1.00
- Job owners have varying degrees of familiarity with Spark1.00
- Business impact by improving the status quo1.00
- Reduce KTLO burden1.00
- Timelier remediation of issues0.99
-
- North Star design1.00
- High-confidence root cause identification1.00
- Focus on identifying one primary root cause1.00
- Evidence cited in the report0.98
- Excel with supported use cases, perform decently for less frequent scenarios, and1.00
- include mechanisms for continuous improvement1.00
- Opinionated, actionable remediation guidance0.99
- Provide prioritized next steps1.00
- Tie suggestions to Pinterest's best practices1.00
- Offer ready-to-use code snippets0.98
- Minimal friction for Spark users1.00
- Meet users where they currently operate1.00
-
- Building the Spark MCP Foundation0.99
- MCP server & deployment1.00
- Spark MCP server and K8s production deployment1.00
- Connecting to data sources0.99
- Job Submission Service for job metadata0.99
- Spark History Server1.00
- Logs1.00
- o0.50
- Driver / Executor0.97
- O0.71
- K8s Pod0.99
- Time series metrics1.00
- MCP Server0.98
- SHS1.00
- JSS1.00
- Logs1.00
- Metrics1.00
-
- Exposed a ReAct Agent0.98
- Manually tuned generic prompt which outlined1.00
- Problem solving approach1.00
- How to present the report1.00
- How to handle certain common failure1.00
- patterns, and their suggested fixes0.99
- MCP Server1.00
- ReAct Agent (Medic)1.00
- SHS1.00
- JSS1.00
- Logs1.00
- Metrics1.00
-
- Agent shortcomings1.00
- Prompt tuning presented several issues1.00
- Unbounded prompt length as we tried to handle edge cases1.00
- Difficulty in steering the agent's reasoning and outcomes1.00
- Limitations with context windows1.00
- Tool calls (e.g., logs, metrics) to the MCP would exceed the LLM models' token0.99
- limits1.00
- Manual testing lacked confidence0.99
- The absence of a regression framework made it impossible to accurately judge0.99
- the agent's quality across supported scenarios1.00
-
- Observability and Testability1.00
- OTel to track traces1.00
- Allows us to see the tool invocations and reasoning steps0.99
- Used Langfuse1.00
- Purpose built E2E test harness1.00
- Needed a way to snapshot production data for future replay0.99
- Used offline evals to grade the Agent's response based on key metrics0.99
- Expanded our corpus of test cases by1.00
- O0.84
- Identifying production use cases1.00
- O0.61
- Recording production state1.00
- O0.68
- Setting expectations in evals1.00
- Tuning prompts until evals pass1.00
-
- Testability1.00
- Playback Mode0.99
- Records Mode1.00
- Report1.00
- SHS1.00
- JSS1.00
- Logs1.00
- Metrics1.00
- Evals1.00
- Fixtures1.00
- Fixtures1.00
- Fixtures0.95
- Fixtures1.00
- 个0.86
- E2E Test Harness1.00
- E2E Test Harness1.00
- MCP Server1.00
- ReAct Agent (Medic)0.99
- ReAct Agent (Medic)0.98
- SHS1.00
- JSS1.00
- Logs1.00
- Metrics1.00
- SHS1.00
- JSS1.00
- Logs1.00
- Metrics1.00
- Fixtures1.00
- Fixtures1.00
- Fixtures1.00
- Fixtures1.00
-
- Testability1.00
- root_cause_check evaluation:1.00
- Score:5/50.95
- Explanation: The response clearly identifies the expected root cause—schema mismatch or incompatible data types in Spark SQL—by specifying that1.00
- the 'volume' column was expected as 'bigint' but found as 'INT32'. It provides detailed supporting evidence from exception messages, stack traces, and0.99
- file paths, and explains why the error occurred. The diagnosis is thorough, with alternative hypotheses considered and ruled out, and actionable fixes0.99
- are provided.0.99
- suggested_fix_check evaluation:0.99
- Score: 5/50.93
- Explanation: The response provides multiple specific, actionable fixes that directly address schema alignment, including explicit casting, enabling1.00
- schema merging, and repairing Parquet files for schema consistency. Each fix includes clear implementation steps and validation guidance, fully1.00
- aligning with the expected fix theme.0.99
- root_cause_justification_check evaluation:1.00
- Score:4/51.00
- Explanation: The justification provides a clear and technically accurate explanation of a schema mismatch between expected and actual column types1.00
- in a Parquet file, with supporting evidence and actionable fixes. However, it does not explicitly connect the failure to joins or operations on tables with1.00
- incompatible schemas.1.00
- Overall: 7/7 evaluations passed0.99
- All evaluations passed!0.99
-
- Handling Logs0.98
- Problems1.00
- Multiple exceptions per job, often benign0.98
- Heuristic if-else rules were brittle & hard to maintain0.99
- Exception classifier pipeline0.98
- Continuous sampling of production jobs1.00
- Fingerprinting and clustering of logs0.99
- Improve signal to noise ratio by filtering out known red-herrings0.99
- Ranking driven by content relevance and time-decay scoring0.99
- MCP tools1.00
- Top-k truncated logs, ordered by their ranking score1.00
- Fully detailed logs by exception id0.98
-
- HandlingMetrics1.00
- Quarantined analysis of metrics data in sub-agent0.99
- Parent Agent provides instructions what observations the sub-agent needs to1.00
- make with the data1.00
- Moved from raw text time series data to encoding the data as a graph, attached1.00
- as an image to the conversation0.98
- Ability to collage multiple time-series datasets into one attached image1.00
-
- From single Agent to Multi Agent architecture1.00
- Built using langraph's deepagents library1.00
- Agents have a dedicated prompt and subset of MCP tools1.00
- Built-in tools to assist with staying on track and for exchange of information0.99
- Decomposed previous systems into a series of logical Agents0.99
- Supervisor that orchestrates the whole flow and return the response to the caller1.00
- Triage agent that determines the lifecycle of a Spark job and generates hypothesis1.00
- Research agents that are assigned a hypothesis to review and validate1.00
- Healer agent that generates suggested fixes for the identified root cause0.99
- Benefits1.00
- Reduced hallucinations by controlling the context window in each LLM convo0.99
- Improved readability of the system0.98
- Paved path for authoring specialized agents1.00
-
- From single Agent to Multi Agent architecture0.99
- User Request1.00
- Job Name Parser - URL.0.97
- Extraction1.00
- Intent Classifier1.00
- Diagnosis Request1.00
- Supervisor Agent -0.98
- Orchestrator0.99
- Phase 1: Triage0.99
- Generic Query0.99
- Triage Agent - Job State0.96
- and Hypotheses1.00
- job_info0.87
- exceptions1.00
- analyze_spark_application_metrics1.00
- Job Info1.00
- Exceptions1.00
- Spark Metrics - Memory,0.98
- GC, 100.97
- Job State?0.99
- RUNNING1.00
- SUCCEEDED1.00
- FALED + Hypotheses0.93
- Recommend Wait0.98
- Optimization Suggstions0.96
- Failed Job Diagnosis - see0.97
- Diagram 20.99
- Generic React Agent0.99
- User Response0.98
-
- From single Agent to Multi Agent architecture0.99
- EAILED + Hypotheses0.91
- Phase 3: Irvestigation0.99
- Research Agent 10.98
- Hypothesis 10.99
- Research Agent 20.98
- Hypothesis 21.00
- Research Agent 30.95
- Hypothesis 31.00
- Confidence Score1.00
- Evidence0.99
- Confidence Score0.97
- Evidence +0.98
- Confidence Score0.99
- Evidence +0.88
- Phase 4: Selectioion0.88
- ALL 30- Tools0.86
- ALL 30= Tools0.92
- ALL. 30+ Tools0.88
- Winting Root Cause0.97
- Highest Confidence1.00
- Phase 5: Healing0.99
- Solution Generation1.00
- Healer Agent1.00
- ALL 30+ Toois0.85
- answer_question_using_spark_documentation0.96
- Spark Analysis0.95
- Tools1.00
- Actionable Flers0.95
- Up to 31.00
- Phase 6: Final Report0.98
- Comprehensive Report0.98
- final_report.md0.94
- User Response0.93
-
- Learnings and What's next0.98
- What worked well1.00
- Multi-Agent architecture with built-in tools1.00
- Investing in ways to boost signal to noise ratio for logs0.99
- Hard-earned lessons0.98
- Using langraph workflows to represent the state of the system is brittle0.99
- Looking ahead0.99
- Moving away from prompt tuning for domain specific knowledge to agentic RAG1.00
- Extending the framework beyond Spark1.00
-
- Thank you1.00
- Project contributors1.00
- Frida1.00
- Yifei1.00
- Vanessa1.00
- Tucker1.00
- Jingyuan1.00
- Drasko1.00
- Special thanks to1.00
- Ang1.00
- Zaheen1.00
- Chen1.00
- Kingsley1.00
Transcript
98 cues· 1,570 words· 9,334 chars
- 0:01 Hi, my name is Drasko Profirovich.
- 0:04 I'm a staff engineer at Pinterest.
- 0:07 Today, I'll cover Medic for Apache Spark, which is our agentech diagnostics tool built to troubleshoot Spark failures.
- 0:16 We'll dive into why we built the Medic, the journey from prototype to the current architecture, lessons learned along the way, and what's next.
- 0:26 A bit of background about myself.
- 0:28 I had the opportunity to work at a few companies under the data platform org.
- 0:33 Despite many differences between those companies, there's at least one similarity.
- 0:38 The high bar for providing quality support to partner teams who rely on the infrastructure owned by the data platform org.
- 0:46 I'm sure I'm not alone when I say that the support rotation feels like a never ending stream of questions or problems to resolve.
- 0:55 Moreover, it's easy to forget how difficult it is to troubleshoot Spark or any distributed system for that matter.
- 1:02 This is particularly true for anyone just getting started with the framework.
- 1:07 The other challenge with supporting a load-bearing system comes down to ambiguous priorities.
- 1:12 Do you focus on helping one team with their failing job or do you unblock another team with a looming deadline?
- 1:20 It's not always straightforward to rank these asks, but as humans, we often have to decide how we'll spend our time.
- 1:28 The same is not true for LLMs.
- 1:30 We can easily scale out knowledge and capabilities on demand.
- 1:36 Our vision for a diagnostics agent was to ask it simply, why did a job fail?
- 1:41 And get back a deep research document, which provides evidence on the root cause of the failure.
- 1:47 The agent would also need to provide suggested fixes that are grounded in the context of the job.
- 1:54 Needless to say, this agent would need to be available in all the surfaces where our users operate today, like Slack or the Airflow UI, to name a few.
- 2:05 We started by exposing our data resources by way of the model context protocol as a way to connect them to the LLMs.
- 2:13 At this point, we could start an LLM conversation with the MCP tools enabled and ask the model to reason about our Spark job.
- 2:22 This worked in practice, but it required a lot of careful prompting from the human operator.
- 2:29 We extended our prototype by creating a single reasoning and acting agent, React for short.
- 2:36 The agent was given a single prompt, which embodied the problem solving approach it would take, how to structure the responses as a report, and specific examples to common failure patterns.
- 2:49 At this point, we had enough capabilities to start trialing the solution with our beta users.
- 2:55 From those early trials, we found a lot of shortcomings with our solution.
- 3:00 Prompt tuning became unsustainable.
- 3:03 One prompt had to do everything and adding detail in one area degraded the behavior in another.
- 3:10 Response quality was inconsistent.
- 3:13 Sometimes analysis was shallow or other times too verbose.
- 3:18 We lacked controls to keep the agent on track.
- 3:22 and we often hit context window issues for production jobs.
- 3:25 As an example, large tool outputs from logs would quickly consume tokens and brought a halt to the agent's reasoning.
- 3:35 Lastly, our end-to-end testing strategy up to this point relied on manual tests from production.
- 3:42 This felt anecdotal since production data would be retention the way.
- 3:47 Overall, it was hard to know if changes broke earlier wins.
- 3:53 To improve the system, we invested in observability and testability.
- 3:57 We used OpenTelemetry to publish traces to LangFuse, and by viewing the agent's execution as a waterfall diagram of steps, we could better understand the cause of lower quality responses.
- 4:11 The reliance on manual end-to-end testing highlighted the need for a more reliable and scalable solution.
- 4:18 We built an end-to-end test harness to snapshot production state,
- 4:23 and we could codify expectations as offline evaluations.
- 4:28 Lastly, this allowed us to tune our prompt based on the results.
- 4:34 In practice, the end-to-end test harness is simple.
- 4:38 In record mode, the agent calls real downstream systems and tool responses are captured as fixtures.
- 4:45 These are then saved to the file system and checked in as code.
- 4:49 Playback mode, the agent runs against fixtures instead of production data, but this time performs the analysis and generates the report.
- 4:59 The test suite then grades the report based on the offline evals we have authored.
- 5:05 For example, an offline eval might check for a limit of three suggested fixes.
- 5:11 The eval would score lower if the agent provided too many fixes towards managing the verbosity of the final report.
loading