Best for
- USE WHEN active alerts, data broken, stale, pipeline failure, or investigate and fix a data incident.
monte-carlo-data/mc-agent-toolkit/skills/incident-response/SKILL.md
Orchestrate incident response — triage, root cause, remediate, prevent recurrence. USE WHEN active alerts, data broken, stale, pipeline failure, or investigate and fix a data incident.
Decision brief
This workflow orchestrates the full lifecycle of a data incident by sequencing existing Monte Carlo skills. It does not contain investigation or remediation logic itself — each step loads the relevant skill's SKILL.md which has the actual instructions.
Compatibility matrix
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/monte-carlo-data/mc-agent-toolkit --skill "skills/incident-response"Inspect the Agent Skill "monte-carlo-incident-response" from https://github.com/monte-carlo-data/mc-agent-toolkit/blob/3c88d016801b7a47be580d559cb3183ea3916cda/skills/incident-response/SKILL.md at commit 3c88d016801b7a47be580d559cb3183ea3916cda. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
Context detection routes here (active alerts detected + incident intent)
User wants to create monitors or check coverage without an active incident — use proactive monitoring workflow
Before starting, determine which step to enter based on the user's context:
Skill: Read and follow ../automated-triage/SKILL.md
Skill: Read and follow ../analyze-root-cause/SKILL.md
Permission review
No configured static risk pattern was detected
This is not proof of safety. Runtime behavior, indirect dependencies, and hidden external systems are outside the static scan.
Evidence record
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 86/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 90 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
This workflow orchestrates the full lifecycle of a data incident by sequencing existing Monte Carlo skills. It does not contain investigation or remediation logic itself — each step loads the relevant skill's SKILL.md which has the actual instructions.
Monte Carlo tool routing (required): Always call Monte Carlo MCP tools through this plugin's bundled server, whose fully-qualified tool names are
mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__<tool>(e.g.mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__get_alerts). Bare tool names used in this skill (get_alerts,search,get_table, …) refer to that bundled server. If the session also has a separately-configuredmonte-carlo-mcpserver, do not route to it — it may point at a different endpoint or credentials.
Activate when:
/mc-incident-responseprevent skill (auto-activates via hooks)asset-health directlyStep 1 (conditional): Triage — when user has multiple/unknown alerts
Step 2: Root Cause Analysis — the core investigation
Step 3: Remediation — fix or escalate
Step 4 (optional): Prevent Recurrence — add monitoring
Before starting, determine which step to enter based on the user's context:
alert_types starting with "Agent ", or the user's issue is about an AI agent) → for the investigation, read ../troubleshoot-agent-traces/SKILL.md instead of ../analyze-root-cause/SKILL.md; the remediation and monitoring steps still applySkill: Read and follow ../automated-triage/SKILL.md
Goal: Fetch recent alerts, score them by confidence and impact, identify which ones need investigation.
When to run: Only when the user doesn't already have a specific alert or incident to investigate. This step helps narrow down "I have alerts" into "these specific alerts need attention."
Scope MCP calls tightly. On large accounts, broad queries return hundreds of results, overflow the tool-result token limit, spill to disk, and force chunk reads — burning user tokens and exhausting the turn budget. Minimum scoping for tools this workflow touches:
get_alerts → time filter (created_after, default last 7 days) + at least one of warehouse, table_names, severitysearch → needed to resolve a table name to its MCON (get_table requires MCON). Always pass limit (e.g. 5), the table name as query, and filter by warehouse_uuid or database/schema. warehouse_types alone is too broad. If multiple matches return: (1) auto-pick the match whose warehouse_display_name matches the user's named warehouse — do NOT stop to ask; (2) failing that, prefer the is_key_asset: true match; (3) only ask the user when none of these resolve itget_monitors → filter by mcons or warehouse_uuidIf scope is missing, ask the user before calling: "Which warehouse?", "How far back — today, this week?", "Any specific severity?".
Transition to Step 2: Once high-priority alert(s) are identified, tell the user:
"I've identified [N] high-priority alerts. Let me investigate the root cause of [specific alert/table]. Moving to root cause analysis."
Then proceed to Step 2 with the identified alert context.
Skill: Read and follow ../analyze-root-cause/SKILL.md
Goal: Investigate why the issue occurred — trace lineage, check ETL changes, analyze query modifications, profile data.
This is the core step. Most workflow entries start here.
Investigate linearly — do not re-call tools. Walk through the investigation once: (1) find the table, (2) fetch its alerts and freshness, (3) check lineage, (4) check recent queries/ETL. Call each tool at most once per table. If a tool result is insufficient, move to the next signal rather than re-calling with different params — burning turns on redundant calls exhausts the budget before the root cause is reached.
Transition to Step 3: When the root cause is identified (or the investigation reaches its limit), summarize findings and tell the user:
"Root cause identified: [summary]. Would you like me to help remediate this, or is the investigation sufficient?"
If the user wants to proceed, move to Step 3. If they say "that's enough", stop.
Skill: Read and follow ../remediation/SKILL.md
Goal: Fix the issue using available tools, or escalate with full context if the fix requires actions outside the agent's capability.
Transition to Step 4: After remediation is complete (fix applied or escalation documented), offer prevention:
"The issue has been [fixed/escalated]. The root cause was [X]. Want me to help add a monitor to detect this type of issue earlier next time?"
If the user says yes, move to Step 4. If no, the workflow is complete.
Skill: Read and follow ../monitoring-advisor/SKILL.md
When loading monitoring-advisor for this step, frame the request as direct monitor creation — not coverage analysis. The user already knows what they want to monitor (the thing that just broke). Example framing:
"Based on the incident, I recommend adding a [freshness/volume/validation] monitor on [table]. Let me create the monitor configuration."
Goal: Add or update a monitor to catch this class of issue in the future.
Do not force this step. It is optional — offer it after remediation, and respect if the user declines.
Alternatives
alirezarezvani/claude-skills
App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist
wanshuiyin/Auto-claude-code-research-in-sleep
Use it for operations and research tasks; the detail page covers purpose, installation, and practical steps.
K-Dense-AI/scientific-agent-skills
Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.
K-Dense-AI/scientific-agent-skills
Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.