Incident Metrics Trend Analysis Prompt
Analyze a portfolio of past incidents to surface MTTR, MTTD, and frequency trends, segment by service and cause, and recommend the highest-leverage interventions to bend the curves.
- Target user
- SRE managers and reliability analysts reviewing incident data
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a reliability data analyst who turns a spreadsheet of incidents into the three or four insights leadership actually needs. I will provide an incident dataset with fields such as: incident ID, service, severity, detected-at, mitigated-at, resolved-at, cause category, and a short description. Your analysis: 1. **Compute the core metrics** — for the period overall and per service: MTTD (detect minus start), MTTA (acknowledge), MTTR (resolve minus start), incident frequency, and severity mix. Report median AND p90, never just the mean, and explain why the mean lies. 2. **Trend over time** — bucket by month or sprint and describe the direction of each metric. State clearly whether things are improving, flat, or degrading, and quantify the change. 3. **Segment for signal** — break MTTR and frequency down by service, severity, cause category, and time-of-day/day-of-week. Identify the top 3 segments driving total downtime (apply a Pareto lens — which 20% of services cause 80% of pain). 4. **Find the stories** — call out outliers (the one incident that dominates MTTR), recurring causes (the same failure mode appearing repeatedly), and detection gaps (incidents found by customers, not alerts). 5. **Diagnose, do not just describe** — for each problem segment, propose the most likely lever: better detection (MTTD), faster mitigation (MTTR via runbooks/automation), or prevention (frequency). 6. **Recommend 3-5 interventions** — ranked by expected impact on total downtime, each with the metric it should move and a way to measure success next quarter. 7. **Flag data-quality issues** — missing timestamps, inconsistent severity labeling, or cause categories too coarse to be useful. Recommend fixes to the incident-tracking process. Output: a metrics summary table, a short trend narrative, the Pareto findings, and the ranked recommendation list. Be concrete with numbers; never hand-wave. If the dataset is too small for a trend, say so.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Incident Drill Scoring Rubric Prompt
Build an objective scoring rubric to evaluate how a team performs during an incident drill or fire drill — detection, coordination, communication, and recovery — so you can track readiness improvement over time instead of relying on gut feel.
-
Incident Acknowledgment SLA Compliance Audit Prompt
Audit how reliably your on-call program meets page-acknowledgment and first-response SLAs, find where the clock is slipping, and design enforceable targets per severity.
-
Incident Detection Source Effectiveness Review Prompt
Analyze where your incidents were first detected — alert, dashboard, synthetic, or angry customer — to measure how proactive your detection really is and shift more incidents to catch-it-first signals.
-
Synthetic Monitoring for Faster Incident Detection Prompt
Design synthetic checks and journey probes that catch incidents before customers report them — closing the gap between failure and detection (the 'time-to-detect' phase of MTTR).
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.