Recurring Pattern Mining Across Postmortems Prompt
Analyze a corpus of past postmortems to surface systemic, recurring failure patterns — the same root cause wearing different hats — and recommend the few structural fixes that would prevent whole classes of incidents.
- Target user
- Reliability leads and SRE managers running quarterly incident reviews
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a reliability analyst who reads dozens of postmortems and sees the shape behind them — the same systemic weakness that keeps producing differently-named incidents. Help me mine a corpus of postmortems for the patterns that matter. I will paste/attach a set of postmortems (or their summaries, timelines, root causes, and action items). Your job: 1. **Normalize first** — extract a structured record per incident: trigger, contributing factors, root cause category, detection source, time-to-detect, time-to-mitigate, and the action items (and whether they were completed). 2. **Cluster by true cause, not symptom** — group incidents by underlying mechanism (e.g., "unbounded retry storm", "missing backpressure", "config change with no canary", "single point of failure in auth"). Two incidents with different services but the same mechanism belong together. 3. **Quantify each cluster** — count, total downtime, customer impact, and trend over time (is this pattern getting worse?). Rank clusters by aggregate pain, not frequency alone. 4. **Find the meta-patterns** — recurring weaknesses across clusters: detection always coming from customers, action items that never shipped, the same service appearing repeatedly, deploys clustering before incidents. 5. **Audit action-item follow-through** — what fraction of prior action items were completed? Which incidents would have been prevented if a prior, never-shipped action item had landed? Name them. 6. **Recommend structural fixes** — for the top 3 clusters, propose the one architectural or process change that neutralizes the whole class, not a per-incident patch. Estimate the blast-radius reduction. 7. **Flag what you can't conclude** — if the corpus is too small or biased toward one team, say so rather than over-generalizing. Output: (a) a normalized incident table, (b) a ranked cluster summary with counts and downtime, (c) a meta-pattern list, (d) an action-item follow-through scorecard, (e) the top 3 structural recommendations with expected impact. Bias toward: causes over symptoms, structural fixes over patches, and honesty about sample size.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Postmortem Latent Risk Extractor Prompt
Mine a postmortem for the latent, systemic risks an incident exposed but didn't directly trigger, so the team fixes the conditions that made the failure possible rather than only the immediate cause.
-
Blameless Root Cause Analysis Facilitation Prompt
Facilitate a rigorous blameless RCA that separates contributing factors from blame, surfaces systemic gaps, and produces durable action items — not a name-and-shame report.
-
Detailed Postmortem Document Generator Prompt
Generate a complete, publication-ready postmortem document from raw incident data — narrative timeline, impact quantification, contributing factors, and tracked action items in a consistent template.
-
Postmortem Runbook Adherence Checker Prompt
Compare what responders actually did during an incident against the documented runbook for that failure, surfacing where the runbook was skipped, wrong, missing, or out of date so the fix lands on the document and tooling, not the people.
More Post Mortems with AI prompts & error guides
Browse every Post Mortems with AI prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.