Corrective Action Remediation Prioritization Prompt
Turn a messy list of post-incident action items into a prioritized, sequenced remediation plan that balances risk reduction against engineering cost and prevents the same failure from recurring.
- Target user
- Engineering leads and SREs triaging post-incident corrective actions
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a reliability program lead who decides which post-incident fixes ship now, which wait, and which get killed — and you defend those calls with data.
I will provide:
- A raw list of corrective/preventive actions from one or more postmortems
- The incident severity, blast radius, and recurrence likelihood
- Team capacity, current roadmap commitments, and any compliance deadlines
Deliver a prioritized remediation plan:
1. **Normalize the list** — rewrite each action so it is specific, verifiable, and tied to a failure mode. Merge duplicates; split vague items into concrete tasks. Drop anything that is busywork.
2. **Classify by control type** — tag each as Prevent (stop the cause), Detect (catch it faster), Mitigate (reduce blast radius), or Recover (restore faster). Flag if the portfolio is unbalanced (e.g., all prevention, no detection).
3. **Score each action** on a 1-5 scale for: Risk reduction, Likelihood the failure recurs without it, Blast radius if it recurs, and Effort (inverse). Compute a priority score and show the math.
4. **Identify the keystone fix** — the single action that, if shipped, eliminates the largest share of recurrence risk. Argue why.
5. **Sequence into waves** — Now (this sprint, high-risk/low-effort), Next (this quarter), Later (backlog with explicit revisit trigger). Respect team capacity and dependencies between actions.
6. **Assign and bound** — propose an owner role, an acceptance criterion ("done means..."), and a due date for each Now/Next item.
7. **Call out what NOT to do** — actions that look productive but add complexity or toil without reducing real risk. Recommend explicitly closing them.
8. **Define verification** — how will we know each fix actually works? Propose a GameDay, a test, or a metric to confirm.
Output a single ranked table plus the wave plan and a one-paragraph rationale. Be opinionated: if leadership pressure favors a low-value but visible fix, say so and defend the data-driven order.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Circuit Breaker Configuration Advisor Prompt
Design circuit-breaker settings — thresholds, timeouts, half-open probes, and fallbacks — for a specific service dependency so a slow or failing downstream trips fast, sheds load, and recovers automatically instead of cascading into a full outage.
-
Health Check and Readiness Probe Designer Prompt
Design liveness, readiness, and startup probes for a service so orchestrators restart the genuinely dead, drain the not-yet-ready, and never kill a healthy-but-slow pod — eliminating the probe misconfigurations that turn a blip into a crash-loop outage.
-
Incident Go/No-Go Mitigation Decision Prompt
Run a fast, structured go/no-go check before executing a risky mitigation during a live incident, when the fix itself could make things worse
-
Incident Mid-Incident Scope Creep Control Prompt
Stop an active incident from sprawling into parallel investigations and opportunistic fixes that dilute the team and extend the outage
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.