Runbook Freshness and Decay Audit Prompt
Audit your runbook library for stale, broken, and untrusted procedures, then design a freshness program so on-call engineers can rely on runbooks instead of working around them.
- Target user
- SRE and on-call teams maintaining a runbook library
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a senior SRE who knows that a runbook nobody trusts is worse than no runbook, because it sends a stressed engineer down a dead end at 3 a.m. I will provide: - A sample of our runbooks (or their structure and metadata) - When each was last edited and last actually used in an incident - Owner / team per runbook - Recent incidents where a runbook was wrong, missing, or ignored Run a runbook freshness and decay audit. Work through these steps: 1. **Define decay signals** — the indicators that a runbook is rotting: stale edit date, dead links, commands referencing retired systems, steps that no longer match the architecture, no owner, never used despite relevant incidents. 2. **Score each runbook** — rate freshness and trustworthiness, and bucket into keep, fix, rewrite, or retire. Flag the dangerous ones (confidently wrong) above the merely outdated. 3. **Find the silent gaps** — incidents that recurred without a runbook, and runbooks that exist but were bypassed (a signal they are not trusted). 4. **Diagnose why they decay** — no ownership, no trigger to update after architecture changes, no validation, write-once-never-read culture. 5. **Design a freshness program** — ownership model, a review cadence tied to usage and to relevant deploys, a "last verified" stamp, and a lightweight validation (dry-run or gameday) for high-stakes runbooks. 6. **Close the loop with incidents** — make "update the runbook" a standard postmortem action item, and make runbooks improve every time they are used. Output: (a) a decay-signal rubric, (b) a per-runbook scorecard with keep/fix/rewrite/retire, (c) the dangerous-runbook shortlist, (d) the freshness program design with cadence and ownership, (e) the postmortem-to-runbook feedback loop. Prioritize trustworthiness over volume; a small set of verified runbooks beats a wiki full of guesses.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
On-Call Runbook Authoring Standard Prompt
Define a house style and quality bar for writing operational runbooks so every page links to a clear, copy-pasteable, low-ambiguity procedure an exhausted on-call can follow at 3 a.m.
-
Runbook Dry-Run Validation Prompt
Stress-test a runbook before you trust it in a real incident — walk each step for ambiguity, missing preconditions, dangerous commands, and dead ends — so it actually works at 3am under pressure.
-
Operational Runbook Generator Prompt
Turn tribal knowledge into a battle-tested operational runbook that a first-time responder can execute safely at 3am — with verification steps, rollback paths, and escalation off-ramps.
-
Runbook Gap Analysis From Incidents Prompt
Mine past incidents to find where responders lacked a runbook, where existing runbooks failed, and produce a prioritized list of runbooks to write or fix — with the specific steps each one needs.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.