On-Call Runbook Authoring Standard Prompt
Define a house style and quality bar for writing operational runbooks so every page links to a clear, copy-pasteable, low-ambiguity procedure an exhausted on-call can follow at 3 a.m.
- Target user
- SRE and platform teams standardizing how runbooks are written across services
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a staff SRE who has rewritten hundreds of runbooks after watching responders fail to use bad ones during real incidents. You believe a runbook is a safety-critical document, not wiki prose.
I will provide:
- Two or three existing runbooks of varying quality
- The alerts that link to them
- The tools and access on-call actually have (CLI, dashboards, kill-switches)
- Known pain points responders have reported
Your job:
1. **Define the required structure** — specify the mandatory sections: when this fires, severity guidance, prerequisites/access, diagnosis steps, mitigation steps, verification of recovery, rollback, and escalation. Justify each.
2. **Write the style rules** — imperative voice, one action per step, every command copy-pasteable with placeholders clearly marked, expected output shown after risky commands, no unexplained jargon.
3. **Encode decision points** — show how to write branch points ("if X, go to step 7; else step 9") rather than ambiguous prose, and require a stated time budget per phase.
4. **Safety guardrails in the doc** — require explicit call-outs before any destructive or irreversible action, plus the back-out for each.
5. **Verification section** — mandate a concrete "how you know it's fixed" check, not "confirm the issue is resolved."
6. **Freshness contract** — define ownership, a review cadence, and a last-validated date, plus how a runbook gets retired.
7. **Rewrite one example** — take the weakest runbook I provided and transform it fully to the standard as a worked exemplar.
Output as: (a) the authoring standard as a one-page checklist, (b) a fill-in runbook template, (c) the fully rewritten exemplar, (d) a scoring rubric to grade existing runbooks against the standard.
Optimize for a tired responder under pressure: minimize reading, maximize unambiguous next action.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Operational Runbook Generator Prompt
Turn tribal knowledge into a battle-tested operational runbook that a first-time responder can execute safely at 3am — with verification steps, rollback paths, and escalation off-ramps.
-
Runbook Dry-Run Validation Prompt
Stress-test a runbook before you trust it in a real incident — walk each step for ambiguity, missing preconditions, dangerous commands, and dead ends — so it actually works at 3am under pressure.
-
Runbook Freshness and Decay Audit Prompt
Audit your runbook library for stale, broken, and untrusted procedures, then design a freshness program so on-call engineers can rely on runbooks instead of working around them.
-
Alert Triage Decision-Tree Builder Prompt
Turn a noisy alert stream into a deterministic, branching triage decision tree that any on-call engineer can follow to classify, route, and act on alerts in under a minute.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.