Runbook and Next-Step Surfacer Prompt
Match the live symptom to the right runbook and surface the exact command to run now — so the responder acts from a known-good procedure instead of improvising, shortening time-to-mitigate.
- Target user
- On-call SREs reaching for the right procedure under pressure
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT, Cursor
The prompt
You are a senior SRE who knows that the slowest part of a fix is often just finding the right runbook and the exact command — not running it. Help me match this incident to a procedure and surface the precise next step. Paste your inputs: - The symptom / working hypothesis: [WHAT IS BROKEN + LIKELY CAUSE] - The alert and service: [ALERT + SERVICE/ENV] - Available runbooks: [PASTE RUNBOOK TITLES + CONTENT, OR LINKS/EXCERPTS] - Constraints: [MAINTENANCE WINDOW? CHANGE FREEZE? APPROVALS NEEDED?] Do this: 1. **Match to runbooks** — from the runbooks I provided, identify the 1-3 that best fit the symptom and hypothesis. For each, give a one-line reason it matches and a confidence. If none clearly fit, say so plainly rather than forcing a match. 2. **Extract the relevant section** — from the top-matched runbook, pull the specific steps that apply to this situation, skipping the parts that don't. Note any prerequisites the runbook assumes. 3. **Surface the exact next command** — state the precise command or action to run now, with the real values from my context filled into the placeholders. Mark clearly whether it is read-only (safe to run) or mutating (changes production). 4. **Pre-flight the mutating step** — for any command that changes production, list what to check first, what the expected effect is, and how to undo it (the rollback command). Flag if it needs an approval or violates a freeze. 5. **Flag runbook gaps** — note anything stale, ambiguous, or missing in the runbook so it can be fixed after the incident. Output format: "BEST MATCH" (runbook + confidence), "STEPS THAT APPLY" (trimmed), "RUN NOW" (the exact command, labeled read-only or MUTATING), and "ROLLBACK". For any mutating command, present it as a proposal with its pre-flight checks and undo path — do not execute it, and make clear the human must run and own it. Rank matches by confidence; never invent a runbook step that isn't in what I gave you.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Why this prompt works
This serves the mitigate phase, where time leaks not in executing the fix but in locating the right procedure and translating its generic steps into the exact command for this service, this environment, this moment. A responder under pressure who has to read three runbooks and mentally substitute placeholders is burning recovery time on retrieval, not action.
The prompt does the retrieval and the substitution: it ranks candidate runbooks by fit, trims the matched one to the steps that actually apply, and fills in the real values so the responder sees a runnable command rather than a template. Confidence scores keep a weak match from masquerading as a strong one, and the explicit “none clearly fit” escape hatch prevents the model from forcing a procedure onto a novel failure.
The mutating-step guardrail is the crux. Mitigation is the one phase where the AI’s output could directly change production, so the prompt hard-separates read-only from mutating actions, requires a pre-flight and a rollback for anything that changes state, and keeps execution firmly with the human. That lets the responder move fast on a known-good procedure while retaining the pause-and-verify that stops a confident wrong runbook from turning a small incident into a large one.
Related prompts
-
Diagnosis Accelerator: Verify-First Hypotheses Prompt
Turn the opening burst of telemetry into a short, ranked list of diagnoses — each paired with a single command to confirm or kill it — so the team tests the likeliest cause first and shortens time-to-diagnose.
-
Post-Fix Verification Checklist Prompt
Build the queries and checks that confirm the fix actually resolved the incident — across the metric, the user, and the dependencies — before calling the all-clear, so you don't reopen later and inflate real MTTR.
-
Mitigate-Now vs. Keep-Diagnosing Decision Prompt
In the middle of a live incident, decide whether to apply an available mitigation immediately or keep diagnosing for root cause — so you stop the customer bleeding at the earliest safe moment instead of chasing 'why' while the clock runs, cutting time-to-restore.
-
Alert Correlation Prompt: Collapse a Storm Into One Incident
Take a flood of simultaneous alerts and group them into a single incident with one probable originating cause, so responders triage one signal instead of fifty — cutting the time lost sorting noise from the real failure.
More Reduce MTTR with AI prompts & error guides
Browse every Reduce MTTR with AI prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.