Alert-Storm Correlation and Triage Prompt
Cut through a flood of simultaneous alerts during an incident to find the originating signal, group symptoms from causes, and tell on-call which single alert actually matters.
- Target user
- On-call engineers drowning in alert storms during cascading failures
- Difficulty
- Beginner
- Tools
- Claude, ChatGPT
The prompt
You are a seasoned SRE who stays calm during alert storms. When fifty alerts fire at once, you know most are downstream symptoms of one upstream cause, and your job is to find that cause fast. I will paste a burst of alerts (names, services, severities, timestamps, labels) plus, if available, our service-dependency map. Your job: 1. **Order by time** — sort the alerts by first-fired timestamp; the earliest firings are likelier to be near the cause than the cascade of symptoms that followed. 2. **Cluster by relationship** — group alerts that share a service, dependency, host, or label, and use the dependency map to separate upstream causes from downstream effects. 3. **Identify the probable origin** — name the one or two alerts most likely to be the originating signal, and explain the chain by which they would produce the rest of the storm. 4. **Separate signal from noise** — flag alerts that are pure symptoms (will clear on their own once the cause is fixed) so on-call ignores them for now. 5. **Customer-impact read** — state which alerts indicate actual user-facing harm versus internal-only noise, to set urgency. 6. **Next action** — recommend the single highest-value thing to investigate first, with the specific dashboard or query to confirm the hypothesis. Mark your confidence. 7. **Watch-list** — list the alerts whose clearing will confirm recovery, so on-call knows what "fixed" looks like. Output as: (a) the probable root signal with the cascade explanation, (b) clustered groups labeled cause / symptom, (c) the one recommended first action with a confirmation query, (d) the recovery watch-list. When evidence is thin, say so and give the safest investigation path rather than a confident guess.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Is-This-Real Page Triage Prompt
Help a freshly paged on-call engineer decide in the first two minutes whether an alert is a real incident worth waking people for, a transient blip, or pure noise — before they over- or under-react.
-
Alert Triage Decision-Tree Builder Prompt
Turn a noisy alert stream into a deterministic, branching triage decision tree that any on-call engineer can follow to classify, route, and act on alerts in under a minute.
-
DNS Resolution Failure Live Diagnosis Prompt
Walk on-call through diagnosing a live DNS-related outage — resolver, authoritative, caching, and propagation layers — to find where name resolution is actually breaking before you start changing records.
-
Incident Alert-to-Owning-Team Router Prompt
Take a freshly fired alert and route it to the team that actually owns the failing component, so the right responder is paged first instead of bouncing through three on-call rotations.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.