Skip to content
DevOps AI ToolKit
Home

Prometheus & Monitoring: Alerts, PromQL and Storage

Write better alert rules, PromQL queries, and Grafana dashboards with AI.

165 copy-paste prompts · 129 in-depth guides Jump to prompts Jump to guides

What actually breaks, and what to check first

Prometheus problems divide cleanly into three: the data is not arriving, the query is wrong, or the alert is right but never reaches anyone. They present very differently and the mistake is treating a routing failure as a query problem — you can spend an afternoon tuning a PromQL expression that was correct all along, while Alertmanager silently drops the notification.

The most consequential failure mode is cardinality. Every distinct label combination is a separate time series, so a single label carrying a user ID, a request path or a container ID can multiply series count by thousands and take the server from healthy to out-of-memory within hours. It rarely announces itself; ingestion just gets slower and memory climbs.

The second is silence. An alert that never fires looks exactly like a healthy system. `for` durations that exceed the condition’s lifetime, expressions returning no series at all rather than zero, and inhibition rules that suppress more than intended all produce the same reassuring nothing.

Triage order

Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.

  1. Is the target actually being scraped?

    Open /targets, or query: up{job="<job>"}

    `up == 0` means the scrape failed and the error is shown next to the target. A target missing entirely means service discovery never produced it — a relabeling problem, not a network one.

  2. Did relabeling drop it?

    Open /service-discovery and compare discovered labels against target labels.

    This page shows targets before and after relabeling. A target present in discovery and absent from /targets was dropped by a relabel rule — the most common reason a service "is not being scraped" despite being reachable.

  3. Does the series exist at all?

    count({__name__=~"<metric>.*"}) — and check the metric name in /metrics on the exporter itself.

    An empty query result and a zero value are different. An expression over a non-existent series returns nothing, and an alert on nothing never fires — this is the quiet failure mode behind alerts that "did not work".

  4. Is cardinality under control?

    topk(10, count by (__name__)({__name__=~".+"}))

    Names the metrics with the most series. A metric with far more series than the number of instances has a high-cardinality label — usually an ID, path or hash that should never have been a label.

  5. Is the alert firing but not delivered?

    Check /alerts in Prometheus, then Alertmanager’s /#/alerts and its logs.

    Firing in Prometheus but absent in Alertmanager means routing or inhibition. Present in Alertmanager but not delivered means the receiver is failing — the log names the reason.

    → Alertmanager notification failures

Diagnostic commands

Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.

promtool check rules /etc/prometheus/rules/*.yml
Read-only

Validate recording and alerting rules before reload.

How to read it: Catches syntax and duplicate-name errors. It does not tell you whether an expression returns data, which is the more common defect — check that separately.

promtool query instant http://localhost:9090 'ALERTS{alertstate="firing"}'
Read-only

List what is firing right now, from the command line.

How to read it: Confirms whether Prometheus believes the alert is active. If it does and nobody was paged, the problem is downstream in Alertmanager.

curl -s localhost:9090/api/v1/status/tsdb | jq .data.seriesCountByMetricName
Read-only

Show series counts per metric directly from the TSDB.

How to read it: The authoritative cardinality view. Compare a metric’s series count against the number of instances exporting it; a large ratio identifies the offending label.

curl -X POST localhost:9090/-/reload
Changes state

Reload configuration without restarting.

How to read it: Requires `--web.enable-lifecycle`. Preferred over a restart because a restart replays the WAL, which on a large TSDB can take minutes during which you are blind.

amtool alert query --alertmanager.url=http://localhost:9093
Read-only

Query Alertmanager for the alerts it currently holds.

How to read it: Distinguishes "Prometheus never sent it" from "Alertmanager received and suppressed it" — two very different bugs with identical symptoms.

amtool config routes test --config.file=alertmanager.yml severity=critical team=infra
Read-only

Show which receiver a given label set resolves to.

How to read it: Tests routing without waiting for a real alert. The correct way to verify a routing tree change before it matters.

Failure modes

These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.

A target disappears from /targets with no error.

Cause:
A relabel rule dropped it, or service discovery stopped returning it.
Fix:
Compare /service-discovery against /targets. Present in the first and absent from the second means relabeling; absent from both means the discovery mechanism.

Prometheus will not start after an unclean shutdown.

Cause:
WAL corruption — the write-ahead log was truncated mid-write.
Fix:
The logs name the corrupt segment. Prometheus can usually recover by discarding it, at the cost of the most recent samples; deleting the whole data directory is data loss and is rarely necessary.

Full walkthrough →

Memory climbs steadily until the process is OOM-killed.

Cause:
Cardinality growth — a label taking unbounded values.
Fix:
Find it with series counts per metric, then drop the label in `metric_relabel_configs` at scrape time. Dropping at query time does not help; the cost is in ingestion and storage.

An alert condition is clearly true but the alert never fires.

Cause:
The expression returns no series, or `for` is longer than the condition persists.
Fix:
Run the bare expression in the query browser. Empty output means it can never fire regardless of threshold — use `absent()` to alert on a metric that has stopped existing.

Alerts fire in Prometheus but nobody is notified.

Cause:
Alertmanager routing sends them to a receiver that is failing, or an inhibition rule suppresses them.
Fix:
Test the label set against the routing tree with `amtool`, then read Alertmanager’s logs for delivery errors. Both are quick, and they separate the two causes definitively.

Grafana shows gaps that Prometheus does not.

Cause:
The query’s range or step does not align with the scrape interval, so windows land between samples.
Fix:
Ensure range selectors span at least twice the scrape interval. `rate()` over a window shorter than two scrapes returns nothing, which renders as a gap.

Common mistakes

  • Putting unbounded values — request paths, user IDs, container IDs — into labels. This is the single most common way to destroy a Prometheus server.
  • Alerting on an expression that returns no series and assuming silence means health. Use `absent()` when a metric disappearing is itself the problem.
  • Restarting Prometheus to apply configuration. Reload instead; a restart replays the WAL and leaves you blind while it does.
  • Using `rate()` over a window shorter than two scrape intervals, which silently returns nothing.
  • Setting `for` durations longer than the incident they are meant to catch, so the alert resolves before it ever fires.
  • Treating recording rules as optional. Expensive dashboard queries executed repeatedly are a common cause of a slow Prometheus.

Frequently asked questions

Why is my Prometheus alert not firing when the condition is clearly true?

Run the alert expression on its own in the query browser. If it returns no series — as opposed to returning zero — then no threshold can ever be crossed, because there is nothing to compare. This is the usual cause, and it happens whenever a metric stops being exported entirely. Wrap it in `absent()` when the disappearance is the thing you care about. The second cause is a `for` duration longer than the condition actually persists.

What causes Prometheus memory to grow until it is killed?

Cardinality. Every unique label combination is a separate time series held in memory, so one label carrying an unbounded value — a user ID, a URL path, a container ID — can create hundreds of thousands of series from a handful of instances. Query series counts per metric to find it, then drop the label in `metric_relabel_configs` so it is discarded at scrape time rather than stored.

A target is reachable but does not appear in /targets. Why?

A relabeling rule dropped it. The /service-discovery page shows targets both before and after relabeling, so if it appears there but not in /targets, a `relabel_configs` rule removed it. If it is missing from both, the problem is upstream in service discovery rather than in Prometheus.

How do I recover from WAL corruption without losing all my data?

Read the startup log — it names the corrupt segment. Prometheus can generally start by discarding that segment, which costs you the most recent samples but preserves the compacted blocks holding your history. Deleting the entire data directory is a much larger loss and is almost never required.

Should I alert from Prometheus or from Grafana?

From Prometheus, in almost all cases. Rules there live in version control, evaluate next to the data, and continue working if Grafana is down — which during an infrastructure incident is a real possibility. Grafana alerting is reasonable for dashboards owned by people who do not deploy rule files, but it should not carry your paging path.

Prompts

Guides

Recommended tools

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week