Triage Prometheus High Cardinality with Three PromQL Queries and Fixes
Use three PromQL queries to find the offending metric, then apply relabels, recording rules, and Prometheus config fixes to stop cardinality spikes before...
Prometheus high cardinality means too many unique time series (metric name plus full label set) are hitting your TSDB at once, and it’s rarely a subtle problem. Memory on the head block spikes, queries slow down or time out, and if nobody catches it, you get an OOM kill in the middle of an incident you’re already trying to debug. Fixing it starts with finding the offending metric, then bounding the labels that created it.
TL;DR:
- High-cardinality metrics can cause memory spikes, query failures, and process crashes, especially if labels like user_id or request_id grow unbounded.
- Detecting spikes involves running quick PromQL queries to count active series, identify top metrics, and monitor scrape health during traffic increases or deployments.
- Fixes should target instrumented code by bounding label values, normalizing dynamic paths, and avoiding unbounded labels, with relabel configs used as an immediate measure.
- Operational safeguards include setting retention size, tuning remote_write buffers, and limiting scrape size with sample and label caps to prevent catastrophic failures.
- Long-term success requires ongoing label review, recording rules, and possibly sharding or federation for high-volume environments, with professional audits recommended for complex issues.
Table of Contents
- What Cardinality Actually Means in Prometheus
- Why High Cardinality Actually Hurts You
- How to Catch a Cardinality Spike Before It Catches You
- Fixing It: Instrumentation, Relabeling, and Recording Rules
- Tuning Prometheus and Remote_write to Contain the Damage
- Histograms Won’t Save You From Bad Labels
- How Devopsaitoolkit Approaches Cardinality Incidents
- What to Actually Do Next, and When to Call for Help
- Get an Observability Review Instead of Guessing at Config Changes
- Sources
- FAQ
What Cardinality Actually Means in Prometheus
Every time series in Prometheus is identified by a metric name plus its complete set of label key/value pairs. http_requests_total{method="GET", handler="/api/users", status="200"} and http_requests_total{method="POST", handler="/api/users", status="200"} are two entirely separate series, not variations of one. Change a single label value, add a new label, or remove one, and you’ve created a new series identity that the TSDB has to track independently.
The math gets ugly fast. A metric with method (4 values), handler (50 routes), status (10 codes), and pod (an autoscaling fleet of 30) doesn’t produce 94 series. It produces the product of those dimensions, potentially tens of thousands. Prometheus’s own naming and label guidance exists precisely because this multiplication is invisible until it isn’t.
Why High Cardinality Actually Hurts You
Cardinality itself isn’t the enemy. As Grafana’s own guidance on managing high-cardinality metrics points out, plenty of high-dimension labels genuinely support dashboards, alerts, and recording rules. The problem is unbounded cardinality, dimensions like user_id or request_id that grow without limit and never stop adding new series.
Here’s what that actually costs you in production:
- Memory pressure on the head block. Prometheus keeps the current head block and its label-value index in RAM, so every new series adds permanent overhead until it’s compacted or expired, per the Prometheus storage docs.
- OOM risk. When the head block outgrows available memory, the Prometheus process gets killed, usually during the exact traffic spike you needed visibility into.
- WAL and disk pressure. High-traffic servers retain more write-ahead log segments, and retention.size needs headroom for WAL and head-chunk peaks, not just steady-state series counts.
- Query timeouts. A selector that should match a handful of series suddenly matches thousands, and PromQL evaluation grinds to a crawl or fails outright.
- Remote_write amplification. Series churn increases the memory needed for the series-to-label mapping cache and queue shards, so a cardinality spike on the local server becomes a queue backlog everywhere downstream.
If you’ve ever watched a dashboard go blank right when you needed it most, this is usually why.
How to Catch a Cardinality Spike Before It Catches You
You don’t need a fancy tool to find the offending metric. You need three PromQL queries and about five minutes.
- Count total active series. Run
count({__name__=~".+"})to get a baseline number for your instance. - Rank metrics by series count.
topk(20, count by (__name__)({__name__=~".+"}))tells you exactly which metric names are eating your budget. This is almost always where you’ll find the culprit. - Check scrape health metrics. Look at
scrape_series_added,scrape_samples_scraped, andscrape_samples_post_metric_relabelingper target. A sudden jump inscrape_series_addedon one job is your smoking gun. - Correlate with
upby job and instance. If a target is also failing scrapes, you may be watchingsample_limitreject the entire scrape rather than just the noisy metric. - Cross-reference with recent deploys. Nine times out of ten, a cardinality spike lines up with a new instrumentation change or a feature flag rollout, not gradual organic growth.
Once you’ve isolated the metric, check which label is driving the explosion by comparing cardinality across label subsets. If it’s clearly a bad deploy, alert your team, and don’t be afraid to coordinate a rollback while you draft the permanent fix.
Pro Tip: Keep a saved PromQL query for topk(20, count by (__name__)({__name__=~".+"})) pinned in your runbook. During an active incident, hunting for the right query syntax wastes minutes you don’t have.

Fixing It: Instrumentation, Relabeling, and Recording Rules
Fix the source before you reach for config band-aids. Bad instrumentation is a code problem, and no amount of relabeling fully compensates for a metric that was designed wrong from the start.
Start with instrumentation discipline:
- Bound every label’s possible value set before you ship it. If you can’t enumerate the values, it doesn’t belong as a label.
- Normalize dynamic paths.
/api/users/12345should become/api/users/{id}, not a distinct label value per user. - Never use
user_id,email,request_id,container_id, orcommit_shaas label values. Prometheus’s own naming practices call these out explicitly as anti-patterns, and for good reason: they’re the single most common cause of runaway cardinality.
When instrumentation is out of your control, immediate or third-party, metric_relabel_configs lets you drop or rewrite offending labels at scrape time. Test relabel rules against a staging target first; a misconfigured relabel rule can silently drop the metric entirely instead of just the label you intended to trim.
Recording rules are your next lever. Pre-aggregating a high-cardinality metric into a coarser recording rule (dropping pod in favor of deployment, for example) lets you keep the useful signal while discarding the series explosion underneath it.
And sometimes the right answer is to stop trying to put the data in Prometheus at all. Request-level identity, like which specific user hit an error, belongs in logs or traces, not metrics. Prometheus is built for aggregate numeric signals over time, not for answering “which request failed.”

Finally, set sample_limit and label_limit as guardrails in your scrape configs. Both default to 0 (unlimited), so an unset limit means an instrumentation bug has no ceiling. Just know the failure mode: exceeding either limit fails the entire scrape, not just the offending series, so alert on scrape failures the moment you turn these on.
Pro Tip: Roll out a new metric_relabel_configs rule to one canary target first. If a regex is slightly off, you want to find out on one instance, not lose visibility fleet-wide.
Tuning Prometheus and Remote_write to Contain the Damage
Even with clean instrumentation, you need operational guardrails sized for reality, not for the good day.
- Retention sizing. Set
retention.sizewith headroom for WAL segments and head-chunk peaks, not just your average series count. A retention limit sized for steady state will get blown through the first time a deploy spikes cardinality. - Remote_write memory. Enabling remote_write commonly increases memory usage by roughly 25%, and series churn adds further load on the mapping cache and queue shards. Size
queue_config.capacity,max_samples_per_send, andmin_shards/max_shardsdeliberately, and watch shard saturation metrics rather than guessing. - Scrape guardrails. Configure
sample_limit,label_limit, andbody_size_limitper job, and wire alerts to scrape failure rates so a limit breach shows up immediately instead of silently. - When to go bigger. If a single Prometheus instance can’t absorb the cardinality your platform legitimately needs, that’s your signal to look at federation, sharding by team or namespace, or a remote storage backend rather than pushing retention and memory limits further.
Our remote_write tuning walkthrough covers shard math in more depth if you’re planning capacity for a growing fleet.
Histograms Won’t Save You From Bad Labels
Classic histograms make cardinality worse by design: every bucket boundary generates its own series, on top of separate _sum and _count series. A histogram with 10 buckets doesn’t cost 1 series, it costs 12.
Native histograms fix that specific problem by storing bucket data as a single composite sample instead of one series per bucket, which meaningfully cuts histogram-driven series counts. But don’t mistake that for a cardinality fix in general. Native histograms do nothing about a user_id label or an unbounded handler path; the label dimension problem is completely separate from the bucket-explosion problem. If you’re switching formats, test the change against a canary target first, since it changes what remote_write ships downstream.
How Devopsaitoolkit Approaches Cardinality Incidents
Devopsaitoolkit’s guides on investigating a cardinality spike with AI-assisted queries and taming metric cardinality before it tames you come out of real triage work, not theory. James, who covers Prometheus retention and TSDB internals for the site, treats every new label the way you’d treat a schema migration: bounded, reviewed, and tested before it ships.
What to Actually Do Next, and When to Call for Help
Short term: run the topk query, isolate the metric, kill or relabel the offending label, and confirm scrape health recovers. Medium term: add recording rules for anything you kept, and push sample_limit/label_limit into every scrape config that doesn’t have them yet. Long term: build label review into your instrumentation checklist so this doesn’t recur every quarter.
Escalate when spikes repeat despite fixes, when active series stay elevated for days rather than hours, or when remote_write queues stay saturated no matter how you tune shards. Those are signs of a structural problem, not a one-off bad deploy, and usually the fastest path forward is a proper capacity plan, relabel strategy, and recording-rule design done in one pass rather than patched incident by incident.
— James
Get an Observability Review Instead of Guessing at Config Changes
Devopsaitoolkit’s Observability Review is built for exactly the mess this article just walked through: a fixed-scope, $300 one-off engagement that audits your Prometheus setup, maps your actual cardinality offenders, and hands you a relabel plan and recording-rule design instead of a generic checklist.

If you’re mid-incident and need hands-on remediation rather than a review, hourly consulting starts at $150 per hour and covers everything from remote_write tuning to retention sizing under real production constraints. Teams that want the toolkit itself, prompt libraries, config validators, and troubleshooting guides for Prometheus and the rest of the stack, can check the Pro plan at $19 per month or $180 per year. Book a review or grab the toolkit, either way you leave with a plan instead of another 2 a.m. page.
FAQ
What Counts as High Cardinality in Prometheus?
There’s no fixed number, but a metric generating tens of thousands of active series, or one whose series count grows unbounded over time, counts as high cardinality. The real test is whether a label’s value set is bounded and enumerable, not how many values it currently has.
Which Labels Should I Never Use in Prometheus?
Avoid user_id, email, request_id, session_id, raw URL paths, and commit_sha as label values, since Prometheus’s naming guidance flags these as classic unbounded-cardinality anti-patterns. Normalize paths and move request-level identity to logs or traces instead.
Does Remote_write Make Cardinality Problems Worse?
Yes. Remote_write typically adds around 25% to memory usage, and series churn from a cardinality spike increases pressure on the mapping cache and queue shards downstream. Tuning queue capacity and shard counts limits how far the damage spreads.
Do Native Histograms Fix High Cardinality?
No. Native histograms reduce the series explosion caused by classic histogram buckets, but they don’t touch unbounded label dimensions like user IDs. Cardinality from bad labels and cardinality from histogram bucket design are separate problems.
How Much Does a Prometheus Observability Review Cost?
Devopsaitoolkit’s Observability Review is a $300 one-off engagement covering cardinality audits, relabel plans, and recording-rule design. Hourly consulting for deeper remediation work starts at $150 per hour.
Recommended
- Investigating a Prometheus Cardinality Spike With AI as Your
- Taming Prometheus Metric Cardinality Before It Tames You
- Prometheus Scrape Config and Relabeling Deep Dive
- Prometheus Recording Rules That Make Slow Queries Fast
Get 500 Battle-Tested DevOps AI Prompts — Free
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.