VictoriaMetrics Cardinality Explorer & TSDB Triage Prompt
Diagnose a VictoriaMetrics cluster suffering from high active time series and churn using the built-in Cardinality Explorer and TSDB status endpoints, then produce a prioritized remediation plan.
- Target user
- SRE and platform engineers running single-node or cluster VictoriaMetrics
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a senior observability engineer who specializes in VictoriaMetrics capacity and cardinality management. I will provide: - Output from /api/v1/status/tsdb and the Cardinality Explorer UI (top metrics by series, top label=value pairs, label-value count) - vm_cache_size_bytes, vm_slow_queries_total, and active series trends - Our ingestion rate (vm_rows_inserted_total) and retention settings Your job: 1. **Baseline** — establish current active time series, churn rate, and how close we are to the RAM-bound series limit. 2. **Offender ranking** — identify the metrics and label keys driving cardinality, distinguishing legitimate growth from unbounded labels (request_id, pod hash, full URLs). 3. **Root cause** — classify each offender as churn (frequent restarts), explosion (high-cardinality label), or duplication (overlapping scrape jobs). 4. **Remediation** — propose relabel_configs, stream aggregation (-streamAggr), or -dropSamplesOnOverload only where appropriate, with the exact metric_relabel_configs snippets. 5. **Guardrails** — recommend -maxLabelsPerTimeseries and -search.maxUniqueTimeseries limits sized to our hardware. 6. **Verification** — define the queries to confirm series reduction without losing needed signals. 7. **Rollback** — describe how to revert each change safely. Output as: (a) ranked offender table, (b) per-offender fix, (c) guardrail config, (d) verification checklist. Flag any change that would silently drop metrics currently used by alerting rules before recommending it.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Prometheus Active Series Cardinality Reduction Triage Prompt
Triage a TSDB active-series and head-memory blowup by finding the offending metric+label, deciding between drop relabeling, label aggregation, or instrumentation fixes, with a measurable before/after series count.
-
Prometheus metric_relabel_configs Drop-List Cardinality Audit Prompt
Audit and generate metric_relabel_configs drop and keep rules that cut high-cardinality series at ingest without dropping metrics your alerts and dashboards depend on.
-
VictoriaMetrics vmagent Stream Aggregation Rules Design Prompt
Design vmagent stream aggregation rules that pre-aggregate high-cardinality metrics at ingest, cutting stored series while preserving the dimensions your queries need.
-
Prometheus node_exporter Collector Selection & Tuning Prompt
Enable, disable, and filter node_exporter collectors so hosts expose the metrics you actually alert on without paying cardinality and CPU cost for the ones you don't.
More Prometheus & Monitoring prompts & error guides
Browse every Prometheus & Monitoring prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.