Ship Prometheus Recording Rules Safely: Naming as API, promtool Tests
Production first guide to Prometheus recording rules: naming as API, promtool unit tests, per group limits, rollout advice, and copy paste examples.
Recording rules precompute PromQL into new time series so expensive or frequently-run queries read those precomputed series instead of recomputing every query. You define them in a rule file, load them with rule_files, and Prometheus evaluates them on a schedule using the same engine that runs your alerts. Use them for dashboard panels that hammer the same heavy expression, and for simplifying alert conditions.
TL;DR:
- Recording rules can significantly speed up heavy queries by precomputing metrics, but careless implementation can cause series explosion and lag issues.
- Use the
limitsetting at the group level to prevent excessive series creation and monitor evaluation duration and missed iterations to catch scaling problems early.- Validate rules thoroughly with
promtooltests, including edge cases like counter resets and missing data, before deploying to production.- Follow consistent naming conventions and avoid renaming recorded metrics mid-life to prevent breaking dashboards and alerts.
- Roll out new rules gradually, starting with low-traffic groups and short intervals, while closely watching the rule engine’s metrics for potential issues.
Table of Contents
- What recording rules actually do to the evaluation pipeline
- How to write, load, and roll out a rule file
- Naming and aggregation patterns that keep rules maintainable
- Testing rules before they ever touch production
- Watching for the failure modes that only show up at scale
- Copy-paste examples you can adapt today
- What actually separates safe rule sets from risky ones
- How DevOps AI ToolKit helps you get recording rules right
- Where to go for the exact syntax and deeper reading
- Sources
- FAQ
What recording rules actually do to the evaluation pipeline
A recording rule takes a PromQL expression, runs it on a schedule, and stores the result under a new metric name, exactly as described in the Prometheus documentation on recording rules. That new series behaves like any other metric: it lives in TSDB, respects your retention settings, and can be queried, graphed, or alerted on.
The evaluation model matters more than most people realize once they start writing rules:
- Rules inside the same group run sequentially at the same evaluation timestamp, so a slow rule early in the group can delay everything after it.
- Prometheus uses a global evaluation interval by default, but each group can override that interval on its own.
- The
query_offsetsetting shifts the evaluation timestamp slightly into the past, which helps when your source samples arrive a little late and would otherwise create gaps.
Group boundaries are scheduling units, but they’re also failure domains. A single overloaded rule can drag down the rest of its group.
How to write, load, and roll out a rule file
Recording rules live in YAML files referenced by rule_files in your Prometheus configuration. The structure is consistent across every rule set you’ll write, and once you’ve done it a few times it stops feeling like ceremony and starts feeling like muscle memory.
- Create a YAML file with a top-level
groupslist; each group has aname, an optionalinterval, an optionallimit, an optionalquery_offset, and optionallabels. - Under each group, add one or more
rules, each with arecordfield (the new metric name) and anexprfield (the PromQL to evaluate). - Run
promtool check rulesagainst the file to catch syntax errors before they reach production. - Write behavior tests with
promtool test rules, which is the same tool the Prometheus unit testing documentation covers in detail. - Reference the file from
rule_filesin your main config, then reload withSIGHUPor a call to the/-/reloadendpoint. - Roll out new rules to one environment or one group at a time rather than shipping a large batch all at once.
Pro Tip: Start every new rule in a low-traffic group with a short interval override, watch it for a day, then merge it into your standard group once you trust the output.
Naming and aggregation patterns that keep rules maintainable
Prometheus’s own practices guide recommends a naming pattern of level:metric:operations, and it’s worth following even when it feels verbose. Per the Prometheus recording rules practices guide, the format tells anyone reading a dashboard what aggregation level they’re looking at (job, instance, or something broader), what the metric represents, and what operation produced it, such as rate or sum.
A few habits prevent the most common mistakes:
- Drop the
_totalsuffix once you applyrate()orirate()to a counter, since the resulting series is a rate, not a counter. - For ratios, aggregate the numerator and the denominator separately, then divide. Averaging a ratio across instances produces a number that looks plausible and means nothing.
- Use
withoutto drop labels you’ve aggregated away while keeping the ones your dashboards or alerts still need, likejobwhen you’re aggregating awayinstance.
Recording rules do not reduce cardinality on their own. Every distinct labelset your expression returns becomes a distinct series under the new name, per the Prometheus documentation on recording rules, so a careless sum by clause can multiply your series count instead of shrinking it.
Testing rules before they ever touch production
promtool gives you two separate checks, and skipping either one is how bad rules end up in production dashboards. promtool check rules validates syntax: it catches typos, malformed YAML, and invalid PromQL. promtool test rules goes further, running your expressions against synthetic input and comparing the output to what you expect, following the pattern in the Prometheus unit testing documentation.
A useful test suite covers more than the happy path:
- Feed in
input_seriesthat include a counter reset, and confirm yourrate()based rule handles it without spiking incorrectly. - Check that aggregated series preserve the labels your dashboards depend on, and drop the ones you intended to remove.
- Test an empty or sparse time range to make sure your rule doesn’t produce misleading output when data is missing.
- Compare your
exp_samplesagainst the actual recorded values at eacheval_timeto confirm the math, not just the shape.
Gate merges on passing rule tests in your CI pipeline the same way you’d gate application code on unit tests. A recording rule with wrong math is worse than no recording rule at all, because dashboards and alerts start trusting a number that’s quietly off.
Watching for the failure modes that only show up at scale
Recording rules are cheap right up until they aren’t, and the warning signs are specific enough that you can monitor for them directly. The per-group limit setting caps how many series a group is allowed to produce; if a rule exceeds it, Prometheus discards every series that rule produced for that evaluation and logs an evaluation error, as described in the Prometheus documentation on recording rules.
A short list of things worth putting on your own dashboard:
- Watch
rule_group_iterations_missed_total. A rising count means a group hasn’t finished before its next scheduled evaluation, and Prometheus is skipping iterations, which creates gaps in your recorded series. - Track evaluation duration per group so you can see which group is closest to blowing its interval before it starts missing iterations.
- Separate expensive, slow-running rules into their own group so they don’t hold up faster, unrelated rules.
- Use
query_offsetwhen your source data tends to arrive late, and expect occasional stale markers if you don’t.
Pro Tip: If a group’s evaluation duration regularly exceeds half its interval, split it before it starts missing iterations, not after.
Copy-paste examples you can adapt today
A rate aggregation rule that rolls request rate up to job level, dropping instance-specific noise, looks like this:
record: job:http_requests:rate5mwithexpr: sum without(instance) (rate(http_requests_total[5m])).- A ratio rule computes numerator and denominator as separate recorded series, then a third rule divides them, rather than averaging a per-instance ratio.
- A minimal
promtool test rulesfixture definesinput_serieswith a metric and a sequence of values, aneval_time, and anexp_samplesblock listing the metric name and expected value at that time.
That pattern, defining a rule file, adding it to rule_files, and reloading, mirrors the walkthrough in the PromLabs recording rules training, which recorded a metric it named path_method:demo_api_request_duration_seconds_count:rate5m. For more worked examples of turning a slow dashboard query into a fast recorded one, see our guide to converting slow PromQL into recording rules.
What actually separates safe rule sets from risky ones

Treat every recorded metric name as a small public API. Someone will build a dashboard or an alert on it, and renaming it later means finding every consumer first. Keep names minimal and predictable, following level:metric:operations rather than whatever felt convenient at the time.
Ship tests with every rule change and gate them in CI, the same way you’d gate a code change. Roll new rules out to one group before trusting them across the fleet, and watch the rule engine’s own metrics, like missed iterations and evaluation duration, instead of just watching for complaints.
— James
How DevOps AI ToolKit helps you get recording rules right
Writing a handful of rules is easy. Keeping fifty of them correct, tested, and cheap to run six months later is where teams lose time, usually while firefighting a stale dashboard at 2 AM. DevOps AI ToolKit’s Observability Review looks at your existing rule files, flags naming and cardinality problems, and hands back a report you can act on immediately.

A typical engagement covers:
- A review of your current recording and alerting rules against the naming and aggregation patterns in this guide.
- Recommended rule changes, including consolidating redundant groups and fixing ratio calculations that average instead of aggregate.
promtool test rulesfixtures for your critical rules, so future changes are gated automatically.- A rollout plan for shipping the changes without breaking dashboards mid-migration.
If you’d rather start smaller, check pricing for the Observability Review and other audit options, or book time directly through work with me.
Where to go for the exact syntax and deeper reading
For the canonical syntax and evaluation semantics, the Prometheus documentation on recording rules is the primary reference, alongside the naming and aggregation practices guide. For a hands-on walkthrough with a working example, the PromLabs recording rules training is worth working through end to end, and the Prometheus unit testing documentation covers every option promtool test rules supports.
Sources
FAQ
How do you write Prometheus alerting rules?
Alerting rules use the same groups and rules structure as recording rules, but with alert and for fields instead of record, plus labels and annotations for routing and messaging. Many teams build alert expressions on top of a recording rule so the alert condition stays simple and the heavy computation happens once, on a schedule.
Is Prometheus free to use commercially?
Yes, Prometheus is open source and free to run in any commercial setting, including production infrastructure. Support, audits, and managed operations around it are separate services that vendors, including DevOps AI ToolKit, offer on top of the free core project.
What are some best practices for using Prometheus effectively?
Name recorded and derived metrics with the level:metric:operations pattern, aggregate ratio numerators and denominators separately, and keep expensive rules in their own evaluation group. Test rule changes with promtool test rules in CI, and watch rule_group_iterations_missed_total to catch groups that are falling behind schedule.
Can you give an example of a Prometheus metric?
A raw counter metric like http_requests_total tracks a running count of requests with labels such as method, status, and instance. A recording rule might turn that into job:http_requests:rate5m, a precomputed five-minute rate aggregated to the job level, following the naming pattern in the Prometheus recording rules practices guide.
Recommended
- Prometheus Recording Rules That Make Slow Queries Fast
- Metric Naming Standards That Keep Prometheus Sane
Get 500 Battle-Tested DevOps AI Prompts — Free
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.