Skip to content
DevOps AI ToolKit
All guides
AI for Automation By James Joyner IV · · 10 min read

Ship Prometheus Recording Rules Safely: Naming as API, promtool Tests

Production first guide to Prometheus recording rules: naming as API, promtool unit tests, per group limits, rollout advice, and copy paste examples.

Ship Prometheus Recording Rules Safely: Naming as API, promtool Tests

Recording rules precompute PromQL into new time series so expensive or frequently-run queries read those precomputed series instead of recomputing every query. You define them in a rule file, load them with rule_files, and Prometheus evaluates them on a schedule using the same engine that runs your alerts. Use them for dashboard panels that hammer the same heavy expression, and for simplifying alert conditions.


TL;DR:

  • Recording rules can significantly speed up heavy queries by precomputing metrics, but careless implementation can cause series explosion and lag issues.
  • Use the limit setting at the group level to prevent excessive series creation and monitor evaluation duration and missed iterations to catch scaling problems early.
  • Validate rules thoroughly with promtool tests, including edge cases like counter resets and missing data, before deploying to production.
  • Follow consistent naming conventions and avoid renaming recorded metrics mid-life to prevent breaking dashboards and alerts.
  • Roll out new rules gradually, starting with low-traffic groups and short intervals, while closely watching the rule engine’s metrics for potential issues.

Table of Contents

What recording rules actually do to the evaluation pipeline

A recording rule takes a PromQL expression, runs it on a schedule, and stores the result under a new metric name, exactly as described in the Prometheus documentation on recording rules. That new series behaves like any other metric: it lives in TSDB, respects your retention settings, and can be queried, graphed, or alerted on.

The evaluation model matters more than most people realize once they start writing rules:

  • Rules inside the same group run sequentially at the same evaluation timestamp, so a slow rule early in the group can delay everything after it.
  • Prometheus uses a global evaluation interval by default, but each group can override that interval on its own.
  • The query_offset setting shifts the evaluation timestamp slightly into the past, which helps when your source samples arrive a little late and would otherwise create gaps.

Group boundaries are scheduling units, but they’re also failure domains. A single overloaded rule can drag down the rest of its group.

How to write, load, and roll out a rule file

Recording rules live in YAML files referenced by rule_files in your Prometheus configuration. The structure is consistent across every rule set you’ll write, and once you’ve done it a few times it stops feeling like ceremony and starts feeling like muscle memory.

  1. Create a YAML file with a top-level groups list; each group has a name, an optional interval, an optional limit, an optional query_offset, and optional labels.
  2. Under each group, add one or more rules, each with a record field (the new metric name) and an expr field (the PromQL to evaluate).
  3. Run promtool check rules against the file to catch syntax errors before they reach production.
  4. Write behavior tests with promtool test rules, which is the same tool the Prometheus unit testing documentation covers in detail.
  5. Reference the file from rule_files in your main config, then reload with SIGHUP or a call to the /-/reload endpoint.
  6. Roll out new rules to one environment or one group at a time rather than shipping a large batch all at once.

Pro Tip: Start every new rule in a low-traffic group with a short interval override, watch it for a day, then merge it into your standard group once you trust the output.

Naming and aggregation patterns that keep rules maintainable

Prometheus’s own practices guide recommends a naming pattern of level:metric:operations, and it’s worth following even when it feels verbose. Per the Prometheus recording rules practices guide, the format tells anyone reading a dashboard what aggregation level they’re looking at (job, instance, or something broader), what the metric represents, and what operation produced it, such as rate or sum.

A few habits prevent the most common mistakes:

  • Drop the _total suffix once you apply rate() or irate() to a counter, since the resulting series is a rate, not a counter.
  • For ratios, aggregate the numerator and the denominator separately, then divide. Averaging a ratio across instances produces a number that looks plausible and means nothing.
  • Use without to drop labels you’ve aggregated away while keeping the ones your dashboards or alerts still need, like job when you’re aggregating away instance.

Recording rules do not reduce cardinality on their own. Every distinct labelset your expression returns becomes a distinct series under the new name, per the Prometheus documentation on recording rules, so a careless sum by clause can multiply your series count instead of shrinking it.

Testing rules before they ever touch production

promtool gives you two separate checks, and skipping either one is how bad rules end up in production dashboards. promtool check rules validates syntax: it catches typos, malformed YAML, and invalid PromQL. promtool test rules goes further, running your expressions against synthetic input and comparing the output to what you expect, following the pattern in the Prometheus unit testing documentation.

A useful test suite covers more than the happy path:

  1. Feed in input_series that include a counter reset, and confirm your rate() based rule handles it without spiking incorrectly.
  2. Check that aggregated series preserve the labels your dashboards depend on, and drop the ones you intended to remove.
  3. Test an empty or sparse time range to make sure your rule doesn’t produce misleading output when data is missing.
  4. Compare your exp_samples against the actual recorded values at each eval_time to confirm the math, not just the shape.

Gate merges on passing rule tests in your CI pipeline the same way you’d gate application code on unit tests. A recording rule with wrong math is worse than no recording rule at all, because dashboards and alerts start trusting a number that’s quietly off.

Watching for the failure modes that only show up at scale

Recording rules are cheap right up until they aren’t, and the warning signs are specific enough that you can monitor for them directly. The per-group limit setting caps how many series a group is allowed to produce; if a rule exceeds it, Prometheus discards every series that rule produced for that evaluation and logs an evaluation error, as described in the Prometheus documentation on recording rules.

A short list of things worth putting on your own dashboard:

  • Watch rule_group_iterations_missed_total. A rising count means a group hasn’t finished before its next scheduled evaluation, and Prometheus is skipping iterations, which creates gaps in your recorded series.
  • Track evaluation duration per group so you can see which group is closest to blowing its interval before it starts missing iterations.
  • Separate expensive, slow-running rules into their own group so they don’t hold up faster, unrelated rules.
  • Use query_offset when your source data tends to arrive late, and expect occasional stale markers if you don’t.

Pro Tip: If a group’s evaluation duration regularly exceeds half its interval, split it before it starts missing iterations, not after.

Copy-paste examples you can adapt today

A rate aggregation rule that rolls request rate up to job level, dropping instance-specific noise, looks like this:

  • record: job:http_requests:rate5m with expr: sum without(instance) (rate(http_requests_total[5m])).
  • A ratio rule computes numerator and denominator as separate recorded series, then a third rule divides them, rather than averaging a per-instance ratio.
  • A minimal promtool test rules fixture defines input_series with a metric and a sequence of values, an eval_time, and an exp_samples block listing the metric name and expected value at that time.

That pattern, defining a rule file, adding it to rule_files, and reloading, mirrors the walkthrough in the PromLabs recording rules training, which recorded a metric it named path_method:demo_api_request_duration_seconds_count:rate5m. For more worked examples of turning a slow dashboard query into a fast recorded one, see our guide to converting slow PromQL into recording rules.

What actually separates safe rule sets from risky ones

What actually separates safe rule sets from risky ones — overview diagram

Treat every recorded metric name as a small public API. Someone will build a dashboard or an alert on it, and renaming it later means finding every consumer first. Keep names minimal and predictable, following level:metric:operations rather than whatever felt convenient at the time.

Ship tests with every rule change and gate them in CI, the same way you’d gate a code change. Roll new rules out to one group before trusting them across the fleet, and watch the rule engine’s own metrics, like missed iterations and evaluation duration, instead of just watching for complaints.

— James

How DevOps AI ToolKit helps you get recording rules right

Writing a handful of rules is easy. Keeping fifty of them correct, tested, and cheap to run six months later is where teams lose time, usually while firefighting a stale dashboard at 2 AM. DevOps AI ToolKit’s Observability Review looks at your existing rule files, flags naming and cardinality problems, and hands back a report you can act on immediately.

Devopsaitoolkit

A typical engagement covers:

  • A review of your current recording and alerting rules against the naming and aggregation patterns in this guide.
  • Recommended rule changes, including consolidating redundant groups and fixing ratio calculations that average instead of aggregate.
  • promtool test rules fixtures for your critical rules, so future changes are gated automatically.
  • A rollout plan for shipping the changes without breaking dashboards mid-migration.

If you’d rather start smaller, check pricing for the Observability Review and other audit options, or book time directly through work with me.

Where to go for the exact syntax and deeper reading

For the canonical syntax and evaluation semantics, the Prometheus documentation on recording rules is the primary reference, alongside the naming and aggregation practices guide. For a hands-on walkthrough with a working example, the PromLabs recording rules training is worth working through end to end, and the Prometheus unit testing documentation covers every option promtool test rules supports.

Sources

FAQ

How do you write Prometheus alerting rules?

Alerting rules use the same groups and rules structure as recording rules, but with alert and for fields instead of record, plus labels and annotations for routing and messaging. Many teams build alert expressions on top of a recording rule so the alert condition stays simple and the heavy computation happens once, on a schedule.

Is Prometheus free to use commercially?

Yes, Prometheus is open source and free to run in any commercial setting, including production infrastructure. Support, audits, and managed operations around it are separate services that vendors, including DevOps AI ToolKit, offer on top of the free core project.

What are some best practices for using Prometheus effectively?

Name recorded and derived metrics with the level:metric:operations pattern, aggregate ratio numerators and denominators separately, and keep expensive rules in their own evaluation group. Test rule changes with promtool test rules in CI, and watch rule_group_iterations_missed_total to catch groups that are falling behind schedule.

Can you give an example of a Prometheus metric?

A raw counter metric like http_requests_total tracks a running count of requests with labels such as method, status, and instance. A recording rule might turn that into job:http_requests:rate5m, a precomputed five-minute rate aggregated to the job level, following the naming pattern in the Prometheus recording rules practices guide.

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week
Free download · 368-page PDF

Get 500 Battle-Tested DevOps AI Prompts — Free

500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.

  • 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
  • Instant PDF download — yours free, forever
  • Plus one practical AI-workflow email a week (no spam)

Single opt-in · unsubscribe anytime · no spam.