Stop 14 Day Cache Evictions: GitLab Cache vs Artifacts for CI Pipelines
Decide when to cache or publish artifacts in GitLab CI. Production tested YAML examples, rollout notes, and how to avoid 14 day cache eviction.
Use cache to speed up rebuildable dependencies, and publish artifacts when a later job needs the exact output from an earlier one. Cache lives on runner storage or a distributed backend and can vanish without warning. Artifacts are GitLab-managed, expire on a schedule you set, and are the thing your next job can actually depend on. The YAML examples below show both in action.
TL;DR:
- Use cache to avoid re-downloading dependencies or recompiling unmodified code, but rely on artifacts for the exact output needed in later stages.
- Artifacts are authoritative, expire on a set schedule, and should contain actual build results, reports, or signed bundles, unlike cache which is disposable.
- Key cache invalidation with lockfile-based keys to prevent stale data, and prefer distributed storage for autoscaled runners to maintain cache effectiveness.
- Keep cache and artifact paths separate to prevent overwriting issues and control downstream artifact downloads with dependencies or needs.
- Regularly review cache expiration, transfer times, and key strategies to balance speed gains against potential build reliability issues.
Table of Contents
- What cache and artifacts actually are
- Key differences at a glance
- When to use cache and when to publish artifacts
- Practical .gitlab-ci.yml examples you can copy and adapt
- Cache keys and invalidation strategies
- Storage backends, availability, and transfer costs
- Common pitfalls and best practices
- Field notes from rolling this out in production
- Balancing speed and determinism in your pipeline
- How DevOps AI ToolKit can help you tune this
- Sources
- FAQ
What cache and artifacts actually are
Cache and artifacts solve two different problems, and mixing them up is where most pipeline slowdowns start. Cache is temporary storage for data you can regenerate: dependency folders, compiled object files, anything expensive to rebuild but disposable if lost. Artifacts are job outputs that another job, or a human, needs to see or use.
A typical cache holds .npm, .cache/pip, or a ccache directory full of precompiled object files. None of that is precious. If the cache disappears, npm install or pip install just runs again from scratch. Artifacts are different: a compiled binary your deploy job needs, a JUnit test report GitLab renders in the merge request, or a signed release bundle someone will actually ship.
Both are defined with paths relative to the project’s working directory, and both show up in different places:
- Cache is stored on runner-local disk or a configured distributed backend, and it’s largely invisible outside the job log.
- Artifacts are uploaded to GitLab itself and appear under the project’s Build artifacts view, per the job artifacts documentation.
- Cache is a private optimization; artifacts are a first-class pipeline output other jobs and reviewers can see.
Key differences at a glance
The fastest way to choose is to ask what happens if the data goes missing. If the answer is “the job reruns a bit slower,” you want cache. If the answer is “the next job breaks,” you want artifacts.
- Lifecycle: artifacts expire according to
expire_in, while cache retention is governed by runner storage policy and eviction rules, not by your pipeline configuration. - Scope: artifacts pass automatically to later-stage jobs in the same pipeline by default, according to GitLab’s job artifacts docs; cache is meant for reuse across jobs and pipelines depending on how your keys are set up.
- Reliability: artifacts are authoritative outputs; cache is best-effort and can be missing or evicted at any time, according to GitLab’s caching documentation.
- Visibility: artifacts live in the GitLab UI for anyone with project access; cache visibility depends entirely on your runner setup and isn’t exposed the same way.
By default, artifacts expire 30 days after upload unless you set a different expire_in value, according to GitLab’s caching documentation. That single setting is worth checking on every pipeline, since the default quietly deletes build outputs you might still need for an audit or a rollback.
When to use cache and when to publish artifacts
The decision rule is short: if a downstream job needs the exact bytes an earlier job produced, that’s an artifact. If you’re just trying to avoid redoing expensive but reproducible work, that’s cache.
- Dependency installs (npm, pip, Composer, Maven): cache the dependency directory so repeat installs skip network downloads.
- Compiler caches (ccache, sccache): cache object files so a clean build doesn’t recompile unchanged code.
- Build outputs (binaries, Docker build context artifacts, installers): publish as artifacts so the deploy or package job gets the real thing.
- Test and coverage reports (JUnit XML, coverage summaries): publish as artifacts so GitLab can render them in the merge request.
- Signed bundles: publish both the file and its signature as artifacts, since a later verification step needs the exact pairing, as shown in GitLab’s signing examples.
A few situations deserve extra care. Large build outputs slow down artifact upload and download, so weigh whether every job downstream actually needs the full payload. Cross-project sharing usually calls for artifacts, since cache isn’t designed to move data between projects. And anything you need to keep for compliance or long-term traceability belongs in artifacts with a deliberate expire_in, never in cache.
Pro Tip: If you’re unsure which one to reach for, ask whether losing the data would fail a build or just slow one down. That answer decides it every time.
Practical .gitlab-ci.yml examples you can copy and adapt
These snippets cover the cases that come up in nearly every pipeline. Adjust paths to match your project structure before using them.
- npm dependency cache keyed on the lockfile:
cache:
key:
files:
- package-lock.json
paths:
- .npm/
policy: pull-push
install:
script:
- npm ci --cache .npm --prefer-offline
Keying on package-lock.json means the cache only gets reused when dependencies haven’t changed, which avoids the classic stale-cache bug.
- Python/pip cache with a project-local directory:
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
cache:
key:
files:
- requirements.txt
paths:
- .cache/pip
Setting PIP_CACHE_DIR inside the project directory keeps the cache path predictable across runners.
- Build job publishing artifacts with reports and expiration:
build:
script:
- make build
- make test
artifacts:
paths:
- dist/
reports:
junit: report.xml
coverage_report:
coverage_format: cobertura
path: coverage.xml
expire_in: 7 days
The reports keys tell GitLab to parse and surface those files in the merge request, not just store them.
- Avoiding path collisions between cache and artifacts:
build:
cache:
paths:
- .npm/
artifacts:
paths:
- dist/
expire_in: 1 week
deploy:
dependencies:
- build
script:
- deploy-script dist/
Keep cache and artifact paths in separate directories, since GitLab’s caching docs note that caches are restored before artifacts, and overlapping paths can cause one to overwrite the other. Use dependencies or needs: [job: build, artifacts: true] to control exactly which artifacts a downstream job pulls in, rather than downloading everything by default.

Cache keys and invalidation strategies
A cache is only as good as its key. A static key reused across every branch and every dependency change will eventually serve stale data, and that’s how “works on my machine, fails in CI” bugs sneak in.
cache:keysets a static or variable-based identifier, useful for simple, branch-scoped caching.cache:key:filesgenerates a key from the content of specific files, so a change topackage-lock.jsonorpoetry.lockautomatically produces a new key, per GitLab’s caching documentation.cache:key:files_commitsties the key to the latest commit touching those files, giving finer-grained invalidation than a content hash alone.- Fallback keys let a job recover a reasonably fresh cache after a miss instead of starting from nothing.
Lockfile-based keys are the safer default because they invalidate exactly when dependencies actually change, not on every commit. A pattern like key: { files: [package-lock.json] } with a fallback key such as $CI_COMMIT_REF_SLUG gives you precision plus a reasonable backup. Branch-only keys are simpler to write but tend to accumulate stale packages, since nothing forces a refresh when the lockfile changes but the branch name doesn’t. For more detail on this approach, see keying caches on lockfiles.
Storage backends, availability, and transfer costs
Where your cache physically lives changes how often it actually helps. A single long-lived runner can keep cache on local disk and get near-perfect hit rates. Autoscaled runners, especially on Kubernetes, spin up fresh instances constantly, so local disk cache rarely survives long enough to matter.
- Runner-local cache works well for static, dedicated runners but disappears the moment the runner is recycled.
- Distributed cache backends, including AWS S3, MinIO, Google Cloud Storage, and Azure Blob Storage, let autoscaled runners share a common cache instead of starting cold every time, according to GitLab Runner’s performance documentation.
- Colocating object storage with your runners cuts transfer latency and improves effective hit rates in practice.
GitLab-hosted runners remove caches that haven’t been updated in 14 days and document a 5 GB maximum compressed upload size for cached content, according to GitLab’s caching documentation. That eviction window matters for low-traffic branches: a cache untouched for two weeks is simply gone on the next run.
Archive creation, upload, download, and extraction all take time, and for small dependency sets that overhead can exceed the cost of just reinstalling from scratch. Before committing to a caching strategy, check whether rebuild time genuinely beats transfer time. If your dependency install takes 20 seconds and cache download takes 15, you’ve saved almost nothing, and you’ve added a new failure mode.
Common pitfalls and best practices
Most caching problems trace back to a handful of repeated mistakes. Here’s what to watch for and how to fix it.
- Overlapping paths: don’t store cache and artifacts in the same directory, since caches restore before artifacts and one can clobber the other.
- Treating cache as a correctness boundary: a pipeline that only works when the cache hits is fragile by design; always confirm a clean, cache-miss build still succeeds, a point echoed in GitLab’s visual guide to caching.
- Oversized caches: keep cached directories focused on what’s actually reused, not entire working trees.
- Unfiltered artifact downloads: use
dependenciesorneedswithartifacts: trueto limit what downstream jobs pull instead of grabbing everything by default. - Careless
expire_invalues: set retention deliberately, since the default 30-day expiration may delete outputs you still need.
Pro Tip: Log cache hit or miss status and artifact transfer time in your pipeline output. A week of that data tells you more about what’s worth caching than any general rule.
Field notes from rolling this out in production
A few patterns show up repeatedly once you start measuring instead of guessing. Baseline your pipeline time before touching anything, then pilot new cache keys on merge request pipelines where a bad key only costs you a rerun, not a broken main branch. Watch hit rate and transfer time together, not separately, since a high hit rate with slow transfer isn’t actually a win.
Autoscaled runners are the most common source of “the cache isn’t working” complaints, because a fresh node has no local cache to restore. Moving to a distributed backend usually fixes it, particularly for projects with many small dependencies rather than one enormous one. For a deeper walkthrough of rollout patterns and metrics worth tracking, see GitLab CI caching strategies and the notes on runner autoscaling behavior.
Balancing speed and determinism in your pipeline
My working rule is simple: artifacts protect correctness, cache buys speed, and you should never let speed decisions quietly become correctness decisions. If a job’s success depends on a cache hit, that’s a bug waiting for a slow week. Reach for cache when rebuild cost is genuinely high and the data is disposable. Reach for artifacts whenever the next job’s behavior depends on getting the exact same bytes. Adopt either incrementally, watch the numbers, and let the data tell you where the real bottleneck is instead of assuming.
— James
How DevOps AI ToolKit can help you tune this
If your pipelines are slow, inconsistent, or randomly rebuilding things they shouldn’t, a second set of eyes on the YAML often finds the fix faster than another round of trial and error. Our Terraform / IaC Audit covers pipeline configuration alongside infrastructure code, and hourly consulting is available for teams that want a focused review of cache keys, artifact retention, and runner setup rather than a full engagement.
We work alongside your existing team to diagnose and fix the pipeline, not to replace the engineers who own it. Start with the work-with-me page to see current services and book a review.
Sources
For the official reference, start with Caching in GitLab CI/CD, Job artifacts, and Speed up job execution. For implementation depth, see our posts on artifacts and reports in merge requests and a Terraform pipeline case study.
FAQ
How does GitLab cache work?
GitLab cache stores files, typically dependency folders, on runner-local disk or a configured distributed backend so later jobs can skip redoing the same work. It’s identified by a cache:key and restored before artifacts, which is why overlapping paths can cause conflicts, according to GitLab’s caching documentation.
What are artifacts in GitLab?
Artifacts are job outputs, such as compiled binaries, test reports, or signed bundles, that GitLab uploads and stores so later jobs or reviewers can use them. They’re available in the project’s Build artifacts view and expire according to expire_in or the project default, per GitLab’s job artifacts documentation.
How long do GitLab caches last?
Cache retention depends on runner storage rather than a fixed setting in your YAML. On GitLab-hosted runners, caches not updated in 14 days are removed, and cached uploads are capped at 5 GB compressed, according to GitLab’s caching documentation.
How do I clear the cache in GitLab pipelines?
The most reliable way is to change the cache:key, since a new key forces GitLab to start fresh rather than reuse the old cache. You can also clear caches manually from the project’s CI/CD settings in the GitLab UI if you need an immediate reset without editing the pipeline configuration.
Why is my GitLab CI cache not working?
The most common cause is an autoscaled runner spinning up a fresh instance with no local cache to restore, which distributed cache backends can fix. Another frequent cause is a cache key that changes too often, such as one tied to the commit SHA instead of a lockfile, which prevents reuse between runs.
Recommended
- GitLab CI Caching Strategies: A Deep Dive That Actually
- GitLab CI Error Guide: ‘Uploading artifacts … too large’
- Keying GitLab CI Caches on Lockfiles With cache:key:files
- Cutting Your GitLab CI Bill: A Practical Guide to Pipeline
Get 500 Battle-Tested DevOps AI Prompts — Free
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.