Skip to content
DevOps AI ToolKit
All guides
AI for Automation By James Joyner IV · · 12 min read

Cluster Operators: 15 Second Check for Evicted Pods in Kubernetes

Operator first on call playbook for cluster operators: confirm an eviction, capture the kubelet eviction log line, and triage node memory, disk, or PID...

Cluster Operators: 15 Second Check for Evicted Pods in Kubernetes

An “Evicted” pod is one that the kubelet or the API has terminated to reclaim node resources. The pod status shows phase Failed with reason Evicted, and that status alone tells you almost nothing about the root cause. Before you delete anything, identify which node hosted the pod, check its node conditions, and pull the kubelet logs from around the time it happened.


TL;DR:

  • The primary causes of pod eviction are node-pressure issues, API-initiated requests, and scheduler preemption, each requiring different diagnosis approaches.
  • Confirming eviction involves checking the pod and node statuses, kubelet logs, and container exit codes to differentiate node-level evictions from container crashes.
  • Accurate resource requests, resource quotas, and setting PodDisruptionBudgets help prevent evictions, but capacity planning remains crucial for avoiding recurring issues.
  • Adjust kubelet eviction thresholds and soft thresholds with grace periods and minimum reclaim settings to fine-tune how quickly and aggressively evictions occur.
  • Collecting detailed evidence such as pod YAML, kubelet logs, and node metrics before cleanup enables faster root cause analysis and targeted resolution.

Table of Contents

What causes pods to get evicted in Kubernetes

Eviction happens through three distinct mechanisms, and mixing them up wastes time during an incident. Node-pressure eviction is the kubelet acting alone when a node runs low on memory, disk, or process IDs. API-initiated eviction happens when something (a person, a controller, the cluster autoscaler) sends an Eviction request through the API, which respects PodDisruptionBudgets and the pod’s grace period unless resource pressure overrides it. Then there’s plain scheduler preemption, where a higher PriorityClass pod bumps a lower-priority one to free up space. I’ve spent more on-call hours than I’d like confusing preemption with node-pressure eviction, and the fix for each is completely different.

The kubelet watches a specific set of eviction signals: memory.available, nodefs.available, imagefs.available, nodefs.inodesFree, and pid.available. Cross a hard threshold and the kubelet reclaims immediately, no grace period.

QoS class decides who goes first when the kubelet needs to free resources under pressure:

  • BestEffort pods (no requests or limits set) get evicted first, every time.
  • Burstable pods (requests below limits) go next if pressure continues.
  • Guaranteed pods (requests equal limits) are evicted last, only under severe pressure.

PriorityClass adds another axis on top of QoS, and the two sometimes pull in different directions, which is worth understanding through our piece on PriorityClass and preemption. One more trap: a container that gets OOMKilled by the kernel looks superficially similar to an eviction in a dashboard, but it’s a container-level event, not a node-level one, and the fix is different too.

How to confirm a pod was actually evicted

Don’t trust a dashboard summary. Run these checks in order before you touch anything:

  1. Run kubectl describe pod <name> and look for status.phase: Failed with reason: Evicted. The message field usually names the exact signal, something like “The node was low on resource: ephemeral-storage.”
  2. Run kubectl describe node <node> and check the Conditions section for MemoryPressure, DiskPressure, or PIDPressure set to True, along with the lastTransitionTime so you can line it up with the eviction timestamp.
  3. SSH into the node and run journalctl -u kubelet filtered around that timestamp, looking for lines containing “eviction manager: attempting to reclaim” followed by the signal name.
  4. Check the pod’s last container state and exit code. Evicted pods generally have no meaningful exit code because the kubelet killed the pod wholesale rather than the container exiting on its own, which is your clearest signal it wasn’t OOMKilled or a CrashLoopBackOff.

If the node conditions never flipped and the kubelet logs are quiet, you’re probably not looking at a real eviction, you’re looking at something the scheduler or a controller did on purpose. That distinction changes your entire next step.

Investigation checklist for finding the resource culprit

Once you’ve confirmed a real eviction, work the node systematically rather than guessing. This is the sequence I run every time, roughly in order of how fast each step pays off:

  1. Grab the evicted pod’s YAML and events with kubectl get pod <name> -o yaml and kubectl describe pod <name> before anyone deletes it.
  2. Pull node conditions and timestamps again with kubectl describe node <node>, this time cross-referencing against your monitoring for a fuller picture of when pressure started.
  3. Read the kubelet logs for the exact reclaim signal and how much it tried to free. This tells you whether it was memory, disk, or PIDs.
  4. Check disk space and inodes directly on the node with df -h and df -i, then look inside /var/log/pods for oversized or unrotated log files.
  5. Hunt down the heavy consumer with du -sh /var/lib/docker/* 2>/dev/null or similar, find / -size +500M, and ps aux --sort=-%mem for memory or PID leaks. Check the image store too, since unused images consuming imagefs are one of the most common causes I see.

A few things worth checking alongside the main sequence:

  • Confirm whether the same pod keeps getting rescheduled and evicted again on the same node, which points to a systemic resource problem rather than a one-off spike.
  • Check dmesg for OOM killer activity if memory pressure is suspected, since it often fires before the kubelet’s own eviction manager does.
  • Snapshot your metrics dashboard for the affected node covering the hour before the eviction, before that data ages out of retention.

Pro Tip: Copy the kubelet’s exact eviction manager log line into your incident notes verbatim. It names the signal and the amount reclaimed, and that single line saves you from re-deriving the cause later.

For a postmortem, collect the pod YAML, the node’s kubelet logs, dmesg output if memory was involved, and a metrics snapshot from just before the event. Do this before the pod garbage collector cleans anything up.

Preventing evictions before they hit production

Most eviction incidents trace back to configuration gaps that were fixable weeks earlier. Start here:

  • Set accurate requests and limits on every workload, and push critical services toward Guaranteed QoS, since requests versus limits is the single highest-leverage fix available to most teams.
  • Assign PriorityClass deliberately so a batch job never accidentally preempts a production database.
  • Apply ResourceQuotas and LimitRanges at the namespace level to stop one misbehaving deployment from starving its neighbors, a practice Kubernetes’ own scheduling documentation recommends directly.
  • Configure PodDisruptionBudgets for replicated workloads, understanding that PDBs govern voluntary API-initiated evictions and offer no protection against node-pressure eviction, which the kubelet performs unilaterally. Our breakdown of PodDisruptionBudgets during cluster maintenance covers where the boundary actually sits.
  • Cap emptyDir volumes with sizeLimit, rotate container logs, and enable periodic image garbage collection so nodefs and imagefs pressure never builds quietly in the background.
  • Alert on the eviction signals themselves, memory.available trending down, nodefs.available approaching the soft threshold, rather than waiting for the eviction event to fire.

Capacity planning matters more than most teams admit. If your autoscaler consistently provisions nodes just barely large enough for average load, you have zero headroom for a single noisy pod, and evictions become a recurring background hum rather than a rare incident.

Pro Tip: Treat repeated evictions on the same node as a sizing problem, not a kubelet problem. The kubelet is doing exactly what it’s supposed to do.

Kubelet eviction thresholds and how to tune them

The kubelet ships with sensible defaults, but production workloads often need adjustment. Default hard eviction thresholds trigger immediate pod termination with no grace period, while soft thresholds give you a configurable buffer.

SignalDefault hard thresholdWhat it means in practice
memory.availableBelow 100Mi for Linux nodesNode is nearly out of usable memory
nodefs.availableBelow the kubelet’s configured thresholdRoot filesystem is nearly full
imagefs.availableBelow the kubelet’s configured thresholdImage storage filesystem is nearly full
nodefs.inodesFreeBelow 5% for Linux nodesFilesystem is out of inodes, even with free space

Soft thresholds pair with eviction-soft-grace-period (how long the condition must persist before the kubelet acts) and eviction-max-pod-grace-period (the cap on how long a pod gets to terminate cleanly once eviction starts). eviction-minimum-reclaim matters more than it looks: without it, the kubelet can reclaim just enough to dip below a threshold, then trigger another eviction moments later. Setting a sensible minimum reclaim value stops that loop.

PID-based eviction, governed by pid.available, catches a class of failure that memory and disk monitoring completely misses: a workload spawning runaway processes or threads. PID limiting documentation covers pod-level PID caps directly. Kubernetes v1.36 also introduced tiered Memory QoS reservation, which changes how memory.min and memory.low get set for Guaranteed and Burstable pods and, by extension, how close to the edge a pod runs before eviction becomes likely.

Cleaning up evicted pods without losing forensic data

Don’t rush to delete Evicted pods the moment you see them. They stick around with phase Failed until the pod garbage collector cleans them up based on terminated-pod-gc-threshold, and until then they carry the events and logs you need for root cause work.

  1. Confirm you’ve captured the pod’s YAML, events, and container logs before deleting anything.
  2. Delete individually with kubectl delete pod <name> once you’ve confirmed root cause, or scope a bulk cleanup with kubectl delete pods --field-selector=status.phase=Failed -n <namespace>.
  3. For automation, use a label-based retention window (keep evicted pods for 24 to 48 hours) run through a CronJob, rather than deleting on sight.
  4. Export anything you need for a postmortem, YAML, logs, and the kubelet eviction line, before the retention window closes.

What actually separates a fast eviction fix from a slow one

The playbook that saves time on-call is short enough to hold in your head: confirm the phase and reason, check node conditions and timestamps, pull the kubelet log line naming the signal, and only then decide whether it’s a sizing problem, a disk problem, or a PID leak. Most of the delay I see in incidents comes from skipping straight to deletion and losing the exact evidence that would have made the diagnosis obvious. For a deeper look at reproducing a similar failure on a distroless image where kubectl exec isn’t available, our guide on debugging distroless pods walks through ephemeral debug containers. Graceful termination behavior, including how endpoints report readiness during pod removal, is also worth understanding alongside eviction, and Singleclic’s overview of Kubernetes in cloud-native apps covers that draining behavior well.

Illustrated Kubernetes eviction diagnosis path

Why most eviction advice skips the part that matters

Most eviction write-ups stop at “set your requests and limits,” which is true and also not the whole story. The bigger issue is that eviction gets treated as a bug to fix rather than the kubelet doing its job correctly under bad conditions. A pod that keeps getting evicted and rescheduled onto the same undersized node isn’t a Kubernetes problem, it’s a capacity and QoS decision someone made months ago and never revisited.

What I think operators underestimate is how much PID pressure and inode exhaustion get ignored compared to memory and disk. Teams monitor memory.available closely and completely miss pid.available until a fork-heavy workload takes down a node. Prioritize reading the exact kubelet signal before you touch requests and limits. Guessing at “probably memory” when it was actually inodes wastes an afternoon that a fifteen-second log check would have saved.

— James

How DevOps AI ToolKit helps you diagnose and prevent evictions

If you’d rather have a second set of eyes on a node that keeps evicting pods, our Kubernetes Health Check is a one-off, fixed-price review built around exactly this kind of resource-pressure diagnosis. Pair it with an Observability Review if your alerting missed the warning signs before the eviction fired, or lean on AI Incident Response tooling to speed up triage the next time it happens.

Devopsaitoolkit

Check current details and book a review on our pricing page.

Where to verify these details yourself

Sources

FAQ

What does it mean for a pod to be evicted?

An evicted pod is one the kubelet terminated to reclaim node resources like memory, disk, or process IDs, or one removed through an API-initiated eviction request. The pod’s status shows phase Failed with reason Evicted, and it stays visible until garbage collected.

How to prevent pod eviction?

Set accurate requests and limits so critical pods run at Guaranteed QoS, apply ResourceQuotas and LimitRanges, and keep an eye on the core eviction signals before they cross a threshold. Rotating logs, capping emptyDir size, and running periodic image garbage collection close off the most common disk-pressure causes.

What does “evicted” mean?

“Evicted” is the status reason Kubernetes assigns to a pod that was forcibly terminated, either by the kubelet responding to node resource pressure or through a voluntary API-initiated eviction. It’s distinct from a container crash or an OOMKill, which happen at the container level rather than the node level.

How to gracefully shut down a pod?

A graceful shutdown relies on PreStop hooks running before the TERM signal and completing within terminationGracePeriodSeconds, as described in Kubernetes’ termination behavior documentation. During this window the pod’s endpoint reports as not ready so load balancers can drain traffic before the container actually stops.

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week
Free download · 368-page PDF

Get 500 Battle-Tested DevOps AI Prompts — Free

500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.

  • 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
  • Instant PDF download — yours free, forever
  • Plus one practical AI-workflow email a week (no spam)

Single opt-in · unsubscribe anytime · no spam.