Skip to content
DevOps AI ToolKit
Home

Linux Production Troubleshooting

Diagnose, automate, and harden Linux servers using AI assistants. Ubuntu, RHEL, Debian, Rocky.

152 copy-paste prompts · 203 in-depth guides Jump to prompts Jump to guides

What actually breaks, and what to check first

Under pressure, the useful skill is not knowing more Linux commands — it is knowing which four to run first, and what their output rules out. A server that is "slow" can be starved of CPU, blocked on I/O, swapping, out of file descriptors, or perfectly healthy while a downstream dependency times out. Those look the same from the application and completely different from the shell.

Load average is the most misread number in Linux. It counts processes in the runnable *and* uninterruptible-sleep states, so a load of 40 on an 8-core box can mean CPU saturation or it can mean forty processes blocked on a failing disk while the CPUs sit idle. Reading it alongside CPU idle and I/O wait is what makes it meaningful.

The other recurring surprise is that "disk full" often is not. A filesystem with free space that still refuses writes has usually exhausted inodes, or is holding space in files that were deleted while a process still had them open — a case where `df` and `du` disagree and only `lsof` explains why.

Triage order

Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.

  1. Is it CPU, I/O, or memory?

    uptime && vmstat 1 5

    High load with high `us`/`sy` is CPU. High load with high `wa` and low CPU is I/O — the processes are blocked, not busy. Non-zero `si`/`so` means swapping, which makes everything slow regardless of the other numbers.

  2. Is memory actually exhausted?

    free -h && dmesg -T | grep -i "out of memory\|killed process" | tail

    Ignore the used column; read `available`, which accounts for reclaimable cache. An OOM kill in dmesg is definitive and names the process — the application log will usually show nothing, because the kernel gave it no chance to write one.

  3. Is the disk full — and is it space or inodes?

    df -h && df -i

    Both must be checked. Free space with exhausted inodes produces "no space left on device" on a filesystem that looks fine, and it is common wherever many small files accumulate.

  4. Is space held by deleted-but-open files?

    lsof +L1 2>/dev/null | head -20

    Explains the case where `df` shows full and `du` does not. A log rotated without signalling the writer leaves the old file deleted but open, holding its space until the process is restarted or the descriptor is closed.

  5. Is the failing unit actually failing, or not running at all?

    systemctl --failed && systemctl status <unit> --no-pager -l

    A unit that exits 0 immediately is "successful" to systemd, which is the classic silent deployment failure. Read the actual exit status and the last log lines rather than the green/red indicator.

  6. Is it DNS?

    resolvectl status; getent hosts <name>; dig +short <name>

    `getent` uses the full NSS stack — the same path your application takes. `dig` bypasses it and queries the resolver directly, so when the two disagree the problem is in NSS configuration rather than in the DNS server.

Diagnostic commands

Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.

vmstat 1 5
Read-only

Sample CPU, memory, swap and I/O together over five seconds.

How to read it: The fastest way to classify a slow server. `wa` high means blocked on I/O; `si`/`so` non-zero means swapping; `r` consistently above core count means genuine CPU saturation.

journalctl -u <unit> -p err --since "1 hour ago" --no-pager
Read-only

Show error-level messages for one unit in a bounded window.

How to read it: Filtering by priority and time turns an unreadable log into a short list. Add `-b` to restrict to the current boot when you suspect a restart is involved.

ss -tulpn | grep LISTEN
Read-only

List listening sockets with the owning process.

How to read it: Answers "is it actually listening, and on which address". A service bound to 127.0.0.1 rather than 0.0.0.0 is reachable locally and refuses every remote connection — a frequent cause of "connection refused" that is not a firewall.

lsof +L1
Read-only

Find deleted files still held open by a process.

How to read it: The explanation whenever `df` and `du` disagree. The space is not returned until the descriptor is closed, so restarting the holding process reclaims it immediately.

journalctl -k --since "30 min ago" | grep -iE "oom|blocked|i/o error|ext4|xfs"
Read-only

Read kernel messages for storage and memory events.

How to read it: Kernel-level I/O errors and OOM kills appear only here. If the application log is silent about a crash, this is usually where the answer is.

systemctl cat <unit>
Read-only

Show the effective unit file including every drop-in override.

How to read it: Prevents editing the wrong file. Drop-ins in `/etc/systemd/system/<unit>.d/` override the packaged unit, and `systemctl cat` is the only view that shows the merged result.

kill -TERM <pid>
Changes state

Ask a process to shut down cleanly.

How to read it: Always before `-9`. SIGKILL prevents flushing buffers and releasing locks, which turns a clean restart into a recovery.

Failure modes

These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.

Load average is very high but CPUs are mostly idle.

Cause:
Processes blocked in uninterruptible sleep on I/O — typically a failing disk or a stalled network filesystem.
Fix:
Confirm with `vmstat` (high `wa`) and check kernel messages for I/O errors. Adding CPU capacity changes nothing here, because nothing is waiting for CPU.

"No space left on device" but `df -h` shows free space.

Cause:
Inode exhaustion, or space held by deleted-but-open files.
Fix:
Check `df -i` first, then `lsof +L1`. Both are quick and between them they explain nearly every instance of this.

A service is running but nothing can connect.

Cause:
Bound to loopback rather than all interfaces, or a firewall rule is dropping traffic.
Fix:
`ss -tulpn` shows the bind address definitively. Check that before investigating the network — a loopback bind looks exactly like a firewall block from the client.

A process was killed with no log entry explaining why.

Cause:
The OOM killer terminated it. The kernel does not give the process an opportunity to log.
Fix:
Search `dmesg -T` for the OOM record; it names the process and the memory state at the time. Then either raise the limit or fix the leak — restarting alone guarantees a repeat.

systemd reports the service started successfully but nothing is running.

Cause:
The unit’s main process exited 0 immediately, which systemd treats as success for `Type=simple`.
Fix:
Read the actual logs rather than the status colour. This is the standard silent deployment failure, and the pattern is a green status with no listening socket.

Name resolution works with `dig` but fails in the application.

Cause:
NSS configuration differs from the resolver — a stub resolver, an `/etc/hosts` entry, or an nsswitch ordering issue.
Fix:
Compare `getent hosts` with `dig`. `getent` follows the same path your application does; when the two differ, the DNS server is not the problem.

Common mistakes

  • Reading load average without also reading CPU idle and I/O wait. Alone it cannot distinguish a busy box from a blocked one.
  • Reaching for `kill -9` first. It skips clean shutdown, so buffers are unflushed and locks are left behind.
  • Judging memory pressure by the used column instead of `available`, and concluding the server is out of memory when it is holding reclaimable cache.
  • Rotating logs without signalling the writing process, which holds the old file open and returns none of the space.
  • Editing a packaged unit file directly instead of adding a drop-in — the change is silently reverted on the next package update.
  • Deleting large files to free space while a process still has them open, which frees nothing until that process restarts.

Frequently asked questions

What does a high load average actually mean?

That many processes are either running or blocked in uninterruptible sleep — and those are very different conditions. On an 8-core machine a load of 40 could be genuine CPU saturation, or forty processes stalled on a failing disk while the CPUs idle. Read it together with CPU idle and I/O wait from `vmstat`; on its own the number cannot distinguish the two.

Why does df say the disk is full when du says it is not?

Because a process still holds a deleted file open. The directory entry is gone, so `du` no longer counts it, but the blocks are not released until the last file descriptor closes — which is why `df` still sees them as used. `lsof +L1` lists exactly these files, and restarting the holding process returns the space immediately.

How do I find out why a process was killed?

Check `dmesg -T` for an OOM record. When the kernel’s OOM killer terminates a process it gets no chance to write a log entry, so the application log shows a clean stop and nothing else — the only account of what happened is in kernel messages, and it includes which process was chosen and the memory state at the time.

Should I edit a systemd unit file directly?

No — use `systemctl edit <unit>` to create a drop-in. Changes made directly to a packaged unit are overwritten on the next package update, silently and usually at an inconvenient time. `systemctl cat <unit>` shows the merged result of the base unit and every override, which is what actually runs.

Why does DNS work with dig but fail from my application?

Because they take different paths. `dig` queries a resolver directly, while your application goes through NSS — `/etc/hosts`, `nsswitch.conf`, a local stub resolver, possibly a container’s own resolver. `getent hosts` uses that same NSS path, so comparing the two tells you immediately whether the DNS server or the local configuration is at fault.

Prompts

Guides

Recommended tools

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week