Skip to content
DevOps AI ToolKit
Home

OpenStack Troubleshooting: Nova, Neutron, Cinder & Keystone

Troubleshoot Nova, Neutron, Cinder, RabbitMQ, and Keystone with AI-assisted workflows.

148 copy-paste prompts · 142 in-depth guides Jump to prompts Jump to guides

What actually breaks, and what to check first

OpenStack failures are almost never confined to the service that reported them. An instance that will not spawn is more often a Neutron port-binding problem, a Cinder backend timing out, or a Placement allocation that cannot be satisfied, than it is a Nova defect. The error surfaces where the request ended, not where it went wrong.

This makes triage order more valuable than any single command. The dependency chain is consistent: Keystone authenticates, Placement decides where capacity exists, Nova schedules, Neutron wires the network, Cinder provides storage, and RabbitMQ carries every one of those conversations. Work down that chain and the real fault appears quickly; start at the symptom and you can spend a day in the wrong service.

RabbitMQ deserves particular attention. It is the message bus for the entire control plane, so when it degrades every service reports timeouts simultaneously and the cluster looks like it has failed everywhere at once. A broad, correlated failure across unrelated services is a bus symptom until proven otherwise.

Triage order

Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.

  1. Is Keystone issuing tokens?

    openstack token issue

    If this fails, everything downstream fails and every other error you see is noise. Check clock skew first — Fernet tokens are time-sensitive, and a compute node drifting by minutes produces authentication errors that look like credential problems.

  2. Are all control-plane services up and reporting?

    openstack compute service list && openstack network agent list && openstack volume service list

    `State: down` on an agent is definitive — that node is not participating. An agent that is `up` but has an old `Updated At` is worse news: it means heartbeats are arriving late, which usually points at the message bus rather than the agent.

  3. Is RabbitMQ healthy?

    rabbitmqctl cluster_status && rabbitmqctl list_queues name messages consumers | sort -k2 -rn | head

    A partitioned cluster or queues with many messages and zero consumers explains simultaneous timeouts across unrelated services. This check ends more multi-service OpenStack incidents than any other.

    → RabbitMQ in OpenStack

  4. Does Placement believe there is capacity?

    openstack allocation candidate list --resource VCPU=1 --resource MEMORY_MB=2048

    Empty output means no host can satisfy the request, and "No valid host was found" is the expected consequence rather than a scheduler bug. Check disabled hosts, allocation ratios and NUMA/pinning constraints before you look at Nova.

    → No valid host was found

  5. What did the instance actually fail on?

    openstack server show <id> -f value -c fault && openstack server event list <id>

    The fault field carries the real exception; the event list shows how far the build progressed. Failure at networking points to Neutron, failure at block-device mapping points to Cinder — this single step tells you which service to open next.

Diagnostic commands

Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.

openstack server show <id> -f value -c fault
Read-only

Retrieve the recorded failure reason for an instance.

How to read it: Contains the underlying exception, often with the responsible service named. Far more informative than the ERROR status shown in the instance list.

openstack network agent list --agent-type l3
Read-only

Check L3 agent health across the cluster.

How to read it: An agent that is down or has stale heartbeats means routers it hosts have no gateway — instances lose external connectivity while remaining fully healthy internally.

openstack port list --device-id <instance-id> -f yaml
Read-only

Inspect an instance’s ports and their binding state.

How to read it: `binding_vif_type: binding_failed` means Neutron could not wire the port — usually an ML2 mechanism-driver or agent mismatch on that specific host.

openstack volume show <id> -f value -c status -c os-vol-host-attr:host
Read-only

Check a volume’s state and which backend owns it.

How to read it: A volume stuck in `attaching` or `detaching` usually means the cinder-volume service for that backend is unresponsive, not that the volume is damaged.

rabbitmqctl list_queues name messages consumers
Read-only

Find queues that are accumulating messages with nothing draining them.

How to read it: High message count with zero consumers is a service that has stopped listening — restart that service, not RabbitMQ.

openstack compute service set --disable --disable-reason "<why>" <host> nova-compute
Changes state

Stop the scheduler placing new instances on a host.

How to read it: The correct first move when a compute node is suspect. Existing instances keep running; only new placement stops. Always record a reason — an undocumented disabled host is found months later.

nova-manage db archive_deleted_rows --max_rows 10000
Destructive

Archive soft-deleted rows out of the live Nova tables.

How to read it: Addresses gradual API slowdown on long-lived clouds. Take a database backup first — this modifies production tables.

Failure modes

These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.

"No valid host was found" when launching an instance.

Cause:
No compute node satisfies the flavour’s requirements — capacity, NUMA topology, CPU pinning, host aggregates or a disabled service.
Fix:
Query Placement for allocation candidates against the flavour’s resources. Empty output confirms the scheduler is correct and moves the investigation to capacity or constraints.

Full walkthrough →

Instance builds, then fails at the networking stage.

Cause:
Port binding failed — the ML2 driver on that host cannot bind the requested network type.
Fix:
Inspect the port’s `binding_vif_type`. `binding_failed` means agent/driver mismatch on the host; confirm the L2 agent is running and configured for the same mechanism driver as the network.

Full walkthrough →

Every service reports timeouts at the same moment.

Cause:
RabbitMQ is partitioned or overloaded; the control plane has lost its message bus.
Fix:
Check `rabbitmqctl cluster_status` for partitions before restarting anything. Restarting OpenStack services against a partitioned broker adds load and prolongs the outage.

Full walkthrough →

Volumes hang in attaching or detaching.

Cause:
The cinder-volume service for that backend is down, or the storage backend is not responding.
Fix:
Check `openstack volume service list` for the owning host. Resetting volume state hides the symptom without fixing the backend and can desynchronise Cinder from what the hypervisor actually has attached.

Authentication fails intermittently across the cloud.

Cause:
Fernet key rotation is inconsistent across nodes, or clocks have drifted.
Fix:
Confirm the key repository is identical on every controller and that NTP is synchronised. Intermittency is the clue: a wholly wrong credential fails every time, whereas key or clock problems fail on some nodes only.

Horizon returns 500 while the CLI works normally.

Cause:
A dashboard-layer problem — a stale session store, a policy file mismatch, or an Apache/WSGI error — rather than an API failure.
Fix:
Read the Horizon error log directly. When the CLI succeeds, the APIs are fine by definition and the fault is confined to the dashboard.

Full walkthrough →

Common mistakes

  • Restarting services in symptom order rather than dependency order, which restarts the bus last and prolongs the outage.
  • Resetting a stuck volume’s state to force progress, desynchronising Cinder from the hypervisor’s actual attachments.
  • Deleting an ERROR instance before capturing `openstack server show --fault`, discarding the only record of the cause.
  • Ignoring clock skew. NTP drift produces authentication and token errors that look nothing like a time problem.
  • Treating "the agent is up" as healthy without checking the heartbeat timestamp — a stale heartbeat is a bus symptom.
  • Rotating Fernet keys on one controller and not the others, which produces intermittent auth failures that are hard to reproduce.

Frequently asked questions

Why does OpenStack report "No valid host was found" when there is clearly free capacity?

Because free capacity and *schedulable* capacity are different things. Placement filters on the flavour’s full requirements — NUMA topology, CPU pinning, host aggregates, allocation ratios and disabled services all remove hosts from consideration. Query allocation candidates directly: if it returns nothing, the scheduler is behaving correctly and the constraint is the thing to investigate.

How do I tell whether a problem is RabbitMQ or the service reporting it?

By its breadth. A single service failing is a service problem; several unrelated services timing out at the same moment is the bus. Confirm with `rabbitmqctl cluster_status` for partitions and `list_queues` for queues with messages and no consumers — a queue nobody is draining names the service that stopped listening.

What should I check first when instances fail to spawn?

The fault field on the instance itself, then the event list. Together they tell you how far the build got and which service rejected it. Reading nova-compute logs before that is guesswork, because the failure is frequently in Neutron or Cinder and Nova is only reporting what it was told.

Is it safe to restart nova-compute on a busy hypervisor?

Generally yes — running instances are not affected, because they are managed by libvirt rather than by the agent. What stops is the ability to act on new requests for that host. Disable the service first so the scheduler stops sending it work, then restart, then re-enable once it reports healthy.

Why do authentication errors appear on some nodes but not others?

That pattern points at Fernet key distribution or clock skew rather than credentials. Tokens are validated against a key repository that must be identical everywhere, and their validity window depends on the node’s clock. A genuinely wrong credential fails uniformly; partial failure means the nodes disagree about keys or time.

Prompts

Guides

Recommended tools

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week