DHCP Resilience for OpenStack Neutron Ops Starts with dhcp_agent.ini
Ops-first OpenStack Neutron DHCP reference: hands-on dhcp_agent.ini tuning, dhcp_agents_per_network HA steps, dnsmasq 2.63+ checks, and OVN migration notes.
The Neutron DHCP agent runs dnsmasq in per-network qdhcp namespaces to provide DHCP leases, DNS forwarding, and metadata services. You need it for DHCP-based IP assignment unless you run OVN native DHCP, which handles that logic in the data plane instead. The primary tuning surface is dhcp_agent.ini, and for high availability you will lean on dhcp_agents_per_network.
TL;DR:
- Multiple DHCP agents per network improve high availability but do not provide IPv6 isolated metadata redundancy due to route injection limitations.
- Proper tuning of resync_interval and resync_throttle prevents DHCP agent slowdowns in environments with high port churn.
- Explicit DNS server configuration on subnets or via dnsmasq_dns_servers is more secure and predictable than relying on local_resolv, which can leak internal DNS infrastructure.
- Running neutron-ovs-cleanup after node reboots is essential to prevent stale Open vSwitch state from blocking DHCP namespace creation.
- DHCPv6 support requires dnsmasq version 2.63 or later, and mismatched ipv6_ra_mode and ipv6_address_mode can cause addressing issues in dual-stack deployments.
Table of Contents
- Role and architecture of the Neutron DHCP agent
- Key dhcp_agent.ini options you must know and how to tune them
- DNS handling and dnsmasq behavior
- High availability and scheduling for DHCP agents
- IPv6 and DHCPv6 considerations
- Deployment and scaling best practices
- Operations and troubleshooting checklist
- Operational lessons from running DHCP at scale
- How Devopsaitoolkit can help with OpenStack DHCP and Neutron operations
- Primary OpenStack docs and internal guides for reference
- Sources
- FAQ
Role and architecture of the Neutron DHCP agent
Every network that needs DHCP gets its own network namespace, named qdhcp-
When you update a subnet, such as adding a host route or changing the DNS servers, the agent detects the change and restarts the relevant dnsmasq process to pick up the new configuration. That restart is quick, but it briefly interrupts DHCP renewals on that network, which matters if you are making changes during business hours.
The DHCP agent does not work in isolation:
- It coordinates with the L3 agent on networks that route through a virtual router.
- It provides the metadata proxy directly on isolated networks that lack a router, since those instances cannot reach the metadata service any other way.
- On networks with a router, metadata typically flows through the L3 agent instead of the DHCP namespace.
OVN deployments can skip this entire agent tier. Instead of dnsmasq processes and namespaces, OVN native DHCP answers DHCP requests directly from the OVN data plane, which removes a layer of processes you would otherwise have to monitor.
Pro Tip: When diagnosing a network with no DHCP response, confirm the qdhcp namespace exists before you touch anything else: ip netns list | grep qdhcp.
Key dhcp_agent.ini options you must know and how to tune them
The Neutron dhcp_agent.ini reference defines the settings that shape how the agent behaves in production. A handful of them deserve real attention:
- dhcp_driver sets which backend manages the network namespace and dnsmasq lifecycle. The default Linux dnsmasq driver covers most deployments.
- dhcp_confs points to the state directory where per-network lease files, host files, and PID files live, useful when you need to inspect a specific network’s DHCP state by hand.
- dnsmasq_config_file lets you inject a custom dnsmasq configuration snippet for options the agent does not expose directly.
- resync_interval and resync_throttle control how often the agent reconciles its state with Neutron and how aggressively it retries. In high-churn environments, setting resync_throttle too low can push the agent into a busy loop that hammers the message bus instead of settling into steady state.
- dnsmasq_lease_max caps the number of leases dnsmasq will track per network, worth raising on large subnets.
- dnsmasq_enable_addr6_list enables address listing behavior needed for certain IPv6 configurations.
resync_interval and resync_throttle exist specifically to prevent agent overload on clouds with heavy port churn, according to the Neutron configuration reference. Getting these two settings wrong is one of the most common causes of a DHCP agent that looks alive but responds slowly.
Metadata behavior comes down to two flags: enable_isolated_metadata turns on the metadata proxy for networks without a router, and force_metadata forces the DHCP namespace to serve metadata even when a router is present, useful in edge cases where the L3 path is unreliable.
DNS handling and dnsmasq behavior
DHCP-assigned instances get their DNS servers through one of two paths, and the choice has real security implications.
The cleanest approach sets dns_nameservers directly on the subnet at creation time. Every instance on that subnet gets exactly the resolvers you specified, with no ambiguity and no dependency on host configuration.
Absent that, the agent falls back to dhcp_agent.ini settings:
- dnsmasq_dns_servers names specific upstream resolvers that dnsmasq forwards queries to, a predictable, centralized choice for the whole deployment.
- dnsmasq_local_resolv=True instead tells dnsmasq to forward using the network node’s own
/etc/resolv.conf. According to the Neutron Pike admin docs, this can leak internal DNS infrastructure to tenant instances if that host resolver points at internal-only name servers.
Within the isolated network, dnsmasq can also resolve instance hostnames locally before forwarding anything upstream, which is handy for east-west name resolution without standing up a separate internal DNS service.
For production, prefer explicit resolvers, either through subnet-level dns_nameservers or a defined dnsmasq_dns_servers list, over local_resolv. You want DNS behavior that does not change depending on what happens to be in a network node’s resolver configuration that week.
Pro Tip: Audit dnsmasq_local_resolv on every DHCP agent host during a security review. It is an easy setting to inherit from a default install and forget about.
High availability and scheduling for DHCP agents
DHCP HA in Neutron is controlled by a single value: dhcp_agents_per_network. Set it above 1, and the scheduler assigns that many DHCP agents to every tenant network, each running its own dnsmasq process and namespace, so a single node failure does not knock out DHCP entirely. This behavior is documented in the Neutron HA admin guide.
Managing this day to day comes down to a few commands:
openstack network agent listshows every registered agent and its alive status, your first stop when DHCP looks unhealthy.openstack network agent add network --dhcp <agent-id> <network>manually assigns an additional DHCP agent to a network.openstack network agent remove network --dhcp <agent-id> <network>pulls an agent off a network, useful before decommissioning a host.
Metadata redundancy is not symmetric between IP versions. The same HA documentation notes that IPv4 isolated metadata is redundant across multiple DHCP agents, but IPv6 isolated metadata is not, because of how the route to 169.254.169.254 gets injected. Plan around that gap rather than assuming IPv6 networks get the same failover behavior for free.
Choose dhcp_agents_per_network based on how many network nodes you actually have available, then test failover by manually removing an agent and confirming leases still renew.
IPv6 and DHCPv6 considerations
Dual-stack networking in Neutron depends on getting ipv6_ra_mode and ipv6_address_mode to agree. SLAAC lets instances self-assign addresses from router advertisements with no DHCP involvement at all. DHCPv6-stateful hands out full addresses through DHCP, closer to familiar IPv4 behavior. DHCPv6-stateless uses router advertisements for addressing but still queries DHCP for options like DNS servers.

Whichever mode you pick, the Neutron admin documentation is explicit that DHCPv6 support requires dnsmasq version 2.63 or later. Older packaged versions silently fail to deliver stateful or stateless DHCPv6 correctly, which makes this one of the first things worth checking on a new deployment.
Before enabling dual-stack subnets, confirm:
- Router advertisements are actually reaching instances, since SLAAC and stateless modes depend on them.
- Prefix delegation is configured correctly if you are handing out delegated prefixes rather than fixed subnets.
- Your dnsmasq package meets the 2.63+ requirement across every DHCP agent host, not just one.
- ipv6_ra_mode and ipv6_address_mode are set consistently on the subnet, since a mismatch produces addressing that half-works.
Deployment and scaling best practices
Where you run DHCP agents affects both performance and blast radius. Dedicated network nodes isolate DHCP load from compute workloads and simplify capacity planning, while collocating agents on compute nodes reduces hardware footprint at the cost of tighter coupling between compute and network failures.
- DVR deployments introduce host_dvr_for_dhcp, and setting it to False reduces router processing overhead but removes DNS service from the DHCP namespace, per the OVS HA and DVR admin guide, so you need subnet-level dns_nameservers as a substitute.
- OVN native DHCP is worth evaluating when namespace sprawl and per-network dnsmasq monitoring have become a real operational burden, since it reduces the number of moving parts you have to track.
- Before migrating, test metadata delivery and IPv6 behavior specifically, since those paths differ from the dnsmasq-based agent.
- Watch agent state (alive/dead), plus the count of networks, ports, and subnets each agent reports, as your baseline health indicators.
Pro Tip: If you’re evaluating OVN, run it in parallel on a non-production network first and compare DHCP lease timing before cutting anything over.
Operations and troubleshooting checklist
When DHCP breaks, work through this in order:
- Confirm the qdhcp namespace exists for the affected network, and that a dnsmasq process is actually running inside it.
- Check
openstack network agent listfor the DHCP agent’s alive status on the relevant host. - Pull details with
openstack network agent show <agent-id>and cross-reference/var/log/neutron/dhcp-agent.logalongsidejournalctl -u neutron-dhcp-agent. - After rebooting a node that hosts a DHCP agent, run neutron-ovs-cleanup before the agent starts. The Neutron admin docs warn that leftover OVS namespace and port state can block new qdhcp namespaces from binding cleanly, and note that Debian-based systems in particular may need this triggered manually rather than automatically.
On distros where the cleanup service is not wired into the boot sequence automatically, a missed neutron-ovs-cleanup run is one of the most common causes of DHCP failing silently after a routine reboot.
If agents look alive but slow, revisit resync_interval and resync_throttle. A debugging walkthrough for Neutron networking covers the broader diagnostic sequence if the DHCP path checks out clean but the network is still misbehaving.
Operational lessons from running DHCP at scale
Two things have saved me more troubleshooting time than anything else: fixing DNS topology explicitly instead of trusting dnsmasq_local_resolv, and never assuming neutron-ovs-cleanup runs on its own after a reboot. Write it into your boot automation if your distro does not handle it. If you are eyeing an OVN migration, budget real test time for metadata and IPv6 flows specifically. Those are the two places dnsmasq-based DHCP and OVN native DHCP diverge in ways that bite you during cutover, not before.
— James
How Devopsaitoolkit can help with OpenStack DHCP and Neutron operations
Working through dhcp_agent.ini tuning, HA scheduling, and an eventual OVN migration is a lot to carry alongside daily incident load. Devopsaitoolkit’s OpenStack / Kolla-Ansible Review ($450 one-off) checks your DHCP and Neutron configuration against production-tested patterns, flags HA gaps, and gives you a concrete remediation list.

- OpenStack / Kolla-Ansible Review is a focused audit of your Neutron agent configuration, dhcp_agents_per_network setup, and DNS handling.
- Managed OpenStack and Kolla-Ansible Services provide ongoing operational ownership for DHCP scheduling, agent health, and OVN migration. OpenStack Operations Toolkit is a reference set of runbook steps for operational checks as described above.
If you want a second set of eyes before your next OVN migration or HA change, book a review or explore managed OpenStack support.
Primary OpenStack docs and internal guides for reference
- Migrating Neutron to OVN and debugging Neutron floating IPs
FAQ
What does the Neutron DHCP agent actually do?
It runs dnsmasq inside a per-network qdhcp namespace to hand out DHCP leases, forward DNS queries, and serve the metadata proxy on isolated networks. It is the default mechanism for DHCP-based IP assignment unless your deployment runs OVN native DHCP instead.
How do I make Neutron DHCP highly available?
Set dhcp_agents_per_network above 1 in your Neutron configuration so the scheduler assigns multiple DHCP agents to each tenant network, as described in the DHCP HA admin guide. Note that IPv4 isolated metadata is redundant under this setup, but IPv6 isolated metadata is not.
Is dnsmasq_local_resolv safe to use in production?
It works, but it forwards DNS queries using the host’s own /etc/resolv.conf, which can leak internal DNS infrastructure to tenant instances according to the Neutron admin docs. Explicit dnsmasq_dns_servers or per-subnet dns_nameservers give you more predictable, auditable behavior.
Why do I need to run neutron-ovs-cleanup after a reboot?
Rebooting a node hosting a DHCP agent can leave stale OVS namespace and port state behind, which blocks new qdhcp namespaces or ports from binding correctly. Running neutron-ovs-cleanup before the agent restarts clears that stale state, and the Neutron admin documentation notes this sometimes needs to be triggered manually on Debian-based systems.
Does DHCPv6 have a minimum dnsmasq version requirement?
Yes, DHCPv6 stateful and stateless support requires dnsmasq version 2.63 or later, per the Neutron admin docs. Running an older dnsmasq package will cause IPv6 DHCP features to behave incorrectly or not work at all.
Recommended
- Debugging Neutron Networking in OpenStack
- Debugging Neutron Floating IPs and NAT in OpenStack
- Migrating Neutron to OVN Networking in OpenStack
- Four Traffic Planes for OpenStack Network Design With OVN
Get 500 Battle-Tested DevOps AI Prompts — Free
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.