Managed OpenStack and Kolla-Ansible Services
Keep your private cloud stable, observable, secure, and ready to scale with specialized OpenStack operations, proactive monitoring, incident response, and lifecycle management.
For production OpenStack environments, private-cloud platforms, and infrastructure engineering teams.
- Nova
- Neutron
- Cinder
- Glance
- Keystone
- Horizon
- Heat
- Placement
- Kolla-Ansible
- RabbitMQ
- MariaDB / Galera
- HAProxy
- Keepalived
- Memcached
- Open vSwitch
- Prometheus
- Alertmanager
- Grafana
- Ubuntu / RHEL
Where production OpenStack actually hurts
These are the failure modes that keep private clouds unstable — control-plane faults, messaging and database problems, capacity surprises, and monitoring that never covered the services that matter. Each one is in scope.
Control-plane instability
API services flap, workers wedge after a restart, and nobody is certain which controller is actually serving traffic.
Nova scheduling and compute issues
Instances fail to build with no valid host found while hypervisors still show free capacity, usually a Placement inventory or allocation-ratio problem.
Neutron networking and agent failures
L3, DHCP, metadata, or Open vSwitch agents drop out of the agent list and take tenant connectivity with them.
Cinder scheduler and storage-backend problems
Volume creation fails or hangs in creating because the scheduler has no eligible backend, or the driver lost its storage path.
RabbitMQ queue growth and RPC timeouts
Unacked messages climb, queues stop draining, and services log RPC timeouts until the cluster is restarted by hand.
MariaDB and Galera cluster issues
Nodes desync, the cluster loses primary component, or connection limits are exhausted under load.
HAProxy and Keepalived availability problems
Backends are marked down, or the VIP fails over unexpectedly and takes the API endpoints with it.
Certificate expiration
Internal or public API certificates expire and take down endpoints that were healthy an hour earlier.
Capacity exhaustion
Compute, storage, or controller headroom runs out without warning because nobody owns forecasting.
Insufficient monitoring
Node-level metrics exist, but OpenStack service health, queue depth, and agent state are invisible.
Noisy or ineffective alerts
Alerts fire constantly and get muted, so the one that mattered is lost in the noise.
Risky upgrades
Upgrades get deferred year after year because no one can validate the path or plan a rollback.
Configuration drift
Hand edits inside containers diverge from globals.yml, and the next reconfigure silently reverts them.
Limited internal OpenStack expertise
One or two people hold all the operational knowledge, and the environment stalls when they are unavailable.
Slow incident diagnosis
Every outage starts from scratch because there are no runbooks and no agreed diagnostic path.
What the managed service covers
Control plane, lifecycle, compute, networking, and storage — operated together, because that is how they fail.
OpenStack control-plane operations
Day-to-day operation of the services your cloud depends on, including the supporting infrastructure that most monitoring ignores.
- Nova
- Neutron
- Cinder
- Glance
- Keystone
- Horizon
- Heat
- Placement
- HAProxy
- Keepalived
- RabbitMQ
- MariaDB / Galera
- Memcached
Kolla-Ansible lifecycle management
Configuration, reconfiguration, and upgrade work driven through Kolla-Ansible rather than by hand-editing containers. Every change is planned, validated, and approved before it touches production.
- Deployment reviews
- Inventory and configuration review
- globals.yml review
- Reconfiguration planning
- Controlled service restarts
- Certificate management
- Container image lifecycle management
- Version compatibility assessment
- Upgrade planning
- Pre-upgrade validation
- Post-upgrade validation
- Rollback planning
Compute and capacity operations
Keeping hypervisors healthy and making sure the capacity story is known well before it becomes an incident.
- Hypervisor health
- Nova service health
- Placement inventory
- CPU and RAM allocation
- Oversubscription analysis
- Disk capacity
- Compute-host maintenance
- Evacuation planning
- Workload distribution
- Capacity forecasting
Networking operations
Agent health and tenant network paths, including the physical assumptions underneath them that cause the hardest failures.
- Neutron server health
- Open vSwitch agents
- L3 agents
- DHCP agents
- Metadata agents
- Router availability
- Provider networks
- VLAN-backed networks
- Bonding and MTU reviews
- HA router troubleshooting
- Network-path diagnosis
Storage operations
Volume services and the backends behind them, including the failure modes that only appear under real load.
- Cinder scheduler
- Cinder volume services
- Storage pools
- LVM backends
- Ceph integrations where applicable
- Capacity monitoring
- Volume creation failures
- Snapshot health
- Backup validation
- Storage latency and error monitoring
OpenStack Monitoring with Prometheus and Grafana
Monitoring is part of the managed service, not a separate product or an add-on you buy alongside it. The same engineer who operates the environment builds the exporters, dashboards, and alert rules, and uses them to run it.
- Prometheus
- Alertmanager
- Grafana
- Node Exporter
- Blackbox Exporter
- OpenStack exporters
- RabbitMQ metrics
- HAProxy metrics
- MariaDB / MySQL metrics
- Linux system metrics
- Container metrics
- Selected Kolla service health metrics
Infrastructure dashboards
- Controller health
- Compute-node health
- Storage capacity
- Network-agent status
- API availability and latency
- RabbitMQ queue depth
- Database cluster health
- HAProxy backend health
- CPU, memory, disk, and network utilization
- Certificate expiration
- OpenStack service availability
Alerting
- OpenStack service down
- Neutron agent unavailable
- Nova compute service unavailable
- Cinder service unavailable
- RabbitMQ queues above defined thresholds
- Database cluster degradation
- API latency or HTTP error increases
- Disk-space exhaustion
- Certificate expiration
- Compute-capacity thresholds
- Monitoring target unavailable
Alert thresholds are tuned to each customer's environment — defaults are a starting point, not a policy.
Reporting
- Monthly infrastructure health reports
- Capacity and utilization summaries
- Alert trends
- Incident summaries
- Recurring-risk identification
- Recommended remediation work
- Upgrade and lifecycle recommendations
A defined path from alert to root cause
Every incident follows the same sequence, so diagnosis starts from evidence instead of guesswork and the follow-up work is tracked rather than forgotten.
Detection
An alert fires from Prometheus or Alertmanager, or you report the problem directly.
Triage
Confirm the signal is real, identify the affected service, and rule out a monitoring fault.
Impact assessment
Establish what tenants and workloads are actually affected, and how badly.
Diagnosis
Work from metrics and logs to a specific cause rather than restarting services and hoping.
Mitigation
Restore service by the safest available route, with the change agreed before it is applied.
Recovery
Bring remaining components back into a known-good state.
Validation
Verify service health from the API, the agents, and the dashboards before closing.
Root-cause analysis
Establish why it happened, in writing, without guessing.
Corrective-action tracking
Track the follow-up work so the same incident does not recur unnoticed.
Incidents covered
- OpenStack API failures
- HTTP 500, 502, 503, and 504 errors
- Nova scheduling failures
- Failed instance builds
- Neutron L3 or DHCP agent outages
- Cinder scheduling failures
- RabbitMQ RPC timeouts
- Database replication issues
- HAProxy backend failures
- Controller capacity problems
- Compute-node outages
- Certificate-related outages
Response targets and coverage hours are defined in your service agreement.
OpenStack Infrastructure Health Assessment
A structured review of your environment that establishes current state, risk, and priority before any operational work begins.
What it covers
- Architecture review
- Controller topology
- Compute topology
- Network architecture
- Storage architecture
- Kolla-Ansible configuration
- OpenStack service health
- Monitoring coverage
- Alerting effectiveness
- Capacity risks
- Backup and recovery procedures
- Certificate lifecycle
- Upgrade readiness
- Security and access review
- Operational documentation
- Runbook maturity
- Major single points of failure
What you receive
- Executive summary
- Current-state findings
- Critical risks
- Prioritized remediation plan
- Monitoring recommendations
- Capacity recommendations
- Upgrade-readiness findings
- Suggested managed-service scope
Three levels of operational coverage
Each level builds on the one before it. The right starting point comes out of the infrastructure assessment, not a menu.
Foundation
For environments that need monitoring, recurring reviews, and scheduled operational support.
- Environment onboarding
- Prometheus monitoring review
- Grafana dashboards
- Alert configuration
- Monthly health review
- Capacity report
- Scheduled operational support
- Configuration recommendations
Operations
Most PopularFor production environments requiring ongoing administration and incident support.
Everything in Foundation, plus:
- OpenStack service operations
- Kolla-Ansible maintenance
- Incident triage
- Configuration changes
- Service recovery assistance
- Patch and certificate planning
- Quarterly architecture review
- Runbook maintenance
Mission Critical
For environments requiring higher-touch operational coverage and customized response commitments.
Everything in Operations, plus:
- Customized response targets
- Escalation procedures
- Expanded incident coverage
- Change-planning assistance
- Upgrade project support
- Capacity forecasting
- Disaster-recovery reviews
- Executive operational reporting
Final scope, coverage hours, response targets, and pricing are defined after the infrastructure assessment.
Who this service is for
A good fit
- Organizations operating an OpenStack private cloud
- Teams using Kolla-Ansible
- Infrastructure teams with limited OpenStack staffing
- Companies preparing for an upgrade
- Organizations experiencing repeated cloud incidents
- Teams lacking effective OpenStack monitoring
- Companies needing senior-level escalation support
- Organizations that need operational help but do not need a full internal OpenStack team
Not a fit
This is specialized private-cloud engineering, not a general IT service desk.
- Desktop support
- Microsoft 365 administration
- End-user help desk
- Printer or workstation support
- Generic web hosting
- Unmanaged application-development requests
- Environments where no authorized infrastructure access can be provided
Specialists, not a general managed-services vendor
Specialized OpenStack and Kolla-Ansible knowledge
Private-cloud operations is the core skill here, not a sideline next to a general managed-services catalogue.
Production infrastructure experience
The same engineer who writes the OpenStack troubleshooting guides on this site does the operational work.
Strong Linux and networking background
Most OpenStack faults are really Linux, networking, or messaging faults wearing an OpenStack error message.
Breadth across the whole stack
Compute, storage, networking, messaging, databases, and observability are diagnosed together, because that is how they fail.
Automation-first operational methods
Changes go through Kolla-Ansible and version control, so the environment stays reproducible instead of drifting.
Troubleshooting based on evidence and metrics
Diagnosis starts from metrics, logs, and service state, not from restarting components until symptoms move.
Clear runbooks and documentation
Operational knowledge is written down and handed over, so your team is not dependent on a single person.
Practical recommendations
Findings come with the specific change to make and the risk of making it, not generic best-practice language.
James Joyner IV
Sr. Systems Software Engineer · San Jose, CA
Engagements are delivered personally by James Joyner IV, a Sr. Systems Software Engineer who runs large-scale Linux systems and OpenStack with Kolla-Ansible in production, along with the Prometheus, VictoriaMetrics, and Grafana observability behind them. You are hiring the same engineer who writes the OpenStack troubleshooting guides on this site.
How onboarding works
Nothing is changed in production before scope, access, and approval are agreed in writing.
Discovery call
Understand the environment, the team, and the operational pain.
Infrastructure assessment
A structured review of architecture, health, monitoring, and risk.
Scope and responsibility definition
Agree in writing what is managed, by whom, and within what hours.
Access and security setup
Named accounts, key-based access, and an audited path into the environment.
Monitoring onboarding
Exporters, Prometheus targets, Alertmanager routing, and Grafana dashboards.
Baseline and risk review
Establish what normal looks like before changing anything.
Runbook creation
Document the diagnostic paths for the failures this environment actually has.
Operational handoff
Agree escalation, change approval, and reporting cadence with your team.
Recurring management and reporting
Ongoing operations, monthly reporting, and tracked corrective actions.
Who owns what
Managed operations only works when the boundaries are written down. This is the starting split, confirmed and adjusted in your service agreement.
Infrastructure access
- Your team
- Provide and revoke authorized access; own the network path in.
- DevOps AI ToolKit
- Use named accounts and key-based authentication; log activity.
- Shared
- Review access periodically.
Monitoring and alerting
- Your team
- Provide host and network access for exporters and scraping.
- DevOps AI ToolKit
- Configure Prometheus, Alertmanager, Grafana, and tune rules.
- Shared
- Agree thresholds and routing for this environment.
Production changes
- Your team
- Approve changes and own the maintenance window.
- DevOps AI ToolKit
- Plan, validate, execute, and document the change.
- Shared
- Agree rollback criteria before the window opens.
Incident response
- Your team
- Report user-visible impact and confirm restoration.
- DevOps AI ToolKit
- Triage, diagnose, mitigate, and produce root-cause analysis.
- Shared
- Agree severity and communication during the incident.
Capacity
- Your team
- Share growth plans and upcoming workload changes.
- DevOps AI ToolKit
- Track utilization and forecast exhaustion.
- Shared
- Plan procurement and expansion timing.
Upgrades
- Your team
- Approve scope, timing, and acceptance criteria.
- DevOps AI ToolKit
- Assess compatibility, validate, and plan rollback.
- Shared
- Define the test plan and the go/no-go decision.
Hardware and facilities
- Your team
- Own physical hardware, power, cooling, and vendor support.
- DevOps AI ToolKit
- Report hardware-level faults surfaced by monitoring.
- Shared
- Coordinate maintenance affecting compute or storage.
Backups and recovery
- Your team
- Own backup infrastructure and retention policy.
- DevOps AI ToolKit
- Validate backup health and review recovery procedures.
- Shared
- Test restores and document the recovery path.
Frequently asked questions
What versions of OpenStack do you support?
Supported versions are confirmed during discovery based on lifecycle status, deployment method, dependencies, and upgrade requirements. Older releases can usually still be operated and monitored, but an upgrade path is discussed where a version is past upstream maintenance.
Do you only support Kolla-Ansible deployments?
Kolla-Ansible is the primary specialty and where the deepest lifecycle support applies. Other deployment methods can be assessed and monitored, and the level of operational support is confirmed during discovery once the deployment tooling is known.
Can you take over an existing OpenStack environment?
Yes. That is the normal starting point. The infrastructure assessment establishes current state, risks, and monitoring gaps first, so operations begin from a documented baseline rather than assumptions.
Do you provide Prometheus and Grafana installation?
Yes. Where no monitoring exists, Prometheus, Alertmanager, Grafana, and the relevant exporters can be deployed and configured as part of onboarding. Where monitoring already exists, the work is usually coverage and alert-quality improvement instead.
Can you integrate with an existing monitoring platform?
Yes. Existing Prometheus, Grafana, or third-party platforms can be extended with OpenStack-specific targets, dashboards, and alert rules. Alertmanager can also route into an existing on-call or ticketing system.
Do you provide 24/7 support?
Coverage hours and response targets are customized and confirmed in the service agreement. They are not fixed in advance, and no round-the-clock commitment is implied until it is written into your agreement.
Can you perform OpenStack upgrades?
Yes, as a planned project rather than routine maintenance. Upgrades include version-compatibility assessment, pre-upgrade validation, a rollback plan, an agreed maintenance window, and post-upgrade validation. Nothing is upgraded without testing and your approval.
Do you manage Ceph?
The service can monitor Ceph and support its integration with OpenStack, including Cinder, Glance, and Nova storage paths. Deep Ceph cluster administration — rebalancing, CRUSH map changes, and capacity operations inside the Ceph cluster itself — must be explicitly included in scope.
Can you help with Nova, Neutron, and Cinder incidents?
Yes. Compute scheduling failures, agent outages, and volume-service problems are core incident categories. Diagnosis works from Placement inventory, agent state, queue depth, and service logs rather than blind restarts.
How is customer access secured?
Access uses named accounts and key-based authentication through a path you control, and you can revoke it at any time. Activity is logged, and access is scoped to what the agreed scope requires.
Do you make production changes without approval?
No. Production changes follow an agreed change process with a defined window and rollback criteria. Emergency actions during an active incident follow the escalation path agreed in your service agreement.
What happens during the infrastructure assessment?
The environment is reviewed across architecture, service health, Kolla-Ansible configuration, monitoring coverage, capacity, certificates, backup and recovery, upgrade readiness, and single points of failure. You receive an executive summary, current-state findings, critical risks, and a prioritized remediation plan.
Can you support multiple OpenStack regions?
Yes. Multi-region and multi-site environments can be supported. Region count affects monitoring topology, onboarding effort, and scope, so it is captured during discovery.
Do you provide capacity-planning reports?
Yes. Capacity and utilization summaries are part of the monthly reporting, covering compute allocation, oversubscription, and storage headroom, with forecasting available at the higher engagement levels.
How quickly can onboarding begin?
Onboarding timing depends on current availability, access provisioning on your side, and the outcome of the assessment. The discovery call is normally the fastest way to get a realistic date for your environment.
Request an Infrastructure Assessment
Tell us about your environment. The more detail you give, the more useful the first conversation is. Fields marked with an asterisk are required.
Ready to talk about your OpenStack environment?
Start with an infrastructure assessment, or book a discovery call to walk through your private cloud, the failures you keep hitting, and where operational support would help most.