Skip to content
DevOps AI ToolKit
Newsletter
Managed services

Managed OpenStack and Kolla-Ansible Services

Keep your private cloud stable, observable, secure, and ready to scale with specialized OpenStack operations, proactive monitoring, incident response, and lifecycle management.

For production OpenStack environments, private-cloud platforms, and infrastructure engineering teams.

  • Nova
  • Neutron
  • Cinder
  • Glance
  • Keystone
  • Horizon
  • Heat
  • Placement
  • Kolla-Ansible
  • RabbitMQ
  • MariaDB / Galera
  • HAProxy
  • Keepalived
  • Memcached
  • Open vSwitch
  • Prometheus
  • Alertmanager
  • Grafana
  • Ubuntu / RHEL
The problems this solves

Where production OpenStack actually hurts

These are the failure modes that keep private clouds unstable — control-plane faults, messaging and database problems, capacity surprises, and monitoring that never covered the services that matter. Each one is in scope.

Control-plane instability

API services flap, workers wedge after a restart, and nobody is certain which controller is actually serving traffic.

Nova scheduling and compute issues

Instances fail to build with no valid host found while hypervisors still show free capacity, usually a Placement inventory or allocation-ratio problem.

Neutron networking and agent failures

L3, DHCP, metadata, or Open vSwitch agents drop out of the agent list and take tenant connectivity with them.

Cinder scheduler and storage-backend problems

Volume creation fails or hangs in creating because the scheduler has no eligible backend, or the driver lost its storage path.

RabbitMQ queue growth and RPC timeouts

Unacked messages climb, queues stop draining, and services log RPC timeouts until the cluster is restarted by hand.

MariaDB and Galera cluster issues

Nodes desync, the cluster loses primary component, or connection limits are exhausted under load.

HAProxy and Keepalived availability problems

Backends are marked down, or the VIP fails over unexpectedly and takes the API endpoints with it.

Certificate expiration

Internal or public API certificates expire and take down endpoints that were healthy an hour earlier.

Capacity exhaustion

Compute, storage, or controller headroom runs out without warning because nobody owns forecasting.

Insufficient monitoring

Node-level metrics exist, but OpenStack service health, queue depth, and agent state are invisible.

Noisy or ineffective alerts

Alerts fire constantly and get muted, so the one that mattered is lost in the noise.

Risky upgrades

Upgrades get deferred year after year because no one can validate the path or plan a rollback.

Configuration drift

Hand edits inside containers diverge from globals.yml, and the next reconfigure silently reverts them.

Limited internal OpenStack expertise

One or two people hold all the operational knowledge, and the environment stalls when they are unavailable.

Slow incident diagnosis

Every outage starts from scratch because there are no runbooks and no agreed diagnostic path.

Services included

What the managed service covers

Control plane, lifecycle, compute, networking, and storage — operated together, because that is how they fail.

OpenStack control-plane operations

Day-to-day operation of the services your cloud depends on, including the supporting infrastructure that most monitoring ignores.

  • Nova
  • Neutron
  • Cinder
  • Glance
  • Keystone
  • Horizon
  • Heat
  • Placement
  • HAProxy
  • Keepalived
  • RabbitMQ
  • MariaDB / Galera
  • Memcached

Kolla-Ansible lifecycle management

Configuration, reconfiguration, and upgrade work driven through Kolla-Ansible rather than by hand-editing containers. Every change is planned, validated, and approved before it touches production.

  • Deployment reviews
  • Inventory and configuration review
  • globals.yml review
  • Reconfiguration planning
  • Controlled service restarts
  • Certificate management
  • Container image lifecycle management
  • Version compatibility assessment
  • Upgrade planning
  • Pre-upgrade validation
  • Post-upgrade validation
  • Rollback planning

Compute and capacity operations

Keeping hypervisors healthy and making sure the capacity story is known well before it becomes an incident.

  • Hypervisor health
  • Nova service health
  • Placement inventory
  • CPU and RAM allocation
  • Oversubscription analysis
  • Disk capacity
  • Compute-host maintenance
  • Evacuation planning
  • Workload distribution
  • Capacity forecasting

Networking operations

Agent health and tenant network paths, including the physical assumptions underneath them that cause the hardest failures.

  • Neutron server health
  • Open vSwitch agents
  • L3 agents
  • DHCP agents
  • Metadata agents
  • Router availability
  • Provider networks
  • VLAN-backed networks
  • Bonding and MTU reviews
  • HA router troubleshooting
  • Network-path diagnosis

Storage operations

Volume services and the backends behind them, including the failure modes that only appear under real load.

  • Cinder scheduler
  • Cinder volume services
  • Storage pools
  • LVM backends
  • Ceph integrations where applicable
  • Capacity monitoring
  • Volume creation failures
  • Snapshot health
  • Backup validation
  • Storage latency and error monitoring
Observability

OpenStack Monitoring with Prometheus and Grafana

Monitoring is part of the managed service, not a separate product or an add-on you buy alongside it. The same engineer who operates the environment builds the exporters, dashboards, and alert rules, and uses them to run it.

  • Prometheus
  • Alertmanager
  • Grafana
  • Node Exporter
  • Blackbox Exporter
  • OpenStack exporters
  • RabbitMQ metrics
  • HAProxy metrics
  • MariaDB / MySQL metrics
  • Linux system metrics
  • Container metrics
  • Selected Kolla service health metrics

Infrastructure dashboards

  • Controller health
  • Compute-node health
  • Storage capacity
  • Network-agent status
  • API availability and latency
  • RabbitMQ queue depth
  • Database cluster health
  • HAProxy backend health
  • CPU, memory, disk, and network utilization
  • Certificate expiration
  • OpenStack service availability

Alerting

  • OpenStack service down
  • Neutron agent unavailable
  • Nova compute service unavailable
  • Cinder service unavailable
  • RabbitMQ queues above defined thresholds
  • Database cluster degradation
  • API latency or HTTP error increases
  • Disk-space exhaustion
  • Certificate expiration
  • Compute-capacity thresholds
  • Monitoring target unavailable

Alert thresholds are tuned to each customer's environment — defaults are a starting point, not a policy.

Reporting

  • Monthly infrastructure health reports
  • Capacity and utilization summaries
  • Alert trends
  • Incident summaries
  • Recurring-risk identification
  • Recommended remediation work
  • Upgrade and lifecycle recommendations
Incident management

A defined path from alert to root cause

Every incident follows the same sequence, so diagnosis starts from evidence instead of guesswork and the follow-up work is tracked rather than forgotten.

1

Detection

An alert fires from Prometheus or Alertmanager, or you report the problem directly.

2

Triage

Confirm the signal is real, identify the affected service, and rule out a monitoring fault.

3

Impact assessment

Establish what tenants and workloads are actually affected, and how badly.

4

Diagnosis

Work from metrics and logs to a specific cause rather than restarting services and hoping.

5

Mitigation

Restore service by the safest available route, with the change agreed before it is applied.

6

Recovery

Bring remaining components back into a known-good state.

7

Validation

Verify service health from the API, the agents, and the dashboards before closing.

8

Root-cause analysis

Establish why it happened, in writing, without guessing.

9

Corrective-action tracking

Track the follow-up work so the same incident does not recur unnoticed.

Incidents covered

  • OpenStack API failures
  • HTTP 500, 502, 503, and 504 errors
  • Nova scheduling failures
  • Failed instance builds
  • Neutron L3 or DHCP agent outages
  • Cinder scheduling failures
  • RabbitMQ RPC timeouts
  • Database replication issues
  • HAProxy backend failures
  • Controller capacity problems
  • Compute-node outages
  • Certificate-related outages

Response targets and coverage hours are defined in your service agreement.

Start here

OpenStack Infrastructure Health Assessment

A structured review of your environment that establishes current state, risk, and priority before any operational work begins.

What it covers

  • Architecture review
  • Controller topology
  • Compute topology
  • Network architecture
  • Storage architecture
  • Kolla-Ansible configuration
  • OpenStack service health
  • Monitoring coverage
  • Alerting effectiveness
  • Capacity risks
  • Backup and recovery procedures
  • Certificate lifecycle
  • Upgrade readiness
  • Security and access review
  • Operational documentation
  • Runbook maturity
  • Major single points of failure

What you receive

  • Executive summary
  • Current-state findings
  • Critical risks
  • Prioritized remediation plan
  • Monitoring recommendations
  • Capacity recommendations
  • Upgrade-readiness findings
  • Suggested managed-service scope
Engagement options

Three levels of operational coverage

Each level builds on the one before it. The right starting point comes out of the infrastructure assessment, not a menu.

Foundation

For environments that need monitoring, recurring reviews, and scheduled operational support.

  • Environment onboarding
  • Prometheus monitoring review
  • Grafana dashboards
  • Alert configuration
  • Monthly health review
  • Capacity report
  • Scheduled operational support
  • Configuration recommendations

Operations

Most Popular

For production environments requiring ongoing administration and incident support.

Everything in Foundation, plus:

  • OpenStack service operations
  • Kolla-Ansible maintenance
  • Incident triage
  • Configuration changes
  • Service recovery assistance
  • Patch and certificate planning
  • Quarterly architecture review
  • Runbook maintenance

Mission Critical

For environments requiring higher-touch operational coverage and customized response commitments.

Everything in Operations, plus:

  • Customized response targets
  • Escalation procedures
  • Expanded incident coverage
  • Change-planning assistance
  • Upgrade project support
  • Capacity forecasting
  • Disaster-recovery reviews
  • Executive operational reporting

Final scope, coverage hours, response targets, and pricing are defined after the infrastructure assessment.

Fit

Who this service is for

A good fit

  • Organizations operating an OpenStack private cloud
  • Teams using Kolla-Ansible
  • Infrastructure teams with limited OpenStack staffing
  • Companies preparing for an upgrade
  • Organizations experiencing repeated cloud incidents
  • Teams lacking effective OpenStack monitoring
  • Companies needing senior-level escalation support
  • Organizations that need operational help but do not need a full internal OpenStack team

Not a fit

This is specialized private-cloud engineering, not a general IT service desk.

  • Desktop support
  • Microsoft 365 administration
  • End-user help desk
  • Printer or workstation support
  • Generic web hosting
  • Unmanaged application-development requests
  • Environments where no authorized infrastructure access can be provided
Why DevOps AI ToolKit

Specialists, not a general managed-services vendor

Specialized OpenStack and Kolla-Ansible knowledge

Private-cloud operations is the core skill here, not a sideline next to a general managed-services catalogue.

Production infrastructure experience

The same engineer who writes the OpenStack troubleshooting guides on this site does the operational work.

Strong Linux and networking background

Most OpenStack faults are really Linux, networking, or messaging faults wearing an OpenStack error message.

Breadth across the whole stack

Compute, storage, networking, messaging, databases, and observability are diagnosed together, because that is how they fail.

Automation-first operational methods

Changes go through Kolla-Ansible and version control, so the environment stays reproducible instead of drifting.

Troubleshooting based on evidence and metrics

Diagnosis starts from metrics, logs, and service state, not from restarting components until symptoms move.

Clear runbooks and documentation

Operational knowledge is written down and handed over, so your team is not dependent on a single person.

Practical recommendations

Findings come with the specific change to make and the risk of making it, not generic best-practice language.

Who you'll be working with

James Joyner IV

Sr. Systems Software Engineer · San Jose, CA

Engagements are delivered personally by James Joyner IV, a Sr. Systems Software Engineer who runs large-scale Linux systems and OpenStack with Kolla-Ansible in production, along with the Prometheus, VictoriaMetrics, and Grafana observability behind them. You are hiring the same engineer who writes the OpenStack troubleshooting guides on this site.

Delivery process

How onboarding works

Nothing is changed in production before scope, access, and approval are agreed in writing.

1

Discovery call

Understand the environment, the team, and the operational pain.

2

Infrastructure assessment

A structured review of architecture, health, monitoring, and risk.

3

Scope and responsibility definition

Agree in writing what is managed, by whom, and within what hours.

4

Access and security setup

Named accounts, key-based access, and an audited path into the environment.

5

Monitoring onboarding

Exporters, Prometheus targets, Alertmanager routing, and Grafana dashboards.

6

Baseline and risk review

Establish what normal looks like before changing anything.

7

Runbook creation

Document the diagnostic paths for the failures this environment actually has.

8

Operational handoff

Agree escalation, change approval, and reporting cadence with your team.

9

Recurring management and reporting

Ongoing operations, monthly reporting, and tracked corrective actions.

Shared responsibility

Who owns what

Managed operations only works when the boundaries are written down. This is the starting split, confirmed and adjusted in your service agreement.

Infrastructure access

Your team
Provide and revoke authorized access; own the network path in.
DevOps AI ToolKit
Use named accounts and key-based authentication; log activity.
Shared
Review access periodically.

Monitoring and alerting

Your team
Provide host and network access for exporters and scraping.
DevOps AI ToolKit
Configure Prometheus, Alertmanager, Grafana, and tune rules.
Shared
Agree thresholds and routing for this environment.

Production changes

Your team
Approve changes and own the maintenance window.
DevOps AI ToolKit
Plan, validate, execute, and document the change.
Shared
Agree rollback criteria before the window opens.

Incident response

Your team
Report user-visible impact and confirm restoration.
DevOps AI ToolKit
Triage, diagnose, mitigate, and produce root-cause analysis.
Shared
Agree severity and communication during the incident.

Capacity

Your team
Share growth plans and upcoming workload changes.
DevOps AI ToolKit
Track utilization and forecast exhaustion.
Shared
Plan procurement and expansion timing.

Upgrades

Your team
Approve scope, timing, and acceptance criteria.
DevOps AI ToolKit
Assess compatibility, validate, and plan rollback.
Shared
Define the test plan and the go/no-go decision.

Hardware and facilities

Your team
Own physical hardware, power, cooling, and vendor support.
DevOps AI ToolKit
Report hardware-level faults surfaced by monitoring.
Shared
Coordinate maintenance affecting compute or storage.

Backups and recovery

Your team
Own backup infrastructure and retention policy.
DevOps AI ToolKit
Validate backup health and review recovery procedures.
Shared
Test restores and document the recovery path.
Questions

Frequently asked questions

What versions of OpenStack do you support?

Supported versions are confirmed during discovery based on lifecycle status, deployment method, dependencies, and upgrade requirements. Older releases can usually still be operated and monitored, but an upgrade path is discussed where a version is past upstream maintenance.

Do you only support Kolla-Ansible deployments?

Kolla-Ansible is the primary specialty and where the deepest lifecycle support applies. Other deployment methods can be assessed and monitored, and the level of operational support is confirmed during discovery once the deployment tooling is known.

Can you take over an existing OpenStack environment?

Yes. That is the normal starting point. The infrastructure assessment establishes current state, risks, and monitoring gaps first, so operations begin from a documented baseline rather than assumptions.

Do you provide Prometheus and Grafana installation?

Yes. Where no monitoring exists, Prometheus, Alertmanager, Grafana, and the relevant exporters can be deployed and configured as part of onboarding. Where monitoring already exists, the work is usually coverage and alert-quality improvement instead.

Can you integrate with an existing monitoring platform?

Yes. Existing Prometheus, Grafana, or third-party platforms can be extended with OpenStack-specific targets, dashboards, and alert rules. Alertmanager can also route into an existing on-call or ticketing system.

Do you provide 24/7 support?

Coverage hours and response targets are customized and confirmed in the service agreement. They are not fixed in advance, and no round-the-clock commitment is implied until it is written into your agreement.

Can you perform OpenStack upgrades?

Yes, as a planned project rather than routine maintenance. Upgrades include version-compatibility assessment, pre-upgrade validation, a rollback plan, an agreed maintenance window, and post-upgrade validation. Nothing is upgraded without testing and your approval.

Do you manage Ceph?

The service can monitor Ceph and support its integration with OpenStack, including Cinder, Glance, and Nova storage paths. Deep Ceph cluster administration — rebalancing, CRUSH map changes, and capacity operations inside the Ceph cluster itself — must be explicitly included in scope.

Can you help with Nova, Neutron, and Cinder incidents?

Yes. Compute scheduling failures, agent outages, and volume-service problems are core incident categories. Diagnosis works from Placement inventory, agent state, queue depth, and service logs rather than blind restarts.

How is customer access secured?

Access uses named accounts and key-based authentication through a path you control, and you can revoke it at any time. Activity is logged, and access is scoped to what the agreed scope requires.

Do you make production changes without approval?

No. Production changes follow an agreed change process with a defined window and rollback criteria. Emergency actions during an active incident follow the escalation path agreed in your service agreement.

What happens during the infrastructure assessment?

The environment is reviewed across architecture, service health, Kolla-Ansible configuration, monitoring coverage, capacity, certificates, backup and recovery, upgrade readiness, and single points of failure. You receive an executive summary, current-state findings, critical risks, and a prioritized remediation plan.

Can you support multiple OpenStack regions?

Yes. Multi-region and multi-site environments can be supported. Region count affects monitoring topology, onboarding effort, and scope, so it is captured during discovery.

Do you provide capacity-planning reports?

Yes. Capacity and utilization summaries are part of the monthly reporting, covering compute allocation, oversubscription, and storage headroom, with forecasting available at the higher engagement levels.

How quickly can onboarding begin?

Onboarding timing depends on current availability, access provisioning on your side, and the outcome of the assessment. The discovery call is normally the fastest way to get a realistic date for your environment.

Infrastructure assessment

Request an Infrastructure Assessment

Tell us about your environment. The more detail you give, the more useful the first conversation is. Fields marked with an asterisk are required.

Request an OpenStack infrastructure health assessment. All information is used only to scope the engagement.

Coverage hours and response targets are confirmed in the service agreement after the assessment.

Schedule a Discovery Call

Your details are used only to scope this engagement — no newsletter signup, no third-party sharing. Prefer email? Write to consulting@devopsaitoolkit.com.

Ready to talk about your OpenStack environment?

Start with an infrastructure assessment, or book a discovery call to walk through your private cloud, the failures you keep hitting, and where operational support would help most.