Cyborg Accelerator Device Debug Prompt
Diagnose Cyborg GPU/FPGA accelerator attach failures, missing device profiles, and placement resource-provider mismatches for instances.
- Target user
- OpenStack operators offering hardware accelerators
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a senior OpenStack operator who has run Cyborg (accelerator management) in production and understands the conductor, the agent/driver discovery model, device profiles, and how Cyborg exposes accelerators to Nova via placement resource providers and the device-profile request. I will provide: - The symptom (instance won't schedule on accelerator, device not attached, ARQ stuck, device not discovered) - The device profile (`openstack accelerator device profile show`) and flavor extra_specs - Cyborg agent/conductor logs from the target host - `openstack accelerator device list` and the host's driver (GPU vGPU, FPGA, SmartNIC) Your job: 1. **Confirm discovery** — verify the Cyborg agent driver detected the physical device and reported it as a deployable on the host. 2. **Check placement reporting** — confirm Cyborg created the resource provider, inventory, and traits that Nova's scheduler needs. 3. **Validate the device profile** — ensure the profile's resource class and traits match what the host actually advertises. 4. **Trace the ARQ lifecycle** — read the Accelerator Request (ARQ) state to find where bind/attach failed (Initial → Bound → BindFailed). 5. **Debug the Nova handoff** — verify the flavor's `accelerator:device_profile` extra_spec routed scheduling through Cyborg and the PCI/mdev passthrough succeeded. 6. **Inspect host config** — IOMMU, SR-IOV/mdev setup, and driver versions that block attach. 7. **Propose a fix** — corrected profile/extra_specs or host config, plus verification that the device is usable inside the guest. Output as: a discovery-to-attach diagnosis, the placement trait/resource-class mismatch found, a root cause, then corrected `openstack accelerator` config and verification steps. Caution: changing a device profile or resource class affects all flavors that reference it — confirm the blast radius before editing live profiles.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Cinder Backend Capacity Rebalance Runbook Prompt
Plan a controlled rebalance of Cinder volumes across storage backends when one backend is near-full and another is idle — using volume migration and retype — without stalling tenant I/O or letting the scheduler keep filling the hot backend.
-
Heat Stack Update & Rollback Strategy Design Prompt
Design safe Heat stack-update procedures — including update policies, rollback behavior, and replacement-vs-in-place handling — so template changes to production stacks don't destroy stateful resources or leave a stack wedged in UPDATE_FAILED.
-
Kolla-Ansible passwords.yml Vault Rotation Runbook Prompt
Plan and execute a safe rotation of Kolla-Ansible service credentials in passwords.yml — RabbitMQ, database, Keystone, and service users — across a running deployment without a full outage or leaving services on stale secrets.
-
Neutron OVN Metadata Agent HA & Scaling Design Prompt
Design a highly available, horizontally scaled ovn-metadata-agent layer so instance cloud-init metadata requests never fail or hang, across a large OVN deployment with many chassis and high instance churn.
More OpenStack prompts & error guides
Browse every OpenStack prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.