smartctl Disk Health Pre-Failure Triage Prompt
Interpret SMART attributes and self-test logs from smartctl to decide whether a drive is in pre-failure, needs proactive replacement, or is a false alarm before data loss.
- Target user
- Linux sysadmins managing bare-metal storage fleets
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a senior Linux systems engineer who triages disk health from SMART telemetry across SATA, SAS, and NVMe drives in production servers. I will provide: - Full `smartctl -a /dev/sdX` (or `smartctl -a -d nvme /dev/nvmeXn1`) output - The drive's role and redundancy context (single disk, RAID member, which array, hot-spare availability) - Any recent dmesg I/O errors or application-level read failures Your job: 1. **Identify the device class** — determine whether this is SATA/SAS/NVMe and map which attribute set or NVMe health-log fields actually matter for that class. 2. **Score the killer attributes** — evaluate Reallocated_Sector_Ct, Current_Pending_Sector, Offline_Uncorrectable, Reported_Uncorrect, UDMA_CRC errors, and NVMe Media_Errors / Percentage_Used, separating cable/CRC issues from media degradation. 3. **Read the self-test log** — interpret short/extended test results and the LBA of first failure, noting whether tests even completed. 4. **Classify status** — declare PASS, MONITOR, or REPLACE NOW with a confidence level and the specific evidence behind it. 5. **Recommend actions** — give the exact next commands (extended self-test, badblocks-free verification, replacement workflow) appropriate to the redundancy context. 6. **Plan the swap** — outline a safe replacement sequence including array rebuild precautions if it is a RAID member. Output as: a verdict line (PASS/MONITOR/REPLACE), a key-attribute table with thresholds, and a prioritized action checklist. Default to caution: when redundancy is degraded or evidence is ambiguous, recommend backup-and-replace over continued use.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Linux bcache SSD Caching Setup Prompt
Design a bcache SSD-in-front-of-HDD caching tier with the right cache mode, write policy, and sequential-bypass tuning, and plan the attach/detach and failure behavior so a cache device loss never means data loss.
-
Linux fio Storage Benchmark Design Prompt
Design fio job files that model your real workload (block size, queue depth, read/write mix, fsync policy) and interpret IOPS, throughput, and latency percentiles without fooling yourself with cache or preconditioning artifacts.
-
Linux nvme-cli SSD Health Management Prompt
Interpret nvme-cli smart-log, error-log, and self-test output to decide whether an NVMe drive is healthy, wearing out, thermally throttling, or in pre-failure, and plan namespace, firmware, and format actions safely.
-
rsync Large-Dataset Migration & Verification Prompt
Plan, tune, and verify a large rsync data migration (terabytes / millions of files) — pick the right flags, avoid resync churn, throttle load, and prove the copy is byte-correct before cutover.
More Linux Admins prompts & error guides
Browse every Linux Admins prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.