Skip to content
DevOps AI ToolKit
All guides
AI for Incident Response By James Joyner IV · · 12 min read

Safe AI Incident Triage: Redact Linux, Docker, and OpenStack Logs

Before you paste logs into an AI assistant, build a small, secret-free evidence bundle. A hands-on workflow for minimizing, redacting, and pseudonymizing Linux, Docker, and OpenStack diagnostics — with a downloadable template and a remove/replace/retain exercise.

  • #incident-response
  • #security-hardening
  • #ai
  • #openstack
  • #docker

An AI assistant is genuinely useful in the first ten minutes of an incident — if you feed it the right evidence. The mistake is pasting a raw journalctl dump or a full .env into a chat box under pressure. Once diagnostic text leaves your environment, you no longer fully control where it’s stored or for how long. So the safe move isn’t “don’t use AI” — it’s prepare a small, correlated, secret-free evidence bundle first.

This guide is the workflow for building that bundle across the three stacks most likely to be on fire: Linux hosts, Docker containers, and OpenStack control planes. You’ll work a synthetic incident from raw logs to a shareable bundle, using a downloadable template and a hands-on exercise. Nothing here sends your data anywhere; the whole process is manual and local by design.

Quick answer: Before using an AI assistant on an incident, (1) minimize — share only the failing services and a tight time window; (2) remove secrets by deleting their structure, not by masking values in place; (3) pseudonymize hosts, IPs, customers, and people with consistent placeholders so the diagnostic relationships survive; (4) keep timestamps, request IDs, status codes, and errors — that’s the evidence; (5) review by hand before sending. Pattern-matching tools help, but they do not guarantee removal of every secret.

Why a bundle, not a dump

Three reasons a curated bundle beats a raw paste:

  1. Data control. You can’t assume a third-party service discards what you send, and different providers and plans differ. Treat anything you paste as potentially retained, and minimize accordingly. (This is about your data-handling posture — don’t rely on inferred provider guarantees.)
  2. Signal. A 4,000-line dump buries the two lines that matter. A tight bundle gets you a better answer faster.
  3. Blast radius. Secrets in a paste don’t just risk exposure — a leaked token or connection string is an incident on top of your incident.

The rest of this article uses a synthetic scenario you can download and follow along with:

Prerequisites: none. Everything is read-only text work. The synthetic secrets in the sample are documentation examples or obvious dummies — there are no real credentials in any file.

The scenario

At 14:02 UTC, a shop API starts returning 500s. On-call grabs whatever’s handy: HAProxy edge logs, nginx access logs, docker logs for the API container, an environment dump “just in case,” and journalctl from an OpenStack controller showing a messaging timeout. It’s a realistic mix — and it contains things that must never be shared:

AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
DATABASE_URL=postgres://shopapi:Hunter2-NOT-REAL@10.20.3.14:5432/shop
INTERNAL_API_TOKEN=Bearer eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiJzdmMtY2hlY2tvdXQi...

…plus internal IPs, a real-looking hostname, a customer ID, and a customer’s email address scattered through the logs. Open sample-incident-raw.log to see the whole thing.

Step 1 — Minimize (the best redaction is omission)

Before redacting anything, throw most of it away. You almost never need a full dump.

  • Narrow the time window to minutes around the event, not hours.
  • Include only the services in the failure path. Here that’s the API and its database dependency; the edge logs add useful context (503s), the OpenStack controller line explains why the DB is unreachable (a messaging/partition problem).
  • Delete whole sections that are all-secret. The environment dump is 100% things that shouldn’t leave your box. Don’t mask it — remove it and leave a note: [env dump withheld: 6 secret variables removed].

Minimizing first means there’s simply less to get wrong in the next steps.

What’s worth collecting per stack

Minimizing is easier when you know where the signal usually is. A starting point for the three stacks in this scenario:

StackUsually worth collectingUsually safe to skip / high-risk
Linux hostthe failing unit’s journalctl -u <unit> around the event, dmesg OOM lines, df -h/df -i outputfull journalctl with no unit filter, env, anything under /etc with credentials
Dockerdocker logs --since/--until for the failing container, exit code + docker inspect state, the relevant Compose snippetdocker inspect in full (mounts, env), .env files, registry credentials
OpenStackthe service log line with the X-Openstack-Request-Id, the specific oslo.messaging / RabbitMQ timeout, agent livenessopenstack ... --debug output (leaks tokens/catalog), clouds.yaml, Fernet keys

Notice the pattern: the evidence is timestamps, IDs, and error strings; the risk is almost always in dumps, env, and config. Collect the first, and consciously exclude the second.

Step 2 — Redaction vs. pseudonymization

These are different tools, and mixing them up is how correlation gets destroyed or identity gets leaked.

  • Redaction removes information you never need back — a password, a private key, an API token. Gone.
  • Pseudonymization replaces an identifier with a consistent placeholder you can map back locally. 10.20.3.14 becomes DB_HOST_1 everywhere it appears. The AI still sees that three log lines refer to the same host — the relationship survives — but never learns the real address. You keep the placeholder→real legend on your machine and never send it.

Rule of thumb: redact secrets, pseudonymize identifiers. If you need the correlation, pseudonymize. If you never need it back, redact.

Step 3 — Remove secrets by removing their structure

Here’s the subtlety most redaction advice misses. Masking a value in place is not enough:

# still leaks that a DB credential exists, for this user, on this host:
DATABASE_URL=postgres://shopapi:********@10.20.3.14:5432/shop

That line still reveals the username, the host, and the fact that a credential exists there. Worse, a masked line still matches secret-detection patterns. I ran the raw sample through this site’s own secret detector (src/lib/docker-auditor/secrets.ts): it found 8 secrets. Then I ran the automatic redactor over it — which masks every value to ******** — and re-scanned. It still reported residual secrets, because postgres://user:********@host and TOKEN=******** still look like secrets structurally.

The fix: drop the structure. In the finished bundle, the DSN becomes:

dsn=postgres://DB_HOST_1:5432/shop?sslmode=require  (user + password removed)

No user:pass@ shape, no masked-value tell. Re-scanned, the redacted sample reports 0 secrets and passes the residual check. That’s the standard to hit — and the reason a human review always comes last.

Step 4 — Pseudonymize identifiers consistently

Walk the remaining evidence and replace identity-bearing values with stable placeholders:

Real (local only)Placeholder
10.20.3.9 (api host)IP_1
10.20.3.14 (db host)DB_HOST_1
prod-api-03.internal.example.netAPI_HOST_1
cust_8f21aa / jane.doe@acme-customer.exampleCUST_1 / CUST_1_EMAIL
10.20.3.0/24SUBNET_1

Consistency is the whole game: IP_1 must mean the same host every time, or you’ve destroyed the correlation that makes the evidence useful. Keep the mapping in the template’s local legend table — never in the bundle you send.

Step 5 — Keep the actual evidence

Don’t over-redact. These stay, untouched:

  • Timestamps — and don’t shift them; sequence is often the diagnosis.
  • Request/trace IDs like req_7c1f9a (pseudonymize only if they encode identity).
  • Status codes, durations, error strings — 503, upstream_response_time=30.001, MessagingTimeout: Timed out waiting for a reply — this is the signal.

Compare the raw and redacted samples side by side and you’ll see the diagnosis is still fully intact: the API times out talking to DB_HOST_1, the edge returns 503s, and the OpenStack controller shows a messaging timeout — pointing at a network/messaging partition, not an app bug. You removed the danger without removing the answer.

Exercise: remove, replace, or retain?

For each synthetic line, decide before you reveal the answer. (These <details> blocks work without JavaScript — just click to expand.)

Authorization: Bearer eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiJzdmMi...

Remove. A JWT is a live credential. Delete it — you never need the token value to diagnose an auth failure; the fact of a 401/403 and the timestamp are the evidence.

order failed for customer cust_8f21aa (jane.doe@acme-customer.example)

Replace (pseudonymize). You may need to see that the same customer recurs across lines, so replace with CUST_1 / CUST_1_EMAIL consistently. Never send the real email or account ID.

2026-09-16T14:02:10Z api ERROR db: connection failed: Connection timed out

Retain. This is pure evidence — no secret, no identity. Keep it exactly, including the timestamp.

retrying with pool=primary host=10.20.3.14

Replace the host, retain the rest. 10.20.3.14 → DB_HOST_1 (consistently); keep pool=primary and the retry behavior — it tells the assistant the app tried and failed to reach one specific host.

AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY

Remove the whole line. Don’t mask it to ******** — drop it and note that secret env vars were withheld. A masked assignment still says “this secret exists here.”

Know the limits

Be honest with yourself about what this process does and doesn’t guarantee:

  • No pattern matcher is complete. Regexes catch common shapes (AKIA…, eyJ…, scheme://user:pass@), but a bespoke token or an internal ID format can slip through. Automated redaction is an assist, not a certification — which is exactly why Step 5 is a manual read.
  • Redaction isn’t reversible; pseudonymization is — locally. If your placeholder legend leaks, pseudonymization offers no protection. Guard it like a secret.
  • Deciding to share at all is a choice. For the most sensitive systems, the right answer may be to keep everything in-house and use AI only on synthetic reproductions.

For a more general treatment of stripping secrets and PII from logs, see redacting secrets and PII from logs with AI; this guide’s focus is the incident evidence bundle specifically.

Put it to work

Once you have a clean bundle, use it:

  1. Start from the evidence-bundle template and the checklist.
  2. Practice on the sample in the Troubleshooting Workspace, which lets you record diagnostic steps and keep your notes local until you choose to act.
  3. When you want a structured triage plan, the AI Incident Response Assistant takes your redacted evidence and returns a risk-classified set of next steps — you decide exactly what to hand it.

The Incident Assistant is on the free plan (with a monthly allowance). If you run incidents regularly, DevOps AI ToolKit Pro lets you save your investigation and return to the evidence later — keeping your redacted bundles, notes, and triage plans organized across incidents instead of re-pasting into a chatbot at 3 a.m. Practice with the sample incident first, then keep your investigations organized in Pro. Your free result is always yours to keep; Pro just remembers it across incidents and devices.


The redaction results in this article are reproducible: run either sample file through the detector in src/lib/docker-auditor/secrets.ts. Verified 2026-09-16 — raw sample: 8 secrets detected; redacted sample: 0 secrets, no residual. All values in the samples are synthetic.

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week
Free download · 368-page PDF

Get 500 Battle-Tested DevOps AI Prompts — Free

500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.

  • 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
  • Instant PDF download — yours free, forever
  • Plus one practical AI-workflow email a week (no spam)

Single opt-in · unsubscribe anytime · no spam.