← Back to Blog
AI Tools June 2026 6 min read

Troubleshooting Production Incidents with Claude

A practical workflow for using an AI assistant during an actual incident — what to paste, what to ask, and where human judgement still has to lead.

The shape of a good incident prompt

The quality of help you get is almost entirely a function of what you give it. A vague "the API is slow, help" produces vague guesses. A good incident prompt includes: the symptom, the timeline, what changed recently, and the actual evidence — logs, metrics, a diff — not a paraphrase of them.

Symptom: p99 latency on /checkout jumped from 180ms to 2.4s at 14:32 UTC
Recent changes: deploy at 14:28 UTC (diff attached), no infra changes
Evidence: [paste actual APM trace / log excerpt]
Already ruled out: not a DB connection pool exhaustion issue (checked pool metrics)
Question: what in this diff could plausibly cause this, and what should I check next?
✓ Including "already ruled out" is one of the highest-leverage additions — it stops the model from re-suggesting things you've already checked, and signals the kind of hypothesis you actually need.

Use it to generate hypotheses, not conclusions

The most reliable use during an incident is hypothesis generation: "given this trace and this diff, what are the three most likely causes, ranked by probability, and what would confirm or rule out each one." You stay the one deciding which hypothesis to chase and in what order — the model is good at generating a broader hypothesis space quickly than you might under incident pressure, not at being the final authority on root cause.

Parsing logs and traces faster

Pasting a large, noisy log excerpt and asking "summarise the error pattern and group by likely root cause" turns a wall of repeated stack traces into an actionable summary in seconds — genuinely useful when you're triaging during a page at 3am and your own pattern-matching is running slow.

Here are 400 lines of application logs from the last 10 minutes.
Group the errors by root cause, show frequency of each, and flag
anything that correlates with the 14:28 deploy timestamp.

Reading unfamiliar code under time pressure

Incidents often land in code paths you didn't write or haven't touched in months. "Walk me through what this function does and what could cause it to throw this specific error" is a faster way into unfamiliar code during an incident than reading it cold, especially across a service boundary you don't own day-to-day.

Drafting the fix — and the postmortem

Once you've identified the likely cause, asking for a minimal, reviewable fix (not a rewrite) keeps the actual incident resolution fast and auditable. Afterwards, the same conversation is a good starting point for a postmortem draft — timeline, root cause, contributing factors, and remediation items — though it still needs a human pass for accuracy and the details only you know (who was paged, what Slack thread has the full context).

Where to be careful

🚫 Don't paste secrets or customer data into any external tool during an incident — redact tokens, connection strings, and PII before pasting logs, even under time pressure.
🚫 Don't apply a suggested fix to production without the review you'd normally require. Incident pressure is exactly when skipping review causes the next incident.
🚫 Don't treat a plausible-sounding explanation as a confirmed root cause. Verify against actual metrics/logs before closing the incident.

Final thoughts

Used well, an AI assistant during an incident is a force multiplier for the parts that are mechanical — summarising logs, generating hypotheses, drafting the postmortem — so you can spend your own judgement on the parts that actually require it: deciding what's safe to ship, and confirming the real root cause against real evidence.