- Home
- Skills
- DevOps & Cloud
- Incident Postmortem with Evidence Trail
Works with the AI tools you already use
Incident Postmortem with Evidence Trail
A blameless draft postmortem where every timeline entry cites its log, alert, or message, and every action item traces to a cause.
$5
Incident Postmortem with Evidence Trail
Example session with this skill installed
Write the postmortem for yesterday's checkout outage. Here's everything I have:
--- Alerts (monitoring tool, times in UTC) ---
2026-09-30 17:42 WARN checkout-api p95 latency > 2s (threshold 800ms)
2026-09-30 17:46 CRIT checkout-api error rate 34% (threshold 5%)
2026-09-30 17:46 page sent to on-call (payments rotation)
2026-09-30 18:31 RESOLVED checkout-api error rate < 1%
--- Deploy log (CI, times in UTC) ---
2026-09-30 17:38 deploy checkout-api v4.18.0 (author: dev on payments team) - "raise DB pool idle timeout, add retry on card-auth"
2026-09-30 18:12 rollback checkout-api to v4.17.3 (manual)
2026-09-30 18:20 config change: payments-db max_connections 200 -> 400
--- #inc-checkout Slack export (times shown in Pacific) ---
10:47 Maya (on-call): ack, looking
10:49 Maya: checkout errors are "connection pool exhausted" on payments-db
10:55 Raj: card-auth provider status page is green
11:02 Maya: retries are multiplying load, each failed auth retries 3x
11:05 Dan (eng manager): who shipped 4.18 on a Tuesday afternoon without a canary? this is on Priya
11:09 Maya: proposing rollback to 4.17.3
11:14 Raj: rolling back now
11:21 Maya: errors dropping but not gone, db still at max connections
11:25 Raj: bumped max_connections to 400
11:31 Maya: error rate under 1%, monitoring
11:58 Maya: calling it resolved. ~3,100 checkout attempts failed per the dashboard. support got 41 tickets.
--- My notes ---
- Canary stage exists in the pipeline but was skipped for this deploy (flag set by hand). Don't know why.
- Retry logic had no backoff or cap.
- No alert on DB connection saturation, only on API errors.
- Revenue impact: finance hasn't given us a number.
- Action items so far: add backoff to retries (Raj?), make canary mandatory, add DB connection alert. No dates yet.
- Read your context and instructions
- Compiled the incident postmortem
- Generated the document
I've written the draft postmortem: postmortem-2026-09-30-checkout-api.md. It's marked as a draft pending review by the incident owner, and nothing has been posted or sent.
Timeline
Timezone: everything is in UTC. Your Slack times were Pacific Daylight Time on Sept 30 (UTC−7). Two Slack messages land exactly where the alerts say they should, which confirms that offset.
Sources: each timeline row is tagged with the line it came from. ALR is the alerts, CI the deploy log, SLK the Slack export, NOTE your notes.
Response times: about 4 min from deploy to first alert, 1 min to acknowledge, 26–28 min to rollback, 45 min until errors were under 1%, 72 min until it was declared resolved.
When customers were affected: at least 49 min. The true start is unknown; it falls somewhere between the 17:38 deploy and the 17:42 alert.
Places the evidence disagrees (I recorded both times rather than picking one):
- Rollback: the deploy log says 18:12 UTC; Raj's "rolling back now" in Slack is 18:14 UTC.
max_connectionschange: the deploy log says 18:20 UTC; Slack says 18:25 UTC.- Date: you said "yesterday", which would be Oct 1. Every log says Sept 30, so the draft uses Sept 30.
- Day of week: Dan said "Tuesday", but Sept 30, 2026 was a Wednesday.
Keeping it blameless
- I left out Dan's 11:05 "this is on Priya." I kept the question he raised (why it shipped without a canary) and treated it as a process issue: the pipeline lets someone skip the canary by setting a flag by hand.
- The draft names people by role and never names anyone as a cause. A check for "failed to", "should have", "careless", "human error", "fault" and Priya's name came back clean.
Contributing factors
- Retries: card-auth retries had no backoff or cap, so they multiplied load on the database.
- Canary skipped: the canary stage was turned off for this deploy.
- No database alert: nothing alerts when payments-db runs out of connections; only the API errors alerted.
Idle timeout (not confirmed): raising the pool idle timeout probably made connections stay held longer. It's labelled as inferred because no connection-count data supports it yet.
Rollback didn't fully clear it: the database stayed at its connection limit after the rollback, and nothing in the evidence explains why.
Action items
- Your three are in: retry backoff (Raj, marked unconfirmed because your note had a "?"), mandatory canary, and the database connection alert.
- I added two proposed ones, labelled so you can drop them: measure the idle-timeout effect before that change ships again, and find out why the database stayed full after the rollback.
- The P1/P2 priorities are my suggestions; your notes didn't give any. Owners other than Raj are unassigned and there are no due dates, as in your notes.
Gaps. The doc ends with 11 open questions for review. The main ones are:
- Severity, incident owner, and whether anyone acted as incident commander were never stated.
- Why the canary was skipped is unknown; the root-cause analysis stops there.
- Whether
max_connectionsstays at 400. I made no recommendation; that's for normal change review. - The revenue figure is still pending from finance, and there's no unique-customer count, SLA impact, or status-page record.
| Deliverable | Last modified | Includes |
|---|---|---|
postmortem-2026-09-30-checkout-api.md (in your working folder) | 2026-10-02 09:40 | Full draft; checked for blame wording, every timeline row has a source, and every time either appears in your inputs or is a Pacific-to-UTC conversion of one |
No other copies exist.
The file it wrote: postmortem-2026-09-30-checkout-api.md
DRAFT - PENDING INCIDENT OWNER REVIEW
Postmortem: checkout-api errors after v4.18.0 deploy (payments-db connection exhaustion)
| Field | Value |
|---|---|
| Incident date | 2026-09-30 (from alert and deploy logs; see Open Questions #1) |
| Type | Outage (partial: peak error rate 34%) |
| Severity | NOT PROVIDED |
| Services affected | checkout-api, payments-db |
| Incident owner / reviewer | NOT PROVIDED |
| Status | Resolved, declared 18:58 UTC |
| Draft prepared | 2026-10-02 |
Summary
At 17:38 UTC on 2026-09-30, checkout-api v4.18.0 was deployed. The release raised the DB pool idle timeout and added retries on card-auth calls. Four minutes later, at 17:42, a p95 latency alert fired. By 17:46 the error rate was 34% and the payments on-call rotation was paged. The on-call engineer found that checkout errors were "connection pool exhausted" on payments-db. Failed card-auth calls were each retried 3 times with no backoff or cap, which multiplied the load on the database. Responders rolled back to v4.17.3 (18:12–18:14 UTC; the two sources disagree) and raised payments-db max_connections from 200 to 400 (18:20–18:25 UTC; the two sources disagree). The error rate fell below 1% at 18:31 UTC, and the incident was declared resolved at 18:58 UTC. About 3,100 checkout attempts failed and support received 41 tickets. Revenue impact is still pending from finance.
Contributing factors (summary): retries with no backoff or cap amplified load on a shared DB connection limit (T2). The canary stage was skipped for this deploy through a flag set by hand (P1). There was no alert on DB connection saturation (D1). The raised pool idle timeout is a likely but
unconfirmed contributor (T1, INFERRED).
Timeline
Reference timezone: UTC. The Slack export is in Pacific time. On 2026-09-30 that is PDT (UTC−7), checked with Intl.DateTimeFormat for America/Los_Angeles. The offset is supported by two independent matches: the on-call ack at 10:47 PDT = 17:47 UTC, one minute after the 17:46 page, and "error rate under 1%" at 11:31 PDT = 18:31 UTC, which matches the RESOLVED alert. The Slack export has no date, so its date is INFERRED as 2026-09-30 from these matches.
Source key (the evidence as supplied, numbered in the order pasted):
- ALR-1..4: monitoring alerts, lines 1–4 (UTC)
- CI-1..3: deploy log, lines 1–3 (UTC)
- SLK-1..11: #inc-checkout Slack export, messages 1–11 (Pacific)
- NOTE-1..5: incident lead's notes, bullets 1–5
| Time (UTC) | Original | Event | Actor (role) | Source |
|---|---|---|---|---|
| 17:38 | 17:38 UTC | checkout-api v4.18.0 deployed: "raise DB pool idle timeout, add retry on card-auth" | CI pipeline; change authored by a payments-team engineer | CI-1 |
| UNKNOWN | — | Canary stage skipped for this deploy by a skip flag set by hand | Unknown | NOTE-1 |
| 17:42 | 17:42 UTC | WARN: checkout-api p95 latency > 2s (threshold 800ms). First evidence of impact. | Monitoring | ALR-1 |
| 17:46 | 17:46 UTC | CRIT: checkout-api error rate 34% (threshold 5%) | Monitoring | ALR-2 |
| 17:46 | 17:46 UTC | Page sent to on-call (payments rotation) | Paging | ALR-3 |
| 17:47 | 10:47 PDT | Page acknowledged, investigation started | On-call engineer, payments (Maya) | SLK-1 |
| 17:49 | 10:49 PDT | Diagnosis: checkout errors are "connection pool exhausted" on payments-db | On-call engineer (Maya) | SLK-2 |
| 17:55 | 10:55 PDT | Card-auth provider status page checked: green (external cause less likely) | Responder (Raj; role not stated) | SLK-3 |
| 18:02 | 11:02 PDT | Diagnosis: retries are multiplying load; each failed auth retries 3x | On-call engineer (Maya) | SLK-4 |
| 18:05 | 11:05 PDT | Engineering manager asks in channel why 4.18 shipped on a weekday afternoon without a canary. (The message also attributed the deploy to a named individual. That attribution is left out of this blameless review. The process question is handled under P1.) | Engineering manager (Dan) | SLK-5 |
| 18:09 | 11:09 PDT | Decision: rollback to v4.17.3 proposed | On-call engineer (Maya) | SLK-6 |
| 18:12 / 18:14 | 18:12 UTC / 11:14 PDT | Mitigation 1: manual rollback of checkout-api to v4.17.3 (CONFLICT, see below) | Responder (Raj) | CI-2 / SLK-7 |
| 18:21 | 11:21 PDT | Errors dropping but not gone; payments-db still at max connections | On-call engineer (Maya) | SLK-8 |
| 18:20 / 18:25 | 18:20 UTC / 11:25 PDT | Mitigation 2: payments-db max_connections 200 → 400 (CONFLICT, see below) | Responder (Raj) | CI-3 / SLK-9 |
| 18:31 | 18:31 UTC | RESOLVED: checkout-api error rate < 1% | Monitoring | ALR-4 |
| 18:31 | 11:31 PDT | Error rate under 1%, monitoring | On-call engineer (Maya) | SLK-10 |
| 18:58 | 11:58 PDT | Incident declared resolved. ~3,100 failed checkout attempts per dashboard; 41 support tickets | On-call engineer (Maya) | SLK-11 |
Conflicts
| Event | Source A time | Source B time | Note |
|---|---|---|---|
| Rollback to v4.17.3 | 18:12 UTC (CI-2, deploy log) | 18:14 UTC / 11:14 PDT (SLK-7, "rolling back now") | Neither is chosen. CI-2 may record when the rollback was triggered and SLK-7 when it was announced, but that is not confirmed. |
max_connections 200 → 400 | 18:20 UTC (CI-3, deploy log) | 18:25 UTC / 11:25 PDT (SLK-9, "bumped max_connections to 400") | 5-minute gap. SLK-8 (18:21) still reports the DB at max connections. If the change took effect at 18:20, that message describes the state just after it was applied. |
| Incident date | 2026-09-30 (ALR-1..4, CI-1..3) | "yesterday" in the request, which is 2026-10-01 relative to today (2026-10-02) | This draft uses the date in the logs. |
| Day of week | Wednesday (2026-09-30 per calendar) | "Tuesday afternoon" (SLK-5) | Recorded as said. It does not change the analysis. |
Response Metrics
| Metric | Value | From | To |
|---|---|---|---|
| Time to detect (change → first alert) | 4 min | 17:38 deploy (CI-1) | 17:42 WARN latency (ALR-1) |
| Time to page (change → page) | 8 min | 17:38 deploy (CI-1) | 17:46 page (ALR-3) |
| Time to detect (impact start → first alert) | NOT PROVIDED | Impact start: UNKNOWN | 17:42 (ALR-1) |
| Time to acknowledge | 1 min | 17:46 page (ALR-3) | 17:47 ack (SLK-1) |
| Time to diagnose (page → pool exhaustion identified) | 3 min | 17:46 page (ALR-3) | 17:49 (SLK-2) |
| Time to first mitigation (page → rollback) | 26–28 min (CONFLICT) | 17:46 page (ALR-3) | 18:12 (CI-2) / 18:14 (SLK-7) |
| Time to recover (page → error rate < 1%) | 45 min | 17:46 page (ALR-3) | 18:31 RESOLVED (ALR-4) |
| Time to resolve (page → declared resolved) | 72 min | 17:46 page (ALR-3) | 18:58 (SLK-11) |
| Observed impact window (first alert → recovery) | 49 min (lower bound; true start UNKNOWN, no earlier than 17:38) | 17:42 (ALR-1) | 18:31 (ALR-4) |
Impact
| Measure | Value | Source |
|---|---|---|
| Failed checkout attempts | ~3,100 (approximate, as stated) | SLK-11, "per the dashboard". The dashboard itself was not supplied. |
| Support tickets | 41 | SLK-11 |
| Peak error rate (alerted) | 34% | ALR-2 |
| p95 latency (alerted) | > 2s (threshold 800ms) | ALR-1 |
| Unique customers affected | NOT PROVIDED | — |
| Revenue impact | NOT PROVIDED (finance has not supplied a figure) | NOTE-4 |
| SLO / SLA impact | NOT PROVIDED | — |
| Customer communication (status page) | NOT PROVIDED | — |
Contributing Factors
| ID | Category | Factor | Evidence |
|---|---|---|---|
| T1 | Technical | v4.18.0 raised the DB pool idle timeout. INFERRED: idle connections were likely held longer, so each checkout-api instance kept more of payments-db's 200 connections. | CI-1 (change description), SLK-2 (pool exhausted). There is no evidence on connection counts before and after the deploy. Unconfirmed. |
| T2 | Technical | The card-auth retry logic added in v4.18.0 retried each failed auth 3x with no backoff and no cap, so every failure became up to 4 DB-touching attempts. Under pool pressure, this turned failures into more load (a retry storm). | SLK-4, NOTE-2, CI-1 |
| T3 | Technical | payments-db stayed at max connections after the rollback, and errors only cleared after max_connections was raised to 400. Why connections were not released on rollback is unknown. | SLK-8, CI-3 / SLK-9, ALR-4 |
| P1 | Process | The pipeline has a canary stage but allowed it to be skipped with a flag set by hand. No reason or approval was recorded (that we have). v4.18.0 therefore went straight to full production traffic. | NOTE-1; the question was also raised in SLK-5 |
| D1 | Detection | There is no alert on payments-db connection saturation. Detection depended on downstream API symptoms (latency, error rate), and the DB-level cause was found by hand. | NOTE-3; SLK-2 shows manual diagnosis |
5 Whys
- Why did checkouts fail? checkout-api requests failed with "connection pool exhausted" on payments-db. (SLK-2)
Why was the pool exhausted? payments-db was at its 200-connection limit (SLK-8, CI-3). Failed card-auth calls were each retried 3x, which multiplied load (SLK-4). INFERRED: the raised idle timeout in the same release made connections stay held longer (CI-1, T1).
3.
Why did retries make it worse instead of riding through? The retry logic had no backoff and no cap, so retries came immediately and at full volume. (NOTE-2)
4. Why did this reach full production traffic before the effect was seen? The canary stage, which would have exposed the release to a fraction of traffic first, was skipped by a flag set by hand. (NOTE-1)
5.
Why could the canary be skipped? The pipeline accepts a manual skip flag.
The evidence stops here. Why it was set for this deploy, and whether skipping needs approval, is not known (Open Questions #3).
What Went Well
Fast detection and response. The first alert fired 4 minutes after the deploy (CI-1 → ALR-1), and the page was acknowledged within 1 minute (ALR-3 → SLK-1).
Fast, accurate diagnosis. Pool exhaustion was identified 2 minutes after ack (SLK-2), and the retry amplification was identified by 18:02 (SLK-4).
External cause ruled out early. The card-auth provider status was checked at 17:55 (SLK-3), so the investigation stayed on internal changes.
- Rollback was available and used. v4.17.3 was restored about 6 minutes after the proposal (SLK-6 → CI-2/SLK-7).
Responders kept checking after the first mitigation. Partial recovery was noticed (SLK-8), and a second mitigation was applied (CI-3/SLK-9).
Recovery was confirmed before closing. The error rate fell below 1% at 18:31 (ALR-4), and the team watched for 27 more minutes before declaring the incident resolved (SLK-11).
What Could Improve
Canary protection was bypassable (P1 → CI-1, NOTE-1). The safeguard existed but could be turned off without, as far as we know, a recorded reason.
No DB-level signal (D1 → SLK-2). Responders had to find connection exhaustion by hand. An alert on payments-db connections would point to the cause directly.
- Retry policy lacked backoff and a cap (T2 → SLK-4). Retries turned a capacity problem into a load amplifier.
Rollback did not fully restore service (T3 → SLK-8). Recovery needed a second change, an emergency max_connections increase. Why connections stayed saturated after rollback is unexplained.
Mid-incident attribution (SLK-5). A "who shipped this" message in the incident channel named an individual while responders were mitigating. Keeping the channel on mitigation, and sending causal questions to the postmortem, keeps responders focused and keeps the review blameless.
Conflicting timestamps between CI and Slack (Conflicts table). Rollback and config-change times differ by 2 and 5 minutes. A single incident log, or bot-posted CI events in the channel, would remove the ambiguity.
Action Items
Items 1–3 come from the incident lead's notes (NOTE-5). Items 4–5 are
proposed by this draft to cover factors that otherwise have no action, so accept or drop them in review. All priorities are proposals: the notes gave none.
| # | Action | Type | Addresses Factor | Priority | Owner | Due | Done When |
|---|---|---|---|---|---|---|---|
| 1 | Add backoff and a retry cap to card-auth retries in checkout-api | Prevent | T2 | P1 (proposed) | Raj (tentative per NOTE-5, unconfirmed) | NO DATE | A test that simulates repeated card-auth failures shows retries spaced by increasing delays and stopping at the configured cap. The change is merged and deployed. |
| 2 | Make the canary stage mandatory for checkout-api production deploys | Prevent | P1 | P1 (proposed) | UNASSIGNED | NO DATE | A production deploy with the canary-skip flag set is blocked by the pipeline (or, if an emergency path is kept, it requires a recorded approval and reason; see Open Questions #4). |
| 3 | Add an alert on payments-db connection saturation, routed to the payments on-call rotation | Detect | D1 | P1 (proposed) | UNASSIGNED | NO DATE | The alert exists with a threshold below max_connections (threshold to be chosen by the owner), and a test firing reaches the payments rotation. |
| 4 | (Proposed) Measure how the v4.18.0 idle-timeout change affects payments-db connection count before that change is reshipped | Prevent | T1 | P2 (proposed) | UNASSIGNED | NO DATE | The connection count with and without the idle-timeout change, at comparable load, is recorded in this postmortem, and T1 is marked confirmed or ruled out. |
| 5 | (Proposed) Determine why payments-db stayed at max connections after the rollback to v4.17.3 | Mitigate | T3 | P2 (proposed) | UNASSIGNED | NO DATE | A written explanation, backed by evidence (e.g., connection source breakdown for 18:12–18:31), is added to this postmortem. |
Similar Incidents
None named in the evidence provided. (See Open Questions #9.)
Open Questions
Date: the request says "yesterday" (2026-10-01), but every timestamped source says 2026-09-30. Confirm 2026-09-30.
2.
Severity, incident owner, and incident commander: none were stated. Was there a separate incident commander, or did the on-call engineer both coordinate and diagnose?
3.
Canary skip (P1): why was the skip flag set for this deploy? Who can set it, and is a reason or approval recorded anywhere? (The 5 Whys stop here.)
4.
Canary "mandatory" scope (Action 2): should an emergency path remain (e.g., for urgent hotfixes), and if so, what approval should it need?
5.
Rollback and config-change conflicts: can CI job logs or the DB config audit log settle 18:12 vs 18:14 (rollback) and 18:20 vs 18:25 (max_connections)?
6.
max_connections = 400: this was an emergency change during the incident. Is it staying, and has anyone confirmed that payments-db is sized for 400 connections? (This draft makes no recommendation. Any change should go through normal change review.)
7.
Impact gaps: what are the revenue figure (pending from finance), the unique customers affected, the SLO/SLA effect, and the source dashboard behind "~3,100"? Was a status page or customer notice posted?
8.
Impact start time: when did errors actually begin, between 17:38 and 17:42? An error-rate graph would let time-to-detect be measured from impact start.
9. Prior incidents: has checkout-api or payments-db had earlier connection-exhaustion or retry-storm incidents?
10. Responder roles: what is Raj's role (secondary on-call, payments engineer, other)?
11. Action owners and dates: confirm Raj for Action 1, and assign owners and due dates for Actions 2–5.
Review and Approval
| Role | Name | Date |
|---|---|---|
| Draft author | Claude (from supplied evidence) | 2026-10-02 |
| Incident owner review | NOT PROVIDED | |
| Technical review (payments) | ||
| Action item approval |
Next review: 30 days after approval, to check action item completion.
incident-postmortem-with-evidence-trail.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
For SREs, on-call engineers, and engineering managers who have to write the postmortem after the outage is over. Point it at the incident-channel export, the alert and paging history, the deploy log, and your notes, and it writes a blameless draft that a reviewer can check line by line rather than take on trust.
What you get
- A timeline with an evidence trail. Every entry cites the log line, alert, or message it came from. One timezone is stated and every source is converted to it.
- Conflicts shown, gaps marked. When two sources disagree on a time, both appear side by side with their sources. An event with no timestamp is marked UNKNOWN, never placed at a guessed time.
- Response metrics from real timestamps only. Time to detect, acknowledge, mitigate, and resolve each name the two timeline rows they were computed from; when an endpoint is missing, the metric reads NOT PROVIDED.
- Impact from your numbers. Customers, failed requests, revenue, and SLO figures each carry their source; anything you call a guess stays labelled ESTIMATE.
- Contributing factors, not one "root cause". Grouped as technical, process, detection, and organizational, with 5 Whys where a factor warrants it.
- Action items that trace back. Each names the factor it addresses, is typed prevent, detect, or mitigate, and has an owner and due date or shows UNASSIGNED / NO DATE.
- Variations for security incidents, near-misses, recurring patterns, and a customer-facing version on request.
What it refuses to do
- Assign blame. People appear by role, and a request for "whose fault" is reframed toward the safeguards that were missing.
- Write the postmortem while the incident is still in progress; you get brief stabilization pointers instead.
- Invent times, durations, impact figures, owners, or causes, or quietly settle a conflict between sources.
- Publish, post, or send anything. The draft stays marked for incident-owner review, and a customer-facing version is marked not for release until reviewed.
- Accept "be more careful" as an action item.
What's in the zip
SKILL.md: the skill.references/recipe.md: the full step-by-step recipe (about 3,900 words) with the analysis prompt, a filled example postmortem, and the variations.evals/: three test cases you can run to check its behavior.LICENSE.txt: single-purchaser license; use it in your own work, including for clients.
Built for the postmortem your reviewers can verify against the evidence, not just read.
The demo below is a real run on a fictional incident: Claude's reply, then the full postmortem it wrote.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- Passed all security checks, Safe to install