Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    Incident Postmortem with Evidence Trail

    1

    A blameless draft postmortem where every timeline entry cites its log, alert, or message, and every action item traces to a cause.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    Incident Postmortem with Evidence Trail

    Incident Postmortem with Evidence Trail

    Example session with this skill installed

    Write the postmortem for yesterday's checkout outage. Here's everything I have:

    --- Alerts (monitoring tool, times in UTC) ---
    2026-09-30 17:42 WARN checkout-api p95 latency > 2s (threshold 800ms)
    2026-09-30 17:46 CRIT checkout-api error rate 34% (threshold 5%)
    2026-09-30 17:46 page sent to on-call (payments rotation)
    2026-09-30 18:31 RESOLVED checkout-api error rate < 1%

    --- Deploy log (CI, times in UTC) ---
    2026-09-30 17:38 deploy checkout-api v4.18.0 (author: dev on payments team) - "raise DB pool idle timeout, add retry on card-auth"
    2026-09-30 18:12 rollback checkout-api to v4.17.3 (manual)
    2026-09-30 18:20 config change: payments-db max_connections 200 -> 400

    --- #inc-checkout Slack export (times shown in Pacific) ---
    10:47 Maya (on-call): ack, looking
    10:49 Maya: checkout errors are "connection pool exhausted" on payments-db
    10:55 Raj: card-auth provider status page is green
    11:02 Maya: retries are multiplying load, each failed auth retries 3x
    11:05 Dan (eng manager): who shipped 4.18 on a Tuesday afternoon without a canary? this is on Priya
    11:09 Maya: proposing rollback to 4.17.3
    11:14 Raj: rolling back now
    11:21 Maya: errors dropping but not gone, db still at max connections
    11:25 Raj: bumped max_connections to 400
    11:31 Maya: error rate under 1%, monitoring
    11:58 Maya: calling it resolved. ~3,100 checkout attempts failed per the dashboard. support got 41 tickets.

    --- My notes ---

    • Canary stage exists in the pipeline but was skipped for this deploy (flag set by hand). Don't know why.
    • Retry logic had no backoff or cap.
    • No alert on DB connection saturation, only on API errors.
    • Revenue impact: finance hasn't given us a number.
    • Action items so far: add backoff to retries (Raj?), make canary mandatory, add DB connection alert. No dates yet.
    • Read your context and instructions
    • Compiled the incident postmortem
    • Generated the document

    I've written the draft postmortem: postmortem-2026-09-30-checkout-api.md. It's marked as a draft pending review by the incident owner, and nothing has been posted or sent.

    Timeline

    Timezone: everything is in UTC. Your Slack times were Pacific Daylight Time on Sept 30 (UTC−7). Two Slack messages land exactly where the alerts say they should, which confirms that offset.

    Sources: each timeline row is tagged with the line it came from. ALR is the alerts, CI the deploy log, SLK the Slack export, NOTE your notes.

    Response times: about 4 min from deploy to first alert, 1 min to acknowledge, 26–28 min to rollback, 45 min until errors were under 1%, 72 min until it was declared resolved.

    When customers were affected: at least 49 min. The true start is unknown; it falls somewhere between the 17:38 deploy and the 17:42 alert.

    Places the evidence disagrees (I recorded both times rather than picking one):

    • Rollback: the deploy log says 18:12 UTC; Raj's "rolling back now" in Slack is 18:14 UTC.
    • max_connections change: the deploy log says 18:20 UTC; Slack says 18:25 UTC.
    • Date: you said "yesterday", which would be Oct 1. Every log says Sept 30, so the draft uses Sept 30.
    • Day of week: Dan said "Tuesday", but Sept 30, 2026 was a Wednesday.

    Keeping it blameless

    • I left out Dan's 11:05 "this is on Priya." I kept the question he raised (why it shipped without a canary) and treated it as a process issue: the pipeline lets someone skip the canary by setting a flag by hand.
    • The draft names people by role and never names anyone as a cause. A check for "failed to", "should have", "careless", "human error", "fault" and Priya's name came back clean.

    Contributing factors

    • Retries: card-auth retries had no backoff or cap, so they multiplied load on the database.
    • Canary skipped: the canary stage was turned off for this deploy.
    • No database alert: nothing alerts when payments-db runs out of connections; only the API errors alerted.

    Idle timeout (not confirmed): raising the pool idle timeout probably made connections stay held longer. It's labelled as inferred because no connection-count data supports it yet.

    Rollback didn't fully clear it: the database stayed at its connection limit after the rollback, and nothing in the evidence explains why.

    Action items

    • Your three are in: retry backoff (Raj, marked unconfirmed because your note had a "?"), mandatory canary, and the database connection alert.
    • I added two proposed ones, labelled so you can drop them: measure the idle-timeout effect before that change ships again, and find out why the database stayed full after the rollback.
    • The P1/P2 priorities are my suggestions; your notes didn't give any. Owners other than Raj are unassigned and there are no due dates, as in your notes.

    Gaps. The doc ends with 11 open questions for review. The main ones are:

    • Severity, incident owner, and whether anyone acted as incident commander were never stated.
    • Why the canary was skipped is unknown; the root-cause analysis stops there.
    • Whether max_connections stays at 400. I made no recommendation; that's for normal change review.
    • The revenue figure is still pending from finance, and there's no unique-customer count, SLA impact, or status-page record.
    DeliverableLast modifiedIncludes
    postmortem-2026-09-30-checkout-api.md (in your working folder)2026-10-02 09:40Full draft; checked for blame wording, every timeline row has a source, and every time either appears in your inputs or is a Pacific-to-UTC conversion of one

    No other copies exist.


    The file it wrote: postmortem-2026-09-30-checkout-api.md

    DRAFT - PENDING INCIDENT OWNER REVIEW

    Postmortem: checkout-api errors after v4.18.0 deploy (payments-db connection exhaustion)

    FieldValue
    Incident date2026-09-30 (from alert and deploy logs; see Open Questions #1)
    TypeOutage (partial: peak error rate 34%)
    SeverityNOT PROVIDED
    Services affectedcheckout-api, payments-db
    Incident owner / reviewerNOT PROVIDED
    StatusResolved, declared 18:58 UTC
    Draft prepared2026-10-02

    Summary

    At 17:38 UTC on 2026-09-30, checkout-api v4.18.0 was deployed. The release raised the DB pool idle timeout and added retries on card-auth calls. Four minutes later, at 17:42, a p95 latency alert fired. By 17:46 the error rate was 34% and the payments on-call rotation was paged. The on-call engineer found that checkout errors were "connection pool exhausted" on payments-db. Failed card-auth calls were each retried 3 times with no backoff or cap, which multiplied the load on the database. Responders rolled back to v4.17.3 (18:12–18:14 UTC; the two sources disagree) and raised payments-db max_connections from 200 to 400 (18:20–18:25 UTC; the two sources disagree). The error rate fell below 1% at 18:31 UTC, and the incident was declared resolved at 18:58 UTC. About 3,100 checkout attempts failed and support received 41 tickets. Revenue impact is still pending from finance.

    Contributing factors (summary): retries with no backoff or cap amplified load on a shared DB connection limit (T2). The canary stage was skipped for this deploy through a flag set by hand (P1). There was no alert on DB connection saturation (D1). The raised pool idle timeout is a likely but

    unconfirmed contributor (T1, INFERRED).


    Timeline

    Reference timezone: UTC. The Slack export is in Pacific time. On 2026-09-30 that is PDT (UTC−7), checked with Intl.DateTimeFormat for America/Los_Angeles. The offset is supported by two independent matches: the on-call ack at 10:47 PDT = 17:47 UTC, one minute after the 17:46 page, and "error rate under 1%" at 11:31 PDT = 18:31 UTC, which matches the RESOLVED alert. The Slack export has no date, so its date is INFERRED as 2026-09-30 from these matches.

    Source key (the evidence as supplied, numbered in the order pasted):

    • ALR-1..4: monitoring alerts, lines 1–4 (UTC)
    • CI-1..3: deploy log, lines 1–3 (UTC)
    • SLK-1..11: #inc-checkout Slack export, messages 1–11 (Pacific)
    • NOTE-1..5: incident lead's notes, bullets 1–5
    Time (UTC)OriginalEventActor (role)Source
    17:3817:38 UTCcheckout-api v4.18.0 deployed: "raise DB pool idle timeout, add retry on card-auth"CI pipeline; change authored by a payments-team engineerCI-1
    UNKNOWN—Canary stage skipped for this deploy by a skip flag set by handUnknownNOTE-1
    17:4217:42 UTCWARN: checkout-api p95 latency > 2s (threshold 800ms). First evidence of impact.MonitoringALR-1
    17:4617:46 UTCCRIT: checkout-api error rate 34% (threshold 5%)MonitoringALR-2
    17:4617:46 UTCPage sent to on-call (payments rotation)PagingALR-3
    17:4710:47 PDTPage acknowledged, investigation startedOn-call engineer, payments (Maya)SLK-1
    17:4910:49 PDTDiagnosis: checkout errors are "connection pool exhausted" on payments-dbOn-call engineer (Maya)SLK-2
    17:5510:55 PDTCard-auth provider status page checked: green (external cause less likely)Responder (Raj; role not stated)SLK-3
    18:0211:02 PDTDiagnosis: retries are multiplying load; each failed auth retries 3xOn-call engineer (Maya)SLK-4
    18:0511:05 PDTEngineering manager asks in channel why 4.18 shipped on a weekday afternoon without a canary. (The message also attributed the deploy to a named individual. That attribution is left out of this blameless review. The process question is handled under P1.)Engineering manager (Dan)SLK-5
    18:0911:09 PDTDecision: rollback to v4.17.3 proposedOn-call engineer (Maya)SLK-6
    18:12 / 18:1418:12 UTC / 11:14 PDTMitigation 1: manual rollback of checkout-api to v4.17.3 (CONFLICT, see below)Responder (Raj)CI-2 / SLK-7
    18:2111:21 PDTErrors dropping but not gone; payments-db still at max connectionsOn-call engineer (Maya)SLK-8
    18:20 / 18:2518:20 UTC / 11:25 PDTMitigation 2: payments-db max_connections 200 → 400 (CONFLICT, see below)Responder (Raj)CI-3 / SLK-9
    18:3118:31 UTCRESOLVED: checkout-api error rate < 1%MonitoringALR-4
    18:3111:31 PDTError rate under 1%, monitoringOn-call engineer (Maya)SLK-10
    18:5811:58 PDTIncident declared resolved. ~3,100 failed checkout attempts per dashboard; 41 support ticketsOn-call engineer (Maya)SLK-11

    Conflicts

    EventSource A timeSource B timeNote
    Rollback to v4.17.318:12 UTC (CI-2, deploy log)18:14 UTC / 11:14 PDT (SLK-7, "rolling back now")Neither is chosen. CI-2 may record when the rollback was triggered and SLK-7 when it was announced, but that is not confirmed.
    max_connections 200 → 40018:20 UTC (CI-3, deploy log)18:25 UTC / 11:25 PDT (SLK-9, "bumped max_connections to 400")5-minute gap. SLK-8 (18:21) still reports the DB at max connections. If the change took effect at 18:20, that message describes the state just after it was applied.
    Incident date2026-09-30 (ALR-1..4, CI-1..3)"yesterday" in the request, which is 2026-10-01 relative to today (2026-10-02)This draft uses the date in the logs.
    Day of weekWednesday (2026-09-30 per calendar)"Tuesday afternoon" (SLK-5)Recorded as said. It does not change the analysis.

    Response Metrics

    MetricValueFromTo
    Time to detect (change → first alert)4 min17:38 deploy (CI-1)17:42 WARN latency (ALR-1)
    Time to page (change → page)8 min17:38 deploy (CI-1)17:46 page (ALR-3)
    Time to detect (impact start → first alert)NOT PROVIDEDImpact start: UNKNOWN17:42 (ALR-1)
    Time to acknowledge1 min17:46 page (ALR-3)17:47 ack (SLK-1)
    Time to diagnose (page → pool exhaustion identified)3 min17:46 page (ALR-3)17:49 (SLK-2)
    Time to first mitigation (page → rollback)26–28 min (CONFLICT)17:46 page (ALR-3)18:12 (CI-2) / 18:14 (SLK-7)
    Time to recover (page → error rate < 1%)45 min17:46 page (ALR-3)18:31 RESOLVED (ALR-4)
    Time to resolve (page → declared resolved)72 min17:46 page (ALR-3)18:58 (SLK-11)
    Observed impact window (first alert → recovery)49 min (lower bound; true start UNKNOWN, no earlier than 17:38)17:42 (ALR-1)18:31 (ALR-4)

    Impact

    MeasureValueSource
    Failed checkout attempts~3,100 (approximate, as stated)SLK-11, "per the dashboard". The dashboard itself was not supplied.
    Support tickets41SLK-11
    Peak error rate (alerted)34%ALR-2
    p95 latency (alerted)> 2s (threshold 800ms)ALR-1
    Unique customers affectedNOT PROVIDED—
    Revenue impactNOT PROVIDED (finance has not supplied a figure)NOTE-4
    SLO / SLA impactNOT PROVIDED—
    Customer communication (status page)NOT PROVIDED—

    Contributing Factors

    IDCategoryFactorEvidence
    T1Technicalv4.18.0 raised the DB pool idle timeout. INFERRED: idle connections were likely held longer, so each checkout-api instance kept more of payments-db's 200 connections.CI-1 (change description), SLK-2 (pool exhausted). There is no evidence on connection counts before and after the deploy. Unconfirmed.
    T2TechnicalThe card-auth retry logic added in v4.18.0 retried each failed auth 3x with no backoff and no cap, so every failure became up to 4 DB-touching attempts. Under pool pressure, this turned failures into more load (a retry storm).SLK-4, NOTE-2, CI-1
    T3Technicalpayments-db stayed at max connections after the rollback, and errors only cleared after max_connections was raised to 400. Why connections were not released on rollback is unknown.SLK-8, CI-3 / SLK-9, ALR-4
    P1ProcessThe pipeline has a canary stage but allowed it to be skipped with a flag set by hand. No reason or approval was recorded (that we have). v4.18.0 therefore went straight to full production traffic.NOTE-1; the question was also raised in SLK-5
    D1DetectionThere is no alert on payments-db connection saturation. Detection depended on downstream API symptoms (latency, error rate), and the DB-level cause was found by hand.NOTE-3; SLK-2 shows manual diagnosis

    5 Whys

    1. Why did checkouts fail? checkout-api requests failed with "connection pool exhausted" on payments-db. (SLK-2)

    Why was the pool exhausted? payments-db was at its 200-connection limit (SLK-8, CI-3). Failed card-auth calls were each retried 3x, which multiplied load (SLK-4). INFERRED: the raised idle timeout in the same release made connections stay held longer (CI-1, T1).
    3.

    Why did retries make it worse instead of riding through? The retry logic had no backoff and no cap, so retries came immediately and at full volume. (NOTE-2)
    4. Why did this reach full production traffic before the effect was seen? The canary stage, which would have exposed the release to a fraction of traffic first, was skipped by a flag set by hand. (NOTE-1)
    5.

    Why could the canary be skipped? The pipeline accepts a manual skip flag.

    The evidence stops here. Why it was set for this deploy, and whether skipping needs approval, is not known (Open Questions #3).


    What Went Well

    Fast detection and response. The first alert fired 4 minutes after the deploy (CI-1 → ALR-1), and the page was acknowledged within 1 minute (ALR-3 → SLK-1).

    Fast, accurate diagnosis. Pool exhaustion was identified 2 minutes after ack (SLK-2), and the retry amplification was identified by 18:02 (SLK-4).

    External cause ruled out early. The card-auth provider status was checked at 17:55 (SLK-3), so the investigation stayed on internal changes.

    • Rollback was available and used. v4.17.3 was restored about 6 minutes after the proposal (SLK-6 → CI-2/SLK-7).

    Responders kept checking after the first mitigation. Partial recovery was noticed (SLK-8), and a second mitigation was applied (CI-3/SLK-9).

    Recovery was confirmed before closing. The error rate fell below 1% at 18:31 (ALR-4), and the team watched for 27 more minutes before declaring the incident resolved (SLK-11).

    What Could Improve

    Canary protection was bypassable (P1 → CI-1, NOTE-1). The safeguard existed but could be turned off without, as far as we know, a recorded reason.

    No DB-level signal (D1 → SLK-2). Responders had to find connection exhaustion by hand. An alert on payments-db connections would point to the cause directly.

    • Retry policy lacked backoff and a cap (T2 → SLK-4). Retries turned a capacity problem into a load amplifier.

    Rollback did not fully restore service (T3 → SLK-8). Recovery needed a second change, an emergency max_connections increase. Why connections stayed saturated after rollback is unexplained.

    Mid-incident attribution (SLK-5). A "who shipped this" message in the incident channel named an individual while responders were mitigating. Keeping the channel on mitigation, and sending causal questions to the postmortem, keeps responders focused and keeps the review blameless.

    Conflicting timestamps between CI and Slack (Conflicts table). Rollback and config-change times differ by 2 and 5 minutes. A single incident log, or bot-posted CI events in the channel, would remove the ambiguity.


    Action Items

    Items 1–3 come from the incident lead's notes (NOTE-5). Items 4–5 are

    proposed by this draft to cover factors that otherwise have no action, so accept or drop them in review. All priorities are proposals: the notes gave none.

    #ActionTypeAddresses FactorPriorityOwnerDueDone When
    1Add backoff and a retry cap to card-auth retries in checkout-apiPreventT2P1 (proposed)Raj (tentative per NOTE-5, unconfirmed)NO DATEA test that simulates repeated card-auth failures shows retries spaced by increasing delays and stopping at the configured cap. The change is merged and deployed.
    2Make the canary stage mandatory for checkout-api production deploysPreventP1P1 (proposed)UNASSIGNEDNO DATEA production deploy with the canary-skip flag set is blocked by the pipeline (or, if an emergency path is kept, it requires a recorded approval and reason; see Open Questions #4).
    3Add an alert on payments-db connection saturation, routed to the payments on-call rotationDetectD1P1 (proposed)UNASSIGNEDNO DATEThe alert exists with a threshold below max_connections (threshold to be chosen by the owner), and a test firing reaches the payments rotation.
    4(Proposed) Measure how the v4.18.0 idle-timeout change affects payments-db connection count before that change is reshippedPreventT1P2 (proposed)UNASSIGNEDNO DATEThe connection count with and without the idle-timeout change, at comparable load, is recorded in this postmortem, and T1 is marked confirmed or ruled out.
    5(Proposed) Determine why payments-db stayed at max connections after the rollback to v4.17.3MitigateT3P2 (proposed)UNASSIGNEDNO DATEA written explanation, backed by evidence (e.g., connection source breakdown for 18:12–18:31), is added to this postmortem.

    Similar Incidents

    None named in the evidence provided. (See Open Questions #9.)


    Open Questions

    Date: the request says "yesterday" (2026-10-01), but every timestamped source says 2026-09-30. Confirm 2026-09-30.
    2.

    Severity, incident owner, and incident commander: none were stated. Was there a separate incident commander, or did the on-call engineer both coordinate and diagnose?
    3.

    Canary skip (P1): why was the skip flag set for this deploy? Who can set it, and is a reason or approval recorded anywhere? (The 5 Whys stop here.)
    4.

    Canary "mandatory" scope (Action 2): should an emergency path remain (e.g., for urgent hotfixes), and if so, what approval should it need?
    5.

    Rollback and config-change conflicts: can CI job logs or the DB config audit log settle 18:12 vs 18:14 (rollback) and 18:20 vs 18:25 (max_connections)?
    6.

    max_connections = 400: this was an emergency change during the incident. Is it staying, and has anyone confirmed that payments-db is sized for 400 connections? (This draft makes no recommendation. Any change should go through normal change review.)
    7.

    Impact gaps: what are the revenue figure (pending from finance), the unique customers affected, the SLO/SLA effect, and the source dashboard behind "~3,100"? Was a status page or customer notice posted?
    8.

    Impact start time: when did errors actually begin, between 17:38 and 17:42? An error-rate graph would let time-to-detect be measured from impact start.
    9. Prior incidents: has checkout-api or payments-db had earlier connection-exhaustion or retry-storm incidents?
    10. Responder roles: what is Raj's role (secondary on-call, payments engineer, other)?
    11. Action owners and dates: confirm Raj for Action 1, and assign owners and due dates for Actions 2–5.


    Review and Approval

    RoleNameDate
    Draft authorClaude (from supplied evidence)2026-10-02
    Incident owner reviewNOT PROVIDED
    Technical review (payments)
    Action item approval

    Next review: 30 days after approval, to check action item completion.

    incident-postmortem-with-evidence-trail.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Build a verified incident timeline from Slack and PagerDuty exports.Calculate TTR and TTD metrics using only timestamped evidence.Draft blameless RCA reports that focus on systems over individuals.Generate traceable action items tied to specific contributing factors.

    About this skill

    For SREs, on-call engineers, and engineering managers who have to write the postmortem after the outage is over. Point it at the incident-channel export, the alert and paging history, the deploy log, and your notes, and it writes a blameless draft that a reviewer can check line by line rather than take on trust.

    What you get

    • A timeline with an evidence trail. Every entry cites the log line, alert, or message it came from. One timezone is stated and every source is converted to it.
    • Conflicts shown, gaps marked. When two sources disagree on a time, both appear side by side with their sources. An event with no timestamp is marked UNKNOWN, never placed at a guessed time.
    • Response metrics from real timestamps only. Time to detect, acknowledge, mitigate, and resolve each name the two timeline rows they were computed from; when an endpoint is missing, the metric reads NOT PROVIDED.
    • Impact from your numbers. Customers, failed requests, revenue, and SLO figures each carry their source; anything you call a guess stays labelled ESTIMATE.
    • Contributing factors, not one "root cause". Grouped as technical, process, detection, and organizational, with 5 Whys where a factor warrants it.
    • Action items that trace back. Each names the factor it addresses, is typed prevent, detect, or mitigate, and has an owner and due date or shows UNASSIGNED / NO DATE.
    • Variations for security incidents, near-misses, recurring patterns, and a customer-facing version on request.

    What it refuses to do

    • Assign blame. People appear by role, and a request for "whose fault" is reframed toward the safeguards that were missing.
    • Write the postmortem while the incident is still in progress; you get brief stabilization pointers instead.
    • Invent times, durations, impact figures, owners, or causes, or quietly settle a conflict between sources.
    • Publish, post, or send anything. The draft stays marked for incident-owner review, and a customer-facing version is marked not for release until reviewed.
    • Accept "be more careful" as an action item.

    What's in the zip

    • SKILL.md: the skill.
    • references/recipe.md: the full step-by-step recipe (about 3,900 words) with the analysis prompt, a filled example postmortem, and the variations.
    • evals/: three test cases you can run to check its behavior.
    • LICENSE.txt: single-purchaser license; use it in your own work, including for clients.

    Built for the postmortem your reviewers can verify against the evidence, not just read.

    The demo below is a real run on a fictional incident: Claude's reply, then the full postmortem it wrote.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • Passed all security checks, Safe to install

    Listedtoday

    What's inside

    Frequently Asked Questions