- Home
- Skills
- DevOps & Cloud
- Incident Postmortem
Works with the AI tools you already use
Incident Postmortem
Structured post-incident reviews: timeline, root cause, action items, and blameless write-ups.
$5
Incident Postmortem
Example session with this skill installed
Draft a postmortem for the API outage on 2023-10-12. Logs show 500 errors started at 14:02 UTC after the v2.1.0 deploy at 13:58. Rollback completed at 14:15. 15% of traffic was affected.
- Read your context and instructions
- Compiled the incident postmortem
Postmortem: API Outage — 2023-10-12
Impact: 15% of requests · 13 mins · SLA Breach
Root Cause
Deploy v2.1.0 → Database connection leak in auth middleware → Connection pool exhaustion
Actions
| # | Category | Action | Owner |
|---|---|---|---|
| 1 | Prevent | Add pool leak detection to CI | @eng-ops |
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
The problem
Writing incident reports after an outage is a chore that often leads to vague descriptions or finger-pointing. Developers struggle to turn messy Slack logs, alert timestamps, and deployment history into a coherent, blameless retrospective.
What it does
- Synthesizes raw evidence like chat logs, alert history, and deploy commits into a structured report.
- Constructs a minute-by-minute chronological timeline of the incident from start to resolution.
- Performs a 5-Whys root cause analysis to identify systemic failures rather than human error.
- Quantifies impact on customers, SLOs, and SLAs using available data or tagged estimates.
- Generates a prioritized list of action items categorized by prevention, detection, and response improvement.
Why this beats prompting it yourself
This skill enforces a strict blameless methodology and a 5-Whys deep dive that generic prompts often skip. It automatically handles the tedious formatting of timelines and action item tables, ensuring your engineering team maintains a consistent standard for every P1 incident.
Use cases
- Creating a formal retrospective after a production outage or service degradation.
- Analyzing a security event or data loss incident to prevent recurrence.
- Standardizing internal documentation for SRE and on-call teams.
- Reviewing a failed deployment to identify gaps in CI/CD or testing pipelines.
Known limitations
This skill is for post-incident review only. Do not use it for real-time incident response, blame attribution, or generating legal liability documentation.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
151 installs
Downloaded by developers to date
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 month ago
- Passed all security checks, Safe to install