Production Reliability & SRE Toolkit
Improve production reliability with a practical SRE workflow spanning audits, resilience analysis, failure testing, incident learning, and targeted repairs. This bundle covers production-readiness review, reliability scoring, SRE improvements, chaos engineering, keep-alive socket leaks, shutdown data loss, incident postmortems, and orphaned-resource reconciliation. Ideal for systems that must survive deployments, failures, traffic spikes, and operational change.
Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.
One payment, lifetime access. 8 skills unlock instantly in your library.
30-day refund guarantee
Instant unlock in your library
Free updates from the creator
What's included
8 skillsIs your system ready for production? This skill audits your system and identifies risks, missing components, and improvements before deployment. ✔ Detect pro...
Evaluate how reliable your system really is. This skill analyzes your architecture and generates a reliability score with insights and improvement suggestion...
Want to improve your system reliability? This skill analyzes your system and provides SRE-level recommendations to improve uptime, resilience, and fault tole...
Test System Resilience Like a Chaos Engineer.
Repair HTTP keep-alive lifecycle bugs that leak sockets, exhaust pools, or cause intermittent resets after idle periods and deploys.
Repair apps that lose jobs, events, or writes when containers restart, deployments roll, or processes receive termination signals.
Need to document a production incident? This skill generates structured incident postmortems with timeline, root cause analysis, impact, and action items. ✔...
Safely detect, verify, quarantine, and clean up orphaned cloud resources after failed provisioning, deletions, or database drift.
See it in action
You say
Audit my Node.js API running on AWS Lambda with a DynamoDB backend for production readiness. It handles user authentication and file uploads.
Your agent does
System Overview: Serverless Node.js API on AWS. Production Risks: Cold start latency, DynamoDB throttling. Missing Components: DLQs for failed uploads, structured logging. Failure Scenarios: Regional AWS outage. Readiness: Needs Work. Implement auto-scaling limits and monitoring.
How to install
Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.
Reviews
No reviews yet
Be one of the first to try it. Every skill in this bundle passes our trust checks.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back