Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Chaos Engineering

    3

    Design rigorous chaos engineering experiments and resilience audits to verify production system reliability.

    Free

    6 installs5.0 (1 review)
    chaos-engineering

    chaos-engineering

    Example session with this skill installed

    Our payment gateway has been slow lately. Can you design a chaos experiment to see how the checkout process handles a latency spike in the payment service?

    • Read your context and instructions
    • Compiled the chaos-engineering

    Hypothesis: P99 latency for /checkout remains <1.2s during payment gateway latency.
    Perturbation: Inject 300ms latency on the 'payment-v2' service for 5% of traffic for 10 mins.
    Abort Condition: Error rate > 2% for 120s.
    Targeted Amplifier: Retry storm and thread-pool exhaustion.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The Science of Controlled Failure

    Moving beyond generic checklists, this skill transforms your AI agent into a senior Chaos Engineer. It addresses the fundamental problem of "theoretical resilience" by replacing vague recommendations with falsifiable, evidence-based experimitalic textents. Instead of suggesting you "add retries," it helps you design the exact stress test needed to prove your system won't collapse under a retry storm.

    What it does

    • Experiment Design: Drafts specific chaos experiments with measurable hypotheses, single-variable perturbations, and defined blast radii.
    • Resilience Auditing: Identifies hidden architectural amplifiers like thundering herds, gray failures, and synchronized backoffs.
    • Operational Rigor: Defines the human roles (Lead, Observer, Abort Authority) and readiness flags required to run experiments safely in production.
    • Post-Mortem Conversion: Analyzes past incidents to create "never again" experiments that verify fixes.

    Why use this skill?

    Standard AI prompting often results in "best practice" lists that are difficult to action. This skill enforces a rigorous four-phase procedure (Hypothesize, Perturb, Minimize, Learn) that treats infrastructure as a laboratory. It focuses on tail-risk (P99/P99.9) rather than averages, ensuring your systems are hardened against the worst-case scenarios that actually cause outages.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    5.0
    1 review
    5
    1
    4
    0
    3
    0
    2
    0
    1
    0

    6 people have installed this skill.

    Trust & safety

    Security scanned

    Verified clean 4 months ago

    • Free to download with an account

    Listed4 months ago

    What's inside

    Frequently Asked Questions