Chaos Experiment and Resilience Testing Strategy

    1

    Plans chaos experiments: steady-state hypotheses, fault injection blast radius, automated abort tripwires, and game days.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Chaos Experiment and Resilience Testing Strategy

    Example session with this skill installed

    Design the chaos engineering experiment portfolio, blast-radius containment, and automated abort tripwires for our Payment Authorization Engine.

    • Subject: payment-authorization-engine v4.0 (run-id: banking-chaos-001) processing 18,000 authorizations/sec across 3 AWS AZs.
    • Deciders: Lead SRE Architect (David O'Reilly) and Head of Payment Infrastructure (Elena Rostova).
    • Scope: Verifying resilience under cross-AZ network partitions, Redis replica crashes, and 250ms WAN latency spikes; p99 latency <= 35 ms; error rate <= 0.05%.
    • Tension: Squads proposed unconstrained chaos in production without automated rollback triggers or steady-state baselines to "test reality"; David O'Reilly and Elena Rostova strictly reject untracked live chaos, citing incident INC-4939 where an unmonitored pod-kill drill crashed database connection pools, dropping $1.8M in checkouts. They mandate a formal chaos testing strategy: steady-state metrics, synthetic traffic canaries, bounded blast radii (max 1 AZ, max 10% traffic), sub-15-second automated abort tripwires, and pre-scheduled staging game days.

    Write the chaos testing strategy specification under docs/.

    • Read your context and instructions
    • Compiled the chaos experiment

    Wrote docs/architecture/tasks/banking-chaos-001/chaos-testing-strategy/chaos-testing-strategy.md. Complete chaos testing strategy establishing steady-state hypotheses, fault injection portfolios across AZs, sub-15s automated abort tripwires, and game day governance.


    skill: chaos-testing-strategy

    Chaos-Testing Strategy: Payment Authorization Engine [CHAOS-BANK-001]

    Summary

    This specification establishes the chaos engineering experiment portfolio, steady-state hypothesis models, blast-radius containment rules, and automated abort tripwires for payment-authorization-engine v4.0 under run ID banking-chaos-001. It governs 18,000 peak authorizations/second distributed across 3 AWS Availability Zones (us-east-1a, us-east-1b, us-east-1c). It decisively eliminates the uncontrolled cascade risks demonstrated in incident INC-4939 (where an unconstrained chaos test killed connection pools and halted $1.8M in checkouts). The strategy enforces measurable steady-state metric baselines (p99 latency <= 35 ms, error rate <= 0.05%), four controlled fault injection scenarios (network partition, pod termination, latency injection, database primary failover), tightly bounded blast-radius limits (<= 10% synthetic traffic canary, single AZ only), and sub-15-second automated abort tripwires.

    Detailed Description

    Injecting infrastructure faults without steady-state baselines and automated circuit breakers turns resilience engineering into uncontrolled self-inflicted outages. When systems experience unexpected dependency degradation—such as a cloud availability zone network partition or silent cache loss—untested retry storms and connection pool exhaustion trigger systemic failures. A formal chaos testing strategy proves fault tolerance under controlled, observable, and reversible experimental conditions.

    Chaos Experiment Invocation (Experiment: `EXP-NET-PARTITION-AZ1`)
                                │
                                ▼
    [ Chaos Controller & Blast-Radius Governor ]
      ├── 1. Asserts Steady-State Baseline: p99 <= 35ms, Errors <= 0.05%
      ├── 2. Isolates Blast Radius: Traffic routed to Synthetic Canary (10% max)
      └── 3. Injects Fault via AWS Fault Injection Simulator (FIS)
                                │
                                ▼ (Cross-AZ Link severed for 180s)
    [ Availability Zone `us-east-1a`: Network Partition Simulation ]
      ├── Envoy Circuit Breaker trips in < 150 ms
      └── Traffic shifts automatically to AZ-b and AZ-c
                                │
           ┌────────────────────┴────────────────────┐
           ▼ (Steady-State Maintained)               ▼ (Error Rate > 0.05% or Latency > 50ms)
    [ Hypothesis Validated: Pass ]           [ AUTOMATED ABORT TRIPWIRE FIRES (< 15s) ]
      Capture Metric Dashboard Artifacts       AWS FIS Aborts; Traffic Restored Instantly
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Blast-Radius Containment (Zero User Loss)Chaos injection must never impact real cardholder checkout transactions (INC-4939).0.40David O'Reilly (Lead SRE Architect)
    Automated Abort Tripwire Velocity (< 15s)If SLA metrics degrade, the experiment must terminate and recover in under 15 seconds.0.30Elena Rostova (Head of Payment Infra)
    Hypothesis Measurability & ObservabilityEvery experiment must evaluate specific Prometheus SLI queries against golden baselines.0.20Enterprise Reliability Engineering
    Production Parity of Injected FaultsInjected failure modes must accurately reflect real AWS cloud outages (partitions, panics).0.10Chaos Engineering Guild Standard

    Comparison

    Chaos Testing Operating ModelFault Target SeamBlast Radius ControlAbort MechanismEvaluation
    Option A: Live Ad-Hoc Chaos in ProdDirect production nodesNone (100% live traffic)Manual operator killRejected: Caused INC-4939 $1.8M outage; reckless and unmonitored.
    Option B: Isolated Dev Environment OnlyStaging mock clustersTotal isolationManual teardownRejected: Dev lacks high-concurrency traffic and real multi-AZ topology.
    Option C: Canary-Bounded FIS Chaos (Chosen)Ephemeral canary (10% load)Strictly bounded (1 AZ)Automated sub-15s Prometheus tripwireSelected: Real-world fidelity, zero customer impact, instant recovery.

    Result

    Option C is selected. AWS FIS executes controlled fault injection; Prometheus tripwires abort experiments automatically; synthetic canary traffic contains customer risk.


    Required Mechanisms

    1. Steady-State Hypothesis & Metrics Verification [MC-SS-01]
    • Steady-State Definition: Prior to fault injection, the system must maintain:
      1. http_request_duration_seconds{quantile="0.99"} <= 0.035 (35 ms).
      2. rate(http_requests_total{status=~"5.."}[1m]) / rate(http_requests_total[1m]) <= 0.0005 (0.05%).
      3. payment_engine_active_connection_pool_saturation <= 0.70 (70%).

    Hypothesis: Severing connectivity to one AZ or killing 30% of Redis replicas will cause p99 latency to rise by <= 8 ms with zero increase in 5xx errors, as traffic seamlessly drains to surviving nodes.

    2. Fault Injection Portfolio [MC-FP-01]
    Experiment IDInjected Fault ActionTarget InfrastructureDurationExpected Resilient Behavior
    EXP-CHAOS-01Sever Cross-AZ Peering RouteAWS TGW / Transit Subnet180 sEnvoy reroutes to AZ-b/c in < 200 ms; 0 dropped payments.
    EXP-CHAOS-02SIGKILL 50% Consumer PodsKubernetes Deployment60 sK8s replaces pods in < 15 s; surviving pods absorb load.
    EXP-CHAOS-03Inject 250ms Latency on DB PoolAurora PostgreSQL Primary120 sApp circuit breaker trips; serves graceful degraded response.
    EXP-CHAOS-04Blackhole Redis Cache ShardsRedis Cluster Node 02180 sApp falls back to read-through DB replicas with < 5 ms penalty.
    3. Automated Abort Tripwires & Rollback Protocol [MC-AT-01]
    • Prometheus Continuous Tripwire Daemon:
      • Evaluates SLIs every 2 seconds during active experiments.
    • Tripwire Abort Triggers:
      • ErrorRate > 0.05% for $> 6$ seconds.
      • p99_Latency > 50ms for $> 6$ seconds.
      • Any customer transaction timeout reported by edge API Gateway.
    • Abort Action:
      • Triggers AWS FIS StopExperiment API in < 3 seconds.
      • Restores all security groups, routes, and Kubernetes deployments in

    < 12 seconds (total recovery duration: < 15 seconds).

    4. Game Day Cadence & Operational Post-Mortem [MC-GD-01]
    • Cadence: Bi-weekly game days in pre-production staging; quarterly dark-launch canary runs in production.
    • Roles:
      • Chaos Commander: Drives injection console and monitors tripwires.
      • SRE Observer: Triage alerts without prior knowledge of specific fault to test runbooks.

    Action Items: Any unexpected cascade triggers an immediate JIRA bug labeled CHAOS-DEFECT with mandatory 14-day remediation SLA.


    Invariants and Contracts

    Mandatory Automated Abort Tripwire [INV-CHAOS-01]
      Chaos experiments must be coupled to automated metric-driven abort tripwires.
      Running experiments without automated sub-15-second rollback triggers is strictly prohibited.
    
    Strict Blast-Radius Ceiling (Ten Percent) [INV-CHAOS-02]
      Fault injection in production environments must never target more than 1 Availability Zone
      or more than 10% of active service replicas simultaneously.
    
    Pre-Experiment Steady-State Gate [INV-CHAOS-03]
      An experiment must not initiate unless the system has demonstrated clean steady-state metrics
      for at least 15 continuous minutes prior to launch.
    

    Explicit Unknowns

    • AWS FIS API call rate limit throttling during emergency abort sequences across 40 concurrent faults (G-1).
    • Cross-region Aurora PostgreSQL read-replica replication lag behavior when primary node CPU is artificially spiked to 100% (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    18,000 authorizations/sec across 3 AZsprovidedTraffic intakeCurrent
    Latency SLA p99 <= 35 ms, error <= 0.05%providedCore Banking SLA contractCurrent
    Incident INC-4939 $1.8M checkout cascadeprovidedPost-mortem incident recordHistorical
    Four controlled fault injection scenariosdecidedDavid O'Reilly & Elena Rostova2026-09-15
    Sub-15-second automated abort tripwiredecidedArchitectural invariant INV-CHAOS-012026-09-15
    10% maximum canary blast-radius ceilingdecidedArchitectural invariant INV-CHAOS-022026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against chaos testing standards:

    • Steady-State Rigor: PASS. Baseline SLIs defined with explicit PromQL queries and thresholds.
    • Blast-Radius Safety: PASS. Capped at 1 AZ and 10% canary traffic; zero full-estate injection.
    • Abort Velocity: PASS. Sub-15s automated rollback ensures immediate recovery upon SLA breach.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-CHAOS-01: David O'Reilly to determine whether automated chaos testing should be integrated directly into post-deployment canary release pipelines (Continuous Chaos) (Owner: David O'Reilly).

    Next steps

    1. Marcus Vance provisions AWS Fault Injection Simulator (FIS) experiment templates via Terraform.
    2. Platform team configures Prometheus continuous tripwire webhook integration.
    3. Conduct staging game day executing EXP-CHAOS-01 (AZ network partition) to verify sub-15s automated abort tripwire.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Plan multi-experiment game days for distributed systemsDefine automated abort triggers and safety gates for fault injectionSequence resilience tests based on risk and failure modelsEstablish evidence custody and execution readiness standards

    About this skill

    What it does

    This skill converts accepted resilience uncertainties into a prioritized, governed portfolio of controlled experiments. It defines coverage, sequencing, readiness, safety, evidence quality and learning gates; individual experiment mechanics remain with chaos-engineering-task.

    Use it when

    Use when multiple resilience hypotheses/failure models require a coherent chaos testing program rather than one experiment.

    For example: “Industrial IoT telemetry ingestion drops sensor messages whenever edge MQTT bridges reconnect. Operations wants to run fault injection on the Kafka broker cluster during peak shift.”

    What you get

    • Chaos Game Day Plan

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/chaos-testing-strategy/.

    What it will not do

    Do not use for one chaos experiment, resilience architecture, load/performance tests, DR/game-day command, incident response or chaos-tool setup.

    How it works

    1. Check steady-state authority exists.
    2. Bound the experiment portfolio.
    3. Establish safety prerequisites and abort triggers.
    4. Enforce target and environment isolation.
    5. Define execution readiness and evidence custody.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions