- Home
- Skills
- Testing & Debugging
- Chaos Experiment and Resilience Testing Strategy
Chaos Experiment and Resilience Testing Strategy
Plans chaos experiments: steady-state hypotheses, fault injection blast radius, automated abort tripwires, and game days.
$5
Works with the AI tools you already use
Chaos Experiment and Resilience Testing Strategy
Example session with this skill installed
Design the chaos engineering experiment portfolio, blast-radius containment, and automated abort tripwires for our Payment Authorization Engine.
- Subject: payment-authorization-engine v4.0 (run-id: banking-chaos-001) processing 18,000 authorizations/sec across 3 AWS AZs.
- Deciders: Lead SRE Architect (David O'Reilly) and Head of Payment Infrastructure (Elena Rostova).
- Scope: Verifying resilience under cross-AZ network partitions, Redis replica crashes, and 250ms WAN latency spikes; p99 latency <= 35 ms; error rate <= 0.05%.
- Tension: Squads proposed unconstrained chaos in production without automated rollback triggers or steady-state baselines to "test reality"; David O'Reilly and Elena Rostova strictly reject untracked live chaos, citing incident INC-4939 where an unmonitored pod-kill drill crashed database connection pools, dropping $1.8M in checkouts. They mandate a formal chaos testing strategy: steady-state metrics, synthetic traffic canaries, bounded blast radii (max 1 AZ, max 10% traffic), sub-15-second automated abort tripwires, and pre-scheduled staging game days.
Write the chaos testing strategy specification under docs/.
- Read your context and instructions
- Compiled the chaos experiment
Wrote docs/architecture/tasks/banking-chaos-001/chaos-testing-strategy/chaos-testing-strategy.md. Complete chaos testing strategy establishing steady-state hypotheses, fault injection portfolios across AZs, sub-15s automated abort tripwires, and game day governance.
skill: chaos-testing-strategy
Chaos-Testing Strategy: Payment Authorization Engine [CHAOS-BANK-001]
Summary
This specification establishes the chaos engineering experiment portfolio, steady-state hypothesis models, blast-radius containment rules, and automated abort tripwires for payment-authorization-engine v4.0 under run ID banking-chaos-001. It governs 18,000 peak authorizations/second distributed across 3 AWS Availability Zones (us-east-1a, us-east-1b, us-east-1c). It decisively eliminates the uncontrolled cascade risks demonstrated in incident INC-4939 (where an unconstrained chaos test killed connection pools and halted $1.8M in checkouts). The strategy enforces measurable steady-state metric baselines (p99 latency <= 35 ms, error rate <= 0.05%), four controlled fault injection scenarios (network partition, pod termination, latency injection, database primary failover), tightly bounded blast-radius limits (<= 10% synthetic traffic canary, single AZ only), and sub-15-second automated abort tripwires.
Detailed Description
Injecting infrastructure faults without steady-state baselines and automated circuit breakers turns resilience engineering into uncontrolled self-inflicted outages. When systems experience unexpected dependency degradation—such as a cloud availability zone network partition or silent cache loss—untested retry storms and connection pool exhaustion trigger systemic failures. A formal chaos testing strategy proves fault tolerance under controlled, observable, and reversible experimental conditions.
Chaos Experiment Invocation (Experiment: `EXP-NET-PARTITION-AZ1`)
│
▼
[ Chaos Controller & Blast-Radius Governor ]
├── 1. Asserts Steady-State Baseline: p99 <= 35ms, Errors <= 0.05%
├── 2. Isolates Blast Radius: Traffic routed to Synthetic Canary (10% max)
└── 3. Injects Fault via AWS Fault Injection Simulator (FIS)
│
▼ (Cross-AZ Link severed for 180s)
[ Availability Zone `us-east-1a`: Network Partition Simulation ]
├── Envoy Circuit Breaker trips in < 150 ms
└── Traffic shifts automatically to AZ-b and AZ-c
│
┌────────────────────┴────────────────────┐
▼ (Steady-State Maintained) ▼ (Error Rate > 0.05% or Latency > 50ms)
[ Hypothesis Validated: Pass ] [ AUTOMATED ABORT TRIPWIRE FIRES (< 15s) ]
Capture Metric Dashboard Artifacts AWS FIS Aborts; Traffic Restored Instantly
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Blast-Radius Containment (Zero User Loss) | Chaos injection must never impact real cardholder checkout transactions (INC-4939). | 0.40 | David O'Reilly (Lead SRE Architect) |
| Automated Abort Tripwire Velocity (< 15s) | If SLA metrics degrade, the experiment must terminate and recover in under 15 seconds. | 0.30 | Elena Rostova (Head of Payment Infra) |
| Hypothesis Measurability & Observability | Every experiment must evaluate specific Prometheus SLI queries against golden baselines. | 0.20 | Enterprise Reliability Engineering |
| Production Parity of Injected Faults | Injected failure modes must accurately reflect real AWS cloud outages (partitions, panics). | 0.10 | Chaos Engineering Guild Standard |
Comparison
| Chaos Testing Operating Model | Fault Target Seam | Blast Radius Control | Abort Mechanism | Evaluation |
|---|---|---|---|---|
| Option A: Live Ad-Hoc Chaos in Prod | Direct production nodes | None (100% live traffic) | Manual operator kill | Rejected: Caused INC-4939 $1.8M outage; reckless and unmonitored. |
| Option B: Isolated Dev Environment Only | Staging mock clusters | Total isolation | Manual teardown | Rejected: Dev lacks high-concurrency traffic and real multi-AZ topology. |
| Option C: Canary-Bounded FIS Chaos (Chosen) | Ephemeral canary (10% load) | Strictly bounded (1 AZ) | Automated sub-15s Prometheus tripwire | Selected: Real-world fidelity, zero customer impact, instant recovery. |
Result
Option C is selected. AWS FIS executes controlled fault injection; Prometheus tripwires abort experiments automatically; synthetic canary traffic contains customer risk.
Required Mechanisms
1. Steady-State Hypothesis & Metrics Verification [MC-SS-01]
- Steady-State Definition: Prior to fault injection, the system must maintain:
http_request_duration_seconds{quantile="0.99"} <= 0.035(35 ms).rate(http_requests_total{status=~"5.."}[1m]) / rate(http_requests_total[1m]) <= 0.0005(0.05%).payment_engine_active_connection_pool_saturation <= 0.70(70%).
Hypothesis: Severing connectivity to one AZ or killing 30% of Redis replicas will cause p99 latency to rise by <= 8 ms with zero increase in 5xx errors, as traffic seamlessly drains to surviving nodes.
2. Fault Injection Portfolio [MC-FP-01]
| Experiment ID | Injected Fault Action | Target Infrastructure | Duration | Expected Resilient Behavior |
|---|---|---|---|---|
EXP-CHAOS-01 | Sever Cross-AZ Peering Route | AWS TGW / Transit Subnet | 180 s | Envoy reroutes to AZ-b/c in < 200 ms; 0 dropped payments. |
EXP-CHAOS-02 | SIGKILL 50% Consumer Pods | Kubernetes Deployment | 60 s | K8s replaces pods in < 15 s; surviving pods absorb load. |
EXP-CHAOS-03 | Inject 250ms Latency on DB Pool | Aurora PostgreSQL Primary | 120 s | App circuit breaker trips; serves graceful degraded response. |
EXP-CHAOS-04 | Blackhole Redis Cache Shards | Redis Cluster Node 02 | 180 s | App falls back to read-through DB replicas with < 5 ms penalty. |
3. Automated Abort Tripwires & Rollback Protocol [MC-AT-01]
- Prometheus Continuous Tripwire Daemon:
- Evaluates SLIs every 2 seconds during active experiments.
- Tripwire Abort Triggers:
ErrorRate > 0.05%for $> 6$ seconds.p99_Latency > 50msfor $> 6$ seconds.- Any customer transaction timeout reported by edge API Gateway.
- Abort Action:
- Triggers AWS FIS
StopExperimentAPI in < 3 seconds. - Restores all security groups, routes, and Kubernetes deployments in
- Triggers AWS FIS
< 12 seconds (total recovery duration: < 15 seconds).
4. Game Day Cadence & Operational Post-Mortem [MC-GD-01]
- Cadence: Bi-weekly game days in pre-production staging; quarterly dark-launch canary runs in production.
- Roles:
- Chaos Commander: Drives injection console and monitors tripwires.
- SRE Observer: Triage alerts without prior knowledge of specific fault to test runbooks.
Action Items: Any unexpected cascade triggers an immediate JIRA bug labeled CHAOS-DEFECT with mandatory 14-day remediation SLA.
Invariants and Contracts
Mandatory Automated Abort Tripwire [INV-CHAOS-01]
Chaos experiments must be coupled to automated metric-driven abort tripwires.
Running experiments without automated sub-15-second rollback triggers is strictly prohibited.
Strict Blast-Radius Ceiling (Ten Percent) [INV-CHAOS-02]
Fault injection in production environments must never target more than 1 Availability Zone
or more than 10% of active service replicas simultaneously.
Pre-Experiment Steady-State Gate [INV-CHAOS-03]
An experiment must not initiate unless the system has demonstrated clean steady-state metrics
for at least 15 continuous minutes prior to launch.
Explicit Unknowns
- AWS FIS API call rate limit throttling during emergency abort sequences across 40 concurrent faults (G-1).
- Cross-region Aurora PostgreSQL read-replica replication lag behavior when primary node CPU is artificially spiked to 100% (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 18,000 authorizations/sec across 3 AZs | provided | Traffic intake | Current |
| Latency SLA p99 <= 35 ms, error <= 0.05% | provided | Core Banking SLA contract | Current |
| Incident INC-4939 $1.8M checkout cascade | provided | Post-mortem incident record | Historical |
| Four controlled fault injection scenarios | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Sub-15-second automated abort tripwire | decided | Architectural invariant INV-CHAOS-01 | 2026-09-15 |
| 10% maximum canary blast-radius ceiling | decided | Architectural invariant INV-CHAOS-02 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against chaos testing standards:
- Steady-State Rigor: PASS. Baseline SLIs defined with explicit PromQL queries and thresholds.
- Blast-Radius Safety: PASS. Capped at 1 AZ and 10% canary traffic; zero full-estate injection.
- Abort Velocity: PASS. Sub-15s automated rollback ensures immediate recovery upon SLA breach.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-CHAOS-01: David O'Reilly to determine whether automated chaos testing should be integrated directly into post-deployment canary release pipelines (Continuous Chaos) (Owner: David O'Reilly).
Next steps
- Marcus Vance provisions AWS Fault Injection Simulator (FIS) experiment templates via Terraform.
- Platform team configures Prometheus continuous tripwire webhook integration.
- Conduct staging game day executing
EXP-CHAOS-01(AZ network partition) to verify sub-15s automated abort tripwire.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill converts accepted resilience uncertainties into a prioritized, governed portfolio of controlled experiments. It defines coverage, sequencing, readiness, safety, evidence quality and learning gates; individual experiment mechanics remain with chaos-engineering-task.
Use it when
Use when multiple resilience hypotheses/failure models require a coherent chaos testing program rather than one experiment.
For example: “Industrial IoT telemetry ingestion drops sensor messages whenever edge MQTT bridges reconnect. Operations wants to run fault injection on the Kafka broker cluster during peak shift.”
What you get
- Chaos Game Day Plan
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/chaos-testing-strategy/.
What it will not do
Do not use for one chaos experiment, resilience architecture, load/performance tests, DR/game-day command, incident response or chaos-tool setup.
How it works
- Check steady-state authority exists.
- Bound the experiment portfolio.
- Establish safety prerequisites and abort triggers.
- Enforce target and environment isolation.
- Define execution readiness and evidence custody.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install