- Home
- Skills
- APIs & Backend
- Controlled Chaos Experiment Design
Controlled Chaos Experiment Design
Designs safe chaos experiments: steady-state metrics, fault injection, blast-radius bounds, and automated abort triggers.
$5
Works with the AI tools you already use
Controlled Chaos Experiment Design
Example session with this skill installed
Design the controlled chaos experiment plan for testing our Payment Tokenization Redis Cluster under network partition.
- Subject: payment-token-service v2.4 (run-id: redis-partition-exp-001) connected to a 6-node Redis cluster (3 primaries, 3 replicas in AWS us-east-1).
- Deciders: SRE Reliability Lead (Marcus Vance) and Security Architecture Lead (Sarah Chen).
- Hypothesis: When primary Redis writer experiences a 60-second network partition (packet drop 100%), the service degrades to local in-memory fallback tokenization within 3 seconds, sustaining payment authorization p95 latency <= 180 ms with error rate < 0.2% across 800 peak transactions/sec.
- Tension: QA team proposed pulling network interfaces in production without automated abort triggers; Marcus Vance strictly forbids unmonitored production testing, mandating execution in a dedicated staging clone with automated abort tripwires (halt experiment if error rate exceeds 1.0% for > 5 seconds), a 5% canary traffic blast radius, and automated network route healing rollback scripts.
- Tooling: Chaos Mesh / AWS Fault Injection Simulator (FIS); 15-minute total experiment window.
Write the chaos experiment plan under docs/.
- Read your context and instructions
- Compiled the controlled chaos experiment
- Generated the UI component
Wrote docs/architecture/tasks/redis-partition-exp-001/chaos-engineering-task/chaos-experiment-plan.md. Complete controlled chaos engineering plan specifying steady-state metrics, 60-second Redis partition injection, automated 1.0% error rate abort tripwires, and route healing rollback scripts.
---
skill: chaos-engineering-task
---
# Chaos Experiment Plan: Payment Redis Cluster Network Partition [CHAOS-PAY-001]
## Summary
This specification defines the controlled chaos engineering experiment for `payment-token-service v2.4` under run ID `redis-partition-exp-001`. The experiment tests the architectural hypothesis that a 60-second complete network partition on the primary Redis writer cluster triggers automatic graceful degradation to local fallback tokenization within 3.0 seconds, maintaining payment authorization latency p95 <= 180 ms and error rates < 0.2% at 800 TPS. It decisively rejects unmonitored production testing, confining the blast radius to a staging replica cluster processing 5% canary traffic. The plan defines hard automated abort gates (halt if error rate >= 1.0% for > 5 seconds), automated traffic shedding, and deterministic network route healing rollbacks.
## Detailed Description
Resilience claims regarding distributed datastores remain speculative until verified through empirical fault injection. In production, Redis primary failovers frequently trigger client connection storm cascades, socket leaks, and cascading HTTP 504 gateway timeouts.
Canary Traffic (800 TPS, 5% Slice)
│
▼
[ payment-token-service v2.4 ]
│
(Injected Fault: 60s AWS FIS Network Partition)
│
▼
[ Redis Primary Writer ] (100% Packet Drop)
│
├─► [ Health Oracle: Lag > 3s ]
▼
[ Fallback Tokenization Engine ] ──► (Sustains p95 <= 180 ms, Errors < 0.2%)
│
(Safety Monitor: If Errors >= 1.0% for 5s ──► TRIGGER AUTO-ABORT & HEAL)
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Safety Blast Radius Containment | Faults must not escape to live customer payments or cross-tenant workloads. | 0.35 | Sarah Chen (Security Lead) |
| Steady-State Hypothesis Precision | Clear, quantitative criteria separating normal operation from degraded failure. | 0.30 | Marcus Vance (SRE Lead) |
| Automated Abort Reaction Time | Automated monitors must terminate fault injection within 5 seconds of an abort condition. | 0.20 | Reliability Engineering Standard |
| Non-Destructive Rollback Healing | Cluster must restore healthy primary-replica replication without operator data salvage. | 0.15 | Disaster Recovery Mandate |
### Comparison
| Candidate Approach | Target Environment | Blast Radius Control | Abort Mechanism | Residual Risk |
|---|---|---|---|---|
| Option A: Live Production Network Pull | Production environment | 100% active customer traffic | Manual operator terminal kill | Critical: Disastrous risk of customer checkout failure during live hours. |
| Option B: Synthetic Local Unit Mock | In-memory mock socket | Isolated developer workstation | Test process exit | High: Fails to reproduce real Linux kernel TCP timeout and socket buffer stalls. |
| Option C: Staged FIS Canary Partition (Chosen) | Staging environment with 5% canary | Dedicated clone; synthetic 800 TPS | Automated Prometheus webhook tripwire | Minimal: Real kernel TCP fault injection with automated 5s killswitch. |
### Result
Option C is selected. AWS Fault Injection Simulator (FIS) executes packet drops in staging against a live 6-node Redis cluster with automated webhook abort monitors.
---
### Required Mechanisms
#### 1. Steady-State Hypothesis [MC-SS-01]
- **Target Workload**: Synthetic load generator pumping 800 transactions/sec matching production payload mix.
- **Normal Steady-State Baseline**:
- `http_requests_success_ratio`: >= 99.95% (HTTP 200/201).
- `p95_authorization_latency`: <= 110 ms.
- `redis_connected_clients`: Stable between 40 and 60.
- **Hypothesis**: Under a 60-second complete network partition of the primary Redis node:
1. Circuit breaker trips within 3.0 seconds of partition onset.
2. Service falls back to local in-memory vault tokens.
3. `http_requests_success_ratio` remains >= 99.80% (error rate < 0.2%).
4. `p95_authorization_latency` remains <= 180 ms throughout the partition.
#### 2. Fault Injection Specification & Blast Radius [MC-FI-01]
- **Tooling**: AWS Fault Injection Simulator (FIS) Action `aws:network:disrupt-connectivity`.
- **Target**: Single EC2/ElastiCache node hosting Redis Primary Writer (`redis-pay-primary-01`).
- **Fault Type**: 100% packet drop on ingress and egress TCP traffic (Port 6379) for exactly 60 seconds.
- **Blast Radius Boundaries**:
- Environment strictly bounded to VPC `vpc-stage-resilience-01`.
- Target cluster tag: `Environment=staging-chaos`. Production traffic routing is physically disjoint.
#### 3. Automated Abort Gates & Safety Tripwires [MC-AG-01]
Continuous Prometheus monitoring queries evaluate every 1.0 second during the test.
- **Abort Condition 1 (Error Spike)**:
```promql
sum(rate(http_requests_total{status=~"5.."}[10s])) / sum(rate(http_requests_total[10s])) >= 0.01
If true for 5 consecutive seconds -> IMMEDIATELY ABORT.
- Abort Condition 2 (Latency Collapse):
If p95 latency exceeds 500 ms for 3 consecutive seconds -> IMMEDIATELY ABORT.histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[10s])) >= 0.500 - Abort Action: AWS FIS webhook triggers
StopExperiment, restoring network security group rules within 1,500 ms.
4. Execution Sequence & Rollback Procedures [MC-EX-01]
T-00:00 to T+05:00: Baseline Verification
- Verify 800 TPS steady-state load. Confirm 0 active alerts.
T+05:00: Fault Injection Onset
- AWS FIS triggers 100% packet disruption on
redis-pay-primary-01. - Verify Redis Sentinel / cluster bus detects primary failure within 3,000 ms.
T+06:00: Fault Termination & Recovery Observation
- AWS FIS terminates disruption. Network connectivity re-established.
- Observe replica promotion, socket pool reconnection, and memory cache repopulation.
- Confirm cluster returns to steady state within 45 seconds of fault removal.
Rollback Script (Manual Override)
If automated abort fails:
aws fis stop-experiment --id exp-0a9b8c7d6e5f4g3h2
aws ec2 authorize-security-group-ingress --group-id sg-redis-chaos --protocol tcp --port 6379 --cidr 10.10.0.0/16
Invariants and Contracts
Automated Abort Tripwire Invariant [INV-CHS-01]
Every chaos experiment must execute with automated webhook abort monitoring active.
If error rates exceed 1.0% or latency exceeds 500 ms for > 5 seconds, the experiment
must self-terminate without human intervention.
Production Physical Disjointness [INV-CHS-02]
Chaos experiments involving network partition or process termination must never target
shared production VPCs, shared multi-tenant Redis clusters, or live payment credentials.
Blast Radius Time Ceiling [INV-CHS-03]
Single-node fault injection durations must not exceed 60 seconds; total active test
observation windows must not exceed 15 minutes.
Explicit Unknowns
- Behavior of legacy Java Lettuce Redis driver connection pools when reconnecting after asymmetric half-open TCP states (G-1).
- CloudWatch metric ingest jitter variance during simultaneous AWS hypervisor stress (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 6-node Redis cluster in us-east-1 | provided | Infrastructure intake | Current |
| 800 TPS load profile | provided | Traffic intake | Current |
| Steady-state latency p95 <= 180 ms | decided | Marcus Vance (SRE Lead) | 2026-09-15 |
| Error rate ceiling < 0.2% | decided | Marcus Vance & Sarah Chen | 2026-09-15 |
| Rejection of live production chaos | decided | Sarah Chen (Security Lead) | 2026-09-15 |
| Automated abort at >= 1.0% error rate | decided | Architectural invariant INV-CHS-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against chaos engineering contracts:
- Hypothesis Formulation: PASS. Exact quantifiable metrics (800 TPS, p95 <= 180 ms, errors < 0.2%).
- Blast Radius Isolation: PASS. Confined to staging VPC with 5% canary slice; 60s hard duration cap.
- Safety Tripwires: PASS. Automated 1.0% error and 500 ms latency abort triggers specified.
- Rollback Completeness: PASS. AWS FIS automated cessation and emergency AWS CLI override documented.
Open Decisions
DEC-CHS-01: Marcus Vance to determine whether Chaos Mesh or native AWS FIS should be the standard orchestrator for upcoming Kubernetes node-drain drills (Owner: Marcus Vance).
Next steps
- Marcus Vance provisions AWS FIS experiment template
fis-redis-network-disrupt.jsonin staging. - SRE team configures Prometheus webhook alert manager endpoint to trigger FIS abort API.
- Conduct dry-run game day review with the on-call payment engineering team before execution.
controlled-chaos-experiment-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps one accepted resilience uncertainty into a falsifiable, bounded and reversible experiment. It defines the control/baseline, target and fault, steady-state oracle, safety envelope, abort path, observations and post-fault recovery evidence independently of a chaos tool.
Use it when
Use when a specific resilience hypothesis can be tested through an authorized controlled perturbation with bounded harm and a named oracle.
For example: “We claim our payment service gracefully handles primary database failover, but we have never tested if checkout requests drop or hang during an actual database leader election.”
What you get
- Chaos Experiment Plan
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/chaos-engineering-task/.
What it will not do
Do not use for reliability architecture, load/performance testing, DR game-day orchestration, security testing, incident response, tool selection/configuration or random failure.
How it works
- Check resilience hypothesis testing is required.
- Define baseline steady-state metrics and oracles.
- Bound the fault injection target and cohort.
- Establish explicit safety guardrails and abort triggers.
- Verify fault rollback and system recovery.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install