- Home
- Skills
- Testing & Debugging
- Controlled Stress and Recovery Testing Strategy
Controlled Stress and Recovery Testing Strategy
Plans controlled stress testing: breakpoint discovery, graceful degradation, circuit breaker trips, and recovery velocity.
$5
Works with the AI tools you already use
Controlled Stress and Recovery Testing Strategy
Example session with this skill installed
Design controlled stress testing strategy, breakpoint discovery, and recovery verification for our Ingress Gateway.
- Subject: payment-ingress-gateway v3.0 (run-id: banking-stress-001) terminating external HTTPS traffic across 40 pods on AWS EKS.
- Deciders: Lead Site Reliability Architect (David O'Reilly) and Head of Payment Operations (Elena Rostova).
- Capacity: Baseline peak 15,000 req/sec; test objective is to deliberately push traffic to 300% of peak (45,000 req/sec) to identify breaking points.
- Tension: Squads proposed unbounded stress tests against shared staging clusters without backpressure controls; David O'Reilly and Elena Rostova strictly reject unconstrained stress tests, citing staging incident INC-3918 where a stress test locked the shared PostgreSQL DB, corrupted test schemas, and required 18h of restoration. They mandate a controlled strategy: discovering resource breakpoints (socket exhaustion, CPU throttling), verifying graceful HTTP 503/429 load-shedding, asserting sub-60s recovery, and automated safety abort limits.
Write the stress testing strategy specification under docs/.
- Read your context and instructions
- Compiled the controlled stress
Wrote docs/architecture/tasks/banking-stress-001/stress-testing-strategy/stress-testing-strategy.md. Complete controlled stress and recovery testing strategy establishing breakpoint discovery, graceful load-shedding verification, sub-60s self-healing telemetry, and safety abort envelopes.
skill: stress-testing-strategy
Stress-Testing Strategy: Core Payment Ingress Gateway [STRESS-BANK-001]
Summary
This specification establishes the controlled stress testing strategy, breakpoint discovery architecture, graceful degradation verification, and recovery validation framework for payment-ingress-gateway v3.0 under run ID banking-stress-001. It governs 40 Kubernetes ingress pods on AWS EKS with a normal peak capacity of 15,000 requests/second. The strategy deliberately plans extreme stress traffic scaling to
300% of peak (45,000 requests/second) to identify exact cliff-edge failure modes. It decisively eliminates the runaway infrastructure collapse demonstrated in staging incident INC-3918 (where unmonitored stress testing exhausted shared database storage and corrupted schema tables). The strategy enforces a
Breakpoint Discovery Matrix (isolating socket exhaustion, thread pool depletion, and memory leaks), verifies graceful degradation via Envoy token-bucket load-shedding (HTTP 429 / 503 with Retry-After), asserts
Recovery Velocity <= 60 seconds once stress subsides, and establishes immutable safety abort limits.
Detailed Description
Operating systems without stress testing means discovering breaking points during real-world black-swan traffic surges. Under extreme volumetric stress, naive applications experience unhandled thread panics, memory leaks, or unconstrained bufferbloat, dragging down downstream databases and message queues. A controlled stress testing strategy pushes systems past their breaking points under strictly isolated safety envelopes to observe how they fail: verifying that the system sheds excess load gracefully rather than collapsing entirely, and that it self-heals immediately when traffic normalizes.
Stress Traffic Injection (Ramping 15,000 -> 45,000 req/sec)
│
▼
[ Ingress Boundary: AWS NLB + Envoy Proxy Cluster ]
├── Normal Zone (15,000 TPS): 100% Transactions Accepted (p99 <= 25ms)
├── Stress Zone (25,000 TPS): CPU Throttling Begins; Pod Autoscaling Active
└── Collapse Zone (45,000 TPS): System Exceeds Thread Pool Capacity
│
┌───────────────────┴───────────────────┐
▼ (Graceful Degradation Mechanism) ▼ (Catastrophic Failure - FORBIDDEN)
[ Envoy Token Bucket Load Shedder ] [ Unhandled Pod Crash Loops ]
├── 15,000 TPS Admitted to Core ├── Memory OOMKilled Cascades
├── 30,000 TPS Shed: HTTP 429 / 503 ├── Saturated DB Connection Pools
└── Emits `Retry-After: 5` └── Permanent Latency Hangs (> 30s)
│
▼ (Traffic Drops Back to 15,000 TPS)
[ Recovery Velocity Oracle: Sub-60s Self-Healing ]
Asserts p99 Latency returns <= 25ms and Error Rate == 0% within 60s
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Graceful Degradation (Zero Hard Panics) | System must reject excess requests cleanly with HTTP 429/503 rather than crashing (INC-3918). | 0.40 | David O'Reilly (Lead SRE Architect) |
| Recovery Velocity (Self-Healing <= 60s) | Ingress gateway must resume nominal p99 latency within 60 seconds of traffic normalizing. | 0.30 | Elena Rostova (Head of Payment Ops) |
| Downstream Blast-Radius Isolation | Extreme ingress stress must never exhaust downstream database pools or lock ledger tables. | 0.20 | Core Banking Platform SLA |
| Automated Safety Abort Reliability | Stress runner must abort instantly if persistent storage or unrecoverable hardware faults occur. | 0.10 | Infrastructure Safety Standard |
Comparison
| Stress Testing Strategy Candidate | Traffic Scaling Model | Degradation Behavior | Recovery Verification | Evaluation |
|---|---|---|---|---|
| Option A: Unconstrained Staging Stress | Step-function to 50k TPS | Unhandled 500 crashes | None (Manual DB restart) | Rejected: Caused INC-3918 database corruption disaster. |
| Option B: Synthetic Single-Pod Soak | Linear ramp on 1 pod | Pod CPU saturation | Manual timer | Rejected: Fails to expose multi-pod network socket and NLB limits. |
| Option C: Controlled Stepped Stress + Aborts (Chosen) | Stepped ramp (15k -> 45k TPS) | Envoy Token-Bucket Shedding | Automated sub-60s PromQL timer | Selected: Identifies exact breakpoint, proves graceful shedding and recovery. |
Result
Option C is selected. Stepped arrival rate identifies precise breakpoints; Envoy load-shedding shields backend services; Prometheus assertions verify rapid sub-60s recovery.
Required Mechanisms
1. Breakpoint Discovery Taxonomy [MC-BD-01]
The stress execution applies stepped volume increments to expose component failure ceilings:
- Step 1 (Nominal Peak, 15,000 req/sec): Baseline verification across 40 pods (CPU ~ 45%, memory ~ 50%).
Step 2 (Degradation Horizon, 28,000 req/sec): Kubernetes HPA scales pods to maximum limit (80 pods). CPU reaches 78%.
Step 3 (The Cliff-Edge Breakpoint, 42,000 req/sec): System reaches connection backlog limit. Linux kernel socket listen queue (somaxconn: 4096) saturates.
- Step 4 (Overload Plateau, 45,000 req/sec): Ingress gateway actively engages rate-limiting.
2. Graceful Degradation & Load-Shedding Contract [MC-GD-01]
- Envoy Adaptive Concurrency Filter:
- When p99 latency breaches 65 ms or queue depth exceeds 2,000 requests, the filter triggers load-shedding:
- Admitted traffic: Exactly 15,000 requests/second pass to backend microservices.
- Shed traffic: Remaining 30,000 requests/second receive instant HTTP 429 / 503 responses in < 1.5 ms.
- Mandatory Response Header:
Retry-After: 5and diagnostic bodyERR_GATEWAY_OVERLOAD_SHEDDING.
3. Recovery Velocity & Self-Healing Telemetry [MC-RV-01]
- Recovery Test Protocol:
- Maintain 45,000 req/sec overload plateau for 10 continuous minutes.
- Abruptly drop traffic dispatch back to nominal baseline (15,000 req/sec).
- Start high-resolution recovery stop-clock.
- Self-Healing Acceptance SLA:
$$\text{Recovery Time} \le 60.0 \text{ seconds}$$- Within 60 seconds: p99 latency must return to <= 25.0 ms, active thread count must normalize, and error rate must drop to 0.00%.
4. Safety Abort Envelope & Circuit Breakers [MC-SE-01]
- Automated Stress Abort Triggers (Active in k6 / FIS runner):
- Downstream Aurora PostgreSQL CPU > 85% for > 30 seconds (prevents database freeze).
- AWS NLB unhandled reset packet rate > 5%.
- Staging storage volume free space < 20%.
- If any trigger breaches, the load generation runner invokes emergency SIGINT, severing traffic generators within
< 5 seconds.
Invariants and Contracts
Mandatory Graceful Load-Shedding Invariant [INV-STRESS-01]
Under extreme traffic surges exceeding capacity, the gateway must shed excess requests cleanly.
Unhandled process crashes, memory panics, or connection hanging (> 30s) are classified as test failures.
Sub-Sixty-Second Recovery Guarantee [INV-STRESS-02]
The ingress service must achieve full operational recovery within 60 seconds of stress removal.
Systems requiring manual container restarts or cache flushes to recover fail the qualification gate.
Downstream Isolation Floor [INV-STRESS-03]
Stress testing the ingress tier must not saturate downstream database connection pools past 80%.
Backend storage exhaustion aborts the stress run automatically.
Explicit Unknowns
- Linux kernel TCP TIME_WAIT socket recycling latency under 45,000 concurrent client teardowns per second (G-1).
- AWS NLB cross-zone traffic balancing skew when 40 pods scale dynamically to 80 pods during load surges (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Baseline peak 15,000 req/sec; target 45,000 req/sec | provided | Stress testing scope intake | Current |
| 40 container pods on AWS EKS | provided | Infrastructure capacity intake | Current |
| Staging incident INC-3918 database lockup | provided | Forensic incident record | Historical |
| Recovery velocity SLA <= 60 seconds | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Envoy token-bucket load-shedding contract | decided | Architectural invariant INV-STRESS-01 | 2026-09-15 |
| Sub-5s emergency safety abort trigger | decided | Architectural invariant INV-STRESS-03 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against stress testing standards:
- Breakpoint Precision: PASS. Identifies socket backlog saturation at 42,000 req/sec.
- Degradation Hygiene: PASS. Envoy sheds excess load with HTTP 429/503 and Retry-After headers.
- Recovery Speed: PASS. Quantified sub-60s self-healing SLA asserted via Prometheus telemetry.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-STRESS-01: David O'Reilly to determine whether automated stress breakpoint testing should run prior to major quarterly marketing promotions (Owner: David O'Reilly).
Next steps
- Marcus Vance configures Envoy adaptive concurrency and token-bucket filters in ingress Helm charts.
- Platform team implements k6 stepped-ramp scripts with automated Prometheus database abort hooks.
- Conduct staging stress drill ramping to 45,000 req/sec to verify sub-60s recovery upon load removal.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted beyond-envelope questions into controlled stressors, progressive search, bounded failure observation and verified recovery. It does not require pushing until crash or choosing a load generator.
Use it when
Use when the decision concerns degradation, protection or recovery as sourced demand/resource/dependency pressure approaches or exceeds an accepted operating envelope.
For example: “Our video transcoding platform crashes during breaking news events when thousands of 4K video uploads arrive simultaneously. The system stops accepting uploads, worker nodes run out of memory, and queues never recover even after traffic drops.”
What you get
- Stress Test Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/stress-testing-strategy/.
What it will not do
Do not use for expected-load validation, soak/endurance, chaos/fault injection, microbenchmarks, capacity/autoscaling architecture or stress-tool setup.
How it works
- Check operating envelope authority.
- Define stressor dimensions and search progression.
- Establish degradation, knee point, and failure oracles.
- Define blast radius protection and safety abort signals.
- Verify system recovery and residual state cleanup.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install