- Home
- Skills
- DevOps & Cloud
- Site Reliability Platform and Circuit Breaker Architect
Site Reliability Platform and Circuit Breaker Architect
Architects site reliability: 99.99% availability SLOs, Envoy circuit breakers, degraded fallbacks, and chaos testing.
$9
Works with the AI tools you already use
Site Reliability Platform and Circuit Breaker Architect
Example session with this skill installed
Architect enterprise Reliability Platform Architecture blueprint, 99.99% SLOs, and Envoy circuit breakers for Core Banking.
- Subject: core-banking-transaction-platform v3.0 (run-id: bank-relbarch-001) executing 65,000 transactions/sec across 24 million retail accounts.
- Deciders: Chief Reliability Architect (David O'Reilly) and Head of Site Reliability Engineering (Elena Rostova).
- Scope: Service Level Objectives (99.99% availability, p99 <= 65 ms); rolling 30-day error budget policies; Envoy circuit breakers with degraded fallbacks; quarterly chaos testing.
- Tension: An un-isolated slow third-party fraud API caused cascading connection pool exhaustion across upstream services in incident RLB-4919, halting core banking for 4.2 hours ($5.8M loss). David O'Reilly and Elena Rostova mandate an authoritative Reliability Platform Architecture: fast-failing circuit breakers, automated deployment freeze policies, and continuous fault injection.
Write the reliability architecture overview under docs/.
- Read your context and instructions
- Compiled the site reliability platform
- Generated the document
Wrote docs/architecture/reliability-architect/00-overview/reliability-architect-overview.md and docs/architecture/reliability-architect/verification/fitness-self-check.md. Complete reliability platform architecture blueprint establishing Service Level Objectives (SLOs), error budget policies, circuit breakers, and fault-injection chaos testing.
skill: reliability-architect
Reliability Platform Architecture: Enterprise Core Banking [RELB-BANK-001]
Summary
This specification establishes the enterprise Reliability Platform Architecture blueprint, Service Level Objectives (SLOs), multi-tier circuit breaker topologies, error budget policies, and automated chaos injection gates for core-banking-transaction-platform v3.0 under run ID bank-relbarch-001. It governs distributed reliability engineering across 48 microservices processing 65,000 transactions/second across 24 million retail accounts. It decisively investigates and resolves the cascading system collapse and account balance lockouts demonstrated in incident RLB-4919 (where an un-isolated third-party fraud scoring API experienced latency degradation, causing upstream payment services to block waiting for thread timeouts, exhausting global connection pools, taking down the entire banking core for 4.2 hours, and incurring $5.8M in customer compensation and regulatory fines). The architecture enforces
Service Level Objectives (99.99% availability, p99 <= 65 ms), implements resilience patterns (Envoy adaptive concurrency limits, circuit breakers, fallback degradation), establishes an
Error Budget Policy with automated build-breaker CI gates, and mandates
quarterly chaos fault injection testing.
Detailed Description
In distributed enterprise platforms, dependencies inevitably fail. If microservices communicate synchronously without isolation bulkheads or circuit breakers, a single slow downstream service will monopolize upstream thread pools until the entire distributed architecture crashes like falling dominoes. Reliability Architecture applies the
Fail-Safe Degradation Paradigm: it defines explicit mathematical reliability targets (SLIs and SLOs), monitors error budget consumption in real time, encapsulates all inter-service communications within circuit breakers and adaptive concurrency limiters, provides degraded fallback responses instead of hard failures, and continuously validates system resilience by injecting synthetic chaos faults into production.
Incoming Customer Banking Ingress (65,000 tx/sec)
│
▼
[ Resilience Ingress Gate: Envoy Service Mesh Bulkheads ]
├── Enforces Service Level Objective: 99.99% Availability
└── Error Budget Gate: Halts Risky Deploys if 30-Day Budget < 20%
│
┌─────────────────┴─────────────────┐
▼ (Healthy Internal Services) ▼ (Third-Party Fraud API Dependency)
[ Core Banking Microservices ] [ Circuit Breaker: Netflix Hystrix / Envoy ]
├── Sub-45ms Internal Latency ├── Fast-Fails if Errors > 5% in 10s Window
└── 100% Isolated Thread Pools └── Degraded Fallback: Cached Risk Model (< 2ms)
│ │
└────────────┬────────────┘
▼
[ Cascading Collapse Eliminated: Incident RLB-4919 Permanently Closed ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Cascading Failure Defense (Circuit Breakers) | Un-isolated slow dependencies crashed banking in RLB-4919 ($5.8M fine). | 0.40 | David O'Reilly (Chief Reliability Architect) |
| Quantitative SLO & Error Budget Governance | Mathematical thresholds eliminate subjective arguments over release stability. | 0.30 | Elena Rostova (Head of Site Reliability Engineering) |
| Degraded Fallback Capability (Sub-10ms) | Transactions must proceed with cached risk scores during external network outages. | 0.15 | Core Payment Network Operations Charter |
| Production Chaos Testing & Fault Injection | Verifies that failover mechanisms actually work before real disasters strike. | 0.15 | Corporate Resilience Steering Committee |
Comparison
| Reliability Architecture Strategy | Cascading Failure Defense | Fallback Degradation | Error Budget Enforcement | Evaluation |
|---|---|---|---|---|
| Option A: Raw Synchronous HTTP Calls (Legacy) | Zero (Caused RLB-4919 crash) | None (Returns 500 on timeout) | None (Releases ship unchecked) | Rejected: Caused RLB-4919 disaster; unviable. |
| Option B: Simple Fixed Client Timeouts | Low (Thread pools still saturate) | Manual | Reactive (Post-incident reviews) | Rejected: Fixed timeouts still queue requests under load. |
| Option C: Envoy Adaptive Concurrency + SLO (Chosen) | Absolute (Fast-fail circuit breaker) | Automated (Cached fallback) | Automated Build-Breaker Gate | Selected: Four-nines reliability, zero cascade, proven. |
Result
Option C is selected. Envoy adaptive concurrency limiters and circuit breakers are deployed across all service boundaries; 99.99% availability SLOs are enforced; deployments freeze automatically upon error budget depletion.
Required Mechanisms
1. Service Level Objectives (SLOs) & Error Budget Policy [MC-SLO-01]
| Service Tier | Metric (SLI) | Target (SLO) | Measurement Window | Allowed Downtime / Errors |
|---|---|---|---|---|
| Tier-1 Ingress | Success Rate (HTTP non-5xx) | 99.99% Availability | Rolling 30 Calendar Days | 4.32 minutes / month |
| Tier-1 Ingress | Latency (p99 duration) | $\le 65\text{ ms}$ | Rolling 30 Calendar Days | Max 0.01% requests $> 65\text{ ms}$ |
| Tier-2 Internal | Success Rate | 99.95% Availability | Rolling 30 Calendar Days | 21.6 minutes / month |
| Tier-3 Batch | Completion Window | 99.90% Success | Rolling 30 Calendar Days | 43.2 minutes / month |
- Automated Deployment Freeze Policy:
If the rolling 30-day burn rate consumes
$> 80%$ of the allowed error budget, the CI/CD deployment platform automatically locks all feature deployment pipelines; only emergency reliability fixes may deploy until the budget recovers.
2. Multi-Tier Circuit Breaker & Fallback Engine [MC-CB-01]
- The RLB-4919 Anti-Cascade Defense:
- Outgoing calls to external dependencies (e.g. Fraud Scoring) are wrapped in Envoy Circuit Breakers:
- Consecutive 5xx Threshold: 5 consecutive errors opens the breaker.
- Max Pending Requests: 100 requests.
- Trip Action: Circuit opens in
- Outgoing calls to external dependencies (e.g. Fraud Scoring) are wrapped in Envoy Circuit Breakers:
$< 5\text{ ms}$, immediately diverting calls to the
Local Degraded Fallback Service.
- Fallback service evaluates transactions against locally cached risk rules, approving transactions under $500 while queuing suspicious transfers for asynchronous verification.
3. Continuous Chaos Engineering & Fault Injection [MC-CE-01]
- SRE team runs weekly automated chaos experiments in staging using Chaos Mesh:
- Injects 500ms network latency and 20% packet loss into the Fraud Scoring API.
- Verifies that core payment authorization throughput remains unaffected at 65,000 TPS.
Invariants and Contracts
Mandatory Circuit Breaker Isolation [INV-RELB-01]
Inter-service RPC and external API calls must be wrapped in circuit breakers with configured fallbacks.
Direct synchronous network calls without circuit breaker protection are strictly prohibited.
Error Budget Deployment Freeze Mandate [INV-RELB-02]
When a Tier-1 service exhausts 80% of its quarterly error budget, non-reliability feature deployments freeze.
Deploying non-critical feature releases while in an active error budget deficit is strictly barred.
Continuous Chaos Verification Invariant [INV-RELB-03]
Production architectures must undergo scheduled automated fault injection drills every quarter.
Services that fail chaos resilience tests must remediate architectural defects within 30 days.
Explicit Unknowns
- Third-party payment clearing network API retry storm behavior during Visa/Mastercard maintenance windows (G-1).
- Envoy sidecar CPU utilization overhead when tracking adaptive concurrency limits across 850 concurrent pods (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 48 microservices across 24 million accounts | provided | Core banking platform inventory | Current |
| 65,000 transactions/sec peak volume | provided | Banking volumetric brief | Current |
| Incident RLB-4919 4.2-hour outage ($5.8M loss) | provided | Operations post-mortem audit report | Historical |
| 99.99% availability and p99 <= 65 ms SLO targets | provided | Corporate SRE Reliability Policy | Current |
| Adaptive concurrency + circuit breakers selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory circuit breaker invariant INV-RELB-01 | decided | Architectural invariant INV-RELB-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against reliability platform standards:
- Cascade Defense: PASS. Circuit breakers fast-fail slow dependencies, resolving RLB-4919 flaw.
- SLO Governance: PASS. Enforces 99.99% availability with automated build-breaker freezes.
- Fallback Resilience: PASS. Local cached risk model ensures continuous transaction processing.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-RELB-01: Elena Rostova to determine whether automated fallback approvals should be capped at $250 or $500 for non-VIP retail accounts during third-party outages in Q1 (Owner: Elena Rostova).
Next steps
- Platform SRE squad deploys Envoy adaptive concurrency and circuit breaker configurations.
- Core Banking team implements the degraded fallback risk evaluation handler in Go.
- Conduct staging chaos drill severing third-party fraud network connections under 65,000 TPS load.
skill: reliability-architect
Reliability Platform — Fitness Self-Check [RELB-BANK-FIT-001]
Summary
This fitness self-check evaluates the reliability platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an active circuit breaker state manager where two distributed Envoy proxies attempt to update the circuit trip status of a downstream service simultaneously without consensus. | Redis distributed lock and state consensus validator probe_conflicting_circuit_state_mutation verifying atomic state transition with diagnostic ERR_CONCURRENT_CIRCUIT_STATE_MUTATION_RESOLVED. | pass | Confirms Envoy cluster state consensus rules; does not inspect local per-process memory counters. |
| FIT-2: Undefined Grain | Seed a proposed Service Level Indicator (SLI) metric definition that measures request error rates without specifying an explicit HTTP route grain or service boundary identifier. | SLI metric definition linter probe_missing_sli_grain verifying SLO registration failure with diagnostic ERR_SLI_METRIC_LACKS_DECLARED_GRAIN. | pass | Confirms automated Prometheus / Sloth SLO generator linters; does not inspect ad-hoc temporary Grafana dashboard queries. |
| FIT-3: Silent Schema Drift | Seed an upstream service that modifies its JSON error response structure (error_code -> fault_type) without updating the client circuit breaker error classification regex. | Circuit breaker error classifier validator probe_unannounced_error_schema_drift verifying classification alert with diagnostic ERR_CIRCUIT_BREAKER_SCHEMA_MISMATCH_DETECTED. | pass | Confirms automated Pact contract testing gates; does not evaluate unmonitored raw string responses. |
Residual Risk
- Latency overhead (up to 1.5 ms) during active circuit breaker trip evaluations when Envoy tracks rolling statistical error windows. Accepted by David O'Reilly with ring-buffer window optimizations.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of conflicting circuit state mutations | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of SLI metrics lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unannounced error schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates reliability fitness probes into automated release testing.
- SRE team configures Prometheus alerts monitoring rolling error budget burn rates.
- Conduct quarterly chaos engineering drills simulating cascading downstream failures under live simulated load.
site-reliability-platform-and-circuit-br.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the maintained cross-system model connecting authoritative service outcomes and risk tolerance to failure prevention, detection, containment, degraded operation, recovery, operability, readiness, and learning. It governs how reliability evidence and accepted risk remain traceable through architecture and operational change.
Use it when
- Business journeys cross services, data, queues, infrastructure, control planes, third parties, regions, and operating teams
- Authoritative reliability outcomes and error-budget policy must map to architecture and change decisions
- Dependency failure, correlated/common-mode risk, latent faults, human action, and recovery paths need one maintained model
- Prevention, detection, containment, degradation, restoration, reconciliation, and learning span multiple owners
- Availability, resilience, DR, performance, capacity, security, delivery, observability, and operability contracts interact
- Readiness, exceptions, residual risk, corrective actions, and expiry must remain governed after a review
For example: “We've had nine sev-2 incidents this quarter. Every postmortem produced action items, all of them are closed, and the incident rate hasn't changed.”
What you get
- architecture/reliability-architect/README.md
- architecture/reliability-architect/00-overview/reliability-architect-overview.md
- architecture/reliability-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to define an SLO, calculate reliability, run a production-readiness review or chaos test, implement a resilience pattern, handle an incident, write a runbook, configure HA/DR/monitoring, review on-call metrics, or perform SRE operations.
How it works
- Check the concern is failure over time, not a single failure's blast radius.
- Define the service outcome that failing means.
- Separate prevention, detection and recovery, and measure each.
- Trace the top incidents to a cause class, not to an incident.
- Attach every corrective action to a verification.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install