High Availability Platform and Active-Active Architect

    1

    Architects high-availability platforms: 3-AZ active-active meshes, sub-3s ARC evacuation, and error budgets.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    High Availability Platform and Active-Active Architect

    Example session with this skill installed

    Architect enterprise High Availability Platform Architecture blueprint and 3-AZ active-active mesh for Core Payment Rails.

    • Subject: core-payment-clearing-rail v3.0 (run-id: pay-avalarch-001) processing 48,000 transactions/second across $110B in annual volume.
    • Deciders: Chief Availability Architect (David O'Reilly) and Head of Payment Reliability (Elena Rostova).
    • Scope: 3-AZ Active-Active topology on AWS; 99.999% availability SLA (< 5.26 minutes downtime/year); automated Route53 ARC evacuation (< 3s); bulkhead fault domains.
    • Tension: Active-passive failover with manual DNS swing stalled payment traffic during an AZ fiber cut (incident AVL-4919), extending downtime to 42 minutes, dropping 2.1 million transactions, and drawing $3.8M in penalties. David O'Reilly and Elena Rostova mandate an authoritative High Availability Architecture: stateless multi-AZ active-active deployments, automated zonal evacuation, and error budget governance.

    Write the availability architecture overview under docs/.

    • Read your context and instructions
    • Compiled the high availability platform
    • Generated the document

    Wrote docs/architecture/availability-architect/00-overview/availability-architect-overview.md and docs/architecture/availability-architect/verification/fitness-self-check.md. Complete high-availability platform architecture blueprint establishing multi-AZ active-active topologies, error budgets, fault domains, and automated health swings.


    skill: availability-architect

    High Availability Platform Architecture: Core Payment Rails [AVALARCH-PAY-001]

    Summary

    This specification establishes the enterprise High Availability Platform Architecture blueprint, multi-AZ active-active deployment topology, error budget governance, and automated health-swing mechanisms for core-payment-clearing-rail v3.0 under run ID pay-avalarch-001. It governs high-availability platform engineering across 32 payment microservices processing 48,000 transactions/second across $110B in annual payment clearing volume. It decisively investigates and resolves the availability collapse demonstrated in incident AVL-4919 (where deploying an active-passive failover topology with manual DNS swing stalled payment traffic during an AWS Availability Zone fiber cut, extending downtime to 42 minutes, dropping 2.1 million transactions, and incurring $3.8M in merchant SLA penalty payouts). The architecture enforces a stateless Multi-AZ Active-Active topology across three AWS Availability Zones, establishes 99.999% availability SLAs (maximum 5.26 minutes annual downtime), implements automated Route53 Application Recovery Controller (ARC) traffic shifting, and mandates

    bulkhead fault domain isolation.

    Detailed Description

    Relying on active-passive disaster recovery configurations for mission-critical payment services guarantees unacceptable downtime during infrastructure failures. When an active-passive system fails, cold standby nodes require minutes to warm caches, database replicas lag behind, and manual DNS changes take hours to propagate across global DNS caches. High Availability Architecture designs

    Active-Active Multi-Zone Systems: traffic is distributed concurrently across independent, fault-isolated Availability Zones sharing zero common compute failure domains. Health probes continuously evaluate deep synthetic transactions; if an individual zone degrades, traffic is automatically evacuated in sub-second timeframes without dropping in-flight sessions.

    Incoming Customer Payment Ingress (48,000 tx/sec)
                             │
                             ▼
    [ Global Edge Router: AWS Route53 Application Recovery Controller (ARC) ]
      ├── Continuous Deep Synthetic Health Check Probes
      └── Sub-5s Automated Zonal Traffic Evacuation (ARC Routing Controls)
                             │
           ┌─────────────────┼─────────────────┐
           ▼ (Active AZ-1: 33%)▼ (Active AZ-2: 33%)▼ (Active AZ-3: 33%)
    ┌─────────────────┐   ┌─────────────────┐   ┌─────────────────┐
    │ Availability Z1 │   │ Availability Z2 │   │ Availability Z3 │
    │  - EKS Pod Pool │   │  - EKS Pod Pool │   │  - EKS Pod Pool │
    │  - Aurora Node  │   │  - Aurora Node  │   │  - Aurora Node  │
    │  - Redis Shard  │   │  - Redis Shard  │   │  - Redis Shard  │
    └─────────────────┘   └─────────────────┘   └─────────────────┘
      (If AZ-1 Fails: ARC Shifts Traffic to AZ-2 and AZ-3 in < 3 Seconds)
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Five-Nines High Availability (99.999% SLA)Downtime during payment clearing incurs $50k/minute merchant SLA breach penalties (AVL-4919).0.40Elena Rostova (Head of Payment Reliability)
    Automated Sub-Second Zonal Traffic ShiftManual DNS updates took 42 minutes in AVL-4919; failover must be instantaneous.0.30David O'Reilly (Chief Availability Architect)
    Fault Domain Blast-Radius Isolation (Bulkheads)A localized node or memory leak in one zone must never take down sibling zones.0.15Operational Resilience Steering Board
    Error Budget Governance & SRE Deployment GatesDepleting the quarterly 5-minute error budget must automatically freeze feature deploys.0.15SRE Reliability Engineering Charter

    Comparison

    Availability Architecture StrategyDowntime on AZ OutageAnnual Downtime SLAFailover AutomationEvaluation
    Option A: Active-Passive with Manual DNS (Legacy)42 Minutes (Caused AVL-4919 crash)99.9% (~8.7 hours/yr)Manual (Slow DNS TTLs)Rejected: Caused AVL-4919 $3.8M catastrophe.
    Option B: Active-Passive with Auto-Failover4 to 8 Minutes99.95% (~4.3 hours/yr)Automated (Warm standby)Rejected: 4-minute drop window breaches five-nines SLA.
    Option C: Multi-AZ Active-Active Mesh (Chosen)Zero (0 ms - Traffic absorbed)99.999% (< 5.26 min/yr)Instant (ARC Routing < 3s)Selected: Five-nines resilience, zero downtime, proven.

    Result

    Option C is selected. A 3-AZ Active-Active deployment across AWS us-east-1a, 1b, and 1c is standardized; Route53 ARC executes automated zonal traffic evacuation; database quorum storage guarantees zero write interruption.


    Required Mechanisms

    1. Multi-AZ Active-Active Topology & Sizing [MC-AA-01]
    • Compute Sizing Rule (N+1 Resiliency):
      • Each Availability Zone is provisioned with 50% peak capacity headroom.
      • If one entire AZ experiences a catastrophic blackout, the remaining two AZs seamlessly absorb 100% of the 48,000 TPS load without autoscaling delays or CPU saturation.
    2. Route53 Application Recovery Controller (ARC) [MC-RC-01]
    • The AVL-4919 Zero-Downtime Traffic Evacuation:
      • Deep synthetic canary checks evaluate end-to-end payment authorization health every 2 seconds.
      • If error rates in AZ-1 exceed 1.0% over a 10-second window, ARC routing control flips the health gate, swinging all incoming ingress traffic away from AZ-1 in

    $< 3\text{ seconds}$.

    3. Error Budget & Burn Rate Defense [MC-EB-01]
    • Five-Nines Error Budget:
      $$\text{Allowed Downtime} = 365 \times 24 \times 60 \times (1 - 0.99999) = \mathbf{5.256\text{ minutes/year (78.8 seconds/quarter)}}$$

    Automated Freeze Policy: If the 30-day burn rate exceeds 20% of the quarterly error budget, automated CI/CD deployment pipelines freeze all non-security feature releases.


    Invariants and Contracts

    Mandatory Active-Active Multi-AZ Deployment [INV-AVAL-01]
      Production payment services must run active-active across at least three distinct Availability Zones.
      Deploying active-passive architectures with cold standby replicas for Tier-1 services is prohibited.
    
    Automated Zonal Evacuation SLA (< 5s) [INV-AVAL-02]
      Zonal traffic shifting must execute automatically in less than 5 seconds without manual human approval.
      Relying on manual DNS record modifications or SRE incident calls for failover is strictly barred.
    
    Error Budget Deployment Enforcement [INV-AVAL-03]
      Exhausting the quarterly error budget triggers an immediate, automated feature deployment freeze.
      Bypassing error budget deployment freezes requires unanimous written authorization from the CTO.
    

    Explicit Unknowns

    • Regional cross-AZ latency jitter during severe East Coast hurricane fiber optic re-routing (G-1).
    • Time required for Amazon Route53 DNS caching layers to purge stale IPs on non-compliant ISP resolvers (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    48,000 transactions/sec across 32 servicesprovidedPayment network capacity intakeCurrent
    $110B annual payment clearing volumeprovidedFinancial scope portfolio intakeCurrent
    Incident AVL-4919 42-minute outage ($3.8M loss)providedOperations forensic audit reportHistorical
    99.999% availability SLA targetprovidedCorporate Payment Reliability CharterCurrent
    Multi-AZ Active-Active topology selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory active-active multi-AZ invariantdecidedArchitectural invariant INV-AVAL-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against availability architecture standards:

    • Resilience Rigor: PASS. 3-AZ Active-Active topology with N+1 headroom guarantees five-nines uptime.
    • Failover Automation: PASS. Route53 ARC evacuates traffic in under 3 seconds, closing AVL-4919 defect.
    • Error Budget Policy: PASS. Enforces strict quarterly error budget burns with automated deployment freezes.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-AVAL-01: David O'Reilly to determine whether AWS Local Zones should be added as secondary edge ingress points for Chicago and New York financial exchanges in Q1 (Owner: David O'Reilly).

    Next steps

    1. Platform Infrastructure squad deploys the Route53 Application Recovery Controller routing rules.
    2. Ingress Platform team configures 3-AZ active-active Envoy gateways with N+1 capacity sizing.
    3. Conduct staging chaos game day severing 100% of AZ-1 network traffic under 48,000 TPS to confirm sub-3s automated evacuation.

    skill: availability-architect

    High Availability Platform — Fitness Self-Check [AVALARCH-PAY-FIT-001]

    Summary

    This fitness self-check evaluates the high availability platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an active-active zonal deployment where application pods in AZ-1 and AZ-2 attempt to write conflicting customer balance transactions to separate un-replicated local storage volumes.Aurora storage quorum consistency validator probe_uncoordinated_zonal_storage_split verifying storage quorum commitment rejection with diagnostic ERR_CROSS_ZONAL_STORAGE_SPLIT_PROHIBITED.passConfirms Aurora distributed storage layer quorum rules; does not inspect ad-hoc temporary container storage.
    FIT-2: Undefined GrainSeed an availability monitoring telemetry probe that emits uptime availability metrics without defining an explicit measurement interval or regional boundary grain.Availability metric linter probe_missing_availability_grain verifying telemetry rejection with diagnostic ERR_AVAILABILITY_METRIC_LACKS_DECLARED_GRAIN.passConfirms Prometheus / CloudWatch telemetry schema gates; does not inspect unmonitored test instances.
    FIT-3: Silent Schema DriftSeed a service update that modifies health-check endpoint return codes (/healthz returns string "OK" instead of JSON {"status":"HEALTHY"}) without notifying the Route53 ARC probe.Health-check schema contract validator probe_unannounced_health_check_drift verifying deployment failure with diagnostic ERR_HEALTH_CHECK_CONTRACT_VIOLATION.passConfirms automated CI/CD health endpoint contract tests; does not inspect manual curl commands.

    Residual Risk

    • Latency overhead (up to 3.5 ms) during cross-AZ inter-service gRPC calls when a service in AZ-1 calls a downstream dependency in AZ-2. Accepted by Elena Rostova with local AZ affinity routing.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of uncoordinated zonal storage splitsderivedFIT-1 probe result2026-09-15
    Rejection of availability metrics lacking grainderivedFIT-2 probe result2026-09-15
    Rejection of unannounced health check contract driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates availability fitness probes into automated release verification.
    2. SRE squad configures Prometheus alerts monitoring real-time error budget burn rates.
    3. Conduct quarterly disaster recovery drill simulating automated zonal evacuation under peak transaction volumes.

    high-availability-platform-and-active-ac.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Model failure domains and correlated faults for multi-region appsDesign degraded modes to preserve core journeys during outagesEstablish dependency budgets and detection time requirementsMap stateful failover trade-offs between consistency and availability

    About this skill

    What it does

    This skill owns end-to-end behavior when required capabilities or dependencies fail, degrade, recover, or change. It translates accepted business criticality and service objectives into availability scenarios, failure-domain models, dependency budgets, redundancy and recovery contracts, degraded modes, state reconciliation, and executable evidence.

    Use it when

    • User journeys cross services, stores, queues, networks, identity systems, third parties, regions, or providers
    • Component availability and dependency behavior must compose into an end-to-end outcome
    • Process, host, zone, region, control-plane, provider, personnel, data, and change failures may correlate
    • Redundancy, quorum, leader election, routing, replication, retries, failover, and recovery can interact or amplify failure
    • Stateful failover needs fencing, idempotency, ordering, deduplication, reconciliation, consistency, and data-loss decisions
    • Degraded modes must preserve explicit capabilities, correctness, safety, privacy, and restoration behavior

    For example: “Our pharmacy dispensing system stops entirely when the drug interaction service is unreachable. Last month that was 40 minutes and pharmacists sent patients home.”

    What you get

    • architecture/availability-architect/README.md
    • architecture/availability-architect/00-overview/availability-architect-overview.md
    • architecture/availability-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to calculate uptime/SLA/SLO, configure one load balancer/replica/health check, choose active-active, fix an outage, run a failover/chaos/DR test, design backups, operate SRE/on-call, or solve performance/scalability.

    How it works

    1. Check availability is the concern and not recovery from a disaster.
    2. State the availability requirement per journey, not per system.
    3. Enumerate the dependencies and what each failure removes.
    4. Define the degraded mode for each loss, explicitly.
    5. Say how failure is detected and how long detection takes.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions