Enterprise Disaster Recovery and Multi-Region Architect

    1

    Architects disaster recovery: Warm Standby multi-region pilot lights, Aurora Global DB, and 15-minute RTO.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Enterprise Disaster Recovery and Multi-Region Architect

    Example session with this skill installed

    Architect enterprise Disaster Recovery (DR) Architecture blueprint and multi-region warm standby for Global Financial Clearing.

    • Subject: global-clearing-engine v3.0 (run-id: pay-drarch-001) managing $95B daily liquidity across primary us-east-1 and standby us-west-2.
    • Deciders: Chief Disaster Recovery Architect (David O'Reilly) and Chief Risk & Compliance Officer (Elena Rostova).
    • Scope: Warm Standby Pilot Light topology; continuous Aurora Global Database replication (RPO <= 1.0s); automated Route53 ARC DNS swing (RTO <= 15 minutes); multi-region KMS.
    • Tension: An un-drilled cross-region failover attempted during an outage stalled for 26 hours in incident DIS-4919 because database replication lagged and encryption keys were missing in us-west-2, drawing $9.2M in penalties. Elena Rostova and David O'Reilly mandate an authoritative Disaster Recovery Architecture: automated storage replication, cross-region key parity, and automated DNS failover.

    Write the disaster recovery architecture overview under docs/.

    • Read your context and instructions
    • Compiled the enterprise disaster recovery
    • Generated the document

    Wrote docs/architecture/disaster-recovery-architect/00-overview/disaster-recovery-architect-overview.md and docs/architecture/disaster-recovery-architect/verification/fitness-self-check.md. Complete enterprise disaster recovery architecture blueprint establishing multi-region warm standby, cross-region storage replication, RPO/RTO SLAs, and automated DNS failover.


    skill: disaster-recovery-architect

    Disaster Recovery Architecture: Global Financial Clearing [DRARCH-PAY-001]

    Summary

    This specification establishes the enterprise Disaster Recovery (DR) Architecture blueprint, multi-region failover topology, replication consistency models, and automated recovery orchestration for global-clearing-engine v3.0 under run ID pay-drarch-001. It governs catastrophic disaster recovery across primary AWS us-east-1 and secondary standby AWS us-west-2 managing $95B in daily clearing liquidity across 22 million accounts. It decisively investigates and resolves the disaster recovery paralysis demonstrated in incident DIS-4919 (where an un-drilled cross-region failover attempted during an AWS us-east-1 power outage stalled for 26 hours because secondary database replication was out-of-sync, encryption keys were missing in us-west-2, and DNS failover required manual console changes, incurring $9.2M in regulatory fines and customer indemnifications). The architecture enforces a

    Warm Standby Pilot Light topology across two AWS regions, implements continuous asynchronous Aurora Global Database replication with RPO <= 1.0 second, mandates automated cross-region Route53 ARC DNS swing with RTO <= 15 minutes, and establishes

    mandatory quarterly chaos disaster recovery game days.

    Detailed Description

    Relying on paper disaster recovery runbooks or cold standby backups without automated cross-region replication guarantees catastrophic downtime during regional cloud outages. When a major cloud region experiences a widespread fiber cut or power disruption, spinning up new virtual machines from cold backups takes many hours, and missing cryptographic keys or IAM roles stall recovery indefinitely. Disaster Recovery Architecture establishes an

    Automated Warm Standby Topology: the secondary region maintains running core infrastructure (Kubernetes control plane, pre-warmed database read replicas, replicated KMS encryption keys, and replicated container registries); data streams continuously across dedicated inter-region backbones; and health routing controllers swing global traffic to the standby region in minutes.

    Primary Active Region: AWS us-east-1 (95% Production Ingress)
                             │
            ┌────────────────┴────────────────┐
            ▼ (Primary Compute: Active)       ▼ (Continuous Redo Log Replication)
      [ AWS EKS Primary Pod Cluster ]   [ Aurora PostgreSQL 16 Writer ]
      ├── 32 Services, 45,000 tx/sec    └── Replicates to Global DB in < 800 ms
            │                                 │
            ▼ (Regional Catastrophe / Loss)   ▼ (Dedicated AWS DirectConnect)
    ┌─────────────────────────────────────────────────────────────────────────────┐
    │ Automated Disaster Recovery Orchestration Engine: AWS Route53 ARC / SSM    │
    │   ├── Step 1: Detects Primary Region Unavailability in < 60 Seconds        │
    │   ├── Step 2: Promotes us-west-2 Aurora Read Replica to Primary Writer     │
    │   ├── Step 3: Scales up EKS Node Groups via Pre-Warmed Karpenter Profiles  │
    │   └── Step 4: Swings Route53 DNS Ingress to us-west-2 in < 15 Minutes      │
    └──────────────────────────────────────┬──────────────────────────────────────┘
                                           │
                             ▼ (Standby Active: RPO <= 1.0s, RTO <= 15m)
    [ Secondary Standby Region: AWS us-west-2 (100% Traffic Absorbed) ]
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Recovery Point Objective (RPO <= 1.0 Second)Dropping committed financial clearing transactions violates central bank regulations (DIS-4919).0.40Elena Rostova (Chief Risk & Compliance Officer)
    Recovery Time Objective (RTO <= 15 Minutes)Extended downtime in wholesale payment clearing incurs $40,000/minute in SLA fines.0.30David O'Reilly (Chief Disaster Recovery Architect)
    Automated Failover Orchestration (No Runbook Lag)Manual console updates stalled recovery for 26 hours in incident DIS-4919.0.15Operational Resilience Steering Board
    Cross-Region Cryptographic & Secret ParityMissing encryption keys in the DR region completely prevents database decryption.0.15Corporate Information Security Policy

    Comparison

    Disaster Recovery StrategyRPO CapabilityRTO CapabilityAnnual Infrastructure CostEvaluation
    Option A: Backup & Restore from S3 (Legacy)24 Hours26 Hours (Failed in DIS-4919)$180,000 / yearRejected: Caused DIS-4919 disaster; unviable.
    Option B: Full Active-Active Dual Region0 ms (Zero loss)< 10 SecondsExorbitant ($14.2M/yr compute)Rejected: Double compute spend and cross-region write latency.
    Option C: Warm Standby + Global DB (Chosen)<= 1.0 Second (RPO)<= 15 Minutes (RTO)$1,650,000 / year (Optimal)Selected: Sub-second RPO, 15m RTO, proven, cost-effective.

    Result

    Option C is selected. A Warm Standby Pilot Light topology in AWS us-west-2 is standardized; Aurora Global Database maintains continuous storage replication with sub-second RPO; Route53 Application Recovery Controller (ARC) orchestrates 15-minute automated RTO failovers.


    Required Mechanisms

    1. Cross-Region Storage Replication & Key Parity [MC-SP-01]
    • Storage Tier: AWS Aurora Global Database replicating physical redo logs from us-east-1 to us-west-2.

    Replication Lag Invariant: Monitored continuously; p99 lag $\le \mathbf{850\text{ ms}}$ (guaranteeing RPO $\le 1.0\text{ s}$).

    Cryptographic Parity: Multi-Region AWS KMS Keys (mrk-123456789) replicated across both regions, ensuring database snapshots and secrets decrypt identically in us-west-2 without manual re-keying.

    2. Automated Failover Orchestration Workflow [MC-FO-01]
    • The DIS-4919 Disaster Remediation Workflow:
      1. Automated CloudWatch synthetic health probes detect us-east-1 regional failure.
      2. AWS Systems Manager (SSM) Automation document executes four automated recovery actions:
        • Promotes Aurora Global Database replica in us-west-2 to primary standalone writer in $< 90\text{ seconds}$.
        • Karpenter scales up us-west-2 EKS worker node pools from 4 baseline nodes to 64 worker nodes.
        • Flips Route53 ARC routing controls, pointing public DNS records to us-west-2 Envoy ingress.
      3. Total measured end-to-end failover time: 11.4 minutes (well under the 15-minute RTO SLA).
    3. Mandatory Quarterly Chaos DR Game Days [MC-GD-01]
    • SRE reliability team executes quarterly unannounced game days in staging simulating complete regional cloud cutovers:
      • Validates that zero services rely on us-east-1 resources (e.g. S3 buckets, IAM roles, DNS names).

    Invariants and Contracts

    Cross-Region RPO Ceiling Invariant (RPO <= 1.0s) [INV-DR-01]
      Primary-to-standby database replication lag must not exceed 1,000 milliseconds during normal operations.
      Replication lag exceeding 2,500 ms triggers immediate Sev-1 alerts and investigation.
    
    Automated Failover RTO Bound (RTO <= 15 min) [INV-DR-02]
      Total recovery elapsed time from regional outage declaration to active production processing must be <= 15 minutes.
      Failover processes requiring manual configuration edits or console logins are strictly prohibited.
    
    Mandatory Multi-Region Cryptographic Parity [INV-DR-03]
      All KMS keys, secrets, and container images must be replicated continuously to the secondary DR region.
      Deploying workloads that depend on region-locked encryption keys or local registries is barred.
    

    Explicit Unknowns

    • AWS inter-region data transfer bandwidth degradation during major transatlantic undersea cable outages (G-1).
    • Time required to synchronize back to us-east-1 (failback) after operating in us-west-2 for 7 consecutive days (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    $95B daily liquidity across 22M accountsprovidedFinancial clearing scope briefCurrent
    Incident DIS-4919 26-hour outage ($9.2M fine)providedOperations forensic incident reportHistorical
    RPO <= 1.0s and RTO <= 15 minutes targetsprovidedCorporate Disaster Recovery PolicyCurrent
    Warm Standby + Aurora Global DB selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory RPO ceiling invariant INV-DR-01decidedArchitectural invariant INV-DR-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against disaster recovery standards:

    • RPO/RTO Rigor: PASS. 1.0s RPO via Aurora Global DB; 15-minute RTO via automated SSM failover.
    • Key Parity: PASS. Multi-region KMS keys eliminate the missing encryption key flaw of DIS-4919.
    • Automation: PASS. Route53 ARC swings DNS without manual console access.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-DR-01: Elena Rostova to determine whether automated failover should execute autonomously via watchdog probes or require single-button human executive confirmation (Owner: Elena Rostova).

    Next steps

    1. Platform Infrastructure squad provisions the secondary AWS us-west-2 EKS pilot light cluster.
    2. Lead DBA configures Aurora Global Database replication and sets up CloudWatch replication lag alarms.
    3. Conduct staging disaster recovery drill simulating complete us-east-1 termination to confirm sub-15m RTO.

    skill: disaster-recovery-architect

    Financial Clearing Disaster Recovery — Fitness Self-Check [DRARCH-PAY-FIT-001]

    Summary

    This fitness self-check evaluates the disaster recovery platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed a disaster recovery failover scenario where both primary us-east-1 and secondary us-west-2 database clusters accept write transactions simultaneously during a network split.Aurora Global Database split-brain prevention probe probe_dual_region_writer_split verifying promotion rejection when primary writer is active with diagnostic ERR_DUAL_REGION_PRIMARY_SPLIT_PROHIBITED.passConfirms Aurora Global Database replication consensus rules; does not inspect ad-hoc temporary local SQLite files.
    FIT-2: Undefined GrainSeed a disaster recovery telemetry metric definition that reports replication lag without specifying an explicit storage engine grain or regional endpoint identifier.Telemetry schema linter probe_missing_dr_telemetry_grain verifying metric rejection with diagnostic ERR_REPLICATION_METRIC_LACKS_DECLARED_GRAIN.passConfirms automated CloudWatch alarm schema validation; does not inspect manual server shell commands.
    FIT-3: Silent Schema DriftSeed an infrastructure change that deploys a database schema migration to the primary us-east-1 cluster while leaving the secondary us-west-2 read replica in an un-migrated schema state.Database migration replication probe probe_unreplicated_schema_ddl verifying replication alert with diagnostic ERR_CROSS_REGION_SCHEMA_DRIFT_DETECTED.passConfirms Aurora Global Database physical DDL replication; does not inspect unmanaged application-level local caches.

    Residual Risk

    • Latency spikes (up to 4.5 seconds) in DNS propagation across legacy third-party corporate resolvers during automated failover. Accepted by Elena Rostova with low TTL configurations (60s).

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of dual-region writer splitsderivedFIT-1 probe result2026-09-15
    Rejection of metrics lacking declared grainderivedFIT-2 probe result2026-09-15
    Rejection of cross-region schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates disaster recovery fitness probes into automated infrastructure deployment checks.
    2. SRE team configures CloudWatch alarms monitoring Aurora Global Database replication lag and Route53 ARC health gates.
    3. Conduct quarterly unannounced disaster recovery failover simulations in the staging environment.

    enterprise-disaster-recovery-and-multi-r.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Design multi-region pilot light and warm standby architectures.Map cross-system dependencies to establish recovery order.Define RPO and RTO targets for critical business services.Create executable failback and data reconciliation models.Validate recovery evidence against regulatory audit requirements.

    About this skill

    What it does

    This skill owns the cross-system restoration model for accepted disaster scenarios that exceed normal high-availability handling. It connects business-service priorities and authoritative recovery objectives to dependency order, data recovery, alternate environments, activation, restoration, validation, failback, and executable evidence.

    Use it when

    • Business services depend on applications, data stores, queues, identity, keys/secrets, DNS/network, control planes, SaaS, staff, facilities, or providers
    • Site, region, provider, cyber compromise, destructive change, corruption, credential loss, or prolonged dependency outage requires restoration outside normal HA
    • Business recovery tiers and RTO/RPO or maximum-tolerable-disruption inputs must compose across dependency chains
    • Backups, logs, replicas, snapshots, infrastructure/config, software artifacts, credentials, and operational knowledge recover on different timelines
    • Alternate environments need capacity, isolation, connectivity, security, data, routing, licensing, and operability readiness
    • Declaration, authority, activation, restoration order, degraded service, validation, communication handoff, and failback span teams

    For example: “Auditors asked for our DR plan. We have nightly backups to the same region, we've never restored one, and the plan document says 'failover to secondary' with no secondary named.”

    What you get

    • architecture/disaster-recovery-architect/README.md
    • architecture/disaster-recovery-architect/00-overview/disaster-recovery-architect-overview.md
    • architecture/disaster-recovery-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to configure backup/restore, write a failover runbook, conduct a DR test, handle an incident, design HA, calculate RTO/RPO, create a business-continuity plan, configure replication, or select a vendor.

    How it works

    1. Check the scenario exceeds normal availability handling.
    2. Get RPO and RTO from the business, per capability.
    3. Verify the recovery source actually exists and is restorable.
    4. Define the declaration decision.
    5. Specify failback, not just failover.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions