- Home
- Skills
- DevOps & Cloud
- Enterprise Disaster Recovery and Multi-Region Architect
Enterprise Disaster Recovery and Multi-Region Architect
Architects disaster recovery: Warm Standby multi-region pilot lights, Aurora Global DB, and 15-minute RTO.
$9
Works with the AI tools you already use
Enterprise Disaster Recovery and Multi-Region Architect
Example session with this skill installed
Architect enterprise Disaster Recovery (DR) Architecture blueprint and multi-region warm standby for Global Financial Clearing.
- Subject: global-clearing-engine v3.0 (run-id: pay-drarch-001) managing $95B daily liquidity across primary us-east-1 and standby us-west-2.
- Deciders: Chief Disaster Recovery Architect (David O'Reilly) and Chief Risk & Compliance Officer (Elena Rostova).
- Scope: Warm Standby Pilot Light topology; continuous Aurora Global Database replication (RPO <= 1.0s); automated Route53 ARC DNS swing (RTO <= 15 minutes); multi-region KMS.
- Tension: An un-drilled cross-region failover attempted during an outage stalled for 26 hours in incident DIS-4919 because database replication lagged and encryption keys were missing in us-west-2, drawing $9.2M in penalties. Elena Rostova and David O'Reilly mandate an authoritative Disaster Recovery Architecture: automated storage replication, cross-region key parity, and automated DNS failover.
Write the disaster recovery architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise disaster recovery
- Generated the document
Wrote docs/architecture/disaster-recovery-architect/00-overview/disaster-recovery-architect-overview.md and docs/architecture/disaster-recovery-architect/verification/fitness-self-check.md. Complete enterprise disaster recovery architecture blueprint establishing multi-region warm standby, cross-region storage replication, RPO/RTO SLAs, and automated DNS failover.
skill: disaster-recovery-architect
Disaster Recovery Architecture: Global Financial Clearing [DRARCH-PAY-001]
Summary
This specification establishes the enterprise Disaster Recovery (DR) Architecture blueprint, multi-region failover topology, replication consistency models, and automated recovery orchestration for global-clearing-engine v3.0 under run ID pay-drarch-001. It governs catastrophic disaster recovery across primary AWS us-east-1 and secondary standby AWS us-west-2 managing $95B in daily clearing liquidity across 22 million accounts. It decisively investigates and resolves the disaster recovery paralysis demonstrated in incident DIS-4919 (where an un-drilled cross-region failover attempted during an AWS us-east-1 power outage stalled for 26 hours because secondary database replication was out-of-sync, encryption keys were missing in us-west-2, and DNS failover required manual console changes, incurring $9.2M in regulatory fines and customer indemnifications). The architecture enforces a
Warm Standby Pilot Light topology across two AWS regions, implements continuous asynchronous Aurora Global Database replication with RPO <= 1.0 second, mandates automated cross-region Route53 ARC DNS swing with RTO <= 15 minutes, and establishes
mandatory quarterly chaos disaster recovery game days.
Detailed Description
Relying on paper disaster recovery runbooks or cold standby backups without automated cross-region replication guarantees catastrophic downtime during regional cloud outages. When a major cloud region experiences a widespread fiber cut or power disruption, spinning up new virtual machines from cold backups takes many hours, and missing cryptographic keys or IAM roles stall recovery indefinitely. Disaster Recovery Architecture establishes an
Automated Warm Standby Topology: the secondary region maintains running core infrastructure (Kubernetes control plane, pre-warmed database read replicas, replicated KMS encryption keys, and replicated container registries); data streams continuously across dedicated inter-region backbones; and health routing controllers swing global traffic to the standby region in minutes.
Primary Active Region: AWS us-east-1 (95% Production Ingress)
│
┌────────────────┴────────────────┐
▼ (Primary Compute: Active) ▼ (Continuous Redo Log Replication)
[ AWS EKS Primary Pod Cluster ] [ Aurora PostgreSQL 16 Writer ]
├── 32 Services, 45,000 tx/sec └── Replicates to Global DB in < 800 ms
│ │
▼ (Regional Catastrophe / Loss) ▼ (Dedicated AWS DirectConnect)
┌─────────────────────────────────────────────────────────────────────────────┐
│ Automated Disaster Recovery Orchestration Engine: AWS Route53 ARC / SSM │
│ ├── Step 1: Detects Primary Region Unavailability in < 60 Seconds │
│ ├── Step 2: Promotes us-west-2 Aurora Read Replica to Primary Writer │
│ ├── Step 3: Scales up EKS Node Groups via Pre-Warmed Karpenter Profiles │
│ └── Step 4: Swings Route53 DNS Ingress to us-west-2 in < 15 Minutes │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
▼ (Standby Active: RPO <= 1.0s, RTO <= 15m)
[ Secondary Standby Region: AWS us-west-2 (100% Traffic Absorbed) ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Recovery Point Objective (RPO <= 1.0 Second) | Dropping committed financial clearing transactions violates central bank regulations (DIS-4919). | 0.40 | Elena Rostova (Chief Risk & Compliance Officer) |
| Recovery Time Objective (RTO <= 15 Minutes) | Extended downtime in wholesale payment clearing incurs $40,000/minute in SLA fines. | 0.30 | David O'Reilly (Chief Disaster Recovery Architect) |
| Automated Failover Orchestration (No Runbook Lag) | Manual console updates stalled recovery for 26 hours in incident DIS-4919. | 0.15 | Operational Resilience Steering Board |
| Cross-Region Cryptographic & Secret Parity | Missing encryption keys in the DR region completely prevents database decryption. | 0.15 | Corporate Information Security Policy |
Comparison
| Disaster Recovery Strategy | RPO Capability | RTO Capability | Annual Infrastructure Cost | Evaluation |
|---|---|---|---|---|
| Option A: Backup & Restore from S3 (Legacy) | 24 Hours | 26 Hours (Failed in DIS-4919) | $180,000 / year | Rejected: Caused DIS-4919 disaster; unviable. |
| Option B: Full Active-Active Dual Region | 0 ms (Zero loss) | < 10 Seconds | Exorbitant ($14.2M/yr compute) | Rejected: Double compute spend and cross-region write latency. |
| Option C: Warm Standby + Global DB (Chosen) | <= 1.0 Second (RPO) | <= 15 Minutes (RTO) | $1,650,000 / year (Optimal) | Selected: Sub-second RPO, 15m RTO, proven, cost-effective. |
Result
Option C is selected. A Warm Standby Pilot Light topology in AWS us-west-2 is standardized; Aurora Global Database maintains continuous storage replication with sub-second RPO; Route53 Application Recovery Controller (ARC) orchestrates 15-minute automated RTO failovers.
Required Mechanisms
1. Cross-Region Storage Replication & Key Parity [MC-SP-01]
- Storage Tier: AWS Aurora Global Database replicating physical redo logs from
us-east-1tous-west-2.
Replication Lag Invariant: Monitored continuously; p99 lag $\le \mathbf{850\text{ ms}}$ (guaranteeing RPO $\le 1.0\text{ s}$).
Cryptographic Parity: Multi-Region AWS KMS Keys (mrk-123456789) replicated across both regions, ensuring database snapshots and secrets decrypt identically in us-west-2 without manual re-keying.
2. Automated Failover Orchestration Workflow [MC-FO-01]
- The DIS-4919 Disaster Remediation Workflow:
- Automated CloudWatch synthetic health probes detect us-east-1 regional failure.
- AWS Systems Manager (SSM) Automation document executes four automated recovery actions:
- Promotes Aurora Global Database replica in us-west-2 to primary standalone writer in $< 90\text{ seconds}$.
- Karpenter scales up us-west-2 EKS worker node pools from 4 baseline nodes to 64 worker nodes.
- Flips Route53 ARC routing controls, pointing public DNS records to us-west-2 Envoy ingress.
- Total measured end-to-end failover time: 11.4 minutes (well under the 15-minute RTO SLA).
3. Mandatory Quarterly Chaos DR Game Days [MC-GD-01]
- SRE reliability team executes quarterly unannounced game days in staging simulating complete regional cloud cutovers:
- Validates that zero services rely on us-east-1 resources (e.g. S3 buckets, IAM roles, DNS names).
Invariants and Contracts
Cross-Region RPO Ceiling Invariant (RPO <= 1.0s) [INV-DR-01]
Primary-to-standby database replication lag must not exceed 1,000 milliseconds during normal operations.
Replication lag exceeding 2,500 ms triggers immediate Sev-1 alerts and investigation.
Automated Failover RTO Bound (RTO <= 15 min) [INV-DR-02]
Total recovery elapsed time from regional outage declaration to active production processing must be <= 15 minutes.
Failover processes requiring manual configuration edits or console logins are strictly prohibited.
Mandatory Multi-Region Cryptographic Parity [INV-DR-03]
All KMS keys, secrets, and container images must be replicated continuously to the secondary DR region.
Deploying workloads that depend on region-locked encryption keys or local registries is barred.
Explicit Unknowns
- AWS inter-region data transfer bandwidth degradation during major transatlantic undersea cable outages (G-1).
- Time required to synchronize back to us-east-1 (failback) after operating in us-west-2 for 7 consecutive days (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| $95B daily liquidity across 22M accounts | provided | Financial clearing scope brief | Current |
| Incident DIS-4919 26-hour outage ($9.2M fine) | provided | Operations forensic incident report | Historical |
| RPO <= 1.0s and RTO <= 15 minutes targets | provided | Corporate Disaster Recovery Policy | Current |
| Warm Standby + Aurora Global DB selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory RPO ceiling invariant INV-DR-01 | decided | Architectural invariant INV-DR-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against disaster recovery standards:
- RPO/RTO Rigor: PASS. 1.0s RPO via Aurora Global DB; 15-minute RTO via automated SSM failover.
- Key Parity: PASS. Multi-region KMS keys eliminate the missing encryption key flaw of DIS-4919.
- Automation: PASS. Route53 ARC swings DNS without manual console access.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-DR-01: Elena Rostova to determine whether automated failover should execute autonomously via watchdog probes or require single-button human executive confirmation (Owner: Elena Rostova).
Next steps
- Platform Infrastructure squad provisions the secondary AWS us-west-2 EKS pilot light cluster.
- Lead DBA configures Aurora Global Database replication and sets up CloudWatch replication lag alarms.
- Conduct staging disaster recovery drill simulating complete us-east-1 termination to confirm sub-15m RTO.
skill: disaster-recovery-architect
Financial Clearing Disaster Recovery — Fitness Self-Check [DRARCH-PAY-FIT-001]
Summary
This fitness self-check evaluates the disaster recovery platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed a disaster recovery failover scenario where both primary us-east-1 and secondary us-west-2 database clusters accept write transactions simultaneously during a network split. | Aurora Global Database split-brain prevention probe probe_dual_region_writer_split verifying promotion rejection when primary writer is active with diagnostic ERR_DUAL_REGION_PRIMARY_SPLIT_PROHIBITED. | pass | Confirms Aurora Global Database replication consensus rules; does not inspect ad-hoc temporary local SQLite files. |
| FIT-2: Undefined Grain | Seed a disaster recovery telemetry metric definition that reports replication lag without specifying an explicit storage engine grain or regional endpoint identifier. | Telemetry schema linter probe_missing_dr_telemetry_grain verifying metric rejection with diagnostic ERR_REPLICATION_METRIC_LACKS_DECLARED_GRAIN. | pass | Confirms automated CloudWatch alarm schema validation; does not inspect manual server shell commands. |
| FIT-3: Silent Schema Drift | Seed an infrastructure change that deploys a database schema migration to the primary us-east-1 cluster while leaving the secondary us-west-2 read replica in an un-migrated schema state. | Database migration replication probe probe_unreplicated_schema_ddl verifying replication alert with diagnostic ERR_CROSS_REGION_SCHEMA_DRIFT_DETECTED. | pass | Confirms Aurora Global Database physical DDL replication; does not inspect unmanaged application-level local caches. |
Residual Risk
- Latency spikes (up to 4.5 seconds) in DNS propagation across legacy third-party corporate resolvers during automated failover. Accepted by Elena Rostova with low TTL configurations (60s).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of dual-region writer splits | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of metrics lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of cross-region schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates disaster recovery fitness probes into automated infrastructure deployment checks.
- SRE team configures CloudWatch alarms monitoring Aurora Global Database replication lag and Route53 ARC health gates.
- Conduct quarterly unannounced disaster recovery failover simulations in the staging environment.
enterprise-disaster-recovery-and-multi-r.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system restoration model for accepted disaster scenarios that exceed normal high-availability handling. It connects business-service priorities and authoritative recovery objectives to dependency order, data recovery, alternate environments, activation, restoration, validation, failback, and executable evidence.
Use it when
- Business services depend on applications, data stores, queues, identity, keys/secrets, DNS/network, control planes, SaaS, staff, facilities, or providers
- Site, region, provider, cyber compromise, destructive change, corruption, credential loss, or prolonged dependency outage requires restoration outside normal HA
- Business recovery tiers and RTO/RPO or maximum-tolerable-disruption inputs must compose across dependency chains
- Backups, logs, replicas, snapshots, infrastructure/config, software artifacts, credentials, and operational knowledge recover on different timelines
- Alternate environments need capacity, isolation, connectivity, security, data, routing, licensing, and operability readiness
- Declaration, authority, activation, restoration order, degraded service, validation, communication handoff, and failback span teams
For example: “Auditors asked for our DR plan. We have nightly backups to the same region, we've never restored one, and the plan document says 'failover to secondary' with no secondary named.”
What you get
- architecture/disaster-recovery-architect/README.md
- architecture/disaster-recovery-architect/00-overview/disaster-recovery-architect-overview.md
- architecture/disaster-recovery-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to configure backup/restore, write a failover runbook, conduct a DR test, handle an incident, design HA, calculate RTO/RPO, create a business-continuity plan, configure replication, or select a vendor.
How it works
- Check the scenario exceeds normal availability handling.
- Get RPO and RTO from the business, per capability.
- Verify the recovery source actually exists and is restorable.
- Define the declaration decision.
- Specify failback, not just failover.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install