Database Backup and Disaster Recovery Contract Design

    1

    Designs database backup architectures: continuous WAL archiving, 5-minute RPO, automated restore drills, and WORM vaults.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Database Backup and Disaster Recovery Contract Design

    Example session with this skill installed

    Design database backup, disaster recovery, Point-in-Time Recovery (PITR), and automated restore verification for Core Payment Ledger.

    • Subject: core-payment-ledger-db v2.0 (run-id: aurora-backup-001) managing $65B in balances across 18 million accounts on AWS Aurora PostgreSQL 16.
    • Deciders: Lead Database Architect (David O'Reilly) and Head of Financial Reliability (Elena Rostova).
    • Scope: Continuous WAL archiving with 5-minute RPO; 30-minute RTO cross-region failover; WORM ransomware air-gaps; weekly automated restore drills.
    • Tension: Automated nightly snapshots were never tested for restore; in incident BKP-4919, volume corruption revealed restore scripts failed due to engine version mismatch, extending downtime to 54 hours, losing 18 hours of transactions, and incurring $5.2M in fines. Elena Rostova and David O'Reilly mandate an authoritative Backup Plan: continuous WAL streaming, sub-30m RTO, immutable WORM vaults, and automated weekly synthetic restore drills.

    Write the database backup plan under docs/.

    • Read your context and instructions
    • Compiled the database backup
    • Generated the UI component

    Wrote docs/architecture/tasks/aurora-backup-001/backup-design/db-backup-plan.md. Complete database backup and restore contract specification establishing point-in-time recovery, RPO/RTO SLAs, automated restore drills, and WORM storage isolation.


    skill: backup-design

    Database Backup & Restore Plan: Core Payment Ledger [BKP-PAY-001]

    Summary

    This specification establishes the database backup, disaster recovery, Point-in-Time Recovery (PITR), and automated restore verification plan for core-payment-ledger-db v2.0 under run ID aurora-backup-001. It governs backup and recovery architecture across primary AWS Aurora PostgreSQL 16 clusters managing $65B in transaction balances across 18 million active credit accounts. It decisively investigates and resolves the catastrophic data loss and recovery paralysis demonstrated in incident BKP-4919 (where automated nightly snapshot backups were never tested for restore, and when a storage volume corrupted, the restore script failed due to an outdated PostgreSQL major version incompatibility, extending downtime to 54 hours, losing 18 hours of transaction history, and incurring $5.2M in regulatory fines and ledger reconciliations). The plan defines continuous write-ahead log (WAL) archiving with a 5-minute RPO, establishes automated multi-AZ cross-region Point-in-Time Recovery with a 30-minute RTO, enforces WORM (Write Once, Read Many) ransomware-isolated snapshot vaults, and mandates

    automated weekly synthetic restore verification drills.

    Detailed Description

    A backup that has never been restored is not a backup; it is merely an unverified assumption. In enterprise financial systems, relying on automated cloud provider snapshot checkmarks without executing regular, automated restore verification guarantees catastrophic downtime when real storage failures or ransomware attacks occur. Comprehensive Backup Design specifies the complete recovery lifecycle: continuous transaction log archiving to achieve near-zero Recovery Point Objectives (RPO), automated snapshot lifecycle tiers, cross-region replication to survive regional cloud outages, immutable storage locks, and automated restore testing pipelines that validate data integrity down to the transaction row.

    Primary Production Database (AWS Aurora PostgreSQL 16)
                             │
            ┌────────────────┴────────────────┐
            ▼ (Continuous Stream: RPO <= 5m)  ▼ (Daily Snapshot: 02:00 UTC)
      [ Write-Ahead Log (WAL) Stream ]   [ Automated Aurora Storage Snapshot ]
            │                                 │
            ▼                                 ▼ (Replicated via AWS Backup)
      [ Amazon S3 Vault (Cross-Region) ] [ Immutable AWS Backup Vault (WORM) ]
            │                                 ├── Object Lock Legal Hold (365 Days)
            │                                 └── Ransomware-Proof Air Gap
            └────────────────┬────────────────┘
                             ▼
    [ Automated Weekly Restore Verification Pipeline (Synthetic Staging) ]
      ├── 1. Provisions Ephemeral Isolated Aurora Cluster from Backup
      ├── 2. Runs Cryptographic Row Checksum & Ledger Integrity Verification
      └── 3. Tears Down Test Cluster & Emits Compliance Attestation Receipt
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Recovery Point Objective (RPO <= 5 Minutes)Losing transaction history destroys financial ledgers and draws fines (BKP-4919).0.40Elena Rostova (Head of Financial Reliability)
    Recovery Time Objective (RTO <= 30 Minutes)Core payment settlement downtime incurs merchant SLA failure penalties of $50k/min.0.30David O'Reilly (Lead Database Architect)
    Automated Restore Testing & Integrity OraclesUn-tested restore scripts failed during BKP-4919; restores must be proven automatically.0.15Operational Resilience Steering Group
    Immutable Ransomware Defense (WORM Lock)Financial regulations mandate protecting transaction archives against malicious deletion.0.15Corporate Information Security Mandate

    Comparison

    Backup Architecture CandidateRPO CapabilityRTO CapabilityRestore VerificationEvaluation
    Option A: Daily Native Snapshots Only (Legacy)24 Hours (Lost 18h in BKP-4919)54 Hours (Failed in BKP-4919)Manual (Never tested)Rejected: Caused BKP-4919 disaster; unviable.
    Option B: Nightly pg_dump Logical Export24 Hours14 Hours (Very slow text import)ScriptedRejected: pg_dump locks tables and takes hours to import at $65B scale.
    Option C: Continuous WAL Archiving + WORM (Chosen)<= 5 Minutes (PITR)<= 30 Minutes (Fast storage clone)Automated Weekly DrillSelected: Zero data loss, sub-30m RTO, proven.

    Result

    Option C is selected. Continuous write-ahead log (WAL) archiving is paired with daily storage snapshots replicated to an immutable cross-region AWS Backup vault; automated weekly restore drills verify recovery pipelines in staging.


    Required Mechanisms

    1. Backup Schedule & Lifecycle Retention Matrix [MC-BL-01]
    Backup TierMechanism & TechnologyFrequencyRetention WindowStorage Target & LocationEncryption & Immutability
    Continuous WALAurora Continuous Backing LogReal-time35 calendar daysAmazon S3 Multi-AZ (Local Region)AWS KMS CMK (AES-256)
    Daily SnapshotAWS Backup Managed SnapshotDaily (02:00 UTC)90 calendar daysAWS Backup Vault (Local us-east-1)KMS CMK; WORM Compliance Lock
    Weekly ArchiveCross-Region Replica SnapshotWeekly (Sun 03:00)365 calendar daysAWS Backup Vault (Remote us-west-2)KMS CMK; WORM Lock (Air-Gapped)
    Monthly RegulatoryImmutable Fiduciary Cold VaultMonthly (1st 04:00)7 years (Statutory)AWS Backup Glacier VaultWORM Legal Hold (21 CFR Part 11)
    2. Recovery Objective Service Levels [MC-RO-01]

    Recovery Point Objective (RPO):

    $\le 5\text{ minutes}$ (Maximum allowable transaction data loss window under catastrophic failure).

    Recovery Time Objective (RTO):

    $\le 30\text{ minutes}$ (Maximum allowable elapsed time from failure declaration to active database read/write availability).

    3. Point-in-Time Recovery (PITR) Execution Procedure [MC-PR-01]
    • The BKP-4919 Prevention Workflow:
      1. Operator or automated orchestrator identifies failure timestamp $T_{\text{fail}}$.
      2. Issues Aurora clone command targeting timestamp $T_{\text{fail}} - 60\text{ seconds}$.
      3. Aurora storage layer recreates cluster using copy-on-write pointers in $< 12\text{ minutes}$.
      4. Database DNS endpoint payment-db.internal swings to the restored cluster via automated Route53 health checks.
    4. Automated Weekly Restore Verification Drill [MC-RV-01]
    • An automated AWS Step Functions workflow executes every Tuesday at 01:00 UTC:
      • Provisions a fresh staging Aurora cluster from the latest cross-region weekly snapshot.
      • Executes data integrity query: SELECT count(*), sum(balance_cents) FROM customer_accounts.
      • Verifies that row count and ledger sums match the production source snapshot within 0.00% tolerance.
      • Emits cryptographic verification attestation receipt to compliance audit bucket; tears down the test cluster.

    Invariants and Contracts

    Mandatory Automated Restore Verification [INV-BKP-01]
      Database backups must undergo automated synthetic restore verification at least once every 7 calendar days.
      Relying on un-tested backup snapshots without automated restore verification is strictly prohibited.
    
    RPO Bound Ceiling Invariant (RPO <= 5 min) [INV-BKP-02]
      Continuous transaction log archiving must not experience replication lag exceeding 300 seconds (5 minutes).
      WAL archive lag exceeding 5 minutes triggers an immediate Sev-1 alert to Database Reliability Engineering.
    
    Immutable WORM Retention Compliance [INV-BKP-03]
      Production backup snapshots must be locked under Write-Once-Read-Many (WORM) compliance mode.
      Deleting or modifying backup snapshots prior to their designated retention expiration is physically blocked.
    

    Explicit Unknowns

    • Cross-region AWS data transfer bandwidth throttling during nationwide fiber optic congestion events (G-1).
    • Time required to restore 45 TB of historical transactional audit archives from Glacier Deep Archive (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    $65B in transaction balances across 18M accountsprovidedCore payment ledger capacity intakeCurrent
    Incident BKP-4919 54-hour outage ($5.2M penalty)providedOperations forensic incident reportHistorical
    RPO <= 5 min and RTO <= 30 min targetsprovidedCorporate Disaster Recovery SLACurrent
    Continuous WAL archiving + WORM vault selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory weekly restore drill invariant INV-BKP-01decidedArchitectural invariant INV-BKP-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against backup design standards:

    • RPO/RTO Realism: PASS. 5-minute RPO achieved via WAL streaming; 30-minute RTO via Aurora fast cloning.
    • Restore Verification: PASS. Automated weekly Step Functions drill prevents repeat of incident BKP-4919.
    • Ransomware Protection: PASS. WORM compliance lock prevents malicious or accidental snapshot deletion.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-BKP-01: Elena Rostova to determine whether cross-region disaster recovery standby should maintain a live hot-standby read replica or rely on 30-minute on-demand PITR restoration (Owner: Elena Rostova).

    Next steps

    1. Database Platform squad configures continuous WAL streaming to S3 and enables AWS Backup WORM locks.
    2. SRE team implements the automated weekly restore verification Step Functions workflow in staging.
    3. Conduct disaster recovery game day simulating total primary AWS AZ failure to verify 30-minute RTO restoration.

    database-backup-and-disaster-recovery-co.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define RPO-aligned WAL archiving and snapshot sequencesSpecify WORM vault and air-gapped encryption requirementsMap restore dependency chains and semantic validation oraclesAudit backup-set integrity and catalog continuity gaps

    About this skill

    What it does

    This skill maps authoritative data scope and recovery obligations into a backup-set, catalog, retention, protection and restore-verification contract. It defines what must be captured, how copy/chain identity composes, and what evidence supports recoverability.

    Use it when

    Use when accepted data owners and recovery consumers need a bounded contract for producing and validating recoverable copies of specified state.

    For example: “Our daily Postgres backups report green status, but during a drill last night we could not restore customer accounts to yesterday 14:00 because WAL logs were missing between 10:00 and 12:00. Also, the restore script failed because it tried to load table data before creating the schema.”

    What you get

    • DB Backup Plan

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/backup-design/.

    What it will not do

    Do not use for DR architecture, storage/database design, replication/snapshot configuration, archive/export, one backup/restore run, vendor selection, migration, implementation or troubleshooting.

    How it works

    1. Check recoverable-copy contract is required.
    2. Bound the protected data set.
    3. Map recovery points to snapshot mechanisms.
    4. Establish backup set and catalog integrity.
    5. Define media placement and key recovery.
    6. Specify restore sequence and validation oracles.
    7. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions