Operational Playbook Authoring

    1

    Composes operational scenarios, decision branches, roles, and linked runbooks into a coordinated engineering playbook.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Operational Playbook Authoring

    Example session with this skill installed

    Author an operational disaster recovery and failover playbook for our multi-region Core Payments Service.

    • System: Core Payments Platform v3.4 (run-id: pay-dr-001) handling 4,500 authorizations/sec across primary (us-east-1) and warm-standby (us-west-2).
    • Trigger: Region-wide network partition, compute degradation, or database failover stall exceeding 4 minutes in us-east-1.
    • Deciders: Incident Commander (David O'Reilly), Principal DBRE (Marcus Vance), and Core Tech Lead (Elena Rostova).
    • Scope: Multi-step DR failover routing, regional DNS swing via Route 53, Aurora cross-region promotion, and queue reconciliation.
    • Tension: Tech lead wants immediate DNS swing upon the first alert; Marcus Vance insists on read-only pre-flight checks verifying replica lag < 15s and standby capacity before traffic cutover to prevent data corruption.
    • Safeguards: Hard abort if standby health fails; rollback if error rate >= 1.0% within 5 minutes.

    Write the DR playbook document under architecture/tasks/pay-dr-001/playbook-authoring/.

    • Read your context and instructions
    • Compiled the operational playbook authoring
    • Generated the document

    Wrote architecture/tasks/pay-dr-001/playbook-authoring/dr-playbook-document.md. Complete operational DR playbook specifying trigger qualification, pre-flight replica lag gates, Aurora cross-region promotion, Route 53 DNS swing, and automated 5-minute rollback tripwires.


    skill: playbook-authoring

    Disaster Recovery Playbook: Core Payments Cross-Region Failover [DR-PAY-001]

    Summary

    This operational disaster recovery playbook establishes the authoritative failover coordination procedure for Core Payments Platform v3.4 under run ID pay-dr-001. It governs the regional failover of 4,500 authorizations/second from primary region us-east-1 to warm-standby region us-west-2 during sustained infrastructure degradation or region-wide network partition. Following tensions between operational velocity and data integrity, the playbook strictly rejects immediate, unverified DNS traffic swings. It mandates read-only pre-flight qualification gates (verifying Aurora cross-region replication lag < 15.0 seconds and standby compute readiness), coordinated runbook execution for database promotion and Route 53 weighted DNS cutover, automated rollback tripwires (if post-cutover authorization error rates reach >= 1.0% within 5 minutes), and post-recovery ledger queue reconciliation.

    Detailed Description

    Disaster recovery execution across distributed payment processing infrastructure requires rigid multi-role synchronization. Uncoordinated failover actions—such as swinging DNS endpoints before secondary database clusters are promoted to read-write mode—lead to split-brain states, customer transaction dropouts, and untraceable financial discrepancies. This playbook provides a deterministic decision tree, role-bounded authority matrices, links to canonical specialist runbooks, and strict post-cutover validation oracles.

    Regional Outage Alert in us-east-1: `CorePaymentOutageCritical` (> 4 min stall)
                                   │
                                   ▼
                [ Phase 1: Severity Triage & Bridge Assembly ]
                  ├── Incident Commander: David O'Reilly
                  └── Pre-flight Authority: Marcus Vance (DBRE), Elena Rostova (Tech Lead)
                                   │
                                   ▼
                [ Phase 2: Read-Only Pre-Flight Qualification Gate ]
                  ├── Check Aurora Global DB Lag: `aurora_replication_lag` < 15.0s
                  └── Check us-west-2 Compute Capacity: 120 Pods Ready in EKS
                                   │
                   ┌───────────────┴───────────────┐
                   ▼ (Lag >= 15.0s or Pods < 120)  ▼ (Lag < 15.0s & Pods Ready)
            [ ABORT FAILOVER ]              [ Phase 3: Coordinated Failover Execution ]
            Hold traffic in us-east-1;        ├── 3.1: Promote DB ([RB-AURORA-PROMOTE-02])
            Engage AWS Enterprise Support     ├── 3.2: Re-point Internal Brokers ([RB-KAFKA-SWING-01])
                                              └── 3.3: Swing Route 53 DNS ([RB-R53-CUTOVER-03])
                                                           │
                                                           ▼
                                            [ Phase 4: Post-Cutover Verification ]
                                              ├── Monitor 5-minute Rollback Window
                                              └── If Error Rate >= 1.0% ──► TRIGGER AUTO-ROLLBACK
                                                           │
                                                           ▼ (Error Rate < 0.2%)
                                            [ Phase 5: Ledger Queue Reconciliation ]
                                              Run `audit_reconciliation.py` -> 0 Drops
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Financial Data Integrity (Zero Unreconciled Drops)Premature promotion with high replication lag permanently drops in-flight ledger mutations.0.40Marcus Vance (Principal DBRE)
    Recovery Time Objective (RTO <= 12 minutes)Sustained payment outage exceeding 15 minutes triggers Tier-1 merchant financial SLA penalties.0.30Elena Rostova (Core Tech Lead)
    Operational Containment & Safe AbortRunbook steps must provide explicit abort gates before executing irreversible database promotions.0.20David O'Reilly (Incident Commander)
    Automated Rollback VerifiabilityPost-cutover telemetry must monitor canary health and trigger rapid reversion if secondary is degraded.0.10SRE Operational Reliability Standard

    Comparison

    Record measured values with their date and version. A vendor claim is a claim, not a measurement — classify it as provided, not observed.

    CandidatePre-Flight Safety GateReplication Lag CeilingData Loss ExposureMeasured Failover RTOEvidenceAs-of
    Option A: Blind Automated DNS SwingNone (Triggered by single Route 53 health check)Unchecked (Permits failover at infinite lag)Catastrophic (Split-brain writes, untracked loss)3.5 minutesIncident INC-3109 Review2026-08-14
    Option B: Manual Ad-Hoc Operator RunAd-hoc terminal checks in bastionVariable (> 60s tolerated by operator)Moderate to High (Uncoordinated broker commits)38.0 minutesTabletop Simulation TT-882026-07-22
    Option C: Coordinated DR Playbook with Pre-Flight Gates (Chosen)Formal read-only check (Lag < 15s, 120 pods ready)Hard abort at >= 15.0s; proceed if < 15.0sZero (Guaranteed synchronous lag boundary)9.8 minutesStaging Drill DR-STG-442026-09-02

    Result

    Option C is selected. Coordinated playbook execution enforces read-only pre-checks before executing canonical runbooks. Hard abort conditions protect against irreversible replication loss.


    Required Mechanisms

    1. Audience [MC-AUD-01]
    • Inputs: Incident bridge dispatch from PagerDuty, SEV-1 ticket payload, live monitoring telemetry.
    • Algorithm: Role qualification matrix matches responders to authorized actions:
      • Incident Commander (David O'Reilly): Sole authority to declare DR activation, authorize Phase 3 failover execution, and order rollback.
      • Principal DBRE (Marcus Vance): Sole authority to evaluate pre-flight database lag and execute cross-region promotion.
      • Core Tech Lead (Elena Rostova): Sole authority to verify EKS compute scaling and validate Route 53 DNS weight changes.
    • Outputs: Authenticated session log on DR incident bridge with timestamped command ownership.
    • Owner: David O'Reilly (Incident Commander).

    Failure Handling: If any primary role holder is unresponsive after 3 minutes, escalation tree pages designated secondary lead.

    Verification: Reviewer self-check confirming named individuals and role boundaries align with incident governance.

    2. Source Authority [MC-SA-01]
    • Inputs: AWS CloudWatch metrics, Route 53 control plane, Aurora Global Database cluster topology.
    • Algorithm: The playbook derives diagnostic state exclusively from authoritative telemetry endpoints:
      • Database lag source: AWS CloudWatch metric AuroraGlobalDBReplicationLag queried via AWS CLI v2 with elevated role AWS-DR-Automation-Role.
      • Ingress health source: CloudWatch Synthetics canary canary-payment-auth-probe querying endpoint https://api.payments.internal/v3/health.
    • Outputs: Cryptographically signed execution tokens passed to linked runbooks.
    • Owner: Marcus Vance (Principal DBRE).
    • Failure Handling: If CloudWatch API throttles, fallback to direct query via read-only RDS bastion endpoint.
    • Verification: Sourced CLI command assertions returning exit code 0 on active telemetry endpoints.
    3. Document Structure [MC-DS-01]
    • Inputs: 5-phase sequential execution lifecycle.
    • Algorithm: The playbook enforces rigid chronological phase progression:
      1. Phase 1: Activation & Triage: Trigger signal verification and war room assembly.
      2. Phase 2: Pre-Flight Qualification: Read-only checks verifying replication lag < 15.0s and warm-standby pod capacity.
      3. Phase 3: Execution: Coordinated invocation of canonical runbooks in sequence:
        • Step 3.1: Promote Aurora replica cluster in us-west-2 (RB-AURORA-PROMOTE-02).
        • Step 3.2: Re-point payment consumer workers to local Kafka cluster in us-west-2 (RB-KAFKA-SWING-01).
        • Step 3.3: Shift Route 53 ARC routing controls 100% to us-west-2 (RB-R53-CUTOVER-03).
      4. Phase 4: Post-Cutover Verification: 5-minute rollback observation window.
      5. Phase 5: Ledger Reconciliation: Asynchronous queue reconciliation confirming zero missing transactions.
    • Outputs: Phase transition log committed to incident record.
    • Owner: Elena Rostova (Core Tech Lead).
    • Failure Handling: Any failed step halts the phase progression immediately and escalates to David O'Reilly.
    • Verification: Phase completion checklist verified prior to next-phase dispatch.
    4. Freshness [MC-FR-01]
    • Inputs: Git commit revisions of referenced runbooks, AWS infrastructure resource IDs, DNS zone identifiers.
    • Algorithm: Playbook binds exact immutable resource hashes:
      • Primary Aurora Cluster: arn:aws:rds:us-east-1:123456789012:cluster:pay-core-primary-cluster (rev 2026-08-30).
      • Standby Aurora Cluster: arn:aws:rds:us-west-2:123456789012:cluster:pay-core-standby-cluster (rev 2026-08-30).
      • Route 53 Control Panel: arn:aws:route53recoverycontrol::123456789012:controlpanel/pay-gate-panel (rev 2026-09-01).
    • Outputs: Validation of runbook link hashes prior to manual command execution.
    • Owner: Elena Rostova (Core Tech Lead).

    Failure Handling: If an upstream runbook hash has diverged by > 30 days without re-certification, playbook flags manual inspection step.

    • Verification: Automated git hook verify_playbook_links.sh checking runbook path existence.

    Adversarial Cases and Routing

    1. Reject Duplicated Truth [ADV-DT-01]

    Vulnerability: Embedding inline bash scripts or raw AWS CLI commands for database promotion directly inside the playbook document. Over time, database engine parameters change in the primary runbook, leaving the playbook with stale, destructive command syntax.

    Enforcement & Diagnostic: The playbook is strictly forbidden from embedding inline database failover scripts or low-level AWS CLI parameters. All execution steps MUST reference canonical runbook documents by stable ID and file path:

    • Database promotion MUST route to RB-AURORA-PROMOTE-02.
    • DNS traffic shift MUST route to RB-R53-CUTOVER-03.
    • If inline CLI commands for database promotion are detected in the playbook, CI linter emits diagnostic ERR_DUPLICATED_TRUTH_DETECTED and rejects the build.

    Forbidden Output Behavior: Documenting inline SQL commands, RDS API parameter blocks, or custom python failover scripts inside this scenario playbook.

    2. Reject Unverifiable Example [ADV-UE-01]

    Vulnerability: Proposing speculative or hypothetical verification steps (e.g. "ensure all transactions look healthy" or "verify traffic seems normal") that lack objective, reproducible telemetry queries.

    Enforcement & Diagnostic: Every verification gate must provide an exact, executable query and numerical threshold:

    • Pre-flight lag query: aws cloudwatch get-metric-data evaluating metric AuroraGlobalDBReplicationLag (must be < 15.0s).
    • Post-cutover error rate query: Prometheus PromQL sum(rate(payment_auth_errors_total[1m])) / sum(rate(payment_auth_total[1m])) (must be < 0.002).
    • Subjective verification statements trigger diagnostic ERR_UNVERIFIABLE_EXAMPLE_REJECTED.

    Forbidden Output Behavior: Unquantified checklist items or subjective visual confirmations in operational decision branches.

    3. Reject Orphan Document [ADV-OD-01]

    Vulnerability: Publishing an operational playbook in an unindexed repository branch without linking it to the system architecture registry, operational tier catalog, or active incident response escalation directories.

    Enforcement & Diagnostic: This playbook must register its path in docs/architecture/playbooks/INDEX.md and bind to the service catalog entry for Core Payments Platform v3.4. If an unindexed playbook is created, documentation audit emits diagnostic ERR_ORPHAN_DOCUMENT_DETECTED.

    Forbidden Output Behavior: Generating detached playbook files without bidirectional links to service architecture and incident management indexes.


    Playbook Execution Procedures & Checkpoints

    Phase 1: Signal Qualification & Entry Conditions
    • Activation Criteria: Playbook activates ONLY when both conditions hold:
      1. Primary region us-east-1 health check fails continuously for $\ge 240$ seconds (4 minutes).
      2. P1 incident declared by Incident Commander David O'Reilly.
    • Authorized Deciders: David O'Reilly (Incident Commander) with consensus from Marcus Vance (DBRE).
    Phase 2: Read-Only Pre-Flight Qualification Checkpoint [CP-PRE-01]

    Execute from secure bastion in us-west-2. Read-only operation. Do not mutate state.

    1. Query Aurora Global Database Replication Lag:
      aws cloudwatch get-metric-data \
        --region us-west-2 \
        --metric-data-queries file://queries/aurora-global-lag.json \
        --start-time $(date -u -d '5 minutes ago' +%Y-%m-%dT%H:%M:%SZ) \
        --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
        --query "MetricDataResults[0].Values[0]" --output text
      

    Abort Gate: If reported lag is $\ge 15.0$ seconds,

    ABORT REGIONAL FAILOVER IMMEDIATELY. Marcus Vance informs Incident Commander of severe data loss risk.

    • Proceed Gate: If reported lag is $< 15.0$ seconds, proceed to compute verification.
    1. Verify EKS Workload Capacity in us-west-2:
      kubectl get deployment pay-core-auth -n payments --context=us-west-2 \
        -o jsonpath='{.status.readyReplicas}'
      

    Proceed Gate: Ready replicas must equal exactly 120 pods. If $< 120$, trigger emergency horizontal scaling before cutover.

    Phase 3: Coordinated Execution & Runbook Dispatch

    Execute sequentially upon David O'Reilly's explicit go-ahead.

    1. Step 3.1: Promote Database Cluster:
      • Dispatch Marcus Vance to execute canonical runbook RB-AURORA-PROMOTE-02.
      • Expected duration: 180 seconds.
      • Confirmation Oracle: aws rds describe-db-clusters in us-west-2 reports Status: available and IsClusterWriter: true.
    2. Step 3.2: Re-point Streaming Consumers:
      • Dispatch Elena Rostova to execute RB-KAFKA-SWING-01.
      • Verify consumer groups bind to local MSK broker cluster msk-payments-uswest2.
    3. Step 3.3: Route 53 Traffic Cutover:
      • Dispatch Elena Rostova to execute RB-R53-CUTOVER-03.
      • Set routing control us-east-1 to Off (0%) and us-west-2 to On (100%).
    Phase 4: Post-Cutover Rollback Tripwires [CP-POST-01]
    • 5-Minute Canary Monitoring Window:
      • Telemetry monitor evaluates PromQL query every 10 seconds:
        sum(rate(payment_authorization_errors_total{region="us-west-2"}[1m]))
        / sum(rate(payment_authorization_requests_total{region="us-west-2"}[1m]))
        

    Automated Rollback Tripwire: If error rate $\ge 0.010$ (1.0%) for 3 consecutive intervals,

    TRIGGER EMERGENCY ROLLBACK:
    1. Elena Rostova reverts Route 53 routing controls to 50/50 split or returns to primary if primary compute has recovered.
    2. Shed non-critical payment lanes via API gateway rate-limiting.

    Phase 5: Ledger Queue Reconciliation & Exit Criteria
    • Scenario Exit Criteria:
      1. Payment authorization traffic sustained at 4,500 req/sec in us-west-2 with p99 latency $\le 120$ ms for 30 consecutive minutes.
      2. Authorization error rate evaluates $< 0.2%$ ($< 0.002$).
      3. Continuous ledger reconciliation script verify_clearing_balances.py confirms

    0 lost financial mutations between us-east-1 replay logs and us-west-2 Aurora tables.


    Explicit Unknowns

    • Exact physical fiber severance duration during major transatlantic or transcontinental carrier outages impacting AWS backbones (G-1).
    • Unpredictable Route 53 client DNS caching behavior on third-party mobile banking apps that ignore 30-second DNS TTLs (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    4,500 authorizations/sec peak workloadprovidedPlatform intake specificationCurrent
    Primary us-east-1 and warm-standby us-west-2providedArchitecture topology intakeCurrent
    4-minute sustained outage triggerprovidedIncident policy rule2026-09-15
    Hard abort ceiling on replication lag >= 15.0sdecidedMarcus Vance (Principal DBRE)2026-09-15
    120-pod warm-standby EKS capacity requirementderived4,500 TPS ÷ 40 TPS/pod capacity limit2026-09-15
    Post-cutover automated rollback at >= 1.0% error ratedecidedDavid O'Reilly & Elena Rostova2026-09-15
    Prohibition of inline SQL/CLI promotion scriptsdecidedArchitectural invariant ADV-DT-012026-09-15

    Verification

    GateCommandExitEvidence time
    Structure & Link Validationbash scripts/verify_playbook_links.sh --playbook dr-playbook-document.md02026-09-15T10:30:00Z
    Prometheus Metric Query Syntax Checkpromtool check metrics queries/payment-error-rate.promql02026-09-15T10:32:00Z
    Pre-Flight CloudWatch Query Syntax Checkaws cloudwatch get-metric-data --cli-input-json file://queries/aurora-global-lag.json --dry-run02026-09-15T10:34:00Z

    Reviewer self-check against playbook authoring standards:

    • Decision Trees & Branching: PASS. Explicit go/abort decision branches based on 15s replication lag.

    Runbook Decoupling: PASS. Complex commands delegated to canonical runbooks ([RB-AURORA-PROMOTE-02], [RB-KAFKA-SWING-01], [RB-R53-CUTOVER-03]).

    • Safety Safeguards: PASS. Automated rollback tripwire defined for 5-minute post-cutover window at >= 1.0% errors.

    Adversarial Rules: PASS. Duplicated truth, unverifiable examples, and orphan documents formally addressed and rejected.

    Open Decisions

    • DEC-DR-01: David O'Reilly to finalize whether Route 53 Application Recovery Controller (ARC) zonal autoshift should be authorized to trigger Phase 3 failover without human IC approval during off-hours (Owner: David O'Reilly).

    Next steps

    1. Marcus Vance executes scheduled dry-run DR simulation DR-STG-45 in staging environment to validate 9.8-minute RTO.
    2. Elena Rostova verifies Route 53 ARC health check routing rules are committed in terraform repository.
    3. SRE team schedules quarterly tabletop rehearsal reviewing Phase 2 abort communication protocols.

    operational-playbook-authoring.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Map incident roles and communication channels for outagesStructure decision trees for complex system failoversLink canonical runbooks into coordinated scenario guidesDefine explicit entry and exit criteria for DR events

    About this skill

    What it does

    This skill composes accepted operational scenarios, decision branches, roles, communications and linked procedures into a safe coordination guide. It does not design incident, security, change or disaster-recovery policy and does not execute operations.

    Use it when

    Use when responders need a coordinated, scenario-oriented guide spanning several decisions, actors or bounded procedures.

    For example: “We need an operational playbook for handling live video transcoding cluster degradation during high-traffic streaming events, including CDN profile fallback branches.”

    What you get

    • DR Playbook Document

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/playbook-authoring/.

    What it will not do

    Does not invent operations policy. Do not use for one narrow command runbook, SOP/checklist, incident/security/DR strategy, policy, tutorial or live execution.

    How it works

    1. Check operational scenario and policy exist.
    2. Bound scenario entry conditions and trigger signals.
    3. Map incident roles, authorities, and communication channels.
    4. Structure decision trees and branching logic.
    5. Link canonical runbooks for execution steps.
    6. Define scenario exit conditions, recovery validation, and handoffs.
    7. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions