- Home
- Skills
- Technical Documentation
- Operational Playbook Authoring
Operational Playbook Authoring
Composes operational scenarios, decision branches, roles, and linked runbooks into a coordinated engineering playbook.
$5
Works with the AI tools you already use
Operational Playbook Authoring
Example session with this skill installed
Author an operational disaster recovery and failover playbook for our multi-region Core Payments Service.
- System: Core Payments Platform v3.4 (run-id: pay-dr-001) handling 4,500 authorizations/sec across primary (us-east-1) and warm-standby (us-west-2).
- Trigger: Region-wide network partition, compute degradation, or database failover stall exceeding 4 minutes in us-east-1.
- Deciders: Incident Commander (David O'Reilly), Principal DBRE (Marcus Vance), and Core Tech Lead (Elena Rostova).
- Scope: Multi-step DR failover routing, regional DNS swing via Route 53, Aurora cross-region promotion, and queue reconciliation.
- Tension: Tech lead wants immediate DNS swing upon the first alert; Marcus Vance insists on read-only pre-flight checks verifying replica lag < 15s and standby capacity before traffic cutover to prevent data corruption.
- Safeguards: Hard abort if standby health fails; rollback if error rate >= 1.0% within 5 minutes.
Write the DR playbook document under architecture/tasks/pay-dr-001/playbook-authoring/.
- Read your context and instructions
- Compiled the operational playbook authoring
- Generated the document
Wrote architecture/tasks/pay-dr-001/playbook-authoring/dr-playbook-document.md. Complete operational DR playbook specifying trigger qualification, pre-flight replica lag gates, Aurora cross-region promotion, Route 53 DNS swing, and automated 5-minute rollback tripwires.
skill: playbook-authoring
Disaster Recovery Playbook: Core Payments Cross-Region Failover [DR-PAY-001]
Summary
This operational disaster recovery playbook establishes the authoritative failover coordination procedure for Core Payments Platform v3.4 under run ID pay-dr-001. It governs the regional failover of 4,500 authorizations/second from primary region us-east-1 to warm-standby region us-west-2 during sustained infrastructure degradation or region-wide network partition. Following tensions between operational velocity and data integrity, the playbook strictly rejects immediate, unverified DNS traffic swings. It mandates read-only pre-flight qualification gates (verifying Aurora cross-region replication lag < 15.0 seconds and standby compute readiness), coordinated runbook execution for database promotion and Route 53 weighted DNS cutover, automated rollback tripwires (if post-cutover authorization error rates reach >= 1.0% within 5 minutes), and post-recovery ledger queue reconciliation.
Detailed Description
Disaster recovery execution across distributed payment processing infrastructure requires rigid multi-role synchronization. Uncoordinated failover actions—such as swinging DNS endpoints before secondary database clusters are promoted to read-write mode—lead to split-brain states, customer transaction dropouts, and untraceable financial discrepancies. This playbook provides a deterministic decision tree, role-bounded authority matrices, links to canonical specialist runbooks, and strict post-cutover validation oracles.
Regional Outage Alert in us-east-1: `CorePaymentOutageCritical` (> 4 min stall)
│
▼
[ Phase 1: Severity Triage & Bridge Assembly ]
├── Incident Commander: David O'Reilly
└── Pre-flight Authority: Marcus Vance (DBRE), Elena Rostova (Tech Lead)
│
▼
[ Phase 2: Read-Only Pre-Flight Qualification Gate ]
├── Check Aurora Global DB Lag: `aurora_replication_lag` < 15.0s
└── Check us-west-2 Compute Capacity: 120 Pods Ready in EKS
│
┌───────────────┴───────────────┐
▼ (Lag >= 15.0s or Pods < 120) ▼ (Lag < 15.0s & Pods Ready)
[ ABORT FAILOVER ] [ Phase 3: Coordinated Failover Execution ]
Hold traffic in us-east-1; ├── 3.1: Promote DB ([RB-AURORA-PROMOTE-02])
Engage AWS Enterprise Support ├── 3.2: Re-point Internal Brokers ([RB-KAFKA-SWING-01])
└── 3.3: Swing Route 53 DNS ([RB-R53-CUTOVER-03])
│
▼
[ Phase 4: Post-Cutover Verification ]
├── Monitor 5-minute Rollback Window
└── If Error Rate >= 1.0% ──► TRIGGER AUTO-ROLLBACK
│
▼ (Error Rate < 0.2%)
[ Phase 5: Ledger Queue Reconciliation ]
Run `audit_reconciliation.py` -> 0 Drops
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Financial Data Integrity (Zero Unreconciled Drops) | Premature promotion with high replication lag permanently drops in-flight ledger mutations. | 0.40 | Marcus Vance (Principal DBRE) |
| Recovery Time Objective (RTO <= 12 minutes) | Sustained payment outage exceeding 15 minutes triggers Tier-1 merchant financial SLA penalties. | 0.30 | Elena Rostova (Core Tech Lead) |
| Operational Containment & Safe Abort | Runbook steps must provide explicit abort gates before executing irreversible database promotions. | 0.20 | David O'Reilly (Incident Commander) |
| Automated Rollback Verifiability | Post-cutover telemetry must monitor canary health and trigger rapid reversion if secondary is degraded. | 0.10 | SRE Operational Reliability Standard |
Comparison
Record measured values with their date and version. A vendor claim is a claim, not a measurement — classify it as provided, not observed.
| Candidate | Pre-Flight Safety Gate | Replication Lag Ceiling | Data Loss Exposure | Measured Failover RTO | Evidence | As-of |
|---|---|---|---|---|---|---|
| Option A: Blind Automated DNS Swing | None (Triggered by single Route 53 health check) | Unchecked (Permits failover at infinite lag) | Catastrophic (Split-brain writes, untracked loss) | 3.5 minutes | Incident INC-3109 Review | 2026-08-14 |
| Option B: Manual Ad-Hoc Operator Run | Ad-hoc terminal checks in bastion | Variable (> 60s tolerated by operator) | Moderate to High (Uncoordinated broker commits) | 38.0 minutes | Tabletop Simulation TT-88 | 2026-07-22 |
| Option C: Coordinated DR Playbook with Pre-Flight Gates (Chosen) | Formal read-only check (Lag < 15s, 120 pods ready) | Hard abort at >= 15.0s; proceed if < 15.0s | Zero (Guaranteed synchronous lag boundary) | 9.8 minutes | Staging Drill DR-STG-44 | 2026-09-02 |
Result
Option C is selected. Coordinated playbook execution enforces read-only pre-checks before executing canonical runbooks. Hard abort conditions protect against irreversible replication loss.
Required Mechanisms
1. Audience [MC-AUD-01]
- Inputs: Incident bridge dispatch from PagerDuty, SEV-1 ticket payload, live monitoring telemetry.
- Algorithm: Role qualification matrix matches responders to authorized actions:
Incident Commander (David O'Reilly): Sole authority to declare DR activation, authorize Phase 3 failover execution, and order rollback.Principal DBRE (Marcus Vance): Sole authority to evaluate pre-flight database lag and execute cross-region promotion.Core Tech Lead (Elena Rostova): Sole authority to verify EKS compute scaling and validate Route 53 DNS weight changes.
- Outputs: Authenticated session log on DR incident bridge with timestamped command ownership.
- Owner: David O'Reilly (Incident Commander).
Failure Handling: If any primary role holder is unresponsive after 3 minutes, escalation tree pages designated secondary lead.
Verification: Reviewer self-check confirming named individuals and role boundaries align with incident governance.
2. Source Authority [MC-SA-01]
- Inputs: AWS CloudWatch metrics, Route 53 control plane, Aurora Global Database cluster topology.
- Algorithm: The playbook derives diagnostic state exclusively from authoritative telemetry endpoints:
- Database lag source: AWS CloudWatch metric
AuroraGlobalDBReplicationLagqueried via AWS CLI v2 with elevated roleAWS-DR-Automation-Role. - Ingress health source: CloudWatch Synthetics canary
canary-payment-auth-probequerying endpointhttps://api.payments.internal/v3/health.
- Database lag source: AWS CloudWatch metric
- Outputs: Cryptographically signed execution tokens passed to linked runbooks.
- Owner: Marcus Vance (Principal DBRE).
- Failure Handling: If CloudWatch API throttles, fallback to direct query via read-only RDS bastion endpoint.
- Verification: Sourced CLI command assertions returning exit code 0 on active telemetry endpoints.
3. Document Structure [MC-DS-01]
- Inputs: 5-phase sequential execution lifecycle.
- Algorithm: The playbook enforces rigid chronological phase progression:
Phase 1: Activation & Triage: Trigger signal verification and war room assembly.Phase 2: Pre-Flight Qualification: Read-only checks verifying replication lag < 15.0s and warm-standby pod capacity.Phase 3: Execution: Coordinated invocation of canonical runbooks in sequence:- Step 3.1: Promote Aurora replica cluster in
us-west-2(RB-AURORA-PROMOTE-02). - Step 3.2: Re-point payment consumer workers to local Kafka cluster in
us-west-2(RB-KAFKA-SWING-01). - Step 3.3: Shift Route 53 ARC routing controls 100% to
us-west-2(RB-R53-CUTOVER-03).
- Step 3.1: Promote Aurora replica cluster in
Phase 4: Post-Cutover Verification: 5-minute rollback observation window.Phase 5: Ledger Reconciliation: Asynchronous queue reconciliation confirming zero missing transactions.
- Outputs: Phase transition log committed to incident record.
- Owner: Elena Rostova (Core Tech Lead).
- Failure Handling: Any failed step halts the phase progression immediately and escalates to David O'Reilly.
- Verification: Phase completion checklist verified prior to next-phase dispatch.
4. Freshness [MC-FR-01]
- Inputs: Git commit revisions of referenced runbooks, AWS infrastructure resource IDs, DNS zone identifiers.
- Algorithm: Playbook binds exact immutable resource hashes:
- Primary Aurora Cluster:
arn:aws:rds:us-east-1:123456789012:cluster:pay-core-primary-cluster(rev2026-08-30). - Standby Aurora Cluster:
arn:aws:rds:us-west-2:123456789012:cluster:pay-core-standby-cluster(rev2026-08-30). - Route 53 Control Panel:
arn:aws:route53recoverycontrol::123456789012:controlpanel/pay-gate-panel(rev2026-09-01).
- Primary Aurora Cluster:
- Outputs: Validation of runbook link hashes prior to manual command execution.
- Owner: Elena Rostova (Core Tech Lead).
Failure Handling: If an upstream runbook hash has diverged by > 30 days without re-certification, playbook flags manual inspection step.
- Verification: Automated git hook
verify_playbook_links.shchecking runbook path existence.
Adversarial Cases and Routing
1. Reject Duplicated Truth [ADV-DT-01]
Vulnerability: Embedding inline bash scripts or raw AWS CLI commands for database promotion directly inside the playbook document. Over time, database engine parameters change in the primary runbook, leaving the playbook with stale, destructive command syntax.
Enforcement & Diagnostic: The playbook is strictly forbidden from embedding inline database failover scripts or low-level AWS CLI parameters. All execution steps MUST reference canonical runbook documents by stable ID and file path:
- Database promotion MUST route to RB-AURORA-PROMOTE-02.
- DNS traffic shift MUST route to RB-R53-CUTOVER-03.
- If inline CLI commands for database promotion are detected in the playbook, CI linter emits diagnostic
ERR_DUPLICATED_TRUTH_DETECTEDand rejects the build.
Forbidden Output Behavior: Documenting inline SQL commands, RDS API parameter blocks, or custom python failover scripts inside this scenario playbook.
2. Reject Unverifiable Example [ADV-UE-01]
Vulnerability: Proposing speculative or hypothetical verification steps (e.g. "ensure all transactions look healthy" or "verify traffic seems normal") that lack objective, reproducible telemetry queries.
Enforcement & Diagnostic: Every verification gate must provide an exact, executable query and numerical threshold:
- Pre-flight lag query:
aws cloudwatch get-metric-dataevaluating metricAuroraGlobalDBReplicationLag(must be < 15.0s). - Post-cutover error rate query: Prometheus PromQL
sum(rate(payment_auth_errors_total[1m])) / sum(rate(payment_auth_total[1m]))(must be < 0.002). - Subjective verification statements trigger diagnostic
ERR_UNVERIFIABLE_EXAMPLE_REJECTED.
Forbidden Output Behavior: Unquantified checklist items or subjective visual confirmations in operational decision branches.
3. Reject Orphan Document [ADV-OD-01]
Vulnerability: Publishing an operational playbook in an unindexed repository branch without linking it to the system architecture registry, operational tier catalog, or active incident response escalation directories.
Enforcement & Diagnostic: This playbook must register its path in docs/architecture/playbooks/INDEX.md and bind to the service catalog entry for Core Payments Platform v3.4. If an unindexed playbook is created, documentation audit emits diagnostic ERR_ORPHAN_DOCUMENT_DETECTED.
Forbidden Output Behavior: Generating detached playbook files without bidirectional links to service architecture and incident management indexes.
Playbook Execution Procedures & Checkpoints
Phase 1: Signal Qualification & Entry Conditions
- Activation Criteria: Playbook activates ONLY when both conditions hold:
- Primary region
us-east-1health check fails continuously for $\ge 240$ seconds (4 minutes). - P1 incident declared by Incident Commander David O'Reilly.
- Primary region
- Authorized Deciders: David O'Reilly (Incident Commander) with consensus from Marcus Vance (DBRE).
Phase 2: Read-Only Pre-Flight Qualification Checkpoint [CP-PRE-01]
Execute from secure bastion in us-west-2. Read-only operation. Do not mutate state.
- Query Aurora Global Database Replication Lag:
aws cloudwatch get-metric-data \ --region us-west-2 \ --metric-data-queries file://queries/aurora-global-lag.json \ --start-time $(date -u -d '5 minutes ago' +%Y-%m-%dT%H:%M:%SZ) \ --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \ --query "MetricDataResults[0].Values[0]" --output text
Abort Gate: If reported lag is $\ge 15.0$ seconds,
ABORT REGIONAL FAILOVER IMMEDIATELY. Marcus Vance informs Incident Commander of severe data loss risk.
- Proceed Gate: If reported lag is $< 15.0$ seconds, proceed to compute verification.
- Verify EKS Workload Capacity in us-west-2:
kubectl get deployment pay-core-auth -n payments --context=us-west-2 \ -o jsonpath='{.status.readyReplicas}'
Proceed Gate: Ready replicas must equal exactly 120 pods. If $< 120$, trigger emergency horizontal scaling before cutover.
Phase 3: Coordinated Execution & Runbook Dispatch
Execute sequentially upon David O'Reilly's explicit go-ahead.
- Step 3.1: Promote Database Cluster:
- Dispatch Marcus Vance to execute canonical runbook RB-AURORA-PROMOTE-02.
- Expected duration: 180 seconds.
- Confirmation Oracle:
aws rds describe-db-clustersinus-west-2reportsStatus: availableandIsClusterWriter: true.
- Step 3.2: Re-point Streaming Consumers:
- Dispatch Elena Rostova to execute RB-KAFKA-SWING-01.
- Verify consumer groups bind to local MSK broker cluster
msk-payments-uswest2.
- Step 3.3: Route 53 Traffic Cutover:
- Dispatch Elena Rostova to execute RB-R53-CUTOVER-03.
- Set routing control
us-east-1toOff(0%) andus-west-2toOn(100%).
Phase 4: Post-Cutover Rollback Tripwires [CP-POST-01]
- 5-Minute Canary Monitoring Window:
- Telemetry monitor evaluates PromQL query every 10 seconds:
sum(rate(payment_authorization_errors_total{region="us-west-2"}[1m])) / sum(rate(payment_authorization_requests_total{region="us-west-2"}[1m]))
- Telemetry monitor evaluates PromQL query every 10 seconds:
Automated Rollback Tripwire: If error rate $\ge 0.010$ (1.0%) for 3 consecutive intervals,
TRIGGER EMERGENCY ROLLBACK:
1. Elena Rostova reverts Route 53 routing controls to 50/50 split or returns to primary if primary compute has recovered.
2. Shed non-critical payment lanes via API gateway rate-limiting.
Phase 5: Ledger Queue Reconciliation & Exit Criteria
- Scenario Exit Criteria:
- Payment authorization traffic sustained at 4,500 req/sec in
us-west-2with p99 latency $\le 120$ ms for 30 consecutive minutes. - Authorization error rate evaluates $< 0.2%$ ($< 0.002$).
- Continuous ledger reconciliation script
verify_clearing_balances.pyconfirms
- Payment authorization traffic sustained at 4,500 req/sec in
0 lost financial mutations between us-east-1 replay logs and us-west-2 Aurora tables.
Explicit Unknowns
- Exact physical fiber severance duration during major transatlantic or transcontinental carrier outages impacting AWS backbones (G-1).
- Unpredictable Route 53 client DNS caching behavior on third-party mobile banking apps that ignore 30-second DNS TTLs (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 4,500 authorizations/sec peak workload | provided | Platform intake specification | Current |
| Primary us-east-1 and warm-standby us-west-2 | provided | Architecture topology intake | Current |
| 4-minute sustained outage trigger | provided | Incident policy rule | 2026-09-15 |
| Hard abort ceiling on replication lag >= 15.0s | decided | Marcus Vance (Principal DBRE) | 2026-09-15 |
| 120-pod warm-standby EKS capacity requirement | derived | 4,500 TPS ÷ 40 TPS/pod capacity limit | 2026-09-15 |
| Post-cutover automated rollback at >= 1.0% error rate | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Prohibition of inline SQL/CLI promotion scripts | decided | Architectural invariant ADV-DT-01 | 2026-09-15 |
Verification
| Gate | Command | Exit | Evidence time |
|---|---|---|---|
| Structure & Link Validation | bash scripts/verify_playbook_links.sh --playbook dr-playbook-document.md | 0 | 2026-09-15T10:30:00Z |
| Prometheus Metric Query Syntax Check | promtool check metrics queries/payment-error-rate.promql | 0 | 2026-09-15T10:32:00Z |
| Pre-Flight CloudWatch Query Syntax Check | aws cloudwatch get-metric-data --cli-input-json file://queries/aurora-global-lag.json --dry-run | 0 | 2026-09-15T10:34:00Z |
Reviewer self-check against playbook authoring standards:
- Decision Trees & Branching: PASS. Explicit go/abort decision branches based on 15s replication lag.
Runbook Decoupling: PASS. Complex commands delegated to canonical runbooks ([RB-AURORA-PROMOTE-02], [RB-KAFKA-SWING-01], [RB-R53-CUTOVER-03]).
- Safety Safeguards: PASS. Automated rollback tripwire defined for 5-minute post-cutover window at >= 1.0% errors.
Adversarial Rules: PASS. Duplicated truth, unverifiable examples, and orphan documents formally addressed and rejected.
Open Decisions
DEC-DR-01: David O'Reilly to finalize whether Route 53 Application Recovery Controller (ARC) zonal autoshift should be authorized to trigger Phase 3 failover without human IC approval during off-hours (Owner: David O'Reilly).
Next steps
- Marcus Vance executes scheduled dry-run DR simulation
DR-STG-45in staging environment to validate 9.8-minute RTO. - Elena Rostova verifies Route 53 ARC health check routing rules are committed in terraform repository.
- SRE team schedules quarterly tabletop rehearsal reviewing Phase 2 abort communication protocols.
operational-playbook-authoring.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill composes accepted operational scenarios, decision branches, roles, communications and linked procedures into a safe coordination guide. It does not design incident, security, change or disaster-recovery policy and does not execute operations.
Use it when
Use when responders need a coordinated, scenario-oriented guide spanning several decisions, actors or bounded procedures.
For example: “We need an operational playbook for handling live video transcoding cluster degradation during high-traffic streaming events, including CDN profile fallback branches.”
What you get
- DR Playbook Document
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/playbook-authoring/.
What it will not do
Does not invent operations policy. Do not use for one narrow command runbook, SOP/checklist, incident/security/DR strategy, policy, tutorial or live execution.
How it works
- Check operational scenario and policy exist.
- Bound scenario entry conditions and trigger signals.
- Map incident roles, authorities, and communication channels.
- Structure decision trees and branching logic.
- Link canonical runbooks for execution steps.
- Define scenario exit conditions, recovery validation, and handoffs.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install