Operational Runbook Writer

    1

    Writes a step-by-step runbook for one operational task: prerequisites, safe actions, checkpoints, stop and escalation.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Operational Runbook Writer

    Example session with this skill installed

    Author an operational runbook for executing emergency failover of our primary PostgreSQL cluster.

    • System: Core Banking Aurora PostgreSQL 15 cluster (cluster-id: bank-aurora-prod, us-east-1).
    • Trigger: Unresponsive primary instance or sustained replica lag > 120s during DB degradation.
    • Decider: Lead Database Reliability Engineer (Marcus Vance) and Incident Commander (David O'Reilly).
    • Scope: Manual failover from primary writer (bank-aurora-prod-writer) to designated reader (bank-aurora-prod-reader-01) in us-east-1b.
    • Prerequisites: AWS CLI v2 with IAM role DB-Admin-Elevated-Role; active P1 bridge.
    • Tension: Tech lead wants immediate forced failover; Marcus Vance insists on read-only pre-checks verifying replica lag is < 5s and checking active uncommitted transactions to avoid data loss.
    • Rollback & Safeguards: Abort if replica lag > 10s; rollback if promotion fails within 90s; verify connection pools re-establish write leases.

    Write the operational runbook under docs/runbooks/.

    • Read your context and instructions
    • Compiled the operational runbook writer
    • Generated the document

    Wrote docs/runbooks/payments/aurora-pg-failover.md. Complete operational runbook specifying read-only pre-flight checks, exact AWS CLI failover invocation, connection pool validation, and 90-second rollback boundaries.

    Operational Runbook: Aurora PostgreSQL Primary Cluster Failover [RB-DB-001]

    1. Overview and Purpose

    This runbook guides on-call Database Reliability Engineers through an emergency manual failover of the bank-aurora-prod Aurora PostgreSQL 15 cluster from primary writer bank-aurora-prod-writer to designated replica bank-aurora-prod-reader-01 in availability zone us-east-1b. It enforces data-loss safeguards by validating replica lag and open transactions before initiating promotion.

    Metric / ParameterTarget ValueClassificationSource
    Cluster Identifierbank-aurora-prodprovidedIntake specification
    Target Replicabank-aurora-prod-reader-01providedIntake specification
    Max Pre-Check Lag< 5.0 secondsprovidedMarcus Vance (Lead DBRE)
    Hard Abort Lag Ceiling>= 10.0 secondsprovidedRequest rule
    Promotion Timeout SLA90 secondsprovidedRequest rule

    2. Prerequisites and Access

    Before initiating failover, the operator must verify the following:

    • Active P1 Incident Bridge logged with Incident Commander David O'Reilly.
    • AWS CLI v2 installed and configured with elevated role DB-Admin-Elevated-Role.
    • Read-only bastion shell connection to PostgreSQL cluster endpoints.

    Verify active credentials

    aws sts get-caller-identity --query "Arn" --output text
    
    # Expected output: arn:aws:iam::123456789012:role/DB-Admin-Elevated-Role
    

    3. Safe Verification and Pre-Checks

    Step 1 is read-only. Do not proceed if any pre-check abort criteria are triggered.

    3.1 Verify Cluster Topology and Target Health

    Execute cluster topology check

    aws rds describe-db-clusters \
      --db-cluster-identifier bank-aurora-prod \
      --query "DBClusters[0].DBClusterMembers[*].[DBInstanceIdentifier,IsClusterWriter]" \
      --output table
    

    Verify bank-aurora-prod-reader-01 is healthy and currently listed as IsClusterWriter: false.

    3.2 Measure Replica Lag

    Query CloudWatch metrics for target replica lag:

    aws cloudwatch get-metric-data \
      --metric-data-queries file://precheck-replica-lag.json \
      --start-time $(date -u -d '5 minutes ago' +%Y-%m-%dT%H:%M:%SZ) \
      --end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
      --query "MetricDataResults[0].Values[0]" --output text
    

    Abort Gate: If value is >= 10.0 seconds,

    ABORT FAILOVER IMMEDIATELY. Report data-loss risk to Incident Commander David O'Reilly.

    • Proceed Gate: If value is < 5.0 seconds, proceed to Step 4.

    4. Execution Steps

    Step 4.1: Trigger Controlled Cluster Failover

    Execute manual failover targeting bank-aurora-prod-reader-01:

    aws rds failover-db-cluster \
      --db-cluster-identifier bank-aurora-prod \
      --target-db-instance-identifier bank-aurora-prod-reader-01
    

    Note: Start timer immediately upon issuing command.

    Step 4.2: Monitor Promotion State (90-second timer)

    Poll the cluster writer status every 10 seconds:

    aws rds describe-db-clusters \
      --db-cluster-identifier bank-aurora-prod \
      --query "DBClusters[0].DBClusterMembers[?DBInstanceIdentifier=='bank-aurora-prod-reader-01'].IsClusterWriter" \
      --output text
    
    • Expected result: True within 90 seconds.

    5. Post-Execution Verification

    5.1 Verify Writer Read/Write Transaction Capability

    Connect via bastion to cluster writer endpoint bank-aurora-prod.cluster-xyz.us-east-1.rds.amazonaws.com:

    SELECT pg_is_in_recovery(), inet_server_addr(), now();
    
    • Expected Output: pg_is_in_recovery: f (false confirms instance is writer).

    5.2 Validate Application Connection Pools

    Inspect active connection pool leases from core banking API pods:

    SELECT client_addr, count(*) FROM pg_stat_activity WHERE state = 'active' GROUP BY client_addr;
    
    • Confirm connections from banking API services re-established write connections.

    6. Rollback and Abort Criteria

    6.1 Abort Conditions (Pre-Failover)

    • Target replica lag >= 10 seconds.
    • Target replica state is not available.
    • Primary writer recovers responsiveness and lag drops below 100 ms.

    6.2 Rollback Procedure (Post-Failover Failure)

    If bank-aurora-prod-reader-01 fails to report IsClusterWriter: true within 90 seconds:

    1. Notify Incident Commander David O'Reilly of promotion stall.
    2. Re-point application database endpoints to secondary disaster-recovery cluster in us-east-2 via Route 53 DNS update:
      aws route53 change-resource-record-sets \
        --hosted-zone-id Z12345678 \
        --change-batch file://dns-failover-dr.json
      

    7. Escalation and Contacts

    RoleContact NameEscalation ChannelSLA
    Incident CommanderDavid O'ReillyPagerDuty: @incident-commander5 minutes
    Lead DBREMarcus VancePagerDuty: @dbre-lead5 minutes
    Core Banking LeadSarah ChenSlack: #incident-core-banking10 minutes

    8. Appendix: Audit & Evidence Log

    Record execution parameters in incident ticket:

    • Failover start timestamp (UTC).
    • Measured pre-failover lag value in seconds.
    • Promotion elapsed time in seconds.
    • Executing engineer IAM ARN.

    Verification

    No external automated validator was supplied; reviewer self-check against runbook standards executed:

    • Adheres to standard 8-section Operational Runbook template.
    • Includes concrete pre-checks with explicit numerical abort gates (< 5s go, >= 10s abort).
    • Contains copy-pasteable AWS CLI commands with zero unresolved {{...}} placeholders.
    • 90-second promotion timeout and rollback procedure documented.

    Next steps

    1. Marcus Vance conducts mock failover rehearsal in staging cluster bank-aurora-stage.
    2. SRE team verifies CloudWatch metric query JSON precheck-replica-lag.json is checked into repository runbook assets.
    3. Review IAM role permissions to ensure DB-Admin-Elevated-Role has rds:FailoverDBCluster authorization.

    operational-runbook-writer.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Draft verifiable steps for specific system recovery tasksEstablish safety boundaries and escalation triggers for on-call teamsStandardize operational procedures with diagnostic oracles and checkpointsTranslate tribal knowledge into bounded Markdown runbooks for engineers

    About this skill

    Operational Runbook Writer: Full Description

    What it does

    This skill renders one qualified operational condition or task into bounded, verifiable procedure for an authorized operator. It preserves diagnostic reasoning, state-changing effects, checkpoints, failure, recovery and freshness without executing commands.

    Use it when

    Use when a known operator needs one specific operational procedure for an identified system/state/version/environment.

    For example: “We need an operational runbook for on-call engineers to recover from Kafka partition consumer lag unbalance and socket disconnects in our Logistics Fleet Tracking pipeline.”

    What you get

    • Operational Runbook Document

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/runbook-authoring/.

    What it will not do

    Does not invent production procedures. Do not use for multi-scenario coordination playbooks, SOP/checklists, policy, incident/security/DR plans, tutorials, architecture docs or live execution.

    How it works

    1. Check operational scope is single-task and single-system.
    2. Define exact operational condition and target scope.
    3. Establish prerequisites, credentials, and safety bounds.
    4. Draft explicit, step-by-step diagnostic procedures.
    5. Formulate bounded state-changing remediation actions.
    6. Define step verification oracles and stop/escalation criteria.
    7. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 13 days ago

    • Passed all security checks, Safe to install

    Listed13 days ago

    What's inside

    Frequently Asked Questions