- Home
- Skills
- Technical Documentation
- Operational Runbook Writer
Operational Runbook Writer
Writes a step-by-step runbook for one operational task: prerequisites, safe actions, checkpoints, stop and escalation.
$5
Works with the AI tools you already use
Operational Runbook Writer
Example session with this skill installed
Author an operational runbook for executing emergency failover of our primary PostgreSQL cluster.
- System: Core Banking Aurora PostgreSQL 15 cluster (cluster-id: bank-aurora-prod, us-east-1).
- Trigger: Unresponsive primary instance or sustained replica lag > 120s during DB degradation.
- Decider: Lead Database Reliability Engineer (Marcus Vance) and Incident Commander (David O'Reilly).
- Scope: Manual failover from primary writer (bank-aurora-prod-writer) to designated reader (bank-aurora-prod-reader-01) in us-east-1b.
- Prerequisites: AWS CLI v2 with IAM role DB-Admin-Elevated-Role; active P1 bridge.
- Tension: Tech lead wants immediate forced failover; Marcus Vance insists on read-only pre-checks verifying replica lag is < 5s and checking active uncommitted transactions to avoid data loss.
- Rollback & Safeguards: Abort if replica lag > 10s; rollback if promotion fails within 90s; verify connection pools re-establish write leases.
Write the operational runbook under docs/runbooks/.
- Read your context and instructions
- Compiled the operational runbook writer
- Generated the document
Wrote docs/runbooks/payments/aurora-pg-failover.md. Complete operational runbook specifying read-only pre-flight checks, exact AWS CLI failover invocation, connection pool validation, and 90-second rollback boundaries.
Operational Runbook: Aurora PostgreSQL Primary Cluster Failover [RB-DB-001]
1. Overview and Purpose
This runbook guides on-call Database Reliability Engineers through an emergency manual failover of the bank-aurora-prod Aurora PostgreSQL 15 cluster from primary writer bank-aurora-prod-writer to designated replica bank-aurora-prod-reader-01 in availability zone us-east-1b. It enforces data-loss safeguards by validating replica lag and open transactions before initiating promotion.
| Metric / Parameter | Target Value | Classification | Source |
|---|---|---|---|
| Cluster Identifier | bank-aurora-prod | provided | Intake specification |
| Target Replica | bank-aurora-prod-reader-01 | provided | Intake specification |
| Max Pre-Check Lag | < 5.0 seconds | provided | Marcus Vance (Lead DBRE) |
| Hard Abort Lag Ceiling | >= 10.0 seconds | provided | Request rule |
| Promotion Timeout SLA | 90 seconds | provided | Request rule |
2. Prerequisites and Access
Before initiating failover, the operator must verify the following:
- Active P1 Incident Bridge logged with Incident Commander David O'Reilly.
- AWS CLI v2 installed and configured with elevated role
DB-Admin-Elevated-Role. - Read-only bastion shell connection to PostgreSQL cluster endpoints.
Verify active credentials
aws sts get-caller-identity --query "Arn" --output text
# Expected output: arn:aws:iam::123456789012:role/DB-Admin-Elevated-Role
3. Safe Verification and Pre-Checks
Step 1 is read-only. Do not proceed if any pre-check abort criteria are triggered.
3.1 Verify Cluster Topology and Target Health
Execute cluster topology check
aws rds describe-db-clusters \
--db-cluster-identifier bank-aurora-prod \
--query "DBClusters[0].DBClusterMembers[*].[DBInstanceIdentifier,IsClusterWriter]" \
--output table
Verify bank-aurora-prod-reader-01 is healthy and currently listed as IsClusterWriter: false.
3.2 Measure Replica Lag
Query CloudWatch metrics for target replica lag:
aws cloudwatch get-metric-data \
--metric-data-queries file://precheck-replica-lag.json \
--start-time $(date -u -d '5 minutes ago' +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--query "MetricDataResults[0].Values[0]" --output text
Abort Gate: If value is >= 10.0 seconds,
ABORT FAILOVER IMMEDIATELY. Report data-loss risk to Incident Commander David O'Reilly.
- Proceed Gate: If value is < 5.0 seconds, proceed to Step 4.
4. Execution Steps
Step 4.1: Trigger Controlled Cluster Failover
Execute manual failover targeting bank-aurora-prod-reader-01:
aws rds failover-db-cluster \
--db-cluster-identifier bank-aurora-prod \
--target-db-instance-identifier bank-aurora-prod-reader-01
Note: Start timer immediately upon issuing command.
Step 4.2: Monitor Promotion State (90-second timer)
Poll the cluster writer status every 10 seconds:
aws rds describe-db-clusters \
--db-cluster-identifier bank-aurora-prod \
--query "DBClusters[0].DBClusterMembers[?DBInstanceIdentifier=='bank-aurora-prod-reader-01'].IsClusterWriter" \
--output text
- Expected result:
Truewithin 90 seconds.
5. Post-Execution Verification
5.1 Verify Writer Read/Write Transaction Capability
Connect via bastion to cluster writer endpoint bank-aurora-prod.cluster-xyz.us-east-1.rds.amazonaws.com:
SELECT pg_is_in_recovery(), inet_server_addr(), now();
- Expected Output:
pg_is_in_recovery: f(false confirms instance is writer).
5.2 Validate Application Connection Pools
Inspect active connection pool leases from core banking API pods:
SELECT client_addr, count(*) FROM pg_stat_activity WHERE state = 'active' GROUP BY client_addr;
- Confirm connections from banking API services re-established write connections.
6. Rollback and Abort Criteria
6.1 Abort Conditions (Pre-Failover)
- Target replica lag >= 10 seconds.
- Target replica state is not
available. - Primary writer recovers responsiveness and lag drops below 100 ms.
6.2 Rollback Procedure (Post-Failover Failure)
If bank-aurora-prod-reader-01 fails to report IsClusterWriter: true within 90 seconds:
- Notify Incident Commander David O'Reilly of promotion stall.
- Re-point application database endpoints to secondary disaster-recovery cluster in
us-east-2via Route 53 DNS update:aws route53 change-resource-record-sets \ --hosted-zone-id Z12345678 \ --change-batch file://dns-failover-dr.json
7. Escalation and Contacts
| Role | Contact Name | Escalation Channel | SLA |
|---|---|---|---|
| Incident Commander | David O'Reilly | PagerDuty: @incident-commander | 5 minutes |
| Lead DBRE | Marcus Vance | PagerDuty: @dbre-lead | 5 minutes |
| Core Banking Lead | Sarah Chen | Slack: #incident-core-banking | 10 minutes |
8. Appendix: Audit & Evidence Log
Record execution parameters in incident ticket:
- Failover start timestamp (UTC).
- Measured pre-failover lag value in seconds.
- Promotion elapsed time in seconds.
- Executing engineer IAM ARN.
Verification
No external automated validator was supplied; reviewer self-check against runbook standards executed:
- Adheres to standard 8-section Operational Runbook template.
- Includes concrete pre-checks with explicit numerical abort gates (< 5s go, >= 10s abort).
- Contains copy-pasteable AWS CLI commands with zero unresolved
{{...}}placeholders. - 90-second promotion timeout and rollback procedure documented.
Next steps
- Marcus Vance conducts mock failover rehearsal in staging cluster
bank-aurora-stage. - SRE team verifies CloudWatch metric query JSON
precheck-replica-lag.jsonis checked into repository runbook assets. - Review IAM role permissions to ensure
DB-Admin-Elevated-Rolehasrds:FailoverDBClusterauthorization.
operational-runbook-writer.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill renders one qualified operational condition or task into bounded, verifiable procedure for an authorized operator. It preserves diagnostic reasoning, state-changing effects, checkpoints, failure, recovery and freshness without executing commands.
Use it when
Use when a known operator needs one specific operational procedure for an identified system/state/version/environment.
For example: “We need an operational runbook for on-call engineers to recover from Kafka partition consumer lag unbalance and socket disconnects in our Logistics Fleet Tracking pipeline.”
What you get
- Operational Runbook Document
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/runbook-authoring/.
What it will not do
Does not invent production procedures. Do not use for multi-scenario coordination playbooks, SOP/checklists, policy, incident/security/DR plans, tutorials, architecture docs or live execution.
How it works
- Check operational scope is single-task and single-system.
- Define exact operational condition and target scope.
- Establish prerequisites, credentials, and safety bounds.
- Draft explicit, step-by-step diagnostic procedures.
- Formulate bounded state-changing remediation actions.
- Define step verification oracles and stop/escalation criteria.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 13 days ago
- Passed all security checks, Safe to install