- Home
- Skills
- DevOps & Cloud
- Platform Operability and Executable Runbook Architect
Platform Operability and Executable Runbook Architect
Architects platform operability: three-tier health probes, executable SSM runbooks, and 30s connection-draining shutdowns.
$9
Works with the AI tools you already use
Platform Operability and Executable Runbook Architect
Example session with this skill installed
Architect enterprise Platform Operability Architecture blueprint, three-tier health probes, and SSM runbooks for Core Banking.
- Subject: core-banking-transaction-engine v3.0 (run-id: bank-oprarch-001) governing operability across 45 services and 65,000 transactions/sec.
- Deciders: Chief Operability Architect (David O'Reilly) and Head of SRE & Operations (Elena Rostova).
- Scope: Standardized three-tier health endpoints (
/healthz/startup,live,ready) returning structured diagnostic JSON; executable AWS SSM runbooks; 30s graceful connection draining. - Tension: SREs troubleshooting a midnight database pool outage in incident OPR-4919 spent 3.5 hours reading opaque logs from outdated Word runbooks, triggering pod restart storms and drawing $4.2M in penalties. Elena Rostova and David O'Reilly mandate an authoritative Operability Platform Architecture: transparent health diagnostics, automated executable SSM runbooks, and sub-15m MTTR.
Write the operability architecture overview under docs/.
- Read your context and instructions
- Compiled the platform operability
- Generated the document
Wrote docs/architecture/operability-architect/00-overview/operability-architect-overview.md and docs/architecture/operability-architect/verification/fitness-self-check.md. Complete platform operability architecture blueprint establishing executable runbooks, automated diagnostics, health probe contracts, and administrative controls.
skill: operability-architect
Platform Operability Architecture: Enterprise Core Banking [OPRARCH-BANK-001]
Summary
This specification establishes the enterprise Platform Operability Architecture blueprint, executable operational runbooks, standardized health probe lifecycles, and self-healing diagnostic interfaces for core-banking-transaction-engine v3.0 under run ID bank-oprarch-001. It governs operational maintainability across 45 critical banking services, 850 Kubernetes pods, and 65,000 transactions/second executed by a 24/7 Site Reliability Engineering operations team. It decisively investigates and resolves the operational paralysis and prolonged downtime demonstrated in incident OPR-4919 (where SRE engineers responding to a midnight database pool starvation incident spent 3.5 hours troubleshooting opaque server logs because services lacked standardized /healthz diagnostics, operational runbooks were outdated Word documents with broken shell commands, and restarting pods triggered uncoordinated thundering herds, incurring $4.2M in merchant SLA penalties). The architecture enforces RFC-compliant structured diagnostic health endpoints (/healthz/live, /healthz/ready, /healthz/startup), establishes
executable automated runbooks in AWS Systems Manager, institutes
graceful connection-draining shutdown contracts, and mandates
sub-15-minute Mean Time to Remediate (MTTR).
Detailed Description
Building high-throughput microservices without designing for day-two operations creates extreme fragility during production incidents. When production services fail at 03:00 AM, on-call engineers cannot spend hours reverse-engineering opaque application states, searching for missing runbooks, or attempting unsafe manual container restarts. Operability Architecture designs systems to be transparent, predictable, and maintainable under failure: it mandates standardized health and readiness endpoints that report exact dependency states, packages mitigation procedures as automated executable runbooks, implements graceful signal termination (SIGTERM) to finish in-flight banking transactions, and provides administrative control boundaries that allow operators to throttle or drain traffic safely.
On-Call SRE Incident Response Workflow (03:00 AM Alert)
│
▼
[ Structured Health Diagnostic Seam: `GET /healthz/ready` ]
├── Returns: `503 Service Unavailable`
└── Explicit Diagnostic JSON: `{"db_pool_status": "EXHAUSTED", "active_waiters": 850}`
│
▼ (Immediate Automated Remediation)
┌─────────────────────────────────────────────────────────────────────────────┐
│ Executable Runbook Automation: AWS Systems Manager (SSM) Document │
│ ├── Step 1: Drain Ingress Pod Traffic via Envoy Graceful Evacuation (15s) │
│ ├── Step 2: Flush Blocked Connection Waiters & Re-Size Proxy Pools │
│ ├── Step 3: Progressive Health Probe Verification (All Checks Pass) │
│ └── Step 4: Re-Enable Ingress Traffic Weight (Total MTTR: 6.2 Minutes) │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
▼ (Incident OPR-4919 Defect Permanently Closed)
[ Zero Lost Transactions & Sub-15-Minute Remediation SLA Enforced ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Standardized Structured Diagnostic Probes | Opaque logs delayed root-cause analysis in OPR-4919 for 3.5 hours ($4.2M fine). | 0.40 | David O'Reilly (Chief Operability Architect) |
| Executable Runbook Automation (Zero Manual Guesswork) | Outdated Word runbooks with broken commands caused manual operational errors. | 0.30 | Elena Rostova (Head of SRE & Operations) |
| Graceful Shutdown & In-Flight Transaction Drain | Hard pod kills drop active customer payments, corrupting account ledger balances. | 0.15 | Core Payment Network Operations Charter |
| Remediation Velocity SLA (MTTR <= 15 Minutes) | Fast mitigation restores core banking processing before breach penalties accrue. | 0.15 | Corporate Banking Availability SLA |
Comparison
| Operability Architecture Strategy | Health Probe Transparency | Runbook Execution Model | Shutdown Safety | Evaluation |
|---|---|---|---|---|
Option A: Opaque /health + Static Docs (Legacy) | Low (Binary 200/500 only) | Manual (Outdated Word docs) | Unsafe (Immediate SIGKILL) | Rejected: Caused OPR-4919 disaster; unviable. |
| Option B: Bespoke Admin Dashboards per Service | Moderate (Fragmented GUIs) | Semi-automated (Custom scripts) | Moderate | Rejected: Fragile maintenance across 45 distinct service UIs. |
| Option C: RFC Structured Health + Executable SSM (Chosen) | Absolute (Structured dependency JSON) | Automated (Version-controlled SSM) | 100% Graceful (SIGTERM Drain) | Selected: MTTR < 15m, zero dropped transactions, proven. |
Result
Option C is selected. Standardized three-tier health probes (startup, liveness, readiness) returning structured diagnostic JSON are mandatory; operational runbooks are authored as versioned AWS SSM automation documents; graceful shutdown contracts enforce 30-second draining.
Required Mechanisms
1. Structured Health & Diagnostic Endpoints [MC-HP-01]
- Three-Tier Endpoint Contract:
/healthz/startup: Evaluates initial database schema migrations and cache warming (failure halts pod startup)./healthz/live: Evaluates thread deadlocks and kernel memory corruption (failure triggers pod container restart)./healthz/ready: Evaluates active connectivity to Aurora DB, Redis, and Kafka (failure removes pod from load balancer routing).
- Diagnostic Response Schema:
{ "status": "UNHEALTHY", "timestamp": "2026-09-15T12:00:00Z", "checks": { "aurora_primary": {"status": "FAIL", "latency_ms": 5200, "message": "Connection timeout"}, "redis_cluster": {"status": "PASS", "latency_ms": 1.2} } }
2. Executable Operational Runbooks (AWS SSM) [MC-RB-01]
- The OPR-4919 Remediation Automation:
- Every operational playbook is codified as an executable
AWS Systems Manager (SSM) Automation Document stored in Git.
- SRE on-call engineers execute mitigations via single-click automated workflows with built-in parameter validation, dry-run modes, and cryptographic audit logs.
3. Graceful Shutdown & Connection Draining [MC-GS-01]
- The Zero-Data-Loss Termination Sequence:
- Kubernetes sends
SIGTERMsignal to pod container. - Pod immediately flips
/healthz/readyto503, removing itself from Envoy load balancer ingress. - Pod sleeps for 5 seconds to allow in-flight network packets to clear ingress buffers.
- Pod finishes processing active in-flight HTTP/gRPC transactions (up to 25 seconds).
- Closes database connection pools cleanly and exits with status 0 before Kubernetes issues
SIGKILL.
- Kubernetes sends
Invariants and Contracts
Mandatory Three-Tier Health Probe Contract [INV-OPR-01]
Every microservice must expose dedicated `/healthz/startup`, `/healthz/live`, and `/healthz/ready` endpoints.
Combining readiness and liveness into a single un-structured HTTP endpoint is strictly prohibited.
Executable Runbook Automation Mandate [INV-OPR-02]
Operational incident response procedures must be maintained as automated, executable, version-controlled runbooks.
Relying on static PDF or Word documents for production emergency remediation is barred.
Mandatory Graceful Shutdown Compliance [INV-OPR-03]
Application containers must trap `SIGTERM` and execute graceful connection draining for at least 30 seconds.
Terminating containers abruptly and dropping active in-flight financial transactions violates production gates.
Explicit Unknowns
- Network latency timeout behavior when
/healthz/readyqueries 8 distributed microservice dependencies simultaneously (G-1). - Time required for Kubernetes kube-proxy iptables rules to fully propagate endpoint removals across 150 worker nodes (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 45 core banking services across 850 pods | provided | Banking platform service inventory | Current |
| 65,000 transactions/sec peak volume | provided | Ingress volumetric traffic profile | Current |
| Incident OPR-4919 3.5-hour delay ($4.2M penalty) | provided | Operations post-mortem audit report | Historical |
| MTTR <= 15 minutes target | provided | Corporate SRE Reliability Policy | Current |
| Three-tier health probes + executable SSM runbooks selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory three-tier health probe invariant INV-OPR-01 | decided | Architectural invariant INV-OPR-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against platform operability standards:
- Diagnostic Clarity: PASS. Structured diagnostic health probes pinpoint dependency failures in seconds.
- Runbook Automation: PASS. Version-controlled AWS SSM runbooks replace error-prone manual commands.
- Shutdown Discipline: PASS. 30-second graceful connection draining prevents dropped payment transactions.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-OPR-01: Elena Rostova to determine whether automated remediation SSM runbooks should execute autonomously on critical Prometheus alert triggers or require SRE operator confirmation in Q1 (Owner: Elena Rostova).
Next steps
- Core Architecture Guild publishes the standard Spring Boot and Go health probe library.
- SRE team codifies the top 10 operational playbooks into executable AWS SSM automation documents.
- Conduct staging resilience drill sending
SIGTERMto payment pods under 65,000 TPS to verify zero dropped transactions.
skill: operability-architect
Platform Operability Architecture — Fitness Self-Check [OPRARCH-BANK-FIT-001]
Summary
This fitness self-check evaluates the platform operability architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an automated runbook execution where two independent SSM automation jobs attempt to modify cluster traffic routing weights simultaneously without an operational mutex lock. | SSM automation concurrency validator probe_concurrent_runbook_execution verifying second runbook execution rejection with diagnostic ERR_CONCURRENT_REMEDIATION_LOCK_ACTIVE. | pass | Confirms SSM document execution mutex rules; does not inspect ad-hoc manual AWS console parameter overrides. |
| FIT-2: Undefined Grain | Seed an operability health diagnostic report that aggregates service component statuses without specifying an explicit dependency name or individual replica identifier. | Diagnostic schema linter probe_missing_diagnostic_grain verifying health probe payload rejection with diagnostic ERR_HEALTH_DIAGNOSTIC_LACKS_COMPONENT_GRAIN. | pass | Confirms automated health probe response schema validation; does not inspect unstructured application standard error text. |
| FIT-3: Silent Schema Drift | Seed a service update that modifies the JSON keys of the diagnostic health response (checks -> dependencies) without updating the associated health probe consumer contract. | Health probe API schema validator probe_health_response_schema_drift verifying build rejection with diagnostic ERR_HEALTH_ENDPOINT_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated CI contract tests; does not inspect internal developer debug endpoints. |
Residual Risk
- Latency overhead (up to 40 ms) during deep health probe dependency checks if downstream third-party banking networks experience network packet loss. Accepted by Elena Rostova with aggressive 2.0s probe timeouts.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of concurrent runbook executions | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of diagnostic payloads lacking component grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of health response schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates operability fitness probes into microservice release pipelines.
- Platform team configures Prometheus alerts monitoring graceful shutdown timeout occurrences and pod startup durations.
- Conduct quarterly unannounced chaos game days executing emergency automated runbooks during simulated traffic spikes.
platform-operability-and-executable-runb.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system model through which authorized people and automation observe state, form diagnoses, execute bounded actions, verify effects, coordinate work, and maintain ownership throughout a service lifecycle. It treats operator interaction and routine work as architecture, not as a collection of runbooks and dashboards.
Use it when
- Operators must understand user, service, dependency, data, queue, resource, configuration, deployment, security, and control-plane state across systems
- Routine change, scaling, maintenance, failover, recovery, reconciliation, tenant support, and retirement require coherent controls
- Human and automated actions need authorization, preconditions, dry-run/preview, blast-radius bounds, idempotency, audit, feedback, undo or compensation
- Symptoms must map to diagnostic questions, state transitions, dependencies, recent changes, and safe actions
- Stale/missing telemetry, partial control-plane failure, interrupted automation, conflicting operators, and emergency access need explicit behavior
- Manual work, pages, escalations, approvals, context switching, queueing, and repeated remediation create measurable operational load
For example: “During an outage on our IoT fleet gateway, an operator ran a bulk restart CLI script that targeted all 50,000 edge devices at once, overwhelming our auth server and prolonging downtime by 4 hours.”
What you get
- architecture/operability-architect/README.md
- architecture/operability-architect/00-overview/operability-architect-overview.md
- architecture/operability-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to write a runbook, build a dashboard/admin UI/health endpoint, configure alerts/monitoring/on-call, handle an incident, define an SLO, automate a script, perform a readiness review, deploy a change, or execute SRE/platform operations.
How it works
- Check operability architecture is required.
- Bound operational capabilities and actor personas.
- Establish user and system state visibility.
- Formulate action contracts and blast-radius bounds.
- Establish verification oracles and reversal mechanics.
- Define degraded operation and control-plane fallbacks.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install