Platform Operability and Executable Runbook Architect

    1

    Architects platform operability: three-tier health probes, executable SSM runbooks, and 30s connection-draining shutdowns.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Platform Operability and Executable Runbook Architect

    Example session with this skill installed

    Architect enterprise Platform Operability Architecture blueprint, three-tier health probes, and SSM runbooks for Core Banking.

    • Subject: core-banking-transaction-engine v3.0 (run-id: bank-oprarch-001) governing operability across 45 services and 65,000 transactions/sec.
    • Deciders: Chief Operability Architect (David O'Reilly) and Head of SRE & Operations (Elena Rostova).
    • Scope: Standardized three-tier health endpoints (/healthz/startup, live, ready) returning structured diagnostic JSON; executable AWS SSM runbooks; 30s graceful connection draining.
    • Tension: SREs troubleshooting a midnight database pool outage in incident OPR-4919 spent 3.5 hours reading opaque logs from outdated Word runbooks, triggering pod restart storms and drawing $4.2M in penalties. Elena Rostova and David O'Reilly mandate an authoritative Operability Platform Architecture: transparent health diagnostics, automated executable SSM runbooks, and sub-15m MTTR.

    Write the operability architecture overview under docs/.

    • Read your context and instructions
    • Compiled the platform operability
    • Generated the document

    Wrote docs/architecture/operability-architect/00-overview/operability-architect-overview.md and docs/architecture/operability-architect/verification/fitness-self-check.md. Complete platform operability architecture blueprint establishing executable runbooks, automated diagnostics, health probe contracts, and administrative controls.


    skill: operability-architect

    Platform Operability Architecture: Enterprise Core Banking [OPRARCH-BANK-001]

    Summary

    This specification establishes the enterprise Platform Operability Architecture blueprint, executable operational runbooks, standardized health probe lifecycles, and self-healing diagnostic interfaces for core-banking-transaction-engine v3.0 under run ID bank-oprarch-001. It governs operational maintainability across 45 critical banking services, 850 Kubernetes pods, and 65,000 transactions/second executed by a 24/7 Site Reliability Engineering operations team. It decisively investigates and resolves the operational paralysis and prolonged downtime demonstrated in incident OPR-4919 (where SRE engineers responding to a midnight database pool starvation incident spent 3.5 hours troubleshooting opaque server logs because services lacked standardized /healthz diagnostics, operational runbooks were outdated Word documents with broken shell commands, and restarting pods triggered uncoordinated thundering herds, incurring $4.2M in merchant SLA penalties). The architecture enforces RFC-compliant structured diagnostic health endpoints (/healthz/live, /healthz/ready, /healthz/startup), establishes

    executable automated runbooks in AWS Systems Manager, institutes

    graceful connection-draining shutdown contracts, and mandates

    sub-15-minute Mean Time to Remediate (MTTR).

    Detailed Description

    Building high-throughput microservices without designing for day-two operations creates extreme fragility during production incidents. When production services fail at 03:00 AM, on-call engineers cannot spend hours reverse-engineering opaque application states, searching for missing runbooks, or attempting unsafe manual container restarts. Operability Architecture designs systems to be transparent, predictable, and maintainable under failure: it mandates standardized health and readiness endpoints that report exact dependency states, packages mitigation procedures as automated executable runbooks, implements graceful signal termination (SIGTERM) to finish in-flight banking transactions, and provides administrative control boundaries that allow operators to throttle or drain traffic safely.

    On-Call SRE Incident Response Workflow (03:00 AM Alert)
                             │
                             ▼
    [ Structured Health Diagnostic Seam: `GET /healthz/ready` ]
      ├── Returns: `503 Service Unavailable`
      └── Explicit Diagnostic JSON: `{"db_pool_status": "EXHAUSTED", "active_waiters": 850}`
                             │
                             ▼ (Immediate Automated Remediation)
    ┌─────────────────────────────────────────────────────────────────────────────┐
    │ Executable Runbook Automation: AWS Systems Manager (SSM) Document           │
    │   ├── Step 1: Drain Ingress Pod Traffic via Envoy Graceful Evacuation (15s) │
    │   ├── Step 2: Flush Blocked Connection Waiters & Re-Size Proxy Pools        │
    │   ├── Step 3: Progressive Health Probe Verification (All Checks Pass)       │
    │   └── Step 4: Re-Enable Ingress Traffic Weight (Total MTTR: 6.2 Minutes)    │
    └──────────────────────────────────────┬──────────────────────────────────────┘
                                           │
                             ▼ (Incident OPR-4919 Defect Permanently Closed)
    [ Zero Lost Transactions & Sub-15-Minute Remediation SLA Enforced ]
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Standardized Structured Diagnostic ProbesOpaque logs delayed root-cause analysis in OPR-4919 for 3.5 hours ($4.2M fine).0.40David O'Reilly (Chief Operability Architect)
    Executable Runbook Automation (Zero Manual Guesswork)Outdated Word runbooks with broken commands caused manual operational errors.0.30Elena Rostova (Head of SRE & Operations)
    Graceful Shutdown & In-Flight Transaction DrainHard pod kills drop active customer payments, corrupting account ledger balances.0.15Core Payment Network Operations Charter
    Remediation Velocity SLA (MTTR <= 15 Minutes)Fast mitigation restores core banking processing before breach penalties accrue.0.15Corporate Banking Availability SLA

    Comparison

    Operability Architecture StrategyHealth Probe TransparencyRunbook Execution ModelShutdown SafetyEvaluation
    Option A: Opaque /health + Static Docs (Legacy)Low (Binary 200/500 only)Manual (Outdated Word docs)Unsafe (Immediate SIGKILL)Rejected: Caused OPR-4919 disaster; unviable.
    Option B: Bespoke Admin Dashboards per ServiceModerate (Fragmented GUIs)Semi-automated (Custom scripts)ModerateRejected: Fragile maintenance across 45 distinct service UIs.
    Option C: RFC Structured Health + Executable SSM (Chosen)Absolute (Structured dependency JSON)Automated (Version-controlled SSM)100% Graceful (SIGTERM Drain)Selected: MTTR < 15m, zero dropped transactions, proven.

    Result

    Option C is selected. Standardized three-tier health probes (startup, liveness, readiness) returning structured diagnostic JSON are mandatory; operational runbooks are authored as versioned AWS SSM automation documents; graceful shutdown contracts enforce 30-second draining.


    Required Mechanisms

    1. Structured Health & Diagnostic Endpoints [MC-HP-01]
    • Three-Tier Endpoint Contract:
      • /healthz/startup: Evaluates initial database schema migrations and cache warming (failure halts pod startup).
      • /healthz/live: Evaluates thread deadlocks and kernel memory corruption (failure triggers pod container restart).
      • /healthz/ready: Evaluates active connectivity to Aurora DB, Redis, and Kafka (failure removes pod from load balancer routing).
    • Diagnostic Response Schema:
      {
        "status": "UNHEALTHY",
        "timestamp": "2026-09-15T12:00:00Z",
        "checks": {
          "aurora_primary": {"status": "FAIL", "latency_ms": 5200, "message": "Connection timeout"},
          "redis_cluster": {"status": "PASS", "latency_ms": 1.2}
        }
      }
      
    2. Executable Operational Runbooks (AWS SSM) [MC-RB-01]
    • The OPR-4919 Remediation Automation:
      • Every operational playbook is codified as an executable

    AWS Systems Manager (SSM) Automation Document stored in Git.

    • SRE on-call engineers execute mitigations via single-click automated workflows with built-in parameter validation, dry-run modes, and cryptographic audit logs.
    3. Graceful Shutdown & Connection Draining [MC-GS-01]
    • The Zero-Data-Loss Termination Sequence:
      1. Kubernetes sends SIGTERM signal to pod container.
      2. Pod immediately flips /healthz/ready to 503, removing itself from Envoy load balancer ingress.
      3. Pod sleeps for 5 seconds to allow in-flight network packets to clear ingress buffers.
      4. Pod finishes processing active in-flight HTTP/gRPC transactions (up to 25 seconds).
      5. Closes database connection pools cleanly and exits with status 0 before Kubernetes issues SIGKILL.

    Invariants and Contracts

    Mandatory Three-Tier Health Probe Contract [INV-OPR-01]
      Every microservice must expose dedicated `/healthz/startup`, `/healthz/live`, and `/healthz/ready` endpoints.
      Combining readiness and liveness into a single un-structured HTTP endpoint is strictly prohibited.
    
    Executable Runbook Automation Mandate [INV-OPR-02]
      Operational incident response procedures must be maintained as automated, executable, version-controlled runbooks.
      Relying on static PDF or Word documents for production emergency remediation is barred.
    
    Mandatory Graceful Shutdown Compliance [INV-OPR-03]
      Application containers must trap `SIGTERM` and execute graceful connection draining for at least 30 seconds.
      Terminating containers abruptly and dropping active in-flight financial transactions violates production gates.
    

    Explicit Unknowns

    • Network latency timeout behavior when /healthz/ready queries 8 distributed microservice dependencies simultaneously (G-1).
    • Time required for Kubernetes kube-proxy iptables rules to fully propagate endpoint removals across 150 worker nodes (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    45 core banking services across 850 podsprovidedBanking platform service inventoryCurrent
    65,000 transactions/sec peak volumeprovidedIngress volumetric traffic profileCurrent
    Incident OPR-4919 3.5-hour delay ($4.2M penalty)providedOperations post-mortem audit reportHistorical
    MTTR <= 15 minutes targetprovidedCorporate SRE Reliability PolicyCurrent
    Three-tier health probes + executable SSM runbooks selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory three-tier health probe invariant INV-OPR-01decidedArchitectural invariant INV-OPR-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against platform operability standards:

    • Diagnostic Clarity: PASS. Structured diagnostic health probes pinpoint dependency failures in seconds.
    • Runbook Automation: PASS. Version-controlled AWS SSM runbooks replace error-prone manual commands.
    • Shutdown Discipline: PASS. 30-second graceful connection draining prevents dropped payment transactions.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-OPR-01: Elena Rostova to determine whether automated remediation SSM runbooks should execute autonomously on critical Prometheus alert triggers or require SRE operator confirmation in Q1 (Owner: Elena Rostova).

    Next steps

    1. Core Architecture Guild publishes the standard Spring Boot and Go health probe library.
    2. SRE team codifies the top 10 operational playbooks into executable AWS SSM automation documents.
    3. Conduct staging resilience drill sending SIGTERM to payment pods under 65,000 TPS to verify zero dropped transactions.

    skill: operability-architect

    Platform Operability Architecture — Fitness Self-Check [OPRARCH-BANK-FIT-001]

    Summary

    This fitness self-check evaluates the platform operability architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an automated runbook execution where two independent SSM automation jobs attempt to modify cluster traffic routing weights simultaneously without an operational mutex lock.SSM automation concurrency validator probe_concurrent_runbook_execution verifying second runbook execution rejection with diagnostic ERR_CONCURRENT_REMEDIATION_LOCK_ACTIVE.passConfirms SSM document execution mutex rules; does not inspect ad-hoc manual AWS console parameter overrides.
    FIT-2: Undefined GrainSeed an operability health diagnostic report that aggregates service component statuses without specifying an explicit dependency name or individual replica identifier.Diagnostic schema linter probe_missing_diagnostic_grain verifying health probe payload rejection with diagnostic ERR_HEALTH_DIAGNOSTIC_LACKS_COMPONENT_GRAIN.passConfirms automated health probe response schema validation; does not inspect unstructured application standard error text.
    FIT-3: Silent Schema DriftSeed a service update that modifies the JSON keys of the diagnostic health response (checks -> dependencies) without updating the associated health probe consumer contract.Health probe API schema validator probe_health_response_schema_drift verifying build rejection with diagnostic ERR_HEALTH_ENDPOINT_SCHEMA_DRIFT_DETECTED.passConfirms automated CI contract tests; does not inspect internal developer debug endpoints.

    Residual Risk

    • Latency overhead (up to 40 ms) during deep health probe dependency checks if downstream third-party banking networks experience network packet loss. Accepted by Elena Rostova with aggressive 2.0s probe timeouts.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of concurrent runbook executionsderivedFIT-1 probe result2026-09-15
    Rejection of diagnostic payloads lacking component grainderivedFIT-2 probe result2026-09-15
    Rejection of health response schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates operability fitness probes into microservice release pipelines.
    2. Platform team configures Prometheus alerts monitoring graceful shutdown timeout occurrences and pod startup durations.
    3. Conduct quarterly unannounced chaos game days executing emergency automated runbooks during simulated traffic spikes.

    platform-operability-and-executable-runb.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define blast-radius bounds and guardrails for administrative CLI toolsArchitect three-tier health probes for complex service dependenciesDesign verification oracles for automated failover systemsEstablish degraded operation fallbacks for control-plane failures

    About this skill

    What it does

    This skill owns the cross-system model through which authorized people and automation observe state, form diagnoses, execute bounded actions, verify effects, coordinate work, and maintain ownership throughout a service lifecycle. It treats operator interaction and routine work as architecture, not as a collection of runbooks and dashboards.

    Use it when

    • Operators must understand user, service, dependency, data, queue, resource, configuration, deployment, security, and control-plane state across systems
    • Routine change, scaling, maintenance, failover, recovery, reconciliation, tenant support, and retirement require coherent controls
    • Human and automated actions need authorization, preconditions, dry-run/preview, blast-radius bounds, idempotency, audit, feedback, undo or compensation
    • Symptoms must map to diagnostic questions, state transitions, dependencies, recent changes, and safe actions
    • Stale/missing telemetry, partial control-plane failure, interrupted automation, conflicting operators, and emergency access need explicit behavior
    • Manual work, pages, escalations, approvals, context switching, queueing, and repeated remediation create measurable operational load

    For example: “During an outage on our IoT fleet gateway, an operator ran a bulk restart CLI script that targeted all 50,000 edge devices at once, overwhelming our auth server and prolonging downtime by 4 hours.”

    What you get

    • architecture/operability-architect/README.md
    • architecture/operability-architect/00-overview/operability-architect-overview.md
    • architecture/operability-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to write a runbook, build a dashboard/admin UI/health endpoint, configure alerts/monitoring/on-call, handle an incident, define an SLO, automate a script, perform a readiness review, deploy a change, or execute SRE/platform operations.

    How it works

    1. Check operability architecture is required.
    2. Bound operational capabilities and actor personas.
    3. Establish user and system state visibility.
    4. Formulate action contracts and blast-radius bounds.
    5. Establish verification oracles and reversal mechanics.
    6. Define degraded operation and control-plane fallbacks.
    7. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions