Health-Check Oracle and Consumer-Effect Design

    1

    Designs health check probes: startup, liveness, readiness separation, cascading failure prevention, and flap damping.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Health-Check Oracle and Consumer-Effect Design

    Example session with this skill installed

    Design health check oracles and consumer effect contracts for our Order Processing Gateway on Kubernetes.

    • Subject: order-gateway-service v2.4 (run-id: orders-health-001) operating on AWS EKS across 40 pods fronted by ALBs.
    • Deciders: Lead SRE Architect (Marcus Vance) and Core Platform Lead (Elena Rostova).
    • Workload: Peak 2,800 orders/sec; service depends on local Redis cache, PostgreSQL database, and payment gateways.
    • Tension: Tech lead proposed a single /health endpoint pinging PostgreSQL, Redis, and payment APIs on every call, wired to both liveness and readiness probes; Marcus Vance warns deep dependency checks on liveness probes trigger cascading restarts (incident INC-3119 where a 10-second DB brownout caused Kubernetes to kill all 40 pods simultaneously, turning a minor database delay into a 45-minute crash-loop outage). He mandates strict probe separation (shallow /healthz/live, dependency-aware /healthz/ready, slow-init /healthz/startup), flap damping (3 consecutive failures to unready, 2 to restore), and immediate removal from ALB rotation without restarting pods.

    Write the health check spec under docs/.

    • Read your context and instructions
    • Compiled the health-check oracle
    • Generated the UI component

    Wrote docs/architecture/tasks/orders-health-001/health-check-design/health-check-spec.md. Complete health check oracle specification establishing startup/liveness/readiness separation, anti-cascading restart protections, hysteresis flap damping, and load balancer traffic shedding.

    ---
    skill: health-check-design
    ---
    
    # Health-Check Spec: Order Processing Gateway [HEALTH-ORD-001]
    
    ## Summary
    
    This specification establishes the health check oracle, probe separation hierarchy, and consumer traffic reaction contract for `order-gateway-service v2.4` under run ID `orders-health-001`, governing 40 Kubernetes pods across 2,800 peak orders/second on AWS EKS. It decisively resolves the cascading cluster collapse demonstrated in incident INC-3119 (where a transient 10-second database brownout caused liveness probes to fail, killing all 40 pods simultaneously and triggering a 45-minute crash-loop outage). The contract enforces strict probe separation: shallow process-only liveness (`/healthz/live`), isolated dependency-aware readiness (`/healthz/ready`), dedicated warm-up startup probes (`/healthz/startup`), flap damping via asymmetric failure/recovery hysteresis, and immediate ALB traffic shedding without pod termination.
    
    ## Detailed Description
    
    Conflating liveness and readiness creates lethal distributed feedback loops. A liveness probe answers: *Is the process deadlocked or corrupted such that restarting it is the only remedy?* A readiness probe answers: *Can this specific instance accept incoming customer requests right now?* Wiring remote database or third-party network pings into liveness probes causes Kubernetes to restart healthy containers when external dependencies degrade, multiplying connection storms upon container reboot.
    
    

    Kubernetes Kubelet Probe Scheduler
    │
    ┌──────────────┼──────────────┐
    ▼ ▼ ▼
    [ Startup Probe ] [ Liveness Probe ] [ Readiness Probe ]
    /healthz/startup /healthz/live /healthz/ready
    │ (Init 45s) │ (Shallow) │ (Dependency Check)
    │ │ │
    ▼ ▼ ├─► Local Redis Cache: Ping (Timeout: 150ms)
    Initial JVM Boot Process Deadlock?├── PostgreSQL Pool: Acquires Lease (< 250ms)
    (Blocks Liveness) (Restart Pod) └── External APIs: Bypassed (Circuit Breaker Decides)
    │
    ┌─────────────────────────┴─────────────────────────┐
    ▼ ▼
    (Readiness PASS: 200 OK) (Readiness FAIL: 503)
    Retain in ALB Target Group [ Consumer Effect: Shed Traffic ]
    ├── Flap Damping: 3 Fails -> Unready
    ├── Drop from EndpointSlice (< 2s)
    └── Pod Remains Running (Zero Restarts)

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Cascading Restart Prevention | External dependency degradation must never trigger container restarts (INC-3119). | 0.40 | Marcus Vance (Lead SRE) |
    | Flap Damping & State Stability | Transient network blips must not rapidly cycle pods between ready and unready states. | 0.25 | Elena Rostova (Platform Lead) |
    | Rapid Traffic Shedding (< 2s) | Unready instances must be excised from ALB routing before customer requests drop. | 0.20 | Order Processing SLA |
    | Cold-Start Buffer & Graceful Warm-Up | Complex JVM class-loading and JIT compilation must not trip premature liveness aborts. | 0.15 | JVM Runtime Engineering |
    
    
    ### Comparison
    
    | Probe Design Candidate | Liveness Scope | Readiness Scope | Reaction to DB Brownout | Restart Cascade Risk |
    |---|---|---|---|---|
    | Option A: Monolithic `/health` (Legacy) | Deep DB + Redis + External | Deep DB + Redis + External | Kubelet terminates all 40 pods | Critical: Recreates INC-3119 crash-loop outage. |
    | Option B: Shallow-Only Probes | Process TCP port 8080 | Process TCP port 8080 | Pods stay in ALB; 500s return to user | High: Blackholes customer traffic to broken pods. |
    | Option C: Tripartite Probe Architecture (Chosen) | Shallow in-memory deadlock | Local Redis + DB pool lease | Pods marked Unready; ALB sheds traffic | Minimal: Pods shed traffic safely; zero restarts. |
    
    
    ### Result
    
    Option C is selected. Three distinct endpoints isolate startup initialization, process survival, and traffic routing eligibility.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Probe Separation Hierarchy [MC-PS-01]
    
    ##### Startup Probe (`GET /healthz/startup`)
    - **Purpose**: Protects slow JVM application bootstrap and initial cache warming.
    - **Evaluation**: Asserts database connection pool initialized and Spring context loaded.
    - **Timing**: `initialDelaySeconds = 10`, `periodSeconds = 5`, `failureThreshold = 18` (Allows up to 100 seconds warm-up).
    - **Behavior**: While startup is active, liveness and readiness probes are completely disabled.
    
    ##### Liveness Probe (`GET /healthz/live`)
    - **Purpose**: Detects internal thread deadlocks, JVM fatal states, or corrupted internal loops.
    - **Evaluation**: Shallow check only: returns HTTP 200 if the HTTP server event loop is responsive and internal health watchdog is kicking.
    - **Strict Prohibition**: Liveness must NEVER execute outbound network calls, database queries, or remote dependency pings.
    - **Timing**: `periodSeconds = 10`, `timeoutSeconds = 2`, `failureThreshold = 3`.
    
    ##### Readiness Probe (`GET /healthz/ready`)
    - **Purpose**: Governs inclusion in the Kubernetes `EndpointSlice` and AWS ALB target group.
    - **Evaluation**:
      1. Internal thread pool capacity: rejected if active work queues >= 90%.
      2. Database connectivity: executes shallow non-locking query (`SELECT 1`) with 250 ms timeout.
      3. Redis connectivity: issues `PING` with 150 ms timeout.
    - **Timing**: `periodSeconds = 5`, `timeoutSeconds = 1`, `failureThreshold = 3`, `successThreshold = 2`.
    
    #### 2. Cascading Failure Prevention & Dependency Isolation [MC-CF-01]
    - **Third-Party API Exclusion**: Outbound third-party payment gateways (e.g. Stripe, Adyen) must be excluded from the readiness probe. Degradation of third-party APIs is handled by internal circuit breakers (`circuit-breaker-design`), not by marking the entire order gateway unready.
    - **Database Partition Protection**: If PostgreSQL is unresponsive, readiness fails (HTTP 503), removing pods from ALB rotation. Containers continue running and immediately resume traffic acceptance once the database recovers, eliminating pod restart storms.
    
    #### 3. Flap Damping & Hysteresis Rules [MC-FD-01]
    - **Hysteresis Thresholds**:
      - Transition from `Ready` -> `Unready`: Requires **3 consecutive failed probes** (15 seconds total).
      - Transition from `Unready` -> `Ready`: Requires **2 consecutive successful probes** (10 seconds total).
    - Eliminates flapping where single dropped packets alternate pods in and out of ALB routing pools every 5 seconds.
    
    #### 4. Consumer Traffic Removal Contract [MC-TR-01]
    - Upon readiness failure (3 consecutive HTTP 503s):
      1. Kubelet removes pod IP from the Kubernetes `Endpoints` / `EndpointSlice` within 1,000 ms.
      2. AWS Target Group Controller detaches the target IP from active ALB traffic.
      3. Existing in-flight requests complete within the 10-second graceful connection draining window.
    
    ---
    
    ### Invariants and Contracts
    
        Liveness Remote Dependency Prohibition [INV-HLT-01]
          Liveness probes must never execute network I/O against remote databases, caches, or third-party
          services. Ingress liveness probes violating this rule fail deployment admission linters.
    
        Asymmetric Hysteresis Requirement [INV-HLT-02]
          Readiness probes must enforce asymmetric hysteresis (`failureThreshold >= 3`, `successThreshold >= 2`).
          Single-check unready transitions are strictly prohibited to prevent route flapping.
    
        Third-Party API Isolation Invariant [INV-HLT-03]
          External vendor APIs must not be included in readiness check evaluations. Vendor outages must
          be contained via internal circuit breakers and fallback degradation mechanisms.
    
    ## Explicit Unknowns
    
    - AWS Load Balancer Controller target registration latency under sudden 40-pod mass unready events (G-1).
    - Memory allocation footprint of frequent Prometheus synthetic probe scraping on `/healthz/ready` (G-2).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | 40 pods across AWS EKS | provided | Infrastructure intake | Current |
    | Peak 2,800 orders/sec | provided | Traffic profile intake | Current |
    | Incident INC-3119 45-minute crash-loop outage | provided | Post-mortem evidence | Historical |
    | 10-second DB brownout killed 40 pods | provided | Historical incident record | Historical |
    | Tripartite probe separation selection | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
    | Asymmetric hysteresis (3 fail / 2 pass) | decided | Architectural invariant INV-HLT-02 | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against health check architecture standards:
    - **Probe Separation**: PASS. Startup (100s window), Liveness (shallow), and Readiness (local dependencies) decoupled.
    - **Anti-Cascade Defense**: PASS. Liveness probe contains zero outbound network calls.
    - **Flap Damping**: PASS. 3-failure / 2-success hysteresis prevents route thrashing.
    - **Markdown Hygiene**: PASS. Conforms strictly to native Markdown rules in `rule_markdown.md`.
    
    ## Open Decisions
    
    - `DEC-HLT-01`: Elena Rostova to determine whether `/healthz/ready` should return a JSON breakdown of component latencies when running in non-production environments (Owner: Elena Rostova).
    
    ## Next steps
    
    1. Marcus Vance configures Kubernetes PodSpec manifests with the tripartite probe definitions.
    2. Platform team implements `/healthz/startup`, `/healthz/live`, and `/healthz/ready` HTTP handlers in Go.
    3. Conduct staging chaos drill injecting 15s PostgreSQL pause to verify zero pod restarts and clean ALB traffic recovery.
    

    health-check-oracle-and-consumer-effect-.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define startup, liveness, and readiness probe contracts.Map dependency criticality to prevent cascading restarts.Specify probe timing and failure threshold parameters.Establish consumer actions for platform-level recovery.Create structured health check specs in Markdown.

    About this skill

    What it does

    This skill maps accepted lifecycle, traffic and restart semantics into bounded health oracles and consumer actions. It defines what each check proves, who observes it and what failure may trigger without implementing endpoints or probes.

    Use it when

    Use when one process/workload/service needs an exact machine-consumed startup, traffic-eligibility, restart-safety or diagnostic check.

    For example: “When our Redis cache restarts, Kubernetes restarts all 20 scheduling pods simultaneously because the /health endpoint checks Redis. The entire appointment system goes down for 10 minutes instead of continuing with direct DB lookups.”

    What you get

    • Health Check Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/health-check-design/.

    What it will not do

    Do not use for general monitoring/alerting, SLOs, load-balancing policy, Kubernetes manifests, implementation or troubleshooting.

    How it works

    1. Check health check contract is required.
    2. Distinguish probe responsibilities.
    3. Map dependency criticality to readiness probes.
    4. Specify probe execution contracts and timing parameters.
    5. Establish consumer actions and restart safety invariants.
    6. Write the health check spec under <output_root>/architecture/tasks/{run-id}/health-check-design/health-check-spec.md.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions