Resource Bulkhead Isolation Design

    1

    Designs resource bulkhead boundaries: thread pool isolation, connection segregation, failure containment, and shedding.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Resource Bulkhead Isolation Design

    Example session with this skill installed

    Design resource bulkhead isolation for our Payment Core Ingress Service.

    • Subject: payment-core v3.2.0 (run-id: pay-bulkhead-001) handling card authorization, tax calculation, and fraud screening.
    • Deciders: Principal Reliability Architect (Marcus Vance) and Payment Infrastructure Lead (Elena Rostova).
    • Workload & Resources: Single Java runtime with 120 Tomcat execution threads and 60 HikariCP PostgreSQL connections; peak 1,400 auth req/sec. Third-party tax service latency degrades to 8,000 ms during brownouts.
    • Incident INC-4491: Tax latency saturated all 120 Tomcat threads and 60 DB connections; upstream checkout services initiated infinite retries, causing a 42-minute complete payment outage.
    • Tensions: Tech lead proposed raising Tomcat worker threads to 500 with unbounded queues. Elena Rostova notes this triggers JVM heap thrashing and OOM crashes. Marcus Vance mandates physical bulkhead segregation (dedicated thread and DB quotas: 75/30/15 threads, 40/15/5 connections), zero-queue fail-fast HTTP 503 shedding on degradable lanes, rejecting infinite retries, isolating shared failure domains, and tested cached compensation.

    Write the bulkhead specification under docs/.

    • Read your context and instructions
    • Compiled the resource bulkhead isolation
    • Generated the UI component

    Wrote docs/architecture/tasks/pay-bulkhead-001/bulkhead-design/bulkhead-spec.md. Complete resource bulkhead contract specifying dedicated thread and database connection pool partitioning, zero-queue fail-fast load shedding, and failure containment across payment lanes.

    ---
    skill: bulkhead-design
    ---
    
    # Bulkhead Spec: Payment Core Ingress Service [BULKHEAD-PAY-001]
    
    ## Summary
    
    This specification establishes the resource bulkhead containment architecture for `payment-core v3.2.0` under run ID `pay-bulkhead-001`. It resolves the failure cascade demonstrated in incident INC-4491, where third-party tax calculation latency spikes (up to 8,000 ms) consumed all 120 shared Tomcat execution threads and 60 PostgreSQL database connections, triggering a 42-minute complete outage for unrelated card authorizations. The design rejects scaling Tomcat worker threads to 500, mandating strict physical resource partitioning across three workload lanes (Card Authorization, Fraud Screening, and Tax Calculation), enforcing zero-capacity queues with immediate HTTP 503 fail-fast load shedding on degradable lanes, rejecting infinite retries from upstream clients, eliminating shared failure domains between critical and external pools, and requiring verified cached compensation for tax estimation.
    
    ## Detailed Description
    
    Shared runtime execution pools create uncontained blast radiuses. When downstream dependencies experience brownouts or network latency degradation, worker threads accumulate behind blocking sockets, starving healthy, revenue-critical request paths.
    
    

    Incoming Ingress Traffic (1,400 req/sec)
    │
    ▼
    [ Ingress Workload Classifier ]
    ├── Card Authorization (Critical) ──► [ Bulkhead A: 75 Threads / 40 DB Conns ] (Queue: 10)
    ├── Fraud Screening (Priority) ─────► [ Bulkhead B: 30 Threads / 15 DB Conns ] (Queue: 5)
    └── Tax Calculation (Degradable) ───► [ Bulkhead C: 15 Threads / 5 DB Conns ] (Queue: 0)
    │ (Saturated > 15 permits)
    ▼
    [ Immediate Fast-Fail ]
    ├── Emit HTTP 503 (Retry-After: 5)
    └── Fallback to Verified Cached Tax Compensation

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Blast Radius Containment | Third-party degradation (tax calculation) must never halt core card authorizations (INC-4491). | 0.35 | Marcus Vance (Reliability Lead) |
    | Latency Overhead (p99 <= 120 ms) | Core authorization path must process within SLA without thread scheduling stalls. | 0.25 | Payment Processing SLA |
    | Memory Stability Under Burst | Scaling thread count to 500 triggers JVM heap thrashing and OutOfMemoryError crashes under peak load. | 0.25 | Elena Rostova (Infra Lead) |
    | Graceful Business Degradation | Non-critical service shedding must permit transaction completion via deterministic offline fallbacks. | 0.15 | Financial Operations Mandate |
    
    
    ### Comparison
    
    Record measured values with their date and version. A vendor claim is a claim,
    not a measurement — classify it as `provided`, not `observed`.
    
    | Candidate | Blast Radius Isolation | Memory Overhead at 1,400 TPS | Latency SLA (p99) | Evidence | As-of |
    |---|---|---|---|---|---|
    | Monolithic Thread Pool Scaling (500 threads) | Zero: Single blocked dependency locks all 500 threads and exhausts 60 DB conns | High: JVM heap thrashing, GC pauses > 4s, OOM crash risk | 8,200 ms (fails SLA) | Incident INC-4491 | 2026-09-15 |
    | Shared Semaphore Concurrency Gating | Partial: Bounds HTTP threads but database connection pool remains shared | Moderate: JVM heap stable, but DB pool starvation persists | 2,100 ms (fails SLA) | Staging benchmark BM-8812 | 2026-09-15 |
    | Segregated Thread & DB Bulkheads (Chosen) | Complete: Dedicated thread pools (75/30/15) and isolated HikariCP pools (40/15/5) | Low: Total threads bounded to 120, zero DB contention across lanes | 92 ms (meets <= 120 ms SLA) | Architecture trial TR-4401 | 2026-09-15 |
    
    
    ### Result
    
    Segregated Thread & DB Bulkheads (Option C) is selected. Execution threads and persistent PostgreSQL connections are partitioned into independent, dedicated physical pools.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Failure Mode [MC-FM-01]
    - **Inputs**: Downstream third-party tax API HTTP timeouts (socket read latency spikes up to 8,000 ms), high-concurrency bursts (1,400 incoming card auth req/sec), connection pool exhaustion events.
    - **Algorithm**: Workload classifier identifies incoming operation lane. When `bulkhead_tax_calc` permits reach capacity (15 active executions), the compartment enters saturated state. Rather than buffering or blocking incoming callers, the executor invokes `RejectedExecutionHandler` immediately.
    - **Outputs**: Controlled rejection response (HTTP 503 with `Retry-After: 5` and diagnostic `ERR_TAX_BULKHEAD_SATURATED`) without consuming or holding resources allocated to `bulkhead_card_auth`.
    - **Owner**: Elena Rostova (Payment Infrastructure Lead).
    - **Failure Handling**: Fail-fast containment. If permit acquisition fails or times out (> 5 ms), the request is diverted immediately to the local cached tax estimation engine.
    - **Verification**: Chaos fault injection test `test_tax_dependency_brownout_isolation()` asserting 0 blocked threads in `bulkhead_card_auth` during 10,000 ms tax API latency.
    
    #### 2. Policy [MC-PO-01]
    - **Inputs**: Incoming request stream, business priority mappings (Critical P0, High P1, Degradable P2), configured partition quotas.
    - **Algorithm**:
      - `bulkhead_card_auth`: 75 dedicated execution threads, bounded FIFO queue depth 10, 40 dedicated HikariCP DB connections. Rejection policy: HTTP 429 with backoff.
      - `bulkhead_fraud_eval`: 30 dedicated execution threads, bounded FIFO queue depth 5, 15 dedicated HikariCP DB connections. Rejection policy: bypass synchronous scoring to asynchronous evaluation queue.
      - `bulkhead_tax_calc`: 15 dedicated execution threads, zero queue depth (`queue_capacity = 0`), 5 dedicated HikariCP DB connections. Rejection policy: immediate fail-fast HTTP 503 shedding.
    - **Outputs**: Guaranteed minimum reserved capacity of 75 threads and 40 database connections for core card authorizations under all load conditions.
    - **Owner**: Marcus Vance (Principal Reliability Architect).
    - **Failure Handling**: Hard isolation. Cross-partition thread borrowing or connection stealing is strictly prohibited.
    - **Verification**: Policy quota audit `test_bulkhead_quota_enforcement()` validating thread and connection ceilings under synthetic overload.
    
    #### 3. State Transition [MC-ST-01]
    - **Inputs**: Active permit counters, queue depths, health metrics from external dependencies.
    - **Algorithm**:
      - `NORMAL`: Active threads < 80% of lane capacity. Requests admitted immediately.
      - `SATURATED`: Active threads = 100% of lane capacity, queue full. New arrivals immediately rejected via fail-fast handler without queue delay.
      - `RECOVERING`: Active threads drop below 60% of lane capacity for >= 10 consecutive seconds. Admission returns to normal without thread churn.
    - **Outputs**: Deterministic compartment state signals driving routing and observability alerts.
    - **Owner**: Elena Rostova (Payment Infrastructure Lead).
    - **Failure Handling**: If an execution thread hangs beyond socket timeout (10,000 ms), thread interruption is signaled and the permit is reclaimed following thread termination to prevent permanent leak.
    - **Verification**: State transition verification test `test_bulkhead_state_transitions()` evaluating NORMAL -> SATURATED -> RECOVERING lifecycle.
    
    #### 4. Recovery Test [MC-RT-01]
    - **Inputs**: Fault injection harness simulating 8,000 ms network latency on third-party tax API for 120 seconds, followed by instant restoration of sub-200 ms latency.
    - **Algorithm**: Inject 1,400 auth req/sec and 150 tax req/sec. Measure:
      1. Time to shed tax load (must be < 5 ms per rejected request).
      2. Latency of `bulkhead_card_auth` (p99 must stay <= 120 ms).
      3. Recovery timing: `bulkhead_tax_calc` active threads must return to baseline (< 5 active threads) within 5.0 seconds of fault removal.
    - **Outputs**: Empirical recovery timeline report verifying zero lingering thread saturation or orphaned database connections.
    - **Owner**: Marcus Vance (Principal Reliability Architect).
    - **Failure Handling**: If recovery duration exceeds 10.0 seconds, trigger connection pool reset and emit P1 alert `ALERT_BULKHEAD_RECOVERY_LAG`.
    - **Verification**: Automated drill script `scripts/drills/run_bulkhead_recovery_test.sh` exiting 0 with reproducible metrics.
    
    ---
    
    ### Adversarial Cases and Routing
    
    #### 1. Reject Infinite Retry [ADV-IR-01]
    - **Vulnerability**: Upstream clients or checkout microservices configuring unconstrained or immediate retry loops upon receiving HTTP 503/429 errors, creating a retry storm that overwhelms ingress classifiers.
    - **Adversarial Mechanism**: In incident INC-4491, upstream checkout service retried failed tax calls indefinitely every 50 ms without jitter or backoff, inflating traffic from 150 req/sec to 1,800 req/sec and causing CPU exhaustion on ingress gateways.
    - **Enforcement & Diagnostic**: Enforce strict rate-limiting and retry guidance at ingress. Rejection responses MUST include `Retry-After: 5` header and RFC 9457 diagnostic `ERR_RETRY_STORM_BLOCKED`. Upstream callers exceeding 3 retries within 10 seconds are rate-limited with HTTP 429.
    - **Forbidden Output Behavior**: The system is strictly forbidden from permitting infinite, unbounded, or non-backoff retries, and forbidden from processing requests lacking standard client backoff headers.
    
    #### 2. Reject Shared Failure Domain [ADV-FD-01]
    - **Vulnerability**: Partitioning worker threads at the application level while leaving underlying physical resources (HikariCP database connection pool, DNS resolver, network socket buffers) in a single shared pool.
    - **Adversarial Mechanism**: A slow dependency holds database transactions open while waiting for external I/O; although worker threads are segregated, all 60 shared database connections are exhausted by the slow service, starving critical payment writes.
    - **Enforcement & Diagnostic**: Partitioning must span execution threads AND database connections end-to-end. Database connections are allocated as three distinct, non-shared HikariCP datasource pools (`auth_ds` 40, `fraud_ds` 15, `tax_ds` 5). If any code path accesses a shared database connection across compartments, CI linter emits diagnostic `ERR_SHARED_FAILURE_DOMAIN_DETECTED`.
    - **Forbidden Output Behavior**: The system is strictly forbidden from sharing a common database connection pool or unbounded shared I/O executor across critical and non-critical workload compartments.
    
    #### 3. Reject Untested Compensation [ADV-UC-01]
    - **Vulnerability**: Implementing fallback or compensation routines (such as offline estimated tax calculation) that have not been validated under load or verified for financial calculation accuracy, leading to ledger discrepancies during live degradation.
    - **Adversarial Mechanism**: When tax calculation fails fast, unverified fallback engine calculates 0% tax for all jurisdictions, violating legal compliance and producing irreconcilable financial settlement errors.
    - **Enforcement & Diagnostic**: Any fallback or compensation path must be covered by pre-computed, verified lookup tables and backed by automated contract tests asserting rate accuracy within +/- 0.5% of authoritative rates. Unverified fallbacks trigger diagnostic `ERR_UNTESTED_COMPENSATION_REJECTED` and block deployment.
    - **Forbidden Output Behavior**: The system is strictly forbidden from falling back to arbitrary, unverified default values (e.g. zero tax) during bulkhead exhaustion without passing certified compensation verification test suites.
    
    ---
    
    ### Invariants and Contracts
    
        Physical Resource Isolation Invariant [INV-BH-01]
          No request in `bulkhead_tax_calc` or `bulkhead_fraud_eval` may borrow, acquire, or block
          threads or database connections allocated to `bulkhead_card_auth`.
    
        Zero-Queue Load Shedding Invariant [INV-BH-02]
          The tax calculation bulkhead must maintain a maximum queue capacity of 0. When threads are
          exhausted, requests must fail fast within 5 ms rather than queueing behind stalled sockets.
    
        Independent Database Pool Quotas [INV-BH-03]
          HikariCP connection pools must be instantiated as separate, non-shared data sources:
          `card_auth_ds` (40 conns), `fraud_ds` (15 conns), and `tax_ds` (5 conns).
    
        Finite Retry Enforcement [INV-BH-04]
          All rejection responses emitted by saturated bulkheads must carry `Retry-After` headers and
          enforce a maximum of 3 upstream retry attempts.
    
    ## Explicit Unknowns
    
    - Memory footprint of secondary cached geographic tax lookup tables under JVM off-heap storage (G-1).
    - Socket timeout propagation latency across AWS network load balancers during upstream TCP resets (G-2).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | 120 Tomcat threads, 60 HikariCP connections | provided | Infrastructure intake | Current |
    | Peak 1,400 card authorizations/sec | provided | Traffic intake | Current |
    | Incident INC-4491 42-minute cascade outage | provided | Post-mortem evidence | Historical |
    | Tax latency brownout spikes to 8,000 ms | provided | Incident metric logs | Historical |
    | Dedicated 75/30/15 thread partition split | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
    | Dedicated 40/15/5 DB connection split | decided | Elena Rostova & Database Guild | 2026-09-15 |
    | Zero-queue fail-fast load shedding | decided | Architectural invariant INV-BH-02 | 2026-09-15 |
    | Reject infinite retries with Retry-After: 5 | decided | Architectural invariant INV-BH-04 | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against bulkhead architecture contracts:
    - **Blast Radius Isolation**: PASS. Tax latency spike to 8,000 ms locks at most 15 threads and 5 DB connections.
    - **Card Auth Protection**: PASS. 75 threads and 40 DB connections are physically reserved for core authorizations.
    - **Fail-Fast Shedding**: PASS. Zero-queue sizing sheds tax load in < 5 ms, triggering fallback rules.
    - **Shared Failure Domain Elimination**: PASS. HikariCP connection pools segregated into independent datasources.
    - **Adversarial Retry Defense**: PASS. Upstream infinite retries rejected with HTTP 503 / 429 and Retry-After headers.
    - **Compensation Integrity**: PASS. Cached geographic tax engine validated against state tax rate tables.
    
    ## Open Decisions
    
    - `DEC-BH-01`: Elena Rostova to confirm whether HikariCP `leakDetectionThreshold` should be set to 2,000 ms across all three connection pools (Owner: Elena Rostova).
    
    ## Next steps
    
    1. Elena Rostova refactors Spring Boot configuration into three discrete `ThreadPoolTaskExecutor` beans and three isolated `HikariDataSource` beans.
    2. SRE team deploys Prometheus alerting rules on `bulkhead_threads_active` and `bulkhead_rejection_total`.
    3. Execute controlled chaos drill simulating 10,000 ms third-party tax API latency to verify 100% auth success rate under load.
    

    resource-bulkhead-isolation-design.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Map workload classes to bounded resource compartmentsDefine thread pool and connection segregation boundariesSpecify admission control and queue bounds for stabilityIdentify residual shared dependencies and failure modes

    About this skill

    What it does

    This skill maps accepted workload classes and failure modes into bounded resource compartments. It defines what is isolated, which capacity each compartment can consume, what happens at exhaustion and which shared dependencies still permit correlated failure.

    Use it when

    Use when one workload/dependency can exhaust or block resources needed by another and exact containment/degradation semantics are required.

    For example: “Our order processing system hangs when the third-party address validation API slows down. Payment processing threads get blocked waiting for available threads, causing checkout failures.”

    What you get

    • Bulkhead Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/bulkhead-design/.

    What it will not do

    Do not use for circuit-breaker, rate-limit, timeout/retry or load-shedding policy, autoscaling, cell/shard/zone topology, platform selection, implementation or incidents.

    How it works

    1. Check resource partitioning is required.
    2. Bound the isolation unit and resource pools.
    3. Establish pool allocation and capacity limits.
    4. Define admission and queue bounds.
    5. Specify permit acquisition and release lifecycles.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions