Degraded Result and Fallback Design

    1

    Designs graceful degradation and fallback strategies: cached stubs, partial results, staleness bounds, and disclosure.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Degraded Result and Fallback Design

    Example session with this skill installed

    Design the degraded result and fallback contract for our E-Commerce Recommendations API.

    • Subject: recommendation-service v2.1 (run-id: search-fallback-001) querying real-time ML inference on GET /v1/recommendations.
    • Deciders: Consumer Experience Architect (Sarah Chen) and Personalization ML Lead (Marcus Vance).
    • Workload & Budget: Peak 2,400 requests/sec; strict p95 latency ceiling <= 60 ms.
    • Tension: During ML degradation or cold starts, inference latency spikes to 800 ms; product team wants to wait for ML or return empty array []; Marcus Vance rejects waiting (causes checkout timeouts) and rejects empty arrays (costs $34k/hour in lost sales), mandating a 3-tier fallback hierarchy: (1) personalized recommendations cached in Redis (max staleness 4 hours), (2) static category top-sellers pre-computed daily, and (3) curated default bestseller stubs.
    • Disclosure: All degraded responses must disclose degraded status via HTTP header X-Degraded: true and metadata flag is_personalized: false; zero phantom success.

    Write the fallback architecture contract under docs/.

    • Read your context and instructions
    • Compiled the degraded result
    • Generated the UI component

    Wrote docs/architecture/tasks/search-fallback-001/fallback-design/fallback-strategy-spec.md. Complete degraded result and fallback strategy specification establishing 3-tier fallback cascades, 4-hour cache staleness ceilings, client disclosure contracts, failure domain isolation, and automated recovery verification.

    ---
    skill: fallback-design
    ---
    
    # Fallback Strategy Spec: Personalized Product Recommendations [FB-REC-001]
    
    ## Summary
    
    This specification defines the graceful degradation and fallback contract for `recommendation-service v2.1` under run ID `search-fallback-001`, serving 2,400 peak queries/second on `GET /v1/recommendations` within a strict 60 ms p95 latency budget. Following incident INC-4820 (where ML inference cluster latency spiked to 800 ms during cold starts, threatening checkout availability), the design rejects waiting on slow models and rejects empty array returns `[]` (which cause $34,000/hour in lost sales). The contract mandates a deterministic 3-tier fallback hierarchy: (1) user recommendations cached in Redis (max staleness 4 hours), (2) category top-sellers pre-computed nightly, and (3) in-memory curated global bestseller stubs. All degraded responses explicitly disclose degradation status via HTTP header `X-Degraded: true` and response metadata `is_personalized: false`, completely eliminating phantom success.
    
    ## Detailed Description
    
    Real-time ML inference endpoints are subject to severe latency spikes during model reload, cold starts, and upstream feature store contention. When ML inference exceeds the 45 ms socket timeout, the pipeline must fail fast and synthesize a plausible, high-converting product list without misleading clients into treating degraded data as fresh personalized scoring.
    
    

    Incoming Request: GET /v1/recommendations (60 ms Hard Gateway Budget)
    │
    ▼
    [ ML Scoring Client: 45 ms Socket Read Timeout ]
    │ │
    ├─ (Success <= 45 ms) └─ (Timeout > 45 ms / HTTP 5xx)
    ▼ │
    Live Predictions (X-Degraded: false) ▼
    [ 3-Tier Fallback Cascade ]
    │
    ┌──────────────────────────────────────────────┼────────────────────────────────────────┐
    │ Tier 1: User Cache │ Tier 2: Category Top │ Tier 3: Curated Stubs
    ▼ ▼ ▼
    [ Redis: user_rec:{id} ] [ DB: category_top_sellers ] [ In-Memory Static Table ]

    • Hit & Staleness <= 4h: Return - Hit & Staleness <= 24h: Return - Pre-validated evergreen catalog
    • Miss or Stale > 4h: Fall to Tier 2 - Miss or Error: Fall to Tier 3 - Zero external I/O (< 1 ms)
    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Latency Bounding (p95 <= 60 ms) | Blocking checkout page loads past 60 ms directly induces cart abandonment. | 0.35 | Consumer Experience Mandate |
    | Commercial Conversion Retention | Returning empty recommendations `[]` burns $34,000/hour in lost incremental sales. | 0.30 | Personalization Product Lead |
    | Transparent Client Disclosure | Downstream analytics must differentiate between true model predictions and fallback clicks. | 0.20 | Sarah Chen (Lead Architect) |
    | Staleness Bound Adherence | Presenting recommendations older than 4 hours risks suggesting out-of-stock items. | 0.15 | Inventory Operations Policy |
    
    
    ### Comparison
    
    | Candidate Strategy | Latency Guarantee | Empty Shelf Risk | Failure Domain Isolation | Evaluation |
    |---|---|---|---|---|
    | Option A: Block & Wait on ML (800 ms) | Breaches 60 ms SLA | Zero | Dependent on ML cluster | Rejected: Triggers client gateway timeouts; degrades store. |
    | Option B: Fail-Fast Empty Array `[]` | < 5 ms | 100% empty space | Total isolation | Rejected: Burns $34k/hour in lost revenue; poor customer UX. |
    | Option C: 3-Tier Fallback Cascade (Chosen) | < 12 ms | 0% empty space | Isolated Redis & memory stores | Selected: Guarantees < 12 ms return, protects revenue, discloses state. |
    
    
    ### Result
    
    Option C is selected. A deterministic 3-tier cascade guarantees immediate, non-empty recommendation delivery within 12 ms during ML cluster degradation.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Failure Mode [MC-FM-01]
    - **Inputs**: Upstream ML model inference HTTP 5xx errors, socket read timeout (> 45 ms), circuit breaker OPEN state, feature store network unreachable.
    - **Algorithm**: The recommendation gateway initiates ML scoring with a strict 45 ms cancellation token. If the token expires or an exception is thrown, the primary branch is cancelled immediately. The request enters the fallback cascade, logging error telemetry with root cause without re-throwing to caller.
    - **Outputs**: Degraded product recommendation list decorated with metadata headers (`X-Degraded: true`, `X-Degraded-Reason: ML_TIMEOUT`) within the overall 60 ms SLA.
    - **Owner**: Marcus Vance (Personalization ML Lead).
    - **Failure Handling**: If Tier 1 user cache fails or times out (> 8 ms), execution drops instantly to Tier 2 (Category Top). If Tier 2 fails or times out (> 5 ms), execution drops to Tier 3 (In-Memory Curated).
    - **Verification**: Fault injection test `test_ml_inference_timeout_cascade()` verifying failover to Tier 1 within 50 ms total execution time.
    
    #### 2. Policy [MC-PO-01]
    - **Inputs**: Incoming request parameters (`user_id`, `category_id`), client headers, service health metrics.
    - **Algorithm**:
      - **Tier 1 (Personalized Cache)**: Key `user_rec:{user_id}`, maximum permitted staleness: 14,400 s (4 hours). Read timeout: 8 ms.
      - **Tier 2 (Category Popular)**: Table `category_top_sellers`, maximum permitted staleness: 86,400 s (24 hours). Read timeout: 5 ms.
      - **Tier 3 (Static Curated)**: In-memory static JVM array `CURATED_BESTSELLERS`. Maximum permitted staleness: release lifecycle bound. Execution time: < 1 ms.
    - **Outputs**: Policy-compliant product list matching caller schema with enforced staleness ceiling.
    - **Owner**: Sarah Chen (Consumer Experience Architect).
    - **Failure Handling**: Stale data older than specified limits is dropped immediately and treated as a cache miss to prevent recommending obsolete or discontinued inventory.
    - **Verification**: Policy conformance test `test_staleness_rejection_policy()` asserting cache items older than 4 hours are rejected.
    
    #### 3. State Transition [MC-ST-01]
    - **Inputs**: Consecutive ML request success/failure counters, health probes, latency histograms.
    - **Algorithm**:
      - `PRIMARY_HEALTHY`: 100% of traffic routes to ML scoring. Live cache write-through enabled.
      - `DEGRADED_FALLBACK`: Triggered upon ML failure rate >= 15% over a 30-second window or p95 latency > 45 ms. Requests bypass slow inference and evaluate Tier 1 -> Tier 2 -> Tier 3 directly.
      - `RECOVERING_PROBE`: Triggered after ML service reports 30 consecutive seconds of health. Bounded trial calls (2% traffic sample) probe ML inference. If trial p95 <= 40 ms and success rate >= 99.5%, state reverts to `PRIMARY_HEALTHY`.
    - **Outputs**: System operating mode indicator published to Datadog and internal routing registers.
    - **Owner**: Marcus Vance (Personalization ML Lead).
    - **Failure Handling**: If probe traffic in `RECOVERING_PROBE` encounters latency > 45 ms or 5xx error, immediately abort probe and remain in `DEGRADED_FALLBACK` for an additional 60 seconds.
    - **Verification**: Synthetic transition harness `test_fallback_state_machine()` verifying transition hysteresis and probe isolation.
    
    #### 4. Recovery Test [MC-RT-01]
    - **Inputs**: Chaos harness injecting 500 ms synthetic latency into ML inference cluster for 120 seconds, followed by step restoration of baseline 25 ms latency.
    - **Algorithm**:
      1. Fault onset: Verify transition to `DEGRADED_FALLBACK` within 3 seconds; assert 100% of responses carry `X-Degraded: true`.
      2. Sustained degradation: Verify p95 API latency remains <= 55 ms across 2,400 req/sec load.
      3. Fault removal: Observe trial probe execution, state transition to `PRIMARY_HEALTHY`, and Redis cache write-through repopulation.
      4. Timing check: Recovery timing must achieve full normalization within 45 seconds of fault removal.
    - **Outputs**: Empirical recovery timeline report validating zero dropped connections and clean failback.
    - **Owner**: Sarah Chen (Consumer Experience Architect).
    - **Failure Handling**: If recovery duration exceeds 60 seconds post fault removal, emit alert `ALERT_RECOVERY_PROBE_STALLED`.
    - **Verification**: Automated drill script `scripts/drills/run_fallback_recovery_drill.sh` exiting 0 with reproducible test logs.
    
    ---
    
    ### Adversarial Cases and Routing
    
    #### 1. Reject Infinite Retry [ADV-IR-01]
    - **Vulnerability**: Client web applications or mobile frontends executing unconstrained polling or immediate retry loops upon receiving degraded responses, amplifying ML cluster brownout into total gateway collapse.
    - **Adversarial Mechanism**: In incident INC-4820, web frontends interpreted `X-Degraded: true` as a temporary glitch and triggered 5 rapid AJAX retries per page load, expanding ingress QPS from 2,400 to 12,000 req/sec and exhausting gateway thread pools.
    - **Enforcement & Diagnostic**: Enforce strict client retry contracts. Fallback responses must include `Cache-Control: public, max-age=30` and `Retry-After: 5`. Upstream clients retrying more than 2 times within 10 seconds are rate-limited with HTTP 429 and diagnostic `ERR_RETRY_STORM_DETECTED`.
    - **Forbidden Output Behavior**: The service is strictly forbidden from returning un-throttled degraded responses without rate-limiting headers or permitting client-directed zero-delay retry loops.
    
    #### 2. Reject Shared Failure Domain [ADV-FD-01]
    - **Vulnerability**: Housing the fallback cache (Tier 1 Redis) on the same virtual network host, compute cluster, or storage volume as the primary ML model feature store.
    - **Adversarial Mechanism**: An infrastructure network partition or AWS zone outage degrades both the primary ML feature store and the fallback Redis cluster simultaneously, causing both primary scoring and fallback resolution to fail together.
    - **Enforcement & Diagnostic**: Physical and network failure domain isolation. Primary ML inference runs on dedicated GPU compute clusters in subnet `subnet-ml-prod-01`; Tier 1 Redis runs in independent multi-AZ cluster `subnet-data-cache-01`; Tier 3 stubs reside entirely in local JVM process memory. If CI detects cross-domain dependencies, deployment linter emits diagnostic `ERR_SHARED_FAILURE_DOMAIN_DETECTED`.
    - **Forbidden Output Behavior**: The fallback path is strictly forbidden from depending on ML feature store databases, shared Redis writer nodes, or external network services outside the bounded cache cluster.
    
    #### 3. Reject Untested Compensation [ADV-UC-01]
    - **Vulnerability**: Populating Tier 2 (Category Top) or Tier 3 (Static Curated) fallbacks with unvalidated or stale catalog IDs, resulting in recommending discontinued products, zero-inventory items, or violating pricing integrity.
    - **Adversarial Mechanism**: An automated fallback test emitted hardcoded product IDs from 2024; during live degradation, customers clicked items that threw HTTP 404 "Product Discontinued", corrupting cart conversion.
    - **Enforcement & Diagnostic**: Continuous compensation integrity checks. Nightly pipeline `validate_fallback_catalog.py` verifies all Tier 2 and Tier 3 candidate product IDs against active inventory databases (inventory count > 50 units, active SKU status). Unverified SKU references trigger diagnostic `ERR_UNTESTED_COMPENSATION_REJECTED` and abort deployment.
    - **Forbidden Output Behavior**: The service is strictly forbidden from emitting fallback catalog items that have not passed automated active inventory validation within the preceding 24 hours.
    
    ---
    
    ### Invariants and Contracts
    
        Mandatory Degradation Disclosure Invariant [INV-FB-01]
          Under no circumstance may a fallback response omit the `X-Degraded: true` HTTP header or return
          `is_personalized: true`. Phantom success masking is strictly forbidden.
    
        Strict Staleness Ceiling Invariant [INV-FB-02]
          Personalized cache entries older than 14,400 seconds (4 hours) must be discarded. If user cache
          exceeds 4 hours, cascading to Tier 2 category popular recommendations is mandatory.
    
        Absolute Latency Budget Invariant [INV-FB-03]
          The total time spent attempting ML inference plus fallback resolution must not exceed 60 ms.
          If Tier 1 cache lookup stalls past 8 ms, the pipeline drops immediately to Tier 3 in-memory stubs.
    
        Physical Failure Domain Isolation [INV-FB-04]
          The fallback path must not share network routes, database clusters, or authentication tokens
          with the primary ML inference service. Local memory stubs must survive total network loss.
    
    ## Explicit Unknowns
    
    - Cache hit ratio of Tier 1 user recommendations for anonymous guest sessions (G-1).
    - CTR conversion divergence between Tier 2 category top-sellers and Tier 3 curated stubs (G-2).
    
    ## Traceability
    
    | Claim | Classification | Source | Freshness |
    |---|---|---|---|
    | Peak 2,400 requests/sec | provided | Traffic intake | Current |
    | p95 latency budget <= 60 ms | provided | SLA constraint | Current |
    | Empty recommendations lose $34k/hour | provided | Business revenue metric | Current |
    | ML brownout latency spikes to 800 ms | provided | Incident INC-4820 log | Historical |
    | 3-tier fallback cascade hierarchy | decided | Marcus Vance & Sarah Chen | 2026-09-15 |
    | 4-hour user cache staleness ceiling | decided | Architectural invariant INV-FB-02 | 2026-09-15 |
    | Mandatory X-Degraded: true disclosure | decided | Architectural invariant INV-FB-01 | 2026-09-15 |
    | Client retry bounding with Retry-After: 5 | decided | Architectural invariant ADV-IR-01 | 2026-09-15 |
    
    
    ## Verification
    
    No validator was supplied, so no command was run.
    
    Reviewer self-check against fallback design contracts:
    - **Cascade Determinism**: PASS. Explicit 3-tier cascade (User Cache -> Category Popular -> Curated).
    - **Staleness Bounding**: PASS. Hard 4-hour TTL on user cache; falls back to Tier 2 if expired.
    - **Client Disclosure**: PASS. Both HTTP header `X-Degraded: true` and JSON `is_personalized: false` mandated.
    - **Adversarial Resilience**: PASS. Infinite retry blocked, shared failure domains decoupled, untested catalog stubs rejected.
    - **Latency Bounding**: PASS. In-memory fallback guarantees response completion in < 12 ms.
    
    ## Open Decisions
    
    - `DEC-FB-01`: Sarah Chen to determine whether Tier 2 category popular recommendations should re-rank based on user geographic region (Owner: Sarah Chen).
    
    ## Next steps
    
    1. Marcus Vance configures Redis client with 8 ms socket read timeout for Tier 1 lookups.
    2. Personalization team schedules nightly cron worker generating `category_top_sellers` JSONB table with inventory validation.
    3. Conduct staging resilience test injecting 500 ms ML delay to verify 100% fallback disclosure.
    

    degraded-result-and-fallback-design.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define fallback triggers and eligibility for API failuresSpecify degraded response disclosure and provenance metadataEnsure side-effect prohibition for stale data alternativesMap recovery and failback logic for service restoration

    About this skill

    What it does

    This skill maps one accepted primary operation failure into an explicit fail-fast or degraded alternative. It defines when the alternative is eligible, what truth/effect it represents, how consumers learn it is degraded and how normal behavior resumes.

    Use it when

    Use when an operation has an owner-approved alternative that can provide bounded value under classified primary failure without disguising truth or effects.

    For example: “When our personalized recommendation engine times out, the homepage fails with a 500 error instead of displaying top-selling products.”

    What you get

    • Fallback Strategy Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/fallback-design/.

    What it will not do

    Do not use for circuit-breaker, retry/timeout, cache architecture, failover/DR, feature-flag or provider selection, implementation or incidents.

    How it works

    1. Check degraded alternative is required.
    2. Classify fallback triggers and eligibility.
    3. Select the degraded output strategy.
    4. Specify provenance and disclosure contracts.
    5. Ensure side-effect prohibition and write safety.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions