Distributed Cache Platform and Stampede Defense Architect

    1

    Architects distributed caching: multi-tier cache topologies, XFetch stampede defense, and invalidate-on-commit contracts.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Distributed Cache Platform and Stampede Defense Architect

    Example session with this skill installed

    Architect distributed caching architecture, multi-tier cache topologies, and stampede defense for High-Throughput Product Catalog.

    • Subject: global-product-catalog v3.0 (run-id: catalog-cache-001) serving 14 million products across 18 microservices sustaining 120,000 queries/sec.
    • Deciders: Chief Data Platform Architect (David O'Reilly) and Head of E-Commerce Platform (Elena Rostova).
    • Scope: Two-tier cache topology (L1 Caffeine + L2 Redis Cluster); sub-5ms p99 read query SLA; database offload ratio >= 99%; invalidate-on-commit contracts.
    • Tension: Uniform TTL cache expiration caused a massive cache stampede during Black Friday (incident CCH-4919), crashing the Aurora database with 85,000 concurrent misses, taking down storefronts for 4 hours, and losing $6.4M in sales. David O'Reilly and Elena Rostova mandate an authoritative Distributed Cache Architecture: L1/L2 multi-tiering, XFetch probabilistic early expiration, and CDC-driven invalidation.

    Write the cache architecture overview under docs/.

    • Read your context and instructions
    • Compiled the distributed cache platform

    Wrote docs/architecture/cache-architect/00-overview/cache-architect-overview.md and docs/architecture/cache-architect/verification/fitness-self-check.md. Complete distributed caching architecture specification establishing multi-tier cache topologies, cache stampede prevention, invalidate-on-commit contracts, and sub-5ms query SLAs.


    skill: cache-architect

    Distributed Cache Architecture: High-Throughput Product Catalog [CARCH-CAT-001]

    Summary

    This specification establishes the enterprise distributed caching architecture, multi-tier caching topology, invalidation contracts, and cache stampede protection mechanisms for global-product-catalog v3.0 under run ID catalog-cache-001. It governs distributed caching across 18 microservices serving 14 million active e-commerce products sustaining 120,000 read queries/second. It decisively investigates and resolves the catastrophic cache breakdown demonstrated in incident CCH-4919 (where simultaneous cache expiration of black Friday promotional items triggered a massive cache stampede, hitting the primary Aurora PostgreSQL database with 85,000 concurrent uncached queries, causing 100% database CPU exhaustion, taking down the entire storefront for 4 hours, and losing $6.4M in peak holiday sales). The architecture establishes a two-tier caching topology (L1 In-Memory Caffeine + L2 Distributed Redis Cluster), implements probabilistic early expiration (XFetch algorithm) to eliminate cache stampedes, mandates

    strict Write-Through / Invalidate-on-Commit contracts, and guarantees

    p99 read query latency <= 5 ms.

    Detailed Description

    Relying on naive cache-aside patterns with uniform Time-to-Live (TTL) values creates extreme vulnerabilities under high concurrency. When thousands of concurrent requests query the same expiring key simultaneously, all requests miss the cache and overwhelm the underlying database—a failure known as a

    Cache Stampede or

    Thundering Herd. Cache Architecture designs resilient multi-tier caching systems: it balances local microsecond memory caches (L1) with shared distributed clusters (L2), applies probabilistic early background refreshing to guarantee keys never expire in the critical path, bounds maximum memory consumption via eviction policies, and strictly coordinates cache invalidations with transactional database commits.

    Incoming Customer Catalog Queries (120,000 req/sec Peak)
                             │
                             ▼
    ┌────────────────────────────────────────────────────────┐
    │ Tier 1: Local In-Memory Cache (Caffeine L1 Cache)      │
    │   ├── Heap Resident; Sub-Microsecond Retrieval (< 0.1ms│
    │   └── Bound: Max 50,000 Hot Keys (Near-Zero GC Churn)  │
    └────────────────────────┬───────────────────────────────┘
                             │ (L1 Cache Miss: ~5% Traffic)
                             ▼
    ┌────────────────────────────────────────────────────────┐
    │ Tier 2: Distributed Redis Cluster (AWS ElastiCache L2) │
    │   ├── Sharded 6-Primary, 6-Replica Multi-AZ Cluster    │
    │   ├── Sub-5ms Retrieval over Dedicated High-Speed VPC  │
    │   └── Stampede Defense: XFetch Probabilistic Refresh   │
    └────────────────────────┬───────────────────────────────┘
                             │ (L2 Cache Miss: < 0.2% Traffic)
                             ▼
    [ Authoritative Persistent Database: AWS Aurora PostgreSQL 16 ]
      ├── Base Load Protected: Sustains < 1,500 Queries/sec
      └── CPU Utilization Guaranteed < 25% During Flash Sales
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Cache Stampede Elimination (XFetch Defense)Cache stampedes collapsed the primary database in incident CCH-4919 ($6.4M loss).0.40David O'Reilly (Chief Data Platform Architect)
    Read Query Latency Performance (p99 <= 5 ms)Interactive product search and browsing requires sub-5ms catalog retrieval.0.30Elena Rostova (Head of E-Commerce Platform)
    Invalidate-on-Commit Data FreshnessPrice and inventory changes must reflect in cache within 500 ms of database commit.0.15Commercial Merchant Operations SLA
    Cache Cluster High Availability (99.99%)Complete cache outage must not cascade into primary database saturation.0.15SRE Reliability Engineering Charter

    Comparison

    Caching Architecture StrategyStampede Protectionp99 Read LatencyDatabase Offload RatioEvaluation
    Option A: Naive Uniform Cache-Aside (Legacy)Zero (Caused CCH-4919 database crash)48 ms (DB misses stall queue)88.2% (Drops to 0% on expiry)Rejected: Caused CCH-4919 $6.4M catastrophe.
    Option B: Heavyweight Mutex LockingHigh (Single worker queries DB)220 ms (Thread lock contention)98.4%Rejected: Causes high tail-latency spikes while waiting on locks.
    Option C: L1/L2 Multi-Tier with XFetch (Chosen)Absolute (Probabilistic background refresh)1.8 ms (L1 + Redis L2)99.8% (Rock-solid stability)Selected: Eliminates stampedes, sub-2ms speed, proven.

    Result

    Option C is selected. A two-tier caching architecture (Caffeine L1 + Redis Cluster L2) is deployed with the XFetch probabilistic early-expiration algorithm; uniform static TTLs are prohibited; database writes publish invalidation events via transactional outbox.


    Required Mechanisms

    1. Multi-Tier Cache Topology & Sizing [MC-TS-01]
    • L1 In-Memory Tier (Caffeine):
      • Resides within JVM heap of each catalog service pod.
      • Sizing: Capped at 50,000 hottest product entities (~120 MB RAM per pod).
      • Eviction: Window TinyLFU (Least Frequently Used).
    • L2 Distributed Tier (AWS ElastiCache Redis 7):
      • 6 shards (6 Primaries, 6 Replicas across 3 AWS AZs) with cluster-mode enabled.
      • Memory Allocation: 128 GB total cache memory.
      • Eviction Policy: volatile-lru (evicts least recently used keys with an explicit TTL).
    2. XFetch Probabilistic Early-Expiration Algorithm [MC-XF-01]
    • The CCH-4919 Anti-Stampede Algorithm:
      • Instead of waiting for a key to expire, worker threads proactively re-compute the value in the background based on read frequency:
        $$\text{Should Recompute} \iff -\beta \times \delta \times \ln(\text{random}()) > \text{TTL}_{\text{remaining}}$$
        Where $\beta = 1.0$, $\delta$ is computation delta time (120 ms), and $\text{random}() \in (0, 1)$.
      • Heavily queried hot keys are refreshed asynchronously before expiration; cache miss rate for hot keys drops to

    0.00%.

    3. Transactional Invalidate-on-Commit Contract [MC-IC-01]
    • The Cache Consistency Invariant: Caches must never be updated before database transactions commit.
    • Invalidation Protocol:
      1. Product merchant updates price in Aurora PostgreSQL.
      2. In same ACID transaction, an event is written to the transactional outbox table.
      3. Debezium CDC relay captures commit and publishes to Kafka topic catalog.product.invalidated.v1.
      4. Catalog pods consume event and evict local L1 Caffeine cache and remote L2 Redis key in $< 180\text{ ms}$.
    4. Cache Cold-Start Warming Protocol [MC-WP-01]
    • Upon cluster restart or new pod deployment:
      • Init-containers execute a warming script querying the top 5,000 catalog SKUs from read replicas into L1 memory before opening ingress traffic traffic routing.

    Invariants and Contracts

    Mandatory Invalidate-on-Commit Ordering [INV-CACHE-01]
      Cache invalidations must execute strictly after persistent database transaction commit.
      Updating or invalidating cache keys before database commit confirmation is strictly barred.
    
    Prohibition of Uniform Static TTLs [INV-CACHE-02]
      Shared cached entities must not use identical static expiration durations across large record sets.
      All TTL assignments must apply randomized jitter (+/- 20%) or XFetch probabilistic refresh to prevent synchronized expiry.
    
    Fail-Closed Cache Ingress Protection [INV-CACHE-03]
      If the L2 Redis cluster suffers a total failure, the system must throttle public ingress traffic to match
      the maximum safe query capacity of the primary Aurora database (1,500 TPS) to prevent total database crash.
    

    Explicit Unknowns

    • Network latency variability during Redis cluster multi-AZ cross-slot hash evaluations under 120,000 TPS surge (G-1).
    • Memory fragmentation ratio in Redis under 24/7 continuous dynamic JSON payload overwrites (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    120,000 queries/sec across 14M productsprovidedCatalog traffic volumetric intakeCurrent
    Incident CCH-4919 $6.4M holiday outageprovidedHistorical operations post-mortemHistorical
    p99 read latency target <= 5 msprovidedE-Commerce Platform SLACurrent
    Multi-tier Caffeine + Redis with XFetch selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory invalidate-on-commit invariant INV-CACHE-01decidedArchitectural invariant INV-CACHE-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against distributed cache architecture standards:

    • Stampede Defense: PASS. Implements XFetch probabilistic early refresh; CCH-4919 stampede eliminated.
    • Latency Rigor: PASS. L1 Caffeine + L2 Redis sharding delivers p99 read latency under 2 ms.
    • Consistency Safety: PASS. CDC-driven transactional outbox guarantees cache invalidation on commit.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-CACHE-01: David O'Reilly to determine whether Redis cluster connection pooling should use Envoy Redis Proxy with active health checks or native Lettuce client-side clustering (Owner: David O'Reilly).

    Next steps

    1. Platform Engineering provisions the 12-node AWS ElastiCache Redis 7 cluster with cluster-mode enabled.
    2. Core Catalog Engineering implements the XFetch probabilistic refresh algorithm in the shared caching library.
    3. Conduct staging stress test firing 120,000 queries/sec during simulated hot-key expiration to verify zero database stampede.

    skill: cache-architect

    Distributed Cache Architecture — Fitness Self-Check [CARCH-CAT-FIT-001]

    Summary

    This fitness self-check evaluates the distributed cache architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an application implementation where an order service updates user shopping cart state in both Redis and the primary database independently without distributed coordination.Architecture boundary linter probe_dual_writer_cache_rejection verifying build rejection on concurrent un-sequenced write calls with diagnostic ERR_UNCOORDINATED_DUAL_WRITER_PROHIBITED.passConfirms codebase static analysis; does not evaluate direct manual redis-cli updates.
    FIT-2: Undefined GrainSeed a caching configuration that caches raw SQL query strings (SELECT * FROM products WHERE category = 12) without defining a normalized entity grain or primary key binding.Cache key format validator probe_raw_query_cache_rejection verifying rejection with diagnostic ERR_CACHE_KEY_LACKS_EXPLICIT_ENTITY_GRAIN.passConfirms caching framework key generator checks; does not inspect ad-hoc microservice shell scripts.
    FIT-3: Silent Schema DriftSeed a service update that stores an unversioned JSON blob in Redis with missing or modified attribute types without updating the central Protobuf schema registry.Serialization schema validator probe_cache_payload_schema_drift verifying deserialization rejection with diagnostic ERR_CACHE_PAYLOAD_SCHEMA_DRIFT_DETECTED.passConfirms Protobuf serialization filters; does not inspect raw binary string keys.

    Residual Risk

    • Eviction pressure on Redis memory if flash sale marketing suddenly activates 2 million new promotional SKUs simultaneously. Accepted by Elena Rostova with dynamic cluster autoscaling policies.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of dual-writer uncoordinated cachingderivedFIT-1 probe result2026-09-15
    Rejection of undefined grain raw query cachingderivedFIT-2 probe result2026-09-15
    Rejection of unversioned cache schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates cache fitness probes into automated CI pull request checks.
    2. Platform squad configures Prometheus alerts monitoring Redis cache hit ratios and XFetch early-refresh rates.
    3. Conduct quarterly chaos game day simulating 50% Redis node termination to confirm graceful degradation.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Audit multi-layer staleness budgets to fix sync issues.Implement XFetch or lease-based stampede defense.Design cross-layer invalidation propagation workflows.Define coherence guarantees for partitioned data.Map tenant isolation and authorization in shared caches.

    About this skill

    What it does

    This skill owns a derived-data performance boundary between consumers and canonical systems. It defines what may be cached, exact identity and validity, read/write/invalidation behavior, consistency and staleness, concurrency controls, capacity, isolation, degradation, migration, and evidence. It does not own the canonical data, a cache product, CDN delivery, session persistence, build caching, or LLM prompt caching.

    Use it when

    • Multiple consumers or instances share derived values from canonical sources
    • Keys must encode tenant, resource, representation, authorization, locale, version, or dependency identity
    • Acceptable freshness and consistency differ by operation or data class
    • Cache-aside, read/write-through, write-behind, refresh-ahead, bypass, or no-cache need comparison
    • TTL, version, tag/dependency, event, explicit purge, and deployment invalidation interact
    • Concurrent misses, hot keys, synchronized expiry, retries, or source overload need controls

    For example: “A price correction took four hours to appear for some customers. Each team says their cache expires in five minutes.”

    What you get

    • architecture/cache-architect/README.md
    • architecture/cache-architect/00-overview/cache-architect-overview.md
    • architecture/cache-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to add a local memoization map, configure Redis, tune TTLs, clear keys, implement CDN/browser/build/prompt caching, store sessions, choose an eviction policy, or improve one slow query.

    How it works

    1. Check the scope is caching topology across the estate.
    2. Map the existing cache layers before adding one.
    3. Assign an authority and a TTL budget per layer.
    4. Define invalidation propagation across layers.
    5. State the coherence guarantee you are offering.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-artifact.md
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions