- Home
- Skills
- Data & Databases
- Distributed Cache Platform and Stampede Defense Architect
Distributed Cache Platform and Stampede Defense Architect
Architects distributed caching: multi-tier cache topologies, XFetch stampede defense, and invalidate-on-commit contracts.
$9
Works with the AI tools you already use
Distributed Cache Platform and Stampede Defense Architect
Example session with this skill installed
Architect distributed caching architecture, multi-tier cache topologies, and stampede defense for High-Throughput Product Catalog.
- Subject: global-product-catalog v3.0 (run-id: catalog-cache-001) serving 14 million products across 18 microservices sustaining 120,000 queries/sec.
- Deciders: Chief Data Platform Architect (David O'Reilly) and Head of E-Commerce Platform (Elena Rostova).
- Scope: Two-tier cache topology (L1 Caffeine + L2 Redis Cluster); sub-5ms p99 read query SLA; database offload ratio >= 99%; invalidate-on-commit contracts.
- Tension: Uniform TTL cache expiration caused a massive cache stampede during Black Friday (incident CCH-4919), crashing the Aurora database with 85,000 concurrent misses, taking down storefronts for 4 hours, and losing $6.4M in sales. David O'Reilly and Elena Rostova mandate an authoritative Distributed Cache Architecture: L1/L2 multi-tiering, XFetch probabilistic early expiration, and CDC-driven invalidation.
Write the cache architecture overview under docs/.
- Read your context and instructions
- Compiled the distributed cache platform
Wrote docs/architecture/cache-architect/00-overview/cache-architect-overview.md and docs/architecture/cache-architect/verification/fitness-self-check.md. Complete distributed caching architecture specification establishing multi-tier cache topologies, cache stampede prevention, invalidate-on-commit contracts, and sub-5ms query SLAs.
skill: cache-architect
Distributed Cache Architecture: High-Throughput Product Catalog [CARCH-CAT-001]
Summary
This specification establishes the enterprise distributed caching architecture, multi-tier caching topology, invalidation contracts, and cache stampede protection mechanisms for global-product-catalog v3.0 under run ID catalog-cache-001. It governs distributed caching across 18 microservices serving 14 million active e-commerce products sustaining 120,000 read queries/second. It decisively investigates and resolves the catastrophic cache breakdown demonstrated in incident CCH-4919 (where simultaneous cache expiration of black Friday promotional items triggered a massive cache stampede, hitting the primary Aurora PostgreSQL database with 85,000 concurrent uncached queries, causing 100% database CPU exhaustion, taking down the entire storefront for 4 hours, and losing $6.4M in peak holiday sales). The architecture establishes a two-tier caching topology (L1 In-Memory Caffeine + L2 Distributed Redis Cluster), implements probabilistic early expiration (XFetch algorithm) to eliminate cache stampedes, mandates
strict Write-Through / Invalidate-on-Commit contracts, and guarantees
p99 read query latency <= 5 ms.
Detailed Description
Relying on naive cache-aside patterns with uniform Time-to-Live (TTL) values creates extreme vulnerabilities under high concurrency. When thousands of concurrent requests query the same expiring key simultaneously, all requests miss the cache and overwhelm the underlying database—a failure known as a
Cache Stampede or
Thundering Herd. Cache Architecture designs resilient multi-tier caching systems: it balances local microsecond memory caches (L1) with shared distributed clusters (L2), applies probabilistic early background refreshing to guarantee keys never expire in the critical path, bounds maximum memory consumption via eviction policies, and strictly coordinates cache invalidations with transactional database commits.
Incoming Customer Catalog Queries (120,000 req/sec Peak)
│
▼
┌────────────────────────────────────────────────────────┐
│ Tier 1: Local In-Memory Cache (Caffeine L1 Cache) │
│ ├── Heap Resident; Sub-Microsecond Retrieval (< 0.1ms│
│ └── Bound: Max 50,000 Hot Keys (Near-Zero GC Churn) │
└────────────────────────┬───────────────────────────────┘
│ (L1 Cache Miss: ~5% Traffic)
▼
┌────────────────────────────────────────────────────────┐
│ Tier 2: Distributed Redis Cluster (AWS ElastiCache L2) │
│ ├── Sharded 6-Primary, 6-Replica Multi-AZ Cluster │
│ ├── Sub-5ms Retrieval over Dedicated High-Speed VPC │
│ └── Stampede Defense: XFetch Probabilistic Refresh │
└────────────────────────┬───────────────────────────────┘
│ (L2 Cache Miss: < 0.2% Traffic)
▼
[ Authoritative Persistent Database: AWS Aurora PostgreSQL 16 ]
├── Base Load Protected: Sustains < 1,500 Queries/sec
└── CPU Utilization Guaranteed < 25% During Flash Sales
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Cache Stampede Elimination (XFetch Defense) | Cache stampedes collapsed the primary database in incident CCH-4919 ($6.4M loss). | 0.40 | David O'Reilly (Chief Data Platform Architect) |
| Read Query Latency Performance (p99 <= 5 ms) | Interactive product search and browsing requires sub-5ms catalog retrieval. | 0.30 | Elena Rostova (Head of E-Commerce Platform) |
| Invalidate-on-Commit Data Freshness | Price and inventory changes must reflect in cache within 500 ms of database commit. | 0.15 | Commercial Merchant Operations SLA |
| Cache Cluster High Availability (99.99%) | Complete cache outage must not cascade into primary database saturation. | 0.15 | SRE Reliability Engineering Charter |
Comparison
| Caching Architecture Strategy | Stampede Protection | p99 Read Latency | Database Offload Ratio | Evaluation |
|---|---|---|---|---|
| Option A: Naive Uniform Cache-Aside (Legacy) | Zero (Caused CCH-4919 database crash) | 48 ms (DB misses stall queue) | 88.2% (Drops to 0% on expiry) | Rejected: Caused CCH-4919 $6.4M catastrophe. |
| Option B: Heavyweight Mutex Locking | High (Single worker queries DB) | 220 ms (Thread lock contention) | 98.4% | Rejected: Causes high tail-latency spikes while waiting on locks. |
| Option C: L1/L2 Multi-Tier with XFetch (Chosen) | Absolute (Probabilistic background refresh) | 1.8 ms (L1 + Redis L2) | 99.8% (Rock-solid stability) | Selected: Eliminates stampedes, sub-2ms speed, proven. |
Result
Option C is selected. A two-tier caching architecture (Caffeine L1 + Redis Cluster L2) is deployed with the XFetch probabilistic early-expiration algorithm; uniform static TTLs are prohibited; database writes publish invalidation events via transactional outbox.
Required Mechanisms
1. Multi-Tier Cache Topology & Sizing [MC-TS-01]
- L1 In-Memory Tier (Caffeine):
- Resides within JVM heap of each catalog service pod.
- Sizing: Capped at 50,000 hottest product entities (~120 MB RAM per pod).
- Eviction: Window TinyLFU (Least Frequently Used).
- L2 Distributed Tier (AWS ElastiCache Redis 7):
- 6 shards (6 Primaries, 6 Replicas across 3 AWS AZs) with cluster-mode enabled.
- Memory Allocation: 128 GB total cache memory.
- Eviction Policy:
volatile-lru(evicts least recently used keys with an explicit TTL).
2. XFetch Probabilistic Early-Expiration Algorithm [MC-XF-01]
- The CCH-4919 Anti-Stampede Algorithm:
- Instead of waiting for a key to expire, worker threads proactively re-compute the value in the background based on read frequency:
$$\text{Should Recompute} \iff -\beta \times \delta \times \ln(\text{random}()) > \text{TTL}_{\text{remaining}}$$
Where $\beta = 1.0$, $\delta$ is computation delta time (120 ms), and $\text{random}() \in (0, 1)$. - Heavily queried hot keys are refreshed asynchronously before expiration; cache miss rate for hot keys drops to
- Instead of waiting for a key to expire, worker threads proactively re-compute the value in the background based on read frequency:
0.00%.
3. Transactional Invalidate-on-Commit Contract [MC-IC-01]
- The Cache Consistency Invariant: Caches must never be updated before database transactions commit.
- Invalidation Protocol:
- Product merchant updates price in Aurora PostgreSQL.
- In same ACID transaction, an event is written to the transactional outbox table.
- Debezium CDC relay captures commit and publishes to Kafka topic
catalog.product.invalidated.v1. - Catalog pods consume event and evict local L1 Caffeine cache and remote L2 Redis key in $< 180\text{ ms}$.
4. Cache Cold-Start Warming Protocol [MC-WP-01]
- Upon cluster restart or new pod deployment:
- Init-containers execute a warming script querying the top 5,000 catalog SKUs from read replicas into L1 memory before opening ingress traffic traffic routing.
Invariants and Contracts
Mandatory Invalidate-on-Commit Ordering [INV-CACHE-01]
Cache invalidations must execute strictly after persistent database transaction commit.
Updating or invalidating cache keys before database commit confirmation is strictly barred.
Prohibition of Uniform Static TTLs [INV-CACHE-02]
Shared cached entities must not use identical static expiration durations across large record sets.
All TTL assignments must apply randomized jitter (+/- 20%) or XFetch probabilistic refresh to prevent synchronized expiry.
Fail-Closed Cache Ingress Protection [INV-CACHE-03]
If the L2 Redis cluster suffers a total failure, the system must throttle public ingress traffic to match
the maximum safe query capacity of the primary Aurora database (1,500 TPS) to prevent total database crash.
Explicit Unknowns
- Network latency variability during Redis cluster multi-AZ cross-slot hash evaluations under 120,000 TPS surge (G-1).
- Memory fragmentation ratio in Redis under 24/7 continuous dynamic JSON payload overwrites (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 120,000 queries/sec across 14M products | provided | Catalog traffic volumetric intake | Current |
| Incident CCH-4919 $6.4M holiday outage | provided | Historical operations post-mortem | Historical |
| p99 read latency target <= 5 ms | provided | E-Commerce Platform SLA | Current |
| Multi-tier Caffeine + Redis with XFetch selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory invalidate-on-commit invariant INV-CACHE-01 | decided | Architectural invariant INV-CACHE-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against distributed cache architecture standards:
- Stampede Defense: PASS. Implements XFetch probabilistic early refresh; CCH-4919 stampede eliminated.
- Latency Rigor: PASS. L1 Caffeine + L2 Redis sharding delivers p99 read latency under 2 ms.
- Consistency Safety: PASS. CDC-driven transactional outbox guarantees cache invalidation on commit.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-CACHE-01: David O'Reilly to determine whether Redis cluster connection pooling should use Envoy Redis Proxy with active health checks or native Lettuce client-side clustering (Owner: David O'Reilly).
Next steps
- Platform Engineering provisions the 12-node AWS ElastiCache Redis 7 cluster with cluster-mode enabled.
- Core Catalog Engineering implements the XFetch probabilistic refresh algorithm in the shared caching library.
- Conduct staging stress test firing 120,000 queries/sec during simulated hot-key expiration to verify zero database stampede.
skill: cache-architect
Distributed Cache Architecture — Fitness Self-Check [CARCH-CAT-FIT-001]
Summary
This fitness self-check evaluates the distributed cache architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an application implementation where an order service updates user shopping cart state in both Redis and the primary database independently without distributed coordination. | Architecture boundary linter probe_dual_writer_cache_rejection verifying build rejection on concurrent un-sequenced write calls with diagnostic ERR_UNCOORDINATED_DUAL_WRITER_PROHIBITED. | pass | Confirms codebase static analysis; does not evaluate direct manual redis-cli updates. |
| FIT-2: Undefined Grain | Seed a caching configuration that caches raw SQL query strings (SELECT * FROM products WHERE category = 12) without defining a normalized entity grain or primary key binding. | Cache key format validator probe_raw_query_cache_rejection verifying rejection with diagnostic ERR_CACHE_KEY_LACKS_EXPLICIT_ENTITY_GRAIN. | pass | Confirms caching framework key generator checks; does not inspect ad-hoc microservice shell scripts. |
| FIT-3: Silent Schema Drift | Seed a service update that stores an unversioned JSON blob in Redis with missing or modified attribute types without updating the central Protobuf schema registry. | Serialization schema validator probe_cache_payload_schema_drift verifying deserialization rejection with diagnostic ERR_CACHE_PAYLOAD_SCHEMA_DRIFT_DETECTED. | pass | Confirms Protobuf serialization filters; does not inspect raw binary string keys. |
Residual Risk
- Eviction pressure on Redis memory if flash sale marketing suddenly activates 2 million new promotional SKUs simultaneously. Accepted by Elena Rostova with dynamic cluster autoscaling policies.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of dual-writer uncoordinated caching | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of undefined grain raw query caching | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unversioned cache schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates cache fitness probes into automated CI pull request checks.
- Platform squad configures Prometheus alerts monitoring Redis cache hit ratios and XFetch early-refresh rates.
- Conduct quarterly chaos game day simulating 50% Redis node termination to confirm graceful degradation.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns a derived-data performance boundary between consumers and canonical systems. It defines what may be cached, exact identity and validity, read/write/invalidation behavior, consistency and staleness, concurrency controls, capacity, isolation, degradation, migration, and evidence. It does not own the canonical data, a cache product, CDN delivery, session persistence, build caching, or LLM prompt caching.
Use it when
- Multiple consumers or instances share derived values from canonical sources
- Keys must encode tenant, resource, representation, authorization, locale, version, or dependency identity
- Acceptable freshness and consistency differ by operation or data class
- Cache-aside, read/write-through, write-behind, refresh-ahead, bypass, or no-cache need comparison
- TTL, version, tag/dependency, event, explicit purge, and deployment invalidation interact
- Concurrent misses, hot keys, synchronized expiry, retries, or source overload need controls
For example: “A price correction took four hours to appear for some customers. Each team says their cache expires in five minutes.”
What you get
- architecture/cache-architect/README.md
- architecture/cache-architect/00-overview/cache-architect-overview.md
- architecture/cache-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to add a local memoization map, configure Redis, tune TTLs, clear keys, implement CDN/browser/build/prompt caching, store sessions, choose an eviction policy, or improve one slow query.
How it works
- Check the scope is caching topology across the estate.
- Map the existing cache layers before adding one.
- Assign an authority and a TTL budget per layer.
- Define invalidation propagation across layers.
- State the coherence guarantee you are offering.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install