- Home
- Skills
- APIs & Backend
- Cache Architecture and Invalidation Design
Cache Architecture and Invalidation Design
Designs caching architectures: cache-aside patterns, TTL freshness models, event invalidations, and stampede defenses.
$5
Works with the AI tools you already use
Cache Architecture and Invalidation Design
Example session with this skill installed
Design cache population strategy, freshness model, invalidation protocol, and stampede protection for our User Profile & Entitlements Service.
- Subject: user-profile-cache-service v2.0 (run-id: banking-cache-001) serving user KYC status, permissions, and accounts across 65 services.
- Deciders: Principal Data Architect (David O'Reilly) and Head of Platform Performance (Elena Rostova).
- Workload: Sustaining peak 42,000 read queries/sec; hit ratio target >= 98.5%; latency budget p99 <= 2.5 ms; 18M active records.
- Tension: Squads proposed cache-aside with static 24h TTL and manual DB triggers; David O'Reilly and Elena Rostova strictly reject static TTLs, citing incident INC-3892 where expired keys triggered a 35,000 req/sec stampede that crashed PostgreSQL for 40 minutes. They mandate an authoritative contract: two-tier caching (Caffeine L1 + Redis Cluster L2), probabilistic early expiration (XFetch algorithm), event-driven Kafka invalidation, jittered TTLs, and a circuit breaker protecting the primary DB.
Write the caching contract specification under docs/.
- Read your context and instructions
- Compiled the cache architecture
- Generated the UI component
Wrote docs/architecture/tasks/banking-cache-001/caching-design/caching-contract-spec.md. Complete cache architecture specification establishing two-tier L1/L2 caching, XFetch probabilistic stampede defense, event-driven Kafka invalidation, and sub-2.5ms p99 read latencies.
skill: caching-design
Cache Architecture & Invalidation Spec: User Profile Service [CACHE-BANK-001]
Summary
This specification establishes the caching architecture, population strategy, freshness lifecycle, invalidation protocol, and stampede mitigation framework for user-profile-cache-service v2.0 under run ID banking-cache-001. It powers user KYC verification, permission sets, and account summaries across 65 downstream banking microservices sustaining 42,000 peak read queries/second. It decisively eliminates the catastrophic database collapse demonstrated in incident INC-3892 (where synchronized TTL expiration triggered a 35,000 req/sec cache stampede that exhausted PostgreSQL connection pools and took down consumer banking for 40 minutes). The architecture enforces a
two-tier caching topology (local in-memory Caffeine L1 + distributed Redis Cluster L2), an
optimal probabilistic early expiration algorithm (XFetch) to refresh hot keys asynchronously before expiration, event-driven invalidation via Apache Kafka, jittered TTL bounds, and sub-2.5ms p99 read latencies.
Detailed Description
Relying on naive cache-aside patterns with fixed TTLs invites thundering herd collapse. When a high-traffic cache key expires under peak load, thousands of concurrent threads simultaneously miss the cache and issue identical expensive queries to the primary database, immediately saturating database connection pools. A resilient caching architecture implements two-tier caching with probabilistic asynchronous pre-warming, ensuring hot keys never expire in the critical read path.
Client Read Request: GET /v1/users/{id}/profile (42,000 req/sec)
│
▼
[ Tier 1: Local In-Memory Caffeine L1 Cache (App Pod RAM) ]
├── Hit (Latency <= 0.05 ms) ──► Returns Cached Profile
└── Miss (Evaluates in < 0.1 ms)
│
▼
[ Tier 2: Distributed Redis Cluster L2 Cache (AWS ElastiCache) ]
├── 1. Hit (Latency <= 1.5 ms):
│ ├── Evaluates XFetch Probabilistic Pre-Warm:
│ │ `rand() * beta * delta * ln(rand()) > (TTL - Now)?`
│ │ └── If True: Spawns async background task to refresh DB
│ └── Returns L2 Value & Warms Local L1
└── 2. Hard Miss (Protected via SingleFlight Distributed Mutex):
├── Exactly ONE thread queries Aurora PostgreSQL Primary
└── Companion threads await Redis key population (< 12 ms)
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Stampede & Thundering Herd Defense (XFetch) | Hot key expiration must never hammer the primary database during traffic spikes (INC-3892). | 0.40 | David O'Reilly (Principal Data Architect) |
| Read Latency SLA Compliance (p99 <= 2.5 ms) | Profile and permission checks sit in the critical path of all 65 microservices. | 0.30 | Elena Rostova (Head of Performance) |
| Cache Hit Ratio Floor (>= 98.5%) | High hit rates insulate backend database storage from high-throughput query churn. | 0.15 | Core Banking Engineering SLA |
| Invalidation Propagation Velocity (< 500 ms) | Security role revocations or KYC status updates must reflect cluster-wide in under 500 ms. | 0.15 | Information Security Policy |
Comparison
| Caching Strategy Candidate | Stampede Protection | Invalidation Mechanism | Cache Topology | Evaluation |
|---|---|---|---|---|
| Option A: Simple Cache-Aside (Legacy) | None (Static TTL) | Manual DB triggers | Single Redis node | Rejected: Caused INC-3892 40-minute database crash; vulnerable to stampedes. |
| Option B: Write-Through with DB Triggers | Distributed lock | Synchronous DB triggers | L2 Redis only | Rejected: Adds 45ms write latency; brittle database trigger dependencies. |
| Option C: Two-Tier L1/L2 + XFetch (Chosen) | Probabilistic XFetch + SingleFlight | Event-Driven Kafka Invalidation | Caffeine L1 + Redis L2 | Selected: Sub-2.5ms speed, zero stampede spikes, sub-500ms invalidation. |
Result
Option C is selected. Caffeine L1 absorbs 85% of traffic in RAM; Redis L2 with XFetch probabilistic pre-warming guarantees zero cache stampedes; Kafka event relays ensure rapid cluster invalidation.
Required Mechanisms
1. Two-Tier Caching Topology & Population Strategy [MC-TP-01]
- Tier 1 (L1 Caffeine in JVM Memory):
- Capacity: 100,000 hottest user entities ($\sim 250\text{ MB RAM}$ per pod).
- Maximum TTL: 60 seconds.
- Tier 2 (L2 Redis 7.2 Cluster):
- 12-shard Redis Cluster on AWS ElastiCache (
cache.r7g.xlargeinstances). - Master Key Schema:
user:profile:{user_uuid}. - Serialization: Google Protocol Buffers v3 (reduces memory footprint by 65% compared to JSON).
- 12-shard Redis Cluster on AWS ElastiCache (
2. XFetch Probabilistic Stampede Protection [MC-ST-01]
- The client library implements the optimal probabilistic early expiration algorithm:
$$\text{Refresh if:} \quad - \beta \times \delta \times \ln(\text{rand}()) > (\text{TTL} - \text{currentTime})$$- $\delta$: Measured computation time to query Aurora DB and serialize profile ($\sim 18\text{ ms}$).
- $\beta$: Aggressiveness tuning constant set to 1.5.
Execution: As a key approaches expiration under heavy load, early worker threads probabilistically trigger an asynchronous background worker to re-fetch and overwrite the Redis key before it expires, eliminating read cache misses.
3. Freshness Model & Jittered TTL Bounds [MC-FM-01]
- Base TTL: 3,600 seconds (1 hour).
- Jitter Range: Uniform random jitter of $\mathbf{\pm 15%}$ ($3,060 \text{ to } 4,140\text{ seconds}$).
- Prevents batch-populated keys from expiring simultaneously.
4. Event-Driven Invalidation Contract (Kafka) [MC-IN-01]
- When user profiles or permissions mutate:
- Service commits database write and writes
UserProfileUpdatedEventto outbox. - Kafka topic
user.profile.invalidationbroadcasts message to all 65 service pods. - Consumer pods evict the key from their local Caffeine L1 cache in $< 150$ ms.
- Worker issues
DEL user:profile:{user_uuid}to Redis L2 in $< 50$ ms.
- Service commits database write and writes
- Total invalidation propagation latency: $< 200$ ms (well under 500 ms SLA).
Invariants and Contracts
Zero Unprotected Hot-Key Expiration [INV-CACHE-01]
Keys identified as critical hot paths must employ probabilistic early expiration (XFetch) or mutex locks.
Static-TTL un-jittered cache population for high-throughput keys (> 1,000 req/sec) is strictly prohibited.
Mandatory TTL Jitter Invariant [INV-CACHE-02]
All batch cache insertions must apply a pseudorandom TTL jitter of at least +- 10%.
Writing batched records with identical expiration timestamps is prohibited.
Sub-Five-Hundred-Millisecond Invalidation SLA [INV-CACHE-03]
Security-sensitive profile mutations must propagate cache eviction across all active application pods
within 500 milliseconds of database commit.
Explicit Unknowns
- Redis cluster network bandwidth saturation when 18 million profiles are bulk pre-warmed after disaster recovery failover (G-1).
- Caffeine in-memory L1 cache GC pause impact on JVM heap when handling 100,000 concurrent cached profiles (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Peak 42,000 read queries/sec | provided | Volumetric traffic profile | Current |
| Hit ratio target >= 98.5% | provided | Performance SLA requirement | Current |
| Incident INC-3892 40-minute database collapse | provided | Historical post-mortem | Historical |
| Latency budget p99 <= 2.5 ms | provided | Performance SLA contract | Current |
| Two-tier Caffeine L1 + Redis L2 topology | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| XFetch probabilistic stampede defense | decided | Architectural invariant INV-CACHE-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against caching architecture standards:
- Stampede Safety: PASS. Probabilistic XFetch pre-warms hot keys, preventing database stampedes.
- Topology Efficiency: PASS. Caffeine L1 + Redis L2 satisfies sub-2.5ms latency at 42,000 TPS.
- Freshness Control: PASS. Jittered TTLs and sub-200ms Kafka event invalidation eliminate stale data.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-CACHE-01: David O'Reilly to determine whether Redis client-side caching (RESP3 tracking) should replace Caffeine L1 to eliminate Kafka invalidation consumers in Q4 (Owner: David O'Reilly).
Next steps
- Marcus Vance provisions 12-shard Redis 7.2 ElastiCache cluster with Graviton processors.
- Platform team implements XFetch algorithm in the shared Java/Kotlin banking data client library.
- Conduct staging resilience drill expiring 5,000 simulated hot keys under 42,000 req/sec load to verify zero database connection spikes.
cache-architecture-and-invalidation-desi.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps one accepted cache opportunity into exact key/value, population, freshness, invalidation, capacity and failure semantics. It preserves the source of truth and caller-visible consistency while bounding stale and sensitive copies.
Use it when
Use when a known object/read path needs one bounded cache contract under accepted freshness and consistency requirements.
For example: “Our product page takes 4s and the database sits at 90% CPU. We added an in-process cache and now two servers show different prices for the same item during a promotion.”
What you get
- Caching Strategy Spec
- Cache Key Design Document
- Eviction Policy Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/caching-design/.
What it will not do
Do not use for system-wide cache architecture, CDN/HTTP/client/database configuration, sessions/persistence, provisioning, query optimization, cache bugs or generic tuning.
How it works
- Check the read is both hot and tolerant of staleness.
- Fix the key and everything that varies the value.
- Choose population and write semantics explicitly.
- Define invalidation by event, not by hope.
- State the behaviour when the cache is down or cold.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install