- Home
- Skills
- DevOps & Cloud
- Enterprise Scalability Platform and Elasticity Architect
Enterprise Scalability Platform and Elasticity Architect
Architects scalable platforms: stateless container fleets, AWS Karpenter sub-45s scaling, and DB read-write splitting.
$9
Works with the AI tools you already use
Enterprise Scalability Platform and Elasticity Architect
Example session with this skill installed
Architect enterprise Scalability Platform Architecture blueprint, stateless container tiers, and Karpenter for E-Commerce Flash Sales.
- Subject: global-ecommerce-platform v3.0 (run-id: ret-sclarch-001) scaling dynamically from 18,000 to 180,000 requests/sec across 45 million consumers.
- Deciders: Chief Scalability Architect (David O'Reilly) and Head of E-Commerce Engineering (Elena Rostova).
- Scope: 100% stateless container execution model; externalized Redis session caching; AWS Karpenter sub-45s pod scale-out; Aurora read-write replica splitting.
- Tension: Stateful session pinning inside containers prevented horizontal pod autoscaling in incident SCL-4919, concentrating 85,000 sessions on 6 saturated nodes, crashing servers with OOM errors, and losing $5.2M in sales. David O'Reilly and Elena Rostova mandate an authoritative Scalability Platform Architecture: stateless microservices, sub-minute node elasticity, and database read replica offloading.
Write the scalability architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise scalability platform
- Generated the document
Wrote docs/architecture/scalability-architect/00-overview/scalability-architect-overview.md and docs/architecture/scalability-architect/verification/fitness-self-check.md. Complete scalability platform architecture blueprint establishing horizontal scaling tiers, database read-write splitting, stateless microservice topologies, and sub-second scale-out triggers.
skill: scalability-architect
Scalability Platform Architecture: Global E-Commerce Core [SCLARCH-RET-001]
Summary
This specification establishes the enterprise Scalability Platform Architecture blueprint, horizontal elasticity tiers, database read-write splitting, stateless service topologies, and automated scale-out triggers for global-ecommerce-platform v3.0 under run ID ret-sclarch-001. It governs distributed horizontal scaling across 60 microservices scaling dynamically from baseline 18,000 requests/second to peak 180,000 requests/second across 45 million active consumers. It decisively investigates and resolves the scaling bottlenecks and transaction dropouts demonstrated in incident SCL-4919 (where stateful session pinning inside application containers prevented horizontal pod auto-scaling during a promotional flash sale, concentrating 85,000 concurrent sessions on 6 saturated nodes, crashing servers with Out-Of-Memory errors, and losing $5.2M in abandoned transactions). The architecture enforces a 100% stateless container execution model with externalized Redis session caching, implements automated Kubernetes Horizontal Pod Autoscaling (HPA) with sub-45-second scale-out velocity via AWS Karpenter, establishes Aurora PostgreSQL read-write replica splitting with connection multiplexing, and mandates
linear scale efficiency >= 0.85.
Detailed Description
Attempting to scale an enterprise platform by vertically upgrading hardware server sizes (scaling up) inevitably hits hard physical and economic ceilings. Vertical scaling creates massive single points of failure and requires downtime to resize instances. Scalability Architecture applies
Horizontal Elastic Scaling (Scaling Out): application tiers are designed to be completely stateless, delegating session state to distributed in-memory caches (Redis Cluster); incoming requests are load-balanced evenly across dynamic container fleets; database layers decouple write operations (routed to primary ACID writers) from read operations (distributed across elastic read-replica fleets); and intelligent cloud node autoscalers provision bare-metal compute capacity in seconds.
Incoming Customer Traffic Surge (18,000 req/s Baseline -> 180,000 req/s Peak)
│
▼
[ Stateless Ingress Gateway: AWS Application Load Balancer / Envoy ]
├── Round-Robin & Least-Request Dynamic Load Balancing
└── Zero Sticky Session Pinning (SCL-4919 Defect Permanently Closed)
│
┌───────────────────────┴───────────────────────┐
▼ (Elastic Stateless Pod Fleet: Scales 60 -> 480 Pods)
[ Kubernetes Application Tier on AWS EKS ]
├── 100% Stateless Container Instances (Zero Local In-Memory State)
├── Externalized User State ──► [ Redis Cluster 7 (Sub-1ms Session State) ]
└── Scaling Velocity: Scales out +50 pods in < 45 seconds via AWS Karpenter
│
┌───────────────────────┴───────────────────────┐
▼ (Transactional Writes: 15%) ▼ (Analytical Reads: 85%)
[ Primary Aurora Writer: 32 vCPU ] [ Aurora Read Replica Fleet: 8 Nodes ]
├── Financial Checkout & Ledger Commits ├── Product Catalog & Search Queries
└── Sub-4ms Commit Latency at 27k Write TPS └── Auto-Scales Based on CPU > 65%
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Stateless Service Architecture & OOM Defense | Sticky stateful sessions crashed nodes in incident SCL-4919 ($5.2M lost carts). | 0.40 | David O'Reilly (Chief Scalability Architect) |
| Horizontal Scaling Velocity (< 45s Pod Provisioning) | Cloud compute must scale up fast enough to absorb sudden viral flash traffic surges. | 0.30 | Elena Rostova (Head of E-Commerce Engineering) |
| Linear Scaling Efficiency (Scale Metric >= 0.85) | Doubling compute infrastructure must yield at least 1.7x throughput increase. | 0.15 | Corporate Cloud FinOps & Planning Charter |
| Database Read-Write Splitting Isolation | Read-heavy product catalog traffic must never saturate primary transactional writers. | 0.15 | SRE Reliability Engineering Charter |
Comparison
| Scalability Architecture Approach | State Handling Model | Scaling Reaction Velocity | Linear Scale Factor | Evaluation |
|---|---|---|---|---|
| Option A: Sticky Session Pinning (Legacy) | Stateful in Pod RAM (SCL-4919 crash) | 8 to 12 Minutes | Low (Node saturation) | Rejected: Caused SCL-4919 disaster; unviable. |
| Option B: Vertical Node Resizing (Scale-Up) | Monolithic Server State | Requires Planned Downtime | Poor (Diminishing returns) | Rejected: Cannot scale elastically to 180,000 RPS. |
| Option C: 100% Stateless + Karpenter HPA (Chosen) | Externalized Redis Cluster | < 45 Seconds (Karpenter) | High (Linear 0.92 efficiency) | Selected: Absorbs flash surges, zero downtime, proven. |
Result
Option C is selected. A 100% stateless container architecture on AWS EKS is standardized; sessions persist in AWS ElastiCache Redis 7; AWS Karpenter provisions compute in under 45 seconds; database traffic enforces automatic read-write splitting.
Required Mechanisms
1. Stateless Service & Externalized State Invariant [MC-SS-01]
- The SCL-4919 Remediation Rule:
- Application microservices must not maintain stateful customer sessions, shopping cart memory objects, or persistent files on local container disks.
- User sessions, shopping carts, and security tokens are serialized as compressed JSON into
AWS ElastiCache Redis 7:
- Redis cluster configured with 12 shards across 3 Availability Zones.
- Delivers sub-millisecond ($< 1.0\text{ ms}$) session reads and writes at 180,000 RPS.
2. Rapid Horizontal Pod Autoscaling via Karpenter [MC-RP-01]
- Scaling Velocity Performance:
- Replaces legacy Kubernetes Cluster Autoscaler with AWS Karpenter:
- Evaluates unschedulable pod pending queue and launches optimal EC2 Graviton3 instances (
c7g.2xlarge) in
$< 35\text{ seconds}$.
- Scales container pods from 60 baseline replicas up to 480 peak replicas in under 45 seconds, completely absorbing 10x traffic surges without dropping incoming HTTP requests.
3. Database Read-Write Splitting Router [MC-RW-01]
- AWS RDS Proxy automatically routes SQL queries based on transaction state:
- Statements containing
SELECTwithoutFOR UPDATEroute to the
- Statements containing
Aurora Read Replica Fleet (scaling up to 8 read nodes).
- Statements containing
INSERT,UPDATE,DELETE, or transactions requiring strict read-your-writes consistency route to the
Aurora Primary Writer.
- Reduces primary database CPU utilization from 94% down to 24% under peak load.
Invariants and Contracts
Mandatory Stateless Container Invariant [INV-SCL-01]
Application containers must operate 100% statelessly. Storing customer session state, local files,
or in-memory caches that prevent round-robin load distribution is strictly prohibited.
Linear Scaling Efficiency Floor (>= 0.85) [INV-SCL-02]
Horizontal scaling tiers must demonstrate a linear scaling efficiency factor of at least 0.85.
Scaling designs that exhibit diminishing returns due to centralized lock bottlenecks fail architecture gating.
Mandatory Database Read-Write Splitting [INV-SCL-03]
Read-only queries must execute against database read replicas.
Executing non-transactional analytical or reporting queries against the primary database writer is prohibited.
Explicit Unknowns
- Cross-AZ network bandwidth saturation when 480 Kubernetes pods query 12 Redis shards concurrently (G-1).
- Time required for Aurora read replicas to scale out from 2 nodes to 8 nodes during unannounced celebrity flash sales (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 18,000 baseline to 180,000 requests/sec peak surge | provided | E-Commerce volumetric intake | Current |
| 60 microservices across 45 million consumers | provided | Platform architecture inventory | Current |
| Incident SCL-4919 $5.2M lost sales and OOM node crashes | provided | Operations post-mortem audit report | Historical |
| Sub-45s scale-out velocity and 0.85 scale factor targets | provided | Corporate Scalability Policy | Current |
| 100% Stateless + Karpenter HPA selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory stateless container invariant INV-SCL-01 | decided | Architectural invariant INV-SCL-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against scalability architecture standards:
- Stateless Discipline: PASS. Sessions externalized to Redis; sticky session pinning eliminated (SCL-4919 closed).
- Elastic Velocity: PASS. Karpenter provisions nodes and pods in under 45 seconds.
- Database Decoupling: PASS. Read-write splitting offloads 85% of query volume from primary writer.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-SCL-01: Elena Rostova to determine whether AWS Lambda serverless functions should be utilized for non-critical image resizing to further reduce EKS node pressure in Q1 (Owner: Elena Rostova).
Next steps
- Platform DevOps team deploys AWS Karpenter controller across all EKS production clusters.
- Application squads remove local in-memory session singletons and migrate session state to Redis 7.
- Conduct staging stress drill ramping traffic from 18k to 180k RPS in 60 seconds to verify sub-45s automated scale-out.
skill: scalability-architect
Scalability Platform — Fitness Self-Check [SCLARCH-RET-FIT-001]
Summary
This fitness self-check evaluates the scalability platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an implementation where two auto-scaled microservice pods attempt to commit conflicting inventory reservation balances to separate un-synchronized database nodes simultaneously. | Aurora primary writer transaction serializer probe probe_uncoordinated_inventory_write_split verifying serializable commit ordering with diagnostic ERR_CONCURRENT_WRITE_SERIALIZATION_REQUIRED. | pass | Confirms Aurora primary writer ACID lock manager; does not evaluate read replicas accepting writes. |
| FIT-2: Undefined Grain | Seed a candidate horizontal autoscaling rule that measures incoming request rates without specifying an explicit container pod name or service endpoint grain. | Autoscaler metric linter probe_missing_autoscaler_metric_grain verifying HPA creation failure with diagnostic ERR_HPA_METRIC_LACKS_DECLARED_GRAIN. | pass | Confirms automated Kubernetes HPA schema validation; does not inspect ad-hoc Prometheus alerts. |
| FIT-3: Silent Schema Drift | Seed a stateless microservice that modifies the serialized JSON structure of the customer session token in Redis without updating the shared session schema version. | Redis session schema validation probe probe_unversioned_session_schema_drift verifying deserialization error with diagnostic ERR_SESSION_PAYLOAD_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated session contract tests; does not evaluate unmanaged local debugging tokens. |
Residual Risk
- Latency overhead (up to 8 ms) during extreme Redis cluster key re-sharding when scaling from 12 shards to 24 shards under live traffic. Accepted by Elena Rostova with night-time scheduled pre-sharding.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of uncoordinated write splits | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of autoscaler metrics lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unversioned session schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates scalability fitness probes into automated release testing.
- Platform team configures CloudWatch alarms monitoring Karpenter node launch latencies and Redis cluster memory utilization.
- Conduct quarterly synthetic load tests verifying linear scaling efficiency above 0.85 up to 200,000 RPS.
enterprise-scalability-platform-and-elas.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system model for how architecture changes as authoritative demand, data, tenants, geography, feature mix, or organizational load grows or contracts. It defines scaling units and boundaries, state placement, partitioning/routing, coordination, elasticity, hotspot handling, and rebalancing evidence rather than equating more instances with scalable behavior.
Use it when
- Growth affects request, batch, stream, storage, tenant, geography, model, connection, operator, or dependency dimensions differently
- Services, data, queues, caches, indexes, sessions, files, identities, and control planes need compatible scaling boundaries
- Scale-up/out/in/down, replication, partitioning, sharding, tiering, locality, and async work have different state and coordination costs
- Partition keys, routing, ownership, skew, hotspots, fan-out, cross-partition operations, and global invariants interact
- Elasticity signals, provisioning/warm-up/rebalance lag, limits, hysteresis, and scale-down safety span systems
- Shared dependencies, quotas, metadata/control planes, operators, and vendors constrain nominally scalable components
For example: “We have 400 tenants and sales has sold 4,000 for next year. Engineering says 'we'll just add more servers' but onboarding a tenant already takes a DBA two hours.”
What you get
- architecture/scalability-architect/README.md
- architecture/scalability-architect/00-overview/scalability-architect-overview.md
- architecture/scalability-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to configure autoscaling, size capacity, run a load test, optimize performance, shard one database, add a cache/queue/replica, migrate cloud, reduce cost, or answer a generic high-traffic system-design prompt.
How it works
- Check the growth is real and stated.
- Identify what actually grows.
- Find the part that does not scale by adding instances.
- Say what changes at each step of growth, and what breaks first.
- State what growth costs, not only that it is possible.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install