Enterprise Scalability Platform and Elasticity Architect

    1

    Architects scalable platforms: stateless container fleets, AWS Karpenter sub-45s scaling, and DB read-write splitting.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Enterprise Scalability Platform and Elasticity Architect

    Example session with this skill installed

    Architect enterprise Scalability Platform Architecture blueprint, stateless container tiers, and Karpenter for E-Commerce Flash Sales.

    • Subject: global-ecommerce-platform v3.0 (run-id: ret-sclarch-001) scaling dynamically from 18,000 to 180,000 requests/sec across 45 million consumers.
    • Deciders: Chief Scalability Architect (David O'Reilly) and Head of E-Commerce Engineering (Elena Rostova).
    • Scope: 100% stateless container execution model; externalized Redis session caching; AWS Karpenter sub-45s pod scale-out; Aurora read-write replica splitting.
    • Tension: Stateful session pinning inside containers prevented horizontal pod autoscaling in incident SCL-4919, concentrating 85,000 sessions on 6 saturated nodes, crashing servers with OOM errors, and losing $5.2M in sales. David O'Reilly and Elena Rostova mandate an authoritative Scalability Platform Architecture: stateless microservices, sub-minute node elasticity, and database read replica offloading.

    Write the scalability architecture overview under docs/.

    • Read your context and instructions
    • Compiled the enterprise scalability platform
    • Generated the document

    Wrote docs/architecture/scalability-architect/00-overview/scalability-architect-overview.md and docs/architecture/scalability-architect/verification/fitness-self-check.md. Complete scalability platform architecture blueprint establishing horizontal scaling tiers, database read-write splitting, stateless microservice topologies, and sub-second scale-out triggers.


    skill: scalability-architect

    Scalability Platform Architecture: Global E-Commerce Core [SCLARCH-RET-001]

    Summary

    This specification establishes the enterprise Scalability Platform Architecture blueprint, horizontal elasticity tiers, database read-write splitting, stateless service topologies, and automated scale-out triggers for global-ecommerce-platform v3.0 under run ID ret-sclarch-001. It governs distributed horizontal scaling across 60 microservices scaling dynamically from baseline 18,000 requests/second to peak 180,000 requests/second across 45 million active consumers. It decisively investigates and resolves the scaling bottlenecks and transaction dropouts demonstrated in incident SCL-4919 (where stateful session pinning inside application containers prevented horizontal pod auto-scaling during a promotional flash sale, concentrating 85,000 concurrent sessions on 6 saturated nodes, crashing servers with Out-Of-Memory errors, and losing $5.2M in abandoned transactions). The architecture enforces a 100% stateless container execution model with externalized Redis session caching, implements automated Kubernetes Horizontal Pod Autoscaling (HPA) with sub-45-second scale-out velocity via AWS Karpenter, establishes Aurora PostgreSQL read-write replica splitting with connection multiplexing, and mandates

    linear scale efficiency >= 0.85.

    Detailed Description

    Attempting to scale an enterprise platform by vertically upgrading hardware server sizes (scaling up) inevitably hits hard physical and economic ceilings. Vertical scaling creates massive single points of failure and requires downtime to resize instances. Scalability Architecture applies

    Horizontal Elastic Scaling (Scaling Out): application tiers are designed to be completely stateless, delegating session state to distributed in-memory caches (Redis Cluster); incoming requests are load-balanced evenly across dynamic container fleets; database layers decouple write operations (routed to primary ACID writers) from read operations (distributed across elastic read-replica fleets); and intelligent cloud node autoscalers provision bare-metal compute capacity in seconds.

    Incoming Customer Traffic Surge (18,000 req/s Baseline -> 180,000 req/s Peak)
                                   │
                                   ▼
    [ Stateless Ingress Gateway: AWS Application Load Balancer / Envoy ]
      ├── Round-Robin & Least-Request Dynamic Load Balancing
      └── Zero Sticky Session Pinning (SCL-4919 Defect Permanently Closed)
                                   │
           ┌───────────────────────┴───────────────────────┐
           ▼ (Elastic Stateless Pod Fleet: Scales 60 -> 480 Pods)
    [ Kubernetes Application Tier on AWS EKS ]
      ├── 100% Stateless Container Instances (Zero Local In-Memory State)
      ├── Externalized User State ──► [ Redis Cluster 7 (Sub-1ms Session State) ]
      └── Scaling Velocity: Scales out +50 pods in < 45 seconds via AWS Karpenter
                                   │
           ┌───────────────────────┴───────────────────────┐
           ▼ (Transactional Writes: 15%)                   ▼ (Analytical Reads: 85%)
    [ Primary Aurora Writer: 32 vCPU ]            [ Aurora Read Replica Fleet: 8 Nodes ]
      ├── Financial Checkout & Ledger Commits       ├── Product Catalog & Search Queries
      └── Sub-4ms Commit Latency at 27k Write TPS   └── Auto-Scales Based on CPU > 65%
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Stateless Service Architecture & OOM DefenseSticky stateful sessions crashed nodes in incident SCL-4919 ($5.2M lost carts).0.40David O'Reilly (Chief Scalability Architect)
    Horizontal Scaling Velocity (< 45s Pod Provisioning)Cloud compute must scale up fast enough to absorb sudden viral flash traffic surges.0.30Elena Rostova (Head of E-Commerce Engineering)
    Linear Scaling Efficiency (Scale Metric >= 0.85)Doubling compute infrastructure must yield at least 1.7x throughput increase.0.15Corporate Cloud FinOps & Planning Charter
    Database Read-Write Splitting IsolationRead-heavy product catalog traffic must never saturate primary transactional writers.0.15SRE Reliability Engineering Charter

    Comparison

    Scalability Architecture ApproachState Handling ModelScaling Reaction VelocityLinear Scale FactorEvaluation
    Option A: Sticky Session Pinning (Legacy)Stateful in Pod RAM (SCL-4919 crash)8 to 12 MinutesLow (Node saturation)Rejected: Caused SCL-4919 disaster; unviable.
    Option B: Vertical Node Resizing (Scale-Up)Monolithic Server StateRequires Planned DowntimePoor (Diminishing returns)Rejected: Cannot scale elastically to 180,000 RPS.
    Option C: 100% Stateless + Karpenter HPA (Chosen)Externalized Redis Cluster< 45 Seconds (Karpenter)High (Linear 0.92 efficiency)Selected: Absorbs flash surges, zero downtime, proven.

    Result

    Option C is selected. A 100% stateless container architecture on AWS EKS is standardized; sessions persist in AWS ElastiCache Redis 7; AWS Karpenter provisions compute in under 45 seconds; database traffic enforces automatic read-write splitting.


    Required Mechanisms

    1. Stateless Service & Externalized State Invariant [MC-SS-01]
    • The SCL-4919 Remediation Rule:
      • Application microservices must not maintain stateful customer sessions, shopping cart memory objects, or persistent files on local container disks.
      • User sessions, shopping carts, and security tokens are serialized as compressed JSON into

    AWS ElastiCache Redis 7:
    - Redis cluster configured with 12 shards across 3 Availability Zones.
    - Delivers sub-millisecond ($< 1.0\text{ ms}$) session reads and writes at 180,000 RPS.

    2. Rapid Horizontal Pod Autoscaling via Karpenter [MC-RP-01]
    • Scaling Velocity Performance:
      • Replaces legacy Kubernetes Cluster Autoscaler with AWS Karpenter:
      • Evaluates unschedulable pod pending queue and launches optimal EC2 Graviton3 instances (c7g.2xlarge) in

    $< 35\text{ seconds}$.

    • Scales container pods from 60 baseline replicas up to 480 peak replicas in under 45 seconds, completely absorbing 10x traffic surges without dropping incoming HTTP requests.
    3. Database Read-Write Splitting Router [MC-RW-01]
    • AWS RDS Proxy automatically routes SQL queries based on transaction state:
      • Statements containing SELECT without FOR UPDATE route to the

    Aurora Read Replica Fleet (scaling up to 8 read nodes).

    • Statements containing INSERT, UPDATE, DELETE, or transactions requiring strict read-your-writes consistency route to the

    Aurora Primary Writer.

    • Reduces primary database CPU utilization from 94% down to 24% under peak load.

    Invariants and Contracts

    Mandatory Stateless Container Invariant [INV-SCL-01]
      Application containers must operate 100% statelessly. Storing customer session state, local files,
      or in-memory caches that prevent round-robin load distribution is strictly prohibited.
    
    Linear Scaling Efficiency Floor (>= 0.85) [INV-SCL-02]
      Horizontal scaling tiers must demonstrate a linear scaling efficiency factor of at least 0.85.
      Scaling designs that exhibit diminishing returns due to centralized lock bottlenecks fail architecture gating.
    
    Mandatory Database Read-Write Splitting [INV-SCL-03]
      Read-only queries must execute against database read replicas.
      Executing non-transactional analytical or reporting queries against the primary database writer is prohibited.
    

    Explicit Unknowns

    • Cross-AZ network bandwidth saturation when 480 Kubernetes pods query 12 Redis shards concurrently (G-1).
    • Time required for Aurora read replicas to scale out from 2 nodes to 8 nodes during unannounced celebrity flash sales (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    18,000 baseline to 180,000 requests/sec peak surgeprovidedE-Commerce volumetric intakeCurrent
    60 microservices across 45 million consumersprovidedPlatform architecture inventoryCurrent
    Incident SCL-4919 $5.2M lost sales and OOM node crashesprovidedOperations post-mortem audit reportHistorical
    Sub-45s scale-out velocity and 0.85 scale factor targetsprovidedCorporate Scalability PolicyCurrent
    100% Stateless + Karpenter HPA selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory stateless container invariant INV-SCL-01decidedArchitectural invariant INV-SCL-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against scalability architecture standards:

    • Stateless Discipline: PASS. Sessions externalized to Redis; sticky session pinning eliminated (SCL-4919 closed).
    • Elastic Velocity: PASS. Karpenter provisions nodes and pods in under 45 seconds.
    • Database Decoupling: PASS. Read-write splitting offloads 85% of query volume from primary writer.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-SCL-01: Elena Rostova to determine whether AWS Lambda serverless functions should be utilized for non-critical image resizing to further reduce EKS node pressure in Q1 (Owner: Elena Rostova).

    Next steps

    1. Platform DevOps team deploys AWS Karpenter controller across all EKS production clusters.
    2. Application squads remove local in-memory session singletons and migrate session state to Redis 7.
    3. Conduct staging stress drill ramping traffic from 18k to 180k RPS in 60 seconds to verify sub-45s automated scale-out.

    skill: scalability-architect

    Scalability Platform — Fitness Self-Check [SCLARCH-RET-FIT-001]

    Summary

    This fitness self-check evaluates the scalability platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed an implementation where two auto-scaled microservice pods attempt to commit conflicting inventory reservation balances to separate un-synchronized database nodes simultaneously.Aurora primary writer transaction serializer probe probe_uncoordinated_inventory_write_split verifying serializable commit ordering with diagnostic ERR_CONCURRENT_WRITE_SERIALIZATION_REQUIRED.passConfirms Aurora primary writer ACID lock manager; does not evaluate read replicas accepting writes.
    FIT-2: Undefined GrainSeed a candidate horizontal autoscaling rule that measures incoming request rates without specifying an explicit container pod name or service endpoint grain.Autoscaler metric linter probe_missing_autoscaler_metric_grain verifying HPA creation failure with diagnostic ERR_HPA_METRIC_LACKS_DECLARED_GRAIN.passConfirms automated Kubernetes HPA schema validation; does not inspect ad-hoc Prometheus alerts.
    FIT-3: Silent Schema DriftSeed a stateless microservice that modifies the serialized JSON structure of the customer session token in Redis without updating the shared session schema version.Redis session schema validation probe probe_unversioned_session_schema_drift verifying deserialization error with diagnostic ERR_SESSION_PAYLOAD_SCHEMA_DRIFT_DETECTED.passConfirms automated session contract tests; does not evaluate unmanaged local debugging tokens.

    Residual Risk

    • Latency overhead (up to 8 ms) during extreme Redis cluster key re-sharding when scaling from 12 shards to 24 shards under live traffic. Accepted by Elena Rostova with night-time scheduled pre-sharding.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of uncoordinated write splitsderivedFIT-1 probe result2026-09-15
    Rejection of autoscaler metrics lacking declared grainderivedFIT-2 probe result2026-09-15
    Rejection of unversioned session schema driftderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates scalability fitness probes into automated release testing.
    2. Platform team configures CloudWatch alarms monitoring Karpenter node launch latencies and Redis cluster memory utilization.
    3. Conduct quarterly synthetic load tests verifying linear scaling efficiency above 0.85 up to 200,000 RPS.

    enterprise-scalability-platform-and-elas.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define scaling units and boundaries for high-traffic fleetsDesign database partitioning and read-write splitting strategiesIdentify non-scaling bottlenecks and technical ceilingsMap elasticity signals to sub-45s scaling requirementsCalculate the economic and technical costs of system growth

    About this skill

    What it does

    This skill owns the cross-system model for how architecture changes as authoritative demand, data, tenants, geography, feature mix, or organizational load grows or contracts. It defines scaling units and boundaries, state placement, partitioning/routing, coordination, elasticity, hotspot handling, and rebalancing evidence rather than equating more instances with scalable behavior.

    Use it when

    • Growth affects request, batch, stream, storage, tenant, geography, model, connection, operator, or dependency dimensions differently
    • Services, data, queues, caches, indexes, sessions, files, identities, and control planes need compatible scaling boundaries
    • Scale-up/out/in/down, replication, partitioning, sharding, tiering, locality, and async work have different state and coordination costs
    • Partition keys, routing, ownership, skew, hotspots, fan-out, cross-partition operations, and global invariants interact
    • Elasticity signals, provisioning/warm-up/rebalance lag, limits, hysteresis, and scale-down safety span systems
    • Shared dependencies, quotas, metadata/control planes, operators, and vendors constrain nominally scalable components

    For example: “We have 400 tenants and sales has sold 4,000 for next year. Engineering says 'we'll just add more servers' but onboarding a tenant already takes a DBA two hours.”

    What you get

    • architecture/scalability-architect/README.md
    • architecture/scalability-architect/00-overview/scalability-architect-overview.md
    • architecture/scalability-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to configure autoscaling, size capacity, run a load test, optimize performance, shard one database, add a cache/queue/replica, migrate cloud, reduce cost, or answer a generic high-traffic system-design prompt.

    How it works

    1. Check the growth is real and stated.
    2. Identify what actually grows.
    3. Find the part that does not scale by adding instances.
    4. Say what changes at each step of growth, and what breaks first.
    5. State what growth costs, not only that it is possible.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions