Infrastructure Autoscaling and Capacity Design

    1

    Designs autoscaling architectures: HPA custom metrics, Karpenter node provisioning, scale-in stabilization, and cost caps.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Infrastructure Autoscaling and Capacity Design

    Example session with this skill installed

    Design infrastructure autoscaling and capacity planning specs for our Order Processing Engine on AWS EKS.

    • Subject: order-processing-service v3.2.0 (run-id: orders-autoscale-001) operating on AWS EKS across 45 worker nodes.
    • Deciders: Lead Cloud Platform Architect (Marcus Vance) and Payment Reliability Director (Elena Rostova).
    • Scale: Baseline 450 orders/sec scaling to peak 4,800 orders/sec during flash sales; avg duration 85 ms.
    • Tension: Tech lead proposed scaling pods purely on CPU (> 70%) with 0s cooldown and generic ASGs; Marcus Vance warns CPU is a lagging metric during flash spikes, causing a 12-minute queue backup in INC-4412, while instant scale-in triggers node thrashing. He mandates custom Prometheus metric scaling (in-flight requests per pod target: 35), rapid scale-out (120 pods in < 60s), Karpenter sub-45s node provisioning with Spot fallback, a 300s scale-in stabilization window, and hard cluster cost ceilings.

    Write the autoscaling architecture specification under docs/.

    • Read your context and instructions
    • Compiled the infrastructure autoscaling
    • Generated the UI component

    Wrote docs/architecture/tasks/orders-autoscale-001/infra-autoscaling-design/autoscaling-architecture-spec.md. Complete infrastructure autoscaling specification establishing custom Prometheus HPA scaling, sub-45s Karpenter node provisioning, 300-second scale-in damping, and cluster cost ceilings.

    ---
    skill: infra-autoscaling-design
    ---
    
    # Autoscaling Architecture Spec: Core Order Processing Engine [SCALE-ORD-001]
    
    ## Summary
    
    This specification establishes the end-to-end infrastructure autoscaling, reactive workload provisioning, and capacity governance architecture for `order-processing-service v3.2.0` under run ID `orders-autoscale-001`, managing burst scaling from 450 to 4,800 orders/second on AWS EKS. It resolves the severe queue saturation and pod thrashing demonstrated in incident INC-4412 (where lagging CPU-only metrics delayed pod scaling by 12 minutes during a flash sale, followed by premature node terminations). The contract enforces custom Prometheus metric-driven Horizontal Pod Autoscaling (targeting 35 in-flight requests/pod), sub-45-second node provisioning via Karpenter using mixed On-Demand and Spot instances, a strict 300-second scale-in stabilization window to prevent flapping, and hard cluster cost caps.
    
    ## Detailed Description
    
    Relying on host CPU utilization as an autoscaling metric fails during rapid traffic bursts because transaction queues saturate long before CPU thresholds register the spike. Furthermore, unbuffered scale-in policies cause rapid pod churn, terminating active connections and evicting pods before burst spikes fully settle.
    
    
    Incoming Ingress Spike (450 -> 4,800 orders/sec in < 60s)
                             │
                             ▼
    

    [ Prometheus Custom Metric Adapter (KEDA) ]
    ├── Evaluates: sum(rate(http_requests_in_flight[1m])) / pod_count
    └── Setpoint Target: 35 concurrent requests / pod
    │
    ▼ (Target Exceeded: Trigger HPA Scale-Out)
    [ Horizontal Pod Autoscaler (HPA v2) ]
    ├── Scale-Out Policy: 100% surge every 15s (Max: 120 Pods)
    └── Pods Scheduled ──► Pending State (Triggers Karpenter)
    │
    ▼
    [ Karpenter Node Autoscaler Controller ] (sub-45s Provisioning)
    ├── Provisions: c6i.2xlarge / c6a.2xlarge (Mixed On-Demand & Spot)
    └── Direct EC2 Fleet API calls (Bypasses slow ASG reconciliation)
    │
    ▼ (Traffic Normalizes: 4,800 -> 450 req/s)
    [ Scale-In Damping Engine ]
    └── 300-second stabilization window; max 10% pod reduction per minute

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Scaling Reaction Velocity (< 60s to Max) | Flash sales require immediate pod and node capacity before client checkout timeouts trip (INC-4412). | 0.40 | Marcus Vance (Lead Cloud Architect) |
    | Workload Stabilization & Flap Prevention | Rapid cycling between scaling up and down evicts in-flight payment settlement transactions. | 0.25 | Elena Rostova (Payment Reliability) |
    | Node Provisioning Latency (< 45s) | Standard ASG launch times (3-5 minutes) leave pods trapped in Pending state during spikes. | 0.20 | Cloud Infrastructure Standard |
    | Cost Governance & Budget Ceilings | Uncontrolled autoscaling during distributed denial-of-service surges risks runaway cloud expenditures. | 0.15 | FinOps Platform Policy |
    
    
    ### Comparison
    
    | Autoscaling Architecture Candidate | Workload Metric Signal | Node Provisioning Engine | Scale-In Cooldown | Evaluation |
    |---|---|---|---|---|
    | Option A: CPU-Only + AWS ASGs (Legacy) | CPU Utilization > 70% | Cluster Autoscaler + ASGs | 0 seconds (Immediate) | Rejected: Lagging metric caused INC-4412 12-minute queue backup. |
    | Option B: Scheduled Static Capacity | Time-based Cron scaling | Pre-warmed static EC2 fleet | None | Rejected: Fails unexpected surges; wastes 65% idle compute budget. |
    | Option C: Custom Metric HPA + Karpenter (Chosen) | In-flight HTTP requests (Target: 35) | Karpenter Just-In-Time Node Pools | 300-second stabilization window | Selected: Immediate burst reaction, sub-45s nodes, zero thrashing. |
    
    
    ### Result
    
    Option C is selected. Prometheus metric adapters signal HPA to scale pods instantly, Karpenter launches optimal compute nodes in under 45 seconds, and a 300-second scale-in window prevents churn.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Custom Metric Oracle & KEDA Adapter [MC-MO-01]
    - **Target Metric**: `http_requests_in_flight` exported by Envoy ingress sidecars.
    - **Evaluation Formula**: Desired Replicas = ceil((Current Replicas * Average In-Flight Requests) / 35).
    - **Metric Scraping Frequency**: Evaluated every 10 seconds via Prometheus Custom Metrics API (Prometheus Adapter / KEDA).
    
    #### 2. Horizontal Pod Autoscaler (HPA v2) Specification [MC-HP-01]
    ```yaml
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: order-processing-hpa
    
    

    namespace: orders-prod
    spec:
    scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: order-processing-service
    minReplicas: 15
    maxReplicas: 120
    metrics:

    
        - type: External
          external:
            metric:
              name: http_requests_in_flight
            target:
              type: AverageValue
              averageValue: 35
      behavior:
        scaleUp:
          stabilizationWindowSeconds: 0
          policies:
            - type: Percent
              value: 100
              periodSeconds: 15
            - type: Pods
              value: 30
              periodSeconds: 15
          selectPolicy: Max
        scaleDown:
          stabilizationWindowSeconds: 300
          policies:
            - type: Percent
              value: 10
              periodSeconds: 60
          selectPolicy: Min
    
    3. Karpenter Node Provisioning Contract [MC-KP-01]
    • NodePool Specification:
      • Provisioner: Karpenter v0.35+ directly interfacing with AWS EC2 Fleet APIs.
      • Instance Families: c6i.2xlarge, c6a.2xlarge, c7i.2xlarge (Compute-optimized).
      • Allocation Strategy: Price-Capacity-Optimized.
      • Capacity Type: 70% Spot instances for worker surge tiers, 30% On-Demand for baseline stability.
      • Node Warm-Up & Ready Latency: Bounded to <= 45.0 seconds from pod Pending event to NodeReady.
    4. Scale-In Damping & Flap Prevention [MC-FP-01]
    • Stabilization Window: Locked to 300 seconds (5 minutes). HPA must observe sustained lowered traffic across the entire 5-minute window before authorizing the first scale-down step.
    • Scale-Down Rate Limit: Capped at maximum 10% pod reduction per 60 seconds, ensuring database connection pools drain gracefully without socket termination spikes.
    5. Capacity Headroom & Cost Ceilings [MC-CC-01]
    • Cluster Hard Ceiling: Maximum 120 pods across maximum 30 c6i.2xlarge nodes.
    • Budget Circuit Breaker: If cluster compute cost rate breaches $1,200/day, Karpenter halts further node expansion and dispatches P1 alert AUTOSCALE_COST_CEILING_REACHED to FinOps SRE on-call.

    Invariants and Contracts

    Lagging Metric Prohibition [INV-SCL-01]
      Production order processing workloads must not scale solely on CPU or memory utilization.
      Scaling policies must incorporate direct traffic demand signals (in-flight requests or queue depth).
    
    Mandatory Five-Minute Scale-In Stabilization [INV-SCL-02]
      Workload HPA configurations must enforce a minimum `stabilizationWindowSeconds` of 300 for scale-down.
      Immediate scale-in (cooldown < 300s) is prohibited to prevent node and pod flapping.
    
    Absolute Compute Cost Ceiling [INV-SCL-03]
      The Karpenter NodePool must define explicit CPU and memory resource limits (`spec.limits.cpu: "240"`).
      Unbounded autoscaling configurations fail platform deployment admission checks.
    

    Explicit Unknowns

    • AWS EC2 Spot interruption frequency in us-east-1 during peak holiday shopping sales (G-1).
    • Prometheus scraping latency overhead when KEDA evaluates custom metrics across 120 pods simultaneously (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    450 baseline to 4,800 peak orders/secprovidedTraffic intakeCurrent
    Average order duration 85 msprovidedWorkload intakeCurrent
    Incident INC-4412 12-minute queue backupprovidedPost-mortem evidenceHistorical
    Custom metric in-flight requests (Target: 35)decidedMarcus Vance & Elena Rostova2026-09-15
    Karpenter sub-45s node provisioningdecidedCloud Infrastructure Standard2026-09-15
    300-second scale-in stabilization windowdecidedArchitectural invariant INV-SCL-022026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against autoscaling architecture standards:

    • Signal Quality: PASS. Custom metric http_requests_in_flight eliminates CPU lag during bursts.
    • Provisioning Velocity: PASS. Karpenter fleet launches compute nodes in < 45s without ASG delays.
    • Flap Prevention: PASS. 300-second scale-in stabilization window prevents pod thrashing.
    • Cost Protection: PASS. Hard ceiling of 120 pods / 240 CPUs stops runaway billing surges.

    Open Decisions

    • DEC-SCL-01: Elena Rostova to determine whether AWS Graviton (ARM64 c7g.2xlarge) should be added to Karpenter instance pool candidates (Owner: Elena Rostova).

    Next steps

    1. Marcus Vance provisions Karpenter NodePool CRD in infra/karpenter/orders-nodepool.yaml.
    2. Platform team deploys KEDA Prometheus ScaledObject targeting in-flight request metrics.
    3. Conduct staging load drill injecting synthetic burst from 450 to 4,800 TPS to verify sub-60s pod scale-up.

    infrastructure-autoscaling-and-capacity-.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define Karpenter NodePool boundaries and instance constraints.Establish consolidation and disruption safety protocols.Simulate node capacity feasibility for pending workloads.Map HPA custom metrics to infrastructure scaling logic.

    About this skill

    What it does

    This skill maps accepted unsatisfied placement demand and infrastructure capacity into bounded node/instance provisioning and removal behavior. It defines capacity states, scheduling feasibility, launch/readiness and disruption safety independently of Karpenter, Cluster Autoscaler or cloud ASGs.

    Use it when

    Use when workload demand cannot be placed on existing infrastructure and a node/instance/pool capacity loop needs exact authority.

    For example: “Our batch video encoding cluster gets stuck with hundreds of pending jobs every night. The cluster autoscaler launches tiny nodes that cannot fit the 16-core jobs, while during off-peak hours node termination forces active rendering tasks to restart without respecting PDBs.”

    What you get

    • Cluster Autoscaler Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/infra-autoscaling-design/.

    What it will not do

    Do not use for workload HPA/VPA/replicas, broad capacity planning, Kubernetes/cloud architecture, implementation or incidents.

    How it works

    1. Check infrastructure node autoscaling is required.
    2. Identify unschedulable placement demand.
    3. Simulate capacity feasibility.
    4. Establish NodePool boundaries.
    5. Enforce disruption and consolidation safety.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions