- Home
- Skills
- DevOps & Cloud
- Infrastructure Autoscaling and Capacity Design
Infrastructure Autoscaling and Capacity Design
Designs autoscaling architectures: HPA custom metrics, Karpenter node provisioning, scale-in stabilization, and cost caps.
$5
Works with the AI tools you already use
Infrastructure Autoscaling and Capacity Design
Example session with this skill installed
Design infrastructure autoscaling and capacity planning specs for our Order Processing Engine on AWS EKS.
- Subject: order-processing-service v3.2.0 (run-id: orders-autoscale-001) operating on AWS EKS across 45 worker nodes.
- Deciders: Lead Cloud Platform Architect (Marcus Vance) and Payment Reliability Director (Elena Rostova).
- Scale: Baseline 450 orders/sec scaling to peak 4,800 orders/sec during flash sales; avg duration 85 ms.
- Tension: Tech lead proposed scaling pods purely on CPU (> 70%) with 0s cooldown and generic ASGs; Marcus Vance warns CPU is a lagging metric during flash spikes, causing a 12-minute queue backup in INC-4412, while instant scale-in triggers node thrashing. He mandates custom Prometheus metric scaling (in-flight requests per pod target: 35), rapid scale-out (120 pods in < 60s), Karpenter sub-45s node provisioning with Spot fallback, a 300s scale-in stabilization window, and hard cluster cost ceilings.
Write the autoscaling architecture specification under docs/.
- Read your context and instructions
- Compiled the infrastructure autoscaling
- Generated the UI component
Wrote docs/architecture/tasks/orders-autoscale-001/infra-autoscaling-design/autoscaling-architecture-spec.md. Complete infrastructure autoscaling specification establishing custom Prometheus HPA scaling, sub-45s Karpenter node provisioning, 300-second scale-in damping, and cluster cost ceilings.
---
skill: infra-autoscaling-design
---
# Autoscaling Architecture Spec: Core Order Processing Engine [SCALE-ORD-001]
## Summary
This specification establishes the end-to-end infrastructure autoscaling, reactive workload provisioning, and capacity governance architecture for `order-processing-service v3.2.0` under run ID `orders-autoscale-001`, managing burst scaling from 450 to 4,800 orders/second on AWS EKS. It resolves the severe queue saturation and pod thrashing demonstrated in incident INC-4412 (where lagging CPU-only metrics delayed pod scaling by 12 minutes during a flash sale, followed by premature node terminations). The contract enforces custom Prometheus metric-driven Horizontal Pod Autoscaling (targeting 35 in-flight requests/pod), sub-45-second node provisioning via Karpenter using mixed On-Demand and Spot instances, a strict 300-second scale-in stabilization window to prevent flapping, and hard cluster cost caps.
## Detailed Description
Relying on host CPU utilization as an autoscaling metric fails during rapid traffic bursts because transaction queues saturate long before CPU thresholds register the spike. Furthermore, unbuffered scale-in policies cause rapid pod churn, terminating active connections and evicting pods before burst spikes fully settle.
Incoming Ingress Spike (450 -> 4,800 orders/sec in < 60s)
│
▼
[ Prometheus Custom Metric Adapter (KEDA) ]
├── Evaluates: sum(rate(http_requests_in_flight[1m])) / pod_count
└── Setpoint Target: 35 concurrent requests / pod
│
▼ (Target Exceeded: Trigger HPA Scale-Out)
[ Horizontal Pod Autoscaler (HPA v2) ]
├── Scale-Out Policy: 100% surge every 15s (Max: 120 Pods)
└── Pods Scheduled ──► Pending State (Triggers Karpenter)
│
▼
[ Karpenter Node Autoscaler Controller ] (sub-45s Provisioning)
├── Provisions: c6i.2xlarge / c6a.2xlarge (Mixed On-Demand & Spot)
└── Direct EC2 Fleet API calls (Bypasses slow ASG reconciliation)
│
▼ (Traffic Normalizes: 4,800 -> 450 req/s)
[ Scale-In Damping Engine ]
└── 300-second stabilization window; max 10% pod reduction per minute
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Scaling Reaction Velocity (< 60s to Max) | Flash sales require immediate pod and node capacity before client checkout timeouts trip (INC-4412). | 0.40 | Marcus Vance (Lead Cloud Architect) |
| Workload Stabilization & Flap Prevention | Rapid cycling between scaling up and down evicts in-flight payment settlement transactions. | 0.25 | Elena Rostova (Payment Reliability) |
| Node Provisioning Latency (< 45s) | Standard ASG launch times (3-5 minutes) leave pods trapped in Pending state during spikes. | 0.20 | Cloud Infrastructure Standard |
| Cost Governance & Budget Ceilings | Uncontrolled autoscaling during distributed denial-of-service surges risks runaway cloud expenditures. | 0.15 | FinOps Platform Policy |
### Comparison
| Autoscaling Architecture Candidate | Workload Metric Signal | Node Provisioning Engine | Scale-In Cooldown | Evaluation |
|---|---|---|---|---|
| Option A: CPU-Only + AWS ASGs (Legacy) | CPU Utilization > 70% | Cluster Autoscaler + ASGs | 0 seconds (Immediate) | Rejected: Lagging metric caused INC-4412 12-minute queue backup. |
| Option B: Scheduled Static Capacity | Time-based Cron scaling | Pre-warmed static EC2 fleet | None | Rejected: Fails unexpected surges; wastes 65% idle compute budget. |
| Option C: Custom Metric HPA + Karpenter (Chosen) | In-flight HTTP requests (Target: 35) | Karpenter Just-In-Time Node Pools | 300-second stabilization window | Selected: Immediate burst reaction, sub-45s nodes, zero thrashing. |
### Result
Option C is selected. Prometheus metric adapters signal HPA to scale pods instantly, Karpenter launches optimal compute nodes in under 45 seconds, and a 300-second scale-in window prevents churn.
---
### Required Mechanisms
#### 1. Custom Metric Oracle & KEDA Adapter [MC-MO-01]
- **Target Metric**: `http_requests_in_flight` exported by Envoy ingress sidecars.
- **Evaluation Formula**: Desired Replicas = ceil((Current Replicas * Average In-Flight Requests) / 35).
- **Metric Scraping Frequency**: Evaluated every 10 seconds via Prometheus Custom Metrics API (Prometheus Adapter / KEDA).
#### 2. Horizontal Pod Autoscaler (HPA v2) Specification [MC-HP-01]
```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: order-processing-hpa
namespace: orders-prod
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: order-processing-service
minReplicas: 15
maxReplicas: 120
metrics:
- type: External
external:
metric:
name: http_requests_in_flight
target:
type: AverageValue
averageValue: 35
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 30
periodSeconds: 15
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60
selectPolicy: Min
3. Karpenter Node Provisioning Contract [MC-KP-01]
- NodePool Specification:
- Provisioner: Karpenter v0.35+ directly interfacing with AWS EC2 Fleet APIs.
- Instance Families:
c6i.2xlarge,c6a.2xlarge,c7i.2xlarge(Compute-optimized). - Allocation Strategy:
Price-Capacity-Optimized. - Capacity Type: 70% Spot instances for worker surge tiers, 30% On-Demand for baseline stability.
- Node Warm-Up & Ready Latency: Bounded to <= 45.0 seconds from pod
Pendingevent toNodeReady.
4. Scale-In Damping & Flap Prevention [MC-FP-01]
- Stabilization Window: Locked to 300 seconds (5 minutes). HPA must observe sustained lowered traffic across the entire 5-minute window before authorizing the first scale-down step.
- Scale-Down Rate Limit: Capped at maximum 10% pod reduction per 60 seconds, ensuring database connection pools drain gracefully without socket termination spikes.
5. Capacity Headroom & Cost Ceilings [MC-CC-01]
- Cluster Hard Ceiling: Maximum 120 pods across maximum 30
c6i.2xlargenodes. - Budget Circuit Breaker: If cluster compute cost rate breaches $1,200/day, Karpenter halts further node expansion and dispatches P1 alert
AUTOSCALE_COST_CEILING_REACHEDto FinOps SRE on-call.
Invariants and Contracts
Lagging Metric Prohibition [INV-SCL-01]
Production order processing workloads must not scale solely on CPU or memory utilization.
Scaling policies must incorporate direct traffic demand signals (in-flight requests or queue depth).
Mandatory Five-Minute Scale-In Stabilization [INV-SCL-02]
Workload HPA configurations must enforce a minimum `stabilizationWindowSeconds` of 300 for scale-down.
Immediate scale-in (cooldown < 300s) is prohibited to prevent node and pod flapping.
Absolute Compute Cost Ceiling [INV-SCL-03]
The Karpenter NodePool must define explicit CPU and memory resource limits (`spec.limits.cpu: "240"`).
Unbounded autoscaling configurations fail platform deployment admission checks.
Explicit Unknowns
- AWS EC2 Spot interruption frequency in
us-east-1during peak holiday shopping sales (G-1). - Prometheus scraping latency overhead when KEDA evaluates custom metrics across 120 pods simultaneously (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 450 baseline to 4,800 peak orders/sec | provided | Traffic intake | Current |
| Average order duration 85 ms | provided | Workload intake | Current |
| Incident INC-4412 12-minute queue backup | provided | Post-mortem evidence | Historical |
| Custom metric in-flight requests (Target: 35) | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
| Karpenter sub-45s node provisioning | decided | Cloud Infrastructure Standard | 2026-09-15 |
| 300-second scale-in stabilization window | decided | Architectural invariant INV-SCL-02 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against autoscaling architecture standards:
- Signal Quality: PASS. Custom metric
http_requests_in_flighteliminates CPU lag during bursts. - Provisioning Velocity: PASS. Karpenter fleet launches compute nodes in < 45s without ASG delays.
- Flap Prevention: PASS. 300-second scale-in stabilization window prevents pod thrashing.
- Cost Protection: PASS. Hard ceiling of 120 pods / 240 CPUs stops runaway billing surges.
Open Decisions
DEC-SCL-01: Elena Rostova to determine whether AWS Graviton (ARM64c7g.2xlarge) should be added to Karpenter instance pool candidates (Owner: Elena Rostova).
Next steps
- Marcus Vance provisions Karpenter NodePool CRD in
infra/karpenter/orders-nodepool.yaml. - Platform team deploys KEDA Prometheus ScaledObject targeting in-flight request metrics.
- Conduct staging load drill injecting synthetic burst from 450 to 4,800 TPS to verify sub-60s pod scale-up.
infrastructure-autoscaling-and-capacity-.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted unsatisfied placement demand and infrastructure capacity into bounded node/instance provisioning and removal behavior. It defines capacity states, scheduling feasibility, launch/readiness and disruption safety independently of Karpenter, Cluster Autoscaler or cloud ASGs.
Use it when
Use when workload demand cannot be placed on existing infrastructure and a node/instance/pool capacity loop needs exact authority.
For example: “Our batch video encoding cluster gets stuck with hundreds of pending jobs every night. The cluster autoscaler launches tiny nodes that cannot fit the 16-core jobs, while during off-peak hours node termination forces active rendering tasks to restart without respecting PDBs.”
What you get
- Cluster Autoscaler Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/infra-autoscaling-design/.
What it will not do
Do not use for workload HPA/VPA/replicas, broad capacity planning, Kubernetes/cloud architecture, implementation or incidents.
How it works
- Check infrastructure node autoscaling is required.
- Identify unschedulable placement demand.
- Simulate capacity feasibility.
- Establish NodePool boundaries.
- Enforce disruption and consolidation safety.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install