- Home
- Skills
- APIs & Backend
- Autoscaling Control-Loop Contract Design
Autoscaling Control-Loop Contract Design
Designs pod autoscaling control loops: custom KEDA metrics, scaling rules, stabilization windows, and flapping prevention.
$5
Works with the AI tools you already use
Autoscaling Control-Loop Contract Design
Example session with this skill installed
Design application autoscaling control-loop contract, KEDA custom metrics, and stabilization window policy for our Payment Ingress.
- Subject: merchant-payment-ingress v3.2 (run-id: banking-autoscale-001) terminating API traffic across AWS EKS.
- Deciders: Lead Site Reliability Architect (David O'Reilly) and Head of Ingress Platform (Elena Rostova).
- Workload & Fluctuations: Baseline traffic 2,500 req/sec ramping to flash surges of 24,000 req/sec within 90s; p99 latency budget <= 25 ms; min replicas 20, max 180.
- Tension: Squads proposed scaling pods based solely on default CPU utilization (target 70%) with zero stabilization cooldown windows; David O'Reilly and Elena Rostova strictly reject raw CPU autoscaling, citing flash surge incident INC-3894 where CPU lag delayed pod scale-out by 4 minutes, causing 35% packet drops, followed by flapping that crashed cluster DNS. They mandate an authoritative KEDA contract: scaling on open-loop arrival rates (http_requests_per_second target 150/pod), 300-second scale-down stabilization, rapid step-scaling (+100% per 15s), and anti-flapping hysteresis guards.
Write the autoscaling contract specification under docs/.
- Read your context and instructions
- Compiled the autoscaling control-loop contract
- Generated the UI component
Wrote docs/architecture/tasks/banking-autoscale-001/autoscaling-design/autoscaling-contract.md. Complete autoscaling control-loop contract establishing KEDA arrival-rate metrics, rapid scale-up policies, 300s scale-down stabilization windows, and anti-flapping hysteresis guards.
skill: autoscaling-design
Autoscaling Control-Loop Contract: Merchant Payment Ingress [AUTOSCALE-PAY-001]
Summary
This specification establishes the application autoscaling control-loop contract, KEDA custom metric pipelines, stabilization windows, and flapping mitigation rules for merchant-payment-ingress v3.2 under run ID banking-autoscale-001. It governs Kubernetes Horizontal Pod Autoscaler (HPA) policies across AWS EKS sustaining traffic fluctuations from 2,500 requests/second to flash surges of 24,000 requests/second within 90 seconds. It decisively resolves the scaling lag and cluster flapping demonstrated in incident INC-3894 (where relying on raw CPU utilization delayed scale-out by 4 minutes, dropping 35% of incoming payment authorizations, followed by rapid scale-down thrashing that crashed cluster CoreDNS). The contract enforces
open-loop request arrival rate scaling via KEDA, sets a target throughput of
150 requests/sec per pod, configures rapid step-scaling (+100% pods per 15 seconds), mandates a
300-second scale-down stabilization window, and restricts replica bounds between 20 (floor) and 180 (ceiling).
Detailed Description
Relying on reactive resource metrics (such as CPU or Memory utilization) for application autoscaling introduces fatal lag during flash crowd arrivals. Pod CPU utilization is a lagging indicator: by the time worker CPU usage crosses an 70% threshold, incoming TCP queues and reverse proxy buffers have already saturated, resulting in dropped connections. Custom application metrics (such as Prometheus request arrival rates or queue depths) provide immediate leading signals, triggering rapid horizontal scaling before resource starvation occurs.
Incoming Flash Surge (2,500 -> 24,000 req/sec in 90s)
│
▼
[ Ingress Gateway: Prometheus Scrape Point ]
├── Emits Leading Metric: `http_requests_per_second`
└── Scraped by Prometheus Operator every 5 seconds
│
▼
[ KEDA Controller: ScaledObject Evaluator ]
├── Compares Ingress Rate to Target: 150 req/sec/pod
└── Computes Desired Replicas: 24,000 / 150 = 160 Pods
│
┌────────────────┴────────────────┐
▼ (Scale-Up: Immediate Action) ▼ (Scale-Down: Stabilization Delay)
[ Fast Step Scaling (+100% / 15s) ] [ 300s Stabilization Window (Cooldown) ]
├── Pods scale 20 ──► 40 ──► 80 ──► 160 ├── Evaluates max recommendation over 5m
└── Pre-warmed readiness in < 18s └── Prevents flapping and DNS thrashing
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Flash Surge Reaction Velocity (< 90s) | Pods must scale from 20 to 160 before gateway buffers overflow (INC-3894). | 0.40 | David O'Reilly (Lead SRE Architect) |
| Elimination of Cluster Flapping (Thrashing) | Rapid scale-up/down oscillations exhaust node resources and crash cluster DNS. | 0.30 | Elena Rostova (Head of Ingress Platform) |
| Latency SLA Compliance (p99 <= 25 ms) | Payment authorizations must not experience queuing delays during traffic spikes. | 0.15 | Core Merchant Banking SLA |
| Cost Efficiency & Over-Provisioning Floor | Baseline replica floor must handle nominal load without burning excess cloud budget. | 0.15 | FinOps Infrastructure Policy |
Comparison
| Autoscaling Control Model | Trigger Metric | Scale-Out Reaction Time | Flapping Mitigation | Evaluation |
|---|---|---|---|---|
| Option A: Standard CPU HPA (Legacy) | CPU Utilization (70%) | 4 to 6 minutes (Lagging) | 60s cooldown | Rejected: Caused INC-3894 35% packet drop disaster. |
| Option B: Memory Utilization HPA | Resident Memory (80%) | Slow (> 10 minutes) | None | Rejected: Java JVM memory does not shrink dynamically; triggers no scale-down. |
| Option C: KEDA Ingress Rate + Step Rules (Chosen) | Leading req/sec (150/pod) | < 45 seconds (Leading) | 300s stabilization window | Selected: Rapid flash-surge scaling, zero flapping, strict SLA. |
Result
Option C is selected. KEDA request-rate triggers scale pods preemptively; HPA behavior blocks enforce a 300-second scale-down stabilization window to eliminate thrashing.
Required Mechanisms
1. Workload Scaling Envelope & Replica Bounds [MC-SE-01]
- Minimum Replica Floor: 20 pods (guarantees baseline absorption of 3,000 req/sec without latency spikes).
Maximum Replica Ceiling:
180 pods (caps compute consumption to prevent downstream database connection exhaustion).
- Workload Density: Target concurrency of 150 requests/second per pod under nominal p99 <= 25 ms.
2. Scaling Metric Pipeline (KEDA Prometheus Scaler) [MC-MP-01]
- KEDA ScaledObject Specification:
apiVersion: keda.sh/v1alpha1 kind: ScaledObject metadata: name: payment-ingress-scaler namespace: payments-prod spec: scaleTargetRef: name: merchant-payment-ingress minReplicaCount: 20 maxReplicaCount: 180 pollingInterval: 5 cooldownPeriod: 300 triggers: - type: prometheus metadata: serverAddress: http://prometheus-k8s.monitoring.svc:9090 metricName: http_requests_per_second threshold: '150' query: sum(rate(http_requests_total{app="merchant-payment-ingress"}[1m]))
3. Stabilization Windows & HPA Step Policies [MC-SP-01]
- Scale-Up Behavior (Aggressive & Unbounded):
Result: Allows doubling replica count every 15 seconds during surges.scaleUp: stabilizationWindowSeconds: 0 policies: - type: Percent value: 100 periodSeconds: 15 - type: Pods value: 30 periodSeconds: 15 selectPolicy: Max - Scale-Down Behavior (Conservative & Stabilized):
Result: Smoothly sheds at most 10% of pods per minute after 5 minutes of quiet traffic.scaleDown: stabilizationWindowSeconds: 300 policies: - type: Percent value: 10 periodSeconds: 60
4. Anti-Flapping Hysteresis Guards [MC-AF-01]
Hysteresis Band: Scaler ignores rate variations within +- 10% of the target threshold (135 to 165 req/sec/pod) to prevent constant pod churning under micro-bursts.
Pod Readiness Gate: Pods configure readiness probes with initial delay of 5 seconds, ensuring new pods accept traffic in < 18 seconds from creation.
Invariants and Contracts
Leading Metric Autoscaling Invariant [INV-AUTOSCALE-01]
Production ingress services subject to flash surges must scale on leading application throughput metrics.
Relying exclusively on lagging CPU or Memory utilization for horizontal pod autoscaling is prohibited.
Mandatory Scale-Down Stabilization Window [INV-AUTOSCALE-02]
Horizontal Pod Autoscalers must enforce a scale-down stabilization window of at least 300 seconds (5 minutes).
Instantaneous scale-down policies that induce pod flapping are strictly barred.
Strict Maximum Replica Bounds Ceiling [INV-AUTOSCALE-03]
Every autoscaling workload must declare an immutable maximum replica ceiling (180).
Unbounded autoscalers that risk cloud quota exhaustion or downstream database failure are prohibited.
Explicit Unknowns
- AWS EKS VPC CNI IP address allocation latency when spinning up 100 pods simultaneously across 3 subnets (G-1).
- Envoy ingress upstream connection draining behavior during 10% scale-down termination steps (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 2,500 to 24,000 req/sec flash surges | provided | Traffic profile intake | Current |
| Latency SLA p99 <= 25 ms | provided | Merchant Banking SLA | Current |
| Incident INC-3894 4-minute scaling lag | provided | Historical post-mortem | Historical |
| KEDA Prometheus request-rate scaling | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| 150 req/sec per pod target threshold | decided | SRE Capacity Modeling | 2026-09-15 |
| 300-second scale-down stabilization window | decided | Architectural invariant INV-AUTOSCALE-02 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against autoscaling control-loop standards:
- Leading Indicator Safety: PASS. KEDA request arrival rate scaling replaces lagging CPU metrics.
- Reaction Speed: PASS. +100% per 15s step policy scales from 20 to 160 pods in < 60 seconds.
- Flapping Mitigation: PASS. 300s stabilization window and 10% step-down rate eliminate thrashing.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-AUTOSCALE-01: David O'Reilly to determine whether Karpenter node autoscaling should be coupled to KEDA metric events to pre-provision EC2 compute nodes ahead of pod scheduling (Owner: David O'Reilly).
Next steps
- Marcus Vance installs KEDA operator v2.13 on production AWS EKS clusters.
- Platform team applies the
ScaledObjectmanifest tomerchant-payment-ingress. - Conduct staging flash-surge drill injecting 24,000 req/sec to verify pod scale-out completes in under 90 seconds.
autoscaling-control-loop-contract-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps an accepted scalable unit and capacity model into a bounded observe-decide-act loop. It defines the signal, target, desired capacity, bounds, transition behavior and evidence independently of Kubernetes or cloud configuration.
Use it when
Use when one known workload/capacity unit needs an exact elastic scaling policy under accepted objectives and constraints.
For example: “Ticket sales open at 10:00 and the site falls over for the first four minutes. By the time the new pods are ready everyone has given up and gone to a reseller.”
What you get
- Autoscaling Configuration Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/autoscaling-design/.
What it will not do
Do not use for broad capacity planning, right-sizing, load-balancing/resilience, Kubernetes/cloud/serverless configuration, cluster provisioning, incidents or generic tuning.
How it works
- Check the load actually varies.
- Name the scalable unit and what it costs to add one.
- Choose a signal that leads demand rather than one that lags it.
- Set the bounds and the floor deliberately.
- Define stabilisation separately in each direction.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install