Autoscaling Control-Loop Contract Design

    1

    Designs pod autoscaling control loops: custom KEDA metrics, scaling rules, stabilization windows, and flapping prevention.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Autoscaling Control-Loop Contract Design

    Example session with this skill installed

    Design application autoscaling control-loop contract, KEDA custom metrics, and stabilization window policy for our Payment Ingress.

    • Subject: merchant-payment-ingress v3.2 (run-id: banking-autoscale-001) terminating API traffic across AWS EKS.
    • Deciders: Lead Site Reliability Architect (David O'Reilly) and Head of Ingress Platform (Elena Rostova).
    • Workload & Fluctuations: Baseline traffic 2,500 req/sec ramping to flash surges of 24,000 req/sec within 90s; p99 latency budget <= 25 ms; min replicas 20, max 180.
    • Tension: Squads proposed scaling pods based solely on default CPU utilization (target 70%) with zero stabilization cooldown windows; David O'Reilly and Elena Rostova strictly reject raw CPU autoscaling, citing flash surge incident INC-3894 where CPU lag delayed pod scale-out by 4 minutes, causing 35% packet drops, followed by flapping that crashed cluster DNS. They mandate an authoritative KEDA contract: scaling on open-loop arrival rates (http_requests_per_second target 150/pod), 300-second scale-down stabilization, rapid step-scaling (+100% per 15s), and anti-flapping hysteresis guards.

    Write the autoscaling contract specification under docs/.

    • Read your context and instructions
    • Compiled the autoscaling control-loop contract
    • Generated the UI component

    Wrote docs/architecture/tasks/banking-autoscale-001/autoscaling-design/autoscaling-contract.md. Complete autoscaling control-loop contract establishing KEDA arrival-rate metrics, rapid scale-up policies, 300s scale-down stabilization windows, and anti-flapping hysteresis guards.


    skill: autoscaling-design

    Autoscaling Control-Loop Contract: Merchant Payment Ingress [AUTOSCALE-PAY-001]

    Summary

    This specification establishes the application autoscaling control-loop contract, KEDA custom metric pipelines, stabilization windows, and flapping mitigation rules for merchant-payment-ingress v3.2 under run ID banking-autoscale-001. It governs Kubernetes Horizontal Pod Autoscaler (HPA) policies across AWS EKS sustaining traffic fluctuations from 2,500 requests/second to flash surges of 24,000 requests/second within 90 seconds. It decisively resolves the scaling lag and cluster flapping demonstrated in incident INC-3894 (where relying on raw CPU utilization delayed scale-out by 4 minutes, dropping 35% of incoming payment authorizations, followed by rapid scale-down thrashing that crashed cluster CoreDNS). The contract enforces

    open-loop request arrival rate scaling via KEDA, sets a target throughput of

    150 requests/sec per pod, configures rapid step-scaling (+100% pods per 15 seconds), mandates a

    300-second scale-down stabilization window, and restricts replica bounds between 20 (floor) and 180 (ceiling).

    Detailed Description

    Relying on reactive resource metrics (such as CPU or Memory utilization) for application autoscaling introduces fatal lag during flash crowd arrivals. Pod CPU utilization is a lagging indicator: by the time worker CPU usage crosses an 70% threshold, incoming TCP queues and reverse proxy buffers have already saturated, resulting in dropped connections. Custom application metrics (such as Prometheus request arrival rates or queue depths) provide immediate leading signals, triggering rapid horizontal scaling before resource starvation occurs.

    Incoming Flash Surge (2,500 -> 24,000 req/sec in 90s)
                             │
                             ▼
    [ Ingress Gateway: Prometheus Scrape Point ]
      ├── Emits Leading Metric: `http_requests_per_second`
      └── Scraped by Prometheus Operator every 5 seconds
                             │
                             ▼
    [ KEDA Controller: ScaledObject Evaluator ]
      ├── Compares Ingress Rate to Target: 150 req/sec/pod
      └── Computes Desired Replicas: 24,000 / 150 = 160 Pods
                             │
            ┌────────────────┴────────────────┐
            ▼ (Scale-Up: Immediate Action)    ▼ (Scale-Down: Stabilization Delay)
    [ Fast Step Scaling (+100% / 15s) ] [ 300s Stabilization Window (Cooldown) ]
      ├── Pods scale 20 ──► 40 ──► 80 ──► 160 ├── Evaluates max recommendation over 5m
      └── Pre-warmed readiness in < 18s   └── Prevents flapping and DNS thrashing
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Flash Surge Reaction Velocity (< 90s)Pods must scale from 20 to 160 before gateway buffers overflow (INC-3894).0.40David O'Reilly (Lead SRE Architect)
    Elimination of Cluster Flapping (Thrashing)Rapid scale-up/down oscillations exhaust node resources and crash cluster DNS.0.30Elena Rostova (Head of Ingress Platform)
    Latency SLA Compliance (p99 <= 25 ms)Payment authorizations must not experience queuing delays during traffic spikes.0.15Core Merchant Banking SLA
    Cost Efficiency & Over-Provisioning FloorBaseline replica floor must handle nominal load without burning excess cloud budget.0.15FinOps Infrastructure Policy

    Comparison

    Autoscaling Control ModelTrigger MetricScale-Out Reaction TimeFlapping MitigationEvaluation
    Option A: Standard CPU HPA (Legacy)CPU Utilization (70%)4 to 6 minutes (Lagging)60s cooldownRejected: Caused INC-3894 35% packet drop disaster.
    Option B: Memory Utilization HPAResident Memory (80%)Slow (> 10 minutes)NoneRejected: Java JVM memory does not shrink dynamically; triggers no scale-down.
    Option C: KEDA Ingress Rate + Step Rules (Chosen)Leading req/sec (150/pod)< 45 seconds (Leading)300s stabilization windowSelected: Rapid flash-surge scaling, zero flapping, strict SLA.

    Result

    Option C is selected. KEDA request-rate triggers scale pods preemptively; HPA behavior blocks enforce a 300-second scale-down stabilization window to eliminate thrashing.


    Required Mechanisms

    1. Workload Scaling Envelope & Replica Bounds [MC-SE-01]
    • Minimum Replica Floor: 20 pods (guarantees baseline absorption of 3,000 req/sec without latency spikes).

    Maximum Replica Ceiling:

    180 pods (caps compute consumption to prevent downstream database connection exhaustion).

    • Workload Density: Target concurrency of 150 requests/second per pod under nominal p99 <= 25 ms.
    2. Scaling Metric Pipeline (KEDA Prometheus Scaler) [MC-MP-01]
    • KEDA ScaledObject Specification:
      apiVersion: keda.sh/v1alpha1
      kind: ScaledObject
      metadata:
        name: payment-ingress-scaler
        namespace: payments-prod
      spec:
        scaleTargetRef:
          name: merchant-payment-ingress
        minReplicaCount: 20
        maxReplicaCount: 180
        pollingInterval: 5
        cooldownPeriod: 300
        triggers:
        - type: prometheus
          metadata:
            serverAddress: http://prometheus-k8s.monitoring.svc:9090
            metricName: http_requests_per_second
            threshold: '150'
            query: sum(rate(http_requests_total{app="merchant-payment-ingress"}[1m]))
      
    3. Stabilization Windows & HPA Step Policies [MC-SP-01]
    • Scale-Up Behavior (Aggressive & Unbounded):
      scaleUp:
        stabilizationWindowSeconds: 0
        policies:
        - type: Percent
          value: 100
          periodSeconds: 15
        - type: Pods
          value: 30
          periodSeconds: 15
        selectPolicy: Max
      
      Result: Allows doubling replica count every 15 seconds during surges.
    • Scale-Down Behavior (Conservative & Stabilized):
      scaleDown:
        stabilizationWindowSeconds: 300
        policies:
        - type: Percent
          value: 10
          periodSeconds: 60
      
      Result: Smoothly sheds at most 10% of pods per minute after 5 minutes of quiet traffic.
    4. Anti-Flapping Hysteresis Guards [MC-AF-01]

    Hysteresis Band: Scaler ignores rate variations within +- 10% of the target threshold (135 to 165 req/sec/pod) to prevent constant pod churning under micro-bursts.

    Pod Readiness Gate: Pods configure readiness probes with initial delay of 5 seconds, ensuring new pods accept traffic in < 18 seconds from creation.


    Invariants and Contracts

    Leading Metric Autoscaling Invariant [INV-AUTOSCALE-01]
      Production ingress services subject to flash surges must scale on leading application throughput metrics.
      Relying exclusively on lagging CPU or Memory utilization for horizontal pod autoscaling is prohibited.
    
    Mandatory Scale-Down Stabilization Window [INV-AUTOSCALE-02]
      Horizontal Pod Autoscalers must enforce a scale-down stabilization window of at least 300 seconds (5 minutes).
      Instantaneous scale-down policies that induce pod flapping are strictly barred.
    
    Strict Maximum Replica Bounds Ceiling [INV-AUTOSCALE-03]
      Every autoscaling workload must declare an immutable maximum replica ceiling (180).
      Unbounded autoscalers that risk cloud quota exhaustion or downstream database failure are prohibited.
    

    Explicit Unknowns

    • AWS EKS VPC CNI IP address allocation latency when spinning up 100 pods simultaneously across 3 subnets (G-1).
    • Envoy ingress upstream connection draining behavior during 10% scale-down termination steps (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    2,500 to 24,000 req/sec flash surgesprovidedTraffic profile intakeCurrent
    Latency SLA p99 <= 25 msprovidedMerchant Banking SLACurrent
    Incident INC-3894 4-minute scaling lagprovidedHistorical post-mortemHistorical
    KEDA Prometheus request-rate scalingdecidedDavid O'Reilly & Elena Rostova2026-09-15
    150 req/sec per pod target thresholddecidedSRE Capacity Modeling2026-09-15
    300-second scale-down stabilization windowdecidedArchitectural invariant INV-AUTOSCALE-022026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against autoscaling control-loop standards:

    • Leading Indicator Safety: PASS. KEDA request arrival rate scaling replaces lagging CPU metrics.
    • Reaction Speed: PASS. +100% per 15s step policy scales from 20 to 160 pods in < 60 seconds.
    • Flapping Mitigation: PASS. 300s stabilization window and 10% step-down rate eliminate thrashing.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-AUTOSCALE-01: David O'Reilly to determine whether Karpenter node autoscaling should be coupled to KEDA metric events to pre-provision EC2 compute nodes ahead of pod scheduling (Owner: David O'Reilly).

    Next steps

    1. Marcus Vance installs KEDA operator v2.13 on production AWS EKS clusters.
    2. Platform team applies the ScaledObject manifest to merchant-payment-ingress.
    3. Conduct staging flash-surge drill injecting 24,000 req/sec to verify pod scale-out completes in under 90 seconds.

    autoscaling-control-loop-contract-design.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define leading signals to replace lagging CPU metrics for faster scaling.Establish stabilization windows to eliminate resource flapping and waste.Calculate cold start impacts on pod readiness and scaling bounds.Document traceable scaling contracts for Kubernetes HPA or cloud ASGs.

    About this skill

    What it does

    This skill maps an accepted scalable unit and capacity model into a bounded observe-decide-act loop. It defines the signal, target, desired capacity, bounds, transition behavior and evidence independently of Kubernetes or cloud configuration.

    Use it when

    Use when one known workload/capacity unit needs an exact elastic scaling policy under accepted objectives and constraints.

    For example: “Ticket sales open at 10:00 and the site falls over for the first four minutes. By the time the new pods are ready everyone has given up and gone to a reseller.”

    What you get

    • Autoscaling Configuration Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/autoscaling-design/.

    What it will not do

    Do not use for broad capacity planning, right-sizing, load-balancing/resilience, Kubernetes/cloud/serverless configuration, cluster provisioning, incidents or generic tuning.

    How it works

    1. Check the load actually varies.
    2. Name the scalable unit and what it costs to add one.
    3. Choose a signal that leads demand rather than one that lags it.
    4. Set the bounds and the floor deliberately.
    5. Define stabilisation separately in each direction.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions