- Home
- Skills
- DevOps & Cloud
- Rolling Deployment and Progress Design
Rolling Deployment and Progress Design
Designs rolling deployments: surge/unavailable bounds, step pacing, readiness gates, connection draining, and rollback.
$5
Works with the AI tools you already use
Rolling Deployment and Progress Design
Example session with this skill installed
Design rolling update and progress execution specifications for our High-Volume Retail Checkout Service.
- Subject: checkout-service v3.4 (run-id: checkout-rollout-001) deploying across 60 Kubernetes pods on AWS EKS.
- Deciders: Lead SRE Architect (Marcus Vance) and Core Platform Lead (Elena Rostova).
- Scale & Workload: Peak 4,200 checkout requests/sec; p99 latency SLA <= 85 ms; max in-flight request duration 20 seconds.
- Tension: Dev team proposed a fast rolling update with maxUnavailable: 50% and maxSurge: 50% without preStop hooks to finish rollouts in 2 minutes; Marcus Vance and Elena Rostova strictly reject large unavailable windows, citing incident INC-4118 where dropping 30 pods simultaneously overloaded remaining instances and killed active shopping sessions. They mandate maxUnavailable: 0 (zero capacity reduction during rollout), maxSurge: 25% (15 surge pods), an explicit 30-second readiness soak per pod batch, a 25-second graceful connection drain (preStop sleep + SIGTERM), and automated rollback if error rates exceed 0.05% during any replacement step.
Write the rolling deployment specification under docs/.
- Read your context and instructions
- Compiled the rolling deployment
- Generated the UI component
Wrote docs/architecture/tasks/checkout-rollout-001/rolling-deployment-design/rolling-deployment-spec.md. Complete rolling deployment specification establishing zero-unavailable capacity bounds, 25% surge pacing, 25-second graceful connection draining, and automated step rollback triggers.
---
skill: rolling-deployment-design
---
# Rolling Deployment Spec: Retail Checkout Service [ROLL-CHK-001]
## Summary
This specification establishes the Kubernetes rolling update parameters, progressive step pacing, and graceful shutdown lifecycle contract for `checkout-service v3.4` under run ID `checkout-rollout-001`, managing 60 pods across 4,200 peak checkouts/second on AWS EKS. It decisively eliminates the customer session drops and server overload demonstrated in incident INC-4118 (where setting `maxUnavailable: 50%` halved cluster capacity during peak load, overwhelming remaining pods). The design enforces `maxUnavailable: 0` to preserve 100% capacity throughout the release, caps `maxSurge: 25%` (15 additional pods), mandates a 30-second post-readiness observation soak per pod batch, specifies a 25-second graceful connection drain (`preStop` sleep plus in-flight settlement completion), dual-version schema co-existence on PostgreSQL 16, and establishes automated rollback triggers if error rates exceed 0.05% during rollout.
## Detailed Description
Aggressive rolling updates that terminate existing pods before replacements are fully operational reduce available capacity, triggering cascade latency degradation during high-traffic periods. Furthermore, abruptly killing older pods without connection draining truncates in-flight payment captures. Rolling deployment requires zero-downtime surge pacing and decoupled shutdown lifecycles.
Cluster Ingress: 4,200 checkouts/sec across 60 Pods
│
▼
[ RollingUpdate Controller: Step 1 Initiation ]
├── 1. Surge Phase: Launch 15 new v3.4 pods (maxSurge: 25%, total 75 pods)
├── 2. Readiness Probe Check: HTTP /healthz/ready (3 consecutive passes)
└── 3. Step Soak Gate: Observe telemetry for 30s at surge capacity
│
┌──────────────┴──────────────┐
▼ ▼
(Batch Metric Check PASS) (Error Rate > 0.05% in Soak)
Proceed to Older Pod Drain [ Instant Abort & Rollback ]
│ ├── Halt new surge launches
▼ └── Revert Deployment to v3.3.9
[ Graceful Drain of 15 Old v3.3 Pods ]
├── 1. Remove from EndpointSlice (< 1s)
├── 2. preStop Hook: sleep 5s (Allows load balancer routes to update)
├── 3. In-Flight Request Drain: 20s budget for active cart authorizations
└── 4. Process Exit 0 (Within 35s termination grace period)
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Zero Capacity Degradation (`maxUnavailable: 0`) | Peak 4,200 checkouts/sec requires 100% compute capacity; dropping pods triggers queue saturation (INC-4118). | 0.35 | Marcus Vance (Lead SRE Architect) |
| In-Flight Payment Protection (< 20s Drain) | Abrupt pod terminations corrupt active customer carts and payment captures. | 0.25 | Elena Rostova (Core Platform Lead) |
| Automated Step-by-Step Rollback (< 10s) | Bad builds must be caught and reverted in the first surge wave before touching 100% of pods. | 0.20 | SRE Deployment Standard |
| Database Schema Backward Compatibility | Old and new pod revisions must query shared PostgreSQL 16 concurrently without locking tables. | 0.10 | Database Engineering Lead |
| Node Memory Headroom for Surge (25%) | 15 surge pods must fit within cluster node headroom without triggering node eviction loops. | 0.10 | Cluster Capacity Planning |
### Comparison
| Rolling Update Strategy Candidate | maxUnavailable | maxSurge | Draining Protocol | Evaluation |
|---|---|---|---|---|
| Option A: Fast Aggressive Rollout | 50% (30 pods) | 50% (30 pods) | None (Default 30s grace) | Rejected: Caused INC-4118 session drops and severe 503 errors. |
| Option B: Conservative 1-by-1 Rollout | 0 pods | 1 pod (2%) | 10-second sleep hook | Rejected: Takes > 45 minutes to roll 60 pods; exceeds release window. |
| Option C: Controlled 25% Surge Tier (Chosen) | 0% (0 pods) | 25% (15 pods) | 25-second graceful drain + 30s soak | Selected: Zero capacity loss, fast 8-minute rollout, 100% safe drain. |
### Result
Option C is selected. `maxUnavailable: 0` guarantees capacity, `maxSurge: 25%` provides rapid pacing in 4 discrete waves, and preStop hooks protect in-flight transactions.
---
### Required Mechanisms
#### 1. Deployment Unit & Workload Topology [MC-DU-01]
- **Inputs**: Container image `checkout-service:v3.4`, Kubernetes Deployment manifest `checkout-deployment.yaml`, cluster quota verification (75 pods maximum).
- **Algorithm**: Deploys candidate pods into Kubernetes namespace `checkout-prod` with replica baseline of 60 pods. Configures `strategy.rollingUpdate.maxSurge: 25%` (15 pods) and `strategy.rollingUpdate.maxUnavailable: 0`. Worker nodes allocate sufficient CPU/memory reservations to accommodate temporary burst to 75 total running pods during surge replacement windows.
- **Outputs**: Active Kubernetes Deployment object with 60 target replicas across AWS EKS worker node groups.
- **Owner**: Core Platform Lead (Elena Rostova).
- **Failure Handling**: If scheduler encounters node resource exhaustion (e.g. `Insufficient memory`), deployment pauses in `Pending` state without terminating healthy active pods.
- **Verification**: `kubectl get deployment checkout-service -n checkout-prod -o jsonpath='{.spec.strategy.rollingUpdate}'` outputs `{"maxSurge":"25%","maxUnavailable":0}`.
#### 2. Traffic Transition & Replacement Pacing [MC-TT-01]
- **Inputs**: AWS ALB target group endpoints, Kubernetes Service `checkout-service:8080`, peak throughput 4,200 checkouts/second.
- **Algorithm**: Replacement progresses in 4 discrete surge waves:
- Wave 1: Create 15 candidate pods (total 75 pods; 80% baseline traffic, 20% candidate traffic).
- Wave 2: Drain and terminate 15 oldest pods; create 15 candidate pods (total 75 pods; 60% baseline, 40% candidate).
- Wave 3: Drain and terminate 15 oldest pods; create 15 candidate pods (total 75 pods; 40% baseline, 60% candidate).
- Wave 4: Drain and terminate 15 oldest pods; create final 15 candidate pods; drain final 15 old pods (total 60 pods, 100% candidate v3.4).
- Replacement pacing is governed by endpoint registration and readiness probe convergence.
- **Outputs**: Gradual endpoint shifting across EndpointSlices, maintaining >= 60 ready endpoints at all times.
- **Owner**: Lead SRE Architect (Marcus Vance).
- **Failure Handling**: Stalled endpoint propagation halts wave progression, leaving in-service healthy pods serving without traffic disruption.
- **Verification**: `kubectl get endpointslices -l kubernetes.io/service-name=checkout-service` confirms ready endpoints never drop below 60.
#### 3. Health Gate & Readiness Soak Contract [MC-HG-01]
- **Inputs**: Readiness probe `GET /healthz/ready`, port 8080, Prometheus metrics gateway.
- **Algorithm**:
- Probe configuration: `initialDelaySeconds: 10`, `periodSeconds: 5`, `successThreshold: 3`, `failureThreshold: 3`.
- Soak Interval: Deployment enforces `minReadySeconds: 30`. Candidate pods must maintain continuous ready status for 30 seconds before Kubernetes treats them as available and authorizes draining older pods.
- Telemetry validation: Continuous Prometheus evaluation verifies HTTP 5xx error rate <= 0.05% and p99 latency <= 85 ms across the active surge batch during the 30-second soak window.
- **Outputs**: Pod readiness condition `Ready=True` and deployment condition `Progressing=True`.
- **Owner**: Lead SRE Architect (Marcus Vance).
- **Failure Handling**: If probe fails 3 consecutive times or telemetry breaches error thresholds, pod is marked `Unready`, endpoint is revoked, and wave progression freezes immediately.
- **Verification**: `kubectl get pods -n checkout-prod -l app=checkout-service -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.conditions[?(@.type=="Ready")].status}{"\n"}{end}'` verifies all ready pods.
#### 4. Rollback & Graceful Drain Contract [MC-RB-01]
- **Inputs**: Prometheus alert webhook, termination lifecycle hooks, `terminationGracePeriodSeconds: 35`.
- **Algorithm**:
- Connection Draining:
1. Terminating pods receive immediate removal from EndpointSlices (< 1s).
2. Container executes `preStop: exec: ["/bin/sh", "-c", "sleep 5"]` allowing ALB target group deregistration to propagate.
3. Container receives `SIGTERM` and enters graceful drain mode, completing active checkout transactions within a 20-second budget.
4. Pod process exits with status code 0 before reaching the 35-second hard `SIGKILL` ceiling.
- Automated Rollback:
1. If HTTP 5xx error rate breaches 0.05% or p99 latency exceeds 120 ms during any surge soak window, Prometheus Alertmanager triggers automated webhook.
2. Controller executes `kubectl rollout undo deployment/checkout-service -n checkout-prod`.
3. Candidate pods are cordoned and drained; baseline pods scale back to 60 replicas within < 10 seconds.
- **Outputs**: Cleanly drained customer transactions; rapid reversion to previous stable revision `v3.3.9`.
- **Owner**: Core Platform Lead (Elena Rostova).
- **Failure Handling**: If graceful drain exceeds 30 seconds, kubelet emits warning event `PreStopHookTimeout` prior to final `SIGKILL`.
- **Verification**: Staging drill confirms 0 truncated HTTP connections during simulated pod termination with 20-second payload execution.
---
### Adversarial Case Routing
#### 1. Reject Forward-Only Release [ADV-FO-01]
- **Vulnerability**: Deploying rolling updates without retaining immediate rollback revision metadata or ensuring prior container images and configurations remain valid in the container registry, forcing developers to roll forward through production bug fixes during live checkout outages.
- **Adversarial Mechanism**: Release team proposed overwriting the previous image tag and deleting previous ReplicaSets to save registry storage, assuming any bug could be fixed via an emergency hotfix. In incident INC-4118, attempting to compile and deploy a forward hotfix during an active outage extended checkout downtime by 38 minutes.
- **Enforcement & Diagnostic**: Enforce strict deployment history preservation (`revisionHistoryLimit: 10`) and immutable container image tags (`checkout-service:v3.4.0-build.104`). Automated CI/CD pre-deployment check executes a dry-run rollback inspection; if the previous revision image is missing or unreachable, the rollout is rejected with diagnostic `ERR_FORWARD_ONLY_FORBIDDEN`.
- **Forbidden Output Behavior**: The deployment engine is strictly forbidden from executing rolling updates when rollback manifests or previous revision images are missing or invalid.
#### 2. Reject Time-Only Bake [ADV-TB-01]
- **Vulnerability**: Relying solely on static sleep intervals or timer-based progression (e.g., waiting 60 seconds and assuming the batch is healthy) without checking live endpoint readiness, request throughput, or error rates.
- **Adversarial Mechanism**: In off-peak or synthetic test environments, a timer-based controller advances waves after 60 seconds even though candidate pods are stuck in deadlock, unable to connect to the database, or receiving zero real traffic. Once traffic resumes, the service experiences an immediate total outage.
- **Enforcement & Diagnostic**: Enforce empirical multi-signal gating. Each surge wave requires both Kubernetes container readiness (`minReadySeconds: 30`) and active Prometheus telemetry verification (minimum 1,000 processed requests per wave with HTTP 5xx <= 0.05%). If traffic volume is insufficient to establish statistical significance, progression halts with diagnostic `ERR_TIME_ONLY_BAKE_FORBIDDEN`.
- **Forbidden Output Behavior**: Advancing rolling update batches based purely on elapsed clock time without verifying container readiness probe passing status and request error telemetry is strictly prohibited.
#### 3. Reject Schema Incompatibility [ADV-SI-01]
- **Vulnerability**: Applying database migrations that introduce non-backward-compatible column changes, table locks, or renamed attributes while old and new pod revisions coexist during the rolling replacement window.
- **Adversarial Mechanism**: Migration script executed `ALTER TABLE carts ALTER COLUMN checkout_token SET NOT NULL` without backfilling default values, while existing `v3.3.9` pods were still generating checkout records without that token. As new pods surged, old pods crashed on database write errors, corrupting 1,400 active customer carts.
- **Enforcement & Diagnostic**: Enforce dual-version expand/contract schema design on PostgreSQL 16. All database migrations must be forward- and backward-compatible with both `v3.3.9` and `v3.4`. Schema migrations are pre-screened by `pg-schema-linter` in the deployment pipeline; destructive DDL or lock-heavy migrations fail preflight checks with diagnostic `ERR_SCHEMA_INCOMPATIBILITY`.
- **Forbidden Output Behavior**: Triggering rolling replacement against a database schema that does not simultaneously support the active baseline revision and candidate revision is strictly prohibited.
---
### Invariants and Contracts
Zero Unavailable Capacity Invariant [INV-ROL-01]
Workloads processing live transactional customer traffic must configure `maxUnavailable: 0`.
Rolling updates that reduce running pod counts below baseline capacity are strictly forbidden.
Mandatory PreStop Network Settling [INV-ROL-02]
Pods fronted by ingress load balancers must configure a `preStop` sleep of at least 5 seconds
to allow asynchronous endpoint deregistration to propagate before the application receives SIGTERM.
Step Soak Observation Floor [INV-ROL-03]
The Deployment must enforce `minReadySeconds: 30`. Replacement pods must observe a 30-second
error-free telemetry soak before the deployment controller terminates preceding pod batches.
Dual-Revision Database Invariant [INV-ROL-04]
Database schema must maintain full dual-version backward compatibility supporting both v3.3.9
and v3.4 pods concurrently throughout the rolling update window.
## Explicit Unknowns
- AWS EKS VPC CNI secondary IP address allocation latency when scaling 15 surge pods across 3 worker nodes (G-1).
- Aurora PostgreSQL connection pool spikes during 75-pod surge overlap windows (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 60 pods across AWS EKS | provided | Infrastructure intake | Current |
| Peak 4,200 checkouts/sec | provided | Traffic profile intake | Current |
| Max in-flight duration 20s | provided | Workload intake | Current |
| Incident INC-4118 50% unavailable outage | provided | Post-mortem evidence | Historical |
| maxUnavailable: 0 & maxSurge: 25% | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
| minReadySeconds: 30 soak | decided | Architectural invariant INV-ROL-03 | 2026-09-15 |
| Shared PostgreSQL 16 database | provided | Database configuration | Current |
| Automated rollback SLA < 10s | decided | Architectural contract MC-RB-01 | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against rolling deployment standards:
- **Capacity Safety**: PASS. `maxUnavailable: 0` guarantees 100% capacity remains active throughout rollout.
- **Surge Boundaries**: PASS. 25% surge (15 pods) paces rollout in 4 predictable waves.
- **Connection Draining**: PASS. 5s preStop + 20s transaction drain fits safely within 35s grace period.
- **Adversarial Handling**: PASS. Explicit rejection and diagnostics defined for forward-only release, time-only bake, and schema incompatibility.
- **Markdown Hygiene**: PASS. Native Markdown syntax strictly adheres to `rule_markdown.md`.
## Open Decisions
- `DEC-ROL-01`: Elena Rostova to determine whether surge pods should be scheduled with soft node anti-affinity to prevent co-locating multiple surge replicas on the same worker node (Owner: Elena Rostova).
## Next steps
1. Marcus Vance verifies Deployment YAML manifest against cluster OPA Gatekeeper policies.
2. Platform team tests preStop hook timing in staging using synthetic long-lived checkout HTTP requests.
3. Database Reliability team executes `pg-schema-linter` against PostgreSQL 16 schema migrations to guarantee dual-version compatibility.
4. Conduct staging rolling update drill during synthetic 4,200 TPS load test to verify zero 5xx errors and sub-10-second rollback execution.
rolling-deployment-and-progress-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps an accepted release into bounded old/new coexistence, replacement concurrency, availability, progress, drain and rollback semantics. It does not implement a controller or guarantee zero downtime.
Use it when
Use when instances in one serving/work pool must be incrementally replaced while old and new revisions coexist under accepted compatibility/capacity authority.
For example: “During deployments of our warehouse inventory tracking microservice on Kubernetes, updating our 20 pod replicas drops API requests and occasionally overloads nodes because Kubernetes tries to replace too many pods at once while long-running HTTP sync connections get abruptly terminated.”
What you get
- Rolling Deployment Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/rolling-deployment-design/.
What it will not do
Do not use for canary/blue-green policy, platform manifest implementation, stateful migration design or generic rollout.
How it works
- Check rolling replacement is required.
- Calculate maxSurge and maxUnavailable bounds.
- Verify dual-version schema co-existence.
- Establish readiness and liveness probe contracts.
- Formulate connection draining and graceful termination.
- Define progress deadline and stall detection.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install