- Home
- Skills
- DevOps & Cloud
- Canary Deployment and Analysis Design
Canary Deployment and Analysis Design
Designs canary releases: progressive traffic steps, statistical metric analysis, automated gates, and rollback tripwires.
$5
Works with the AI tools you already use
Canary Deployment and Analysis Design
Example session with this skill installed
Design the canary deployment specification for our Core Checkout Service on Kubernetes.
- Subject: checkout-service v3.4.0 (run-id: checkout-canary-001) deploying via Argo Rollouts fronted by Istio.
- Deciders: Lead SRE Architect (Marcus Vance) and Core Platform Lead (Elena Rostova).
- Traffic & Scale: Peak 3,500 checkouts/sec; baseline error rate 0.04%; baseline p95 latency 110 ms; shared PostgreSQL 16 db.
- Tension: Product team proposed a forward-only 2-step rollout (10% with a 5-minute time-only bake, then 100% cutover) without automated rollback. Marcus Vance and Elena Rostova reject forward-only and time-only rollouts, citing incident INC-4019 where schema breakage corrupted carts. They mandate: 4 progressive steps (1%, 5%, 20%, 50%), 15-minute statistical soak per step, automated Prometheus health gates (abort if HTTP 5xx > 0.2% or p95 latency > 160 ms for 3 consecutive checks), backward-compatible schema co-existence, and automated 15-second rollback tripwires restoring 100% traffic to stable v3.3.2.
- Tooling: Prometheus endpoint (http://prometheus-k8s.monitoring.svc:9090); Argo Rollouts controller.
Write the canary deployment specification under docs/.
- Read your context and instructions
- Compiled the canary deployment
- Generated the UI component
Wrote docs/architecture/tasks/checkout-canary-001/canary-design/canary-deployment-spec.md. Complete canary deployment specification establishing 4-step progressive exposure, Prometheus statistical health gates, database co-existence controls, and automated 15-second rollback tripwires.
---
skill: canary-design
---
# Canary Deployment Specification: Core Checkout Service [CANARY-CHK-001]
## Summary
This specification establishes the canary progressive exposure, statistical health gates, database schema co-existence rules, and automated rollback contract for `checkout-service v3.4.0` under run ID `checkout-canary-001`, managing production deployments across 3,500 peak checkouts/second via Argo Rollouts and Istio service mesh. It decisively resolves the cart corruption outage seen in incident INC-4019 by rejecting forward-only deployments, time-only bake intervals, and unvalidated database migrations. The contract enforces an empirical 4-stage progressive exposure ladder (1%, 5%, 20%, 50%) with a mandatory 15-minute statistical observation soak per step, continuous Prometheus metric analysis (`http://prometheus-k8s.monitoring.svc:9090`), dual-schema backward compatibility across PostgreSQL 16, and automated rollback tripwires restoring 100% traffic to stable `v3.3.2` within 15 seconds upon error or latency threshold breaches.
## Detailed Description
Abrupt binary promotions and arbitrary time-only bake pauses expose 100% of live customer transactions to latent defects, regressions, and capacity bottlenecks. In incident INC-4019, an untested forward-only rollout with breaking PostgreSQL schema alterations corrupted active customer carts, causing a 35-minute checkout outage. Progressive canary exposure boundaries isolate blast radius while validating candidate revisions against real production traffic and baseline telemetry.
Incoming Customer Ingress (3,500 checkouts/sec)
│
▼
[ Istio VirtualService Ingress Router ]
├── Stable Pods (v3.3.2): 95% Traffic (3,325 req/s)
└── Canary Pods (v3.4.0): 5% Traffic (175 req/s)
│
▼
[ Argo Rollouts Metric Analysis Gate ]
├── Provider: http://prometheus-k8s.monitoring.svc:9090
├── Metric 1: HTTP 5xx Error Ratio <= 0.002 (0.2%)
└── Metric 2: Latency p95 <= 160 ms
│
┌────────────────┴────────────────┐
▼ ▼
(Pass 5 consecutive checks) (Fail >= 3 checks)
Promote to Step 3 (20% traffic) [ Instant Rollback Trigger ]
├── Set Canary Weight = 0% (< 5s)
└── Route 100% to Stable v3.3.2
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Blast Radius Minimization | Faulty releases must affect at most 1% to 5% of checkout transactions before automated detection. | 0.35 | Marcus Vance (Lead SRE) |
| Automated Abort Reaction (< 15s) | Rollback execution must be instantaneous upon threshold breach, eliminating operator panic. | 0.25 | Platform Reliability Standard |
| Database Schema Backward Compatibility | Candidate and stable must co-exist concurrently against shared PostgreSQL 16 without locks or errors. | 0.25 | Elena Rostova (Core Platform Lead) |
| Statistical Significance & Soak Time | 15-minute observation windows ensure metric convergence and capture low-frequency database deadlocks. | 0.15 | INC-4019 Post-Mortem |
### Comparison
| Deployment Strategy | Traffic Exposure Steps | Evaluation Soak Window | Rollback Automation | Schema Co-existence |
|---|---|---|---|---|
| Option A: Forward-Only 2-Step (Product Proposal) | 10% -> 100% | 5-minute timer only | Manual operator rollback | Destructive migration; breaks stable v3.3.2 queries. |
| Option B: Rolling Replacement (MaxSurge 25%) | Pod-by-pod rolling replacement | None (Kubelet ready probe only) | Manual `kubectl rollout undo` | Assumes instant schema cutover; locks checkout table. |
| Option C: Progressive 4-Step Analysis (Chosen) | 1% -> 5% -> 20% -> 50% -> 100% | 15 minutes per step (metric-gated) | Automated Prometheus trigger (< 15s) | Phased expand/contract; backward-compatible columns. |
### Result
Option C is selected. Argo Rollouts manages progressive Istio weight shifts, executing Prometheus metrics analysis at each step before authorizing advancement.
---
### Required Mechanisms
#### 1. Deployment Unit & Workload Topology [MC-DU-01]
- **Inputs**: Container image `checkout-service:v3.4.0`, Helm chart revision 3.4.0-build.82, Argo Rollout manifest `rollout-checkout.yaml`.
- **Algorithm**: Deploys candidate pods into Kubernetes namespace `checkout-prod` with explicit labels (`role: canary`, `app: checkout-service`). Pod replica autoscaler maintains minimum 4 canary replicas (Step 1) scaling up to 40 replicas (Step 4) to ensure CPU utilization parity with stable baseline pods (80 pods, `v3.3.2`).
- **Outputs**: Active stable `Service` (`checkout-stable:8080`) and candidate `Service` (`checkout-canary:8080`).
- **Owner**: Platform Engineering (Elena Rostova).
- **Failure Handling**: If candidate pods fail readiness probes (`GET /healthz/ready`) within 180 seconds, rollout halts in `Degraded` state without traffic routing.
- **Verification**: `kubectl get pods -n checkout-prod -l role=canary` reports all pods in `Running` state and passing health probes.
#### 2. Traffic Transition & Progressive Exposure Ladder [MC-TT-01]
- **Inputs**: Istio `VirtualService` ingress routing definition, peak throughput 3,500 checkouts/second.
- **Algorithm**: Progressive 4-step weight advancement executed via Argo Rollouts traffic routing plugin for Istio:
- Step 1: Weight 1% (35 req/s), minimum soak 15 minutes, requires 5 consecutive passing checks.
- Step 2: Weight 5% (175 req/s), minimum soak 15 minutes, requires 5 consecutive passing checks.
- Step 3: Weight 20% (700 req/s), minimum soak 15 minutes, requires 5 consecutive passing checks.
- Step 4: Weight 50% (1,750 req/s), minimum soak 15 minutes, requires 5 consecutive passing checks.
- Promotion GA: Weight 100% (3,500 req/s), retire stable `v3.3.2` after 60-minute clean run.
- **Outputs**: Istio `VirtualService` weight distribution updates applied dynamically.
- **Owner**: Lead SRE Architect (Marcus Vance).
- **Failure Handling**: Stalled or rejected weight transitions pause rollout progression and alert SRE on-call.
- **Verification**: PromQL query `sum(rate(istio_requests_total{destination_workload="checkout-service-canary"}[3m])) / sum(rate(istio_requests_total{app="checkout-service"}[3m]))` matches configured step weight within ±0.2%.
#### 3. Health Gate & Telemetry Analysis Contract [MC-HG-01]
- **Inputs**: Prometheus metrics endpoint (`http://prometheus-k8s.monitoring.svc:9090`), evaluation interval 180 seconds.
- **Algorithm**: Argo Rollouts `AnalysisTemplate` continuously runs 2 PromQL queries comparing candidate against baseline:
- Metric 1: HTTP 5xx Error Rate:
```promql
sum(rate(istio_requests_total{reporter="destination", destination_workload="checkout-service-canary", response_code=~"5.."}[3m]))
/
sum(rate(istio_requests_total{reporter="destination", destination_workload="checkout-service-canary"}[3m]))
```
Threshold: <= 0.002 (0.2% error ceiling; baseline is 0.04%).
- Metric 2: p95 Request Latency:
```promql
histogram_quantile(0.95, sum(rate(istio_request_duration_milliseconds_bucket{reporter="destination", destination_workload="checkout-service-canary"}[3m])) by (le))
```
Threshold: <= 160.0 ms (baseline p95 is 110 ms; >160 ms indicates pool contention).
- **Outputs**: `AnalysisRun` status updates (`Successful`, `Failed`, or `Inconclusive`).
- **Owner**: Core Platform Lead (Elena Rostova).
- **Failure Handling**: If Prometheus returns query timeout or empty results, status is marked `Inconclusive` and rollout holds current weight without advancing.
- **Verification**: `kubectl get analysisrun -n checkout-prod` displays passing status across all completed evaluation cycles.
#### 4. Rollback & Automated Tripwires [MC-RB-01]
- **Inputs**: Consecutive metric breach count from active `AnalysisRun`.
- **Algorithm**:
- Tripwire: If either metric breaches threshold for 3 consecutive checks (9 minutes total) or if error rate exceeds 1.0% on any single check:
1. Argo Rollouts controller triggers instant abort transition.
2. Istio `VirtualService` traffic weight reset to `stable: 100%, canary: 0%` within 5 seconds.
3. Canary pod deployment scaled to 0 replicas.
4. Emits PagerDuty P1 incident notification with metric diagnostics.
5. Maximum total rollback duration: <= 15.0 seconds.
- **Outputs**: Restored 100% routing to stable `checkout-service v3.3.2`.
- **Owner**: Lead SRE Architect (Marcus Vance).
- **Failure Handling**: If Istio VirtualService modification stalls, fallback script executes direct `kubectl patch virtualservice` via CI/CD break-glass runner.
- **Verification**: Rollback staging game day verifies full traffic shedding and zero canary routing within 11.4 seconds.
---
### Adversarial Case Routing
#### 1. Reject Forward-Only Release [ADV-FO-01]
- **Vulnerability**: Shipping deployment manifests without retaining previous revision images, configuration maps, or standby capacity, forcing teams to roll forward through production bug fixes during an ongoing customer-impacting outage.
- **Adversarial Mechanism**: Product team proposed deleting `v3.3.2` deployment manifests immediately upon Step 1 launch to save cluster compute costs. During incident INC-4019, rolling forward required a 22-minute hotfix compile while cart checkouts failed.
- **Enforcement & Diagnostic**: Enforce strict rollback readiness. Stable deployment `checkout-service-stable` (`v3.3.2`) must remain pinned at full production capacity (80 replicas) throughout all canary steps. CI/CD pipeline linter checks for rollback capacity reservation; if absent, it emits diagnostic `ERR_FORWARD_ONLY_FORBIDDEN` and rejects deployment.
- **Forbidden Output Behavior**: The deployment pipeline is strictly forbidden from terminating stable baseline replicas or marking a release successful prior to Step 4 promotion completion.
#### 2. Reject Time-Only Bake [ADV-TB-01]
- **Vulnerability**: Relying solely on elapsed clock time (e.g., "wait 5 minutes then promote") without validating actual request volume, error distributions, or latency histograms against statistical thresholds.
- **Adversarial Mechanism**: In low-traffic overnight hours or during upstream DNS outages, a 5-minute timer elapses with zero checkout transactions processed. The candidate version is automatically promoted to 100%, only to crash during the morning peak load spike.
- **Enforcement & Diagnostic**: Enforce sample-size and metric-driven gating. Each 15-minute bake step requires a minimum of 5 consecutive passing Prometheus analysis cycles, with a minimum sample size of 10,000 recorded requests per evaluation window. If request count falls below threshold, Argo Rollouts pauses progression with diagnostic `ERR_INSUFFICIENT_TELEMETRY_SAMPLE`.
- **Forbidden Output Behavior**: Rollout progression based purely on static time delays without evaluating statistical PromQL queries is strictly prohibited.
#### 3. Reject Schema Incompatibility [ADV-SI-01]
- **Vulnerability**: Executing non-backward-compatible database DDL changes (dropping columns, renaming columns, adding non-null constraints without default values) ahead of canary exposure, immediately breaking queries from the active stable version.
- **Adversarial Mechanism**: Feature branch migration executed `ALTER TABLE carts DROP COLUMN legacy_cart_token`. Active `v3.3.2` pods reading that column immediately failed with SQL syntax errors, crashing 95% of live checkouts while the canary was at only 5% traffic.
- **Enforcement & Diagnostic**: Enforce expand/contract dual-schema compatibility on PostgreSQL 16. Candidate `v3.4.0` must read and write schemas in a manner fully compatible with `v3.3.2`. Migration scripts must be verified by `pg-schema-linter`; destructive column drops are blocked until `v3.3.2` is fully retired. Breaches emit diagnostic `ERR_BREAKING_SCHEMA_INCOMPATIBILITY`.
- **Forbidden Output Behavior**: Generating or applying database migrations that fail concurrent execution against both stable and candidate revisions is strictly prohibited.
---
### Invariants and Contracts
Automated Abort Invariant [INV-CAN-01]
Canary deployments must execute with automated metric evaluation active. If the error rate
exceeds 0.2% or p95 latency exceeds 160 ms for 3 cycles, the system must self-terminate
and restore 100% traffic to stable within 15 seconds without human intervention.
Progressive Exposure Boundary [INV-CAN-02]
Canary traffic must not jump directly from 0% to > 5%. The release must strictly execute
the 1% initial smoke phase for a minimum of 15 minutes.
Physical Separation of Canary Workloads [INV-CAN-03]
Canary pods must be deployed as a distinct Kubernetes Deployment resource with explicit labels
(`role: canary`) to ensure metrics isolation from stable baselines.
Dual-Version Database Compatibility [INV-CAN-04]
Database schemas must remain fully compatible with both stable v3.3.2 and canary v3.4.0
throughout all progression steps. Destructive DDL migrations are strictly forbidden.
## Explicit Unknowns
- Client-side browser cookie session stickiness impact on Istio traffic splitting distribution (G-1).
- Prometheus scraping interval lag during Kubernetes cluster node autoscaling events (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Peak 3,500 checkouts/sec | provided | Traffic intake specification | Current |
| Baseline error rate 0.04%, latency p95 110 ms | provided | Intake specification | Current |
| Shared PostgreSQL 16 database | provided | Database configuration | Current |
| Incident INC-4019 cart corruption outage | provided | Post-mortem incident record | Historical |
| 4-step progressive exposure ladder (1/5/20/50) | decided | Marcus Vance (Lead SRE) | 2026-09-15 |
| Error ceiling 0.2%, latency ceiling 160 ms | decided | Architectural decision MC-HG-01 | 2026-09-15 |
| 15-second automated rollback SLA | decided | Architectural invariant INV-CAN-01 | 2026-09-15 |
| Prometheus provider URL | provided | Internal cluster DNS | Current |
| Expand/contract schema co-existence | decided | Elena Rostova (Core Platform Lead) | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against canary deployment standards:
- **Exposure Ladder**: PASS. 4-step progressive exposure (1%, 5%, 20%, 50%) prevents abrupt traffic jumps.
- **Metric Rigor**: PASS. PromQL queries evaluate destination workload error rates and p95 latency against baselines.
- **Abort Speed**: PASS. 3-failure threshold triggers Istio 0% cutover and pod drain within 15 seconds.
- **Adversarial Checks**: PASS. Rejection mechanisms defined for forward-only rollouts, time-only bake, and schema breakage.
- **Format Integrity**: PASS. Follows native Markdown rules from `rule_markdown.md`.
## Open Decisions
- `DEC-CAN-01`: Marcus Vance to determine whether canary analysis should evaluate database connection pool saturation metrics alongside HTTP status codes (Owner: Marcus Vance).
## Next steps
1. Elena Rostova verifies Prometheus AnalysisTemplate manifest against cluster Prometheus server.
2. Marcus Vance configures Argo Rollout object and Istio VirtualService weights in staging repository.
3. Database Reliability team certifies PostgreSQL 16 schema compatibility using `pg-schema-linter`.
4. Conduct staging game day injecting 0.5% synthetic HTTP 500 errors to certify automated 15-second rollback execution.
canary-deployment-and-analysis-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps an accepted release into a bounded candidate exposure, observation and promotion/hold/abort contract. It separates desired traffic from actual cohorts and analysis evidence without implementing a rollout controller.
Use it when
Use when controlled real exposure is needed to learn release risk before broader promotion and authoritative comparison/decision evidence exists.
For example: “We upgraded our streaming video recommendation engine, but pushing it to 100% of users caused a cache stampede that crashed Redis. We want an automated Argo Rollouts canary strategy with Prometheus metrics that shifts traffic 5% -> 20% -> 50% -> 100%, checks p99 latency and error rates at each step, and automatically aborts if recommendation latency spikes.”
What you get
- Canary Deployment Spec
- Argo Rollouts YAML
- Rollback Trigger Conditions Doc
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/canary-design/.
What it will not do
Do not use for blue-green/rolling strategy, A/B product experiments, platform implementation or generic post-deploy monitoring.
How it works
- Check canary analysis is required.
- Bound exposure units and cohorts.
- Establish baseline vs candidate signal pairs.
- Formulate step progression and bake intervals.
- Configure automated analysis gates and abort triggers.
- Verify database migration co-existence.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install