- Home
- Skills
- DevOps & Cloud
- Enterprise Deployment Platform and GitOps Architect
Enterprise Deployment Platform and GitOps Architect
Architects continuous deployment: ArgoCD GitOps, progressive canary analysis, and sub-60s automated rollbacks.
$9
Works with the AI tools you already use
Enterprise Deployment Platform and GitOps Architect
Example session with this skill installed
Architect enterprise Deployment Platform Architecture blueprint, declarative GitOps, and canaries for Continuous Delivery.
- Subject: enterprise-continuous-deployment v3.0 (run-id: corp-dplarch-001) executing 450 deployments/week across 340 microservices and 12 clusters.
- Deciders: Chief Deployment Architect (David O'Reilly) and Head of Platform Engineering (Elena Rostova).
- Scope: Declarative GitOps using ArgoCD on AWS EKS; progressive canary rollouts via Argo Rollouts; Prometheus metric analysis templates; sub-60s automated rollback.
- Tension: An un-monitored direct rolling update introduced a latent connection leak in incident DEP-4919, crashing 14 downstream services for 45 minutes, corrupting 16,000 orders, and costing $2.8M in lost sales. David O'Reilly and Elena Rostova mandate an authoritative Deployment Platform Architecture: declarative GitOps, automated canary analysis, and instant sub-30s rollbacks.
Write the deployment architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise deployment platform
- Generated the document
Wrote docs/architecture/deployment-architect/00-overview/deployment-architect-overview.md and docs/architecture/deployment-architect/verification/fitness-self-check.md. Complete deployment platform architecture blueprint establishing GitOps delivery pipelines, canary analysis, automated rollback triggers, and release boundaries.
skill: deployment-architect
Enterprise Deployment Platform Architecture: Continuous GitOps [DPLARCH-CORP-001]
Summary
This specification establishes the enterprise Deployment Platform Architecture blueprint, declarative GitOps continuous delivery pipelines, progressive canary release strategies, and automated rollback triggers for enterprise-continuous-deployment v3.0 under run ID corp-dplarch-001. It governs deployment execution across 340 microservices, 1,200 software engineers, and 12 Kubernetes production clusters executing 450 deployments/week. It decisively investigates and resolves the deployment paralysis and production outages demonstrated in incident DEP-4919 (where deploying an un-monitored direct rolling update to core checkout microservices introduced a latent database connection leak that escaped undetected for 45 minutes, crashing 14 downstream payment services, corrupting 16,000 in-flight orders, and incurring $2.8M in lost merchant transactions). The architecture enforces
declarative GitOps using ArgoCD on Kubernetes, implements automated progressive canary analysis with Prometheus metrics and Argo Rollouts, mandates sub-60-second automated rollback upon error budget degradation, and institutes
strict segregation of duties.
Detailed Description
Relying on imperative deployment scripts (kubectl apply or Jenkins shell pipelines) with manual verification leads to frequent production outages. When deployment scripts execute without automated canary analysis, defective code is pushed to 100% of production traffic simultaneously; by the time human operators notice error alarms in dashboards, thousands of customer transactions have failed. Deployment Platform Architecture establishes
Automated Progressive Delivery: all infrastructure and application states are declared in Git repositories (GitOps), deployments shift traffic incrementally (1% -> 5% -> 25% -> 100%), automated analysis monitors real-time telemetry (HTTP 5xx error rates and p99 latency), and the deployment controller automatically rolls back within seconds if anomalies are detected.
Git Push: Approved Pull Request Merged to `main` Branch
│
▼
[ Declarative GitOps Controller: ArgoCD on AWS EKS ]
├── 1. Continuously Reconciles Git Desired State with Cluster State
└── 2. Deploys Progressive Canary: Argo Rollouts Custom Controller
│
┌─────────────────┴─────────────────┐
▼ (Canary Phase 1: 5% Traffic) ▼ (Baseline Phase: 95% Traffic)
┌─────────────────┐ ┌─────────────────┐
│ Canary Pod Pool │ │ Baseline Pods │
│ - New App v3.2 │ │ - Stable v3.1 │
└────────┬────────┘ └────────┬────────┘
│ │
▼ (Real-Time Prometheus Analysis) │
[ Automated Analysis Template: 5-Minute Run ]│
├── Rule 1: HTTP 5xx Error Rate <= 0.05% │
└── Rule 2: p99 Latency <= 65 ms │
│ │
├───────────────────────────────────┤
▼ (Analysis Passes: 100% Promotion) ▼ (Analysis Fails: DEP-4919 Fix)
[ Full Automated Cutover Certified ] [ Instant Automated Rollback (< 30s) ]
└── 0 Downtime Deployment Achieved ├── Traffic Diverted to Baseline
└── Diagnostic: `ERR_CANARY_ANALYSIS_FAILED`
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Automated Canary Analysis & Instant Rollback | Undetected rollout leaks caused incident DEP-4919 ($2.8M lost merchant sales). | 0.40 | David O'Reilly (Chief Deployment Architect) |
| Declarative GitOps Single Source of Truth | Eliminates configuration drift between staging and production clusters. | 0.30 | Elena Rostova (Head of Platform Engineering) |
| Deployment Frequency & Cadence (Lead Time < 1h) | Engineering squads must ship independent releases without cross-squad coordination. | 0.15 | Core Developer Experience Charter |
| Segregation of Operational Duties | Developers cannot hold direct write access to production Kubernetes cluster APIs. | 0.15 | Corporate Information Security Mandate |
Comparison
| Deployment Platform Strategy | Failure Detection Latency | Rollback Automation | Human Error Exposure | Evaluation |
|---|---|---|---|---|
| Option A: Imperative Jenkins Shell Scripts (Legacy) | 45 Minutes (Manual alerts in DEP-4919) | Manual (kubectl rollout undo) | Extreme (Hardcoded shell scripts) | Rejected: Caused DEP-4919 disaster; unviable. |
| Option B: Blue/Green Deployment with Manual Review | 15 Minutes (Human tester sign-off) | Manual (Load balancer switch) | Moderate (Waiting on human approvals) | Rejected: Manual gates stall delivery lead times. |
| Option C: Declarative GitOps + Argo Rollouts (Chosen) | < 30 Seconds (Prometheus analysis) | Instant (< 30s Automated) | Zero (GitOps desired-state engine) | Selected: Sub-minute rollback, automated, proven. |
Result
Option C is selected. Declarative GitOps via ArgoCD combined with progressive canary analysis via Argo Rollouts is standardized across all production clusters; manual deployment scripts are barred; deployments roll back automatically upon metric breach.
Required Mechanisms
1. Progressive Canary Rollout Matrix [MC-PR-01]
| Step | Traffic Weight | Duration | Analysis Template & Verification Metric | Rollback Trigger Threshold |
|---|---|---|---|---|
| Step 1 | 5% Canary | 5 minutes | Prometheus HTTP 5xx query across canary pods | HTTP 5xx error rate $> 0.05%$ |
| Step 2 | 20% Canary | 10 minutes | p99 Latency & Connection Pool Utilization | p99 latency $> 65\text{ ms}$ OR pool sat $> 80%$ |
| Step 3 | 50% Canary | 10 minutes | Synthetic end-to-end checkout completion rate | Successful checkout rate $< 99.9%$ |
| Step 4 | 100% Full Cutover | Immediate | Final cluster health and resource consumption | Any pod restart or OOM kill |
2. Declarative Analysis Template & Rollback Oracle [MC-AO-01]
- The DEP-4919 Automated Defense Engine:
- Argo Rollout references an automated
AnalysisTemplate:apiVersion: argoproj.io/v1alpha1 kind: AnalysisTemplate metadata: name: payment-success-rate spec: metrics: - name: success-rate interval: 30s successCondition: result[0] >= 0.9995 failureLimit: 2 provider: prometheus: address: http://prometheus.monitoring:9090 query: sum(rate(http_requests_total{status!~"5.*"}[1m])) / sum(rate(http_requests_total[1m])) - If the success condition fails twice consecutively, Argo Rollouts aborts the canary, swings Envoy traffic back to the stable baseline, and scales down canary pods in
- Argo Rollout references an automated
$< 30\text{ seconds}$.
3. Strict Segregation of Duties & Access Boundaries [MC-SD-01]
- Developers possess read-only access to production Kubernetes clusters.
- All deployment modifications occur strictly via Git pull requests merged into deployment repositories.
- ArgoCD service accounts hold exclusive rights to mutate cluster state.
Invariants and Contracts
Mandatory Progressive Canary Analysis [INV-DPL-01]
Production service deployments must execute progressive canary rollouts with automated metric analysis.
Deploying direct 100% all-at-once rolling updates to Tier-1 production services is strictly prohibited.
Sub-60-Second Automated Rollback SLA [INV-DPL-02]
When canary analysis detects metric degradation, the deployment platform must revert traffic in < 60 seconds.
Requiring manual human intervention or incident calls to authorize deployment rollback is barred.
Declarative GitOps Exclusivity [INV-DPL-03]
Production cluster state must derive strictly from version-controlled Git repositories via ArgoCD.
Manual mutations via direct `kubectl` commands in production environments are physically blocked.
Explicit Unknowns
- Prometheus metrics scrape latency jitter during extreme cluster network congestion events (G-1).
- Time required for long-lived WebSocket connections to gracefully drain during canary pod termination (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 340 microservices across 12 clusters | provided | Cloud platform asset inventory | Current |
| 450 deployments/week across 1,200 engineers | provided | Engineering delivery telemetry brief | Current |
| Incident DEP-4919 $2.8M loss and 45-minute outage | provided | Operations forensic audit report | Historical |
| Sub-60-second automated rollback target | provided | Corporate SRE Reliability Policy | Current |
| Declarative GitOps + Argo Rollouts selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory progressive canary invariant INV-DPL-01 | decided | Architectural invariant INV-DPL-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against deployment architecture standards:
- Progressive Safety: PASS. 4-step canary rollout verifies real-time Prometheus error rates.
- Rollback Automation: PASS. Reverts traffic in under 30 seconds, resolving the root cause of DEP-4919.
- GitOps Integrity: PASS. Enforces declarative Git desired-state reconciliation with zero manual kubectl access.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-DPL-01: Elena Rostova to determine whether canary analysis should include automatic business metric tracking (e.g. checkout conversion dollar amounts) alongside technical latency metrics in Q1 (Owner: Elena Rostova).
Next steps
- Platform DevOps team deploys ArgoCD and Argo Rollouts controllers across all production EKS clusters.
- Ingress Platform squad configures Envoy Gateway routing integrations for dynamic canary traffic weighting.
- Conduct staging chaos game day injecting synthetic 5% HTTP 500 errors to verify automated rollback under 30 seconds.
skill: deployment-architect
Continuous GitOps Deployment Platform — Fitness Self-Check [DPLARCH-CORP-FIT-001]
Summary
This fitness self-check evaluates the enterprise deployment platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed a deployment pipeline where an imperative CI script attempts to execute kubectl apply directly while ArgoCD is simultaneously synchronizing the deployment from Git. | Kubernetes API server RBAC admission validator probe_direct_kubectl_write_rejection verifying API rejection with diagnostic ERR_DIRECT_API_MUTATION_PROHIBITED_GITOPS_ONLY. | pass | Confirms cluster RBAC permissions; does not evaluate root administrative break-glass access. |
| FIT-2: Undefined Grain | Seed a canary analysis metric definition that queries Prometheus error rates without specifying an explicit container namespace or service label filter. | Argo Rollout AnalysisTemplate linter probe_missing_metric_label_grain verifying template compilation rejection with diagnostic ERR_ANALYSIS_METRIC_LACKS_SERVICE_GRAIN. | pass | Confirms automated Helm linting and Argo schema validation; does not inspect ad-hoc Prometheus web console queries. |
| FIT-3: Silent Schema Drift | Seed a Git repository update that alters an environment variable name in the deployment manifest without updating the associated container configuration schema. | GitOps pull request validation probe probe_unannounced_deployment_schema_drift verifying PR check failure with diagnostic ERR_DEPLOYMENT_CONFIG_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated Conftest / Kubeval CI pull request gates; does not evaluate unmonitored test clusters. |
Residual Risk
- Latency overhead (up to 45 seconds) in canary promotion if Prometheus scraper experiences intermittent network packet drops. Accepted by David O'Reilly with redundant multi-endpoint Prometheus querying.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of direct kubectl imperative writes | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of analysis metrics lacking service grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unannounced deployment schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates deployment fitness probes into automated GitOps pull request checks.
- Platform team configures Prometheus alerts monitoring ArgoCD sync status and canary rollback counts.
- Conduct quarterly disaster recovery drill simulating deployment controller failure during active canary rollouts.
enterprise-deployment-platform-and-gitop.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system change-delivery model connecting immutable release units, environments, compatibility, promotion, rollout, data/infrastructure transitions, recovery, and retirement. It specifies how independently deployable changes coexist and advance safely; it does not prescribe a pipeline product or execute production changes.
Use it when
- One release spans application, API/message, schema/data, configuration, infrastructure, policy, identity, or control-plane versions
- Independently deployed producers/consumers require explicit backward/forward compatibility and mixed-version behavior
- Environment identity, equivalence, promotion lineage, configuration, secrets, data, and dependencies affect validity
- Rollout cohorts, traffic/work allocation, observation, pause, abort, completion, and cleanup must compose across systems
- Database/data migrations, backfills, dual reads/writes, feature controls, and code rollout have ordered compatibility windows
- Rollback may be impossible or unsafe after external effects, irreversible schema/state changes, or mixed-version exposure
For example: “Our checkout service deployment failed halfway through a multi-pod rollout because the database column was dropped before the new code finished starting up, causing 500 errors for 15 minutes.”
What you get
- architecture/deployment-architect/README.md
- architecture/deployment-architect/00-overview/deployment-architect-overview.md
- architecture/deployment-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/topology.md, {module}/provisioning.md, {module}/networking.md, {module}/secrets.md, {module}/cost.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to build a CI/CD pipeline, write a Kubernetes manifest, run a deploy command, configure cloud/GitOps/progressive delivery, execute a release, add a feature flag, perform one database migration, roll back an incident, or operate a platform.
How it works
- Check deployment architecture is required.
- Bound the release units.
- Establish backward and forward compatibility.
- Define environment identity and promotion lineage.
- Establish rollout exposure and health oracles.
- Define recovery and rollback state semantics.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-diagram.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install