- Home
- Skills
- DevOps & Cloud
- Capacity Planning Platform and Demand Architect
Capacity Planning Platform and Demand Architect
Architects capacity planning: M/M/c queueing models, scheduled pre-warming, and 30% headroom governance.
$9
Works with the AI tools you already use
Capacity Planning Platform and Demand Architect
Example session with this skill installed
Architect enterprise Capacity Planning Architecture blueprint, M/M/c queueing models, and headroom for E-Commerce Flash Sales.
- Subject: global-retail-ecommerce-platform v3.0 (run-id: ret-caparch-001) scaling from baseline 25,000 to peak 180,000 requests/second across 48 million shoppers.
- Deciders: Chief Capacity Planning Architect (David O'Reilly) and Head of Retail Platform Operations (Elena Rostova).
- Scope: M/M/c queueing theory demand models; scheduled pre-warming automation 30 minutes before sales; minimum 30% cluster headroom floor; synthetic peak load testing.
- Tension: Reactive CPU auto-scaling lagged behind a 400% traffic surge in incident CAP-4919 because 6-minute VM provisioning delays caused thread pool exhaustion, cascading gateway crashes, and $8.4M in lost carts. David O'Reilly and Elena Rostova mandate an authoritative Capacity Planning Architecture: predictive queueing models, pre-warmed pod pools, and strict headroom governance.
Write the capacity planning architecture overview under docs/.
- Read your context and instructions
- Compiled the capacity planning platform
- Generated the document
Wrote docs/architecture/capacity-planning-architect/00-overview/capacity-planning-architect-overview.md and docs/architecture/capacity-planning-architect/verification/fitness-self-check.md. Complete capacity planning platform architecture blueprint establishing peak demand forecasting, queueing models, cloud resource provisioning, and head-room governance.
skill: capacity-planning-architect
Capacity Planning Architecture: Global Retail E-Commerce [CAPARCH-RET-001]
Summary
This specification establishes the enterprise Capacity Planning Architecture blueprint, mathematical demand modeling, auto-scaling thresholds, and infrastructure headroom governance for global-retail-ecommerce-platform v3.0 under run ID ret-caparch-001. It governs compute, memory, database, and network capacity across 120 microservices sustaining baseline 25,000 requests/second and scaling to Black Friday peak surges of 180,000 requests/second across 48 million active shoppers. It decisively resolves the catastrophic infrastructure collapse demonstrated in incident CAP-4919 (where relying on naive reactive CPU-based auto-scaling during a flash sale failed because cloud provider VM provisioning latency of 6 minutes lagged behind a 400% traffic surge in 90 seconds, causing thread pool exhaustion, cascading gateway crashes, and $8.4M in abandoned shopping carts). The architecture enforces predictive workload demand modeling using Erlang-C and M/M/c queueing theory, implements
scheduled pre-warmed capacity provisioning, bounds
cluster resource headroom at >= 30%, and institutes
rigorous synthetic load testing gates.
Detailed Description
Relying exclusively on reactive auto-scaling policies (such as Kubernetes Horizontal Pod Autoscaler scaling on average CPU > 70%) guarantees production failure during sudden traffic spikes. When traffic spikes 4x in two minutes, existing pods saturate before new virtual machine nodes can boot, join the cluster, pull container images, and pass readiness probes. Capacity Planning Architecture establishes proactive, quantitative infrastructure governance: it models request arrivals as stochastic queueing networks, projects seasonal peak envelopes, pre-provisions capacity headrooms ahead of anticipated marketing campaigns, and enforces automated load tests to verify system breaking points.
Incoming Customer Traffic (Baseline 25k req/s -> Flash Surge 180k req/s)
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ Capacity Modeling & Demand Forecasting Engine [CAPARCH-RET-001] │
│ ├── M/M/c Queueing Model: Calculates Required Concurrency & Buffer Headroom│
│ ├── Pre-Warming Orchestrator: Provisions Nodes 30 Minutes Before Sales │
│ └── Resource Quota Governance: Enforces Cluster Headroom Floor >= 30% │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
┌─────────────────────────────┼─────────────────────────────┐
▼ (Stateless Ingress Pods) ▼ (Database Persistent IOPS) ▼ (Cache Memory Fleet)
[ 450 Kubernetes Pods on EKS ] [ AWS Aurora Provisioned IOPS ] [ 128 GB Redis Cluster ]
├── Pre-Warmed Pod Capacity ├── 80,000 Provisioned IOPS ├── Capped Memory Footprint
└── Eliminates CAP-4919 Crash └── Zero Write Queue Stalls └── Sub-2ms Hit Latencies
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Peak Demand Burst Absorption (Zero Dropped Traffic) | Reactive scaling lag crashed the platform in incident CAP-4919 ($8.4M lost carts). | 0.40 | Elena Rostova (Head of Retail Platform Ops) |
| Mathematical Capacity Rigor (Queueing Models) | Empirical Erlang-C modeling eliminates guesswork in infrastructure sizing. | 0.30 | David O'Reilly (Chief Capacity Planning Architect) |
| Cloud Infrastructure Cost Efficiency & Sizing | Over-provisioning static infrastructure 24/7 wastes $4.2M annually in idle cloud spend. | 0.15 | Corporate FinOps & Planning Charter |
| Predictive Headroom Governance (Floor >= 30%) | Sustaining a 30% headroom buffer protects against unexpected viral marketing surges. | 0.15 | SRE Reliability Engineering Charter |
Comparison
| Capacity Planning Strategy | Surge Reaction Time | Headroom Protection | Cost Optimization | Evaluation |
|---|---|---|---|---|
| Option A: Reactive CPU Autoscaling Only (Legacy) | 6 to 8 Minutes (Lag crashed CAP-4919) | Zero (Saturates before scale) | High (Idle between peaks) | Rejected: Caused CAP-4919 $8.4M disaster. |
| Option B: Permanent Static Peak Sizing | Instant (Always running) | High (Wasted capacity) | Terrible ($4.2M wasted compute) | Rejected: Economically unsustainable 365 days/year. |
| Option C: Predictive Modeling + Pre-Warming (Chosen) | Instant (Pre-warmed before surge) | Guaranteed >= 30% Floor | Optimal (Scheduled elasticity) | Selected: Absorbs surges, zero downtime, cost-efficient. |
Result
Option C is selected. Predictive M/M/c queueing models combined with scheduled pre-warming pools are standardized; reactive autoscalers are tuned to secondary buffer roles; a minimum 30% headroom floor is mandatory across all tier-1 services.
Required Mechanisms
1. M/M/c Queueing Model Sizing Formula [MC-QM-01]
- The CAP-4919 Anti-Saturation Sizing:
- Target Peak Throughput ($\lambda$): 180,000 requests/second.
- Service Rate per Pod ($\mu$): 500 requests/second.
- Target Concurrency ($c$):
$$c = \frac{\lambda}{\mu} \times \text{Headroom Factor} = \frac{180,000}{500} \times 1.30 = \mathbf{468\text{ pods}}$$ - Sized to ensure queue waiting time $W_q \approx 0\text{ ms}$ with $99.9%$ probability.
2. Scheduled Pre-Warming Orchestrator [MC-PW-01]
- For planned marketing events, flash sales, and Black Friday promotional hours:
- Automated Kubernetes CronJob scales the node groups and pod deployments to
470 pods exactly 30 minutes prior to event launch.
- Pre-pulls container images, pre-warms database connection pools, and primes local L1 caches before public traffic hits.
3. Continuous Headroom Monitoring & Quota Gates [MC-HM-01]
- The 30% Headroom Floor Invariant:
$$\text{Effective Headroom} = \frac{\text{Allocated Capacity} - \text{Current Peak Demand}}{\text{Allocated Capacity}} \ge 30.0%$$ - If effective headroom drops below 20.0%, automated alerting dispatches to Capacity Engineering, triggering emergency node scaling.
Invariants and Contracts
Mandatory 30% Headroom Floor Invariant [INV-CAP-01]
Tier-1 mission-critical services must maintain at least 30.0% unutilized capacity headroom during peak hours.
Operating clusters at over 75% steady-state utilization during scheduled events is strictly prohibited.
Mandatory Scheduled Pre-Warming Protocol [INV-CAP-02]
Workloads anticipating greater than a 200% traffic surge must execute pre-warmed capacity provisioning.
Relying exclusively on reactive horizontal pod autoscalers for scheduled flash sales is barred.
Continuous Load Verification Certification [INV-CAP-03]
Production architectures must undergo synthetic peak load testing at 120% of projected volume quarterly.
Passing synthetic peak load benchmarks is an absolute prerequisite for major promotional releases.
Explicit Unknowns
- AWS EC2 Spot instance termination frequency during region-wide Cyber Monday cloud hardware demand spikes (G-1).
- Database write IOPS queue depth growth rate when 180,000 concurrent shopping carts execute checkout transactions (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 180,000 requests/sec peak surge | provided | E-commerce peak capacity intake | Current |
| 120 microservices across 48M shoppers | provided | Platform architecture inventory | Current |
| Incident CAP-4919 $8.4M cart abandonment loss | provided | Operations post-mortem audit report | Historical |
| 30% headroom floor target | provided | Corporate SRE Reliability Policy | Current |
| Predictive queueing model + pre-warming selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory 30% headroom invariant INV-CAP-01 | decided | Architectural invariant INV-CAP-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against capacity planning standards:
- Mathematical Sizing: PASS. M/M/c queueing theory sizes 468 pods to absorb 180k RPS surge.
- Pre-Warming Protection: PASS. Scheduled scaling eliminates the 6-minute provisioning lag of CAP-4919.
- Headroom Governance: PASS. Enforces a strict 30% capacity headroom floor across all compute and database tiers.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-CAP-01: David O'Reilly to determine whether AWS Karpenter should replace the legacy Cluster Autoscaler to reduce node provisioning times from 6 minutes to 45 seconds in Q1 (Owner: David O'Reilly).
Next steps
- Platform SRE team deploys the scheduled pre-warming automation controller on AWS EKS.
- Capacity Engineering executes a 200,000 RPS synthetic stress test on staging using distributed k6 runners.
- FinOps team configures automated Slack alerts tracking cluster headroom metrics against the 30% floor.
skill: capacity-planning-architect
Capacity Planning Architecture — Fitness Self-Check [CAPARCH-RET-FIT-001]
Summary
This fitness self-check evaluates the capacity planning architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an autoscaling topology where two uncoordinated autoscaling controllers (e.g. KEDA and native HPA) attempt to mutate pod replica counts on the same deployment simultaneously without consensus. | Kubernetes deployment admission validator probe_conflicting_autoscaler_rejection verifying manifest rejection with diagnostic ERR_DUAL_AUTOSCALER_CONTROLLER_COLLISION. | pass | Confirms Kubernetes API admission webhook checks; does not evaluate direct manual kubectl scale commands. |
| FIT-2: Undefined Grain | Seed a capacity demand forecasting dataset that aggregates compute resource utilization without declaring an explicit time-window grain or regional cluster identifier. | Capacity metrics linter probe_missing_capacity_metrics_grain verifying telemetry rejection with diagnostic ERR_CAPACITY_DATASET_LACKS_DECLARED_GRAIN. | pass | Confirms automated Prometheus metrics schema validation; does not inspect ad-hoc temporary log files. |
| FIT-3: Silent Schema Drift | Seed a workload autoscaling rule that references an obsolete or renamed metric (custom_cpu_ratio -> container_cpu_usage_seconds) without updating the HPA resource specification. | Kubernetes HPA resource validator probe probe_invalid_metric_reference verifying deployment rejection with diagnostic ERR_AUTOSCALER_METRIC_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated Helm linting and kubectl dry-run gates; does not inspect unmonitored test namespaces. |
Residual Risk
- Cloud provider regional hardware quota limits if a simultaneous nationwide data center hardware supply shortage occurs during Black Friday. Accepted by Elena Rostova with multi-region reserve capacity agreements.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of conflicting dual autoscalers | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of capacity datasets lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of obsolete autoscaler metric references | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates capacity fitness probes into automated Helm chart release pipelines.
- Platform team configures Prometheus alerts monitoring Kubernetes node CPU allocation ratios and headroom floors.
- Conduct quarterly capacity planning reviews evaluating actual holiday peaks against predictive queueing models.
capacity-planning-platform-and-demand-ar.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system model that converts authoritative demand scenarios and service constraints into resource requirements, bottleneck forecasts, headroom decisions, scaling lead times, and acquisition/release triggers. It preserves uncertainty, provenance, failure capacity, and observed model error rather than emitting one deterministic size.
Use it when
- User journeys drive synchronous requests, asynchronous work, storage growth, network transfer, model inference, batch jobs, or human operations
- Business forecasts, historical demand, launches, migrations, seasonality, incidents, and uncertainty require multiple scenarios
- CPU, memory, storage, IOPS, bandwidth, connections, workers, partitions, quotas, licenses, and human capacity constrain different stages
- Bottlenecks shift as one resource scales and shared dependencies serve many products or tenants
- Redundancy, failure, maintenance, deployments, rebalancing, recovery, and geographic placement consume capacity
- Static provisioning, vertical/horizontal scaling, queues, load shedding, reservations, and flexible capacity have different lead times
For example: “Enrolment opens in six weeks and we're onboarding four new universities. Last year the portal fell over on day one and nobody can tell me how many servers we need.”
What you get
- architecture/capacity-planning-architect/README.md
- architecture/capacity-planning-architect/00-overview/capacity-planning-architect-overview.md
- architecture/capacity-planning-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to size one instance/database, configure autoscaling, optimize cloud cost, run a load test, tune performance, request quota, purchase infrastructure, or operate SRE.
How it works
- Check the question is how much, not how fast.
- Get the demand scenario from its owner, in business units.
- Establish the conversion from business unit to resource, and measure it.
- Model the headroom you need and why.
- Name the constraint that binds first.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install