- Home
- Skills
- DevOps & Cloud
- Kubernetes Workload Resource Design
Kubernetes Workload Resource Design
Designs Kubernetes workload resources: Deployments, resource QoS classes, disruption budgets, and pod security contexts.
$5
Works with the AI tools you already use
Kubernetes Workload Resource Design
Example session with this skill installed
Design the production Kubernetes workload resource specification for our Core Payment Ledger Microservice.
- Subject: payment-ledger-service v2.4 (run-id: payment-workload-001) running on AWS EKS across 3 availability zones.
- Deciders: Lead Kubernetes Architect (Marcus Vance) and Payment Reliability Director (Elena Rostova).
- Scale & Workload: Peak 2,400 ledger write requests/sec; 24 container replicas; latency SLA p99 <= 45 ms.
- Tension: Dev team proposed omitting CPU/memory limits to allow burst scaling, running with default root privileges, and skipping disruption budgets; Marcus Vance and Elena Rostova strictly reject unconstrained pods and missing budgets, citing incident INC-3810 where an unconstrained noisy neighbor starved node memory and forced node drain evictions that took down payment settlement. They mandate a Guaranteed QoS class (requests equal limits: 2 CPU, 4 GiB memory), a PodDisruptionBudget (minAvailable: 80%), read-only root filesystem with non-root UID 10001, drop ALL capabilities, topologySpreadConstraints across 3 AZs (maxSkew: 1), and a 35-second termination grace period.
Write the workload contract specification under docs/.
- Read your context and instructions
- Compiled the kubernetes workload resource
- Generated the UI component
Wrote docs/architecture/tasks/payment-workload-001/k8s-design/k8s-workload-contract.md. Complete Kubernetes workload resource contract specifying Guaranteed QoS resource allocations, PodDisruptionBudgets, topology spread constraints across 3 AZs, and least-privilege security contexts.
---
skill: k8s-design
---
# Kubernetes Workload Contract: Payment Ledger Service [K8S-PAY-001]
## Summary
This specification establishes the production Kubernetes workload resource contract and pod runtime isolation policies for `payment-ledger-service v2.4` under run ID `payment-workload-001`, managing 24 replicas across 3 AWS availability zones on EKS. It decisively eliminates the noisy neighbor evictions and availability outages demonstrated in incident INC-3810 (where unconstrained container memory allocations triggered Linux kernel OOM kills and uncoordinated node drains). The contract enforces a Kubernetes `Deployment` controller with a Guaranteed Quality of Service (QoS) tier (`requests == limits`), a PodDisruptionBudget maintaining >= 80% active replicas during cluster upgrades, multi-AZ `topologySpreadConstraints` with `maxSkew: 1`, hardened container security contexts (`runAsNonRoot: true`, `readOnlyRootFilesystem: true`, dropping all Linux capabilities), and a 35-second graceful connection drain budget.
## Detailed Description
Unconstrained container workloads in multi-tenant Kubernetes clusters cause severe memory starvation, CPU throttling, and unpredictable evictions. When pods omit explicit resource requests or set memory limits far exceeding requests, the Kubernetes scheduler overcommits node memory, exposing critical financial workloads to the Linux kernel Out-Of-Memory (OOM) killer.
Cluster Scheduler / Kubelet Admission
│
▼
[ Pod Admission: payment-ledger-service ]
├── 1. Guaranteed QoS Check: Requests == Limits (2 CPU, 4 GiB RAM)
├── 2. SecurityContext Gate: UID 10001, drop ALL capabilities, read-only root
└── 3. Topology Spread Filter: Distributes 24 pods evenly across 3 AZs (8/AZ)
│
▼ (Scheduled to Worker Nodes)
[ Active Runtime Execution ]
├── PodDisruptionBudget: minAvailable: 80% (Guarantees >= 20 pods running)
├── Graceful Termination: terminationGracePeriodSeconds: 35
└── Ingress Health: Probes validate database pool before receiving traffic
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Anti-Eviction Immunity (Guaranteed QoS) | Financial ledger transactions must never be terminated by the kernel OOM killer during memory contention (INC-3810). | 0.40 | Marcus Vance (Lead K8s Architect) |
| High Availability During Cluster Maintenance | Node upgrades and Karpenter consolidation must never drop ledger processing capacity below 80%. | 0.25 | Elena Rostova (Payment Reliability) |
| Container Defense-in-Depth & Hardening | Non-root users and read-only filesystems neutralize post-exploitation persistence and malware drops. | 0.20 | PCI-DSS Compliance Mandate |
| Even Multi-AZ Blast Radius Distribution | Pods must not concentrate on a single AWS AZ, preventing outages during single-datacenter loss. | 0.15 | Cloud Infrastructure Standard |
### Comparison
| Workload Configuration Dimension | Dev Proposal | Best-Effort Burst Model | Guaranteed QoS Contract (Chosen) | Evaluation |
|---|---|---|---|---|
| QoS Class | Burstable (No Limits) | Best-Effort | Guaranteed (`requests == limits`) | Guarantees pod is last in line for kernel OOM termination. |
| Pod Disruption Budget | Omitted (0% protected) | `maxUnavailable: 50%` | `minAvailable: 80%` (PDB) | Prevents node drains from reducing cluster below 20 healthy pods. |
| Security Context | `root` (UID 0), writable root | `runAsNonRoot`, writable root | `UID 10001`, `readOnlyRootFilesystem: true` | Complete prevention of unauthorized binary execution on disk. |
| AZ Distribution | Kube-scheduler default | Soft node affinity | `topologySpreadConstraints` (maxSkew: 1) | Strict mathematical balance (8 pods per availability zone). |
### Result
Guaranteed QoS with strict PDB enforcement is selected. The workload receives top scheduling priority, immune to kernel eviction, while security contexts eliminate write access to container root partitions.
---
### Required Mechanisms
#### 1. Workload Deployment Specification [MC-WD-01]
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-ledger-service
namespace: payments-prod
labels:
app.kubernetes.io/name: payment-ledger-service
app.kubernetes.io/part-of: core-banking
spec:
replicas: 24
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25%
maxUnavailable: 0
selector:
matchLabels:
app: payment-ledger
template:
metadata:
labels:
app: payment-ledger
spec:
terminationGracePeriodSeconds: 35
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: payment-ledger
containers:
- name: ledger-app
image: registry.bank.internal/payments/payment-ledger:v2.4.0
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: "2000m"
memory: "4096Mi"
limits:
cpu: "2000m"
memory: "4096Mi"
volumeMounts:
- name: tmp-volume
mountPath: /tmp
volumes:
- name: tmp-volume
emptyDir:
medium: Memory
sizeLimit: 128Mi
2. PodDisruptionBudget (PDB) Contract [MC-PDB-01]
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payment-ledger-pdb
namespace: payments-prod
spec:
minAvailable: 80%
selector:
matchLabels:
app: payment-ledger
- Operational Effect: Eviction API calls (
kubectl drain) block if active pods drop to 19 or fewer. Node drains execute sequentially, maintaining at least 20 active replicas serving customer traffic.
3. Graceful Termination & Drain Lifecycle [MC-GT-01]
- Upon receipt of
SIGTERM:- Pod enters
Terminatingstate; Kubelet immediately detaches pod IP fromEndpointSlices. - PreStop hook executes a 5-second sleep to allow external load balancer routes to update:
preStop: { exec: { command: ["/bin/sleep", "5"] } } - Go server flushes in-flight database ledger mutations to PostgreSQL (budget: 25 seconds).
- Total execution finishes well within the 35-second
terminationGracePeriodSecondswindow.
- Pod enters
Invariants and Contracts
Guaranteed QoS Invariant [INV-K8S-01]
Payment ledger pod specifications must configure identical values for CPU/memory requests and limits.
Omitting limits or creating Burstable/BestEffort QoS pods is rejected by admission webhooks.
Mandatory Pod Disruption Protection [INV-K8S-02]
Workloads must be covered by a PodDisruptionBudget declaring `minAvailable >= 80%`.
Cluster maintenance operations that breach this availability ceiling are rejected by the API server.
Strict Least-Privilege Isolation [INV-K8S-03]
Pods must enforce `readOnlyRootFilesystem: true`, `runAsNonRoot: true`, and `capabilities.drop: ["ALL"]`.
Containers attempting to run as root (UID 0) are blocked by cluster Kyverno policies.
Explicit Unknowns
- Kubelet memory cgroup v2 page cache writeback latency during sustained 2,400 TPS SQLite WAL operations (G-1).
- Karpenter node consolidation disruption frequency when packing multi-AZ guaranteed pods (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 24 replicas across 3 AWS AZs | provided | Capacity intake | Current |
| Peak 2,400 ledger write requests/sec | provided | Traffic intake | Current |
| Latency SLA p99 <= 45 ms | provided | Performance SLA | Current |
| Incident INC-3810 noisy neighbor eviction | provided | Post-mortem evidence | Historical |
| Guaranteed QoS allocation (2 CPU, 4 GiB) | decided | Marcus Vance (Lead K8s Architect) | 2026-09-15 |
| PDB minAvailable: 80% | decided | Elena Rostova (Payment Reliability) | 2026-09-15 |
| Topology spread maxSkew: 1 across AZs | decided | Architectural invariant INV-K8S-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against Kubernetes workload standards:
- QoS Hardening: PASS.
requests == limitsachieves Guaranteed QoS class, eliminating OOM evictions. - Maintenance Safety: PASS. PDB
minAvailable: 80%keeps >= 20 pods running during node drains. - Security Boundaries: PASS. Read-only root filesystem, unprivileged UID 10001, and capabilities dropped.
- Zone Balancing: PASS.
topologySpreadConstraintsenforces strict multi-AZ distribution (8 pods/zone).
Open Decisions
DEC-K8S-01: Marcus Vance to determine whether Kubernetes In-Place Pod Vertical Scaling (alpha in 1.30) should be evaluated for dynamic memory resizing without container restarts (Owner: Marcus Vance).
Next steps
- Marcus Vance verifies the Deployment and PDB YAML manifests against cluster Gatekeeper policies.
- Platform team configures CI test running
kubeconformto validate Kubernetes 1.30 API compatibility. - Conduct staging node drain test (
kubectl drain) under active 2,400 TPS traffic to verify zero dropped transactions.
kubernetes-workload-resource-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted application and platform contracts into a minimum Kubernetes workload resource graph. It defines controller, Pod, Service, config, storage, identity, resource, health, security and placement semantics without choosing cluster architecture or applying manifests.
Use it when
Use when a bounded application workload already has authoritative runtime and platform contracts and needs exact Kubernetes resource semantics.
For example: “Our financial ledger worker service repeatedly gets killed during peak transaction processing because memory limits are missing, and deployments trigger downtime because Kubernetes routes traffic to new Pods before database connections are initialized.”
What you get
- Kubernetes YAML Specs
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/k8s-design/.
What it will not do
Do not use for cluster architecture, Helm/Argo packaging/reconciliation, Ingress, autoscaling, rollout strategy, NetworkPolicy, implementation or incidents.
How it works
- Check workload resource mapping is required.
- Select controller type.
- Establish Pod spec and container lifecycle.
- Formulate resource requests and QoS class.
- Configure health probes and termination grace.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install