- Home
- Skills
- DevOps & Cloud
- Kubernetes Platform Architect
Kubernetes Platform Architect
Designs a shared multi-tenant Kubernetes platform: topology, node pools, isolation, upgrades, capacity, recovery.
$12
Works with the AI tools you already use
Kubernetes Platform Architect
Example session with this skill installed
Architect the shared multi-tenant Kubernetes platform for our Core Banking and Payments workloads.
- Subject: Internal Kubernetes Platform (EKS v1.30, cluster name: k8s-prod-useast1) hosting 12 microservices across 3 engineering teams (Payments, Accounts, Risk).
- Decider: Cloud Architecture Review Board (Lead: Alex Mercer) and CISO SecOps (Lead: Sarah Chen).
- Workload & Capacity: 240 worker nodes across 3 availability zones (AWS us-east-1a/b/c); peak 1,800 pods; 65,000 requests/sec aggregate ingress traffic.
- Tension: Payments engineering insists on running database pods and cache daemons on shared general worker nodes with host networking for ultra-low latency; SecOps strictly forbids host networking and mandates network segmentation, pod security standard 'restricted', and dedicated tainted node pools for PCI-scoped payments services.
- Policies & Upgrades: Zero downtime during quarterly Kubernetes minor version upgrades; Karpenter autoscaling with 15-minute consolidation delay; Cilium CNI enforcing strict default-deny NetworkPolicies between namespaces.
Write the platform architecture overview under docs/.
- Read your context and instructions
- Compiled the kubernetes platform architect
Wrote docs/architecture/kubernetes-architect/00-overview/kubernetes-architect-overview.md. Complete Kubernetes platform architecture specifying multi-tenant isolation, tainted node pools, Cilium NetworkPolicies, and zero-downtime rolling upgrade lifecycles.
---
skill: kubernetes-architect
---
# Kubernetes Platform Architecture: Core Banking Cluster [K8S-BANK-001]
## Summary
This specification defines the multi-tenant Kubernetes platform architecture for `k8s-prod-useast1` (Amazon EKS v1.30), hosting 12 core banking and payment microservices across 3 availability zones. It resolves the conflict between low-latency application demands and PCI-DSS isolation requirements by rejecting shared host networking in favor of dedicated, tainted PCI node pools, Cilium eBPF network segmentation with default-deny policies, Pod Security Standard `restricted`, and Karpenter-driven capacity autoscaling. The architecture enforces zero-downtime rolling worker upgrades across 240 nodes and 1,800 pods.
## Detailed Description
Core banking workloads require strict multi-tenant isolation combined with high-throughput network performance. Attempting to run stateful or PCI-scoped workloads on general shared node pools with host networking exposes neighbor pods to network eavesdropping and violates compliance baselines.
Incoming Ingress (65,000 req/sec)
│
▼
[ Cilium eBPF Ingress Gateway ]
│
┌────────┴────────┐
▼ ▼
[ General Pool ] [ Tainted PCI Pool ]
(Accounts / Risk) (Payments Services)
No Taints Taint: pci=true:NoSchedule
PSS: Restricted PSS: Restricted + NetworkPolicy Default-Deny
### Alternatives rejected
| Option | Why it was not taken | Under what evidence it would win |
|---|---|---|
| Shared Node Pools with Host Networking | Severe security hazard; bypasses network policies, exposes host interfaces, and violates PCI-DSS requirements. | Never. Multi-tenant regulatory environments strictly disallow host network namespace sharing. |
| Separate Physical Clusters per Team | Triples control-plane licensing costs and fragments infrastructure utilization across small development teams. | If cross-team compliance mandates complete hardware cryptographic isolation. |
| Cluster Autoscaler with Fixed ASGs | Slower scaling response time (5+ minutes) compared to Karpenter (sub-60 seconds) during flash transaction bursts. | If cluster is deployed on bare-metal or environments without Karpenter provider support. |
## Contracts and Invariants
PCI Node Pool Isolation [INV-K8S-01]
All payments workloads subject to PCI scope must run on dedicated node pools tainted with
`pci=true:NoSchedule`. Workloads must specify matching tolerations and node affinity:
`nodeSelector: tier=pci`. Co-locating non-PCI pods on these nodes is blocked by admission controllers.
Default-Deny Ingress and Egress [INV-K8S-02]
Every tenant namespace must enforce CiliumNetworkPolicy default-deny for both ingress
and egress. Cross-namespace communication requires explicit mutual policy authorization
referencing service identities.
Pod Security Standard Enforcement [INV-K8S-03]
All tenant namespaces must enforce Pod Security Standard `restricted`. Pods attempting
to run as root (`runAsNonRoot: true` required), mount host paths, or use `hostNetwork: true`
are deterministically rejected by the admission webhook.
Quarterly Upgrade Budget [INV-K8S-04]
Quarterly EKS minor version upgrades must execute node draining with PodDisruptionBudgets
(PDBs) ensuring `minAvailable: 75%` per deployment. Eviction timeout is capped at 300 seconds.
## Ownership and Handoffs
| Concern | Owner | Handoff payload | Blocked until |
|---|---|---|---|
| Cluster Control Plane & EKS Upgrades | Platform Engineering (Alex Mercer) | `k8s_platform_contract` | Terraform infrastructure approval |
| Security Policy & Cilium Rules | CISO SecOps (Sarah Chen) | `k8s_network_security_spec` | PSS & NetworkPolicy audit pass |
| Payments Namespace Deployment | Core Payments Engineering | Application manifests with tolerations | PCI node pool provisioning |
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 240 nodes, 1,800 pods across 3 AZs | provided | Capacity intake | Current |
| 65,000 req/sec aggregate ingress | provided | Ingress traffic specification | Current |
| Dedicated tainted PCI node pools | decided | Architectural decision INV-K8S-01 | 2026-09-15 |
| Host networking rejection | decided | Sarah Chen & Alex Mercer | 2026-09-15 |
| Pod Security Standard `restricted` | provided | CISO security requirement | Current |
| Karpenter autoscaler selection | decided | Infrastructure team standard | Current |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against Kubernetes platform contracts:
- **Tenant Isolation**: PASS. Taints, tolerations, and Cilium default-deny policies isolate payments.
- **Compliance Baseline**: PASS. PSS `restricted` admission profile rejects host networking and root execution.
- **Availability Bounds**: PASS. PDBs and 3-AZ spread protect 65k req/sec ingress against disruption during upgrades.
- **Autoscaling Mechanics**: PASS. Karpenter provisioner configured with consolidation delay.
## Open Decisions
- `DEC-K8S-01`: SecOps to confirm whether Cilium WireGuard encryption in transit is mandatory for cross-AZ worker node traffic (Owner: Sarah Chen).
## Next steps
1. Alex Mercer provisions tainted worker node pools via Terraform Karpenter provisioner manifests.
2. Sarah Chen deploys cluster-wide admission webhook enforcing PSS `restricted` across all namespaces.
3. Validate rolling node pool drain and replacement in staging under synthetic 20,000 TPS load.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the shared Kubernetes substrate boundary for one or more clusters serving multiple workloads or tenants. It integrates API/control-plane topology, nodes, scheduling, isolation, networking/storage interfaces, extensions, upgrades, capacity, security, observability, and recovery without absorbing manifest authoring, workload troubleshooting, operators, GitOps, service mesh, containers, or cloud-provider architecture.
Use it when
- Multiple workloads/tenants require a cluster or fleet boundary and ownership model
- Control-plane, cluster, region/zone, node-pool, namespace, and workload identities need stable contracts
- Failure/upgrade/security/compliance boundaries determine cluster/fleet topology
- Workload placement, requests/limits, priority, disruption, affinity, taints, topology, and admission interact
- Tenant isolation spans API/RBAC, admission, scheduling, network, storage, secrets, observability, quotas, and failures
- Cluster networking, ingress/egress, DNS, service discovery, load balancing, and mesh need explicit handoffs
For example: “One team's batch job consumed every node and took down the payments API in the same cluster. We also can't upgrade because nobody knows which workloads use the removed APIs.”
What you get
- architecture/kubernetes-architect/README.md
- architecture/kubernetes-architect/00-overview/kubernetes-architect-overview.md
- architecture/kubernetes-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/topology.md, {module}/provisioning.md, {module}/networking.md, {module}/secrets.md, {module}/cost.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to write manifests/Helm/Kustomize, deploy or debug one workload/pod, install a cluster, build an operator, configure GitOps/service mesh, tune one HPA/VPA, secure one namespace, or choose a managed service.
How it works
- Check Kubernetes is already the decision.
- Decide the cluster boundary and count.
- Choose the multi-tenancy model and its enforcement.
- Fix the networking and storage plugins with their consequences.
- State the upgrade path and its cadence.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-diagram.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 13 days ago
- Passed all security checks, Safe to install