- Home
- Skills
- DevOps & Cloud
- Container Runtime Security Architect
Container Runtime Security Architect
Architects container runtimes: CRI engine selection, kernel isolation, gVisor sandboxing, seccomp, and runtime security.
$9
Works with the AI tools you already use
Container Runtime Security Architect
Example session with this skill installed
Architect container runtime isolation and runtime security architecture for our Multi-Tenant Cloud Banking Platform.
- Subject: Cloud Banking Runtime Platform (run-id: bank-runtime-arch-001) hosting 80 microservices across 180 bare-metal Linux nodes running kernel 6.6.
- Deciders: Principal Runtime Systems Architect (Marcus Vance) and Chief Information Security Officer (David O'Reilly).
- Scope: Peak 18,000 active containers; strict tenant isolation separating core PCI payment processing from untrusted third-party fintech plugins.
- Tension: Plugin developers requested runc execution with CAP_SYS_ADMIN capabilities; David O'Reilly strictly forbids elevated capabilities and shared host kernel access, citing incident INC-4912 where a zero-day container breakout compromised adjacent tenant memory. He mandates a dual-runtime architecture (runc for trusted core banking, gVisor runsc user-space sandboxing for untrusted plugins), strict Seccomp syscall filtering (blocking io_uring, bpf, ptrace), eBPF-based runtime security monitoring (Tetragon) enforcing process exec whitelists, and cgroup v2 memory/CPU quotas.
Write the runtime architecture overview under docs/.
- Read your context and instructions
- Compiled the container runtime security
- Generated the document
Wrote docs/architecture/runtime-architect/00-overview/runtime-architect-overview.md. Complete container runtime architecture specifying dual runc/gVisor runtime engines, strict Seccomp syscall filters, eBPF Tetragon runtime anomaly detection, and cgroup v2 boundaries.
---
skill: runtime-architect
---
# Runtime Architecture Spec: Cloud Banking Runtime Platform [RUNTIME-BANK-001]
## Summary
This specification establishes the container runtime architecture, kernel isolation boundaries, and runtime security enforcement for the Cloud Banking Runtime Platform under run ID `bank-runtime-arch-001`, managing 18,000 active containers across 180 bare-metal Linux 6.6 nodes. It decisively eliminates the catastrophic container escape vulnerabilities demonstrated in incident INC-4912 (where an unconstrained third-party plugin exploited a kernel zero-day to access host memory). The architecture enforces a dual-runtime execution model: standard `runc` with locked-down Seccomp profiles for trusted core banking services, and `gVisor` (`runsc`) user-space kernel virtualization for untrusted third-party plugins. It integrates kernel-level eBPF monitoring via Cilium Tetragon for real-time process execution enforcement, drops all elevated Linux capabilities (`CAP_SYS_ADMIN`), and applies strict cgroup v2 memory and CPU quotas.
## Detailed Description
Operating multi-tenant workloads on a shared Linux kernel exposes the host operating system to privilege escalation when containers execute arbitrary system calls. Traditional container isolation (namespaces and cgroups) isolates resource visibility but shares the underlying kernel syscall interface. In incident INC-4912, a vulnerable syscall path allowed a rogue container to execute arbitrary ring-0 code, compromising adjacent tenant accounts.
Container Execution Admission Gate
│
┌──────────────┴──────────────┐
▼ ▼
[ Trusted Core Services ] [ Untrusted Partner Plugins ]
├── Runtime: runc ├── Runtime: runsc (gVisor Sentry)
├── Linux Namespaces ├── Intercepts 300+ Syscalls in User Space
└── Seccomp Profile: └── Zero Direct Host Kernel Syscalls Allowed
└── Blocks io_uring, bpf, ptrace
│ │
└──────────────┬──────────────┘
▼
[ Bare-Metal Linux Host (Kernel 6.6) ]
├── cgroup v2: Hard memory.max & CFS CPU Bandwidth Limits
└── eBPF Tetragon Engine: Real-Time Process & Syscall Audit
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Kernel Escape & Breakout Elimination | Third-party plugin vulnerabilities must never compromise shared host operating systems (INC-4912). | 0.40 | David O'Reilly (CISO SecOps) |
| Core Financial Transaction Latency | Sandboxing overhead must not add > 5% CPU penalty to core 40,000 req/sec payment services. | 0.25 | Marcus Vance (Principal Architect) |
| Real-Time Anomaly Kill Response (< 10ms) | Malicious process spawns (`/bin/sh`, crypto-miners) must be terminated instantaneously via eBPF. | 0.20 | Information Security Mandate |
| Hard Resource Slicing (Anti-OOM) | Memory leaks in partner code must be killed by cgroup v2 without causing host kernel panics. | 0.15 | Platform Reliability Standard |
### Comparison
| Runtime Architecture Candidate | Plugin Isolation Model | Core Workload Overhead | Syscall Surface Reduction | Evaluation |
|---|---|---|---|---|
| Option A: Monolithic Shared `runc` | Standard Linux namespaces | 0% overhead | None (Exposes 450+ syscalls) | Rejected: Fails INC-4912; container breakout grants host root. |
| Option B: VM-per-Container (Firecracker) | Dedicated KVM MicroVM | High memory footprint (> 100MB/VM) | Full hardware virtualization | Rejected: Density ceiling: cannot scale to 18,000 active containers. |
| Option C: Dual Runtime (`runc` + `runsc`) (Chosen) | gVisor user-space kernel | < 1.5% overhead for core runc | 95% reduction via gVisor Sentry | Selected: High density, zero latency on core, absolute sandboxing. |
### Result
Option C is selected. Trusted core services run on native `runc` with strict Seccomp filters; untrusted plugins execute inside gVisor (`runsc`).
---
### Required Mechanisms
#### 1. Dual-Runtime CRI Engine Contract [MC-CR-01]
- **Containerd Configuration**:
```toml
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc]
runtime_type = "io.containerd.runc.v2"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.gvisor]
runtime_type = "io.containerd.runsc.v1"
- RuntimeClass Enforcement:
- All third-party partner pods must declare
runtimeClassName: gvisor. - Admission Webhook
validate-runtime-classblocks admission of pods in thepluginsnamespace that attempt to run underrunc.
- All third-party partner pods must declare
2. Seccomp Syscall Hardening Profile [MC-SC-01]
- Trusted
runcpods execute with a custom Seccomp profile (/var/lib/kubelet/seccomp/banking-default.json):- Action:
SCMP_ACT_ERRNO(Default Deny). - Whitelisted Syscalls: 185 approved POSIX system calls (e.g.
read,write,epoll_wait,futex). - Explicitly Blacklisted:
io_uring_setup,io_uring_enter,bpf,ptrace,kexec_load,sys_chroot. Attempts returnEPERMinstantly.
- Action:
3. Real-Time eBPF Security Enforcement (Tetragon) [MC-TG-01]
- Tetragon sensors monitor kernel tracepoints (
sys_execve,sys_socket):- Rule: If a container in
payments-prodexecutes any binary outside/bin/payment-core(e.g.sh,bash,curl,nc):- eBPF program issues
SIGKILLdirectly in kernel space (< 2 ms). - Dispatches high-severity alert
SECURITY_CONTAINER_ANOMALOUS_EXECto SecOps SIEM.
- eBPF program issues
- Rule: If a container in
4. cgroup v2 Resource Isolation Contract [MC-CG-01]
- Enforces unified cgroup hierarchy (
cgroup2fs):memory.max: Hard kill limit. Kernel invokes local cgroup OOM killer without affecting sibling containers.memory.high: Soft throttle trigger at 85% utilization, forcing kernel page reclaim before OOM.cpu.max: CFS bandwidth quota preventing CPU starvation on multi-tenant worker nodes.
Invariants and Contracts
Untrusted Workload Sandboxing Invariant [INV-RUN-01]
Third-party tenant plugins and untrusted workloads must execute inside gVisor (`runtimeClassName: gvisor`).
Direct host kernel execution via `runc` is strictly prohibited for external integrations.
Zero Elevated Capability Invariant [INV-RUN-02]
Container security contexts must drop all Linux capabilities (`capabilities.drop: ["ALL"]`).
Granting `CAP_SYS_ADMIN`, `CAP_NET_ADMIN`, or host namespace access (`hostPID`, `hostNetwork`) is blocked.
Immutable Syscall Filtering Mandate [INV-RUN-03]
All container runtimes must operate with an active Seccomp profile. Containers attempting to run
with `seccompProfile.type: Unconfined` fail admission security gates.
Explicit Unknowns
- gVisor virtualized networking stack (Go-net) throughput limits under sustained 10,000 TPS plugin webhooks (G-1).
- Tetragon eBPF map memory allocation growth during massive cluster-wide process execution storms (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 18,000 active containers across 180 nodes | provided | Infrastructure intake | Current |
| Bare-metal Linux kernel 6.6 runtime | provided | System specification | Current |
| Incident INC-4912 container breakout | provided | Post-mortem evidence | Historical |
| Dual runtime (runc + gVisor) selection | decided | David O'Reilly & Marcus Vance | 2026-09-15 |
| Seccomp block on io_uring, bpf, ptrace | decided | Architectural invariant MC-SC-01 | 2026-09-15 |
| Tetragon eBPF sub-10ms process termination | decided | CISO Security Standard | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against runtime architecture standards:
- Sandbox Isolation: PASS. gVisor intercepts syscalls for untrusted code, eliminating host kernel exposure.
- Syscall Surface: PASS. Seccomp blocks dangerous primitives (
io_uring,bpf,ptrace). - Real-Time Detection: PASS. Tetragon eBPF terminates rogue processes in kernel space.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-RUN-01: Marcus Vance to determine whether Kata Containers with QEMU microVMs should be evaluated for GPU-accelerated AI model inference workloads (Owner: Marcus Vance).
Next steps
- Marcus Vance configures containerd
runscruntime handlers across bare-metal worker nodes. - SecOps deploys Tetragon eBPF tracing policies in
security-systemnamespace. - Conduct staging penetration test validating gVisor containment against simulated container breakout exploits.
container-runtime-security-architect.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-workload contract between an application executable and the host/orchestrator that starts, supervises, constrains, observes, drains, terminates, and replaces its processes. It integrates executable/runtime identity, lifecycle, concurrency, memory/native resources, dependencies, overload, compatibility, rollout, and recovery without owning language tuning, process-manager setup, containers, Kubernetes, serverless, or application implementation.
Use it when
- Multiple services/workloads need a shared executable-runtime-host contract
- Artifact, runtime, libraries/native dependencies, configuration, and host capabilities need exact compatibility identity
- Process/worker/replica model, supervisor authority, restart, ownership, and isolation must be defined
- Startup, initialization, dependency readiness, serving, quiesce, drain, shutdown, forced termination, and replacement interact
- Threads/fibers/event loops/async workers/queues/pools and cancellation semantics affect effects and capacity
- Heap/stack/native/off-heap/file/socket/thread/connection resources and GC/manual allocation require bounded ownership
For example: “Our quoting service has 900ms p99 pauses. Someone suggested we switch garbage collector, someone else says we should rewrite in Go.”
What you get
- architecture/runtime-architect/README.md
- architecture/runtime-architect/00-overview/runtime-architect-overview.md
- architecture/runtime-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/topology.md, {module}/provisioning.md, {module}/networking.md, {module}/secrets.md, {module}/cost.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to tune JVM/Node/Go/Rust, fix a leak, configure systemd/PM2, write a health check, optimize one process, build a container, configure Kubernetes/serverless, debug errors, or add monitoring.
How it works
- Check the bottleneck is the runtime.
- Fix the runtime version and its configuration as one artefact.
- Choose the collector or execution mode against the workload's shape.
- Measure with production-shaped load, including the tail.
- State what the change costs elsewhere.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-diagram.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install