- Home
- Skills
- DevOps & Cloud
- Internal Infrastructure Platform Architect
Internal Infrastructure Platform Architect
Architects shared infrastructure capability platforms: provider-consumer contracts, tenancy, lifecycle APIs, and quotas.
$9
Works with the AI tools you already use
Internal Infrastructure Platform Architect
Example session with this skill installed
Architect the shared infrastructure platform capability contracts and control-plane tenancy for our FinTech Engineering Organization.
- Subject: FinTech Infrastructure Capability Platform (run-id: fintech-infra-plat-001) serving 600 engineers across 45 squads.
- Deciders: Principal Platform Architect (Marcus Vance) and Head of Infrastructure Engineering (Elena Rostova).
- Scope & Scale: 110 active microservices; 2 shared capabilities (Managed PostgreSQL, Managed Kafka Streaming); 250 self-service operations/week.
- Tension: Feature squads demanded ungoverned direct AWS IAM access and unmetered streaming clusters; an unconstrained replay in INC-5520 saturated Kafka broker bandwidth and starved payment settlement. Elena Rostova and Marcus Vance strictly reject agreement-only limits and direct cloud access. They mandate an infrastructure control plane with explicit capability contracts (standard vs critical tiers), broker-enforced tenancy and throughput quotas, declarative lifecycle APIs (provision, tier upgrade, safe decommission), and quota-based cost attribution.
Write the platform architecture overview under docs/.
- Read your context and instructions
- Compiled the internal infrastructure platform
Wrote docs/architecture/infra-platform-architect/00-overview/infra-platform-architect-overview.md. Complete infrastructure capability platform specification defining provider-consumer contracts, broker-enforced tenancy, declarative lifecycle APIs, and quota-based cost attribution.
---
skill: infra-platform-architect
---
# Infrastructure Capability Platform Architecture: FinTech Platform [INFRA-PLAT-001]
## Summary
This specification establishes the shared infrastructure capability platform architecture, control-plane mediation contracts, and multi-tenant tenancy boundaries for the FinTech Engineering Organization under run ID `fintech-infra-plat-001`. It supports 600 software engineers and 110 active microservices across 45 product squads executing 250 self-service provisioning and lifecycle operations weekly. The design decisively eliminates un-governed cloud console access and agreement-only resource sharing, resolving the root cause of incident INC-5520 where an unconstrained consumer replay exhausted shared Kafka cluster network bandwidth and disrupted real-time payment settlement. The architecture specifies authoritative capability contracts across two core infrastructure offerings (Managed PostgreSQL Database and Managed Kafka Event Streaming), broker-enforced isolation and quota ceilings, declarative lifecycle operations (provision, tier change, safe decommission), and quota-based cost attribution.
## Detailed Description
Operating shared infrastructure without explicit capability contracts and broker-enforced quotas invites noisy-neighbor failures and unowned resource sprawl. In incident INC-5520, an unmetered analytics replay saturated shared broker egress bandwidth, starving real-time transaction consumers. Relying on social agreements or honor-system limits fails under operational pressure. Furthermore, allowing direct cloud console or IAM write access creates configuration drift, orphaned datastores, and unmanaged blast radiuses.
Consumer Interaction (Declarative Manifest / Platform API / Idempotency Key)
│
▼
[ Infrastructure Platform Control Plane (fintech-infra-plat-001) ]
├── Tenancy & Quota Admission Gate (Validates per-squad quotas before provisioning)
├── Policy Authority Engine (Validates tier constraints, KMS encryption, private endpoints)
└── Orchestration State Machine (Queued ──► Executing ──► Provisioned ──► Bound)
│
┌─────────────────────────┴─────────────────────────┐
▼ ▼
[ Capability: Managed PostgreSQL (CC-PG-01) ] [ Capability: Managed Kafka Streaming (CC-KF-01) ]
├── Standard Tier: Burstable, 99.9% avail ├── Standard Tier: 20 MB/s, 7d retention, 99.9%
└── Critical Tier: Multi-AZ HA, 99.99% avail └── Critical Tier: 100 MB/s, 30d retention, 99.95%
│ │
▼ (Executor: Managed Cloud Provider) ▼ (Broker-Enforced Quota)
AWS RDS PostgreSQL Private Subnet Dedicated Cluster per Tier + Client Quotas
### Alternatives rejected
| Option | Why it was not taken | Under what evidence it would win |
|---|---|---|
| Direct Cloud IAM Access & Ad-hoc IaC | High cognitive load, unmonitored security drift, and zero lifecycle auditability across 45 squads. | Never in regulated fintech environments requiring strict SOC2/PCI data isolation. |
| Single Undifferentiated Shared Cluster | Proved catastrophic in INC-5520: noisy-neighbor workload consumed entire cluster bandwidth, halting payments. | Only viable in low-traffic non-production development environments with single-tenant workloads. |
| Agreement-Only / Social Honor Limits | Fails deterministically under pressure; software engineers bypass guidelines during urgent batch replays. | Reversal requires mathematically provable client-side rate limiters that cannot be reconfigured. |
| Centralized Manual Infrastructure Ticket Queue | Average provisioning lead time exceeds 14 days, creating engineering bottlenecks and blocking feature delivery. | Only if organizational change volume drops below 2 operations per month. |
## Contracts and Invariants
### 1. Capability Catalogue Contracts
#### Managed PostgreSQL Database Capability Contract [CC-PG-01]
- **Offered Outcome**: Governed relational database instance with automated backup, KMS encryption at rest, and private VPC endpoint bindings.
- **Tiers & Guarantees**:
- `standard`: Single-AZ instance, 99.9% monthly availability, daily automated snapshot, 7-day retention, max 500 connections.
- `critical`: Multi-AZ synchronized instance with read replica, 99.99% monthly availability, continuous WAL archiving, 35-day point-in-time recovery, max 2,500 connections.
- **Consumer Obligations**: Connection pooling via PgBouncer; queries bounded by statement timeouts (`statement_timeout = 30s`).
#### Managed Kafka Event Streaming Capability Contract [CC-KF-01]
- **Offered Outcome**: Partitioned topic streams with broker-level authentication (SASL_SSL / mTLS) and consumer group isolation.
- **Tiers & Guarantees**:
- `standard`: 99.9% broker availability, 7-day retention ceiling, max 20 MB/s ingress/egress per consumer, office-hours operational support.
- `critical`: 99.95% broker availability across 3 AZs, 30-day retention ceiling, max 100 MB/s ingress/egress per consumer, 24/7 SRE on-call support.
### 2. Tenancy, Isolation, and Quotas [TI-01]
- **Physical Isolation**: `critical` streaming tier operates on dedicated hardware broker clusters physically separated from `standard` and batch workloads.
- **Broker-Level Enforcement**: Throughput limits are programmatically enforced at the Kafka broker level via client-quota policies (`client-id` and `user` quotas), strictly rejecting excess traffic with `QuotaViolationException` rather than relying on consumer restraint.
- **Database Tenancy**: Separate virtual database instances and isolated tenant schemas; cross-squad access is blocked at network and database role boundaries.
### 3. Declarative Request and Lifecycle State Machine [LC-01]
- **Request Semantics**: Every API operation requires a unique `request_id`, capability version, target environment, and deterministic idempotency token to eliminate duplicate provisioning.
- **Supported Operations**:
- `provision`: Validates allocated squad quota, evaluates OPA security constraints, provisions underlying provider resources, and emits binding credentials.
- `tier_upgrade`: Requires capacity pre-check and financial cost center approval before rolling migration.
- `decommission`: Automated orphan detection; topics or database instances with zero connections/traffic over 90 days are flagged `idle` and safely archived after owner confirmation.
### 4. Quota and Cost Attribution Model [CM-01]
- **Attribution Driver**: Costs are charged back to squad cost centers based on *provisioned quota allocation* rather than transient actual usage.
- **Economic Rationale**: Charging purely by actual consumption on shared infrastructure perversely incentivizes uncoordinated burst behavior, directly causing the cluster contention observed in INC-5520.
Broker Bandwidth Saturation Invariant [INV-PLAT-01]
Shared infrastructure capabilities must programmatically enforce tenant quotas at the broker/engine layer.
No tenant request or batch replay may exceed provisioned throughput quotas regardless of available cluster headroom.
Zero Direct Console Provisioning Invariant [INV-PLAT-02]
All production infrastructure resources across the 110 microservices must be created, mutated, and
retired exclusively through platform control-plane APIs. Direct manual AWS console mutations are blocked by IAM SCPs.
Orphaned Resource Reclamation Invariant [INV-PLAT-03]
Shared capabilities must implement automated lifecycle tracking. Resources showing zero consumer activity
for 90 consecutive days must be marked for decommission to prevent unowned infrastructure cost sprawl.
## Ownership and Handoffs
| Concern | Owner | Handoff payload | Blocked until |
|---|---|---|---|
| Platform Architecture & Capability Contracts | Principal Platform Lead (Marcus Vance) | `infra_platform_capability_contract` | Head of Infra (Elena Rostova) review |
| Infrastructure Operations & SRE Support | Platform SRE Operations | `infra_platform_service_contract` | 24/7 on-call shift rotation schedule |
| Broker & Cloud Security Governance | CISO SecOps (Sarah Chen) | KMS policy & IAM least-privilege contracts | Security audit sign-off |
| Squad Infrastructure Adoption | 45 Feature Squad Leads | Declarative workload manifest schemas | Control-plane API v1 deployment |
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 600 engineers across 45 product squads | provided | Organizational intake | Current |
| 110 active microservices | provided | Workload intake | Current |
| 250 self-service operations/week | provided | Scale intake | Current |
| 2 core capabilities: PostgreSQL & Kafka | provided | Platform scope intake | Current |
| Incident INC-5520 replay outage | provided | Incident post-mortem report | Current |
| Rejection of direct AWS console/IAM access | decided | Marcus Vance & Elena Rostova | 2026-09-15 |
| Broker-enforced client quotas (CC-KF-01) | decided | Architectural invariant INV-PLAT-01 | 2026-09-15 |
| Provisioned quota-based chargeback (CM-01) | decided | Financial governance model | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against platform architecture fitness criteria:
- **Snowflake Infrastructure Probe**: PASS. All capabilities provisioned via declarative versioned schemas; manual snowflake resources strictly rejected by root SCPs.
- **Process-Equals-Readiness Probe**: PASS. Platform does not infer readiness from ticket completion; broker-level health probes and verified credential bindings gate usable state.
- **Unowned Shared Platform Probe**: PASS. Explicit squad ownership tagging, orphan detection (90-day idle rule), and quota chargeback ensure zero unowned shared infrastructure.
- **Markdown Conformance**: PASS. Follows native Markdown rules from `rule_markdown.md`.
## Open Decisions
- `DEC-PLAT-01`: Elena Rostova to determine whether Strimzi Kafka Operator on EKS or AWS Managed Streaming for Kafka (MSK) will be standardized as the execution backend for `CC-KF-01` (Owner: Elena Rostova).
## Next steps
1. Marcus Vance publishes OpenAPI and JSON schema specifications for `CC-PG-01` and `CC-KF-01` capability contracts.
2. Platform SRE team deploys Kafka broker client-quota enforcement policies in staging to validate rejection semantics.
3. Conduct pilot integration with 2 high-volume payment squads to validate self-service topic lifecycle operations.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the shared control-plane architecture that exposes infrastructure capabilities as supported products to multiple consumer teams and delegates effects to canonical cloud, cluster, network, storage, IaC, GitOps, and security systems. It defines capability contracts, tenancy, request/lifecycle semantics, orchestration, support, adoption, and recovery without owning every implementation or the platform-team operating model.
Use it when
- Multiple consumer teams need a coherent catalog of shared infrastructure capabilities
- Provider and consumer responsibilities, decision rights, support, and escape paths are ambiguous
- Requests must map to versioned capability contracts and provider-owned execution backends
- Tenant/project/environment identity and isolation span control, execution, resource, data, and evidence planes
- Asynchronous provisioning/update/delete operations need stable request, operation, resource, and effect identities
- Policy, approval, exception, quota, capacity, cost, and lifecycle authorities must remain distinct
For example: “Our shared Kafka is used by everyone. One team's replay consumed the whole cluster, and nobody could tell me what we actually promise anyone about it.”
What you get
- architecture/infra-platform-architect/README.md
- architecture/infra-platform-architect/00-overview/infra-platform-architect-overview.md
- architecture/infra-platform-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/topology.md, {module}/provisioning.md, {module}/networking.md, {module}/secrets.md, {module}/cost.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to provision one resource, write IaC, install Kubernetes/GitOps, choose cloud topology, build a developer portal, define a golden path, or reorganize a platform team.
How it works
- Check the subject is the shared capability itself.
- Define the capability catalogue and the contract per capability.
- Fix tenancy and isolation.
- State the request and lifecycle path per capability.
- Set the capacity and cost model.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-diagram.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install