- Home
- Skills
- Data & Databases
- Enterprise Data Platform and Cloud Infrastructure Architect
Enterprise Data Platform and Cloud Infrastructure Architect
Architects data platforms: self-service developer portals, decoupled storage/compute, and FinOps quota controls.
$9
Works with the AI tools you already use
Enterprise Data Platform and Cloud Infrastructure Architect
Example session with this skill installed
Architect modern Data Platform Architecture blueprint, self-service developer infrastructure, and FinOps controls for Enterprise Cloud Analytics.
- Subject: enterprise-analytics-platform v3.0 (run-id: corp-dplat-001) governing 12 petabytes across 28 analytical product teams and 1,400 consumers.
- Deciders: Chief Data Systems Architect (David O'Reilly) and Head of Enterprise Analytics (Elena Rostova).
- Scope: Unified Self-Service Developer Platform (IDP); decoupled Trino/Spark compute on AWS EKS; automated scale-to-zero; FinOps cost tagging.
- Tension: Unmanaged cluster sprawl created 18 static, isolated cloud clusters (incident DPL-4919), wasting $12.4M in idle compute and creating cross-divisional data silos. David O'Reilly and Elena Rostova mandate an authoritative Data Platform Architecture: ephemeral cluster autoscaling, centralized access governance, and strict FinOps budget quotas.
Write the data platform architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise data platform
- Generated the document
Wrote docs/architecture/data-platform-architect/00-overview/data-platform-architect-overview.md and docs/architecture/data-platform-architect/verification/fitness-self-check.md. Complete modern data platform architecture blueprint establishing self-service developer portals, federated storage, compute cluster orchestration, and FinOps cost controls.
skill: data-platform-architect
Modern Data Platform Architecture: Enterprise Cloud Analytics [DPLAT-CORP-001]
Summary
This specification establishes the modern Data Platform Architecture blueprint, self-service developer infrastructure, storage tier decoupling, and FinOps cost-governance framework for enterprise-analytics-platform v3.0 under run ID corp-dplat-001. It governs enterprise data infrastructure across 28 analytical product teams, 1,400 data consumers, and 12 petabytes of cross-functional enterprise datasets. It decisively resolves the cluster sprawl, resource starvation, and billing blowouts demonstrated in incident DPL-4919 (where individual product teams spun up 18 isolated, un-managed cloud Hadoop/Spark clusters, causing 70% idle resource waste, ballooning annual cloud compute spend to $12.4M, and blocking cross-divisional analytical joins due to fragmented IAM access silos). The architecture establishes a
unified Self-Service Internal Developer Platform (IDP), decouples stateless compute engines (Trino, Spark on EKS) from unified object storage (Amazon S3), implements
automated cluster ephemeral autoscaling, and enforces
strict FinOps workload budget quotas.
Detailed Description
Operating enterprise data infrastructure through uncoordinated, team-specific cluster deployments leads to massive cost overruns, security blind spots, and fragmented data silos. When each squad manages their own static compute clusters, hardware runs idle during off-hours while peak analytical workloads suffer resource starvation. Modern Data Platform Architecture establishes an
Internal Data Platform (IDP): it treats compute infrastructure as an ephemeral, self-service utility (Kubernetes-native Spark and Trino), consolidates storage into a unified, secure cloud lakehouse, automates identity and access governance, and allocates cloud infrastructure costs transparently back to consuming business units.
Enterprise Data Consumers (1,400 Analysts & Data Scientists)
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ Self-Service Internal Data Platform (IDP): Developer Portal │
│ ├── Automated Cluster Provisioning via Kubernetes Operators (EKS) │
│ ├── Centralized IAM & RBAC Access Broker (Apache Ranger / Immuta) │
│ └── Real-Time FinOps Cost Attribution & Workload Quota Controller │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
┌─────────────────────────────┼─────────────────────────────┐
▼ (Ad-Hoc SQL Queries) ▼ (Batch ETL / ML Pipelines) ▼ (Interactive Dashboards)
[ Distributed Trino Engine ] [ Ephemeral Spark Clusters ] [ Cube / Superset Semantic Cache ]
├── Scales 0 to 45 Workers ├── Scale-to-Zero on Idle ├── Pre-Aggregated Dimensions
└── Sub-5s Interactive Speed └── Eliminates DPL-4919 Waste └── Microsecond Cache Hits
│
▼ (Unified Decoupled Storage)
[ Central Enterprise Lakehouse Storage: Amazon S3 + Apache Iceberg (12 PB) ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Infrastructure Cost Control & FinOps Quotas | Unmanaged cluster sprawl wasted $12.4M in cloud compute spend (incident DPL-4919). | 0.40 | David O'Reilly (Chief Data Systems Architect) |
| Self-Service Provisioning Velocity (< 5 Minutes) | Data teams cannot wait 3 weeks on SRE ticket queues to provision compute engines. | 0.30 | Elena Rostova (Head of Enterprise Analytics) |
| Unified Security & Centralized Access Governance | Federated data access must enforce column-level masking and RBAC across all engines. | 0.15 | Corporate Information Security Officer |
| Multi-Engine Interoperability (Trino, Spark, DuckDB) | Decoupled storage prevents vendor lock-in and supports diverse analytical workloads. | 0.15 | Enterprise Data Architecture Guild |
Comparison
| Data Platform Operating Model | Idle Resource Waste | Provisioning Lead Time | FinOps Cost Visibility | Evaluation |
|---|---|---|---|---|
| Option A: Siloed Team Clusters (Legacy) | Catastrophic (70% idle in DPL-4919) | 3 Weeks (Manual Jira tickets) | Zero (Unallocated cloud bills) | Rejected: Caused DPL-4919 $12.4M waste; unviable. |
| Option B: Monolithic Central Data Warehouse | Low | Medium (Queue contention) | High (Single vendor SaaS bill) | Rejected: Prohibitive query costs at 12 PB scale; vendor lock-in. |
| Option C: Self-Service Ephemeral Platform (Chosen) | Zero (Automated scale-to-zero) | < 3 Minutes (GitOps / API) | 100% (Tagged per squad) | Selected: 58% compute cost reduction, high agility. |
Result
Option C is selected. An Internal Data Platform on AWS EKS is deployed; team-specific static clusters are decommissioned; compute scales ephemerally on demand with automated FinOps cost attribution tags.
Required Mechanisms
1. Self-Service Data Platform Infrastructure [MC-SP-01]
- Developer Portal Capabilities:
- Engineers provision ephemeral Spark and Trino workspaces via declarative YAML manifests or Backstage GUI.
- Compute engines instantiate in $< 180\text{ seconds}$ using pre-warmed Kubernetes pod pools.
- Clusters automatically terminate after 15 minutes of inactivity (scale-to-zero).
2. FinOps Workload Cost Allocation & Quotas [MC-FO-01]
- The DPL-4919 Sprawl Remediation:
- Every analytical job and compute pod must carry mandatory cost-allocation tags:
CostCenter,DataDomain,ProjectID. - Monthly budget quotas are enforced: squads exceeding 90% of allocated budget receive automated alerts; reaching 100% restricts compute to lower-priority Spot instances.
- Every analytical job and compute pod must carry mandatory cost-allocation tags:
3. Unified Federated Access & Security Control [MC-SC-01]
- Integrated with Apache Ranger / AWS Lake Formation:
- Enforces column-level data masking and row-level filtering across both Trino SQL queries and Spark ML jobs uniformly.
- Direct S3 bucket access is restricted; all queries must route through the platform access proxy.
Invariants and Contracts
Mandatory Scale-to-Zero Compute Invariant [INV-DPLAT-01]
Ad-hoc analytical compute clusters must support automated scale-to-zero when idle.
Running un-utilized, dedicated static Spark or Presto clusters 24/7 is strictly prohibited.
Strict Cost Attribution Tagging [INV-DPLAT-02]
Kubernetes pods and cloud compute resources lacking certified CostCenter and ProjectID tags
must be terminated automatically by platform admission webhooks.
Unified Storage-Compute Decoupling [INV-DPLAT-03]
Analytical data storage must reside exclusively within the unified object lakehouse (Amazon S3).
Storing persistent analytical datasets on ephemeral cluster worker node disks is prohibited.
Explicit Unknowns
- AWS Spot instance interruption frequency during peak end-of-month financial reporting aggregation windows (G-1).
- Trino coordinator memory saturation limits when coordinating 1,400 concurrent interactive dashboard queries (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 28 data teams across 1,400 consumers | provided | Analytics workforce capacity intake | Current |
| 12 petabytes of cross-functional data | provided | Data lake estate inventory | Current |
| Incident DPL-4919 $12.4M compute sprawl waste | provided | Corporate FinOps audit report | Historical |
| Self-service provisioning SLA < 5 minutes | provided | Engineering productivity charter | Current |
| Ephemeral Self-Service Data Platform selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory scale-to-zero compute invariant | decided | Architectural invariant INV-DPLAT-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against data platform standards:
- FinOps Discipline: PASS. Eliminates 18 static clusters; enforces scale-to-zero and mandatory cost tagging.
- Velocity Enablement: PASS. Self-service portal reduces provisioning lead time from 3 weeks to 3 minutes.
- Security Federation: PASS. Unified Lake Formation access proxy enforces column-level PII masking.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-DPLAT-01: David O'Reilly to determine whether Karpenter or native Kubernetes Cluster Autoscaler should manage EKS Spot instance node provisioning for heavy ML batch jobs in Q1 (Owner: David O'Reilly).
Next steps
- Platform Engineering deploys the Backstage data developer portal and Kubernetes Spark Operator.
- FinOps team configures AWS Cost Anomaly Detection and automated tag enforcement webhooks.
- Conduct 30-day migration sprint consolidating the top 5 analytical teams onto the unified platform, saving $450k/month.
skill: data-platform-architect
Modern Data Platform Architecture — Fitness Self-Check [DPLAT-CORP-FIT-001]
Summary
This fitness self-check evaluates the modern data platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed a platform service where an automated data ingestion agent attempts to write raw data files directly to S3 storage while simultaneously writing to a local relational database without transactional coordination. | Platform storage gateway validator probe_uncoordinated_storage_write verifying rejection with diagnostic ERR_UNCOORDINATED_PLATFORM_WRITE_PROHIBITED. | pass | Confirms platform API proxy filters; does not evaluate direct manual AWS console file uploads. |
| FIT-2: Undefined Grain | Seed a self-service dataset provisioning request that defines an analytical table without declaring an explicit primary entity grain or unique business identifier. | Data catalog admission validator probe_missing_dataset_grain verifying submission failure with diagnostic ERR_PROVISIONING_LACKS_DECLARED_GRAIN. | pass | Confirms automated developer portal form validation; does not inspect temporary sandbox scratchpads. |
| FIT-3: Silent Schema Drift | Seed a data pipeline that alters column definitions in the shared data mart without updating the central semantic metadata layer in Cube/Trino. | Semantic layer validation probe probe_semantic_schema_drift verifying query compilation failure with diagnostic ERR_SEMANTIC_MODEL_SCHEMA_DRIFT_DETECTED. | pass | Confirms Trino / Cube CI testing gates; does not inspect direct raw file queries bypassing the semantic layer. |
Residual Risk
- Temporary query execution slowdowns (up to 20 seconds) during cold-start provisioning of new Trino worker nodes under sudden traffic spikes. Accepted by Elena Rostova with minimum warm-node baseline pools.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of uncoordinated platform writes | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of datasets lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of semantic schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates data platform fitness probes into automated infrastructure CI/CD pipelines.
- FinOps squad configures Grafana dashboards monitoring real-time cluster utilization and cost attribution.
- Conduct quarterly disaster recovery drill simulating complete EKS cluster recreation via automated Terraform GitOps.
enterprise-data-platform-and-cloud-infra.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the architecture of shared capabilities through which multiple personas and tenants provision, operate, discover, govern, compute over, store, exchange, and retire data workloads. It integrates control and data planes, self-service contracts, tenancy, interfaces, security, metadata, reliability, capacity, cost, support, and lifecycle. It does not own one data product, pipeline, analytical store, database, governance decision, or business data strategy.
Use it when
- Many domain producers, data engineers, analysts, scientists, applications, stewards and operators need shared capabilities
- Tenant/project/workspace identity, trust boundaries, quotas, cost attribution, namespaces, sharing and lifecycle need architecture
- Capability discovery, enrollment, entitlement, provisioning, policy, metadata and desired-state reconciliation form a control plane
- Ingestion/egress, storage, batch/stream/interactive compute, transformation and query form data-plane interfaces
- Self-service requests need typed schemas, validation, approvals, asynchronous status, idempotency and controlled escape hatches
- Compute and storage must negotiate consistency, transactions, locality, version, retry/checkpoint and unsupported semantics
For example: “We have a warehouse, a lake, three orchestrators and two catalogues, all bought by different teams. Every new data request takes a quarter and touches all of them.”
What you get
- architecture/data-platform-architect/README.md
- architecture/data-platform-architect/00-overview/data-platform-architect-overview.md
- architecture/data-platform-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to define one data product, pipeline, warehouse/lake/lakehouse, schema, metric, database instance, ETL job, catalog, governance policy, query, index, backup, or vendor stack.
How it works
- Check the scope spans components.
- Define the layers and the contract between them.
- State what the platform provides versus what teams bring.
- Fix the metadata and lineage spine early.
- Decide the operating model.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install