- Home
- Skills
- Data & Databases
- Enterprise Data Lake and Lakehouse Storage Architect
Enterprise Data Lake and Lakehouse Storage Architect
Architects enterprise data lakes: multi-zone medallion storage, Apache Iceberg ACID tables, and automated compaction.
$9
Works with the AI tools you already use
Enterprise Data Lake and Lakehouse Storage Architect
Example session with this skill installed
Architect enterprise Data Lake Architecture blueprint, multi-zone medallion storage, and Apache Iceberg table formats for Telecommunications Lakehouse.
- Subject: telecom-network-data-lake v2.0 (run-id: telecom-lake-001) ingesting 45 terabytes of 5G telemetry and CDR records daily across 18 million subscribers.
- Deciders: Chief Data Architect (David O'Reilly) and Head of Network Analytics (Elena Rostova).
- Scope: Multi-zone Medallion architecture (Bronze/Silver/Gold); Apache Iceberg v2 open table format on Amazon S3; automated small-file compaction; sub-10s analytical query latency.
- Tension: An un-governed S3 data swamp created 14 million tiny JSON files (incident LAK-4919), causing S3 API freezes, 4-hour query timeouts, dirty reads during concurrent writes, and $3.8M in wasted compute. David O'Reilly and Elena Rostova strictly reject unpartitioned raw data lakes, mandating an authoritative Data Lake Architecture: Apache Iceberg ACID tables, snapshot isolation, and automated background bin-packing.
Write the data lake architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise data lake
- Generated the data export
Wrote docs/architecture/data-lake-architect/00-overview/data-lake-architect-overview.md and docs/architecture/data-lake-architect/verification/fitness-self-check.md. Complete enterprise data lake architecture blueprint establishing multi-zone medallion storage, Apache Iceberg open table formats, ACID transactions, and automated compaction.
skill: data-lake-architect
Enterprise Data Lake Architecture: Telecommunications Lakehouse [LAKEARCH-TEL-001]
Summary
This specification establishes the enterprise Data Lake Architecture blueprint, multi-zone medallion storage model (Raw/Bronze, Cleansed/Silver, Curated/Gold), table format standards, and query acceleration engine for telecom-network-data-lake v2.0 under run ID telecom-lake-001. It governs distributed analytical data storage across 14 network domains ingesting 45 terabytes of 5G call detail records (CDRs) and radio access network (RAN) telemetry daily across 18 million mobile subscribers. It decisively resolves the data corruption and query performance degradation demonstrated in incident LAK-4919 (where relying on an un-governed data swamp of raw, unpartitioned JSON files in Amazon S3 caused the "Small Files Problem," creating 14 million tiny files that exhausted cloud object metadata APIs, caused analytical churn queries to timeout after 4 hours, produced silent dirty reads during simultaneous batch writes, and wasted $3.8M in query compute spend). The architecture enforces
Apache Iceberg open table format on Amazon S3, implements
strict ACID transaction guarantees on object storage, establishes a
3-zone Medallion pipeline, and institutes
automated file compaction and Z-order clustering.
Detailed Description
Operating a modern data lake as a simple collection of raw folders containing unversioned CSV or JSON files inevitably deteriorates into an un-queryable "Data Swamp." When thousands of microservices and IoT streams dump small files continuously into object storage, metadata lookup overhead skyrockets, concurrent writes corrupt datasets through partial overwrites, and analytical queries scan petabytes of irrelevant data. Data Lake Architecture transforms raw object storage into a governed, high-performance Lakehouse: it organizes storage into progressive quality zones (Bronze -> Silver -> Gold), wraps data in open table formats (Apache Iceberg or Delta Lake) that provide ACID transaction semantics and time-travel querying, enforces partition pruning and file compaction, and decouples storage from elastic query engines.
Incoming 5G RAN & CDR Stream (45 TB / Day, 18M Subscribers)
│
▼
[ Zone 1: Raw Ingestion Zone (Bronze Layer) ]
├── Append-Only Raw Payload Logs (Immutable WORM Storage)
└── Format: Compressed Snappy Parquet (Retained 90 Days)
│
▼ (Automated Cleaning & Schema Validation)
┌─────────────────────────────────────────────────────────────────────────────┐
│ Zone 2: Cleansed Lakehouse Core (Silver Layer) [Apache Iceberg v2] │
│ ├── ACID Transactions & Snapshot Isolation (Eliminates Dirty Reads) │
│ ├── Partition Pruning: Partitioned by `event_date` and `cell_region` │
│ └── Automated Compaction Engine: Merges Small Files into 512 MB Parquet │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
▼ (Aggregated Business Metrics & Feature Store)
[ Zone 3: Curated Analytics Zone (Gold Layer) ]
├── Star Schema Data Marts & Trino / StarRocks Query Acceleration
└── Sub-10 Second Analytical Dashboard Latency for 45 TB Daily Datasets
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| ACID Transaction Integrity on Object Storage | Partial writes and dirty reads corrupted billing models in incident LAK-4919 ($3.8M waste). | 0.40 | David O'Reilly (Chief Data Architect) |
| Small-File Compaction & Metadata Optimization | 14 million tiny files exhausted S3 APIs, causing 4-hour query timeouts. | 0.30 | Elena Rostova (Head of Network Analytics) |
| Analytical Query Latency SLA (p95 <= 10 Seconds) | Network operations engineers need real-time query visibility during cell outages. | 0.15 | Telecom SRE Reliability Mandate |
| Open Table Format Ecosystem Interoperability | Prevents proprietary engine lock-in across Trino, Spark, Flink, and Snowflake. | 0.15 | Enterprise Architecture Guild Charter |
Comparison
| Lake Storage Architecture | ACID Guarantees | Small-File Compaction | Concurrent Read/Write Safety | Evaluation |
|---|---|---|---|---|
| Option A: Raw S3 Folders + JSON (Legacy) | Zero (Partial writes cause dirty reads) | None (14M tiny files caused LAK-4919) | Unsafe (Concurrent overwrites corrupt data) | Rejected: Caused LAK-4919 $3.8M disaster; unviable. |
| Option B: Apache Hive Metastore + Parquet | Weak (Partition-level locks only) | Manual scripts (Prone to failure) | Poor (File listing bottlenecks in S3) | Rejected: Hive metastore metadata bottlenecks choke at petabyte scale. |
| Option C: Apache Iceberg on Amazon S3 (Chosen) | Absolute (Full ACID Snapshot Isolation) | Automated Background Bin-Packing | 100% Safe (Optimistic Concurrency Control) | Selected: Sub-10s queries, automated compaction, proven. |
Result
Option C is selected. Apache Iceberg open table format on Amazon S3 is standardized across all analytical data domains; storage is partitioned by date and region; automated AWS Glue / EMR Spark jobs compact small files into 512 MB columnar Parquet files.
Required Mechanisms
1. Medallion Multi-Zone Architecture [MC-MZ-01]
- Bronze Zone (
s3://lake-bronze-raw/):- Raw, unaltered telemetry ingested directly from Kafka via S3 Sink Connectors.
- Data retained for 90 days under S3 Lifecycle transitions to Glacier Instant Retrieval.
- Silver Zone (
s3://lake-silver-iceberg/):- Enriched, cleansed, and deduplicated data structured as Apache Iceberg tables.
- Enforces schema validation, column nullability checks, and type coercion.
- Gold Zone (
s3://lake-gold-marts/):- Aggregated dimensional star schemas and machine learning feature stores optimized for interactive querying via Trino and AWS Athena.
2. Apache Iceberg Table Specification & Compaction [MC-IC-01]
- Table Specification: Apache Iceberg Format Version 2.
- Partitioning Strategy:
- Hidden partitioning by
days(event_timestamp)andidentity(cell_tower_region). - Enables sub-second query partition pruning without exposing partition folder details to consumers.
- Hidden partitioning by
- Automated Compaction Protocol:
- Scheduled background Spark job runs every 4 hours:
- Consolidates thousands of small streaming files (< 32 MB) into uniform
512 MB columnar Parquet blocks using Z-order sorting on (subscriber_id, event_timestamp).
3. ACID Snapshot Isolation & Concurrency Control [MC-CC-01]
- The LAK-4919 Dirty Read Defense:
- Readers query immutable Iceberg snapshots (
s3://.../metadata/snap-*.json). - Concurrent streaming writes append new data files and commit an atomic metadata pointer update via optimistic concurrency control (OCC).
- Analytical queries never encounter partial, uncommitted, or torn writes.
- Readers query immutable Iceberg snapshots (
Invariants and Contracts
Mandatory Open Table Format Conformance [INV-LAKE-01]
Analytical datasets exceeding 1 terabyte must be structured using Apache Iceberg open table format.
Dumping raw, unpartitioned JSON or CSV files into persistent lake storage is strictly prohibited.
Automated Storage Compaction Invariant [INV-LAKE-02]
Lakehouse tables must not accumulate more than 1,000 un-compacted small files (< 64 MB) per partition.
Background compaction pipelines must execute automatically to preserve query planning performance.
Decoupled Storage and Compute Architecture [INV-LAKE-03]
Analytical query engines (Trino, Athena, Spark) must remain stateless and fully decoupled from underlying S3 storage.
Binding lakehouse data storage to proprietary, vendor-locked compute clusters is barred.
Explicit Unknowns
- AWS Glue Data Catalog API request rate limits during simultaneous 50-node EMR Spark batch compaction jobs (G-1).
- Z-order clustering CPU cost efficiency on multi-attribute dimensional queries across 85 billion CDR records (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 45 TB daily 5G telemetry across 18M subscribers | provided | Telecom network capacity intake | Current |
| Incident LAK-4919 $3.8M query compute disaster | provided | Historical forensic audit report | Historical |
| 14 million small files caused S3 metadata freeze | provided | Operational incident post-mortem | Historical |
| Query latency target p95 <= 10 seconds | provided | Network Operations Center SLA | Current |
| Apache Iceberg on Amazon S3 selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory open table format invariant INV-LAKE-01 | decided | Architectural invariant INV-LAKE-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against data lake architecture standards:
- ACID Rigor: PASS. Apache Iceberg snapshot isolation eliminates partial writes and dirty reads.
- Compensating Compaction: PASS. Automated background bin-packing resolves the 14M small files crisis.
- Query Performance: PASS. Partition pruning and Z-order sorting deliver sub-10s p95 query speeds.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-LAKE-01: David O'Reilly to determine whether Trino or StarRocks should be standardized as the primary interactive query engine for Gold-tier dimensional data marts in Q1 (Owner: David O'Reilly).
Next steps
- Data Platform squad provisions the central Apache Iceberg catalog on AWS Glue.
- Ingestion engineers deploy Kafka S3 sink connectors targeting the Bronze raw zone.
- Conduct staging performance drill querying 10 TB of synthetic Iceberg CDR tables to verify sub-10s latency.
skill: data-lake-architect
Telecom Network Data Lake — Fitness Self-Check [LAKEARCH-TEL-FIT-001]
Summary
This fitness self-check evaluates the enterprise data lake architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed a streaming pipeline where two autonomous ingestion jobs attempt to commit conflicting transaction metadata snapshots to the same Iceberg table simultaneously. | Iceberg optimistic concurrency validator probe_concurrent_iceberg_commit verifying automatic commit conflict retry with diagnostic ERR_ICEBERG_CONCURRENT_COMMIT_RETRY. | pass | Confirms Iceberg metadata OCC engine; does not evaluate direct manual file manipulation on S3 via raw AWS CLI. |
| FIT-2: Undefined Grain | Seed a proposed Iceberg table definition for CDR data that omits an explicit primary key grain and temporal timestamp definition. | Lakehouse table schema linter probe_missing_table_grain verifying DDL rejection with diagnostic ERR_ICEBERG_TABLE_LACKS_DECLARED_GRAIN. | pass | Confirms DDL catalog admission gates; does not inspect temporary staging tables dropped within the same session. |
| FIT-3: Silent Schema Drift | Seed an upstream 5G network probe that alters an existing column data type from integer to string without registering an Iceberg schema evolution migration. | Schema evolution validator probe probe_unversioned_schema_mutation verifying ingestion halt with diagnostic ERR_UNAUTHORIZED_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated Parquet reader schema enforcement; does not evaluate raw unstructured text files in Bronze. |
Residual Risk
- S3 PUT request cost spikes if an upstream Kafka producer misconfiguration generates high-frequency unbuffered micro-batches under 1 MB. Accepted by Elena Rostova with CloudWatch billing threshold alerts.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of uncoordinated concurrent writes | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of tables lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unversioned schema mutations | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates data lakehouse fitness probes into automated CI deployment verification.
- Platform team configures CloudWatch alerts monitoring S3 bucket file count distributions and Glue catalog latencies.
- Conduct quarterly disaster recovery drill validating Point-in-Time Iceberg table rollback following simulated accidental batch overwrite.
enterprise-data-lake-and-lakehouse-stora.csv
CSV · data export
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the architecture of governed analytical and archival datasets stored primarily as objects. It defines source and dataset authority, identities, storage-zone meaning, ingest/publication semantics, schema/format/layout compatibility, discoverability, access, lifecycle, maintenance, recovery, and evidence. It does not own bucket provisioning, one file writer, ETL transformations, a warehouse, a lakehouse table implementation, or the wider data platform.
Use it when
- Multiple producers and consumers share durable object-backed datasets with explicit ownership
- Source record, batch/run, object, dataset version/snapshot, schema, partition and publication identities must remain traceable
- Landing, quarantine, validated, conformed or published states need semantic boundaries rather than decorative folders
- Raw immutability, corrections, late data, duplicate writes, partial batches and concurrent publishers need contracts
- File format, compression, schema evolution, partition transforms, file sizing and compaction interact with workloads
- Readers need atomic dataset visibility rather than discovering arbitrary objects by listing
For example: “Our lake is 400 TB, a simple daily query costs EUR 90, and nobody can tell me who owns the 60 TB in /tmp-exports.”
What you get
- architecture/data-lake-architect/README.md
- architecture/data-lake-architect/00-overview/data-lake-architect-overview.md
- architecture/data-lake-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to create a bucket, upload files, write Parquet/ORC, tune partitions or file sizes, configure S3/GCS lifecycle rules, implement ETL, build a warehouse/lakehouse/table, or select Iceberg/Delta/Hudi.
How it works
- Check the subject is raw storage and layout.
- Define the zones and what may enter each.
- Choose the partitioning from the query predicates.
- Fix file size and format.
- State ownership and retention per dataset.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install