Enterprise Data Lake and Lakehouse Storage Architect

    1

    Architects enterprise data lakes: multi-zone medallion storage, Apache Iceberg ACID tables, and automated compaction.

    $9

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Enterprise Data Lake and Lakehouse Storage Architect

    Example session with this skill installed

    Architect enterprise Data Lake Architecture blueprint, multi-zone medallion storage, and Apache Iceberg table formats for Telecommunications Lakehouse.

    • Subject: telecom-network-data-lake v2.0 (run-id: telecom-lake-001) ingesting 45 terabytes of 5G telemetry and CDR records daily across 18 million subscribers.
    • Deciders: Chief Data Architect (David O'Reilly) and Head of Network Analytics (Elena Rostova).
    • Scope: Multi-zone Medallion architecture (Bronze/Silver/Gold); Apache Iceberg v2 open table format on Amazon S3; automated small-file compaction; sub-10s analytical query latency.
    • Tension: An un-governed S3 data swamp created 14 million tiny JSON files (incident LAK-4919), causing S3 API freezes, 4-hour query timeouts, dirty reads during concurrent writes, and $3.8M in wasted compute. David O'Reilly and Elena Rostova strictly reject unpartitioned raw data lakes, mandating an authoritative Data Lake Architecture: Apache Iceberg ACID tables, snapshot isolation, and automated background bin-packing.

    Write the data lake architecture overview under docs/.

    • Read your context and instructions
    • Compiled the enterprise data lake
    • Generated the data export

    Wrote docs/architecture/data-lake-architect/00-overview/data-lake-architect-overview.md and docs/architecture/data-lake-architect/verification/fitness-self-check.md. Complete enterprise data lake architecture blueprint establishing multi-zone medallion storage, Apache Iceberg open table formats, ACID transactions, and automated compaction.


    skill: data-lake-architect

    Enterprise Data Lake Architecture: Telecommunications Lakehouse [LAKEARCH-TEL-001]

    Summary

    This specification establishes the enterprise Data Lake Architecture blueprint, multi-zone medallion storage model (Raw/Bronze, Cleansed/Silver, Curated/Gold), table format standards, and query acceleration engine for telecom-network-data-lake v2.0 under run ID telecom-lake-001. It governs distributed analytical data storage across 14 network domains ingesting 45 terabytes of 5G call detail records (CDRs) and radio access network (RAN) telemetry daily across 18 million mobile subscribers. It decisively resolves the data corruption and query performance degradation demonstrated in incident LAK-4919 (where relying on an un-governed data swamp of raw, unpartitioned JSON files in Amazon S3 caused the "Small Files Problem," creating 14 million tiny files that exhausted cloud object metadata APIs, caused analytical churn queries to timeout after 4 hours, produced silent dirty reads during simultaneous batch writes, and wasted $3.8M in query compute spend). The architecture enforces

    Apache Iceberg open table format on Amazon S3, implements

    strict ACID transaction guarantees on object storage, establishes a

    3-zone Medallion pipeline, and institutes

    automated file compaction and Z-order clustering.

    Detailed Description

    Operating a modern data lake as a simple collection of raw folders containing unversioned CSV or JSON files inevitably deteriorates into an un-queryable "Data Swamp." When thousands of microservices and IoT streams dump small files continuously into object storage, metadata lookup overhead skyrockets, concurrent writes corrupt datasets through partial overwrites, and analytical queries scan petabytes of irrelevant data. Data Lake Architecture transforms raw object storage into a governed, high-performance Lakehouse: it organizes storage into progressive quality zones (Bronze -> Silver -> Gold), wraps data in open table formats (Apache Iceberg or Delta Lake) that provide ACID transaction semantics and time-travel querying, enforces partition pruning and file compaction, and decouples storage from elastic query engines.

    Incoming 5G RAN & CDR Stream (45 TB / Day, 18M Subscribers)
                             │
                             ▼
    [ Zone 1: Raw Ingestion Zone (Bronze Layer) ]
      ├── Append-Only Raw Payload Logs (Immutable WORM Storage)
      └── Format: Compressed Snappy Parquet (Retained 90 Days)
                             │
                             ▼ (Automated Cleaning & Schema Validation)
    ┌─────────────────────────────────────────────────────────────────────────────┐
    │ Zone 2: Cleansed Lakehouse Core (Silver Layer) [Apache Iceberg v2]          │
    │   ├── ACID Transactions & Snapshot Isolation (Eliminates Dirty Reads)       │
    │   ├── Partition Pruning: Partitioned by `event_date` and `cell_region`      │
    │   └── Automated Compaction Engine: Merges Small Files into 512 MB Parquet   │
    └──────────────────────────────────────┬──────────────────────────────────────┘
                                           │
                             ▼ (Aggregated Business Metrics & Feature Store)
    [ Zone 3: Curated Analytics Zone (Gold Layer) ]
      ├── Star Schema Data Marts & Trino / StarRocks Query Acceleration
      └── Sub-10 Second Analytical Dashboard Latency for 45 TB Daily Datasets
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    ACID Transaction Integrity on Object StoragePartial writes and dirty reads corrupted billing models in incident LAK-4919 ($3.8M waste).0.40David O'Reilly (Chief Data Architect)
    Small-File Compaction & Metadata Optimization14 million tiny files exhausted S3 APIs, causing 4-hour query timeouts.0.30Elena Rostova (Head of Network Analytics)
    Analytical Query Latency SLA (p95 <= 10 Seconds)Network operations engineers need real-time query visibility during cell outages.0.15Telecom SRE Reliability Mandate
    Open Table Format Ecosystem InteroperabilityPrevents proprietary engine lock-in across Trino, Spark, Flink, and Snowflake.0.15Enterprise Architecture Guild Charter

    Comparison

    Lake Storage ArchitectureACID GuaranteesSmall-File CompactionConcurrent Read/Write SafetyEvaluation
    Option A: Raw S3 Folders + JSON (Legacy)Zero (Partial writes cause dirty reads)None (14M tiny files caused LAK-4919)Unsafe (Concurrent overwrites corrupt data)Rejected: Caused LAK-4919 $3.8M disaster; unviable.
    Option B: Apache Hive Metastore + ParquetWeak (Partition-level locks only)Manual scripts (Prone to failure)Poor (File listing bottlenecks in S3)Rejected: Hive metastore metadata bottlenecks choke at petabyte scale.
    Option C: Apache Iceberg on Amazon S3 (Chosen)Absolute (Full ACID Snapshot Isolation)Automated Background Bin-Packing100% Safe (Optimistic Concurrency Control)Selected: Sub-10s queries, automated compaction, proven.

    Result

    Option C is selected. Apache Iceberg open table format on Amazon S3 is standardized across all analytical data domains; storage is partitioned by date and region; automated AWS Glue / EMR Spark jobs compact small files into 512 MB columnar Parquet files.


    Required Mechanisms

    1. Medallion Multi-Zone Architecture [MC-MZ-01]
    • Bronze Zone (s3://lake-bronze-raw/):
      • Raw, unaltered telemetry ingested directly from Kafka via S3 Sink Connectors.
      • Data retained for 90 days under S3 Lifecycle transitions to Glacier Instant Retrieval.
    • Silver Zone (s3://lake-silver-iceberg/):
      • Enriched, cleansed, and deduplicated data structured as Apache Iceberg tables.
      • Enforces schema validation, column nullability checks, and type coercion.
    • Gold Zone (s3://lake-gold-marts/):
      • Aggregated dimensional star schemas and machine learning feature stores optimized for interactive querying via Trino and AWS Athena.
    2. Apache Iceberg Table Specification & Compaction [MC-IC-01]
    • Table Specification: Apache Iceberg Format Version 2.
    • Partitioning Strategy:
      • Hidden partitioning by days(event_timestamp) and identity(cell_tower_region).
      • Enables sub-second query partition pruning without exposing partition folder details to consumers.
    • Automated Compaction Protocol:
      • Scheduled background Spark job runs every 4 hours:
      • Consolidates thousands of small streaming files (< 32 MB) into uniform

    512 MB columnar Parquet blocks using Z-order sorting on (subscriber_id, event_timestamp).

    3. ACID Snapshot Isolation & Concurrency Control [MC-CC-01]
    • The LAK-4919 Dirty Read Defense:
      • Readers query immutable Iceberg snapshots (s3://.../metadata/snap-*.json).
      • Concurrent streaming writes append new data files and commit an atomic metadata pointer update via optimistic concurrency control (OCC).
      • Analytical queries never encounter partial, uncommitted, or torn writes.

    Invariants and Contracts

    Mandatory Open Table Format Conformance [INV-LAKE-01]
      Analytical datasets exceeding 1 terabyte must be structured using Apache Iceberg open table format.
      Dumping raw, unpartitioned JSON or CSV files into persistent lake storage is strictly prohibited.
    
    Automated Storage Compaction Invariant [INV-LAKE-02]
      Lakehouse tables must not accumulate more than 1,000 un-compacted small files (< 64 MB) per partition.
      Background compaction pipelines must execute automatically to preserve query planning performance.
    
    Decoupled Storage and Compute Architecture [INV-LAKE-03]
      Analytical query engines (Trino, Athena, Spark) must remain stateless and fully decoupled from underlying S3 storage.
      Binding lakehouse data storage to proprietary, vendor-locked compute clusters is barred.
    

    Explicit Unknowns

    • AWS Glue Data Catalog API request rate limits during simultaneous 50-node EMR Spark batch compaction jobs (G-1).
    • Z-order clustering CPU cost efficiency on multi-attribute dimensional queries across 85 billion CDR records (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    45 TB daily 5G telemetry across 18M subscribersprovidedTelecom network capacity intakeCurrent
    Incident LAK-4919 $3.8M query compute disasterprovidedHistorical forensic audit reportHistorical
    14 million small files caused S3 metadata freezeprovidedOperational incident post-mortemHistorical
    Query latency target p95 <= 10 secondsprovidedNetwork Operations Center SLACurrent
    Apache Iceberg on Amazon S3 selecteddecidedDavid O'Reilly & Elena Rostova2026-09-15
    Mandatory open table format invariant INV-LAKE-01decidedArchitectural invariant INV-LAKE-012026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against data lake architecture standards:

    • ACID Rigor: PASS. Apache Iceberg snapshot isolation eliminates partial writes and dirty reads.
    • Compensating Compaction: PASS. Automated background bin-packing resolves the 14M small files crisis.
    • Query Performance: PASS. Partition pruning and Z-order sorting deliver sub-10s p95 query speeds.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-LAKE-01: David O'Reilly to determine whether Trino or StarRocks should be standardized as the primary interactive query engine for Gold-tier dimensional data marts in Q1 (Owner: David O'Reilly).

    Next steps

    1. Data Platform squad provisions the central Apache Iceberg catalog on AWS Glue.
    2. Ingestion engineers deploy Kafka S3 sink connectors targeting the Bronze raw zone.
    3. Conduct staging performance drill querying 10 TB of synthetic Iceberg CDR tables to verify sub-10s latency.

    skill: data-lake-architect

    Telecom Network Data Lake — Fitness Self-Check [LAKEARCH-TEL-FIT-001]

    Summary

    This fitness self-check evaluates the enterprise data lake architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.

    Detailed Description

    Criterion [FIT-n]ProbeEvidenceResultLimits of the claim
    FIT-1: Dual WriterSeed a streaming pipeline where two autonomous ingestion jobs attempt to commit conflicting transaction metadata snapshots to the same Iceberg table simultaneously.Iceberg optimistic concurrency validator probe_concurrent_iceberg_commit verifying automatic commit conflict retry with diagnostic ERR_ICEBERG_CONCURRENT_COMMIT_RETRY.passConfirms Iceberg metadata OCC engine; does not evaluate direct manual file manipulation on S3 via raw AWS CLI.
    FIT-2: Undefined GrainSeed a proposed Iceberg table definition for CDR data that omits an explicit primary key grain and temporal timestamp definition.Lakehouse table schema linter probe_missing_table_grain verifying DDL rejection with diagnostic ERR_ICEBERG_TABLE_LACKS_DECLARED_GRAIN.passConfirms DDL catalog admission gates; does not inspect temporary staging tables dropped within the same session.
    FIT-3: Silent Schema DriftSeed an upstream 5G network probe that alters an existing column data type from integer to string without registering an Iceberg schema evolution migration.Schema evolution validator probe probe_unversioned_schema_mutation verifying ingestion halt with diagnostic ERR_UNAUTHORIZED_SCHEMA_DRIFT_DETECTED.passConfirms automated Parquet reader schema enforcement; does not evaluate raw unstructured text files in Bronze.

    Residual Risk

    • S3 PUT request cost spikes if an upstream Kafka producer misconfiguration generates high-frequency unbuffered micro-batches under 1 MB. Accepted by Elena Rostova with CloudWatch billing threshold alerts.

    Traceability

    ClaimClassificationSourceFreshness
    Rejection of uncoordinated concurrent writesderivedFIT-1 probe result2026-09-15
    Rejection of tables lacking declared grainderivedFIT-2 probe result2026-09-15
    Rejection of unversioned schema mutationsderivedFIT-3 probe result2026-09-15

    Verification

    No validator was supplied, so no command was run.

    Open Decisions

    None.

    Next steps

    1. Architecture Guild incorporates data lakehouse fitness probes into automated CI deployment verification.
    2. Platform team configures CloudWatch alerts monitoring S3 bucket file count distributions and Glue catalog latencies.
    3. Conduct quarterly disaster recovery drill validating Point-in-Time Iceberg table rollback following simulated accidental batch overwrite.

    enterprise-data-lake-and-lakehouse-stora.csv

    CSV · data export

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Design multi-zone medallion storage architectures for raw and curated dataOptimize partitioning strategies to reduce cloud storage query costsDefine Apache Iceberg table layouts for ACID compliance and compactionEstablish data ownership and retention policies for large scale lakes

    About this skill

    What it does

    This skill owns the architecture of governed analytical and archival datasets stored primarily as objects. It defines source and dataset authority, identities, storage-zone meaning, ingest/publication semantics, schema/format/layout compatibility, discoverability, access, lifecycle, maintenance, recovery, and evidence. It does not own bucket provisioning, one file writer, ETL transformations, a warehouse, a lakehouse table implementation, or the wider data platform.

    Use it when

    • Multiple producers and consumers share durable object-backed datasets with explicit ownership
    • Source record, batch/run, object, dataset version/snapshot, schema, partition and publication identities must remain traceable
    • Landing, quarantine, validated, conformed or published states need semantic boundaries rather than decorative folders
    • Raw immutability, corrections, late data, duplicate writes, partial batches and concurrent publishers need contracts
    • File format, compression, schema evolution, partition transforms, file sizing and compaction interact with workloads
    • Readers need atomic dataset visibility rather than discovering arbitrary objects by listing

    For example: “Our lake is 400 TB, a simple daily query costs EUR 90, and nobody can tell me who owns the 60 TB in /tmp-exports.”

    What you get

    • architecture/data-lake-architect/README.md
    • architecture/data-lake-architect/00-overview/data-lake-architect-overview.md
    • architecture/data-lake-architect/verification/fitness-self-check.md

    Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.

    All paths are relative to the output folder you choose.

    What it will not do

    Do not use merely to create a bucket, upload files, write Parquet/ORC, tune partitions or file sizes, configure S3/GCS lifecycle rules, implement ETL, build a warehouse/lakehouse/table, or select Iceberg/Delta/Hudi.

    How it works

    1. Check the subject is raw storage and layout.
    2. Define the zones and what may enter each.
    3. Choose the partitioning from the query predicates.
    4. Fix file size and format.
    5. State ownership and retention per dataset.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-artifact.md
    • assets/output-template-contract.md
    • assets/output-template-domain.md
    • assets/output-template-fitness.md
    • assets/output-template-mechanism.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions