Most teams don’t fail at building a data lake. They fail at running one. Small files pile up, schemas drift, nobody trusts the numbers, and query costs keep growing. Open table formats like Apache Iceberg fix many of these problems, but only if you design around them from day one.
This post walks through the blueprint we use for lakehouses on AWS.
The reference architecture
| Layer | Purpose | Typical AWS services |
|---|---|---|
| Ingestion | Land data reliably | AWS DMS, Kinesis Data Streams, Amazon MSK, AppFlow |
| Raw (bronze) | Immutable copy of source data | Amazon S3 |
| Curated (silver) | Cleaned, conformed Iceberg tables | AWS Glue / EMR (Spark), Iceberg |
| Serving (gold) | Business-ready models and aggregates | dbt, Athena, Redshift Spectrum |
| Governance | Catalog, access, lineage | Glue Data Catalog, Lake Formation |
| Orchestration | Scheduling and dependencies | Amazon MWAA (Airflow), Step Functions |
Why Iceberg?
Iceberg brings warehouse-style guarantees to files on S3:
- ACID transactions, so readers never see half-written data
- Schema evolution without rewriting tables
- Hidden partitioning, so analysts don’t need to know the partition layout
- Time travel for audits, debugging, and reproducible ML training sets
- Engine independence: Spark, Athena, Redshift, Snowflake, Flink and Trino can all read the same tables
Five operational details that matter
1. Plan compaction from the start
Streaming and frequent micro-batches create many small files. Schedule compaction (rewrite_data_files) and snapshot expiration, or turn on the automatic table optimization in the Glue Data Catalog.
CALL glue_catalog.system.rewrite_data_files(
table => 'curated.orders',
options => map('target-file-size-bytes', '536870912')
);
2. Use MERGE for CDC, not full reloads
With DMS or Debezium change data, use MERGE INTO on Iceberg tables. It is cheaper, faster, and keeps history intact.
3. Centralize permissions in Lake Formation
Grant table-level and column-level access in one place instead of spreading S3 bucket policies around. Your security team will thank you.
4. Treat data quality as a pipeline stage
Add checks (Glue Data Quality, Great Expectations, or dbt tests) between layers and fail fast. Bad data that reaches the gold layer costs far more to fix.
5. Watch costs per workload
Tag Glue jobs, EMR clusters and Athena workgroups by domain. Set Athena per-query scan limits. Most “the lake is too expensive” conversations are really “we can’t see where the money goes” conversations.
Getting started
Start small: pick one high-value domain, build it end to end through gold, and put a dashboard or ML feature on top. Proving value on one domain beats building a perfect platform nobody uses.
Planning a lakehouse or migrating from a legacy warehouse? Get in touch. We are happy to review your architecture.