The most common data pattern on Databricks is just three honest questions about your tables. Here is what each layer is for, when it helps, and when it is overkill.
Imagine you run a small bakery. Every morning a delivery truck drops off flour, eggs, sugar, and butter. You would not serve a bag of flour to a customer. You also would not bake a cake the moment the truck arrives and hope someone orders exactly that cake.
Instead you do three things. You store the ingredients as they arrived. You prep them: measure, sift, crack the eggs into bowls. Then, when an order comes in, you assemble the final product quickly, because the hard work is already done.
Medallion architecture is the same idea, applied to data. It is a way of organizing tables into three layers: bronze, silver, and gold. Each layer answers one honest question. That is the whole concept.
This pattern is not tied to any product, vendor, or year. It is simply a naming convention for a habit good data teams have always had: keep the raw copy, clean it once, and serve it in the shape people actually need. Databricks popularized the names, but the thinking is older than the tooling, and it will outlast whatever tool comes next.

The bronze layer stores data exactly as it arrived. No fixing, no filtering, no judgment.
If a CSV lands with a column called amt and values like 49.99 and free, bronze keeps both. If the same record shows up twice because a system retried, bronze keeps both copies.
This feels wrong at first. Beginners want to clean data at the door. Resist that urge, for one simple reason: you cannot fix a mistake you have already thrown away.
When a silver table looks odd three weeks later, the first question is always "was the source data wrong, or was our cleaning wrong?" Bronze is how you answer that. It is your receipt. It is also how you rebuild everything downstream when you discover a bug in your cleaning logic, which you will.
Bronze tables are append-only, schema-flexible, and boring. Boring is the goal.
The silver layer is where data becomes trustworthy. This is where you do the unglamorous work: parse the dates, remove the duplicates, fix the types, standardize the values, and enforce the rules you believe should always be true.
amt becomes amount, a proper decimal. free becomes 0.00 or gets quarantined, depending on your business rule. The duplicated retry is removed. 5/1/24 becomes a real date.
A good silver table has a contract: one row means one thing, columns have stable names and types, and anyone on the team can query it without knowing the messy history of how it arrived.
Most data quality work lives here, and it is worth doing deliberately. If you want to go deeper on this step, our data quality lesson walks through the checks that matter, and the duplicates article covers the most common silver-layer bug in practice.
The gold layer is where data becomes answers. Gold tables are shaped for a specific audience or question, not for completeness.
"Daily sales by city" for the dashboard. "Customer lifetime value" for the marketing team. "Features for the churn model." Each gold table exists because someone needs to make a decision with it.
This is the layer where aggregation, joining across silver tables, and business-friendly naming happen. Gold tables can be rebuilt from silver at any time, which means you can change your mind about what the business cares about without re-ingesting anything.
Notice the direction of trust: bronze is trusted by engineers, silver is trusted by analysts, gold is trusted by the business.
Say an e-commerce site exports orders every night. Here is the same data flowing through the three layers.

Bronze keeps the raw export, plus a column recording when it was loaded:
CREATE TABLE bronze_orders AS
SELECT
*,
current_timestamp() AS loaded_at
FROM raw_orders_export;Silver applies the cleaning rules once, in one place:
CREATE TABLE silver_orders AS
SELECT DISTINCT
CAST(order_id AS INT) AS order_id,
TO_DATE(order_date, 'M/d/yy') AS order_date,
INITCAP(TRIM(city)) AS city,
CAST(amount AS DECIMAL(10,2)) AS amount
FROM bronze_orders
WHERE order_id IS NOT NULL;Gold answers a question someone actually asked:
CREATE TABLE gold_daily_sales_by_city AS
SELECT
order_date,
city,
COUNT(*) AS orders,
SUM(amount) AS revenue
FROM silver_orders
GROUP BY order_date, city;Three small tables. Each one has a single job. When the dashboard number looks wrong, you know exactly where to look first: gold logic, silver rules, or bronze receipts.
If you prefer watching to reading, this walkthrough covers the same flow end to end:
https://youtu.be/3rOEktLcLOA
Cleaning at ingest. If you transform before bronze, you have no receipt. When someone asks "why is March different?", you will be guessing.
Gold tables read straight from bronze. This happens when a dashboard is needed by Friday. It works once, then the cleaning logic gets copied into five gold tables, and the five copies quietly disagree. Clean once in silver, reuse everywhere.
Treating the layers as a law. The medallion pattern is a vocabulary, not a religion. Its real value is that when someone says "this is a silver table," everyone knows what to expect from it.
For a personal learning project, a single pipeline with a raw table and one clean table is completely fine. Two layers, same thinking.
Some teams add a fourth layer, sometimes called platinum, for highly shared data products. That is also fine.
The number of layers matters far less than the discipline behind them: never lose the raw copy, clean once, and shape data for the question being asked. A two-layer design with that discipline beats a five-layer design without it.
Tools change every few years. File formats change. Compute engines change. But the three questions behind medallion architecture do not change, because they are not about technology. They are about how humans recover from mistakes, build trust in numbers, and turn data into decisions.
That is why this pattern shows up everywhere, from a student's first Delta Lake tables in Databricks Free Edition to the largest lakehouses in production. Learn it once and it transfers to every data platform you will ever touch.
If you want to build this pattern with your own hands, start here: