A source team changed one column type. Nothing failed. The revenue number was wrong for four weeks.
The pipeline had been green for four months. Nobody thought about it. It ran at 2 a.m., the bronze table filled, silver cleaned it up, gold rolled it into the revenue dashboard the finance team opened every Monday.
Then one Monday the number was low. Not broken. Not zero. Just low enough that someone said "that looks off" and low enough that nobody could prove it.
No job failed. No alert fired. The pipeline had run perfectly.
What had changed was upstream. A source team shipped a small release. They changed amount from a number to a string, because one of their integrations occasionally sends "Not Available" instead of a value. To them it was a tolerant field. To our pipeline it was a silent hole. Rows with text landed as NULL, the sum ignored them, and the dashboard reported a smaller month than the business actually had.
This is schema evolution. And the interesting part is not the failure. It is the silence.
Start with the table your pipeline expects.
orders
order_id int
customer_id int
amount decimal(10,2)
order_date dateFour columns. Your transform reads them, cleans them, aggregates them.
Now imagine three things the source team might do.
They add discount_amt. Usually harmless. Your code never asked for it, so your code does not care. You choose later whether you want it.
They rename amount to total_amount. Your transform still looks for amount, so the job fails at once. That sounds bad, and it feels bad at 6 a.m., but it is the good case. You know immediately. Nothing wrong reached the dashboard.
They change amount from decimal to string. Nothing fails. Most rows still parse. One row says "Not Available". The engine turns it into NULL, or your cast quietly drops it, and the pipeline stays green while the number goes wrong.
That is the shape of the whole problem. A schema change either breaks your pipeline or breaks your data. Breaking the pipeline is a scheduling problem. Breaking the data is a trust problem, and trust takes far longer to rebuild.
We wrote about the loud version of this in the schema changed overnight, where a rename took down a bronze load. This article is about the quiet version, and about designing so the quiet version does not happen.

1. A column is added. The safest change. Your existing readers ignore it. In a bronze layer you usually want to capture it anyway, because you cannot go back and collect data you threw away.
2. A column is removed. Anything that depended on it breaks. Views, joins, dashboards, feature tables. The blast radius is however many things referenced that name, which is often more than the source team imagines.
3. A column is renamed. To a human this is tidying up. To a pipeline it is two events at once: a column disappeared and a stranger arrived. Delta Lake has no way to know they are the same field unless you tell it.
4. A type changes. This one has more failure modes than people expect. An int turned into a string breaks joins, because 1001 and "1001" are not the same key. A timestamp format change breaks incremental loads, because your watermark comparison stops making sense. A decimal turned into text breaks math without complaining, which is how you get the Monday morning story above.
5. The meaning changes and the schema does not. The most dangerous change on the list, and the only one no tool can detect for you.
Picture a status column. Yesterday "active" meant a customer who purchased in the last 90 days. Today the source team decided 30 days is a better definition. Same column name. Same type. Same row count. No pipeline failure anywhere. And every retention number, every churn model feature, every board slide that used that field just shifted, silently, on a Tuesday.
You cannot catch that with schema enforcement. You catch it with a conversation, or with a contract.

A schema is a technical description: these columns, these types, this nesting. Useful, and not enough.
A data contract is the agreement about what must stay true:
Notice that only the first two items live in a schema. The rest live in the relationship between two teams. That is why schema evolution is only half a technical problem. The other half is communication, and no amount of Delta configuration fixes it.
A short, boring, written contract prevents more incidents than any clever pipeline.
The most expensive assumption in data engineering is that a source schema is fixed. It is not. Sources are owned by teams with their own roadmap, and your table is downstream of their sprint. Design as if a column will move next quarter, because it will.
When something arrives that you did not expect, name it first. Is this backward compatible, meaning old readers still work, or is it breaking, meaning something must be updated before the next run? A new nullable column is the first kind. A rename or a narrowing type change is the second. Most bad decisions come from treating every schema surprise as the same emergency.
This is the highest value habit on the list. Check the incoming shape at the door, before anything reaches a trusted table.
from pyspark.sql import functions as F
expected_columns = {"order_id", "customer_id", "amount", "order_date"}
incoming = spark.read.option("header", True).csv(
"/Volumes/workspace/default/book_data/orders_landing"
)
incoming_columns = set(incoming.columns)
missing_columns = expected_columns - incoming_columns
new_columns = incoming_columns - expected_columns
if missing_columns:
raise ValueError(f"Contract violation. Missing columns: {sorted(missing_columns)}")
if new_columns:
print(f"New columns arrived, review before using: {sorted(new_columns)}")Six lines of set arithmetic. Missing columns stop the load. New columns are logged for a human, not silently trusted. That is the whole idea.
For type safety, do the cast yourself and count what fails instead of letting the engine decide.
typed = incoming.withColumn("amount_num", F.col("amount").cast("decimal(10,2)"))
bad_amount_count = typed.filter(
F.col("amount").isNotNull() & F.col("amount_num").isNull()
).count()
if bad_amount_count > 0:
raise ValueError(f"{bad_amount_count} rows have a non numeric amount")That check is exactly what would have caught "Not Available" on the first night instead of the fourth Monday.
This is why the medallion pattern exists, and schema change is the clearest argument for it.
Bronze stores what the source actually sent, including the odd column and the messy type. Silver applies your rules: correct types, deduplication, standard names. Gold is the business view. When a source changes, bronze absorbs it and silver decides what to do about it. You still have the original rows to reprocess, which means a mistake in your interpretation is recoverable.
If your pipeline casts and renames on the way in, and stores only the cleaned result, then every interpretation error is permanent. Our medallion architecture lesson walks through the layering, and schema evolution covers the Delta side of it in the workspace.
For file ingestion, Auto Loader gives you a good default: track the schema, add new columns automatically, and park anything unexpected in a rescue column instead of failing.
orders_stream = (
spark.readStream
.format("cloudFiles")
.option("cloudFiles.format", "json")
.option("cloudFiles.schemaLocation", "/Volumes/workspace/default/book_data/_schema/orders")
.option("cloudFiles.schemaEvolutionMode", "addNewColumns")
.option("rescuedDataColumn", "_rescued_data")
.load("/Volumes/workspace/default/book_data/orders_raw")
)Then monitor _rescued_data. A rescue column nobody looks at is just a nicer place to lose data.
You want to know about a schema change before a dashboard owner does. Column count, column names, type signature, and null rates per column are enough to catch most of it. Store yesterday's signature, compare today's, alert on the diff.
Delta constraints are the other half, because they refuse bad rows at write time rather than reporting them later.
ALTER TABLE silver_orders
ADD CONSTRAINT amount_is_positive CHECK (amount >= 0);We went deeper on this in the pipeline was green, the numbers were wrong, which is the same failure mode seen from the data quality side. Late data has a similar shape too, and we covered it in the event arrived two days late.
One table rarely has one consumer. It feeds dashboards, a machine learning feature set, an internal API, an export, and increasingly an agent. A change that is trivial for one consumer is a production incident for another. The person making the change usually cannot see the full list, which is precisely why the announcement matters more than the migration.
For years the last mile of a data platform was a human. Someone read a dashboard, noticed the number looked odd, and asked a question. Human doubt was our final quality check.
Agents do not doubt. An agent reads a field, uses it to make a recommendation, and takes an action. If status quietly changed meaning last Tuesday, the agent keeps acting on last week's definition with complete confidence, at machine speed, across every case it touches.
That is the real cost of silent schema drift now. It is not a wrong chart. It is a wrong decision, repeated.
We wrote about this shift in the next database may be built as much for agents as for developers and in the data stack is being redesigned for AI agents. The short version: when a machine is the consumer, your column names, comments, constraints, and documented meaning stop being nice hygiene and become the interface. If you want that idea in long form, it is the argument behind our companion book, The Context Advantage.
Schema evolution is normal. Sources will change, and a platform that cannot absorb change is not stable, it is brittle in a different way.
Uncontrolled schema evolution is the dangerous part. The goal is not to freeze your sources. The goal is to make change visible, testable, and safe.
So the next time a source team ships a release, the question to ask is not "did the pipeline fail?" The question is "did anything change that my pipeline accepted without understanding?"
Because sometimes a pipeline breaks when the source changes. And sometimes it keeps running.
The second one is the bigger problem.