Pipeline latency is easy to measure. Decision latency is the honest number, and it is where the next decade of this work happens.
Every Monday at nine, the pricing team met to talk about the weekend.
The numbers were ready. The dashboard was green. The pipeline had run on time, every night, without a single failure. And yet the conversation always sounded the same: "So this happened two days ago. Is it still happening?"
Nobody knew. The data was correct. It was just late for the decision it was supposed to serve.
That gap is the real story of the next phase of data engineering. It is not just about faster pipelines. It is about faster decisions.
Over the last few years we got very good at the machinery. Auto Loader picks up files as they land. Delta Lake gives us reliable tables. Lakeflow Declarative Pipelines handle dependencies and retries for us. A job that used to take an engineer a week now takes an afternoon.
So why does the business still feel slow?
Because a pipeline finishing is not the same as a decision happening. Between the event and the action there is a chain, and most of the delay lives in the parts we never measure.

Most teams optimise the second box and ignore the last two.
An order is placed. Ten minutes later the row lands in bronze. Two hours later the nightly batch rolls it into a silver table. The next morning someone opens a dashboard. Two days later a decision gets made in a meeting.
We shaved the ten minutes down to two and called it a win. The decision still took two days.
Pipeline latency is easy to measure. Decision latency is the honest one.
Decision latency is the time between something happening in the real world and someone or something acting on it. It includes your ingestion. It also includes the dashboard nobody opens, the alert that goes to a channel nobody watches, and the weekly meeting that acts as a queue.
When you start measuring that number, priorities change. Making a job run in four minutes instead of nine matters very little if the decision waits for Monday. Moving a decision from a weekly meeting to a daily alert matters enormously, even on the same pipeline.
This is why I think of data engineering less as plumbing and more as shortening the distance between a fact and an action.
Most pipelines are designed bottom up. We look at the source, ingest it, clean it, model it, and then ask who wants a dashboard. The tip of that pyramid is a question mark.
Design in the other direction.

Ask four questions before you write any code.
What decision does this data serve? Not "the marketing team wants sales data" but "we decide each morning which products to promote."
Who or what makes that decision? A person, a scheduled report, or increasingly an agent.
How fresh does the data need to be for that decision to be right? Hourly and daily are very different systems, and so are their costs.
What happens when the data is wrong? If the answer is "we lose money quietly," you need quality checks in the pipeline, not in a review meeting.
Those four answers tell you your ingestion pattern, your table design, and your freshness target. They also stop you building an expensive streaming pipeline for a number somebody looks at once a week.
There is a habit in our field of treating real time as a compliment. It is not. It is a cost.
Free Edition is a good place to feel this. You can build the same table three ways and watch what each one buys you.
A daily batch is the cheapest and simplest. It is right for anything reviewed on a human rhythm: monthly finance reporting, weekly cohort analysis, quarterly planning.
An incremental micro batch, running every fifteen minutes with availableNow=True, covers most operational needs. Stock levels, support queues, campaign spend. It looks like streaming to the business and costs like batch.
from pyspark.sql import functions as F
orders_stream = (
spark.readStream
.format("cloudFiles")
.option("cloudFiles.format", "csv")
.option("cloudFiles.schemaLocation", "/Volumes/workspace/default/book_data/_schema/orders")
.load("/Volumes/workspace/default/book_data/orders/")
)
daily_orders = (
orders_stream
.withColumn("order_date", F.to_date("order_timestamp"))
.groupBy("order_date", "product_id")
.agg(F.sum("order_amount").alias("revenue"))
)Continuous streaming is right when the decision itself is machine speed: fraud checks, personalisation, an agent choosing what to do next. Very few dashboards belong here.
If you want the mechanics behind these choices, the incremental processing lesson walks through the trade-offs, and streaming tables versus materialized views covers when each object type fits.
Here is the part that changes the job.
For twenty years the last mile of every data platform was a human reading a chart. That assumption shaped everything: daily refreshes because people work in days, dashboards because people read pictures, wide tables because people like to explore.
Agents do not work that way. An agent can ask a question every second, and it will act on whatever it finds without pausing to wonder if the number looks odd. When the consumer is a machine, freshness stops being a nicety and semantics stop being optional.
That is why Databricks has been pushing in this direction with Lakebase for operational state and the Electric acquisition for keeping data close to where reasoning happens. It is also why databases are starting to be designed for agents rather than only for developers.
The engineering consequence is simple. Your tables need to explain themselves. Column names, descriptions, and grain need to be unambiguous, because nobody is there to interpret them. The Unity Catalog lesson covers the governance side of that, and if you want the reasoning behind writing context that machines can use well, that is the whole argument of The Context Advantage.
If you want to hear this argument in a shorter form, here is a take on what changes when decisions become the center of the pipeline:
https://youtu.be/tz28AQNNQjc
None of these need a paid workspace.
Write down the decision each of your top three tables serves. If you cannot name one, you have found a table to retire.
Measure decision latency for one important number. Timestamp the source event, the table update, and the moment someone acted. The last gap is usually the biggest, and it is usually not technical.
Move one decision closer to the data. Replace one weekly review with a threshold alert. That single change often beats a month of tuning. The observability and alerts article shows how to wire it.
Put quality checks where the decision is made, not after. Delta constraints and pipeline expectations stop a wrong number from reaching a decision at all. See data quality checks in Databricks.
Faster pipelines were the last decade. Faster decisions are this one.
It changes what a good data engineer looks like. Less "I made the job run in four minutes," more "the team now acts on this the same morning instead of the following Monday." The second sentence is harder to earn and much harder to replace.
The tools are ready for it. Free Edition gives you Delta Lake, Auto Loader, Unity Catalog, and pipelines with expectations. What is missing on most teams is not capability. It is the habit of starting from the decision.
If you want to build this way from the ground up, start with Jobs and Pipelines to get the scheduling rhythm right, then Data Quality so the numbers can be trusted, then Incremental Processing to pick the right freshness for each decision.
And if you learn best by building, your first data engineering project in Databricks Free Edition takes you from a raw CSV to a table someone could actually decide from, in two evenings.