The next database may be built as much for agents as for developers

Agents read your data platform through metadata. That changes what a good database looks like.

Every database you have ever used was designed for someone who could ask a question out loud.

Think about what a developer actually does before writing a query. They open the catalog. They squint at three tables with similar names. They ping someone on Slack. They look at last month's dashboard to see which table the finance team trusted. Then they write the SQL.

None of that happens in the database. It happens around it, in human memory and human conversation. The database only ever had to be correct. The meaning lived somewhere else.

Agents do not have somewhere else.

An agent gets one shot at your platform and it sees exactly what the platform exposes: table names, column names, comments, tags, constraints, lineage, run history. That is its whole world. If the meaning is not written down inside the system, the agent invents it.

That is why the next generation of databases will be shaped as much by agents as by developers. Not because agents are more important, but because agents are stricter readers.

Developer path and agent path into the same lakehouse

The old contract was correctness. The new one is legibility

For a long time we judged a data platform on a small list of things. Does it store data safely. Does it answer queries fast. Does it scale. Does it enforce permissions.

All of that still matters. But an agent adds a question we never really had to answer before.

Can the system explain itself?

A developer can survive a table called tbl_pmt_f_v2 with no comments, because a developer can find the person who built it. An agent reads the name, guesses, writes a query, and returns a confident number that happens to be wrong. Nobody sees the mistake, because the pipeline was green.

We wrote about that exact failure in The pipeline was green. The numbers were wrong. Agents make the quiet failure more common, because they operate faster than review does.

So legibility becomes a real engineering property. Not documentation. Not a wiki nobody opens. Something the system itself carries.

Four things an agent reads

It helps to picture a data platform as four layers, stacked from the physical to the conceptual.

The four layers an agent reads: storage, meaning, rules, memory

Storage is the part we always worked on. Files, formats, Delta tables. If you are still deciding between formats, CSV, JSON, Parquet, and Delta: which one should you use? covers the trade-offs.

Meaning is names, comments, and tags. Humans treat this as optional polish. An agent treats it as the API.

Rules are permissions and constraints. For a human, rules are a wall you occasionally bump into. For an agent, rules are the only thing standing between a helpful suggestion and a rewritten fact table.

Memory is lineage and run history. Who produced this table, from what, when, and what happened the last few times it ran. A person builds this memory in their head over months. An agent needs it available on request.

Most teams are strong on storage, decent on rules, and thin on meaning and memory. That gap is exactly where agent work goes wrong.

What this looks like in Databricks today

You do not need a new product to start. Unity Catalog already exposes the meaning layer, and you can improve it with plain SQL in an afternoon.

COMMENT ON TABLE workspace.default.orders IS
  'One row per confirmed customer order. Source: checkout events. Grain: order_id.';

ALTER TABLE workspace.default.orders
  ALTER COLUMN order_status
  COMMENT 'One of placed, shipped, cancelled. Set by the fulfilment system.';

ALTER TABLE workspace.default.orders
  SET TAGS ('domain' = 'sales', 'certified' = 'true');

Three statements, and an agent reading that table now knows the grain, the source, the vocabulary of a confusing column, and whether a human has blessed it.

A good weekly habit is to read your own catalog the way an agent would, and look for silence.

from pyspark.sql import functions as F

described_df = spark.sql("DESCRIBE TABLE EXTENDED workspace.default.orders")

# A column with no comment is a column your agent will guess about.
missing_comments_df = described_df.filter(
    (F.col("col_name") != "") & (F.col("comment").isNull())
)

missing_comments_df.select("col_name").show(truncate=False)

If that list is long, you do not have an AI problem. You have a description problem, and description problems are fixable this week.

Then there is the rules layer. When something other than a person can write to your tables, constraints stop being paperwork and start being a seatbelt.

ALTER TABLE workspace.default.orders
  ADD CONSTRAINT order_amount_positive CHECK (order_amount > 0);

ALTER TABLE workspace.default.orders
  ADD CONSTRAINT order_status_known
  CHECK (order_status IN ('placed', 'shipped', 'cancelled'));

A constraint is a sentence about the world that the database will defend for you, at 3 a.m., against any writer, human or not.

Comments tell an agent what the data means. Constraints tell it what it is not allowed to break.

The habits that carry over

Here is the reassuring part. Almost nothing in this shift is new engineering. It is old engineering that suddenly pays better.

Make every step repeatable, so a retry is boring rather than dangerous. That is the argument in The job ran twice. Why safe reruns decide if a pipeline is production ready.

Keep a clear separation between raw, cleaned, and business tables, so an agent has an obvious place to read from and an obvious place it must never touch. The medallion layout does that work quietly.

Write down what you assume. A comment that says "amounts are in cents" saves a person ten minutes and saves an agent from being wrong.

Let the platform hold the memory instead of your team's group chat. Lineage and history are not reporting features anymore. They are how a machine learns what normal looks like.

If you want the wider picture of where this is heading, The data stack is being redesigned for AI agents and Agent discovers. Agent applies. What that changes for data engineers both go deeper on the read side and the write side of the same story.

What changes for you

The job title does not change. The centre of gravity does.

Less time spent being the person who knows what the table means. More time spent making sure the table can say what it means without you in the room.

That is a better job, honestly. Explaining the same column to a new teammate every quarter was never the interesting part. Designing a system that explains itself is.

And the skill has a name now. Deciding what context a system should carry, how much of it, and where it lives, is turning into a distinct part of the work. We looked at that in Context engineering is becoming a real job skill for data engineers, and it is the whole subject of our sister book, The Context Advantage.

Where to start this week

Pick one table that matters. Give it a table comment that states the grain and the source. Comment the three columns people always ask about. Add one CHECK constraint you know is true. Tag it with a domain and a certified flag.

Then ask a colleague who has never seen that table to explain it using only what the catalog shows. If they can, an agent probably can too.

That is the test. Not clever prompting. Just a table that speaks for itself.

Continue learning

If you are building these habits from the ground up, our book walks the same path in order, from your first table to pipelines you can trust.