Your AI agent remembers the refund. The business still says pending.

Memory helps an agent reason. Business state tells the organization what is actually true. Reliable agents need both, but they must never confuse them.

Everyone is talking about giving AI agents memory.

The idea is easy to understand. An agent should remember your preferences, your earlier questions, the documents it has read, and the work it has already completed. Without memory, every conversation starts from zero.

But the moment an agent starts acting inside a real organization, memory is not enough.

Imagine asking a customer-service agent about a delayed refund. The agent reads an earlier conversation and replies:

Your refund has been approved. The money should arrive soon.

That sounds reassuring. There is only one problem.

The payment system still shows the refund as pending.

The agent remembers a discussion about approval. The business has not recorded an approved refund. Those are two different facts.

This distinction will shape whether enterprise agents become dependable coworkers or confident sources of confusion.

Memory tells the agent what it knows. Business state tells the organization what is true.

Both matter. They solve different problems. They should never be treated as the same thing.

Watch the explainer

This short BricksNotes video introduces the distinction before we examine the systems, controls, and examples behind it.

https://youtu.be/_oA3AVOYfAE

A BricksNotes watercolor diagram comparing the contextual role of AI agent memory with authoritative business state

What agent memory actually does

Memory gives an agent continuity.

It may contain the user's preferred language, a summary of an earlier conversation, retrieved passages from a policy document, the result of a previous tool call, or a note that a customer was frustrated yesterday.

This information helps the agent interpret the next request. It makes the experience feel connected rather than repetitive.

Memory is especially useful for questions such as:

These are reasoning questions. The answers can help the agent decide what to look at next.

They do not automatically authorize an action.

A remembered sentence saying "the manager approved this" is not the same as an approval event recorded by the workflow system. A summary saying "the invoice was paid" is not the same as a settled payment record. A retrieved document saying an employee is eligible for leave is not the same as an approved leave request.

Memory is context from the past. It can be incomplete, stale, compressed, misunderstood, or simply wrong.

That is not a reason to avoid memory. It is a reason to give it the right job.

Business state is the organization's current truth

Business state describes where a real process stands now.

An order may be created, paid, shipped, delivered, cancelled, or refunded. A support case may be open, waiting for customer, escalated, or resolved. A data pipeline run may be queued, running, failed, or completed.

Those labels are not conversational hints. They are operational facts.

Business state normally belongs in a system of record such as a payment platform, an order database, a workflow engine, a human-resources system, PostgreSQL, or a governed lakehouse table.

The exact technology matters less than the contract around it. The state must be:

This is familiar territory for data engineers. We already design systems that preserve facts, process changes safely, recover from failure, and explain what happened.

The agent does not replace those responsibilities. It makes them more important.

The refund that exists only in memory

Return to the refund example.

A customer asks for a refund. The agent checks the conversation and sees that a support representative said, "This looks eligible. I will submit it for approval."

A weak design compresses that exchange into memory as "refund approved."

The next agent trusts the summary and tells the customer the refund is on its way.

A reliable design follows a different path.

First, the agent uses memory to understand the request and avoid asking the customer to repeat everything.

Then it reads the current refund record from the payment or order system. The record says pending review.

The agent can now answer honestly:

I found the earlier request. It has been submitted and is still waiting for approval. I have not found a completed refund yet.

If the agent is allowed to help move the case forward, it can propose the next action. But the actual transition from pending to approved must pass through the business workflow.

The workflow may verify the amount, check the return policy, confirm the user's permission, request human approval, write the new status, record the actor, and emit an audit event.

Only after that succeeds should the agent say the refund was approved.

This is a small example of a large design rule:

Never let fluent language get ahead of committed state.

State transitions are business facts

Many business processes can be understood as a set of allowed states and transitions.

A simple request might move like this:

PENDING → APPROVED → COMPLETED

It may also move to REJECTED or FAILED.

The arrows matter as much as the labels. They define what is allowed to happen.

Can a failed payment jump directly to refunded? Can an employee approve their own access request? Can an agent close a high-risk incident without a human review? Can a pipeline be marked completed before its quality checks pass?

These are policy questions, not language-model questions.

An agent can interpret intent and recommend a transition. A governed service should validate and commit it.

A BricksNotes watercolor process showing permissions, approval, durable state, audit history, and observation around an agent action

A safe transition usually includes several steps:

  1. Read the latest state.
  2. Confirm that the requested transition is valid.
  3. Check the identity and permissions of the actor.
  4. Gather any required approval.
  5. Write the change atomically.
  6. Record an audit event and reason.
  7. Observe the outcome.
  8. Return the committed result to the agent.

If the write fails after step four, the agent must not invent success. It should report that the request could not be completed and leave the record in its last valid state.

This is where idempotency matters too. A retry should not create a second refund, submit a second payroll change, or run the same destructive operation twice. The principles in our incremental processing explainer apply here: checkpoints, unique operation identifiers, and careful retry logic turn failure from a surprise into a designed path.

Three places where the distinction becomes real

HR: a conversation is not an employment record

An HR agent may remember that an employee asked to change their bank account, take parental leave, or receive access to a sensitive system.

That memory helps continue the conversation. It should not directly change payroll, benefits, or permissions.

A bank-account change may require identity verification. Leave approval may depend on policy and a manager decision. Access may require a role check, separation of duties, and an expiry date.

The agent can collect the right information and explain the process. The HR system or approval workflow must remain the source of truth.

There is also a privacy reason to keep the boundary clear. An agent's memory should not become an uncontrolled copy of every sensitive employee fact. The Unity Catalog lesson explains the deeper principle: access should follow identity, policy, and purpose, not convenience.

Data operations: saying fixed is not the same as being healthy

Imagine an operations agent investigating a failed pipeline.

It remembers that a similar failure last week was solved by refreshing a credential. It may reasonably propose the same check.

But it cannot declare the pipeline healthy because the memory says the fix worked before.

The current run state belongs in the orchestration system. The output row counts belong in the data platform. Quality results belong in monitored tables. The final status should change to completed only after the run and its checks actually succeed.

This is why observability is part of business state. A status without evidence is only a label.

The debugging and monitoring lesson shows how to inspect what a system did instead of trusting what we hoped it did. For agents, that mindset is essential. Tool calls, approvals, writes, failures, and retries should be traceable as one connected operation.

Customer service: empathy needs current facts

Customer-service agents benefit greatly from memory. Remembering a customer's preference or earlier frustration can make support feel human.

But empathy does not change an order record.

Before promising a replacement, credit, cancellation, or refund, the agent should read current inventory, payment, shipment, and policy state. When several systems disagree, it should surface the uncertainty rather than select the most convenient answer.

This is a quiet but important form of trust. The agent can be warm without pretending certainty.

Why memory can be wrong

Agent memory can fail in several ways.

A summary may remove a qualification. "Eligible for review" may become "eligible." A retrieval system may return an older policy. A user may state something that was never verified. Two records may describe different customers with similar names. A tool call may have timed out after the agent assumed it succeeded.

Memory can also become stale. An address remembered yesterday may have changed this morning. A support case may have been reopened. An access grant may have expired.

Even technically correct memory can be inappropriate for the current user. An agent serving one department should not recall restricted information from another.

The response is not to make memory authoritative. The response is to attach provenance and limits to it.

Useful memory should answer questions such as:

This connects to context engineering. Good context is not simply more text. It is the right evidence, delivered at the right moment, with enough structure to judge its reliability.

Many agents need one shared truth

The distinction becomes even more important when multiple agents work together.

A service agent may gather a request. A policy agent may check eligibility. A finance agent may authorize an amount. An operations agent may execute the change.

If each agent carries its own private memory of the process, they can quickly disagree.

One remembers that approval was requested. Another remembers that approval was expected. A third interprets an optimistic message as completion.

Shared business state prevents this drift.

Every agent can read the same current status, version, owner, approval record, and operation identifier. Each transition can be conditional on the version it read. If another agent changed the record first, the stale write can be rejected and reconsidered.

This is coordination through facts rather than coordination through conversation.

Our article on managing a fleet of AI agents explores the broader control problem. The practical foundation is simple: agents may have different skills and memories, but they need a common place to learn what the organization has actually committed.

What Databricks contributes

Databricks does not replace payment platforms, order databases, workflow engines, or transactional applications. Those systems should continue to own the operational facts they are designed to manage.

The lakehouse becomes valuable when an organization needs a governed, historical, cross-system view of those facts.

Bronze tables can preserve raw events from agent tools and source systems. Silver tables can standardize identities, operation IDs, timestamps, state names, and outcomes. Gold tables can provide trusted views such as refund performance, approval latency, failed actions, or repeated agent retries.

That pattern is explained in the medallion architecture lesson. The aim is not to create decorative layers. It is to move from raw evidence toward dependable meaning.

Delta Lake adds durable history and transactional writes. Its version history can help investigators reconstruct which state was visible when an agent acted. Start with the Delta Lake lesson to understand why reliable tables are more than files in a folder.

Unity Catalog can govern which identities may read or change particular data. This matters when an HR agent, finance agent, and support agent should not see the same fields.

Data quality rules can catch impossible transitions, missing operation IDs, duplicated refunds, or completed records without completion timestamps. The data quality lesson is a practical place to build that habit.

AI Gateway and Serving can help govern model access and observe AI usage around the decision path. These capabilities may require a paid Databricks workspace. The AI Gateway and Serving lesson explains the concepts and clearly marks what is outside Free Edition.

The larger lesson is that the model is only one part of the system. Reliable action comes from the model, the data platform, the systems of record, and the control path working together.

A practical architecture for trustworthy agents

A useful design separates five responsibilities.

Memory layer: stores conversation summaries, preferences, retrieved context, and prior reasoning clues.

Truth layer: reads current facts from systems of record and governed analytical tables.

Policy layer: decides whether the requested action is allowed and whether approval is required.

Execution layer: performs an idempotent write through a narrow, validated tool or service.

Evidence layer: records the request, identity, input state, decision, approval, write result, and final state.

The agent can move between these layers, but it should not collapse them into one prompt.

Before a meaningful action, ask:

  1. What does the agent remember?
  2. What is currently true?
  3. Who is allowed to change it?
  4. Which transition is valid?
  5. What evidence will prove the outcome?

Those five questions are more valuable than adding another page of instructions to the agent.

The handover is the real design problem

The hardest moment is not when the agent forms an idea. It is when reasoning becomes action.

That handover should be explicit.

The agent proposes: "Approve refund R-1042 for $79 because the return was received and policy P-7 applies."

The business service checks the current record, policy version, amount, user permission, prior operations, and approval requirement. It either commits the state change or returns a precise rejection.

Then the agent reads the committed result.

The agent is good at understanding messy intent and explaining outcomes. The business service is good at enforcing deterministic rules and preserving truth. A reliable architecture lets each do the work it is suited for.

Build the data foundation before the illusion

It is tempting to judge an agent by the quality of its conversation.

But an enterprise agent should also be judged by quieter questions.

Can it distinguish a suggestion from a committed change? Can it recover when a tool fails? Can two agents agree on the same state? Can an auditor reconstruct the action? Can a person stop or reverse a dangerous transition? Can the system say, "I do not know yet"?

These questions bring us back to the central BricksNotes belief: the future of AI depends on the quality of the data systems beneath it.

Our article on the agentic data stack explains why data architecture is being redesigned for agents. Our data foundation article goes deeper into what models need before they can act reliably. And Thinking in Data Engineering with Databricks builds the foundation step by step, from tables and transformations to quality, history, governance, and operations.

You do not need to begin with a large multi-agent platform.

Begin with one process. Name its states. Identify the system of record. Define one allowed transition. Add one permission check, one durable operation ID, and one audit event. Then make the agent read the result before it speaks.

That is how trustworthy systems grow.

Memory guides reasoning. Business state guides action.

Keep those two ideas separate, and agents can become far more useful without asking the organization to surrender control.

Brick by brick.