From retrieval to reasoning to action — the complete architecture of enterprise AI agents, mapped to the data engineering skills you already have.
Most AI agents fail in production.
Not because the models are bad. Not because the prompts are wrong. They fail because the data underneath them is messy, ungoverned, and disconnected from the business context that makes decisions meaningful.
Databricks knows this. That is why they built Agent Bricks, not as another chatbot framework, but as a system that treats AI agents the same way data engineers treat pipelines: with structure, testing, governance, and observability baked in from the start.
If you have been following the Agentic Enterprise conversation, Agent Bricks is where theory becomes practice. With the launch of Genie One and Genie Ontology, agents now have a live business context layer to draw from.
Agent Bricks is Databricks' product for building domain-specific AI agents that are grounded in enterprise data. It sits within the Mosaic AI platform and integrates with everything you already know: Delta Lake, Unity Catalog, MLflow, and Spark.
But here is the key insight most people miss.
Agent Bricks does not just wrap a language model in an API. It builds what Databricks calls a Compound AI System, multiple components (retrieval, reasoning, tools, memory, guardrails) working together as a pipeline.
Sound familiar? It should.
A compound AI system is architecturally identical to a data pipeline. The same principles that make your Bronze-Silver-Gold layers reliable are what make an AI agent trustworthy.
If you have worked through the Medallion Architecture chapter, you already understand the foundation.
Let us break down what Agent Bricks actually builds when you create an agent. Every component maps directly to a data engineering concept you can learn and practice.
graph TD
A[User Query] --> B[Retrieval Layer]
B --> C[Context Assembly]
C --> D[Reasoning Engine]
D --> E{Decision Point}
E -->|Need More Data| B
E -->|Ready to Act| F[Tool Execution]
F --> G[Response Generation]
G --> H[Guardrails Check]
H --> I[Output to User]
J[Unity Catalog] --> B
K[Vector Search Index] --> B
L[Delta Tables] --> B
M[Feature Store] --> D
N[MLflow Tracking] --> D
style A fill:#1a365d,color:#fff
style I fill:#1a365d,color:#fff
style J fill:#2d5016,color:#fff
style K fill:#2d5016,color:#fff
style L fill:#2d5016,color:#fff
style M fill:#744210,color:#fff
style N fill:#744210,color:#fffEvery agent begins with retrieval. When a user asks a question, the agent needs to find the right data to reason about.
In Agent Bricks, retrieval pulls from three sources:
Delta Tables, Your structured, governed, versioned data. The same tables you build in your Delta Lake pipelines serve as the agent's knowledge base.
Vector Search Indexes, Unstructured data (documents, logs, support tickets) embedded as vectors and indexed for semantic search. These indexes are built on top of Delta tables and governed by Unity Catalog. Note that Unity Catalog now supports a native FILE type to govern unstructured assets like PDFs and video.
Feature Store, Pre-computed entity features (customer risk scores, product embeddings, usage patterns) that give agents contextual understanding. This is covered in detail in the Machine Learning chapter.
Here is what retrieval looks like in practice:
Note: The code examples below are illustrative of the Agent Bricks pattern. Refer to the latest Databricks documentation for current SDK syntax, as the API evolves with new releases.
from databricks.agents import AgentBricks
from databricks.vector_search import VectorSearchClient
# Create a vector search index on your Delta table
vsc = VectorSearchClient()
index = vsc.create_delta_sync_index(
endpoint_name="agent-retrieval",
source_table_name="gold.support_tickets",
index_name="gold.support_ticket_index",
pipeline_type="TRIGGERED",
primary_key="ticket_id",
embedding_source_column="description",
embedding_model_endpoint_name="databricks-bge-large-en"
)Notice what is happening here. The vector index syncs from a Delta table, which means your data quality rules, schema evolution, and access controls all carry forward. If you have read the Data Quality chapter, you know why this matters.
Bad data in your Delta table means bad retrieval for your agent. There is no shortcut around this.
Raw retrieval results are like Bronze-layer data. They are relevant but noisy. Context assembly is the Silver layer, cleaning, deduplicating, ranking, and structuring the retrieved information before it reaches the reasoning engine.
Agent Bricks handles this with configurable retrieval strategies:
from databricks.agents import RetrievalConfig
retrieval_config = RetrievalConfig(
vector_search_index="gold.support_ticket_index",
max_results=10,
min_similarity_score=0.7,
filters={
"status": "open",
"priority": ["high", "critical"]
},
reranking_model="databricks-reranker-v1"
)The filtering here uses the same column-level logic you practice in Transformations and Spark SQL. The reranking step is essentially a quality gate, the same pattern you apply when validating data between pipeline stages.
This is where the language model lives. But in Agent Bricks, reasoning is not just "send prompt, get response." It is a structured decision loop.
The agent can:
from databricks.agents import AgentConfig, ToolConfig
agent = AgentConfig(
model="databricks-meta-llama-3-1-70b-instruct",
retrieval=retrieval_config,
tools=[
ToolConfig(
name="query_sales_data",
description="Query sales metrics from the Gold layer",
function=query_gold_sales_table
),
ToolConfig(
name="check_inventory",
description="Check current inventory levels",
function=check_inventory_api
)
],
system_prompt="""You are a business analyst agent for Acme Corp.
Use the available tools to answer questions about sales performance
and inventory status. Always cite the data source in your response.""",
max_iterations=5
)That max_iterations parameter is important. It prevents the agent from entering infinite reasoning loops, a production concern that mirrors the circuit-breaker patterns you learn about in Debugging and Monitoring.
This is what separates agents from chatbots. Tools let the agent take action: query databases, call APIs, trigger workflows, update records.
In Agent Bricks, tools are Python functions registered with the agent. They are governed by Unity Catalog, meaning access control, lineage, and audit logging apply automatically. Governance is further enhanced by the Unity AI Gateway for runtime tool and agent monitoring.
def query_gold_sales_table(region: str, date_range: str) -> str:
"""Query the Gold sales table for regional metrics."""
from pyspark.sql import functions as F
sales_df = spark.table("gold.regional_sales")
result = (
sales_df
.filter(F.col("region") == region)
.filter(F.col("sale_date").between(*parse_date_range(date_range)))
.agg(
F.sum("revenue").alias("total_revenue"),
F.count("order_id").alias("order_count"),
F.avg("order_value").alias("avg_order_value")
)
.collect()[0]
)
return f"Region {region}: ${result.total_revenue:,.2f} revenue, " \
f"{result.order_count} orders, ${result.avg_order_value:,.2f} avg"This is just PySpark. The same DataFrames and Joins and Aggregations patterns you practice in BricksNotes chapters. The agent calls this function when it decides it needs sales data, and the result flows back into its reasoning loop.
Every production pipeline has quality gates. Agent Bricks applies the same principle to AI outputs with guardrails:
from databricks.agents import GuardrailConfig
guardrails = GuardrailConfig(
input_guardrails=[
"no_pii_in_queries",
"topic_relevance_check"
],
output_guardrails=[
"no_hallucination_check",
"no_pii_in_responses",
"toxicity_filter"
]
)Think of guardrails as the agent equivalent of Data Quality checks. Input guardrails are like schema validation on ingestion. Output guardrails are like assertion tests before writing to Gold tables. The pattern is identical, only the data type changes from rows to natural language.
One of the most innovative aspects of Agent Bricks is ALHF, Agent Learning from Human Feedback.
Traditional ML requires retraining models. ALHF lets domain experts provide natural-language guidance that the system automatically translates into technical optimizations:
"When answering questions about quarterly revenue, always include year-over-year comparison."
The agent incorporates this feedback without any code changes. It is like having a business user define transformation rules in plain English, which then get compiled into the agent's reasoning pipeline.
This connects directly to the philosophy we explore in our Genie deep-dive article: the most powerful AI systems are the ones that learn from the people closest to the data, now supported by Genie Ontology for live business context.
Agent Bricks uses MLflow for everything you would expect, and some things you might not.
Experiment Tracking: Every agent configuration, prompt variation, and retrieval strategy is logged as an MLflow experiment. You can compare agent versions the same way you compare model versions.
Evaluation: Agent Bricks includes built-in evaluation metrics:
import mlflow
with mlflow.start_run(run_name="support_agent_v2"):
mlflow.log_params({
"model": "llama-3-1-70b",
"max_retrieval_results": 10,
"min_similarity": 0.7,
"max_iterations": 5
})
eval_results = mlflow.evaluate(
model=agent,
data=eval_dataset,
targets="expected_response",
model_type="agent",
evaluators=[
"relevance",
"groundedness",
"safety",
"latency"
]
)
mlflow.log_metrics(eval_results.metrics)Model Serving: Agents deploy as Model Serving endpoints, the same infrastructure that serves ML models. Auto-scaling, A/B testing, and monitoring come free.
This entire workflow is an extension of what you learn in the Machine Learning chapter. The tools are the same. The patterns are the same. The data types are different.
Every component of an Agent Bricks agent is registered in Unity Catalog:
| Component | Unity Catalog Object | Governance |
|---|---|---|
| Source data | Delta tables | Row/column ACLs, lineage |
| Vector indexes | Vector search indexes | Access control, audit |
| Tools | Functions | Permission grants, logging |
| Models | Registered models | Versioning, approval workflows |
| Agent endpoints | Serving endpoints | Rate limiting, access control |
| Evaluation data | Delta tables | Data sharing, lineage |
This is not optional decoration. It is the reason agents built on Databricks can operate in regulated industries. Unity Catalog now includes Glossary and Domains to further organize these assets.
When an agent accesses a customer record, Unity Catalog logs who accessed what, when, and why. When an agent's tool queries a table, column-level masking ensures it only sees what it should. When an agent is promoted from staging to production, the approval workflow tracks every change.
If you have completed the Unity Catalog chapter, you understand these governance patterns already. Agent Bricks simply extends them to AI.
Here is how the pieces fit together in practice. Follow this path to build a production-ready agent:
Before writing a single line of agent code, your data must be ready.
For unstructured data (documents, support tickets, product descriptions), create vector search indexes that sync from your Delta tables.
For entity-level context (customer profiles, product attributes, risk scores), use Feature Store tables. These give your agent structured, pre-computed knowledge.
Write Python functions that the agent can call. Each tool should do one thing well, query a table, call an API, perform a calculation. Register them in Unity Catalog.
Set up the agent configuration, run evaluation suites, iterate on prompts and retrieval strategies using MLflow.
Deploy via Model Serving with input and output guardrails. Monitor using the same Debugging and Monitoring patterns you use for pipelines.
Collect human feedback, incorporate domain expertise, and continuously improve agent behavior without retraining.
This is the part that matters most. Every skill you are building with BricksNotes maps directly to Agent Bricks.
| Data Engineering Skill | Agent Bricks Application | BricksNotes Chapter |
|---|---|---|
| Delta Lake fundamentals | Agent knowledge base storage | Delta Lake |
| Medallion Architecture | Bronze-Silver-Gold for agent data | Medallion Architecture |
| Data quality checks | Input/output guardrails | Data Quality |
| Schema evolution | Handling changing data sources | Schema Evolution |
| Incremental processing | Efficient index updates | Incremental Processing |
| Spark SQL | Agent tool functions | Spark SQL |
| DataFrames | Data retrieval and transformation | DataFrames |
| Joins and aggregations | Complex tool queries | Joins and Aggregations |
| UDFs | Custom agent tool logic | UDFs |
| Streaming | Real-time agent data feeds | Streaming |
| File formats | Optimized storage for retrieval | File Formats |
| Partitioning | Fast agent data access | Partitioning and Performance |
| Unit testing | Agent evaluation and testing | Unit Testing |
| Workflows | Agent deployment pipelines | Workflows |
| Debugging | Agent observability | Debugging and Monitoring |
| ML pipelines | Model serving and tracking | Machine Learning |
| Unity Catalog | Agent governance | Unity Catalog |
| SCD patterns | Historical context for agents | SCD Patterns |
Agent Bricks inherits the lakehouse cost model: storage and compute are decoupled.
Your Delta tables store data cheaply on object storage. Vector indexes are computed once and updated incrementally. Model Serving auto-scales to zero when not in use.
Compare this to traditional approaches where you pay for a vector database, a separate model hosting service, and a retrieval infrastructure, all disconnected from your governance layer.
If you read our Cost-Efficient Pipelines article, you already understand why compute-storage decoupling matters. Agent Bricks extends that principle to AI workloads.
Companies using Agent Bricks are seeing transformative results:
Mastercard automated merchant onboarding with agents that process applications, verify compliance, and route decisions, reducing onboarding time by 70%.
AT&T built fraud detection agents that analyze call patterns, cross-reference customer history, and flag suspicious activity in real time.
AstraZeneca deployed research agents that search clinical trial databases, summarize findings, and suggest experimental directions, grounded in governed, auditable data.
These are not experimental demos. They are production systems handling real business decisions.
The Agentic Enterprise is not replacing data engineers. It is promoting them.
The engineers who understand how Delta Lake handles transactions, why data quality matters at every layer, how Unity Catalog governs access, and why the Medallion Architecture creates reliable data products, those engineers are the ones building the AI systems that actually work.
Agent Bricks is not a new discipline. It is data engineering with a new output format.
The pipeline you are learning to build today is the same pipeline that will power autonomous agents tomorrow. The fundamentals do not change. The stakes just get higher.
The best AI agents are built by people who understand data, not just models. Start with the foundation. The agents will follow.
Ready to build that foundation? Start with Chapter 1 and work through the complete BricksNotes certification path. Every chapter is a building block for the Agentic Enterprise.
Want to see how the broader Agentic Enterprise framework connects these concepts? Read The Agentic Enterprise: Why Your Data Engineering Skills Are the Foundation of Autonomous AI. For a deep dive into how Genie uses similar patterns for natural-language analytics, check out Databricks Genie: The AI Layer That Turns Your Lakehouse Into a Conversation. And to see how these patterns are already defending enterprises against AI-driven threats, explore Databricks Lakewatch: How the Lakehouse Architecture Is Reshaping Security Operations.