Agent Bricks: Build Production AI Agents on Databricks

From retrieval to reasoning to action — the complete architecture of enterprise AI agents, mapped to the data engineering skills you already have.

Databricks Agent Bricks: How to Build Production-Grade AI Agents That Actually Understand Your Data

Most AI agents fail in production.

Not because the models are bad. Not because the prompts are wrong. They fail because the data underneath them is messy, ungoverned, and disconnected from the business context that makes decisions meaningful.

Databricks knows this. That is why they built Agent Bricks, not as another chatbot framework, but as a system that treats AI agents the same way data engineers treat pipelines: with structure, testing, governance, and observability baked in from the start.

If you have been following the Agentic Enterprise conversation, Agent Bricks is where theory becomes practice. With the launch of Genie One and Genie Ontology, agents now have a live business context layer to draw from.

What Are Agent Bricks, Really?

Agent Bricks is Databricks' product for building domain-specific AI agents that are grounded in enterprise data. It sits within the Mosaic AI platform and integrates with everything you already know: Delta Lake, Unity Catalog, MLflow, and Spark.

But here is the key insight most people miss.

Agent Bricks does not just wrap a language model in an API. It builds what Databricks calls a Compound AI System, multiple components (retrieval, reasoning, tools, memory, guardrails) working together as a pipeline.

Sound familiar? It should.

A compound AI system is architecturally identical to a data pipeline. The same principles that make your Bronze-Silver-Gold layers reliable are what make an AI agent trustworthy.

If you have worked through the Medallion Architecture chapter, you already understand the foundation.

The Anatomy of a Production Agent

Let us break down what Agent Bricks actually builds when you create an agent. Every component maps directly to a data engineering concept you can learn and practice.

graph TD
    A[User Query] --> B[Retrieval Layer]
    B --> C[Context Assembly]
    C --> D[Reasoning Engine]
    D --> E{Decision Point}
    E -->|Need More Data| B
    E -->|Ready to Act| F[Tool Execution]
    F --> G[Response Generation]
    G --> H[Guardrails Check]
    H --> I[Output to User]
    
    J[Unity Catalog] --> B
    K[Vector Search Index] --> B
    L[Delta Tables] --> B
    M[Feature Store] --> D
    N[MLflow Tracking] --> D
    
    style A fill:#1a365d,color:#fff
    style I fill:#1a365d,color:#fff
    style J fill:#2d5016,color:#fff
    style K fill:#2d5016,color:#fff
    style L fill:#2d5016,color:#fff
    style M fill:#744210,color:#fff
    style N fill:#744210,color:#fff

Layer 1: Retrieval, Where Data Engineering Starts

Every agent begins with retrieval. When a user asks a question, the agent needs to find the right data to reason about.

In Agent Bricks, retrieval pulls from three sources:

Delta Tables, Your structured, governed, versioned data. The same tables you build in your Delta Lake pipelines serve as the agent's knowledge base.

Vector Search Indexes, Unstructured data (documents, logs, support tickets) embedded as vectors and indexed for semantic search. These indexes are built on top of Delta tables and governed by Unity Catalog. Note that Unity Catalog now supports a native FILE type to govern unstructured assets like PDFs and video.

Feature Store, Pre-computed entity features (customer risk scores, product embeddings, usage patterns) that give agents contextual understanding. This is covered in detail in the Machine Learning chapter.

Here is what retrieval looks like in practice:

Note: The code examples below are illustrative of the Agent Bricks pattern. Refer to the latest Databricks documentation for current SDK syntax, as the API evolves with new releases.

from databricks.agents import AgentBricks
from databricks.vector_search import VectorSearchClient

# Create a vector search index on your Delta table
vsc = VectorSearchClient()

index = vsc.create_delta_sync_index(
    endpoint_name="agent-retrieval",
    source_table_name="gold.support_tickets",
    index_name="gold.support_ticket_index",
    pipeline_type="TRIGGERED",
    primary_key="ticket_id",
    embedding_source_column="description",
    embedding_model_endpoint_name="databricks-bge-large-en"
)

Notice what is happening here. The vector index syncs from a Delta table, which means your data quality rules, schema evolution, and access controls all carry forward. If you have read the Data Quality chapter, you know why this matters.

Bad data in your Delta table means bad retrieval for your agent. There is no shortcut around this.

Layer 2: Context Assembly, The Silver Layer for Agents

Raw retrieval results are like Bronze-layer data. They are relevant but noisy. Context assembly is the Silver layer, cleaning, deduplicating, ranking, and structuring the retrieved information before it reaches the reasoning engine.

Agent Bricks handles this with configurable retrieval strategies:

from databricks.agents import RetrievalConfig

retrieval_config = RetrievalConfig(
    vector_search_index="gold.support_ticket_index",
    max_results=10,
    min_similarity_score=0.7,
    filters={
        "status": "open",
        "priority": ["high", "critical"]
    },
    reranking_model="databricks-reranker-v1"
)

The filtering here uses the same column-level logic you practice in Transformations and Spark SQL. The reranking step is essentially a quality gate, the same pattern you apply when validating data between pipeline stages.

Layer 3: Reasoning, Where the Model Thinks

This is where the language model lives. But in Agent Bricks, reasoning is not just "send prompt, get response." It is a structured decision loop.

The agent can:

from databricks.agents import AgentConfig, ToolConfig

agent = AgentConfig(
    model="databricks-meta-llama-3-1-70b-instruct",
    retrieval=retrieval_config,
    tools=[
        ToolConfig(
            name="query_sales_data",
            description="Query sales metrics from the Gold layer",
            function=query_gold_sales_table
        ),
        ToolConfig(
            name="check_inventory",
            description="Check current inventory levels",
            function=check_inventory_api
        )
    ],
    system_prompt="""You are a business analyst agent for Acme Corp.
    Use the available tools to answer questions about sales performance
    and inventory status. Always cite the data source in your response.""",
    max_iterations=5
)

That max_iterations parameter is important. It prevents the agent from entering infinite reasoning loops, a production concern that mirrors the circuit-breaker patterns you learn about in Debugging and Monitoring.

Layer 4: Tool Execution, Agents That Do Things

This is what separates agents from chatbots. Tools let the agent take action: query databases, call APIs, trigger workflows, update records.

In Agent Bricks, tools are Python functions registered with the agent. They are governed by Unity Catalog, meaning access control, lineage, and audit logging apply automatically. Governance is further enhanced by the Unity AI Gateway for runtime tool and agent monitoring.

def query_gold_sales_table(region: str, date_range: str) -> str:
    """Query the Gold sales table for regional metrics."""
    from pyspark.sql import functions as F
    
    sales_df = spark.table("gold.regional_sales")
    
    result = (
        sales_df
        .filter(F.col("region") == region)
        .filter(F.col("sale_date").between(*parse_date_range(date_range)))
        .agg(
            F.sum("revenue").alias("total_revenue"),
            F.count("order_id").alias("order_count"),
            F.avg("order_value").alias("avg_order_value")
        )
        .collect()[0]
    )
    
    return f"Region {region}: ${result.total_revenue:,.2f} revenue, "  \
           f"{result.order_count} orders, ${result.avg_order_value:,.2f} avg"

This is just PySpark. The same DataFrames and Joins and Aggregations patterns you practice in BricksNotes chapters. The agent calls this function when it decides it needs sales data, and the result flows back into its reasoning loop.

Layer 5: Guardrails, The Data Quality of AI

Every production pipeline has quality gates. Agent Bricks applies the same principle to AI outputs with guardrails:

from databricks.agents import GuardrailConfig

guardrails = GuardrailConfig(
    input_guardrails=[
        "no_pii_in_queries",
        "topic_relevance_check"
    ],
    output_guardrails=[
        "no_hallucination_check",
        "no_pii_in_responses",
        "toxicity_filter"
    ]
)

Think of guardrails as the agent equivalent of Data Quality checks. Input guardrails are like schema validation on ingestion. Output guardrails are like assertion tests before writing to Gold tables. The pattern is identical, only the data type changes from rows to natural language.

Agent Learning from Human Feedback (ALHF)

One of the most innovative aspects of Agent Bricks is ALHF, Agent Learning from Human Feedback.

Traditional ML requires retraining models. ALHF lets domain experts provide natural-language guidance that the system automatically translates into technical optimizations:

"When answering questions about quarterly revenue, always include year-over-year comparison."

The agent incorporates this feedback without any code changes. It is like having a business user define transformation rules in plain English, which then get compiled into the agent's reasoning pipeline.

This connects directly to the philosophy we explore in our Genie deep-dive article: the most powerful AI systems are the ones that learn from the people closest to the data, now supported by Genie Ontology for live business context.

The MLflow Integration: Experiment Tracking for Agents

Agent Bricks uses MLflow for everything you would expect, and some things you might not.

Experiment Tracking: Every agent configuration, prompt variation, and retrieval strategy is logged as an MLflow experiment. You can compare agent versions the same way you compare model versions.

Evaluation: Agent Bricks includes built-in evaluation metrics:

import mlflow

with mlflow.start_run(run_name="support_agent_v2"):
    mlflow.log_params({
        "model": "llama-3-1-70b",
        "max_retrieval_results": 10,
        "min_similarity": 0.7,
        "max_iterations": 5
    })
    
    eval_results = mlflow.evaluate(
        model=agent,
        data=eval_dataset,
        targets="expected_response",
        model_type="agent",
        evaluators=[
            "relevance",
            "groundedness",
            "safety",
            "latency"
        ]
    )
    
    mlflow.log_metrics(eval_results.metrics)

Model Serving: Agents deploy as Model Serving endpoints, the same infrastructure that serves ML models. Auto-scaling, A/B testing, and monitoring come free.

This entire workflow is an extension of what you learn in the Machine Learning chapter. The tools are the same. The patterns are the same. The data types are different.

Unity Catalog: The Governance Backbone

Every component of an Agent Bricks agent is registered in Unity Catalog:

ComponentUnity Catalog ObjectGovernance
Source dataDelta tablesRow/column ACLs, lineage
Vector indexesVector search indexesAccess control, audit
ToolsFunctionsPermission grants, logging
ModelsRegistered modelsVersioning, approval workflows
Agent endpointsServing endpointsRate limiting, access control
Evaluation dataDelta tablesData sharing, lineage

This is not optional decoration. It is the reason agents built on Databricks can operate in regulated industries. Unity Catalog now includes Glossary and Domains to further organize these assets.

When an agent accesses a customer record, Unity Catalog logs who accessed what, when, and why. When an agent's tool queries a table, column-level masking ensures it only sees what it should. When an agent is promoted from staging to production, the approval workflow tracks every change.

If you have completed the Unity Catalog chapter, you understand these governance patterns already. Agent Bricks simply extends them to AI.

Building Your First Agent: A Step-by-Step Path

Here is how the pieces fit together in practice. Follow this path to build a production-ready agent:

Step 1: Build the Data Foundation

Before writing a single line of agent code, your data must be ready.

Step 2: Create Vector Search Indexes

For unstructured data (documents, support tickets, product descriptions), create vector search indexes that sync from your Delta tables.

Step 3: Build Feature Tables

For entity-level context (customer profiles, product attributes, risk scores), use Feature Store tables. These give your agent structured, pre-computed knowledge.

Step 4: Define Tools

Write Python functions that the agent can call. Each tool should do one thing well, query a table, call an API, perform a calculation. Register them in Unity Catalog.

Step 5: Configure and Test

Set up the agent configuration, run evaluation suites, iterate on prompts and retrieval strategies using MLflow.

Step 6: Deploy with Guardrails

Deploy via Model Serving with input and output guardrails. Monitor using the same Debugging and Monitoring patterns you use for pipelines.

Step 7: Iterate with ALHF

Collect human feedback, incorporate domain expertise, and continuously improve agent behavior without retraining.

Where Agent Bricks Connects to Your Existing Skills

This is the part that matters most. Every skill you are building with BricksNotes maps directly to Agent Bricks.

Data Engineering SkillAgent Bricks ApplicationBricksNotes Chapter
Delta Lake fundamentalsAgent knowledge base storageDelta Lake
Medallion ArchitectureBronze-Silver-Gold for agent dataMedallion Architecture
Data quality checksInput/output guardrailsData Quality
Schema evolutionHandling changing data sourcesSchema Evolution
Incremental processingEfficient index updatesIncremental Processing
Spark SQLAgent tool functionsSpark SQL
DataFramesData retrieval and transformationDataFrames
Joins and aggregationsComplex tool queriesJoins and Aggregations
UDFsCustom agent tool logicUDFs
StreamingReal-time agent data feedsStreaming
File formatsOptimized storage for retrievalFile Formats
PartitioningFast agent data accessPartitioning and Performance
Unit testingAgent evaluation and testingUnit Testing
WorkflowsAgent deployment pipelinesWorkflows
DebuggingAgent observabilityDebugging and Monitoring
ML pipelinesModel serving and trackingMachine Learning
Unity CatalogAgent governanceUnity Catalog
SCD patternsHistorical context for agentsSCD Patterns

The Cost Dimension

Agent Bricks inherits the lakehouse cost model: storage and compute are decoupled.

Your Delta tables store data cheaply on object storage. Vector indexes are computed once and updated incrementally. Model Serving auto-scales to zero when not in use.

Compare this to traditional approaches where you pay for a vector database, a separate model hosting service, and a retrieval infrastructure, all disconnected from your governance layer.

If you read our Cost-Efficient Pipelines article, you already understand why compute-storage decoupling matters. Agent Bricks extends that principle to AI workloads.

Real-World Impact

Companies using Agent Bricks are seeing transformative results:

Mastercard automated merchant onboarding with agents that process applications, verify compliance, and route decisions, reducing onboarding time by 70%.

AT&T built fraud detection agents that analyze call patterns, cross-reference customer history, and flag suspicious activity in real time.

AstraZeneca deployed research agents that search clinical trial databases, summarize findings, and suggest experimental directions, grounded in governed, auditable data.

These are not experimental demos. They are production systems handling real business decisions.

What This Means for You

The Agentic Enterprise is not replacing data engineers. It is promoting them.

The engineers who understand how Delta Lake handles transactions, why data quality matters at every layer, how Unity Catalog governs access, and why the Medallion Architecture creates reliable data products, those engineers are the ones building the AI systems that actually work.

Agent Bricks is not a new discipline. It is data engineering with a new output format.

The pipeline you are learning to build today is the same pipeline that will power autonomous agents tomorrow. The fundamentals do not change. The stakes just get higher.

The best AI agents are built by people who understand data, not just models. Start with the foundation. The agents will follow.

Ready to build that foundation? Start with Chapter 1 and work through the complete BricksNotes certification path. Every chapter is a building block for the Agentic Enterprise.


Want to see how the broader Agentic Enterprise framework connects these concepts? Read The Agentic Enterprise: Why Your Data Engineering Skills Are the Foundation of Autonomous AI. For a deep dive into how Genie uses similar patterns for natural-language analytics, check out Databricks Genie: The AI Layer That Turns Your Lakehouse Into a Conversation. And to see how these patterns are already defending enterprises against AI-driven threats, explore Databricks Lakewatch: How the Lakehouse Architecture Is Reshaping Security Operations.