A calm, step-by-step guide to CI/CD on Databricks using Declarative Automation Bundles, formerly known as Databricks Asset Bundles.
Most data engineers learn pipelines first and deployment later. Sometimes much later. You build something in a notebook, it works, you click "Run All" in production, and you call it done.
That works until it does not. One untested change reaches production. A pipeline breaks at 3 AM. Now you are reading logs in the dark and trying to remember what you changed.
CI/CD is the calm answer to this. Declarative Automation Bundles are how you do CI/CD on Databricks. This guide is for engineers who have never done either.
CI stands for Continuous Integration. Every time you change code, an automated process runs your tests before the change is allowed in. If the tests fail, the change does not move forward.
CD stands for Continuous Deployment. Once the tests pass, the change is deployed to the next environment automatically. No one clicks a button. No one runs a notebook.
For a data engineer this means a pipeline change goes from your laptop, through a Git commit, through automated tests, into a dev workspace, then a staging workspace, then production. Each step is checked. Each step is repeatable.
This is not just for software engineers. If you build pipelines that other teams depend on, you need CI/CD. The cost of one bad deploy in production is almost always larger than the time it takes to set this up properly.
Declarative Automation Bundles are the official way to package a Databricks project. You may have heard them called Databricks Asset Bundles, or DABs. Same thing, newer name.
A bundle is a folder with:
databricks.yml that describes what to deployYou describe what you want. Databricks figures out how to deploy it. You do not write API calls to create jobs. You do not click around the UI. You write the YAML, run one command, and the bundle deploys.
Think of a bundle as your pipeline plus its configuration plus its deployment rules, all in one versioned package that lives in Git.
Here is a minimal databricks.yml for a project with one pipeline:
bundle:
name: customer-events
variables:
catalog:
description: Unity Catalog to write to
default: dev_catalog
targets:
dev:
workspace:
host: https://my-workspace.cloud.databricks.com
variables:
catalog: dev_catalog
prod:
mode: production
workspace:
host: https://my-workspace.cloud.databricks.com
variables:
catalog: prod_catalog
resources:
jobs:
customer_events_job:
name: customer_events_${bundle.target}
tasks:
- task_key: load_events
notebook_task:
notebook_path: ./notebooks/load_events.py
job_cluster_key: main_cluster
job_clusters:
- job_cluster_key: main_cluster
new_cluster:
spark_version: 15.4.x-scala2.12
num_workers: 2
node_type_id: i3.xlargeLine by line:
bundle.name is the project name. Databricks uses this to prefix every resource it creates so dev and prod do not collide.
variables are values that change between environments. Here catalog lets the same pipeline write to dev_catalog in dev and prod_catalog in production without changing any code.
targets are the environments. Each target can override variables and workspace settings. mode: production enables extra safety checks like blocking deploys that would skip approvals.
resources is where you list what to create. Jobs, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Unity Catalog schemas, ML experiments. The example creates one job with one task that runs a notebook.
The ${bundle.target} interpolation means the job will be named customer_events_dev in dev and customer_events_prod in prod. You will never confuse the two.
This works in any Databricks workspace you have access to. The single-target flow runs on Free Edition. Multi-target promotion needs a paid workspace per environment.
1. Install the Databricks CLI.
curl -fsSL https://raw.githubusercontent.com/databricks/setup-cli/main/install.sh | sh2. Authenticate.
databricks configureIt asks for your workspace URL and a personal access token. Paste both.
3. Create a minimal project.
mkdir customer-events && cd customer-events
mkdir notebooks
echo 'print("hello from a bundle")' > notebooks/load_events.py4. Write the databricks.yml from the example above, with your workspace host.
5. Validate before you deploy.
databricks bundle validateThis catches typos, missing files, and broken references. Run it every time. Never skip it.
6. Deploy to dev.
databricks bundle deploy -t devDatabricks uploads your files, creates the job, and reports back the URL.
7. Run it.
databricks bundle run -t dev customer_events_jobYou will see the run id and status. Open the workspace and you will find a job named customer_events_dev with the run in its history.
That is the full loop. Validate, deploy, run.
Real teams use three environments: dev, staging, production. The flow looks like this:
feature/ branch.databricks bundle validate and any unit tests.main. CI deploys to staging with databricks bundle deploy -t staging.databricks bundle deploy -t prod.Each promotion is the same bundle, just a different target. The variables swap. The workspace swaps. The code does not.
A minimal GitHub Actions workflow looks like:
name: deploy
on:
push:
branches: [main]
jobs:
deploy-staging:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: databricks/setup-cli@main
- run: databricks bundle deploy -t staging
env:
DATABRICKS_HOST: ${{ secrets.DATABRICKS_HOST }}
DATABRICKS_TOKEN: ${{ secrets.DATABRICKS_STAGING_TOKEN }}The same pattern repeats for production, gated behind a manual approval step.
CI/CD is a domain on the Databricks Data Engineer Associate exam. You should be able to:
databricks.yml file and read its structure.bundle validate, bundle deploy, bundle run, bundle destroy.You will not be asked to write a full databricks.yml from scratch. You will be asked which command to run, what a target does, and why CI/CD matters for a pipeline.
Skipping bundle validate. It takes a second. It catches the typos that would otherwise fail mid-deploy and leave you with half-created resources.
Hardcoding paths. If your notebook has dev_catalog.sales.orders hardcoded, your bundle is not portable. Use variables.
One target for everything. If dev and prod point at the same workspace and write to the same tables, you do not have environments. You have one environment with a confusing config.
Deploying straight to prod. Use staging. Always. The cost of one extra deploy step is much smaller than the cost of one production incident.
Forgetting mode: production. This flag enables guardrails that block accidental destructive changes in your prod target. Turn it on and leave it on.