Declarative Automation Bundles for Data Engineers Who Have Never Done CI/CD

A calm, step-by-step guide to CI/CD on Databricks using Declarative Automation Bundles, formerly known as Databricks Asset Bundles.

Most data engineers learn pipelines first and deployment later. Sometimes much later. You build something in a notebook, it works, you click "Run All" in production, and you call it done.

That works until it does not. One untested change reaches production. A pipeline breaks at 3 AM. Now you are reading logs in the dark and trying to remember what you changed.

CI/CD is the calm answer to this. Declarative Automation Bundles are how you do CI/CD on Databricks. This guide is for engineers who have never done either.

What CI/CD actually means for a data engineer

CI stands for Continuous Integration. Every time you change code, an automated process runs your tests before the change is allowed in. If the tests fail, the change does not move forward.

CD stands for Continuous Deployment. Once the tests pass, the change is deployed to the next environment automatically. No one clicks a button. No one runs a notebook.

For a data engineer this means a pipeline change goes from your laptop, through a Git commit, through automated tests, into a dev workspace, then a staging workspace, then production. Each step is checked. Each step is repeatable.

This is not just for software engineers. If you build pipelines that other teams depend on, you need CI/CD. The cost of one bad deploy in production is almost always larger than the time it takes to set this up properly.

What Declarative Automation Bundles are

Declarative Automation Bundles are the official way to package a Databricks project. You may have heard them called Databricks Asset Bundles, or DABs. Same thing, newer name.

A bundle is a folder with:

You describe what you want. Databricks figures out how to deploy it. You do not write API calls to create jobs. You do not click around the UI. You write the YAML, run one command, and the bundle deploys.

Think of a bundle as your pipeline plus its configuration plus its deployment rules, all in one versioned package that lives in Git.

The databricks.yml file

Here is a minimal databricks.yml for a project with one pipeline:

bundle:
  name: customer-events

variables:
  catalog:
    description: Unity Catalog to write to
    default: dev_catalog

targets:
  dev:
    workspace:
      host: https://my-workspace.cloud.databricks.com
    variables:
      catalog: dev_catalog

  prod:
    mode: production
    workspace:
      host: https://my-workspace.cloud.databricks.com
    variables:
      catalog: prod_catalog

resources:
  jobs:
    customer_events_job:
      name: customer_events_${bundle.target}
      tasks:
        - task_key: load_events
          notebook_task:
            notebook_path: ./notebooks/load_events.py
          job_cluster_key: main_cluster
      job_clusters:
        - job_cluster_key: main_cluster
          new_cluster:
            spark_version: 15.4.x-scala2.12
            num_workers: 2
            node_type_id: i3.xlarge

Line by line:

bundle.name is the project name. Databricks uses this to prefix every resource it creates so dev and prod do not collide.

variables are values that change between environments. Here catalog lets the same pipeline write to dev_catalog in dev and prod_catalog in production without changing any code.

targets are the environments. Each target can override variables and workspace settings. mode: production enables extra safety checks like blocking deploys that would skip approvals.

resources is where you list what to create. Jobs, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Unity Catalog schemas, ML experiments. The example creates one job with one task that runs a notebook.

The ${bundle.target} interpolation means the job will be named customer_events_dev in dev and customer_events_prod in prod. You will never confuse the two.

Your first deployment, step by step

This works in any Databricks workspace you have access to. The single-target flow runs on Free Edition. Multi-target promotion needs a paid workspace per environment.

1. Install the Databricks CLI.

curl -fsSL https://raw.githubusercontent.com/databricks/setup-cli/main/install.sh | sh

2. Authenticate.

databricks configure

It asks for your workspace URL and a personal access token. Paste both.

3. Create a minimal project.

mkdir customer-events && cd customer-events
mkdir notebooks
echo 'print("hello from a bundle")' > notebooks/load_events.py

4. Write the databricks.yml from the example above, with your workspace host.

5. Validate before you deploy.

databricks bundle validate

This catches typos, missing files, and broken references. Run it every time. Never skip it.

6. Deploy to dev.

databricks bundle deploy -t dev

Databricks uploads your files, creates the job, and reports back the URL.

7. Run it.

databricks bundle run -t dev customer_events_job

You will see the run id and status. Open the workspace and you will find a job named customer_events_dev with the run in its history.

That is the full loop. Validate, deploy, run.

Promoting code across environments

Real teams use three environments: dev, staging, production. The flow looks like this:

Each promotion is the same bundle, just a different target. The variables swap. The workspace swaps. The code does not.

A minimal GitHub Actions workflow looks like:

name: deploy
on:
  push:
    branches: [main]

jobs:
  deploy-staging:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: databricks/setup-cli@main
      - run: databricks bundle deploy -t staging
        env:
          DATABRICKS_HOST: ${{ secrets.DATABRICKS_HOST }}
          DATABRICKS_TOKEN: ${{ secrets.DATABRICKS_STAGING_TOKEN }}

The same pattern repeats for production, gated behind a manual approval step.

What the exam expects

CI/CD is a domain on the Databricks Data Engineer Associate exam. You should be able to:

You will not be asked to write a full databricks.yml from scratch. You will be asked which command to run, what a target does, and why CI/CD matters for a pipeline.

Common mistakes

Skipping bundle validate. It takes a second. It catches the typos that would otherwise fail mid-deploy and leave you with half-created resources.

Hardcoding paths. If your notebook has dev_catalog.sales.orders hardcoded, your bundle is not portable. Use variables.

One target for everything. If dev and prod point at the same workspace and write to the same tables, you do not have environments. You have one environment with a confusing config.

Deploying straight to prod. Use staging. Always. The cost of one extra deploy step is much smaller than the cost of one production incident.

Forgetting mode: production. This flag enables guardrails that block accidental destructive changes in your prod target. Turn it on and leave it on.

Continue learning