Databricks Lakeflow Connect: Simplifying Data Ingestion from SaaS and Databases

How managed connectors eliminate custom ingestion code and bring governed, incremental data into your lakehouse.

Every data engineering team eventually faces the same problem. You have data in Salesforce. You have data in PostgreSQL. You have data in ServiceNow, SharePoint, Google Analytics, and a dozen other systems.

Getting that data into your lakehouse reliably, incrementally, and with proper governance is where most teams spend a disproportionate amount of time.

Databricks Lakeflow Connect was built to solve exactly this. It provides fully-managed connectors that handle authentication, incremental reads, schema evolution, and retries. You configure the source and destination. Lakeflow Connect handles everything in between.

What Lakeflow Connect Actually Does

Lakeflow Connect is a managed ingestion layer built directly into the Databricks platform. Every pipeline it creates is governed by Unity Catalog and powered by serverless compute through Lakeflow Declarative Pipelines.

There are two categories of connectors:

SaaS connectors ingest data from enterprise applications like Salesforce, ServiceNow, Google Analytics, HubSpot, Workday, and others. These connectors handle API pagination, rate limiting, incremental reads, and schema changes automatically.

Database connectors ingest data from relational databases like MySQL, PostgreSQL, and SQL Server using change data capture (CDC). These connectors track changes at the source and apply them incrementally to Delta tables, keeping your lakehouse in sync without full reloads.

The key difference from writing custom ingestion code: you do not manage compute, handle retries, build checkpointing logic, or worry about schema drift. Lakeflow Connect handles all of it.

Supported Connectors

As of June 2026, Lakeflow Connect is generally available and supports a growing list of enterprise sources across both categories.

SaaS Connectors

ConnectorStatusWhat It Ingests
SalesforceGACRM records, customer data
ServiceNowGAIT service management records, tickets
Google AnalyticsGAWebsite traffic, user behavior data
Workday ReportsGACustom Workday reports
Microsoft Dynamics 365GACRM and ERP data
NetSuiteGAERP and financial data
HubSpotGACRM, marketing, sales activity
ConfluenceGAPages, spaces, comments
JiraGAProject management, issue tracking
SharePointGAFiles and list data
Google AdsGACampaign performance metrics
Meta AdsGAFacebook and Instagram Ads data
TikTok AdsGAAdvertising performance data
Workday HCMGAHR, payroll, workforce data
Zendesk SupportGASupport tickets, users, organizations

Database Connectors (CDC)

ConnectorStatus
SQL ServerGA
MySQLGA
PostgreSQLGA

Query-based connectors for over 100 sources are also generally available as of May 29 2026.

How SaaS Connectors Work

A SaaS connector consists of three components:

  1. Connection — A Unity Catalog securable object that stores authentication details for the source application. This means credentials are governed by the same access controls as your tables and volumes.
  1. Ingestion pipeline — A serverless pipeline that copies data from the application into destination tables. You do not provision or manage compute. The pipeline handles incremental reads, retries, and scheduling.
  1. Destination tables — Streaming tables in Delta Lake format. These are regular Delta tables with built-in support for incremental processing, which means downstream transformations can efficiently read only new or changed records.

The setup is straightforward. You create a connection, point it at the source objects you want, specify a destination catalog and schema, and start the pipeline. No Spark code. No custom API clients. No checkpoint management.

-- Example: Creating a Salesforce connection
CREATE CONNECTION salesforce_crm
TYPE salesforce
OPTIONS (
  login_url = 'https://login.salesforce.com',
  client_id = SECRET('scope', 'sf_client_id'),
  client_secret = SECRET('scope', 'sf_client_secret')
);

Once the connection exists, you create an ingestion pipeline that references it:

-- Example: Creating an ingestion pipeline from Salesforce
CREATE OR REFRESH STREAMING TABLE salesforce_accounts
AS SELECT * FROM READ_SALESFORCE(
  connection => 'salesforce_crm',
  object => 'Account'
);

Lakeflow Connect handles the OAuth flow, API pagination, rate limit backoff, and incremental state tracking. If the source schema changes (new fields added, fields removed), the connector adapts automatically based on your schema evolution settings.

How Database Connectors Work

Database connectors are architecturally different because they use change data capture (CDC) to track modifications at the source.

The components are:

  1. Connection — Same Unity Catalog securable object for authentication.
  1. Ingestion gateway — A continuously running pipeline that extracts snapshots, change logs, and metadata from the source database. The gateway runs on classic compute because it needs to maintain a persistent connection to the source database's transaction log.
  1. Staging storage — A Unity Catalog volume that temporarily stores extracted data before it is applied to destination tables. Data is automatically purged after 30 days. This decouples extraction from loading, so your ingestion pipeline can run on whatever schedule you choose.
  1. Ingestion pipeline — A serverless pipeline that moves data from staging into destination tables.
  1. Destination tables — Streaming tables in Delta Lake.

This architecture means the gateway captures every change as it happens, while the ingestion pipeline can batch-apply those changes on a schedule. You never lose data because the staging layer acts as a buffer.

Key Database Connector Features

-- Example: Creating a PostgreSQL CDC connection
CREATE CONNECTION postgres_prod
TYPE postgresql
OPTIONS (
  host = 'prod-db.example.com',
  port = '5432',
  user = SECRET('scope', 'pg_user'),
  password = SECRET('scope', 'pg_password')
);

Unity Catalog Governance

This is what separates Lakeflow Connect from third-party ingestion tools.

Every connection is a Unity Catalog securable object. This means:

When a Lakeflow Connect pipeline writes data, the destination tables are automatically tracked in Unity Catalog's lineage graph. You can see exactly which source system produced which table, when it was last updated, and who has access to it.

For organizations that need to demonstrate data provenance for compliance (GDPR, SOC 2, HIPAA), this built-in governance eliminates the need for separate lineage tracking tools. Delta Sharing has also evolved into OpenSharing, an open protocol for sharing these governed assets across platforms.

When to Use Lakeflow Connect vs Custom Ingestion

Lakeflow Connect is the right choice when:

Custom ingestion code is still necessary when:

For most enterprise data engineering teams, Lakeflow Connect handles 60 to 80 percent of ingestion workloads. The remaining sources that need custom code can use Auto Loader for files or custom Spark structured streaming jobs.

How It Fits Into the Medallion Architecture

Lakeflow Connect naturally maps to the Bronze layer of a medallion architecture.

The destination tables from Lakeflow Connect are your Bronze tables. They contain raw, incrementally updated data from source systems. From there, your Lakeflow Declarative Pipelines (formerly Delta Live Tables) handle Silver and Gold transformations.

The flow looks like this:

  1. Lakeflow Connect ingests raw data into Bronze streaming tables
  2. Lakeflow Declarative Pipelines clean, validate, and enrich data into Silver tables
  3. Gold tables aggregate Silver data into business-ready marts

Because both Lakeflow Connect and Lakeflow Declarative Pipelines are governed by Unity Catalog and run on serverless compute, the entire pipeline from source to BI is managed, governed, and cost-efficient.

Production Considerations

Monitoring

Lakeflow Connect pipelines emit metrics and events that you can monitor through the Databricks UI or programmatically. Key metrics to watch:

Cost

SaaS connectors run on serverless compute, so you pay only for processing time. Database connectors require the ingestion gateway to run on classic compute (continuously), but the actual data loading runs serverless.

For high-volume CDC scenarios, the gateway compute cost is the primary consideration. Size it based on your source database's change rate, not the total data volume.

Schema Evolution Strategy

Decide upfront how to handle schema changes:

For most production setups, auto-merge with monitoring alerts is the recommended approach. You want schema changes to flow through, but you also want to know when they happen.

Getting Started

If you are evaluating Lakeflow Connect for your organization:

  1. Start with a GA connector (Salesforce, ServiceNow, Google Analytics, or SQL Server)
  2. Create a connection using Unity Catalog
  3. Set up a small ingestion pipeline for one or two source objects
  4. Verify the data lands correctly in your destination catalog
  5. Expand to more objects and sources once you are comfortable with the pattern

The investment is minimal. You do not need to write code, manage infrastructure, or build custom monitoring. The pipeline is governed from day one.


This article connects to concepts covered in the Data Sources, Incremental Processing, Unity Catalog, and Workflows chapters of BricksNotes. For a hands-on guide to building downstream transformations after ingestion, see our Lakeflow Declarative Pipelines with Medallion Architecture article.