How managed connectors eliminate custom ingestion code and bring governed, incremental data into your lakehouse.
Every data engineering team eventually faces the same problem. You have data in Salesforce. You have data in PostgreSQL. You have data in ServiceNow, SharePoint, Google Analytics, and a dozen other systems.
Getting that data into your lakehouse reliably, incrementally, and with proper governance is where most teams spend a disproportionate amount of time.
Databricks Lakeflow Connect was built to solve exactly this. It provides fully-managed connectors that handle authentication, incremental reads, schema evolution, and retries. You configure the source and destination. Lakeflow Connect handles everything in between.
Lakeflow Connect is a managed ingestion layer built directly into the Databricks platform. Every pipeline it creates is governed by Unity Catalog and powered by serverless compute through Lakeflow Declarative Pipelines.
There are two categories of connectors:
SaaS connectors ingest data from enterprise applications like Salesforce, ServiceNow, Google Analytics, HubSpot, Workday, and others. These connectors handle API pagination, rate limiting, incremental reads, and schema changes automatically.
Database connectors ingest data from relational databases like MySQL, PostgreSQL, and SQL Server using change data capture (CDC). These connectors track changes at the source and apply them incrementally to Delta tables, keeping your lakehouse in sync without full reloads.
The key difference from writing custom ingestion code: you do not manage compute, handle retries, build checkpointing logic, or worry about schema drift. Lakeflow Connect handles all of it.
As of June 2026, Lakeflow Connect is generally available and supports a growing list of enterprise sources across both categories.
| Connector | Status | What It Ingests |
|---|---|---|
| Salesforce | GA | CRM records, customer data |
| ServiceNow | GA | IT service management records, tickets |
| Google Analytics | GA | Website traffic, user behavior data |
| Workday Reports | GA | Custom Workday reports |
| Microsoft Dynamics 365 | GA | CRM and ERP data |
| NetSuite | GA | ERP and financial data |
| HubSpot | GA | CRM, marketing, sales activity |
| Confluence | GA | Pages, spaces, comments |
| Jira | GA | Project management, issue tracking |
| SharePoint | GA | Files and list data |
| Google Ads | GA | Campaign performance metrics |
| Meta Ads | GA | Facebook and Instagram Ads data |
| TikTok Ads | GA | Advertising performance data |
| Workday HCM | GA | HR, payroll, workforce data |
| Zendesk Support | GA | Support tickets, users, organizations |
| Connector | Status |
|---|---|
| SQL Server | GA |
| MySQL | GA |
| PostgreSQL | GA |
Query-based connectors for over 100 sources are also generally available as of May 29 2026.
A SaaS connector consists of three components:
The setup is straightforward. You create a connection, point it at the source objects you want, specify a destination catalog and schema, and start the pipeline. No Spark code. No custom API clients. No checkpoint management.
-- Example: Creating a Salesforce connection
CREATE CONNECTION salesforce_crm
TYPE salesforce
OPTIONS (
login_url = 'https://login.salesforce.com',
client_id = SECRET('scope', 'sf_client_id'),
client_secret = SECRET('scope', 'sf_client_secret')
);Once the connection exists, you create an ingestion pipeline that references it:
-- Example: Creating an ingestion pipeline from Salesforce
CREATE OR REFRESH STREAMING TABLE salesforce_accounts
AS SELECT * FROM READ_SALESFORCE(
connection => 'salesforce_crm',
object => 'Account'
);Lakeflow Connect handles the OAuth flow, API pagination, rate limit backoff, and incremental state tracking. If the source schema changes (new fields added, fields removed), the connector adapts automatically based on your schema evolution settings.
Database connectors are architecturally different because they use change data capture (CDC) to track modifications at the source.
The components are:
This architecture means the gateway captures every change as it happens, while the ingestion pipeline can batch-apply those changes on a schedule. You never lose data because the staging layer acts as a buffer.
-- Example: Creating a PostgreSQL CDC connection
CREATE CONNECTION postgres_prod
TYPE postgresql
OPTIONS (
host = 'prod-db.example.com',
port = '5432',
user = SECRET('scope', 'pg_user'),
password = SECRET('scope', 'pg_password')
);This is what separates Lakeflow Connect from third-party ingestion tools.
Every connection is a Unity Catalog securable object. This means:
When a Lakeflow Connect pipeline writes data, the destination tables are automatically tracked in Unity Catalog's lineage graph. You can see exactly which source system produced which table, when it was last updated, and who has access to it.
For organizations that need to demonstrate data provenance for compliance (GDPR, SOC 2, HIPAA), this built-in governance eliminates the need for separate lineage tracking tools. Delta Sharing has also evolved into OpenSharing, an open protocol for sharing these governed assets across platforms.
Lakeflow Connect is the right choice when:
Custom ingestion code is still necessary when:
For most enterprise data engineering teams, Lakeflow Connect handles 60 to 80 percent of ingestion workloads. The remaining sources that need custom code can use Auto Loader for files or custom Spark structured streaming jobs.
Lakeflow Connect naturally maps to the Bronze layer of a medallion architecture.
The destination tables from Lakeflow Connect are your Bronze tables. They contain raw, incrementally updated data from source systems. From there, your Lakeflow Declarative Pipelines (formerly Delta Live Tables) handle Silver and Gold transformations.
The flow looks like this:
Because both Lakeflow Connect and Lakeflow Declarative Pipelines are governed by Unity Catalog and run on serverless compute, the entire pipeline from source to BI is managed, governed, and cost-efficient.
Lakeflow Connect pipelines emit metrics and events that you can monitor through the Databricks UI or programmatically. Key metrics to watch:
SaaS connectors run on serverless compute, so you pay only for processing time. Database connectors require the ingestion gateway to run on classic compute (continuously), but the actual data loading runs serverless.
For high-volume CDC scenarios, the gateway compute cost is the primary consideration. Size it based on your source database's change rate, not the total data volume.
Decide upfront how to handle schema changes:
For most production setups, auto-merge with monitoring alerts is the recommended approach. You want schema changes to flow through, but you also want to know when they happen.
If you are evaluating Lakeflow Connect for your organization:
The investment is minimal. You do not need to write code, manage infrastructure, or build custom monitoring. The pipeline is governed from day one.
This article connects to concepts covered in the Data Sources, Incremental Processing, Unity Catalog, and Workflows chapters of BricksNotes. For a hands-on guide to building downstream transformations after ingestion, see our Lakeflow Declarative Pipelines with Medallion Architecture article.