Databricks Auto Loader Best Practices: Schema Inference, Rescue Columns, and Cloud Notifications

How to build reliable file-based ingestion pipelines that handle schema changes, recover unexpected data, and scale to millions of files.

Files keep arriving. New columns appear without warning. Data types change between batches. A field that was an integer last week is suddenly a string.

This is the reality of file-based ingestion. And it is exactly what Auto Loader was designed to handle.

Auto Loader is Databricks' built-in solution for incrementally processing new files as they land in cloud storage. It tracks which files have been processed, detects schema changes, rescues unexpected data, and scales from hundreds to millions of files.

But using Auto Loader effectively requires understanding how its features interact. Schema inference, rescue columns, schema hints, evolution modes, and file detection strategies all work together. Getting the configuration right means the difference between a pipeline that silently drops data and one that handles every edge case gracefully.

How Schema Inference Works

When Auto Loader first encounters data, it samples files to determine the schema. By default, it samples the first 50 GB or 1,000 files, whichever limit is reached first.

The inferred schema is stored in a _schemas directory at the location you specify with cloudFiles.schemaLocation. This means the schema is persistent across stream restarts. Auto Loader does not re-infer the schema every time your pipeline runs.

# Basic Auto Loader with schema inference
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.schemaLocation", "/mnt/checkpoints/orders_schema")
  .load("/mnt/landing/orders/")
  .writeStream
  .option("checkpointLocation", "/mnt/checkpoints/orders")
  .option("mergeSchema", "true")
  .toTable("bronze.orders")
)

Default Type Behavior

How Auto Loader infers types depends on the file format:

FormatDefault Behavior
JSONAll columns inferred as strings
CSVAll columns inferred as strings
XMLAll columns inferred as strings
ParquetTypes from Parquet schema
AvroTypes from Avro schema

For JSON and CSV, inferring everything as strings is intentional. It prevents type mismatch failures when source data is inconsistent. You handle type casting downstream in your Silver layer transformations.

If you want Auto Loader to infer actual types from JSON or CSV data (integers, doubles, timestamps), set cloudFiles.inferColumnTypes to true. But be careful. This can cause pipeline failures if the same field contains mixed types across files.

# Enable type inference (use with caution for JSON/CSV)
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.inferColumnTypes", "true")
  .option("cloudFiles.schemaLocation", "/mnt/checkpoints/events_schema")
  .load("/mnt/landing/events/")
)

Best practice: For most production pipelines, keep the default string inference for JSON and CSV. Cast types explicitly in your Silver layer where you control the logic and can handle edge cases.

The Rescued Data Column

This is one of Auto Loader's most important features, and one that many teams overlook.

When schema inference is enabled, Auto Loader automatically adds a _rescued_data column to every record. This column captures any data that does not match the inferred schema.

Data ends up in the rescue column for three reasons:

  1. Missing from schema — A new column appears that was not in the original inferred schema
  2. Type mismatch — A value does not match the expected data type
  3. Case mismatch — A column name differs in casing (e.g., userId vs userid)

The rescued data is stored as a JSON blob containing the mismatched columns and the source file path.

# The _rescued_data column is added automatically
# You can also rename it:
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("rescuedDataColumn", "_rescue")
  .option("cloudFiles.schemaLocation", "/mnt/checkpoints/schema")
  .load("/mnt/landing/data/")
)

Why This Matters in Production

Without the rescue column, unexpected data is silently dropped. You would never know that a source system added a new field, changed a data type, or started sending column names with different casing.

With the rescue column, nothing is lost. You can:

Best practice: Always keep the rescue column enabled. In your Silver layer, add a data quality check that flags records where _rescued_data IS NOT NULL. This gives you visibility into schema drift without breaking your pipeline.

# Silver layer: flag records with rescued data
from pyspark.sql import functions as F

silver_df = (
  spark.readStream
    .table("bronze.orders")
    .withColumn(
      "has_schema_drift",
      F.col("_rescued_data").isNotNull()
    )
)

Schema Hints

Schema hints let you override specific columns in the inferred schema without providing the entire schema yourself. This is useful when you know certain fields should be specific types.

# Force specific column types during inference
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.schemaHints",
    "order_id long, amount decimal(10,2), tags map<string,string>")
  .option("cloudFiles.schemaLocation", "/mnt/checkpoints/schema")
  .load("/mnt/landing/orders/")
)

Schema hints work alongside inference. Auto Loader infers the schema from sample data, then applies your hints on top. If a hint conflicts with the inferred type, your hint wins.

You can also use schema hints to add columns that might not exist in the initial sample but will appear later:

# Pre-declare a column that will appear in future files
.option("cloudFiles.schemaHints",
  "loyalty_tier string, discount_pct double")

Best practice: Use schema hints for columns where you know the correct type and the source is inconsistent. Common examples include ID fields (force to string or long), monetary amounts (force to decimal), and timestamp fields.

Schema Evolution Modes

When Auto Loader encounters new columns that were not in the original schema, its behavior depends on the schema evolution mode.

ModeWhat Happens
addNewColumns (default without schema)Stream stops. New columns are added to schema. Restart to continue.
rescueStream continues. New columns go to _rescued_data. Schema never changes.
failOnNewColumnsStream stops. Does not update schema. You must manually intervene.
noneStream continues. New columns are silently ignored (unless rescue column is enabled).
addNewColumnsWithTypeWideningStream stops. New columns added, and type changes are widened automatically. (Public Preview in Runtime 16.4+)
# Set schema evolution mode
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.schemaEvolutionMode", "addNewColumns")
  .option("cloudFiles.schemaLocation", "/mnt/checkpoints/schema")
  .load("/mnt/landing/events/")
)

Choosing the Right Mode

Use addNewColumns when you want your schema to grow automatically but want to be notified (via stream restart) when changes happen. Pair this with Databricks Workflows that auto-restart on failure.

Use rescue when you need maximum stability. The stream never stops due to schema changes. New data is captured in the rescue column for later analysis. This is ideal for Bronze layers where uptime matters more than immediate schema alignment.

Use failOnNewColumns when schema changes require human review. This is common in regulated environments where schema modifications need approval.

Use none when you have a fixed, known schema and want to ignore any unexpected columns entirely.

Best practice: For most production Bronze pipelines, use addNewColumns with auto-restart in Workflows. For ultra-stable pipelines where downtime is unacceptable, use rescue mode and process rescued data in a separate monitoring pipeline.

File Detection Modes

Auto Loader discovers new files using one of two modes: directory listing or file notification.

Directory Listing Mode (Default)

Auto Loader periodically lists the contents of the input directory and identifies new files by comparing against previously processed files. It uses an internal RocksDB-based state store to track processed files efficiently.

This mode works out of the box with no additional cloud infrastructure. It is suitable for directories with up to a few million files.

# Directory listing is the default
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.useNotifications", "false")  # default
  .load("/mnt/landing/events/")
)

Limitations: As the number of files grows into the tens of millions, directory listing becomes slower because it must list all files to find new ones. The cost and latency of listing scale with the total number of files in the directory.

File Notification Mode

In file notification mode, Auto Loader automatically sets up cloud notification services to receive events when new files arrive. On AWS, this uses SNS and SQS. On Azure, it uses Event Grid and Queue Storage. On GCP, it uses Pub/Sub.

# Enable file notification mode
(spark.readStream
  .format("cloudFiles")
  .option("cloudFiles.format", "json")
  .option("cloudFiles.useNotifications", "true")
  .load("s3://my-bucket/landing/events/")
)

Advantages:

Requirements:

Which Mode to Choose

ScenarioRecommended Mode
Less than 1 million files totalDirectory listing
More than 1 million filesFile notification
Low-latency requirementsFile notification
Cannot grant notification setup permissionsDirectory listing
Using Lakeflow Declarative PipelinesDirectory listing works well; state is managed automatically

Best practice: Start with directory listing. Switch to file notification when directory listing latency becomes noticeable or when your directory exceeds a few million files. The switch is seamless since Auto Loader maintains exactly-once guarantees across mode changes.

Production Configuration Template

Here is a complete production configuration that combines the best practices discussed above:

from pyspark.sql import functions as F

# Production Auto Loader configuration
raw_df = (
  spark.readStream
    .format("cloudFiles")
    .option("cloudFiles.format", "json")
    .option("cloudFiles.schemaLocation",
      "/mnt/checkpoints/orders/schema")
    .option("cloudFiles.schemaEvolutionMode", "addNewColumns")
    .option("cloudFiles.schemaHints",
      "order_id string, amount decimal(10,2), created_at timestamp")
    .option("rescuedDataColumn", "_rescued_data")
    .option("cloudFiles.inferColumnTypes", "false")
    .load("/mnt/landing/orders/")
)

# Add ingestion metadata
enriched_df = (
  raw_df
    .withColumn("_ingested_at", F.current_timestamp())
    .withColumn("_source_file", F.input_file_name())
)

# Write to Bronze table
(enriched_df
  .writeStream
    .option("checkpointLocation",
      "/mnt/checkpoints/orders/checkpoint")
    .option("mergeSchema", "true")
    .trigger(availableNow=True)
    .toTable("bronze.orders")
)

This configuration:

Auto Loader with Lakeflow Declarative Pipelines

If you are using Lakeflow Declarative Pipelines (formerly Delta Live Tables), Auto Loader integration is even simpler. DLT manages schema location and checkpoints automatically.

import dlt

@dlt.table(
  comment="Raw order events from landing zone"
)
@dlt.expect_or_drop("valid_order_id", "order_id IS NOT NULL")
def bronze_orders():
  return (
    spark.readStream
      .format("cloudFiles")
      .option("cloudFiles.format", "json")
      .option("cloudFiles.inferColumnTypes", "false")
      .option("cloudFiles.schemaHints",
        "order_id string, amount decimal(10,2)")
      .load("/mnt/landing/orders/")
      .withColumn("_ingested_at", F.current_timestamp())
  )

With Lakeflow Declarative Pipelines, you do not need to specify schemaLocation or checkpointLocation. The platform handles both. You also get built-in expectations for data quality and automatic lineage tracking through Unity Catalog.

Common Pitfalls

1. Not monitoring the rescue column. If you never check _rescued_data, you are blind to schema drift. Build a monitoring query that runs regularly.

2. Using inferColumnTypes with inconsistent JSON sources. A field that is 42 in one file and "forty-two" in another will cause type conflicts. Default string inference is safer.

3. Forgetting mergeSchema on the write side. Auto Loader can evolve the read schema, but the destination table also needs to accept new columns. Set mergeSchema to true on the write stream.

4. Not setting schema hints for known ID fields. Without hints, Auto Loader might infer an ID as integer from the sample, then fail when it encounters values that exceed integer range. Hint IDs as string or long.

5. Using file notification mode without cleanup. Auto Loader creates cloud notification resources. If you delete a pipeline without cleaning up, orphaned SNS topics or SQS queues may accumulate. Use cloudFiles.cleanSource thoughtfully.


This article connects to concepts covered in the Data Sources, Incremental Processing, Schema Evolution, and Data Quality chapters of BricksNotes. For ingestion from SaaS and database sources instead of files, see our Lakeflow Connect guide.