A streaming pipeline kept creating tiny Parquet files. The cluster was not the problem. The data layout was.
Recently, a query on a 100 GB Delta Lake table in Databricks was taking close to 47 minutes to finish. The first instinct, like most teams, was to blame the cluster. Maybe more workers. Maybe a bigger driver. Maybe a Spark config to tune.
None of that was the real issue.
The table was being written continuously by a Structured Streaming job. Over time, it had quietly produced hundreds of tiny Parquet files. The data volume was not large, but Spark had to open and track a huge number of files before it could do any real work. The bottleneck was not compute. It was the layout of the data on storage.
This is the small file problem, and it shows up almost everywhere streaming meets Delta Lake. If you are new to how Delta organizes data, the Delta Lake lesson walks through how files, transaction logs, and snapshots fit together.
Every Parquet file Spark reads has a cost beyond its bytes. The driver has to list files. The executors have to open them, read footers, and plan tasks. Data skipping depends on file-level statistics, so a thousand tiny files give Spark a thousand tiny decisions to make instead of a few clean ones.
When the file count grows faster than the data, three things slow down at the same time:
You can see this in the Spark UI. If the job spends most of its time in planning and scan stages with tiny task durations, file layout is the suspect. The debugging and monitoring lesson shows what to look for.
The first fix was the most direct one. Run OPTIMIZE on the table to combine the small files into fewer, larger ones.
OPTIMIZE my_catalog.my_schema.events;OPTIMIZE rewrites small files into target-sized files (around 1 GB by default) without changing the data. Queries that used to scan thousands of files now scan a handful. This alone cut a large chunk of the runtime.
OPTIMIZE is not free. It is a maintenance operation that costs compute and writes new files. The right pattern is to run it on a schedule, not after every micro-batch. The Delta Lake best practices article on VACUUM, OPTIMIZE, and time travel covers when to run it and how it interacts with retention.
Compaction fixed the file count. The second fix was about file content.
Most queries on this table filtered by account_id. So data was reorganized using Z-ORDER on that column:
OPTIMIZE my_catalog.my_schema.events
ZORDER BY (account_id);Z-Ordering co-locates rows that share similar values in the chosen column. After Z-ORDER, queries that filter by account_id can skip entire files instead of scanning the whole table. This is data skipping in its most useful form.
A note for 2026: on new tables, Liquid Clustering is now the recommended default over partitioning and Z-ORDER. It adapts as data and queries evolve, instead of locking in a layout. The small file problem and Liquid Clustering article explains when to choose which.
If you want to see partitioning, file sizing, and Z-ORDER in one place with hands-on examples, the Partitioning and Performance lesson is the focused chapter.
Compaction and Z-ORDER make a slow table fast today. They do not stop the table from getting slow again tomorrow. The third fix was upstream.
Structured Streaming was committing micro-batches very frequently. Each commit produced new files. Three small changes restored a healthier write pattern:
processingTime or a Trigger.AvailableNow schedule writes fewer, larger files per commit.OPTIMIZE (and VACUUM with a safe retention) as a separate job, not inside the streaming query.The Structured Streaming lesson covers trigger choices and checkpointing, and the incremental processing lesson goes deeper into write patterns that stay healthy under load. For a focused walkthrough on checkpointing and watermarks specifically, the streaming best practices article is the companion piece.
If you are designing a new streaming pipeline and want sensible trigger and maintenance defaults out of the box, Databricks Lakeflow is now GA, offering unified ingestion and Lakeflow Declarative Pipelines (formerly Delta Live Tables).
Same cluster. Same data. Same query.
No bigger machines. No clever Spark config. Just a better layout, a better write pattern, and a small maintenance job to keep it that way.
The lesson is one most data engineers learn the hard way: performance is a layout problem before it is a compute problem. Compute makes a bad layout cheaper to suffer. It does not make it fast.