A calm walk through the data engineering cheat sheet, with the lesson behind every line
A cheat sheet is a map, not the territory.
It is one of the most shared things in data engineering. A single page that lists every Databricks feature next to the job it does. Save it, pin it, screenshot it. It is genuinely useful when you already know the ideas and just need a quick reminder of the name.
But a cheat sheet is weak at one thing. It cannot teach you why a feature exists, or when reaching for it is the wrong move. The grid says "Time Travel: VERSION AS OF" and stops there. It does not tell you that time travel quietly stops working once VACUUM removes the old files.
So this article does the opposite of a cheat sheet. We walk through the same map, section by section, and for each line we link the lesson where you actually practice the idea on real data. Use the cheat sheet to remember. Use this to understand.
Everything here runs on Databricks Free Edition unless we say otherwise.
Before features, the platform itself. Six pieces show up on every cheat sheet.
The Workspace is where you live: notebooks, files, and collaboration. The Cluster is the compute that runs your code. Delta Lake is the reliable storage layer that makes a data lake behave like a database. The Notebook is where you write Python, SQL, and Scala together. Jobs & Pipelines schedule and automate your work. Unity Catalog governs who can see and touch what.
If these words feel abstract, start at Workspace Essentials. It is the calmest place to get oriented before any of the feature names mean anything.
The middle of most cheat sheets is a column of small code snippets. Read a DataFrame, write it, filter, select, group by, order by. These are the verbs of daily work.
They look obvious on paper. The understanding comes from running them and watching what Spark actually does. Lazy evaluation, for example, means none of your transformations run until an action asks for a result. You cannot feel that from a snippet.
Practice the reading and writing in DataFrames, the SQL side in Spark SQL, and the transformation verbs in Transformations. When you just need the syntax fast, the PySpark Cheat Sheet and the SQL Cheat Sheet are the quick-reference pages on this site.
Cheat sheets often show a dbfs:/mnt/... path and mention mounting cloud storage. That style is older. On a modern Unity Catalog workspace, your data lives in Volumes, with paths like /Volumes/workspace/default/book_data/. The examples across our lessons use Volumes for this reason.
The choice of file format also matters more than the grid suggests. Parquet, Delta, and the trade-offs between them are covered in File Formats. How data arrives in the first place is in Data Sources. Unity Catalog now also supports Apache Iceberg v3 as a production-ready catalog.
This is the part of the cheat sheet worth slowing down on. ACID transactions, schema enforcement, time travel, versioning, and MERGE are the reasons Delta Lake exists.
ACID means a write either fully happens or does not happen at all. No half-written tables after a job fails. Schema enforcement stops a malformed batch from quietly corrupting a clean table. Time travel lets you query the table as it looked at an earlier version, which turns "what did this look like before the bad load?" into a one-line query.
MERGE is the quiet hero. It handles insert, update, and delete in a single statement, which is how upserts and change data work in practice.
All of this is hands-on in Delta Lake. For the production habits around it, read Delta Lake Best Practices: VACUUM, OPTIMIZE, and Time Travel.
The longer cheat sheet maps fifty use cases to fifty features. Fifty rows is a lot to hold in your head. It is easier when you group them into the themes a data engineer actually thinks in.
Auto Loader, COPY INTO, Lakeflow Connect, and Change Data Feed all answer one question: how does new data arrive reliably?
Auto Loader watches a location and picks up new files as they land. COPY INTO is the simpler, idempotent batch load. Lakeflow Connect (now GA) pulls from SaaS apps and databases directly. They overlap, which is exactly why people pick the wrong one.
We compared them side by side in Lakeflow Connect vs Auto Loader vs COPY INTO, went deeper on ingestion at scale in Let the Files Come to You, and the lesson is Data Sources. For incremental and CDC patterns, see Incremental Processing.
MERGE, UPDATE, DELETE, and VERSION AS OF are the operations that older data lakes could not do. They are the everyday Delta verbs. For managed Postgres workloads directly in the lakehouse, Lakebase is now generally available.
The most common real use is slowly changing dimensions, where you need to keep history as records change. That pattern is built carefully in SCD Patterns, with the underlying mechanics in Delta Lake.
OPTIMIZE, VACUUM, ZORDER, Liquid Clustering, and Photon are the performance column. This is where cheat sheets cause the most confusion, because the names sound interchangeable and are not.
OPTIMIZE compacts many small files into fewer large ones. VACUUM removes old, unreferenced files to reclaim storage, and it is also what eventually limits how far back time travel can go. Liquid Clustering is the modern replacement for rigid partitioning, and on new tables it is usually the better default. Photon is the faster query engine underneath.
The clearest story we have written on this is From 47 Minutes to 5 Minutes: Fixing the Small File Problem. Photon is explained plainly in Databricks Photon Engine. The lessons are Partitioning and Performance and File Formats. And if you want generated, production-shaped PySpark with these habits already in place, try the PySpark Pipeline Generator.
Bronze, Silver, and Gold are not features. They are a way of arranging tables so raw data, cleaned data, and business-ready data each have a clear home. That structure is the Medallion Architecture, covered in Medallion Architecture. This is often implemented using Lakeflow Declarative Pipelines (formerly Delta Live Tables).
Structured Streaming, watermarks, and checkpointing handle data that never stops arriving. The hard parts are state and late data, not the syntax.
Learn the model in Streaming, then read Structured Streaming Best Practices: Checkpointing and Watermarks for the habits that keep a streaming job healthy.
Expectations validate data as it flows. Schema evolution lets a table change shape over time without breaking. Both protect the people downstream who trust your tables.
Practice quality checks in Data Quality and safe schema change in Schema Evolution.
Unity Catalog, lineage, Catalog Explorer, and external tables all sit under one idea: knowing what data exists, who owns it, and who is allowed to use it. OpenSharing, the evolution of Delta Sharing, now allows for vendor-neutral sharing of these assets.
Start with Unity Catalog, then read Unity Catalog Best Practices for how to lay out namespaces, access, and lineage in a way that scales.
Jobs & Pipelines, Workflows, Repos, and parameters turn a working notebook into something that runs on its own, on a schedule, and survives changes safely.
The lesson is Workflows. For the part most people find intimidating, version control and deployment, read Declarative Automation Bundles for Data Engineers Who Have Never Done CI/CD.
ML Runtime and Feature Engineering appear near the bottom of the cheat sheet. They are where data engineering hands off to data science, and understanding the boundary makes you better at both. See Machine Learning.
The "best practices" box on a cheat sheet is a list of commands. Here is the same list with the reason attached, because the reason is what you actually remember.
Use Delta Lake by default, because reliability should not be optional. Cluster large tables with Liquid Clustering so reads only touch the data they need. Run OPTIMIZE and VACUUM on a schedule, so small files and old files never pile up silently. Orchestrate with Jobs & Pipelines instead of running notebooks by hand, so work is repeatable. Monitor with the Spark UI and query history, so you find problems before users do. Govern with Unity Catalog, so access is a decision, not an accident.
None of these are tricks. They are just the calm habits of people who have run pipelines long enough to know what breaks.
Keep it for recall, not for learning. When you see a feature name you cannot explain in one sentence, that is your signal to open the lesson and run it once. After that, the cheat sheet does its real job: a quick nudge for something you already understand.