A thinking-first walk through the questions that come up most in 2026 data engineering interviews
Most interview lists hand you answers to memorize. That is not what gets you the job.
What gets you the job is showing how you think. An interviewer rarely wants the textbook definition. They want to see that you understand why a system behaves the way it does, and what you would actually do when it breaks at two in the morning.
So this is not a flashcard pile. For each question you will find three things: what the interviewer is really testing, a short answer you can say out loud, and a stronger answer that shows depth. Where a topic connects to something you can practice, there is a link to the matching chapter so you can go from reading to doing.
If you want a structured way to rehearse out loud, our Interview Readiness trainer is built for exactly this rhythm of short answer, then stronger answer.
The best interview answers sound like someone explaining a system they have lived with, not someone reciting a page.
What they are testing: whether you understand that the table is the log, not the files.
Short answer: Delta Lake is Parquet files plus a transaction log. The _delta_log records every commit as an ordered set of add and remove actions. Readers build the current table state from the log, which gives you ACID guarantees on plain object storage.
Stronger answer: A classic data lake is just files in a folder. Two writers can corrupt each other, a failed job can leave half-written files, and readers can see partial data. Delta fixes this by making the log the source of truth. Each transaction writes a JSON entry listing which files were added and removed. Because commits are ordered and atomic, you get isolation, consistent reads, and the ability to roll back. The Parquet files alone are not the table. The log is. Note that as of 2026, Apache Iceberg v3 is also GA on Databricks, offering similar reliability through Unity Catalog.
Walk through the reliability story in Chapter: Delta Lake, where you can open the _delta_log and see the commits for yourself in Free Edition.
What they are testing: whether you treat data as something with history.
Short answer: Every commit creates a new version. You can query an older version by number or timestamp using VERSION AS OF or TIMESTAMP AS OF. Use it for audits, debugging, and recovering from a bad write.
Stronger answer: Versioning falls out naturally from the transaction log. Each commit is a version, so you can read the table as it was at any point. The real value shows up in three moments. First, a bad upsert ships and you need to compare before and after. Second, an audit asks what a report looked like last quarter. Third, a downstream model needs a stable, reproducible snapshot. Time travel is bounded by your retention and by VACUUM, so do not treat it as infinite backup.
This is covered hands-on in Chapter: Delta Lake.
What they are testing: real upsert experience, not just the definition.
Short answer: Use MERGE INTO. Close the current row by setting an end date and an active flag to false when a tracked column changes, then insert the new version as the current row.
Stronger answer: SCD Type 2 keeps history by never overwriting. You compare incoming records to the current active rows. When a tracked attribute changes, you expire the old row with an end_date and is_current = false, then insert a fresh row with is_current = true and an open end date. MERGE handles the match and not-matched branches in one statement. Watch for two traps: a wide source makes the merge expensive, so narrow it first, and you need a stable business key, not a surrogate, to detect change correctly.
The full pattern with runnable code is in Chapter: SCD Patterns.
What they are testing: whether you debug by evidence or by guessing.
Short answer: Find the cause first. Usually it is small files, missing data skipping, or a poor layout. Run OPTIMIZE, adopt Liquid Clustering on the columns you filter on, and confirm data skipping is working.
Stronger answer: Do not start by adding compute. Start by measuring. Open the query profile and look for many small file reads, full scans where a filter should prune, or skew. The most common cause on a large append-heavy table is the small file problem, where streaming or frequent writes create thousands of tiny files. Compact them with OPTIMIZE. For layout, prefer Delta Lake Liquid Clustering over manual partitioning on high-cardinality columns, because it clusters by your query keys without the partition explosion. Then verify that data skipping statistics actually prune files for your common predicates.
We wrote a full case study on this, From 47 Minutes to 5 Minutes: Fixing the Small File Problem, and the tuning levers live in Chapter: Partitioning and Performance. If you want to see optimized write settings generated for your own use case, try the PySpark Pipeline Generator.
What they are testing: file-size awareness and safe cleanup.
Short answer: OPTIMIZE compacts many small files into fewer larger ones to speed up reads. VACUUM deletes old data files that are no longer referenced by the log, to reclaim storage.
Stronger answer: They solve different problems. OPTIMIZE is about read performance. Frequent appends leave small files, and OPTIMIZE rewrites them to a target size, and with Liquid Clustering it also clusters by your keys. VACUUM is about housekeeping. When you update or delete, old files stay around so time travel keeps working. VACUUM removes files older than the retention threshold, which defaults to seven days. The risk is running VACUUM with too short a retention, because that breaks time travel and can pull files out from under a running reader.
Both commands are explained in Chapter: File Formats and applied in Chapter: Partitioning and Performance.
What they are testing: whether you can assign responsibility to each layer.
Short answer: Bronze holds raw, faithful source data. Silver applies structure, types, and quality. Gold shapes data for analytics and business use. The advantage is clear ownership and easy reprocessing.
Stronger answer: Each layer has one job. Bronze is an honest copy of the source, so you can always reprocess from a clean starting point. Silver is where deduplication, type casting, joins to reference data, and quality checks belong. Gold is shaped for consumption, often aggregated and denormalized for BI. The big advantages are debuggability, because you can inspect data at each stage, and reprocessability, because Bronze never loses fidelity. The common mistake is doing business logic in Bronze, which destroys the ability to rebuild. Databricks Lakeflow (now GA) unifies this ingestion and transformation under Unity Catalog.
Build all three layers in Chapter: Medallion Architecture.
What they are testing: whether you understand exactly-once is a design choice, not a default.
Short answer: Prevent duplicates with idempotent writes. Use a checkpoint per query, foreachBatch with MERGE on a business key, or dropDuplicatesWithinWatermark for near-duplicates in a time window.
Stronger answer: Duplicates usually come from one of two places. Either the source delivers at-least-once, so retries replay events, or a job restarted without a stable checkpoint. To prevent them, give each stream its own checkpoint so offsets are tracked. For real deduplication, do not just append. Use foreachBatch and MERGE on a unique key so a replayed record updates instead of inserting. For event streams with a natural time window, dropDuplicatesWithinWatermark removes near-duplicates without holding unbounded state. For automated pipelines, Lakeflow Declarative Pipelines (formerly Delta Live Tables) handles much of this state management for you.
The streaming mechanics are in Chapter: Streaming and the upsert side is in Chapter: Incremental Processing.
What they are testing: a clear mental model of how Spark runs work over time.
Short answer: Batch runs once over a bounded dataset. Micro-batch, the default for Structured Streaming, runs many small batches on a trigger. Continuous processing aims for very low latency by processing records as they arrive.
Stronger answer: Batch is a single job over a fixed input, ideal for nightly loads. Micro-batch treats a stream as a series of small batch jobs, which gives you most of the benefits of streaming with the reliability and tooling of batch, and it is what almost everyone uses. The availableNow trigger lets you run a streaming pipeline as a scheduled batch, which is a great middle ground. Continuous processing trades some guarantees for lower latency. For millisecond latency on governed tables, the new Reyden engine in Lakehouse//RT is the modern choice.
See the trigger options in action in Chapter: Streaming.
What they are testing: isolation, governance, and cost awareness at scale.
Short answer: Isolate tenants with the Unity Catalog three-level namespace, usually a catalog per tenant or per environment, enforce access with grants, and separate compute so one tenant cannot starve another.
Stronger answer: Multi-tenancy is mostly about boundaries. Use Unity Catalog to give each tenant its own catalog, then schemas and tables underneath, with grants that inherit downward. This keeps data and permissions cleanly separated. For compute, isolate workloads so a heavy tenant does not affect others, and tag resources for cost attribution. Standardize ingestion and the medallion layers so every tenant follows the same shape, which keeps operations sane as you add tenants. Lineage and audit from Unity Catalog give you the visibility to prove isolation. You can now also use Domains in Unity Catalog to further organize these boundaries.
Governance foundations are in Chapter: Unity Catalog and orchestration in Chapter: Workflows.
What they are testing: a calm, evidence-led debugging method.
Short answer: Check the streaming metrics for input rate versus processing rate and batch duration. Look for skew, expensive operations, small files, and too few shuffle partitions. Scale the bottleneck, not everything.
Stronger answer: Open the streaming query metrics and compare input rows per second with processed rows per second. If processing time per batch keeps climbing, you have a real bottleneck. Common causes are a stateful operation growing without bounds, skew where one partition does most of the work, an expensive join that should be a broadcast, or a sink writing many small files. Tune the trigger interval, set maxFilesPerTrigger or maxBytesPerTrigger to control intake, right-size shuffle partitions, and bound state with a watermark. Add capacity only after you know which stage is slow.
The method generalizes from Chapter: Debugging and Monitoring, paired with Chapter: Streaming.
What they are testing: whether governance is a habit or an afterthought.
Short answer: Use the three-level namespace, grant access at the right level so it inherits down, and rely on built-in lineage and audit. Govern tables, views, volumes, and functions in one place.
Stronger answer: Unity Catalog centralizes governance across workspaces. The three-level naming, catalog then schema then object, gives you a clean place to apply grants that flow downward, so you set policy once. Beyond table grants, you can govern volumes for files and functions for logic. Column and row level controls handle sensitive data. Lineage shows where data came from and what depends on it, which is what makes change safe. The mindset shift is governing by policy and inheritance, not by chasing individual permissions. Note that Unity Catalog now also governs the new FILE type for unstructured data like PDFs and images.
We told the before and after story in Before and After Unity Catalog, and the hands-on version is Chapter: Unity Catalog.
What they are testing: whether you treat data pipelines as software.
Short answer: Keep everything in version control, test transformation logic, and deploy with Databricks Asset Bundles so notebooks, jobs, and pipelines move between environments as one unit.
Stronger answer: Treat pipelines like any other software. Code lives in Git. Transformation logic has unit tests so a bad change fails before it reaches data. Databricks Asset Bundles describe your jobs, pipelines, and notebooks declaratively, so the same definition deploys to dev, staging, and production with environment-specific values. A pull request runs the tests, and a merge triggers the deployment. The goal is that a release is boring and repeatable, not a manual click-through.
We wrote a gentle on-ramp for engineers new to this, Declarative Automation Bundles for Data Engineers Who Have Never Done CI/CD, and testing is covered in Chapter: Unit Testing with orchestration in Chapter: Workflows.
Use Auto Loader for incremental file ingestion at scale, because it tracks which files it has already seen. Use COPY INTO for simpler, idempotent loads of a known set of files. Use Lakeflow Connect (now GA) for managed ingestion from 100+ common sources without building the plumbing yourself. We compared all three in Lakeflow Connect vs Auto Loader vs COPY INTO, and the foundations are in Chapter: Data Sources.
Partitioning by a high-cardinality column creates too many small partitions and hurts performance. Liquid Clustering organizes data by the columns you actually filter on, without the partition explosion, and you can change the clustering keys later. Default to Liquid Clustering for most large tables. See Chapter: Partitioning and Performance.
Decide per source whether you allow new columns. For trusted sources, enable schema evolution on the write so new fields flow through. For everything else, validate and fail loudly so a silent schema drift does not corrupt downstream logic. The trade-offs are in Chapter: Schema Evolution.
Open the Spark UI and find the longest stage. Look for one task far slower than the rest, which means skew, large shuffle volumes, or many small files. Fix the biggest cause, then measure again. This method is the heart of Chapter: Debugging and Monitoring.
These questions reward understanding over memorization. The fastest way to sound fluent is to practice the answers out loud and to have actually run the code.
You do not need to memorize a hundred answers. You need to understand a handful of systems well enough to reason about anything they ask.