Five-hour rerun to redo a four-minute failure
Symptoms
- The nightly run fails on the vendor export near the end and starts over from the top
- Total pipeline time doubles on every failure
- Nobody can say which part of the pipeline was slow
Interactive Databricks interview questions on Workflows: a DAG of tasks, not one notebook. Practice with the matching lesson. A job is a DAG of tasks with real parallelism, per-task retries, and a repair run that skips what already passed.
Question 1 of 4
A teammate's production pipeline is one notebook with 40 cells on a schedule. Sell me on breaking it up.
Answer it out loud, then reveal. Play steps through like the simulators.
Case 1 of 2 · symptoms
Symptoms
Indexed as FAQ. Open any item if you prefer a list to Play.
A teammate's production pipeline is one notebook with 40 cells on a schedule. Sell me on breaking it up.
Workflows: a DAG of tasks, not one notebook · tap to open the answer
Short: A failure in cell 37 re-runs all 40 cells, and you get no per-step retry, no parallelism, and no per-step timing.
Detailed: A Databricks Job is a DAG of tasks wired with depends_on, so independent steps fan out in parallel, each task carries its own retry policy and timeout, and the run timeline gives you duration per task. Inside one notebook Spark sees a single linear sequence, so recovery means rerunning the expensive ingest you already completed.
Common mistake: Calling the monolith 'simpler' because it is one file, while every failure costs a full rerun.
Follow-up: Where would you put the task boundaries in that notebook?
A six-task job failed on task 4; tasks 5 and 6 show Skipped and someone is about to hit Run now. What do you tell them?
Workflows: a DAG of tasks, not one notebook · tap to open the answer
Short: Fix the cause and use Repair run — it re-executes the failed task and its skipped downstream branch while tasks 1 to 3 stay successful.
Detailed: Downstream tasks skip rather than fail when a dependency fails, so the run keeps an accurate record of what completed. Repair run reuses that state and restarts from the failure point, which is why task granularity matters: coarse tasks make repair redo work that was already fine.
Common mistake: Clicking Run now and paying for the first three tasks again — or double-writing output that is not idempotent.
Follow-up: How would you make that task safe to retry more than once?
The nightly job runs on the team's all-purpose cluster because it is already warm. Make the cost argument.
Workflows: a DAG of tasks, not one notebook · tap to open the answer
Short: All-purpose DBUs are a more expensive SKU than jobs compute, and you also pay for the cluster's idle hours; a job cluster exists for the run and terminates.
Detailed: In Databricks pricing, Jobs Compute is a cheaper DBU rate than All-Purpose Compute for the same instance type, and a job cluster starts per run and shuts down at the end so uptime matches work. Shared interactive clusters also leak state — installed libraries, cached data, someone else's runaway cell — into production runs, which is how a green pipeline fails only on Mondays.
Senior: There are two separate costs: the DBU rate and uptime you did not need. I would move scheduled work to job clusters or serverless, add a cluster policy so nobody can schedule production onto all-purpose, and keep interactive clusters for humans. If per-run startup latency is the genuine objection, serverless job compute answers it without the idle bill — the warm shared cluster answers it by paying for the cluster all night.
Common mistake: Arguing only about startup convenience and never comparing the DBU rate or the idle hours on the bill.
Follow-up: When is serverless job compute the better answer than a job cluster?
Two runs of the same hourly job overlapped and the target table got duplicate rows. What is wrong beyond 'the job was slow'?
Workflows: a DAG of tasks, not one notebook · tap to open the answer
Short: Concurrent runs were allowed against a non-idempotent write, so two runs processed overlapping input and both appended.
Detailed: A job's max concurrent runs setting decides whether a scheduled run starts while the previous one is still going; at 1 the new run is skipped or queued instead of racing. The durable fix is making the write idempotent — MERGE on a business key, or a replaceWhere overwrite bounded to the run's window — so an overlap or a retry converges instead of duplicating.
Senior: Overlap is the symptom; the non-idempotent append is the bug. I would set max concurrent runs to 1 with a timeout and an alert so a slow run becomes visible rather than doubled, pass the window explicitly through job parameters and task values instead of each task calling now(), and make the sink a MERGE or a bounded replaceWhere so re-running a window is a no-op. Then define the job in Databricks Asset Bundles so concurrency, retries, and schedule are code, not something each workspace gets clicked differently.
Common mistake: Just lengthening the schedule interval, which hides the race until one run is slow again.
Follow-up: How would you pass the window boundaries between tasks so each task knows exactly what it processes?