databricks/pipelines
intermediate
Connecting…

Workflows: a DAG of tasks, not one notebook

A job is a DAG of tasks with real parallelism, per-task retries, and a repair run that skips what already passed.

Lesson 13 of 14 · Databricks path

Explain it at my level

  1. Task DAG
  2. depends_on
  3. Compute
  4. Retries
  5. Repair
Watch the canvas:ingest tasktransform taskserving taskjob cluster / serverlessretry + run_iffailed / all-purposeskippedrepair runLive simulation
Job DAG — orders_dailybronze_ingestAuto Loader → bronzeNBsilver_ordersdedupe + type · MERGENBsilver_customersdbt build · dim_customerDBTgold_metricsgold.daily_metricsSQLCompute per taskJob clustercreated per run · dies at the endJobs Compute rate · job_cluster_keyAll-purpose clusterinteractive · stays alive afterpremium rate · the quiet mistakeServerless jobsno provisioning waitper task · nothing to sizeEach task has its own type, parameters, and compute — that is the whole point.
Tasks, dependencies, per-task compute, per-task retries — and a repair that re-runs one branch.