Interview/Spark

Spark beginner interview questions

Partitions, actions, lazy evaluation, and what the driver actually does.

Lesson

Walk a Spark action from code to rows without skipping a layer.

Answer it out loud, then reveal. Play steps through like the simulators.

All questions on this page

Indexed as FAQ. Open any item if you prefer a list to Play.

beginner

Walk a Spark action from code to rows without skipping a layer.

What happens when Spark executes a job? · tap to open the answer

Short: Code → logical plan → job → stages → tasks → (optional shuffle) → result.

Detailed: Transformations are lazy. The action builds a plan, splits it at shuffles into stages, runs one task per partition, then returns a small result or writes files.

Common mistake: Saying 'Spark runs the DataFrame line by line'.

Follow-up: Where does a shuffle sit in that chain?

Lesson · Simulation

intermediate

Why does one action create multiple stages?

What happens when Spark executes a job? · tap to open the answer

Short: Each wide transformation (shuffle) is a stage boundary.

Detailed: Narrow steps (filter, map, broadcast join) stay in one stage. groupBy, join-without-broadcast, repartition cut a new stage because data must move.

Common mistake: Counting jobs instead of stages when debugging runtime.

Follow-up: How do you see stage boundaries in the physical plan?

Lesson · Simulation

beginner

What is Apache Spark, in one sentence a hiring manager wants?

What is Apache Spark? · tap to open the answer

Short: A distributed engine that processes large datasets in parallel across a cluster.

Detailed: Spark splits data into partitions, runs a task per partition on executors, and only ships results when an action runs. It is not a database and it does not store your lake.

Common mistake: Calling Spark a 'database' or saying it always holds data in memory.

Follow-up: What is lazy about a DataFrame transformation?

Lesson · Simulation

intermediate

Why can a 20-line notebook do nothing for minutes, then explode when you call show()?

What is Apache Spark? · tap to open the answer

Short: Transformations build a plan. Actions execute it.

Detailed: groupBy / filter / join are lazy. Spark waits for show, count, write, collect. That is when jobs, stages, and tasks appear in Spark UI.

Common mistake: Thinking Spark is slow at parse time because the notebook looks busy.

Follow-up: Name three actions that trigger a job.

Lesson · Simulation

beginner

Name the three Spark cluster roles and what each one must not do.

Spark cluster architecture · tap to open the answer

Short: Driver coordinates. Executors compute. Cluster manager places the JVMs.

Detailed: The driver builds the DAG and tracks tasks. Executors run tasks and hold cache. The cluster manager (YARN, K8s, Databricks) allocates machines. Big data stays on executors and storage — not in the driver heap.

Common mistake: Saying the driver 'processes the data'.

Follow-up: Where does collect() put the result?

Lesson · Simulation

intermediate

If the driver dies, why does the whole job die even if executors are healthy?

Spark cluster architecture · tap to open the answer

Short: The driver owns the SparkSession, DAG, and task scheduler. Executors are workers, not the brain.

Detailed: Lose the driver JVM and there is no one to retry tasks or return the action result. Databricks job clusters fail the run; notebooks disconnect.

Common mistake: Restarting one executor to 'fix a driver OOM'.

Follow-up: What Spark UI page is served from the driver?

Lesson · Simulation

beginner

What is an RDD, and why do we still mention it?

RDDs: lineage and partitions · tap to open the answer

Short: The original distributed collection. DataFrames compile down toward RDD/Tungsten execution.

Detailed: You rarely write RDD code now. Interviewers want: partitions, lineage, and that DataFrames are the API you should use.

Common mistake: Starting a new ETL in RDD map/reduce.

Follow-up: What is lineage on an RDD?

Lesson · Simulation

intermediate

Why is Dataset/DataFrame usually faster than hand-rolled RDDs?

RDDs: lineage and partitions · tap to open the answer

Short: Catalyst optimizes the plan; Tungsten uses off-heap/binary rows. RDDs are Java objects.

Detailed: RDD map on Row objects boxes everything. DataFrame expressions can whole-stage codegen. You also get predicate pushdown on files.

Common mistake: rdd.filter because 'it's more control'.

Follow-up: When is an RDD the right tool?

Lesson · Simulation

beginner

What is a DataFrame in Spark?

DataFrames and schemas · tap to open the answer

Short: A distributed table with a schema and a lazy query plan.

Detailed: Rows are partitioned across executors. Operations return a new plan, not a local pandas object. spark.createDataFrame on a huge Python list still starts on the driver.

Common mistake: Treating it like pandas (df[i] loops, collect to 'feel the data').

Follow-up: How do you print the schema without executing a job?

Lesson · Simulation

intermediate

When should you drop to RDDs from a DataFrame?

DataFrames and schemas · tap to open the answer

Short: Almost never for ETL. DataFrames get Catalyst and Tungsten. RDDs skip that.

Detailed: Use RDDs only for custom partitioning or APIs that still need them. groupBy/join/filter belong on DataFrames/Datasets.

Common mistake: rdd.map for a column expression Catalyst could optimize.

Follow-up: What do you lose when you call rdd.map on a DataFrame?

Lesson · Simulation

beginner

Two teammates write the same aggregation — one in spark.sql(), one in the DataFrame API. Which one runs faster?

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: Neither. Both go through Catalyst and end up as the same physical plan.

Detailed: A SQL string and a chain of DataFrame calls both become an unresolved logical plan, then get analyzed, optimized, and planned by the same code path. Run explain('formatted') on both and you see the same FileScan, Exchange, and HashAggregate nodes.

Common mistake: Claiming the DataFrame API is 'closer to the engine' so it skips a layer that costs runtime.

Follow-up: So what would actually make the two versions differ?

Lesson · Simulation

intermediate

Your notebook created a temp view; the next notebook on the same cluster can't see it. Explain.

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: createOrReplaceTempView is session-scoped, and the other notebook has its own SparkSession.

Detailed: Temp views live in the session catalog and disappear with the session. createGlobalTempView registers into the global_temp database, shared across sessions of the same application, and must be referenced as global_temp.name. Neither is in the metastore — only saveAsTable creates something another job can find tomorrow.

Common mistake: Assuming a temp view is a table because SELECT * worked in the same notebook.

Follow-up: When would you reach for a global temp view instead of just writing a table?

Lesson · Simulation

beginner

What is a transformation versus an action?

Narrow vs wide transformations · tap to open the answer

Short: Transformation: new DataFrame, lazy. Action: job, side effect or result.

Detailed: filter, select, groupBy, join are transformations. show, count, write, collect, take are actions. groupBy alone does not shuffle until an action.

Common mistake: Saying groupBy 'runs the shuffle immediately'.

Follow-up: Is cache a transformation or an action?

Lesson · Simulation

intermediate

Narrow vs wide transformation — give one example of each and the cost.

Narrow vs wide transformations · tap to open the answer

Short: Narrow: filter/map — no shuffle. Wide: groupBy/join — shuffle and a new stage.

Detailed: Narrow tasks read only their partition. Wide tasks wait on Exchange. Broadcast join is wide in API but not a shuffle of the fact table.

Common mistake: Calling join always a shuffle join.

Follow-up: Why can a filter after a groupBy not reduce shuffle bytes of that groupBy?

Lesson · Simulation

beginner

Name four Spark actions and what each returns to the driver.

Actions that trigger jobs · tap to open the answer

Short: count → a number. show → a few printed rows. collect → all rows. write → nothing (files on storage).

Detailed: take/limit also pull a small result. foreach runs on executors. The dangerous one is collect on a large frame.

Common mistake: Using collect() as the default way to 'see data'.

Follow-up: Which action is safest to check a pipeline ran?

Lesson · Simulation

intermediate

Why can count() be expensive even though it returns one integer?

Actions that trigger jobs · tap to open the answer

Short: It still executes the full plan, including shuffles.

Detailed: count after a wide transformation pays the shuffle. It is a cheap result, not a cheap job. Sometimes Spark can optimize count on a scan; not after a join.

Common mistake: Sprinkling count() after every transform in production.

Follow-up: When is approx_count_distinct the better interview answer?

Lesson · Simulation

beginner

What does lazy evaluation mean for a DataFrame?

Why Spark is lazy · tap to open the answer

Short: Spark records the recipe and waits for an action.

Detailed: That lets Catalyst optimize the whole chain (predicate pushdown, join reorder) instead of running each line. It also means errors can appear late.

Common mistake: Evaluating each transformation when the line runs, like pandas.

Follow-up: How do you force execution without collect()?

Lesson · Simulation

intermediate

You persist a DataFrame, then add a filter, then count. Why is cache unused?

Why Spark is lazy · tap to open the answer

Short: The cached plan is the unfiltered one; the new plan does not match, or you never materialized the cache.

Detailed: cache() is lazy. Need an action to populate Storage. A new filter is a different lineage unless you cache after the filter.

Common mistake: Assuming persist survives arbitrary extra operators.

Follow-up: How do you confirm a scan hit cache in Spark UI?

Lesson · Simulation

beginner

What is a Spark partition?

Partitions: the unit of parallelism · tap to open the answer

Short: A chunk of a DataFrame that one task reads in a stage.

Detailed: Input partitions often follow files. After a shuffle, spark.sql.shuffle.partitions (or AQE) decides how many. One partition = one task, not one executor.

Common mistake: One partition per executor as a rule.

Follow-up: Who decides input partitions for a Parquet scan?

Lesson · Simulation

intermediate

repartition vs coalesce — when do you use each?

Partitions: the unit of parallelism · tap to open the answer

Short: repartition shuffles to a new count. coalesce reduces without a full shuffle (can stay unbalanced).

Detailed: repartition(n) for even parallelism or a join key. coalesce(n) after a filter that dropped most rows, to avoid a shuffle. Don't coalesce to 1 on a huge frame except for a tiny output.

Common mistake: coalesce(1) to 'make a single CSV' on a 400 GB table.

Follow-up: What does partitionBy on write control versus RDD partitions?

Lesson · Simulation

beginner

repartition(200) and coalesce(200) both leave me with 200 partitions. What is actually different?

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: repartition does a full shuffle and evens out the data; coalesce merges existing partitions with no shuffle and keeps them uneven.

Detailed: repartition inserts an Exchange — round-robin, or hash if you pass columns — so partitions come out roughly equal in size. coalesce is a narrow operation that glues partitions together on the same executor, which is cheap but inherits whatever imbalance the input had.

Common mistake: Trying to use coalesce to increase the partition count — it can only reduce.

Follow-up: When is the uneven result of coalesce perfectly fine?

Lesson · Simulation

intermediate

Someone put coalesce(1) before the write to get a single output file, and now the entire pipeline runs single-threaded. Explain the mechanism.

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: coalesce has no shuffle boundary, so the reduced partition count pushes back and the upstream transformations run with one task too.

Detailed: Because coalesce does not create an Exchange, the join or aggregation feeding the write ends up inside the same stage at one partition — one task on one executor doing everything. repartition(1) adds an Exchange instead, so upstream work keeps its parallelism and only the final write is serial.

Common mistake: Believing coalesce only affects the write because that is where the call sits in the code.

Follow-up: So how do you actually produce one output file without wrecking the job?

Lesson · Simulation

beginner

What does a shuffle do?

What a shuffle actually does · tap to open the answer

Short: Moves records so equal keys land on the same reducer. Disk + network + a stage boundary.

Detailed: Map tasks write shuffle files. Reducers fetch their slice. groupBy, distinct, join (non-broadcast), and repartition all shuffle.

Common mistake: Thinking shuffle is 'Spark being slow' rather than a specific Exchange.

Follow-up: Which Spark UI numbers prove a shuffle?

Lesson · Simulation

intermediate

Why can a shuffle take longer than the compute?

What a shuffle actually does · tap to open the answer

Short: Network fetch, disk spill, and waiting on the slowest reducer.

Detailed: Shuffle write/read bytes, fetch wait, spill. A skewed key makes one reducer read most of the data. Compression and AQE help only if the plan is sane.

Common mistake: Adding CPU cores when fetch wait is the bottleneck.

Follow-up: What is a shuffle block?

Lesson · Simulation

beginner

Broadcast join vs shuffle join — when do you pick each?

Join strategies in Spark · tap to open the answer

Short: Broadcast if one side fits in memory on every executor. Shuffle (sort-merge) if both sides are large.

Detailed: Broadcast replicates the small table. SMJ partitions both by key and sorts. A wrong broadcast of a 'small' 10 GB table OOMs executors.

Common mistake: Broadcasting the fact table because 'joins should be broadcast'.

Follow-up: How do you force a broadcast in Spark SQL?

Lesson · Simulation

intermediate

The plan says SortMergeJoin but you expected broadcast. Why?

Join strategies in Spark · tap to open the answer

Short: Statistics said both sides were large, or AQE/broadcast threshold was below the build side.

Detailed: Check table stats, spark.sql.autoBroadcastJoinThreshold, and filters that Catalyst didn't push (so size is wrong). AQE can switch later if enabled.

Common mistake: Hints without checking stats.

Follow-up: What happens if stats are stale after a 10× load?

Lesson · Simulation

beginner

What does cache() actually do?

Cache and persist · tap to open the answer

Short: Marks the DataFrame to be stored after the next action. It is lazy.

Detailed: First action computes and stores partitions (MEMORY_AND_DISK by default for DataFrames). Second action can skip recomputation if the plan matches and data still fits.

Common mistake: cache() as an action that runs immediately.

Follow-up: How do you uncache?

Lesson · Simulation

intermediate

When is cache the wrong answer?

Cache and persist · tap to open the answer

Short: When you read the data once, or when the cached set is bigger than memory and thrashes disk.

Detailed: Cache for reuse in the same session (ML iterative, branching QA). For a single write, cache adds memory pressure. Checkpoint if lineage is huge.

Common mistake: Caching every intermediate table in a 30-step ETL.

Follow-up: MEMORY_ONLY vs MEMORY_AND_DISK — which fails harder?

Lesson · Simulation

beginner

Same rows, same cluster: why does the Parquet version of a table read so much faster than the CSV version?

File formats: why Parquet reads less · tap to open the answer

Short: Parquet is columnar, typed, and carries statistics — a CSV scan has to read and parse every byte of every line.

Detailed: A Parquet scan reads only the columns in your select and can skip entire row groups using min/max statistics. CSV has no schema, no column boundaries, and no stats, so Spark parses the whole file. The plan shows FileScan parquet with a ReadSchema listing just the columns you asked for.

Common mistake: Saying Parquet is faster 'because it's compressed' and stopping there.

Follow-up: What is still a fair use case for CSV or JSON?

Lesson · Simulation

intermediate

You read a CSV with inferSchema=true and a job appears in Spark UI before your action ever runs. What is it?

File formats: why Parquet reads less · tap to open the answer

Short: Schema inference is an eager extra pass over the data to guess column types.

Detailed: inferSchema reads the file (all of it for CSV by default) to decide types, which submits its own job ahead of your transformation. Passing an explicit StructType removes that pass; Parquet and ORC never need it because the schema is in the file footer.

Common mistake: Blaming that first slow cell on cluster warm-up instead of the inference job.

Follow-up: What else can inferSchema get wrong besides being slow?

Lesson · Simulation

beginner

You need a per-customer running total but the report still needs every transaction row. Window or groupBy?

Window functions: keep the row, add the answer · tap to open the answer

Short: Window — it computes over related rows but returns a value per row, so nothing collapses.

Detailed: groupBy + agg reduces each key to one row; a window function keeps input cardinality and adds a column. In the plan you get a WindowExec with a Sort under it instead of a HashAggregate, and no join-back step to reattach the aggregate.

Common mistake: Doing groupBy then joining the aggregate back to the detail — two shuffles for what a window does in one.

Follow-up: How would you use a window to grab the latest row per key?

Lesson · Simulation

intermediate

rowsBetween(-2, 0) and rangeBetween(-2, 0) on the same ordered window return different numbers. Why?

Window functions: keep the row, add the answer · tap to open the answer

Short: rowsBetween counts physical rows; rangeBetween compares values of the ordering column.

Detailed: rowsBetween is positional — the two preceding rows regardless of their values. rangeBetween is a value frame: every row whose ordering value falls inside the offset is included, so gaps shrink the frame and ties pull in all peers. With duplicate ordering values, rangeBetween rows share identical frames.

Common mistake: Using rangeBetween on a timestamp and expecting exactly three rows in the frame.

Follow-up: Which one do you need for a true 7-day rolling sum over daily data with missing days?

Lesson · Simulation

beginner

You wrote a Python UDF to uppercase a column. Why does a reviewer push back?

UDFs: the box Catalyst cannot see into · tap to open the answer

Short: Every row leaves the JVM for a Python process and comes back; upper() never leaves the JVM at all.

Detailed: A Python UDF serializes rows to a Python worker, evaluates them one at a time, and serializes results back — it shows up as BatchEvalPython in the plan. Built-ins like upper, regexp_replace, and to_timestamp compile into whole-stage codegen and operate on Tungsten binary rows directly.

Common mistake: Assuming the UDF is fine 'because it's just string formatting'.

Follow-up: Name a built-in you would use instead of a UDF you have written before.

Lesson · Simulation

intermediate

You moved a filter into a UDF and the scan went from 4 GB to 900 GB. Explain.

UDFs: the box Catalyst cannot see into · tap to open the answer

Short: Catalyst cannot see inside a UDF, so the predicate is not pushed into the scan and no partitions get pruned.

Detailed: A UDF is an opaque expression: it can never become a PushedFilter or a PartitionFilter, so Spark reads every file and filters afterwards above BatchEvalPython. explain('formatted') shows PushedFilters empty and the FileScan's Input Size / Records covering the whole table.

Common mistake: Blaming 'Python is slow' when the real cost is the extra 896 GB of I/O.

Follow-up: How do you keep the custom logic but restore pruning?

Lesson · Simulation

beginner

What runs on the driver versus an executor?

Driver vs executors · tap to open the answer

Short: Driver: SparkSession, Catalyst, DAG, task scheduling. Executor: tasks, cache, shuffle files.

Detailed: Your notebook code until an action is driver-side. After the action, work is split into tasks that executors run on partitions. Results come back only if the action asks (show, collect).

Common mistake: Thinking each executor has its own SparkSession you should create.

Follow-up: Why is creating a SparkSession inside a foreach a bug?

Lesson · Simulation

intermediate

Why does write() not send the table through the driver?

Driver vs executors · tap to open the answer

Short: Executors write partitions in parallel to storage. The driver only coordinates the commit.

Detailed: Each task writes its slice (Parquet/Delta files). The driver never materializes the full dataset. That is why write is safe and collect is not.

Common mistake: Using collect() 'to inspect' a 200 GB frame before writing.

Follow-up: What does show() send back?

Lesson · Simulation

beginner

Define job, stage, and task in one breath.

Jobs, stages, and tasks · tap to open the answer

Short: Job = one action. Stage = slice of the DAG between shuffles. Task = one partition in that stage.

Detailed: This is the vocabulary interviewers use to see if you've opened Spark UI. Mixing them is an instant no-hire for senior roles.

Common mistake: Calling an executor a stage.

Follow-up: How many tasks in a stage with 400 partitions?

Lesson · Simulation

intermediate

A SQL query shows 3 jobs. Is that a bug?

Jobs, stages, and tasks · tap to open the answer

Short: Not necessarily — multiple actions, AQE, or Databricks SQL extra jobs.

Detailed: count + write is two jobs. AQE can add. Temporary views plus display add more. Map jobs to actions in the notebook.

Common mistake: Assuming one SQL string is always one job.

Follow-up: Where in Spark UI do you map SQL to jobs?

Lesson · Simulation

beginner

What is the Spark DAG?

The Spark DAG · tap to open the answer

Short: The graph of RDD/DataFrame dependencies the scheduler uses to run stages.

Detailed: Narrow edges pipeline. Wide edges (shuffles) cut stages. Lineage lets Spark recompute lost partitions.

Common mistake: DAG as a Databricks workflow graph.

Follow-up: What happens to the DAG when an executor loses a cached partition?

Lesson · Simulation

intermediate

Why can a long lineage make a job fragile?

The Spark DAG · tap to open the answer

Short: Recompute after failure replays the whole chain; the DAG is huge.

Detailed: Checkpoint or write Delta to cut lineage. Spark UI DAG visualization gets unreadable after dozens of wide steps — that's a smell.

Common mistake: More cache() to 'shorten' a DAG without an action to materialize.

Follow-up: What's the difference between DAG Scheduler and Task Scheduler?

Lesson · Simulation

beginner

What is Catalyst?

Catalyst optimizer · tap to open the answer

Short: Spark SQL's optimizer: analysis → logical optimization → physical planning.

Detailed: It resolves names, pushes filters, reorders joins, then picks physical operators (broadcast vs SMJ). DataFrame/SQL get Catalyst; raw RDDs do not.

Common mistake: Catalyst as a storage format.

Follow-up: At which phase do unresolved attributes fail?

Lesson · Simulation

intermediate

How do you read explain('formatted') in an interview?

Catalyst optimizer · tap to open the answer

Short: Bottom is the scan. Look for PushedFilters, PartitionFilters, Exchange, and the join type.

Detailed: If the filter isn't in PushedFilters, you're scanning extra files. Exchange means shuffle. BroadcastHashJoin vs SortMergeJoin is the money line.

Common mistake: Pasting explain() without pointing at an operator.

Follow-up: What does AdaptiveSparkPlan mean?

Lesson · Simulation

beginner

What is Tungsten?

Tungsten execution engine · tap to open the answer

Short: Spark's execution engine: binary rows, off-heap memory, whole-stage codegen.

Detailed: It avoids Java object overhead on the hot path. You get it with DataFrame/SQL. Python UDFs jump back to objects and kill the benefit.

Common mistake: Tungsten as a Databricks SKU.

Follow-up: What operator in explain() shows whole-stage codegen?

Lesson · Simulation

intermediate

Why does a Python UDF disable the fast path?

Tungsten execution engine · tap to open the answer

Short: Rows must be deserialized into Python, one call at a time (or batches for pandas UDFs).

Detailed: Plan shows BatchEvalPython / PythonUDF. Photon also bails out. Rewrite with Spark functions or Scala.

Common mistake: Pandas UDF as 'basically native'.

Follow-up: When is a pandas UDF acceptable?

Lesson · Simulation

beginner

An executor has 32 GB. How much of that can you actually cache into?

Executor memory: execution wins, cache loses · tap to open the answer

Short: Nowhere near all of it — the heap is split into reserved memory, a unified execution/storage region, and user memory.

Detailed: spark.memory.fraction decides how much of the usable heap the unified region gets, and inside that spark.memory.storageFraction is the share cache is guaranteed to hold onto. Whatever sits outside the fraction holds your own objects and Spark internals. Execution memory is what shuffles, sorts, joins, and aggregations consume.

Common mistake: Treating spark.executor.memory as the cache budget.

Follow-up: What happens when execution needs memory that cache is currently holding?

Lesson · Simulation

intermediate

Execution and storage share one pool. Who is allowed to evict whom?

Executor memory: execution wins, cache loses · tap to open the answer

Short: Execution can evict cached blocks; storage can never push out execution memory.

Detailed: Either side can borrow free space, but execution wins the tug-of-war and will evict cached partitions down to the storageFraction floor. That is why a DataFrame you cached quietly disappears and the next action rescans the source — the Storage tab shows the cached fraction dropping and the plan goes back to a FileScan.

Common mistake: Assuming cache() guarantees the data stays in memory for the rest of the session.

Follow-up: How would you spot that silent eviction in the UI?

Lesson · Simulation

beginner

Your streaming query looks like ordinary DataFrame code. What is the engine actually doing under it?

Structured Streaming: an unbounded table · tap to open the answer

Short: Running a series of micro-batches: each batch reads a new offset range and executes the same incremental plan.

Detailed: The engine records which offsets it has processed in the checkpoint's offset log, plans a batch over the new range, and writes a commit when the batch finishes. That is why Spark UI shows a steady stream of small jobs rather than one long-running job.

Common mistake: Describing it as row-by-row streaming, like a message-queue consumer loop.

Follow-up: What guarantees the same row is not processed twice after a restart?

Lesson · Simulation

intermediate

On-call deleted the checkpoint directory to 'clear a stuck stream'. What did they just do?

Structured Streaming: an unbounded table · tap to open the answer

Short: Threw away the offset log, the commit log, and all state — the stream now restarts with no memory of what it processed.

Detailed: The checkpoint location holds offsets, commits, and the state store, and exactly-once depends on those lining up with the sink. Deleting it means reprocessing from the configured starting position — duplicates into an append sink, or silently skipped data if it starts at latest — and every aggregation's state is gone.

Common mistake: Treating the checkpoint directory as a cache you can safely clear when something looks stuck.

Follow-up: When is deleting a checkpoint actually the correct move?

Lesson · Simulation

beginner

What is data skew in Spark?

Data skew · tap to open the answer

Short: One key (or partition) has far more data than the others, so one task defines stage time.

Detailed: Classic: null country, 'US', or one customer_id. Median task is fine; max task is the SLA.

Common mistake: Skew as 'the cluster is unbalanced' without a key.

Follow-up: Which UI view shows skew in 10 seconds?

Lesson · Simulation

intermediate

Name three skew fixes and when each applies.

Data skew · tap to open the answer

Short: Filter/isolate hot keys, salt the key, AQE skew join.

Detailed: If nulls are junk, drop them. If one key is real (US), process it separately or salt. AQE split can help SMJ skew. Broadcast if the other side is small — skew on the fact may not matter.

Common mistake: repartition(1000) as a skew fix — the hot key still hashes to one partition.

Follow-up: Why doesn't more shuffle partitions fix a single hot key?

Lesson · Simulation

beginner

What is a driver OOM?

Driver OOM · tap to open the answer

Short: The driver JVM ran out of heap. Executors may still be healthy.

Detailed: Classic causes: collect, toPandas, createDataFrame from a huge list, broadcast of a large table. The notebook kernel dies.

Common mistake: Adding executors to fix a driver OOM.

Follow-up: Which Spark UI tab shows driver memory?

Lesson · Simulation

intermediate

How do you rewrite collect() on a 2 TB frame?

Driver OOM · tap to open the answer

Short: Don't. Aggregate, sample, write, or take(n).

Detailed: If you need a local ML sample, sample() then limit, or write a sampled Delta table and read elsewhere. Never collect the grain.

Common mistake: Increasing driver memory from 8 GB to 16 GB as the plan.

Follow-up: What's the Databricks-specific cousin of collect()?

Lesson · Simulation

beginner

How is executor OOM different from driver OOM?

Executor OOM · tap to open the answer

Short: A worker JVM dies. The driver usually stays up. You see ExecutorLostFailure / container killed.

Detailed: Causes: fat task (skew), huge partition, MEMORY_ONLY cache, big broadcast on the executor, explode() blowup.

Common mistake: Restarting the driver to fix an executor OOM.

Follow-up: What does Spark do after an executor is lost?

Lesson · Simulation

intermediate

A task explodes a JSON array and dies. What's the fix?

Executor OOM · tap to open the answer

Short: The partition became huge after explode — not the scan.

Detailed: explode multiplies rows in one task. Repartition before explode, filter arrays, or process hot keys separately. Memory fraction / spark.memory.fraction is a last resort.

Common mistake: spark.executor.memory 64g as the first change.

Follow-up: How do you see the task that died?

Lesson · Simulation

beginner

How does a broadcast join work?

Broadcast joins · tap to open the answer

Short: The driver (then each executor) gets a copy of the small table; the big table stays put.

Detailed: No shuffle of the fact table. Threshold is spark.sql.autoBroadcastJoinThreshold (default 10 MB, often raised). Too-large broadcast OOMs the driver or executors.

Common mistake: Broadcasting whichever table is mentioned first in SQL.

Follow-up: Which table should be broadcast?

Lesson · Simulation

intermediate

Broadcast join OOM — driver or executor?

Broadcast joins · tap to open the answer

Short: Either: driver collects the build side; executors hold the hashed relation.

Detailed: Huge broadcast: driver collect of the dimension, then per-executor copy. Spark UI SQL shows broadcast exchange size. If the 'small' side is 8 GB, you chose wrong.

Common mistake: Raising executor memory without measuring broadcast size.

Follow-up: What's the hint syntax and when do you undo it?

Lesson · Simulation

beginner

What is Adaptive Query Execution?

Adaptive Query Execution · tap to open the answer

Short: Spark re-plans later stages at runtime using real sizes.

Detailed: Three headlines: coalesce shuffle partitions, switch join strategy, handle skewed joins. Enabled by default in modern Spark / Databricks.

Common mistake: AQE as a replacement for designing a good join key.

Follow-up: Does AQE run before the first stage?

Lesson · Simulation

intermediate

When will AQE not save you?

Adaptive Query Execution · tap to open the answer

Short: A single skewed key, a Python UDF, or a broadcast that's already too big.

Detailed: Coalesce won't split a hot key. Join conversion needs a truly small side. Skew join has limits. You still need a sane plan and files.

Common mistake: Turning every knobs to true and calling it architecture.

Follow-up: Which AQE feature shows up as extra stages in the UI?

Lesson · Simulation

Practice by topic