Interview/Spark

Spark interview questions: Repartition vs coalesce vs partitionBy

Interactive Spark interview questions on Repartition vs coalesce vs partitionBy. Practice with the matching lesson. repartition shuffles for even partitions, coalesce only merges, and write.partitionBy is a different thing entirely.

Lesson · Simulation

repartition(200) and coalesce(200) both leave me with 200 partitions. What is actually different?

Answer it out loud, then reveal. Play steps through like the simulators.

Production scenario

Repartition vs coalesce vs partitionBy

coalesce(1) before write turned a 40-executor cluster into one busy core

Symptoms

  • One executor working while 39 sit idle for hours
  • Job went from 8 minutes to 3 hours after 'we need a single CSV'
  • One task in the write stage — and in the stage feeding it

All questions on this page

Indexed as FAQ. Open any item if you prefer a list to Play.

beginner

repartition(200) and coalesce(200) both leave me with 200 partitions. What is actually different?

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: repartition does a full shuffle and evens out the data; coalesce merges existing partitions with no shuffle and keeps them uneven.

Detailed: repartition inserts an Exchange — round-robin, or hash if you pass columns — so partitions come out roughly equal in size. coalesce is a narrow operation that glues partitions together on the same executor, which is cheap but inherits whatever imbalance the input had.

Common mistake: Trying to use coalesce to increase the partition count — it can only reduce.

Follow-up: When is the uneven result of coalesce perfectly fine?

Lesson · Simulation

intermediate

Someone put coalesce(1) before the write to get a single output file, and now the entire pipeline runs single-threaded. Explain the mechanism.

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: coalesce has no shuffle boundary, so the reduced partition count pushes back and the upstream transformations run with one task too.

Detailed: Because coalesce does not create an Exchange, the join or aggregation feeding the write ends up inside the same stage at one partition — one task on one executor doing everything. repartition(1) adds an Exchange instead, so upstream work keeps its parallelism and only the final write is serial.

Common mistake: Believing coalesce only affects the write because that is where the call sits in the code.

Follow-up: So how do you actually produce one output file without wrecking the job?

Lesson · Simulation

senior

You have a job that joins on customer_id and then windows on customer_id. When is repartition('customer_id') worth paying for?

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: When one shuffle on that key gets reused by more than one downstream operator.

Detailed: repartition('customer_id') hash-partitions rows so matching keys are co-located, and a following join or window on the same key can then skip its own Exchange. If only one operator needs the key, it would have shuffled once anyway — you have just paid for two shuffles instead of one.

Senior: explain('formatted') and count Exchange nodes before and after the change. If the count did not drop, the repartition bought you nothing.

Common mistake: Adding repartition(col) before every join out of habit and doubling the shuffle count.

Follow-up: How do you verify the second shuffle actually disappeared?

Lesson · Simulation

architect

spark.sql.shuffle.partitions is 200 and the code also calls repartition(1000). Which one wins, and where?

Repartition vs coalesce vs partitionBy · tap to open the answer

Short: They have different scopes: the config sizes the engine's own Exchanges, repartition(1000) sizes that one explicit shuffle.

Detailed: Every groupBy or sort-merge join Exchange targets spark.sql.shuffle.partitions unless AQE coalesces it at runtime; repartition(1000) is its own Exchange with its own target, and a later groupBy re-shuffles back to the configured number. df.write.partitionBy is a third, unrelated thing — it controls the on-disk directory layout, not in-memory partition count.

Senior: repartition on the same columns you pass to write.partitionBy, so each task owns one output directory and writes one decent-sized file instead of every task dropping a fragment into every directory.

Common mistake: Conflating write.partitionBy with repartition and expecting one to control the other.

Follow-up: How do you get a sane directory layout and reasonable file sizes at the same time?

Lesson · Simulation