spark/core
intermediate
Connecting…

Repartition vs coalesce vs partitionBy

repartition shuffles for even partitions, coalesce only merges, and write.partitionBy is a different thing entirely.

Lesson 12 of 29 · Spark path

Explain it at my level

  1. Skew
  2. Repartition
  3. Coalesce
  4. Hash by key
  5. Directories
Watch the canvas:skew / shuffleeven partitionscoalesce (merge)hash by keystorage layoutLive simulation
Start · 4 skewed partitionsP01.2 GB — hot keyP180 MBP260 MBP340 MBrepartition(8) · full shuffle → 8 even partitionsP0P1P2P3P4P5P6P7coalesce(2) · merge in place, no networkP0 + P1 merged1.28 GB — still skewed, nothing movedP2 + P3 merged100 MB — cheap, but upstream parallelism drops to 2repartition("region") · hash(key) % nhash(US) → P0every row with this key, one taskhash(EU) → P1every row with this key, one taskhash(IN) → P2every row with this key, one taskwrite.partitionBy("region") · directories on storage/region=US/part-0000*.snappy.parquet/region=EU/part-0000*.snappy.parquet/region=IN/part-0000*.snappy.parquetA stage ends when its slowest task ends. P0 is the stage.
repartition shuffles, coalesce merges, write.partitionBy lays out directories. Three different mechanisms with confusingly similar names.