Simulation/Spark

Educational simulation
Spark

Repartition vs coalesce

Full shuffle, cheap merge, or a directory layout on write.

Watch the canvas:skew / shuffleeven partitionscoalesce (merge)hash by keystorage layoutLive simulation
Start · 4 skewed partitionsP01.2 GB — hot keyP180 MBP260 MBP340 MBrepartition(8) · full shuffle → 8 even partitionsP0P1P2P3P4P5P6P7coalesce(2) · merge in place, no networkP0 + P1 merged1.28 GB — still skewed, nothing movedP2 + P3 merged100 MB — cheap, but upstream parallelism drops to 2repartition("region") · hash(key) % nhash(US) → P0every row with this key, one taskhash(EU) → P1every row with this key, one taskhash(IN) → P2every row with this key, one taskwrite.partitionBy("region") · directories on storage/region=US/part-0000*.snappy.parquet/region=EU/part-0000*.snappy.parquet/region=IN/part-0000*.snappy.parquetA stage ends when its slowest task ends. P0 is the stage.
repartition shuffles, coalesce merges, write.partitionBy lays out directories. Three different mechanisms with confusingly similar names.