databricks
intermediate
Connecting…

OPTIMIZE, Z-ORDER, Liquid Clustering

Small files from MERGE and streams get compacted. Z-ORDER and liquid clustering make filters skip files.

Explain it at my level

  1. Small files
  2. OPTIMIZE
  3. Z-ORDER
  4. Liquid
Watch the canvas:tiny filecompactedclustered by regionLive simulation
2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB2 MB16 files × 2 MB = 16 tiny tasks. Compact before you tune Spark.
Layout is a performance feature. OPTIMIZE and clustering change how files sit.