databricks/lakehouse
advanced
Connecting…

Liquid Clustering: change your mind later

CLUSTER BY replaces rigid directories and full Z-ORDER rewrites — and lets you change the keys later.

Lesson 10 of 14 · Databricks path

Explain it at my level

  1. Hive dirs
  2. Z-ORDER
  3. CLUSTER BY
  4. Incremental
  5. Re-key
Watch the canvas:hive directoriesz-ordered filesliquid clusterednewly written dataLive simulation
Physical layout — directory treeevent_date=09-08/country=US · DE · IN …280 dirs · 2 MB filesevent_date=09-09/country=US · DE · IN …280 dirs · 2 MB filesevent_date=09-10/country=US · DE · IN …280 dirs · 2 MB filesevent_date=09-11/country=US · DE · IN …280 dirs · 2 MB filesFiles scanned — WHERE country='US' AND event_date='2026-09-10'Hive240 files · 3 MB each · task overhead winsWhat it costs to change your mindPARTITIONED BYevent_date, countrypaths are the layoutOPTIMIZE ZORDER BYcountry, event_datemust be re-run nightlyCLUSTER BYtable property · no dirsreplaces bothHigh-cardinality partition columns produce tiny files you cannot un-choose.
Three ways to lay out the same table — and only one of them lets you change your mind.