Short: High-cardinality partition columns create a directory per value so every write drops tiny files; recreate the table with CLUSTER BY (customer_id, event_date).
Detailed: Hive partitioning makes the physical layout a function of cardinality, so a million customers means a million directories and a scan that launches a task per tiny file. Liquid clustering decouples key choice from directory structure: files are written at a normal target size and the keys only decide which rows sit together, with file-level min/max statistics doing the skipping.
Senior: Partitioning is a physical commitment; clustering is a layout the engine maintains for you. I would migrate with a CTAS into a CLUSTER BY table, cut readers over, then verify in the query profile — files pruned versus files read, and bytes scanned before and after. The real win is not that clustering is magic, it is that re-keying later costs an ALTER instead of a full rewrite, so the choice stops being permanent.
Common mistake: Adding more OPTIMIZE runs while keeping the partition column, so compaction fights the partition boundaries forever.
Follow-up: How do you prove skipping actually improved after the migration?
Lesson · Simulation