Interview/Spark

Spark interview questions: File formats: why Parquet reads less

Interactive Spark interview questions on File formats: why Parquet reads less. Practice with the matching lesson. Why Parquet reads a fraction of the bytes CSV does: column pruning, row-group stats, and embedded schema.

Lesson · Simulation

Same rows, same cluster: why does the Parquet version of a table read so much faster than the CSV version?

Answer it out loud, then reveal. Play steps through like the simulators.

Production scenario

File formats: why Parquet reads less

Dashboard query got 30x slower after a 'harmless' upstream format change

Symptoms

  • Upstream team switched a landing table from Parquet to JSON
  • Same row count, same cluster, query time up roughly 30x
  • The scan is now most of the wall clock

All questions on this page

Indexed as FAQ. Open any item if you prefer a list to Play.

beginner

Same rows, same cluster: why does the Parquet version of a table read so much faster than the CSV version?

File formats: why Parquet reads less · tap to open the answer

Short: Parquet is columnar, typed, and carries statistics — a CSV scan has to read and parse every byte of every line.

Detailed: A Parquet scan reads only the columns in your select and can skip entire row groups using min/max statistics. CSV has no schema, no column boundaries, and no stats, so Spark parses the whole file. The plan shows FileScan parquet with a ReadSchema listing just the columns you asked for.

Common mistake: Saying Parquet is faster 'because it's compressed' and stopping there.

Follow-up: What is still a fair use case for CSV or JSON?

Lesson · Simulation

intermediate

You read a CSV with inferSchema=true and a job appears in Spark UI before your action ever runs. What is it?

File formats: why Parquet reads less · tap to open the answer

Short: Schema inference is an eager extra pass over the data to guess column types.

Detailed: inferSchema reads the file (all of it for CSV by default) to decide types, which submits its own job ahead of your transformation. Passing an explicit StructType removes that pass; Parquet and ORC never need it because the schema is in the file footer.

Common mistake: Blaming that first slow cell on cluster warm-up instead of the inference job.

Follow-up: What else can inferSchema get wrong besides being slow?

Lesson · Simulation

senior

Your predicate is on a non-partition column and the job still only touches 3% of the files. What made that work, and when does it stop working?

File formats: why Parquet reads less · tap to open the answer

Short: Parquet row-group min/max statistics let Spark skip groups whose value range cannot match the predicate.

Detailed: Every row group stores per-column min/max; with pushdown the reader drops groups outside the predicate range, which shows up as a small Input Size / Records relative to the table. It only works when data is physically clustered on that column — if values are scattered, every row group's min/max spans the whole domain and nothing is pruned.

Senior: Sort or cluster on the filter column at write time (Z-ORDER or liquid clustering on Delta). Statistics are only as useful as the physical layout — this is a layout decision, not a config flag.

Common mistake: Calling row-group skipping 'an index' and expecting it to work on randomly ordered data.

Follow-up: How do you make the data cooperate?

Lesson · Simulation

architect

The team wants gzip because it compresses the smallest. What do you tell them?

File formats: why Parquet reads less · tap to open the answer

Short: A gzip text file is not splittable, so one file becomes one task no matter how large it is.

Detailed: Gzip must be decompressed from the beginning, so Spark cannot split a gzip CSV or JSON file — a 20 GB file is a single task and the rest of the cluster idles. Snappy and zstd inside Parquet are applied per page and row group, so the file stays splittable; zstd trades some CPU for noticeably smaller files than snappy.

Senior: It is the opposite failure: thousands of tiny Parquet files make per-file open and metadata cost dominate, with the task duration histogram piled near zero. Compact on write or with OPTIMIZE and aim for files in the hundreds of MB.

Common mistake: Comparing codecs only on compressed size and ignoring splittability and decode cost.

Follow-up: Where does the small-file problem fit into this conversation?

Lesson · Simulation