Simulation/Spark

Educational simulation
Spark

File formats

Bytes read: CSV reads everything, Parquet reads a column.

Watch the canvas:row-oriented costcolumnar layoutqueryencodingbytes savedLive simulation
Same 4 GB table, two layoutsorders.csv — row oriented40 columns interleaved on every lineuntyped text · no footer · no statisticsorders.parquet — columnarrow groups → per-column chunksschema + min/max in the footerQueryselect("region", "amount").where("amount > 500")2 of 40 columns · one predicateinferSchemaextra full scanParquet row groups · footer min/maxRow group 0amount 12–380Row group 1amount 410–990Row group 2amount 505–870Row group 3amount 5–260Per-column encoding + compressionDictionarylow-cardinality strings → codesRLE + bit-packingrepeated values collapseSnappy / ZSTD per chunkstill splittable, unlike gzip CSVBytes read for this queryCSV4.0 GB · every byte read and parsedParquet4.0 GB · nothing pruned yetRow layout means reaching field 7 costs you fields 1 through 6.
Column pruning, row-group skipping and per-column encoding compound. File format usually beats cluster size.