Main feature
Apache Spark and Databricks simulators
Step through Spark shuffle, joins, Catalyst, Delta Lake, and Unity Catalog on an animated canvas. Educational simulations — not a live cluster.
Spark
- beginnerSpark
Spark Execution Simulator
Five labs from lazy code → plan → job → stages → tasks → shuffle → result.
- beginnerSpark
Cluster Architecture
Driver, cluster manager, and executors — who does what.
- beginnerSpark
Lazy Evaluation
Watch transformations pile up until an action finally runs.
- beginnerSpark
Partitions & Tasks
How a table splits into partitions and one task per partition.
- intermediateSpark
Shuffle Simulator
Watch keys hash to reducers across the network.
- intermediateSpark
Join Simulator
Broadcast vs shuffle join — when the big table should stay put.
- intermediateSpark
Caching & Persist
First action computes; second action hits the cache.
- intermediateSpark
Catalyst Optimizer
Unresolved → analyzed → optimized → physical plan.
- seniorSpark
Data Skew Simulator
One fat key vs a salted plan — stage time is the max.
- advancedSpark
AQE Simulator
Runtime coalesce, join switch, and skew split.
- seniorSpark
Driver OOM
collect() piles rows onto the driver. Executors stay green. The notebook dies.
- seniorSpark
Executor OOM
One fat task fills Spark memory. The driver stays up. That worker is gone.
- beginnerSpark
DataFrames & schemas
A distributed table with a schema — not an RDD of tuples, not pandas on the driver.
- beginnerSpark
Narrow vs wide transforms
filter stays put. groupBy shuffles. Neither runs until an action.
- beginnerSpark
Actions
show, count, collect, write — four actions, four driver costs.
- intermediateSpark
RDD lineage
Partitions plus lineage. DataFrames sit on top and get Catalyst.
- intermediateSpark
Spark DAG
Operator graph. Wide edges cut stages. Not a Databricks workflow.
- intermediateSpark
Jobs, stages, tasks
One action, shuffle cuts, one task per partition.
- advancedSpark
Tungsten
Binary rows and whole-stage codegen — until a Python UDF.
- intermediateSpark
Broadcast join
Copy the small side. The fact table does not shuffle.
- beginnerSpark
Spark SQL & the catalog
A SQL string and a DataFrame converge on one physical plan.
- intermediateSpark
Window functions
Keep every row, add a column. partitionBy is the shuffle.
- intermediateSpark
UDFs vs built-ins
Watch rows cross the JVM to Python boundary, then batch them.
- intermediateSpark
File formats
Bytes read: CSV reads everything, Parquet reads a column.
- intermediateSpark
Repartition vs coalesce
Full shuffle, cheap merge, or a directory layout on write.
- advancedSpark
Executor memory
Execution evicts cache, spills to disk, then the heap dies.
- advancedSpark
Structured Streaming
Micro-batches, a checkpoint, and a watermark dropping state.
Databricks
- beginnerDatabricks
Databricks Architecture
Control plane vs data plane: workspace in Databricks SaaS, clusters and files in your cloud.
- beginnerDatabricks
Compute types
All-purpose vs jobs cluster vs SQL warehouse.
- seniorDatabricks
Unity Catalog
catalog.schema.table, grants, and lineage across layers.
- intermediateDatabricks
Delta Lake
Parquet files + _delta_log commits, ACID, and time travel.
- beginnerDatabricks
Medallion Architecture
Bronze → silver → gold. Three layers, three contracts.
- intermediateDatabricks
Delta MERGE
Match keys, rewrite files, commit one atomic snapshot.
- intermediateDatabricks
OPTIMIZE & Clustering
Tiny files compact, then Z-ORDER / CLUSTER BY skip filters.
- intermediateDatabricks
Auto Loader
New files notify, ingest, checkpoint — no full bucket list.
- advancedDatabricks
Delta Live Tables
Declared tables, inferred graph, expectations, ordered commits.
- advancedDatabricks
Photon
Native operators vs a UDF that falls back to the JVM.
- beginnerDatabricks
SQL Warehouses
Serverless start-up, the cache layers, and the two size dials.
- intermediateDatabricks
Time travel & VACUUM
Replay the commit log, RESTORE forward, then VACUUM the past.
- advancedDatabricks
Liquid Clustering
Files scanned: partitions vs Z-ORDER vs CLUSTER BY.
- intermediateDatabricks
Workflows
A task DAG with fan-out, per-task retries, and skipped children.
