Curriculum
Learn Apache Spark and Databricks
Apache Spark tutorial and Databricks tutorial in one curriculum. Every lesson follows the same path: concept, visualization, PySpark code, internals, a production failure, and interview questions.
Simulation
Prefer watching it run? Open the interactive simulators for Spark and Databricks — step through shuffle, joins, Delta, Unity Catalog, and more.
- Spark Execution SimulatorSpark · beginner
- Cluster ArchitectureSpark · beginner
- Lazy EvaluationSpark · beginner
- Partitions & TasksSpark · beginner
- Shuffle SimulatorSpark · intermediate
- Join SimulatorSpark · intermediate
- Databricks ArchitectureDatabricks · beginner
- Compute typesDatabricks · beginner
- Unity CatalogDatabricks · senior
- Delta LakeDatabricks · intermediate
- Medallion ArchitectureDatabricks · beginner
- Delta MERGEDatabricks · intermediate
- OPTIMIZE & ClusteringDatabricks · intermediate
- Auto LoaderDatabricks · intermediate
- Delta Live TablesDatabricks · advanced
- PhotonDatabricks · advanced
- SQL WarehousesDatabricks · beginner
- Time travel & VACUUMDatabricks · intermediate
- Liquid ClusteringDatabricks · advanced
- WorkflowsDatabricks · intermediate
41 simulators total · see all
Learn Spark
- beginnerFundamentals
Spark Overview
What Spark is, why it exists, and the mental model of driver, executors, and lazy plans.
- beginnerFundamentals
Spark Architecture
Cluster manager, driver JVM, executor JVMs, and how a SparkSession talks to a cluster.
- beginnerFundamentals
Driver & Executors
The driver plans. Executors run tasks. Never ship a huge collect() back to the driver.
- beginnerFundamentals
Execution Flow
Trace groupBy + show() from lazy code through the plan, job, stages, tasks, shuffle, and result — one picture.
- beginnerFundamentals
DataFrame
Structured data with a schema. Catalyst can optimize DataFrames in ways RDDs cannot.
- beginnerFundamentals
Spark SQL
SQL strings and the DataFrame API compile to the same plan. Temp views, the catalog, and when names actually resolve.
- beginnerFundamentals
Transformations
Narrow vs wide transformations, and why nothing runs until an action arrives.
- beginnerFundamentals
Actions
show, count, collect, write — the calls that force Spark to produce a result.
- beginnerFundamentals
Lazy Evaluation
Spark builds a DAG of transformations and waits. That wait is a feature, not a bug.
- intermediateFundamentals
RDD
Resilient Distributed Datasets: partitions, lineage, and when you still need the RDD API.
- beginnerCore Concepts
Partitions
A partition is a chunk of data. One task processes one partition. Size them for healthy runtime.
- intermediateCore Concepts
Repartition vs Coalesce
repartition shuffles for even partitions, coalesce only merges, and write.partitionBy is a different thing entirely.
- intermediateCore Concepts
Shuffle
Records move across the network so matching keys land on the same reducer.
- intermediateCore Concepts
Joins
Sort-merge, shuffle hash, and broadcast. The strategy Spark picks changes the cost by orders of magnitude.
- intermediateCore Concepts
Caching
persist and cache store computed partitions so later actions do not recompute the whole DAG.
- intermediateCore Concepts
File Formats
Why Parquet reads a fraction of the bytes CSV does: column pruning, row-group stats, and embedded schema.
- intermediateSQL & Functions
Window Functions
Rank, lag, and running totals keep every row. partitionBy is a shuffle — and forgetting it funnels everything to one task.
- intermediateSQL & Functions
UDFs
A Python UDF is a black box Catalyst cannot optimize. Built-ins, higher-order functions, then pandas_udf.
Spark Internals
- intermediateExecution
Job / Stage / Task
An action creates a job. Shuffle boundaries split jobs into stages. Stages split into tasks.
- intermediateExecution
DAG
The directed acyclic graph of RDDs / DataFrame operators Spark will execute.
- advancedQuery Engine
Catalyst
Logical plan → analyzed → optimized → physical plan. Exchange nodes mark shuffles.
- advancedQuery Engine
Tungsten
Off-heap memory, whole-stage codegen, and cache-aware layout that make Spark SQL fast.
- advancedMemory
Memory Management
Execution and storage share one pool. Execution can evict your cache, then spill to disk, then die.
- advancedStreaming
Structured Streaming
A stream is an unbounded table read in micro-batches. Checkpoints give exactly-once; watermarks bound state.
Performance
- intermediatePerformance
Broadcast Join
Replicate a small table to every executor and skip the shuffle on the large side.
- seniorPerformance
Data Skew
One key owns most of the data. One task runs for 40 minutes while 199 finish in 2.
- advancedPerformance
AQE
Adaptive Query Execution coalesces shuffle partitions, switches join strategy, and handles skew at runtime.
- seniorPerformance
Driver OOM
collect() and toPandas() pull the cluster into one JVM. Executors look fine. The driver is dead.
- seniorPerformance
Executor OOM
Skew, MEMORY_ONLY cache, or a fat broadcast blows an executor heap. The driver is still alive.
Databricks
- intermediatePlatform
Databricks Architecture
Control plane vs data plane: workspace in Databricks SaaS, clusters and files in your cloud.
- beginnerPlatform
Compute types
All-purpose clusters, job clusters that terminate, and SQL warehouses for BI.
- beginnerPlatform
SQL Warehouses
Serverless vs Pro vs Classic, the caches behind a fast dashboard, and why size and cluster count are different dials.
- seniorPlatform
Unity Catalog
Three-level namespace, governed storage, and how permissions attach to catalogs, schemas, and tables.
- intermediateLakehouse
Delta Lake
ACID transactions on object storage: the _delta_log, Parquet data files, time travel, and OPTIMIZE.
- intermediateLakehouse
Time Travel & VACUUM
Read an old version from the commit log, RESTORE as a new forward commit, and the VACUUM that ends time travel.
- beginnerLakehouse
Medallion Architecture
Bronze keeps the raw files. Silver is typed and joined. Gold is the metric table BI queries.
- intermediateLakehouse
Delta MERGE
Upserts a change set into Delta by rewriting only the files the join key touches.
- intermediateLakehouse
OPTIMIZE & Clustering
Compact tiny files, Z-ORDER for skipping, and liquid clustering on write.
- advancedLakehouse
Liquid Clustering
CLUSTER BY replaces rigid directories and full Z-ORDER rewrites — and lets you change the keys later.
- intermediatePipelines
Auto Loader
Incremental ingest from object storage with notifications, checkpoint, and schema evolution.
- advancedPipelines
Delta Live Tables
Declare medallion tables and expectations. DLT infers the graph and stops downstream on failure.
- intermediatePipelines
Workflows
A job is a DAG of tasks with real parallelism, per-task retries, and a repair run that skips what already passed.
- advancedEngine
Photon
Native vectorized engine for eligible operators — and the UDF fallback that cancels it.
https://www.datayatri.com
