spark/fundamentals
beginner
Connecting…

DataFrames and schemas

A DataFrame is a distributed table with a schema. Catalyst can optimize it; an RDD of tuples cannot.

Lesson 5 of 29 · Spark path

Explain it at my level

  1. Schema
  2. Partitions
  3. Lazy plan
  4. show()
Watch the canvas:schemapartitioned rowsactionLive simulation
orders DataFrameorder_id: long · region: string · amount: decimalPartition 0executor-1Partition 1executor-2Partition 2executor-3filter(region == 'US').select(...)still a plan — no scan yetshow()action — few rows to driverNamed columns. This is the contract Catalyst will optimize.
DataFrame = distributed table + schema. That schema is why Catalyst exists.