spark/core
intermediate
Connecting…

Cache and persist

Caching stores computed partitions so later actions do not recompute the DAG. Unpersist when you are done.

Lesson 15 of 29 · Spark path

Explain it at my level

  1. Compute
  2. Store
  3. Reuse
Watch the canvas:computed stagecached blockLive simulation
First action — count()scanread 1 TBshufflenetwork moveaggregatesum by regionPartitions discardedlineage is all that remains2nd count()would recompute everythingunpersist()give the memory back
Caching buys the second action a shortcut, using the first action's work.