Container killed with no OutOfMemoryError anywhere in the logs
Symptoms
- Executors lost repeatedly, with a kill reason from the cluster manager rather than the JVM
- No OutOfMemoryError in executor stderr
- It always dies during the same pandas UDF stage
Interactive Spark interview questions on Executor memory: execution wins, cache loses. Practice with the matching lesson. Execution and storage share one pool. Execution can evict your cache, then spill to disk, then die.
Question 1 of 4
An executor has 32 GB. How much of that can you actually cache into?
Answer it out loud, then reveal. Play steps through like the simulators.
Case 1 of 2 · symptoms
Symptoms
Indexed as FAQ. Open any item if you prefer a list to Play.
An executor has 32 GB. How much of that can you actually cache into?
Executor memory: execution wins, cache loses · tap to open the answer
Short: Nowhere near all of it — the heap is split into reserved memory, a unified execution/storage region, and user memory.
Detailed: spark.memory.fraction decides how much of the usable heap the unified region gets, and inside that spark.memory.storageFraction is the share cache is guaranteed to hold onto. Whatever sits outside the fraction holds your own objects and Spark internals. Execution memory is what shuffles, sorts, joins, and aggregations consume.
Common mistake: Treating spark.executor.memory as the cache budget.
Follow-up: What happens when execution needs memory that cache is currently holding?
Execution and storage share one pool. Who is allowed to evict whom?
Executor memory: execution wins, cache loses · tap to open the answer
Short: Execution can evict cached blocks; storage can never push out execution memory.
Detailed: Either side can borrow free space, but execution wins the tug-of-war and will evict cached partitions down to the storageFraction floor. That is why a DataFrame you cached quietly disappears and the next action rescans the source — the Storage tab shows the cached fraction dropping and the plan goes back to a FileScan.
Common mistake: Assuming cache() guarantees the data stays in memory for the rest of the session.
Follow-up: How would you spot that silent eviction in the UI?
One task spilled 60 GB to disk and finished. A different task's container was killed. Same memory problem?
Executor memory: execution wins, cache loses · tap to open the answer
Short: No — spill is Spark managing execution memory; a killed container is the cluster manager seeing total process memory exceed its limit.
Detailed: Spark tracks execution memory and spills sorted runs to local disk when it runs short, so the task survives slowly and you see Spill (Memory) and Spill (Disk) on the stage. YARN or Kubernetes knows nothing about that accounting — it sees the whole process: JVM heap plus off-heap buffers plus Python workers. Cross the container limit and it is killed with no Spark-level OutOfMemoryError anywhere.
Senior: spark.executor.memoryOverhead — the non-heap headroom for Python workers, off-heap buffers, and shuffle. On PySpark and pandas UDF workloads that is usually what is actually being exceeded.
Common mistake: Raising spark.executor.memory to fix a container kill, which just asks for a limit it still overshoots.
Follow-up: Which setting do people forget in that second case?
Why does Tungsten's binary row format belong in a memory conversation, not just a CPU one?
Executor memory: execution wins, cache loses · tap to open the answer
Short: Binary rows outside the Java object model use far less space per row and stop the GC from walking your data.
Detailed: Tungsten keeps rows in compact buffers with explicit offsets instead of JVM objects, so the same records occupy much less memory and the collector ignores them. Turning on spark.memory.offHeap.enabled with a size moves execution memory off the heap entirely, which keeps GC pauses down — but that memory still counts against the container's total limit.
Senior: Python UDFs — rows leave the binary format and become Python objects, so you pay the conversion and the memory back. That is a memory argument for built-ins, not only a CPU one.
Common mistake: Enabling off-heap memory without raising the container's overall allowance.
Follow-up: What destroys the Tungsten advantage in a PySpark pipeline?