spark/internals
advanced
Connecting…

Tungsten execution engine

Tungsten stores rows off-heap and generates JVM bytecode for whole stages so Spark SQL is not a naive iterator of JVM objects.

Lesson 22 of 29 · Spark path

Explain it at my level

  1. Binary rows
  2. Codegen
  3. UDF hole
Watch the canvas:Tungsten / codegenPython UDF fallbackLive simulation
Unsafe / binary rowoff-heap bytesWholeStageCodegenfused loopPythonUDFobjects againNot Java objects on the hot path.
Tungsten = binary rows + codegen. UDFs throw you back to the slow path.