Interview/Spark

Spark interview questions: Spark SQL: the same engine, a different keyboard

Interactive Spark interview questions on Spark SQL: the same engine, a different keyboard. Practice with the matching lesson. SQL strings and the DataFrame API compile to the same plan. Temp views, the catalog, and when names actually resolve.

Lesson · Simulation

Two teammates write the same aggregation — one in spark.sql(), one in the DataFrame API. Which one runs faster?

Answer it out loud, then reveal. Play steps through like the simulators.

Production scenario

Spark SQL: the same engine, a different keyboard

SQL rewrite 'for performance' shipped an identical plan and identical runtime

Symptoms

  • Team converted 300 lines of DataFrame code to spark.sql()
  • Runtime unchanged within noise
  • The ticket promised a 30% gain

All questions on this page

Indexed as FAQ. Open any item if you prefer a list to Play.

beginner

Two teammates write the same aggregation — one in spark.sql(), one in the DataFrame API. Which one runs faster?

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: Neither. Both go through Catalyst and end up as the same physical plan.

Detailed: A SQL string and a chain of DataFrame calls both become an unresolved logical plan, then get analyzed, optimized, and planned by the same code path. Run explain('formatted') on both and you see the same FileScan, Exchange, and HashAggregate nodes.

Common mistake: Claiming the DataFrame API is 'closer to the engine' so it skips a layer that costs runtime.

Follow-up: So what would actually make the two versions differ?

Lesson · Simulation

intermediate

Your notebook created a temp view; the next notebook on the same cluster can't see it. Explain.

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: createOrReplaceTempView is session-scoped, and the other notebook has its own SparkSession.

Detailed: Temp views live in the session catalog and disappear with the session. createGlobalTempView registers into the global_temp database, shared across sessions of the same application, and must be referenced as global_temp.name. Neither is in the metastore — only saveAsTable creates something another job can find tomorrow.

Common mistake: Assuming a temp view is a table because SELECT * worked in the same notebook.

Follow-up: When would you reach for a global temp view instead of just writing a table?

Lesson · Simulation

senior

A managed table and an external table both show up in SHOW TABLES. What changes the moment someone runs DROP TABLE?

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: Dropping a managed table deletes the data files; dropping an external table only removes the metadata.

Detailed: A managed table's location is owned by the catalog, so DROP removes the metastore entry and the underlying directory. An external table keeps its files at the LOCATION you declared — DROP leaves the data and you can re-register it. DESCRIBE EXTENDED shows Type: MANAGED vs EXTERNAL along with the Location.

Senior: DESCRIBE EXTENDED before any DROP in a runbook. Managed is fine for data the pipeline owns; external is for data other systems also write.

Common mistake: Saying DROP TABLE never deletes data because 'Spark doesn't own storage'.

Follow-up: How do you verify which kind you are about to drop?

Lesson · Simulation

architect

A scheduled SQL pipeline fails with an AnalysisException and Spark UI shows zero jobs for that run. Where do you look, and what is it definitely not?

Spark SQL: the same engine, a different keyboard · tap to open the answer

Short: It failed in Catalyst on the driver — name and type resolution — so it is not an executor or data-volume problem.

Detailed: Analysis resolves tables, columns, and functions against the catalog before a single stage is submitted, so a resolution failure means no job, no stage, and nothing in executor logs. spark.sql() builds and analyzes its plan when called, so the guilty statement is the first one referencing the renamed column or dropped table — not the write at the bottom of the notebook.

Senior: File permissions, corrupt Parquet footers, and Delta schema-on-write mismatches — those need a task to run. Resolution failures never get that far.

Common mistake: Digging through executor stderr for a query that never submitted a task.

Follow-up: Which failures do get deferred until execution time?

Lesson · Simulation