feat: env-configurable spark-submit binary name

Hardcoding 'spark-submit' as cmd[0] in build_spark_submit_command
breaks for hosts where:
  - both Spark 1.x and 2.x/3.x are installed and 'spark-submit' resolves
    to the wrong one (use 'spark2-submit' or 'spark3-submit' explicitly)
  - the user wants to launch via the PySpark entrypoint ('pyspark')
  - a custom wrapper script sits on PATH (e.g. a credentials-injecting
    'spark-submit-wrapper')

New env var SPARK_EXECUTOR_SPARK_SUBMIT_BIN. Default is 'spark-submit'
(preserves the current behavior for everyone). Override in .env /
docker-compose.yml to change.

common/config.py:
  - new Settings.spark_submit_bin field
  - env-var resolution in from_env() with default 'spark-submit'
  - included in reload() so tests work

spark_executor/core/spark_submit.py:
  - cmd[0] reads settings.spark_submit_bin (was hardcoded 'spark-submit')

.env.example: new section with the override and example values.
docker-compose.yml: forwards the var with the standard 'spark-submit'
default.

Tests: 2 new (settings.spark_submit_bin='spark2-submit', ='pyspark')
plus existing tests updated to use the settings-restore fixture so
mutations don't leak between tests.

165/163 still pass.
This commit is contained in:
Claude
2026-06-25 16:46:59 +08:00
parent cb909f7fea
commit 1c9e4a321d
5 changed files with 87 additions and 2 deletions
+4
View File
@@ -52,6 +52,10 @@ services:
# Loguru verbosity for stderr + info file. DEBUG | INFO.
SPARK_EXECUTOR_LOG_LEVEL: ${SPARK_EXECUTOR_LOG_LEVEL:-DEBUG}
# Spark CLI binary name. 'spark-submit' by default; override to
# 'spark2-submit' on mixed-version hosts, or to a wrapper path.
SPARK_EXECUTOR_SPARK_SUBMIT_BIN: ${SPARK_EXECUTOR_SPARK_SUBMIT_BIN:-spark-submit}
# --- Runtime paths (consumed by docker-entrypoint.sh, not common/config.py) ---
# Override to point at a different JDK install or pre-mounted Spark
# distribution. The entrypoint re-derives PATH from these at every