The function had hardcoded --master / --deploy-mode / --queue /
--executor-memory / --executor-cores / --num-executors plus a
`spark_conf` dict for --conf, with no way to pass other flags like
--jars, --py-files, --files, --driver-memory, --name, etc.
Add `extra_args: dict[str, str] | None = None` that emits `--{key}
{value}` pairs, placed after the --conf loop and before the script
path. Default None preserves the existing cmd shape exactly.
Structured (dict) instead of raw list[str] so LLM agents writing
tool-call payloads can't accidentally pack flag+value into a single
string. Key order in the dict is preserved.
Tests: 4 cases — None default, single pair, multiple pairs in order,
no collision with spark_conf. uv run pytest -> 175 passed.
Hardcoding 'spark-submit' as cmd[0] in build_spark_submit_command
breaks for hosts where:
- both Spark 1.x and 2.x/3.x are installed and 'spark-submit' resolves
to the wrong one (use 'spark2-submit' or 'spark3-submit' explicitly)
- the user wants to launch via the PySpark entrypoint ('pyspark')
- a custom wrapper script sits on PATH (e.g. a credentials-injecting
'spark-submit-wrapper')
New env var SPARK_EXECUTOR_SPARK_SUBMIT_BIN. Default is 'spark-submit'
(preserves the current behavior for everyone). Override in .env /
docker-compose.yml to change.
common/config.py:
- new Settings.spark_submit_bin field
- env-var resolution in from_env() with default 'spark-submit'
- included in reload() so tests work
spark_executor/core/spark_submit.py:
- cmd[0] reads settings.spark_submit_bin (was hardcoded 'spark-submit')
.env.example: new section with the override and example values.
docker-compose.yml: forwards the var with the standard 'spark-submit'
default.
Tests: 2 new (settings.spark_submit_bin='spark2-submit', ='pyspark')
plus existing tests updated to use the settings-restore fixture so
mutations don't leak between tests.
165/163 still pass.