Harden confirm_submit_job for unreliable test environments and transient
spark-submit failures:
- Add confirm_max_retries and confirm_retry_delay_seconds settings
(env-configurable).
- SUBMITTED pending returns cached result without re-running spark-submit.
- FAILED pending is reset to PENDING and retried.
- CANCELLED pending is still rejected.
- On success, persist pending as SUBMITTED before creating the in-memory Job
so a job_store failure cannot leave the record as PENDING while the YARN
app is already running.
- Persist tracking_url on PendingSubmission for idempotent returns.
Tests cover retry-then-success, max-retries-exceeded, idempotency,
FAILED reset, and pending-saved-before-job-store.
Remove defaults from queue, executor_memory, executor_cores,
num_executors, and app_name in prepare_submit_job. All must now be
explicitly confirmed by the caller.
Also add extra_args to the PendingSubmission snapshot so users can
confirm non-conf spark-submit flags (e.g. --jars, --py-files) at
prepare time; confirm_submit_job passes them through to
build_spark_submit_command.
This is an intentional breaking change to the MCP tool contract:
callers can no longer rely on implicit defaults.
- Add optional app_name to PendingSubmission and prepare_submit_job,
passed through PrepareSubmitJobRequest for user-provided tracking.
- Rewrite PendingStore to shard records by UTC date under
data/pending_jobs/YYYY-MM-DD.json instead of a single monolithic
pending_jobs.json.
- Legacy single-file pending_jobs.json remains readable; first save()
migrates it and renames the old file to avoid double reads.
- Tests cover date sharding, multi-date list, legacy read, legacy
migration, and app_name round-trip.