Spark-submit can fail AFTER YARN has already accepted the application
(lost RM connection, PySpark script exited non-zero on the cluster
side, auth expired mid-submit, ...). In those cases the stderr still
contains 'Submitted application <id>', but the previous
confirm_submit_job implementation threw away that signal and recorded
the pending as FAILED with no application_id. The user had no way to
fetch the YARN logs of the failed job.
Fix:
- run_spark_submit now attaches the failed CompletedProcess to the
SparkSubmitError as .result, so callers can still inspect stderr.
- confirm_submit_job's failure branch tries to parse
(application_id, tracking_url) from exc.result.stderr. If found,
it writes them onto the pending record before re-raising. The
status is still FAILED and the error message is still set — the
user sees the failure normally, but the pending now also has a
log target they can pass to get_application_logs.
- If the stderr has no application_id (e.g. spark-submit failed
locally before reaching YARN), the parse attempt returns None and
the pending is FAILED with no log target — exactly the previous
behavior in that case.
Tests:
- test_confirm_marks_failed_on_spark_submit_error: existing test
now also asserts application_id is None (no result attached).
- test_confirm_recovers_application_id_from_failed_spark_submit:
failed run with 'Submitted application X' in stderr → pending is
FAILED + application_id is set.
- test_confirm_failed_spark_submit_without_application_id_keeps_it_none:
failed run with no application_id in stderr → pending is FAILED
+ application_id is None.
Full suite 248 passed (+2 from 246).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>