feat(result): get_job_result tool for terminal view of Spark job
`get_job_status` returns YARN state + the raw response blob, so the terminal fields (finalStatus, diagnostics, trackingUrl, startedTime, finishedTime) are buried inside `raw` and not surfaced in a structured form. Add a new tool that parses them. `get_job_result(job_id)` reuses `yarn_client.get_application_status` and extracts: - finalStatus (SUCCEEDED / FAILED / KILLED / UNDEFINED) - diagnostics (YARN final message) - tracking_url (Spark Web UI) - started_time / finished_time (epoch ms) All five fields are optional: running jobs have no `finishedTime`, and older YARN versions (CDH 5 / H2) may omit some fields. Missing fields stay None — never raise. Coexistence with `get_job_status` is intentional: the latter is for polling the running YARN state, the former is the terminal view. Tests: 4 cases — happy path, running job (no finishedTime), bare-minimum raw (all optionals None), unknown job_id raises KeyError. uv run pytest -> 171 passed.
This commit is contained in:
@@ -24,6 +24,16 @@ class JobStatus(BaseModel):
|
||||
raw: str = Field(default="")
|
||||
|
||||
|
||||
class JobResult(BaseModel):
|
||||
application_id: str
|
||||
state: str
|
||||
final_status: str | None = None
|
||||
diagnostics: str | None = None
|
||||
tracking_url: str | None = None
|
||||
started_time: int | None = None
|
||||
finished_time: int | None = None
|
||||
|
||||
|
||||
class SubmitResult(BaseModel):
|
||||
job_id: str
|
||||
application_id: str
|
||||
|
||||
Reference in New Issue
Block a user