fix: drop file tools 1 MB cap + add YarnError -> 502 handler
Two follow-ups to the previous error-handling fix (16fe011): 1) Drop the 1 MB cap on read_job_file and update_job_file. The cap was originally there to keep MCP responses bounded, but it blocks legitimate use of large PySpark scripts and large update payloads. With the new model (LLM composes scripts via write_job_file and edits them via read+update), a fixed 1 MB cap is more hindrance than protection. MCP response size is already bounded by the JSON transport and httpx; the tool itself doesn't need a second limit. Changes: - tools/job_file.py: delete MAX_FILE_BYTES constant, drop the two size checks in read_job_file and update_job_file, update module docstring. - tests/unit/test_job_file.py: delete test_update_rejects_ oversized_content and test_update_rejects_1mb_plus_1_byte (the two tests that asserted the cap), replace with test_update_accepts_content_larger_than_former_1mb_cap. - server.py: drop "Caps reads at 1 MB" and "Caps writes at 1 MB" from the two route descriptions. 2) Add YarnError -> 502 handler. yarn_client wraps every httpx call: on connect / TLS / timeout / 4xx / 5xx / parse failure it raises YarnError. Previously this was unhandled, so all six external job tools (get_external_job_*, list_applications, plus anything else that hits YARN) returned 500 "Internal Server Error" with no detail — the LLM couldn't tell whether the cluster was down or the request was bad. Same fix as16fe011(which did this for fetch_url directly). The new handler returns HTTP 502 Bad Gateway with the YarnError message in the response detail. 502 because the MCP service is acting as a gateway to YARN — 502 is the standard status for "upstream didn't respond correctly". Changes: - server.py: import YarnError, add @app.exception_handler returning 502 + the YarnError message. - tests/integration/test_mcp_routes.py: new test asserts that when get_application_status raises YarnError, /get_job_status returns 502 with the YarnError message in the detail. Note: ValueError (request was bad) is still 400, KeyError (job not in JobStore) is still 404. The three handlers form a clean 3-way classification of tool-layer errors. Tests: 401 passed (was 401, +1 YarnError test, -2 cap tests = net -1). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -6,6 +6,7 @@
|
||||
from fastapi import FastAPI, Request
|
||||
from fastapi.responses import JSONResponse
|
||||
|
||||
from spark_executor.core.yarn_client import YarnError
|
||||
from spark_executor.tools.connections import (
|
||||
delete_connection,
|
||||
get_connection,
|
||||
@@ -83,6 +84,20 @@ async def _valueerror_handler(_request: Request, exc: ValueError) -> JSONRespons
|
||||
return JSONResponse(status_code=400, content={"detail": str(exc)})
|
||||
|
||||
|
||||
@app.exception_handler(YarnError)
|
||||
async def _yarnerror_handler(_request: Request, exc: YarnError) -> JSONResponse:
|
||||
"""YARN RM unreachable or returned an error.
|
||||
|
||||
Surfaces as HTTP 502 Bad Gateway (we are a gateway to YARN). The
|
||||
detail includes whatever yarn_client put in the YarnError message
|
||||
(host:port unreachable, HTTP status from YARN, parse failure,
|
||||
etc). The LLM can act on this to distinguish "YARN is down" from
|
||||
"the request was bad" (the latter would be a 400 from
|
||||
_valueerror_handler instead).
|
||||
"""
|
||||
return JSONResponse(status_code=502, content={"detail": str(exc)})
|
||||
|
||||
|
||||
@app.get("/health")
|
||||
def health_check():
|
||||
return {"status": "ok"}
|
||||
@@ -513,10 +528,9 @@ def _write_job_file(req: WriteJobFileRequest):
|
||||
summary="Read the contents of an existing PySpark script",
|
||||
description=(
|
||||
"Returns the text content of an existing script file at the given "
|
||||
"path. Caps reads at 1 MB. Typical use: after write_job_file "
|
||||
"returns a path, call read_job_file on that path to inspect what "
|
||||
"was actually written, before deciding to prepare_submit_job or "
|
||||
"update_job_file."
|
||||
"path. No size cap. Typical use: after write_job_file returns a "
|
||||
"path, call read_job_file on that path to inspect what was actually "
|
||||
"written, before deciding to prepare_submit_job or update_job_file."
|
||||
),
|
||||
)
|
||||
def _read_job_file(req: ReadJobFileRequest):
|
||||
@@ -531,9 +545,9 @@ def _read_job_file(req: ReadJobFileRequest):
|
||||
"Replaces the entire content of an existing script file. Path must "
|
||||
"be under SPARK_EXECUTOR_JOBS_DIR (the dir write_job_file writes "
|
||||
"to) — protects against overwriting host-mounted configs or other "
|
||||
"non-script files. Caps writes at 1 MB. Typical use: read_job_file, "
|
||||
"edit the content (LLM or human), update_job_file, then "
|
||||
"prepare_submit_job with the same path."
|
||||
"non-script files. No size cap. Typical use: read_job_file, edit "
|
||||
"the content (LLM or human), update_job_file, then prepare_submit_job "
|
||||
"with the same path."
|
||||
),
|
||||
)
|
||||
def _update_job_file(req: UpdateJobFileRequest):
|
||||
|
||||
Reference in New Issue
Block a user