feat: add read_job_file and update_job_file MCP tools
Closes the review-and-edit loop for LLM-generated PySpark code:
generate_job_file(code=...) -> {script_path}
read_job_file(script_path=...) -> {content, path, size}
update_job_file(path, content) -> {path, bytes_written}
prepare_submit_job(path) -> {pending_id, ...}
Or, in a single edit cycle:
1. generate (LLM writes initial draft)
2. read (LLM or human inspects)
3. update (overwrite with edited version)
4. prepare (submit for two-step confirmation)
Safety:
- read_job_file has no path restriction (read-only; useful for
inspecting any file the agent can see: scripts, logs/, hadoop-conf/)
- update_job_file is sandboxed to settings.jobs_dir (the same dir
generate_job_file writes to). Rejects paths outside that tree,
including ../-traversal attempts. This protects host-mounted
configs (/etc/passwd, hadoop-conf/*) from being overwritten by
the agent.
- 1 MB cap on both reads and writes so MCP responses stay bounded.
Pydantic body models (ReadJobFileRequest, UpdateJobFileRequest) follow
the Stage 1 pattern so tools/call roundtrips long code strings without
the FastAPI query-length 422.
Tests (13 new):
- 9 unit tests: read success/missing/empty/dir, update success/outside/
relative-escape/missing/oversize/1mb+1, full edit cycle round-trip
- 4 integration tests: read via MCP, missing file 400, write+readback
via MCP, outside-jobs_dir rejection via MCP
163/146 still pass. Live verified end-to-end: generate -> read v1
-> update -> read v2; update /etc/passwd correctly 400'd with
'script_path must be under ... data/jobs/'.
This commit is contained in:
@@ -13,6 +13,7 @@ from spark_executor.tools.connections import (
|
||||
save_connection,
|
||||
)
|
||||
from spark_executor.tools.generate import generate_job_file
|
||||
from spark_executor.tools.job_file import read_job_file, update_job_file
|
||||
from spark_executor.tools.kill import kill_job
|
||||
from spark_executor.tools.logs import get_job_logs
|
||||
from spark_executor.tools.requests import (
|
||||
@@ -23,7 +24,9 @@ from spark_executor.tools.requests import (
|
||||
JobIdRequest,
|
||||
PendingIdRequest,
|
||||
PrepareSubmitJobRequest,
|
||||
ReadJobFileRequest,
|
||||
SaveConnectionRequest,
|
||||
UpdateJobFileRequest,
|
||||
)
|
||||
from spark_executor.tools.status import get_job_status
|
||||
from spark_executor.tools.submit import (
|
||||
@@ -226,3 +229,34 @@ def _delete_connection(req: ConnectionNameRequest):
|
||||
)
|
||||
def _generate_job_file(req: GenerateJobFileRequest):
|
||||
return generate_job_file(req.code)
|
||||
|
||||
|
||||
@app.post(
|
||||
"/read_job_file",
|
||||
summary="Read the contents of an existing PySpark script",
|
||||
description=(
|
||||
"Returns the text content of an existing script file at the given "
|
||||
"path. Caps reads at 1 MB. Typical use: after generate_job_file "
|
||||
"returns a path, call read_job_file on that path to inspect what "
|
||||
"was actually written, before deciding to prepare_submit_job or "
|
||||
"update_job_file."
|
||||
),
|
||||
)
|
||||
def _read_job_file(req: ReadJobFileRequest):
|
||||
return read_job_file(req.script_path)
|
||||
|
||||
|
||||
@app.post(
|
||||
"/update_job_file",
|
||||
summary="Overwrite an existing PySpark script with new content",
|
||||
description=(
|
||||
"Replaces the entire content of an existing script file. Path must "
|
||||
"be under SPARK_EXECUTOR_JOBS_DIR (the dir generate_job_file writes "
|
||||
"to) — protects against overwriting host-mounted configs or other "
|
||||
"non-script files. Caps writes at 1 MB. Typical use: read_job_file, "
|
||||
"edit the content (LLM or human), update_job_file, then "
|
||||
"prepare_submit_job with the same path."
|
||||
),
|
||||
)
|
||||
def _update_job_file(req: UpdateJobFileRequest):
|
||||
return update_job_file(req.script_path, req.content)
|
||||
|
||||
Reference in New Issue
Block a user