The MCP tool was named `generate_job_file` from Stage 2 but it does
NOT generate PySpark code — the calling LLM writes the code in its own
context, and this tool only persists it to a file under
SPARK_EXECUTOR_JOBS_DIR so `spark-submit` can see it. The misleading
`generate_` prefix sent agents (and humans) looking for a code
generator that doesn't exist.
This commit folds three related polish changes into one (split later
with rebase -i if you want them as separate history):
1. The rename itself:
- `tools/generate.py` → `tools/write_job.py`
- `generate_job_file` → `write_job_file`
- `GenerateJobFileRequest` → `WriteJobFileRequest`
- `/generate_job_file` route → `/write_job_file`
- `operation_id="generate_job_file"` → `operation_id="write_job_file"`
The internal helper `core.job_writer.write_job_file` (which just
writes bytes to disk with no SQL guard) is imported with an
`_write_to_disk` alias to avoid the name collision with the
MCP-exposed function in the same module.
The description for the tool now explicitly states 'this tool
does NOT generate PySpark code. The calling LLM is expected to
have already written the code; this tool only persists it.'
2. Skill for LLM agents operating the service
(`docs/superpowers/skills/spark-executor-mcp-operate/SKILL.md`,
449 lines). Covers the 16 tools, the two-step prepare/confirm
flow, the dual-ID contract (job_id vs application_id), the
PendingSubmission state machine, the Connection profile, the
job-file workflow, the error reference, common pitfalls, and a
full end-to-end word-count example.
3. Default `executor_memory` lowered 4G → 2G
(`_DEFAULTS_TO_CONFIRM` in `server.py`). Mirrors the matching
change in `test_mcp_routes.py` and the 5 unit tests that
reference the default. Aligns with the lighter workloads the
service is sized for in its current container profile.
Also tracked in git for the first time:
- `docs/superpowers/plans/2026-06-24-spark-executor-mcp.md`
(the original Stage 1/2/3 design plan, updated to use the new
tool name throughout).
Test rename:
- `tests/unit/test_generate_tool.py` → `test_write_job_tool.py`
- the new test file picks up an extra assertion that the SQL guard
rejects a `DROP TABLE` statement at write time.
243 tests pass (was 242; +1 new SQL-guard assertion). Zero regressions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
112 lines
4.1 KiB
Python
112 lines
4.1 KiB
Python
# coding=utf-8
|
|
"""
|
|
@Time :2026/6/24
|
|
@Author :tao.chen
|
|
|
|
Read + update the contents of an existing PySpark script file. These two
|
|
tools close the review-and-edit loop:
|
|
|
|
write_job_file(code=...) -> {script_path}
|
|
read_job_file(script_path=...) -> {content, path} <-- inspect
|
|
update_job_file(path, content) -> {path, bytes_written} <-- edit
|
|
prepare_submit_job(path) -> {pending_id, ...}
|
|
|
|
Safety:
|
|
- read_job_file: any existing regular file. Path-existence only.
|
|
- update_job_file: must be under SPARK_EXECUTOR_JOBS_DIR
|
|
(settings.jobs_dir) so the agent cannot overwrite host-mounted
|
|
configs or arbitrary files on the container FS.
|
|
- 1 MB cap on both read and write payloads to keep MCP responses bounded.
|
|
"""
|
|
import os
|
|
from pathlib import Path
|
|
|
|
from common.config import settings
|
|
from common.logging import logger
|
|
|
|
MAX_FILE_BYTES = 1 * 1024 * 1024 # 1 MB
|
|
|
|
|
|
class ScriptFileError(ValueError):
|
|
"""Raised when read/update fails. -> HTTP 400 via the FastAPI ValueError
|
|
handler in server.py.
|
|
"""
|
|
pass
|
|
|
|
|
|
def _check_readable(script_path: str) -> None:
|
|
if not script_path or not os.path.isfile(script_path):
|
|
raise ScriptFileError(
|
|
f"script_path does not exist or is not a file: {script_path!r}"
|
|
)
|
|
|
|
|
|
def _check_writable(script_path: str) -> None:
|
|
"""update_job_file is restricted to files under settings.jobs_dir
|
|
(the same dir write_job_file writes to). This prevents the agent
|
|
from overwriting arbitrary host-mounted files or the app's own code.
|
|
"""
|
|
if not script_path or not os.path.isfile(script_path):
|
|
raise ScriptFileError(
|
|
f"script_path does not exist or is not a file: {script_path!r}. "
|
|
f"update_job_file can only edit existing files. "
|
|
f"Use write_job_file to create a new one."
|
|
)
|
|
jobs_root = Path(settings.jobs_dir).resolve()
|
|
target = Path(script_path).resolve()
|
|
try:
|
|
target.relative_to(jobs_root)
|
|
except ValueError:
|
|
raise ScriptFileError(
|
|
f"script_path must be under {jobs_root} (the directory "
|
|
f"write_job_file writes to). Got {script_path!r}. "
|
|
f"This restriction protects host-mounted configs and other "
|
|
f"non-script files from being overwritten by the agent."
|
|
)
|
|
|
|
|
|
def read_job_file(script_path: str) -> dict[str, object]:
|
|
"""Return the text content of an existing script file.
|
|
|
|
Caps the read at 1 MB to keep MCP responses bounded; raises
|
|
ScriptFileError (-> 400) if the file is missing or too large.
|
|
"""
|
|
_check_readable(script_path)
|
|
size = os.path.getsize(script_path)
|
|
if size > MAX_FILE_BYTES:
|
|
raise ScriptFileError(
|
|
f"Script is too large to read back ({size} bytes > {MAX_FILE_BYTES} "
|
|
f"byte cap). Edit it via a host volume mount instead."
|
|
)
|
|
logger.debug(f"read_job_file enter script_path={script_path} size={size}")
|
|
with open(script_path, encoding="utf-8") as f:
|
|
content = f.read()
|
|
logger.info(f"read_job_file ok script_path={script_path} size={size}")
|
|
return {"path": script_path, "content": content, "size": size}
|
|
|
|
|
|
def update_job_file(script_path: str, content: str) -> dict[str, object]:
|
|
"""Overwrite an existing script file with new content.
|
|
|
|
Restricted to paths under settings.jobs_dir. Caps writes at 1 MB.
|
|
Raises ScriptFileError (-> 400) if the path is missing, outside
|
|
the allowed dir, or the content is too large.
|
|
"""
|
|
_check_writable(script_path)
|
|
encoded_size = len(content.encode("utf-8"))
|
|
if encoded_size > MAX_FILE_BYTES:
|
|
raise ScriptFileError(
|
|
f"content is too large ({encoded_size} bytes > {MAX_FILE_BYTES} "
|
|
f"byte cap). Split the script into multiple files."
|
|
)
|
|
logger.debug(
|
|
f"update_job_file enter script_path={script_path} "
|
|
f"new_bytes={encoded_size}"
|
|
)
|
|
with open(script_path, "w", encoding="utf-8") as f:
|
|
written = f.write(content)
|
|
logger.info(
|
|
f"update_job_file ok script_path={script_path} bytes_written={written}"
|
|
)
|
|
return {"path": script_path, "bytes_written": written}
|