Files
mcp-server/CLAUDE.md
T
ClaudeandClaude Fable 5 0fef77c0b8 fix(job_store): persist Jobs to disk and accept either ID in job tools
Two user-reported bugs, same root cause: the in-memory JobStore + the
'job_id must be the 12-char hex' tool contract.

Bug 1: 'Unknown job_id' reported frequently
  JobStore was a process-local dict (spark_executor/core/job_store.py).
  Under gunicorn workers > 1, a job created by confirm_submit_job
  landing on worker A was invisible to worker B, so a follow-up
  get_job_status / get_job_result / get_job_logs / kill_job landing on
  a different worker returned 'Unknown job_id'. Same multi-worker
  problem that bit the MCP session layer; only the affected data was
  different.

Bug 2: 'get_job_logs frequently confuses job_id and application_id'
  confirm_submit_job returns BOTH identifiers in SubmitResult, but
  get_job_logs (and friends) only accepted the local 12-char job_id
  and never said so in their description. When the agent passed the
  YARN application_id, the error message itself was misleading:
  'Unknown job_id: application_17400000001_0001' — the agent had
  passed an id, just the wrong kind.

This change fixes both at the root:

  * JobStore is now JSON-backed at data/jobs.json (atomic tempfile +
    os.replace), with cross-process safety via fcntl.flock on a sibling
    .lock file. Stage 3's SQLite migration is still planned; the file
    format is intentionally simple so it is a straight
    'for j in read_all(): db.insert(j)'.

  * New JobStore.get_either(uid) looks up by job_id first, then
    application_id. All four job-lifecycle tools (get_job_status,
    get_job_result, get_job_logs, kill_job) call get_either instead
    of get(job_id), so the agent can pass either identifier and get
    the same answer.

  * The 'neither matched' KeyError now spells out both id forms and
    what they look like, so the agent isn't left guessing.

  * server.py tool descriptions for the four job tools explicitly
    state 'job_id accepts BOTH identifiers' so this is visible to the
    LLM at tool-selection time, not only at error time.

Tests:
  * test_job_store.py: tmp_path isolation, persistence across
    instances, human-readable JSON, corrupt-file resilience,
    get_either (by job_id, by application_id, collision preference,
    unknown), put idempotency.
  * test_{logs,status,kill,result}_tool.py: per-test tmp_path fixture,
    'accepts application_id' regression for each tool, and an
    explicit assertion that the unknown-id error message mentions
    BOTH id forms. test_result_raises_keyerror_for_unknown_job's
    match pattern updated for the new message.

242 tests pass (was 226; +16 new). Zero regressions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-29 18:57:31 +08:00

8.0 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

spark-executor-mcp is a Python 3.12+ service that exposes Spark-on-YARN operations as MCP tools via fastapi-mcp, so an LLM agent can submit, monitor, fetch logs from, and kill PySpark jobs. Submission is a deliberate two-step flow (prepare_submit_job then confirm_submit_job on second confirmation), with a saved Connection registry (so the agent can pick which YARN/standalone cluster to use).

Common Commands

The project uses uv for dependency management. Dependencies are pinned in pyproject.toml and uv.lock. A virtualenv already exists at .venv/. PyPI index is configured to the Tsinghua mirror in pyproject.toml.

# Install/sync dependencies
uv sync

# Run the server (binds 0.0.0.0:8000; MCP endpoint at /spark-executor-mcp)
uv run main.py

# Run the full test suite
uv run pytest

# Run a single test file
uv run pytest tests/unit/test_submit_tool.py -v

# Run tests matching a name
uv run pytest -v -k cancel_pending

# Activate venv and run directly
source .venv/bin/activate
python main.py

The MCP transport needs an MCP initialize handshake first to get a mcp-session-id header; only then do tools/list / tools/call work.

Architecture

main.py                      # Root FastAPI app; lifespan wires in the MCP server
common/
  factory.py                 # init_mcp_server(app) -> FastApiMCP wrapper
  logging.py                 # loguru: stderr + data/logs/{debug,info}/*.log, rotated 30d
spark_executor/
  __init__.py                # Re-exports `app`
  server.py                  # FastAPI app; 12 tool routes + exception handlers
  models.py                  # Pydantic: Job, JobStatus, SubmitResult, Connection, PendingSubmission
  tools/
    submit.py                # prepare / confirm / list / get / cancel pending + job_store
    status.py logs.py kill.py   # job-lifecycle tools
    connections.py           # save / list / get / delete connection tools
    requests.py              # Pydantic body models for every FastAPI route
  core/
    spark_submit.py          # builds & runs spark-submit commands
    yarn_client.py           # wraps yarn application / yarn logs
    log_parser.py            # extracts application_id from spark-submit output
    job_store.py             # in-memory dict: job_id -> Job (Stage 3 -> SQLite)
    connection_store.py      # JSON CRUD over ./data/connections.json
    pending_store.py         # JSON CRUD over ./data/pending_jobs.json
data/                        # gitignored: connections.json, pending_jobs.json, logs/
tests/                       # unit/ + integration/; conftest adds repo root to sys.path
docs/superpowers/plans/      # implementation plans

How a request flows:

  1. main.py creates the root FastAPI(title="Main App") and registers an asynccontextmanager lifespan.
  2. At startup, the lifespan calls init_mcp_server(spark_executor_app) (common/factory.py) which returns a FastApiMCP instance bound to the spark executor's FastAPI app.
  3. That FastApiMCP is mounted onto the root app at /spark-executor-mcp via mount_http(...). fastapi-mcp inspects the spark executor's routes and registers each one as an MCP tool.
  4. An MCP client calls initialize (gets a mcp-session-id), then tools/list (sees all 12 tools), then tools/call (sends args as a JSON body, FastAPI validates via the Pydantic model in tools/requests.py).

Persistence layout (under ./data/, overridable via SPARK_EXECUTOR_DATA_DIR):

  • data/connections.jsonConnection records keyed by name (atomic write via tempfile + os.replace)
  • data/pending_jobs.jsonPendingSubmission records keyed by pending_id
  • data/jobs.jsonJob records keyed by job_id (atomic write via tempfile + os.replace, cross-process safety via fcntl.flock on a sibling .lock file). The four job-lifecycle tools (get_job_status / get_job_result / get_job_logs / kill_job) accept EITHER the local job_id (12-char hex) OR the YARN application_idJobStore.get_either(...) looks up by job_id first, then by application_id, so an agent that confuses the two (a common LLM mistake) still gets a sensible result.
  • data/logs/debug/YYYY-MM-DD.log — DEBUG sink, gzipped, 30-day retention
  • data/logs/info/YYYY-MM-DD.log — INFO sink, gzipped, 30-day retention

Why JobStore is now file-backed (not in-memory): in-memory only broke under multi-worker gunicorn (workers > 1) for the same reason MCP sessions do — the dict lives in worker A's process, so a follow-up get_job_logs landing on worker B returns "Unknown job_id". File-backed JSON is a Step-1 fix; Stage 3 will still move to SQLite for queryable / transactional semantics, and the file format is intentionally simple so the migration is a straight for j in read_all(): db.insert(j).

Key conventions:

  • All modules include the file header # coding=utf-8 plus a @Time / @Author docstring. Match this when adding new files.
  • common/logging.py configures a single process-wide loguru logger (stderr + two rotated file sinks). Import from common.logging import logger rather than creating new loggers.
  • spark_executor/__init__.py re-exports app so callers can from spark_executor import app — keep this re-export when adding to the package.
  • The FastApiMCP instance is created per-app in the lifespan; do not cache it at module import time.
  • FastAPI routes use Pydantic body models (from tools/requests.py), not query parameters. fastapi-mcp passes tool args as a JSON body, and dict-typed query params arrive as strings and 422. Every new MCP tool needs a request model in tools/requests.py.
  • Exception handlers in server.py: KeyError → 404, ValueError → 400, Pydantic validation → 422. Use the existing handlers — don't add try/except in route bodies.
  • Tool functions raise KeyError for "unknown id" and ValueError for invalid state transitions (e.g. confirming a CANCELLED pending). The handlers translate these to clean HTTP statuses.
  • The two-step submit flow is core, not optional. prepare_submit_job snapshots the connection's master / deploy_mode / spark_conf into the PendingSubmission; confirm_submit_job is the only place spark-submit is invoked. Editing a connection between prepare and confirm does not retarget the pending job.
  • MCP session affinity: keep GUNICORN_WORKERS=1 (or front with a sticky-session LB). The mcp library stores each session in a per-process dict (StreamableHTTPSessionManager._server_instances); gunicorn round-robins requests across workers, so a multi-worker deploy returns "Session not found" / "Invalid or expired session ID" for the same mcp-session-id whenever it lands on a worker that didn't create it. fastapi-mcp hardcodes stateless=False, so there is no in-process workaround. The on_starting hook in gunicorn.conf.py logs a WARNING whenever workers > 1 so the misconfig is loud, not silent. See the comment block above workers = in gunicorn.conf.py for full context and the nginx-sticky escape hatch.
  • Test fixtures rebind module-level singletons (connections.store, submit.conn_store, submit.pending_store) in monkeypatch.setattr because the tool modules captured the originals at import time. See tests/integration/test_mcp_routes.py for the pattern.
  • 86 tests pass as of the last Stage 1 cleanup; run uv run pytest after any change. (242 tests as of the JobStore persistence fix.)

Stage Status

  • Stage 1: complete — 12 MCP tools, 2-step submit, connection management, structured logging, Pydantic body models, exception handlers. 22 commits on feat/stage-1.
  • Stage 2: not startedgenerate_job_file for LLM-written PySpark code (job_writer + one new tool).
  • Stage 3: deferred — async submission, SQLite JobRegistry, status poller, offset-based LogCache, multi-tenant owner. Plan lives at docs/superpowers/plans/2026-06-24-spark-executor-mcp.md.