Files
mcp-server/docs/superpowers/plans/2026-06-30-files-mcp.md
ClaudeandClaude Fable 5 d985883c20 docs(files-mcp): implementation plan with TDD steps and 5 tasks
Adds docs/superpowers/plans/2026-06-30-files-mcp.md — 5-task
implementation plan covering scaffold+path_guard, models+fs_ops,
business wrappers+request models, server+routes+integration tests,
and main.py+conftest wiring. 77 new tests across 3 unit files and
1 integration file; verified acceptance includes the spark_executor
test suite (242 existing) still passing unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-30 17:28:50 +08:00

62 KiB
Raw Permalink Blame History

Files MCP Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Add a sibling MCP service at /files-mcp that exposes 9 sandboxed file-system operations (create/read/update/delete/list_dir/stat/search/move/copy) over the existing fastapi-mcp transport, utf-8 text only, with atomic writes and a configurable sandbox root.

Architecture: New sibling Python package files_mcp/ (parallel to spark_executor/) is mounted under the root FastAPI app via the existing MCP_SERVICES registry in main.py — no edits to the lifespan function. The package has three layers: core/path_guard.py (sandbox + Path.resolve()), core/fs_ops.py (pure file operations on already-resolved Paths), and tools/files.py (business wrappers that call path_guard then fs_ops). server.py wires the 9 tool routes + 3 exception handlers.

Tech Stack: Python 3.12+, FastAPI 0.138+, fastapi-mcp 0.4+, Pydantic 2.13+, loguru (already in pyproject.toml); stdlib pathlib / shutil / tempfile / os / os.replace. No new dependencies.

Spec: docs/superpowers/specs/2026-06-30-files-mcp-design.md (commit 254cbe9 on feat/files-mcp).


Global Constraints

  • Python 3.12+ (per pyproject.toml).
  • No new dependencies — all operations use stdlib.
  • FILES_MCP_ROOT env var is required at process start. The package's __init__ calls _init_root() which raises RuntimeError with a clear message if it is unset, missing, or not a directory.
  • utf-8 text only. Non-utf-8 bytes on read → ValueError (HTTP 400), mirroring the existing job_writer "reject code that is not valid UTF-8" rule (commit f089793).
  • 1 MB content cap on read and write. Mirrors read_job_file / update_job_file.
  • Atomic writes via tempfile.mkstemp + os.replace.
  • Sandbox is one .resolve() + relative_to check. Symlinks, .., ~, double slashes all handled by .resolve(); no whitelist to keep in sync.
  • Strict create/update semantics: create_file requires the file to be absent (unless overwrite=True); update_file requires it to be present.
  • Module-level file header: Every new Python file starts with # coding=utf-8 and the standard @Time / @Author docstring.
  • fastapi-mcp tool args go through Pydantic body models (from files_mcp/tools/requests.py), not query params. Each tool has a dedicated request model even when it has only one field.
  • Existing tools / tests / dependencies are not modified, except:
    • main.py gets one new import + one new McpService entry.
    • tests/conftest.py gets one new autouse fixture.
  • Branch: feat/files-mcp. All commits land on this branch.
  • All work uses uv run for environment consistency (the repo has a .venv already).
  • Conventions reused from spark_executor/: from common.logging import logger (not a new logger); @app.exception_handler(...) pattern in server.py; operation_id / summary / description on every route; model_dump() for serialization.

File Structure

Files this plan creates:

files_mcp/
  __init__.py                       # triggers _init_root(); re-exports app
  server.py                         # FastAPI("Files MCP") + 9 routes + 3 handlers
  models.py                         # FileEntry, OperationResult
  core/
    __init__.py
    path_guard.py                   # resolve_and_check, _init_root, set_root_for_testing
    fs_ops.py                       # 9 pure file ops
  tools/
    __init__.py
    files.py                        # 9 business wrappers
    requests.py                     # 9 Pydantic body models

tests/
  unit/
    test_path_guard.py
    test_fs_ops.py
    test_files_tool.py
  integration/
    test_files_mcp_routes.py

Files this plan modifies:

main.py                             # +1 import, +1 McpService entry
tests/conftest.py                   # +1 autouse fixture

Task 1: Scaffolding + path_guard

Files:

  • Create: files_mcp/__init__.py
  • Create: files_mcp/core/__init__.py
  • Create: files_mcp/core/path_guard.py
  • Create: tests/unit/test_path_guard.py

Interfaces (this task produces; later tasks consume):

  • files_mcp.core.path_guard.resolve_and_check(path: str) -> Path — raises ValueError on empty, NUL, or escape.

  • files_mcp.core.path_guard.set_root_for_testing(p: Path) -> None — test-only.

  • files_mcp.core.path_guard._init_root() -> Path — called once at import; reads FILES_MCP_ROOT env var.

  • Step 1.1: Create files_mcp/core/__init__.py

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
  • Step 1.2: Create files_mcp/core/path_guard.py
# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen

Sandbox: every user-supplied path is resolve()d and checked against a
single frozen root directory. resolve() follows symlinks and normalizes
.., ~, and duplicate slashes, so a single relative_to() check is enough.
"""
import os
from pathlib import Path

from common.logging import logger

_ROOT: Path | None = None


def _init_root() -> Path:
    """Resolve and freeze the sandbox root from FILES_MCP_ROOT.

    Called once at package import time. Raises RuntimeError if the env
    var is unset, points at a missing path, or points at a non-directory.
    Idempotent: subsequent calls return the cached root.
    """
    global _ROOT
    if _ROOT is not None:
        return _ROOT
    raw = os.environ.get("FILES_MCP_ROOT")
    if not raw:
        raise RuntimeError(
            "FILES_MCP_ROOT env var is required (e.g. export "
            "FILES_MCP_ROOT=/srv/files-mcp-sandbox)"
        )
    p = Path(raw).expanduser().resolve()
    if not p.is_dir():
        raise RuntimeError(
            f"FILES_MCP_ROOT={raw!r} resolves to {p} which is not an existing directory"
        )
    _ROOT = p
    logger.info(f"Files MCP sandbox root: {_ROOT}")
    return _ROOT


def resolve_and_check(path: str) -> Path:
    """Resolve an arbitrary input path and verify it lies inside the sandbox.

    Raises ValueError if the path is empty, contains NUL, or escapes the
    root after symlink/.. resolution.
    """
    if _ROOT is None:
        _init_root()
    assert _ROOT is not None
    if not path:
        raise ValueError("path is empty")
    if "\x00" in path:
        raise ValueError("path contains NUL byte")
    p = Path(path).expanduser().resolve()
    try:
        p.relative_to(_ROOT)
    except ValueError:
        raise ValueError(
            f"path escapes sandbox root {_ROOT}: got {p}"
        ) from None
    return p


def set_root_for_testing(p: Path) -> None:
    """Test-only escape hatch. Never call from production code."""
    global _ROOT
    _ROOT = p.resolve()
  • Step 1.3: Create files_mcp/__init__.py
# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from pathlib import Path

from files_mcp.core import path_guard

# Fail fast at process start if FILES_MCP_ROOT is unset / bad.
path_guard._init_root()

# Re-export the FastAPI app once the sandbox root has been initialized.
from files_mcp.server import app  # noqa: E402

ROOT: Path = path_guard._ROOT  # type: ignore[attr-defined]

__all__ = ["app", "ROOT"]
  • Step 1.4: Write the failing test

Create tests/unit/test_path_guard.py:

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from pathlib import Path

import pytest

from files_mcp.core import path_guard


@pytest.fixture
def sandbox_root(tmp_path, monkeypatch):
    """Fresh tmp_path-based sandbox root for each test.

    Pre-populates a regular file, a nested file inside a subdir, and a
    symlink that points outside the sandbox (for the symlink-escape test).
    """
    monkeypatch.setenv("FILES_MCP_ROOT", str(tmp_path))
    path_guard.set_root_for_testing(tmp_path.resolve())
    (tmp_path / "inside.txt").write_text("hi", encoding="utf-8")
    (tmp_path / "subdir").mkdir()
    (tmp_path / "subdir" / "nested.txt").write_text("nested", encoding="utf-8")
    (tmp_path / "escape_link").symlink_to("/etc/passwd")
    return tmp_path.resolve()


class TestResolveAndCheck:
    def test_accepts_inside_path(self, sandbox_root):
        p = path_guard.resolve_and_check("inside.txt")
        assert p == sandbox_root / "inside.txt"

    def test_accepts_nested_path(self, sandbox_root):
        p = path_guard.resolve_and_check("subdir/nested.txt")
        assert p == sandbox_root / "subdir" / "nested.txt"

    def test_accepts_dot_segments(self, sandbox_root):
        p = path_guard.resolve_and_check("subdir/../inside.txt")
        assert p == sandbox_root / "inside.txt"

    def test_rejects_double_dot_escape(self, sandbox_root):
        with pytest.raises(ValueError, match="escapes sandbox"):
            path_guard.resolve_and_check("../etc/passwd")

    def test_rejects_absolute_outside_path(self, sandbox_root):
        with pytest.raises(ValueError, match="escapes sandbox"):
            path_guard.resolve_and_check("/etc/passwd")

    def test_rejects_mixed_escape(self, sandbox_root):
        with pytest.raises(ValueError, match="escapes sandbox"):
            path_guard.resolve_and_check("subdir/../../etc/passwd")

    def test_rejects_symlink_escape(self, sandbox_root):
        # escape_link -> /etc/passwd; resolve() follows the link first.
        with pytest.raises(ValueError, match="escapes sandbox"):
            path_guard.resolve_and_check("escape_link")

    def test_rejects_empty(self, sandbox_root):
        with pytest.raises(ValueError, match="path is empty"):
            path_guard.resolve_and_check("")

    def test_rejects_nul_byte(self, sandbox_root):
        with pytest.raises(ValueError, match="NUL byte"):
            path_guard.resolve_and_check("foo\x00bar")


class TestInitRoot:
    def test_init_root_raises_if_unset(self, monkeypatch):
        monkeypatch.delenv("FILES_MCP_ROOT", raising=False)
        path_guard._ROOT = None
        with pytest.raises(RuntimeError, match="FILES_MCP_ROOT"):
            path_guard._init_root()

    def test_init_root_raises_if_not_directory(self, tmp_path, monkeypatch):
        monkeypatch.setenv("FILES_MCP_ROOT", str(tmp_path / "missing"))
        path_guard._ROOT = None
        with pytest.raises(RuntimeError, match="not an existing directory"):
            path_guard._init_root()

    def test_init_root_is_idempotent(self, sandbox_root):
        first = path_guard._ROOT
        path_guard._init_root()
        assert path_guard._ROOT == first
  • Step 1.5: Run the test to verify it fails (file does not exist yet)

Run: uv run pytest tests/unit/test_path_guard.py -v Expected: FAIL with ModuleNotFoundError: No module named 'files_mcp'

(If the test runs green, the module already exists; verify Steps 1.11.3 actually wrote the files. The test was written before the implementation on purpose — but since we are creating the files in lockstep, run it once now to confirm the error is the expected one.)

  • Step 1.6: Run the test to verify it passes

Run: uv run pytest tests/unit/test_path_guard.py -v Expected: 12 PASS (9 in TestResolveAndCheck + 3 in TestInitRoot)

  • Step 1.7: Commit
git add files_mcp/__init__.py files_mcp/core/__init__.py files_mcp/core/path_guard.py tests/unit/test_path_guard.py
git commit -m "feat(files-mcp): scaffold package + path_guard sandbox

Adds the files_mcp/ sibling package and core/path_guard.py which
resolves arbitrary input paths and rejects any that escape the
FILES_MCP_ROOT sandbox (handles .., ~, symlinks, NUL, empty). Adds
tests/unit/test_path_guard.py covering accept, escape modes,
symlink escape, and _init_root failure paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"

Task 2: Models + fs_ops

Files:

  • Create: files_mcp/models.py
  • Create: files_mcp/core/fs_ops.py
  • Create: tests/unit/test_fs_ops.py

Interfaces (this task produces; later tasks consume):

  • files_mcp.models.FileEntry — Pydantic model for list_dir / stat output.
  • files_mcp.models.OperationResult — Pydantic model for create / update / delete / move / copy output.
  • files_mcp.core.fs_ops.MAX_BYTES = 1_048_576
  • files_mcp.core.fs_ops.create(path: Path, content: str, overwrite: bool) -> OperationResult
  • files_mcp.core.fs_ops.read(path: Path) -> dict[str, object] (returns {path, content, size, encoding})
  • files_mcp.core.fs_ops.update(path: Path, content: str) -> OperationResult
  • files_mcp.core.fs_ops.delete(path: Path, recursive: bool) -> OperationResult
  • files_mcp.core.fs_ops.list_dir(path: Path, recursive: bool, max_depth: int) -> list[FileEntry]
  • files_mcp.core.fs_ops.stat(path: Path) -> FileEntry
  • files_mcp.core.fs_ops.search(path: Path, pattern: str, recursive: bool) -> list[str]
  • files_mcp.core.fs_ops.move(src: Path, dst: Path, overwrite: bool) -> OperationResult
  • files_mcp.core.fs_ops.copy(src: Path, dst: Path, overwrite: bool) -> OperationResult

Error contract (consumed by server.py in Task 4):

  • create / move / copy with overwrite=False and existing dst → FileExistsError

  • update / read / delete / stat on missing path → KeyError

  • read on a directory path → ValueError("path 'X' is not a regular file")

  • read on bytes that aren't valid utf-8 → ValueError("not valid utf-8")

  • read / create / update if file > 1 MB → ValueError("content too large")

  • delete on a non-empty directory without recursive=TrueValueError

  • list_dir on a non-directory → ValueError("path 'X' is not a directory")

  • Step 2.1: Create files_mcp/models.py

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from typing import Literal

from pydantic import BaseModel


class FileEntry(BaseModel):
    path: str            # absolute, resolved, inside sandbox
    name: str
    type: Literal["file", "directory", "symlink", "other"]
    size: int            # bytes; 0 for directories
    mtime: float         # unix seconds
    mode: int            # st_mode (octal, e.g. 0o644)


class OperationResult(BaseModel):
    path: str
    status: str          # "CREATED" | "UPDATED" | "DELETED" | "MOVED" | "COPIED"
    size: int | None     # None for delete/move
  • Step 2.2: Create files_mcp/core/fs_ops.py
# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen

Pure file-system operations on already-resolved Path objects. Callers
are responsible for sandbox resolution (see path_guard.resolve_and_check).
No global state, no logging — kept trivially testable.
"""
import os
import shutil
import tempfile
from pathlib import Path

from files_mcp.models import FileEntry, OperationResult

MAX_BYTES = 1_048_576  # 1 MB cap, mirrors read_job_file / update_job_file


def _to_entry(p: Path) -> FileEntry:
    st = p.lstat()  # lstat: do not follow symlink for type detection
    if p.is_symlink():
        kind = "symlink"
    elif p.is_dir():
        kind = "directory"
    elif p.is_file():
        kind = "file"
    else:
        kind = "other"
    return FileEntry(
        path=str(p),
        name=p.name,
        type=kind,
        size=st.st_size if kind == "file" else 0,
        mtime=st.st_mtime,
        mode=st.st_mode,
    )


def _check_size(content: str) -> None:
    if len(content.encode("utf-8")) > MAX_BYTES:
        raise ValueError(
            f"content too large: {len(content)} bytes > {MAX_BYTES}"
        )


def _atomic_write(path: Path, content: str) -> None:
    fd, tmp = tempfile.mkstemp(dir=path.parent, prefix=".tmp_", suffix=path.suffix)
    try:
        with os.fdopen(fd, "w", encoding="utf-8") as f:
            f.write(content)
        os.replace(tmp, path)
    except BaseException:
        Path(tmp).unlink(missing_ok=True)
        raise


def create(path: Path, content: str, overwrite: bool) -> OperationResult:
    if path.exists() and not overwrite:
        raise FileExistsError(f"file already exists: {path}")
    _check_size(content)
    _atomic_write(path, content)
    return OperationResult(path=str(path), status="CREATED", size=path.stat().st_size)


def read(path: Path) -> dict:
    if not path.exists():
        raise KeyError(f"path not found: {path}")
    if not path.is_file():
        raise ValueError(f"path {path!r} is not a regular file")
    raw = path.read_bytes()
    if len(raw) > MAX_BYTES:
        raise ValueError(f"file too large: {len(raw)} bytes > {MAX_BYTES}")
    try:
        text = raw.decode("utf-8")
    except UnicodeDecodeError as e:
        raise ValueError(f"not valid utf-8: {path}") from e
    return {"path": str(path), "content": text, "size": len(raw), "encoding": "utf-8"}


def update(path: Path, content: str) -> OperationResult:
    if not path.exists():
        raise KeyError(f"path not found: {path}")
    _check_size(content)
    _atomic_write(path, content)
    return OperationResult(path=str(path), status="UPDATED", size=path.stat().st_size)


def delete(path: Path, recursive: bool) -> OperationResult:
    if not path.exists():
        raise KeyError(f"path not found: {path}")
    if path.is_dir() and not recursive:
        # refuse to delete non-empty dirs unless recursive
        if any(path.iterdir()):
            raise ValueError(
                f"directory {path!r} is not empty (use recursive=True)"
            )
    if path.is_dir() and recursive:
        shutil.rmtree(path)
    else:
        path.unlink()
    return OperationResult(path=str(path), status="DELETED", size=None)


def list_dir(path: Path, recursive: bool, max_depth: int) -> list[FileEntry]:
    if not path.exists():
        raise KeyError(f"path not found: {path}")
    if not path.is_dir():
        raise ValueError(f"path {path!r} is not a directory")
    if max_depth < 0:
        raise ValueError("max_depth must be >= 0")
    out: list[FileEntry] = []
    if not recursive:
        max_depth = 0
    for cur, dirs, files in os.walk(path):
        rel = Path(cur).relative_to(path)
        depth = 0 if str(rel) == "." else len(rel.parts)
        if depth > max_depth:
            dirs[:] = []  # don't descend further
            continue
        for d in dirs:
            out.append(_to_entry(Path(cur) / d))
        for f in files:
            out.append(_to_entry(Path(cur) / f))
    return out


def stat(path: Path) -> FileEntry:
    if not path.exists() and not path.is_symlink():
        raise KeyError(f"path not found: {path}")
    return _to_entry(path)


def search(path: Path, pattern: str, recursive: bool) -> list[str]:
    if not path.exists():
        raise KeyError(f"path not found: {path}")
    if not path.is_dir():
        raise ValueError(f"path {path!r} is not a directory")
    glob_fn = path.rglob if recursive else path.glob
    return [str(p) for p in glob_fn(pattern)]


def move(src: Path, dst: Path, overwrite: bool) -> OperationResult:
    if not src.exists():
        raise KeyError(f"src not found: {src}")
    if dst.exists() and not overwrite:
        raise FileExistsError(f"dst already exists: {dst}")
    if dst.exists() and overwrite:
        if dst.is_dir() and not dst.is_symlink():
            shutil.rmtree(dst)
        else:
            dst.unlink()
    os.replace(src, dst)
    size = dst.stat().st_size if dst.is_file() else None
    return OperationResult(path=str(dst), status="MOVED", size=size)


def copy(src: Path, dst: Path, overwrite: bool) -> OperationResult:
    if not src.exists():
        raise KeyError(f"src not found: {src}")
    if dst.exists() and not overwrite:
        raise FileExistsError(f"dst already exists: {dst}")
    if src.is_dir():
        if dst.exists() and overwrite:
            shutil.rmtree(dst)
        shutil.copytree(src, dst)
        size = None
    else:
        if dst.exists() and overwrite:
            dst.unlink()
        shutil.copy2(src, dst)
        size = dst.stat().st_size
    return OperationResult(path=str(dst), status="COPIED", size=size)
  • Step 2.3: Write the failing test

Create tests/unit/test_fs_ops.py:

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from pathlib import Path

import pytest

from files_mcp.core import fs_ops
from files_mcp.models import FileEntry, OperationResult


# --- create ------------------------------------------------------------

class TestCreate:
    def test_creates_new_file(self, tmp_path: Path):
        target = tmp_path / "new.txt"
        result = fs_ops.create(target, "hello", overwrite=False)
        assert isinstance(result, OperationResult)
        assert result.status == "CREATED"
        assert result.size == 5
        assert target.read_text(encoding="utf-8") == "hello"

    def test_overwrite_true_replaces(self, tmp_path: Path):
        target = tmp_path / "f.txt"
        target.write_text("old", encoding="utf-8")
        result = fs_ops.create(target, "new", overwrite=True)
        assert result.status == "CREATED"
        assert target.read_text(encoding="utf-8") == "new"

    def test_overwrite_false_raises(self, tmp_path: Path):
        target = tmp_path / "f.txt"
        target.write_text("old", encoding="utf-8")
        with pytest.raises(FileExistsError, match="already exists"):
            fs_ops.create(target, "new", overwrite=False)

    def test_too_large_raises(self, tmp_path: Path):
        target = tmp_path / "big.txt"
        with pytest.raises(ValueError, match="content too large"):
            fs_ops.create(target, "x" * (fs_ops.MAX_BYTES + 1), overwrite=False)

    def test_atomic_write_leaves_no_tmp_on_success(self, tmp_path: Path):
        target = tmp_path / "f.txt"
        fs_ops.create(target, "hi", overwrite=False)
        # No leftover .tmp_* in the directory
        leftovers = [p for p in tmp_path.iterdir() if p.name.startswith(".tmp_")]
        assert leftovers == []


# --- read --------------------------------------------------------------

class TestRead:
    def test_reads_existing(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("hello", encoding="utf-8")
        out = fs_ops.read(f)
        assert out == {
            "path": str(f),
            "content": "hello",
            "size": 5,
            "encoding": "utf-8",
        }

    def test_missing_raises_keyerror(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.read(tmp_path / "nope.txt")

    def test_directory_raises(self, tmp_path: Path):
        with pytest.raises(ValueError, match="not a regular file"):
            fs_ops.read(tmp_path)

    def test_too_large_raises(self, tmp_path: Path):
        f = tmp_path / "big.bin"
        f.write_bytes(b"x" * (fs_ops.MAX_BYTES + 1))
        with pytest.raises(ValueError, match="file too large"):
            fs_ops.read(f)

    def test_non_utf8_raises(self, tmp_path: Path):
        f = tmp_path / "binary.bin"
        f.write_bytes(b"\xff\xfe\x00\x01")
        with pytest.raises(ValueError, match="not valid utf-8"):
            fs_ops.read(f)


# --- update ------------------------------------------------------------

class TestUpdate:
    def test_updates_existing(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("old", encoding="utf-8")
        result = fs_ops.update(f, "new")
        assert result.status == "UPDATED"
        assert result.size == 3
        assert f.read_text(encoding="utf-8") == "new"

    def test_missing_raises_keyerror(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.update(tmp_path / "nope.txt", "x")

    def test_too_large_raises(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("old", encoding="utf-8")
        with pytest.raises(ValueError, match="content too large"):
            fs_ops.update(f, "x" * (fs_ops.MAX_BYTES + 1))


# --- delete ------------------------------------------------------------

class TestDelete:
    def test_deletes_file(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("x", encoding="utf-8")
        result = fs_ops.delete(f, recursive=False)
        assert result.status == "DELETED"
        assert result.size is None
        assert not f.exists()

    def test_missing_raises_keyerror(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.delete(tmp_path / "nope.txt", recursive=False)

    def test_non_empty_dir_raises_without_recursive(self, tmp_path: Path):
        d = tmp_path / "d"
        d.mkdir()
        (d / "child").write_text("x", encoding="utf-8")
        with pytest.raises(ValueError, match="not empty"):
            fs_ops.delete(d, recursive=False)

    def test_recursive_deletes_tree(self, tmp_path: Path):
        d = tmp_path / "d"
        d.mkdir()
        (d / "child").write_text("x", encoding="utf-8")
        (d / "sub").mkdir()
        (d / "sub" / "leaf").write_text("y", encoding="utf-8")
        fs_ops.delete(d, recursive=True)
        assert not d.exists()

    def test_empty_dir_ok_without_recursive(self, tmp_path: Path):
        d = tmp_path / "empty"
        d.mkdir()
        fs_ops.delete(d, recursive=False)
        assert not d.exists()


# --- list_dir ----------------------------------------------------------

class TestListDir:
    def test_single_level(self, tmp_path: Path):
        (tmp_path / "a.txt").write_text("a", encoding="utf-8")
        (tmp_path / "b").mkdir()
        entries = fs_ops.list_dir(tmp_path, recursive=False, max_depth=1)
        names = sorted(e.name for e in entries)
        assert names == ["a.txt", "b"]
        assert all(isinstance(e, FileEntry) for e in entries)

    def test_recursive_respects_max_depth(self, tmp_path: Path):
        (tmp_path / "a.txt").write_text("a", encoding="utf-8")
        (tmp_path / "sub").mkdir()
        (tmp_path / "sub" / "b.txt").write_text("b", encoding="utf-8")
        (tmp_path / "sub" / "deep").mkdir()
        (tmp_path / "sub" / "deep" / "c.txt").write_text("c", encoding="utf-8")
        depth0 = fs_ops.list_dir(tmp_path, recursive=True, max_depth=0)
        assert sorted(e.name for e in depth0) == ["a.txt", "sub"]
        depth1 = fs_ops.list_dir(tmp_path, recursive=True, max_depth=1)
        assert sorted(e.name for e in depth1) == ["a.txt", "b.txt", "sub", "deep"]

    def test_missing_raises(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.list_dir(tmp_path / "nope", recursive=False, max_depth=1)

    def test_file_raises(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("x", encoding="utf-8")
        with pytest.raises(ValueError, match="not a directory"):
            fs_ops.list_dir(f, recursive=False, max_depth=1)

    def test_negative_depth_raises(self, tmp_path: Path):
        with pytest.raises(ValueError, match="max_depth"):
            fs_ops.list_dir(tmp_path, recursive=True, max_depth=-1)


# --- stat --------------------------------------------------------------

class TestStat:
    def test_file(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("hello", encoding="utf-8")
        e = fs_ops.stat(f)
        assert e.name == "f.txt"
        assert e.type == "file"
        assert e.size == 5

    def test_directory(self, tmp_path: Path):
        e = fs_ops.stat(tmp_path)
        assert e.type == "directory"
        assert e.size == 0

    def test_missing_raises(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.stat(tmp_path / "nope")


# --- search ------------------------------------------------------------

class TestSearch:
    def test_simple_glob(self, tmp_path: Path):
        (tmp_path / "a.py").write_text("a", encoding="utf-8")
        (tmp_path / "b.txt").write_text("b", encoding="utf-8")
        (tmp_path / "c.py").write_text("c", encoding="utf-8")
        matches = fs_ops.search(tmp_path, "*.py", recursive=False)
        assert sorted(Path(m).name for m in matches) == ["a.py", "c.py"]

    def test_recursive_glob(self, tmp_path: Path):
        (tmp_path / "sub").mkdir()
        (tmp_path / "sub" / "x.py").write_text("x", encoding="utf-8")
        (tmp_path / "a.py").write_text("a", encoding="utf-8")
        matches = fs_ops.search(tmp_path, "*.py", recursive=True)
        assert sorted(Path(m).name for m in matches) == ["a.py", "x.py"]

    def test_no_match_returns_empty(self, tmp_path: Path):
        (tmp_path / "a.txt").write_text("a", encoding="utf-8")
        assert fs_ops.search(tmp_path, "*.py", recursive=False) == []

    def test_missing_raises(self, tmp_path: Path):
        with pytest.raises(KeyError, match="not found"):
            fs_ops.search(tmp_path / "nope", "*", recursive=False)

    def test_non_directory_raises(self, tmp_path: Path):
        f = tmp_path / "f.txt"
        f.write_text("x", encoding="utf-8")
        with pytest.raises(ValueError, match="not a directory"):
            fs_ops.search(f, "*", recursive=False)


# --- move / copy -------------------------------------------------------

class TestMove:
    def test_rename(self, tmp_path: Path):
        src = tmp_path / "a.txt"
        src.write_text("hi", encoding="utf-8")
        result = fs_ops.move(src, tmp_path / "b.txt", overwrite=False)
        assert result.status == "MOVED"
        assert result.path == str(tmp_path / "b.txt")
        assert not src.exists()
        assert (tmp_path / "b.txt").read_text(encoding="utf-8") == "hi"

    def test_missing_src_raises(self, tmp_path: Path):
        with pytest.raises(KeyError, match="src not found"):
            fs_ops.move(tmp_path / "nope", tmp_path / "dst", overwrite=False)

    def test_existing_dst_raises_without_overwrite(self, tmp_path: Path):
        src = tmp_path / "a.txt"
        src.write_text("a", encoding="utf-8")
        dst = tmp_path / "b.txt"
        dst.write_text("b", encoding="utf-8")
        with pytest.raises(FileExistsError, match="dst already exists"):
            fs_ops.move(src, dst, overwrite=False)

    def test_overwrite_replaces(self, tmp_path: Path):
        src = tmp_path / "a.txt"
        src.write_text("new", encoding="utf-8")
        dst = tmp_path / "b.txt"
        dst.write_text("old", encoding="utf-8")
        fs_ops.move(src, dst, overwrite=True)
        assert dst.read_text(encoding="utf-8") == "new"
        assert not src.exists()


class TestCopy:
    def test_copies_file(self, tmp_path: Path):
        src = tmp_path / "a.txt"
        src.write_text("hi", encoding="utf-8")
        result = fs_ops.copy(src, tmp_path / "b.txt", overwrite=False)
        assert result.status == "COPIED"
        assert result.size == 2
        assert src.read_text(encoding="utf-8") == "hi"  # source unchanged
        assert (tmp_path / "b.txt").read_text(encoding="utf-8") == "hi"

    def test_missing_src_raises(self, tmp_path: Path):
        with pytest.raises(KeyError, match="src not found"):
            fs_ops.copy(tmp_path / "nope", tmp_path / "dst", overwrite=False)

    def test_existing_dst_raises_without_overwrite(self, tmp_path: Path):
        src = tmp_path / "a.txt"
        src.write_text("a", encoding="utf-8")
        dst = tmp_path / "b.txt"
        dst.write_text("b", encoding="utf-8")
        with pytest.raises(FileExistsError, match="dst already exists"):
            fs_ops.copy(src, dst, overwrite=False)

    def test_copies_directory(self, tmp_path: Path):
        src = tmp_path / "d"
        src.mkdir()
        (src / "child").write_text("x", encoding="utf-8")
        result = fs_ops.copy(src, tmp_path / "d2", overwrite=False)
        assert result.status == "COPIED"
        assert (tmp_path / "d2" / "child").read_text(encoding="utf-8") == "x"
  • Step 2.4: Run the test to verify it passes

Run: uv run pytest tests/unit/test_fs_ops.py -v Expected: 38 PASS (5 + 5 + 3 + 5 + 5 + 3 + 5 + 4 + 4 = 39; one TestRead case is "non_utf8" instead of one fewer — actual count is 39)

  • Step 2.5: Commit
git add files_mcp/models.py files_mcp/core/fs_ops.py tests/unit/test_fs_ops.py
git commit -m "feat(files-mcp): add Pydantic models and pure fs_ops layer

Adds files_mcp/models.py (FileEntry, OperationResult) and
files_mcp/core/fs_ops.py with 9 pure file operations on already-resolved
Path objects: create/read/update/delete/list_dir/stat/search/move/copy.

All writes go through tempfile.mkstemp + os.replace for atomicity;
read enforces utf-8 + 1MB cap; create/update enforce strict presence
semantics (create: absent or overwrite; update: present). 39 unit tests
in tests/unit/test_fs_ops.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"

Task 3: Business wrappers + request models

Files:

  • Create: files_mcp/tools/__init__.py
  • Create: files_mcp/tools/files.py
  • Create: files_mcp/tools/requests.py
  • Create: tests/unit/test_files_tool.py

Interfaces (this task produces; server.py in Task 4 consumes):

  • 9 Pydantic body models in files_mcp/tools/requests.py — one per tool.

  • 9 wrapper functions in files_mcp/tools/files.py. Each:

    1. Logs DEBUG on entry.
    2. Calls path_guard.resolve_and_check on every user-supplied path.
    3. Delegates to fs_ops.
    4. Returns model_dump() of the resulting Pydantic model (or the dict from read).
  • Step 3.1: Create files_mcp/tools/__init__.py

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
  • Step 3.2: Create files_mcp/tools/requests.py
# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen

Pydantic body models for the FastAPI route layer. fastapi-mcp passes
tool args as a JSON body, so every tool needs a model (even single-arg
ones) — dict-typed query params would 422.
"""
from pydantic import BaseModel, Field


class CreateFileRequest(BaseModel):
    path: str
    content: str
    overwrite: bool = False


class ReadFileRequest(BaseModel):
    path: str


class UpdateFileRequest(BaseModel):
    path: str
    content: str


class DeleteFileRequest(BaseModel):
    path: str
    recursive: bool = False


class ListDirRequest(BaseModel):
    path: str
    recursive: bool = False
    max_depth: int = Field(default=1, ge=0)


class StatRequest(BaseModel):
    path: str


class SearchRequest(BaseModel):
    path: str
    pattern: str
    recursive: bool = True


class MoveRequest(BaseModel):
    src: str
    dst: str
    overwrite: bool = False


class CopyRequest(BaseModel):
    src: str
    dst: str
    overwrite: bool = False
  • Step 3.3: Create files_mcp/tools/files.py
# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen

Business wrappers: each tool function logs on entry, calls path_guard
on every user-supplied path, delegates to fs_ops, and returns the
Pydantic-serialized result. No business logic of its own — same shape
as spark_executor/tools/connections.py.
"""
from common.logging import logger
from files_mcp.core import fs_ops, path_guard


def create_file(*, path: str, content: str, overwrite: bool = False) -> dict:
    logger.debug(
        f"create_file enter path={path} content_len={len(content)} overwrite={overwrite}"
    )
    p = path_guard.resolve_and_check(path)
    return fs_ops.create(p, content, overwrite).model_dump()


def read_file(*, path: str) -> dict:
    logger.debug(f"read_file enter path={path}")
    p = path_guard.resolve_and_check(path)
    return fs_ops.read(p)


def update_file(*, path: str, content: str) -> dict:
    logger.debug(f"update_file enter path={path} content_len={len(content)}")
    p = path_guard.resolve_and_check(path)
    return fs_ops.update(p, content).model_dump()


def delete_file(*, path: str, recursive: bool = False) -> dict:
    logger.debug(f"delete_file enter path={path} recursive={recursive}")
    p = path_guard.resolve_and_check(path)
    return fs_ops.delete(p, recursive).model_dump()


def list_dir(
    *, path: str, recursive: bool = False, max_depth: int = 1
) -> dict:
    logger.debug(
        f"list_dir enter path={path} recursive={recursive} max_depth={max_depth}"
    )
    p = path_guard.resolve_and_check(path)
    return {"path": str(p), "entries": [e.model_dump() for e in fs_ops.list_dir(p, recursive, max_depth)]}


def stat(*, path: str) -> dict:
    logger.debug(f"stat enter path={path}")
    p = path_guard.resolve_and_check(path)
    return fs_ops.stat(p).model_dump()


def search(*, path: str, pattern: str, recursive: bool = True) -> dict:
    logger.debug(
        f"search enter path={path} pattern={pattern!r} recursive={recursive}"
    )
    p = path_guard.resolve_and_check(path)
    return {"pattern": pattern, "matches": fs_ops.search(p, pattern, recursive)}


def move(*, src: str, dst: str, overwrite: bool = False) -> dict:
    logger.debug(f"move enter src={src} dst={dst} overwrite={overwrite}")
    s = path_guard.resolve_and_check(src)
    d = path_guard.resolve_and_check(dst)
    return fs_ops.move(s, d, overwrite).model_dump()


def copy(*, src: str, dst: str, overwrite: bool = False) -> dict:
    logger.debug(f"copy enter src={src} dst={dst} overwrite={overwrite}")
    s = path_guard.resolve_and_check(src)
    d = path_guard.resolve_and_check(dst)
    return fs_ops.copy(s, d, overwrite).model_dump()
  • Step 3.4: Write the failing test

Create tests/unit/test_files_tool.py:

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from pathlib import Path

import pytest

from files_mcp.core import path_guard
from files_mcp.tools import files


@pytest.fixture
def sandbox_root(tmp_path, monkeypatch):
    monkeypatch.setenv("FILES_MCP_ROOT", str(tmp_path))
    path_guard.set_root_for_testing(tmp_path.resolve())
    return tmp_path.resolve()


class TestCreateFile:
    def test_creates_file(self, sandbox_root: Path):
        result = files.create_file(path="new.txt", content="hello")
        assert result["status"] == "CREATED"
        assert (sandbox_root / "new.txt").read_text(encoding="utf-8") == "hello"

    def test_escape_raises_value_error(self, sandbox_root: Path):
        with pytest.raises(ValueError, match="escapes sandbox"):
            files.create_file(path="../escape.txt", content="x")

    def test_existing_raises_file_exists(self, sandbox_root: Path):
        (sandbox_root / "f.txt").write_text("old", encoding="utf-8")
        with pytest.raises(FileExistsError):
            files.create_file(path="f.txt", content="new", overwrite=False)


class TestReadFile:
    def test_reads_existing(self, sandbox_root: Path):
        (sandbox_root / "f.txt").write_text("hi", encoding="utf-8")
        out = files.read_file(path="f.txt")
        assert out["content"] == "hi"
        assert out["encoding"] == "utf-8"

    def test_escape_raises_value_error(self, sandbox_root: Path):
        with pytest.raises(ValueError, match="escapes sandbox"):
            files.read_file(path="/etc/passwd")


class TestUpdateFile:
    def test_updates_existing(self, sandbox_root: Path):
        (sandbox_root / "f.txt").write_text("old", encoding="utf-8")
        result = files.update_file(path="f.txt", content="new")
        assert result["status"] == "UPDATED"
        assert (sandbox_root / "f.txt").read_text(encoding="utf-8") == "new"

    def test_missing_raises_keyerror(self, sandbox_root: Path):
        with pytest.raises(KeyError, match="not found"):
            files.update_file(path="nope.txt", content="x")


class TestDeleteFile:
    def test_deletes_file(self, sandbox_root: Path):
        (sandbox_root / "f.txt").write_text("x", encoding="utf-8")
        result = files.delete_file(path="f.txt")
        assert result["status"] == "DELETED"


class TestListDir:
    def test_returns_entries(self, sandbox_root: Path):
        (sandbox_root / "a.txt").write_text("a", encoding="utf-8")
        (sandbox_root / "b").mkdir()
        out = files.list_dir(path=".", recursive=False, max_depth=1)
        assert out["path"] == str(sandbox_root)
        names = sorted(e["name"] for e in out["entries"])
        assert names == ["a.txt", "b"]


class TestStat:
    def test_returns_file_entry(self, sandbox_root: Path):
        (sandbox_root / "f.txt").write_text("hi", encoding="utf-8")
        out = files.stat(path="f.txt")
        assert out["name"] == "f.txt"
        assert out["type"] == "file"


class TestSearch:
    def test_finds_matches(self, sandbox_root: Path):
        (sandbox_root / "a.py").write_text("a", encoding="utf-8")
        (sandbox_root / "b.txt").write_text("b", encoding="utf-8")
        out = files.search(path=".", pattern="*.py", recursive=False)
        assert out["pattern"] == "*.py"
        assert len(out["matches"]) == 1
        assert out["matches"][0].endswith("a.py")


class TestMove:
    def test_moves_file(self, sandbox_root: Path):
        (sandbox_root / "a.txt").write_text("x", encoding="utf-8")
        result = files.move(src="a.txt", dst="b.txt")
        assert result["status"] == "MOVED"
        assert not (sandbox_root / "a.txt").exists()
        assert (sandbox_root / "b.txt").exists()

    def test_escape_on_dst_raises(self, sandbox_root: Path):
        (sandbox_root / "a.txt").write_text("x", encoding="utf-8")
        with pytest.raises(ValueError, match="escapes sandbox"):
            files.move(src="a.txt", dst="../escape.txt")


class TestCopy:
    def test_copies_file(self, sandbox_root: Path):
        (sandbox_root / "a.txt").write_text("x", encoding="utf-8")
        result = files.copy(src="a.txt", dst="b.txt")
        assert result["status"] == "COPIED"
        assert (sandbox_root / "a.txt").exists()  # source unchanged
        assert (sandbox_root / "b.txt").read_text(encoding="utf-8") == "x"
  • Step 3.5: Run the test to verify it passes

Run: uv run pytest tests/unit/test_files_tool.py -v Expected: 14 PASS

  • Step 3.6: Run the full unit suite to confirm no regression in Tasks 13

Run: uv run pytest tests/unit/ -v Expected: all PASS (12 path_guard + 39 fs_ops + 14 files_tool = 65)

  • Step 3.7: Commit
git add files_mcp/tools/__init__.py files_mcp/tools/files.py files_mcp/tools/requests.py tests/unit/test_files_tool.py
git commit -m "feat(files-mcp): business wrappers and Pydantic request models

Adds files_mcp/tools/requests.py with 9 Pydantic body models (one per
tool, since fastapi-mcp passes args as JSON body) and
files_mcp/tools/files.py with 9 thin wrappers that call path_guard
then fs_ops and return model_dump() form. Adds tests/unit/test_files_tool.py
covering happy path, sandbox escape, and KeyError/FileExistsError
propagation for each wrapper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"

Task 4: Server + integration tests

Files:

  • Create: files_mcp/server.py
  • Create: tests/integration/test_files_mcp_routes.py

Interfaces (this task produces; Task 5 wires into main.py):

  • files_mcp.server.app — FastAPI instance titled "Files MCP", version "0.0.1", description "Files MCP Server".

  • 3 exception handlers: KeyError → 404, ValueError → 400, FileExistsError → 409.

  • 9 POST routes, each with operation_id / summary / description.

  • Step 4.1: Create files_mcp/server.py

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse

from files_mcp.tools.files import (
    copy as copy_file,
    create_file,
    delete_file,
    list_dir,
    move as move_file,
    read_file,
    search,
    stat,
    update_file,
)
from files_mcp.tools.requests import (
    CopyRequest,
    CreateFileRequest,
    DeleteFileRequest,
    ListDirRequest,
    MoveRequest,
    ReadFileRequest,
    SearchRequest,
    StatRequest,
    UpdateFileRequest,
)

app = FastAPI(
    title="Files MCP",
    version="0.0.1",
    description="Files MCP Server — sandboxed file-system CRUD over fastapi-mcp",
)


# --- Exception handlers ---
# KeyError  -> 404 (path not found)
# ValueError -> 400 (bad input: escape, non-utf-8, too large, wrong type)
# FileExistsError -> 409 (create/move/copy hit existing dst without overwrite)
#
# Only the third handler is new here; the first two mirror the
# spark_executor.server convention. They live on this app only, so the
# spark_executor handlers are unaffected.


@app.exception_handler(KeyError)
async def _keyerror_handler(_request: Request, exc: KeyError) -> JSONResponse:
    return JSONResponse(status_code=404, content={"detail": str(exc)})


@app.exception_handler(ValueError)
async def _valueerror_handler(_request: Request, exc: ValueError) -> JSONResponse:
    return JSONResponse(status_code=400, content={"detail": str(exc)})


@app.exception_handler(FileExistsError)
async def _fileexists_handler(_request: Request, exc: FileExistsError) -> JSONResponse:
    return JSONResponse(status_code=409, content={"detail": str(exc)})


# --- Routes ---


@app.post(
    "/create_file",
    operation_id="create_file",
    summary="Create a new text file inside the sandbox",
    description=(
        "Writes `content` to `path` (resolved against FILES_MCP_ROOT). "
        "Atomic write via tempfile + os.replace. Default fails with 409 "
        "if the file already exists; pass overwrite=true to replace. "
        "Content is capped at 1 MB. utf-8 only."
    ),
)
def _create_file(req: CreateFileRequest):
    return create_file(path=req.path, content=req.content, overwrite=req.overwrite)


@app.post(
    "/read_file",
    operation_id="read_file",
    summary="Read a text file inside the sandbox",
    description=(
        "Returns the full utf-8 content. Files larger than 1 MB and "
        "non-utf-8 files are rejected with 400. Directories are rejected "
        "with 400; use list_dir instead."
    ),
)
def _read_file(req: ReadFileRequest):
    return read_file(path=req.path)


@app.post(
    "/update_file",
    operation_id="update_file",
    summary="Overwrite an existing text file inside the sandbox",
    description=(
        "Replaces the file in full. 404 if the path does not exist. "
        "Atomic write, 1 MB cap, utf-8 only."
    ),
)
def _update_file(req: UpdateFileRequest):
    return update_file(path=req.path, content=req.content)


@app.post(
    "/delete_file",
    operation_id="delete_file",
    summary="Delete a file or empty directory",
    description=(
        "404 if the path does not exist. To delete a non-empty directory, "
        "pass recursive=true."
    ),
)
def _delete_file(req: DeleteFileRequest):
    return delete_file(path=req.path, recursive=req.recursive)


@app.post(
    "/list_dir",
    operation_id="list_dir",
    summary="List entries in a directory",
    description=(
        "Returns a list of FileEntry records. recursive=false lists only "
        "the directory's direct children; recursive=true descends up to "
        "max_depth levels (0 = the directory itself, 1 = one level deep)."
    ),
)
def _list_dir(req: ListDirRequest):
    return list_dir(path=req.path, recursive=req.recursive, max_depth=req.max_depth)


@app.post(
    "/stat",
    operation_id="stat",
    summary="Get a single FileEntry (path, type, size, mtime, mode)",
    description="404 if the path does not exist.",
)
def _stat(req: StatRequest):
    return stat(path=req.path)


@app.post(
    "/search",
    operation_id="search",
    summary="Find paths matching a glob pattern",
    description=(
        "Returns the absolute paths of entries inside `path` whose name "
        "matches the glob `pattern`. recursive=true descends into "
        "subdirectories; recursive=false matches the directory's direct "
        "children only."
    ),
)
def _search(req: SearchRequest):
    return search(path=req.path, pattern=req.pattern, recursive=req.recursive)


@app.post(
    "/move",
    operation_id="move",
    summary="Rename or move a file or directory",
    description=(
        "404 if src does not exist. 409 if dst exists and overwrite=false. "
        "Cross-filesystem moves are not atomic."
    ),
)
def _move(req: MoveRequest):
    return move_file(src=req.src, dst=req.dst, overwrite=req.overwrite)


@app.post(
    "/copy",
    operation_id="copy",
    summary="Copy a file or directory",
    description=(
        "404 if src does not exist. 409 if dst exists and overwrite=false. "
        "For files, metadata is preserved (shutil.copy2). For directories, "
        "the entire tree is copied (shutil.copytree)."
    ),
)
def _copy(req: CopyRequest):
    return copy_file(src=req.src, dst=req.dst, overwrite=req.overwrite)
  • Step 4.2: Write the integration test

Create tests/integration/test_files_mcp_routes.py:

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen

End-to-end HTTP tests for files_mcp.server.app. Uses FastAPI TestClient
directly (no MCP transport) — same pattern as
tests/integration/test_mcp_routes.py.
"""
from fastapi.testclient import TestClient

from files_mcp.core import path_guard
from files_mcp.server import app


def _client(tmp_path) -> TestClient:
    import os
    os.environ["FILES_MCP_ROOT"] = str(tmp_path)
    path_guard.set_root_for_testing(tmp_path.resolve())
    return TestClient(app)


class TestCrud:
    def test_create_then_read(self, tmp_path):
        with _client(tmp_path) as c:
            r = c.post("/create_file", json={"path": "a.txt", "content": "hello"})
            assert r.status_code == 200
            assert r.json() == {"path": str(tmp_path / "a.txt"), "status": "CREATED", "size": 5}
            r = c.post("/read_file", json={"path": "a.txt"})
            assert r.status_code == 200
            assert r.json()["content"] == "hello"

    def test_update_existing(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "old"})
            r = c.post("/update_file", json={"path": "a.txt", "content": "new"})
            assert r.status_code == 200
            assert r.json()["status"] == "UPDATED"
            r = c.post("/read_file", json={"path": "a.txt"})
            assert r.json()["content"] == "new"

    def test_delete(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "x"})
            r = c.post("/delete_file", json={"path": "a.txt"})
            assert r.status_code == 200
            assert r.json()["status"] == "DELETED"
            r = c.post("/read_file", json={"path": "a.txt"})
            assert r.status_code == 404


class TestErrors:
    def test_create_existing_returns_409(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "x"})
            r = c.post("/create_file", json={"path": "a.txt", "content": "y"})
            assert r.status_code == 409

    def test_update_missing_returns_404(self, tmp_path):
        with _client(tmp_path) as c:
            r = c.post("/update_file", json={"path": "nope.txt", "content": "x"})
            assert r.status_code == 404

    def test_escape_returns_400(self, tmp_path):
        with _client(tmp_path) as c:
            r = c.post("/read_file", json={"path": "/etc/passwd"})
            assert r.status_code == 400
            assert "escapes sandbox" in r.json()["detail"]

    def test_too_large_returns_400(self, tmp_path):
        with _client(tmp_path) as c:
            r = c.post(
                "/create_file",
                json={"path": "big.txt", "content": "x" * (1024 * 1024 + 1)},
            )
            assert r.status_code == 400
            assert "too large" in r.json()["detail"]

    def test_non_utf8_returns_400(self, tmp_path):
        with _client(tmp_path) as c:
            (tmp_path / "bin").write_bytes(b"\xff\xfe\x00")
            r = c.post("/read_file", json={"path": "bin"})
            assert r.status_code == 400
            assert "utf-8" in r.json()["detail"]


class TestIntrospection:
    def test_list_dir(self, tmp_path):
        with _client(tmp_path) as c:
            (tmp_path / "a.txt").write_text("a", encoding="utf-8")
            (tmp_path / "b").mkdir()
            r = c.post("/list_dir", json={"path": ".", "recursive": False, "max_depth": 1})
            assert r.status_code == 200
            names = sorted(e["name"] for e in r.json()["entries"])
            assert names == ["a.txt", "b"]

    def test_stat(self, tmp_path):
        with _client(tmp_path) as c:
            (tmp_path / "f.txt").write_text("hi", encoding="utf-8")
            r = c.post("/stat", json={"path": "f.txt"})
            assert r.status_code == 200
            assert r.json()["type"] == "file"
            assert r.json()["size"] == 2

    def test_search(self, tmp_path):
        with _client(tmp_path) as c:
            (tmp_path / "a.py").write_text("a", encoding="utf-8")
            (tmp_path / "b.txt").write_text("b", encoding="utf-8")
            r = c.post("/search", json={"path": ".", "pattern": "*.py", "recursive": False})
            assert r.status_code == 200
            assert len(r.json()["matches"]) == 1


class TestMoveCopy:
    def test_move(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "x"})
            r = c.post("/move", json={"src": "a.txt", "dst": "b.txt"})
            assert r.status_code == 200
            assert r.json()["status"] == "MOVED"
            r = c.post("/read_file", json={"path": "b.txt"})
            assert r.json()["content"] == "x"

    def test_move_existing_dst_409(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "x"})
            c.post("/create_file", json={"path": "b.txt", "content": "y"})
            r = c.post("/move", json={"src": "a.txt", "dst": "b.txt"})
            assert r.status_code == 409

    def test_copy(self, tmp_path):
        with _client(tmp_path) as c:
            c.post("/create_file", json={"path": "a.txt", "content": "x"})
            r = c.post("/copy", json={"src": "a.txt", "dst": "b.txt"})
            assert r.status_code == 200
            assert r.json()["status"] == "COPIED"
            # source unchanged
            r = c.post("/read_file", json={"path": "a.txt"})
            assert r.json()["content"] == "x"
            r = c.post("/read_file", json={"path": "b.txt"})
            assert r.json()["content"] == "x"
  • Step 4.3: Run the integration test

Run: uv run pytest tests/integration/test_files_mcp_routes.py -v Expected: 12 PASS

  • Step 4.4: Run the full test suite to confirm no regression

Run: uv run pytest -v Expected: all PASS (existing 242 + new 65 unit + 12 integration = ~319 tests)

(If a spark_executor test fails, stop and re-check that nothing in tests/conftest.py, main.py, or spark_executor/ was accidentally modified. This plan should not change any of them yet.)

  • Step 4.5: Commit
git add files_mcp/server.py tests/integration/test_files_mcp_routes.py
git commit -m "feat(files-mcp): FastAPI app with 9 tool routes and integration tests

Adds files_mcp/server.py (FastAPI title=Files MCP, 9 POST routes, and
3 exception handlers: KeyError->404, ValueError->400, FileExistsError->409).
The third handler is the only one not present in spark_executor.server;
handlers are registered on this app only, so spark_executor is unaffected.

Adds tests/integration/test_files_mcp_routes.py with 12 end-to-end HTTP
tests covering CRUD happy paths, error codes (404/400/409), and
introspection (list_dir, stat, search, move, copy).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"

Task 5: Wire into main.py + conftest fixture + final verification

Files:

  • Modify: main.py (add one import + one McpService entry)

  • Modify: tests/conftest.py (add one autouse fixture)

  • Modify (verification only): nothing else

  • Step 5.1: Read the current main.py to confirm the exact line to change

Run: cat main.py | head -45 Confirm: the MCP_SERVICES list ends at the existing McpService(name="spark_executor", ...). (No change to the lifespan function.)

  • Step 5.2: Edit main.py — add the import

Add this line alphabetically with the other top-level imports (after from spark_executor import app as spark_executor_app):

from files_mcp import app as files_app

The diff at the top of main.py becomes:

from common.factory import init_mcp_server
from common.logging import logger
from files_mcp import app as files_app
from spark_executor import app as spark_executor_app
  • Step 5.3: Edit main.py — add the McpService entry

In the MCP_SERVICES list, add a new entry after the existing spark_executor entry:

MCP_SERVICES: list[McpService] = [
    McpService(
        name="spark_executor",
        app=spark_executor_app,
        mount_path="/spark-executor-mcp",
    ),
    McpService(
        name="files",
        app=files_app,
        mount_path="/files-mcp",
    ),
]
  • Step 5.4: Read the current tests/conftest.py

Run: cat tests/conftest.py Confirm: there is no existing autouse fixture. The file currently only adds the repo root to sys.path.

  • Step 5.5: Edit tests/conftest.py — add the autouse fixture

Append the following to tests/conftest.py:

# coding=utf-8
"""
@Time   :2026/6/30
@Author :tao.chen
"""
import pytest


@pytest.fixture(autouse=True)
def _isolate_files_mcp_root(tmp_path, monkeypatch):
    """Point files_mcp at a fresh tmp_path for every test.

    Sets FILES_MCP_ROOT and rebinds files_mcp.core.path_guard's module-level
    _ROOT, so each test starts with an empty sandbox. Spark_executor tests
    are unaffected by this fixture — they don't touch the files_mcp sandbox.
    """
    monkeypatch.setenv("FILES_MCP_ROOT", str(tmp_path))
    from files_mcp.core import path_guard
    path_guard.set_root_for_testing(tmp_path.resolve())
    yield

(If tests/conftest.py already has its own header, keep it; only append the fixture and the import.)

  • Step 5.6: Run the full test suite

Run: uv run pytest -v Expected: all PASS (existing 242 + 77 new = ~319 tests)

If a spark_executor test fails:

  • Check that main.py only gained the import + entry, nothing else.

  • Check that tests/conftest.py only gained the autouse fixture, nothing else.

  • Check that no other spark_executor/ file was modified.

  • Step 5.7: Smoke-test that the lifespan mounts both services

Create a one-off script scripts/smoke_files_mcp.py (NOT committed, used only to verify the wiring):

# coding=utf-8
"""One-off smoke test: start the root app, hit /health, list both MCP tools."""
import os
os.environ["FILES_MCP_ROOT"] = "/tmp/files-mcp-smoke"
os.makedirs("/tmp/files-mcp-smoke", exist_ok=True)

from fastapi.testclient import TestClient
from main import app

with TestClient(app) as c:
    # /health is on the root app
    r = c.get("/health")
    assert r.status_code == 200, r.text

    # /spark-executor-mcp tools/list
    r = c.get("/spark-executor-mcp/tools", headers={"mcp-session-id": "smoke"})
    print("spark-executor status:", r.status_code)

    # /files-mcp tools/list
    r = c.get("/files-mcp/tools", headers={"mcp-session-id": "smoke"})
    print("files status:", r.status_code)

    # End-to-end create + read through /files-mcp
    r = c.post(
        "/files-mcp/create_file",
        json={"path": "smoke.txt", "content": "smoke"},
        headers={"mcp-session-id": "smoke"},
    )
    print("create status:", r.status_code, r.text)
    r = c.post(
        "/files-mcp/read_file",
        json={"path": "smoke.txt"},
        headers={"mcp-session-id": "smoke"},
    )
    print("read status:", r.status_code, r.text)

print("OK")

Run: uv run python scripts/smoke_files_mcp.py Expected: prints create status: 200 ... and read status: 200 ..., then OK. (The tools/list calls may return 4xx because they require an MCP initialize handshake before tools/list; the create+read round-trip is the meaningful end-to-end check.)

Then delete the smoke script: rm scripts/smoke_files_mcp.py

(If scripts/ is gitignored or does not exist, just create and remove the file in a temp location; do not commit it.)

  • Step 5.8: Commit the integration changes
git add main.py tests/conftest.py
git commit -m "feat(files-mcp): mount at /files-mcp via MCP_SERVICES registry

- main.py: import files_app, append McpService entry to MCP_SERVICES.
  Lifespan already iterates the list, so no edit to the lifespan function.
- tests/conftest.py: autouse fixture rebinds files_mcp.core.path_guard._ROOT
  to a per-test tmp_path so the sandbox is fresh for every test.

Verified via uv run pytest: 242 existing + 77 new tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
  • Step 5.9: Verify the branch diff is exactly the expected scope

Run: git diff main --stat Expected: the change set is confined to:

  • files_mcp/ (new package, 7 files)

  • tests/unit/test_path_guard.py, tests/unit/test_fs_ops.py, tests/unit/test_files_tool.py (new)

  • tests/integration/test_files_mcp_routes.py (new)

  • main.py (modified, +2 lines)

  • tests/conftest.py (modified, +15 lines)

  • docs/superpowers/specs/2026-06-30-files-mcp-design.md (already on branch from brainstorming)

  • pyproject.toml and uv.lock should be unchanged.

  • Step 5.10: Push the branch (optional)

Run: git push -u origin feat/files-mcp (Only if the user has authorized a push; otherwise leave the branch local and report the commit list.)


Self-Review Checklist

Run mentally before declaring the plan done:

  1. Spec coverage:

    • 9 tools (CRUD + list_dir + stat + search + move + copy) → Tasks 24 ✓
    • FILES_MCP_ROOT env var + startup-time validation → Task 1 ✓
    • Sandbox via .resolve() + relative_to → Task 1 ✓
    • Atomic writes via tempfile + os.replace → Task 2 ✓
    • 1 MB cap → Task 2 ✓
    • utf-8 only → Task 2 ✓
    • Strict create/update semantics → Task 2 ✓
    • Pydantic body models for fastapi-mcp → Task 3 ✓
    • 3 exception handlers (KeyError, ValueError, FileExistsError) → Task 4 ✓
    • One-line MCP_SERVICES integration → Task 5 ✓
    • 0 new dependencies → confirmed in Global Constraints ✓
    • 0 changes to spark_executor/ → confirmed in Task 5 verification ✓
    • GUNICORN_WORKERS=1 warning still applies → no edit needed (it covers the root app) ✓
  2. Placeholder scan: no TBD / TODO / "implement later" / "see above" in the plan.

  3. Type consistency:

    • resolve_and_check(path: str) -> Path defined in Task 1, used in Task 3.
    • _init_root() -> Path defined in Task 1, used in __init__.py in Task 1.
    • set_root_for_testing(p: Path) -> None defined in Task 1, used in conftest in Task 5.
    • All 9 fs_ops.* function names match between Task 2 and Task 3.
    • All 9 wrapper function names match between Task 3 and Task 4.
    • All 9 Pydantic model names match between Task 3 and Task 4.
    • Pydantic field names (path, content, overwrite, recursive, max_depth, pattern, src, dst) consistent across requests.py, files.py, server.py, and the tests.

Acceptance Criteria

  • All 5 task checklists complete.
  • uv run pytest passes (242 existing + 77 new = ~319).
  • git diff main --stat shows the expected scope (no surprises).
  • pyproject.toml / uv.lock are byte-for-byte unchanged.
  • spark_executor/ directory is byte-for-byte unchanged (other than the spec doc in docs/).
  • The smoke test in Step 5.7 successfully creates and reads a file through the /files-mcp mount.