Files
mcp-server/spark_executor/tools/generate.py
T
Claude 34e6208f54 feat: SQL safety policy (SELECT/INSERT only) at submit time
The agent can now write PySpark that runs DROP/DELETE/UPDATE/etc. on
production tables. Add a static guard that rejects anything other than
SELECT and INSERT at two enforcement points:

  1. generate_job_file: validates BEFORE writing to disk. Agent gets
     immediate feedback ('rewrite to use only SELECT/INSERT') rather
     than learning at submit time.

  2. prepare_submit_job: re-validates the script content (reads the
     file) as a defense-in-depth check. Catches host-mounted files,
     manually-edited files, anything that bypassed generate_job_file.

How it works:
  - common/sql_guard.py extracts Python string literals whose first
    keyword is a SQL verb (catches spark.sql('...'), f-strings, and any
    raw SQL literal)
  - sqlparse splits each literal into statements; we check the first
    keyword against the policy (SELECT/INSERT/WITH allowed; DROP,
    DELETE, UPDATE, TRUNCATE, ALTER, CREATE, REPLACE, MERGE, GRANT,
    REVOKE, SET, SHOW, KILL, EXEC, etc. forbidden)
  - WITH recurses into the CTE body to catch WITH x AS (DROP ...) ...
  - The MCP layer maps ValueError -> HTTP 400 (existing handler)

Test coverage:
  - 28 unit tests in test_sql_guard.py cover: extraction (single/double/
    f-string, English false positives, multi-literal), statement
    classification (SELECT, INSERT, DROP, DELETE, UPDATE, TRUNCATE,
    ALTER, CREATE, multi-statement, CTE bodies, comments)
  - 2 integration tests verify MCP layer returns 400 with the policy
    explanation at both generate_job_file and prepare_submit_job

Limitations (documented in sql_guard.py docstring):
  - f-strings where the SQL is built at runtime (e.g. f'SELECT * FROM
    {user_input}') look like SELECTs at static-analysis time. The
    guard catches the static literal; the runtime substitution is the
    caller's responsibility.
  - pyspark.sql.functions.expr('...') accepts SQL inline; not currently
    caught. (Future work.)

146/146 still pass. Live verified: DROP TABLE -> MCP 400 with policy
explanation; SELECT -> MCP 200 + file written to ./data/jobs/.
2026-06-25 12:52:16 +08:00

44 lines
1.5 KiB
Python

# coding=utf-8
"""
@Time :2026/6/24
@Author :tao.chen
"""
from common.logging import logger
from common.sql_guard import validate_pyspark_code
from spark_executor.core.job_writer import write_job_file
class SqlGuardViolation(ValueError):
"""Raised when generate_job_file receives code that violates the SQL
safety policy. -> HTTP 400 via the FastAPI ValueError handler.
"""
pass
def generate_job_file(code: str) -> dict[str, str]:
"""Write a PySpark code string to disk; return its absolute path.
Validates the code against the SQL safety policy (SELECT/INSERT only)
BEFORE writing, so the agent gets immediate feedback rather than
learning at prepare_submit_job time. Use the returned path as the
`script_path` argument of prepare_submit_job.
The output directory is controlled by the SPARK_EXECUTOR_JOBS_DIR env var
(default: ./data/jobs/).
"""
logger.debug(f"generate_job_file enter code_bytes={len(code)}")
offenses = validate_pyspark_code(code)
if offenses:
logger.warning(
f"generate_job_file rejected: SQL policy violation(s): {offenses}"
)
raise SqlGuardViolation(
f"PySpark code violates SQL safety policy. "
f"Only SELECT and INSERT statements are allowed. "
f"Found forbidden statement(s): {offenses}. "
f"Rewrite the code to use only SELECT/INSERT (or DataFrame DSL)."
)
path = write_job_file(code)
logger.info(f"generate_job_file ok script_path={path}")
return {"script_path": path}