dab4cd062e2b127d9400399b585680df80193693
对应代码审查发现的问题(#1, #2, #4-#8, #10): #1 CRITICAL:spark_submit 在 argv 中重复拼接 binary。此前 cmd := []string{binary} 后又把 cmd 作为 Args 传给 executor.Run, 而 executor 会再拼一次 Binary,导致 OS argv 为 [binary, binary, ...], spark-submit 会把自身当作应用 jar。现在 cmd 从 --master 开始, executor.Run 使用 Binary + Args,argv 正确。 #2:Validate 曾接受 .meta.json 路径本身。现在显式拒绝 sidecar 路径, 要求传入数据文件路径。 #4:Sweep 对 sidecar 损坏的数据文件跳过清理。现在损坏 sidecar 会回退 到数据文件 mtime,超期即删除。 #5:upload_file 描述仍引用已移除的 args 字段,已改为引用 script_path 及结构化字段。 #6:Deps.UploadStore 改为值类型 uploads.Store,避免 nil 绕过上传校验; 移除 spark_submit/upload_file 中的 nil 检查。 #7:master/queue/executor_memory 增加空字符串校验。 #8:提取 buildSparkSubmitCommand 构建 argv,消除双写参数的结构性根因。 #10:Validate 失败时记录 slog.Warn("spark_submit.unminted_path_rejected")。 新增测试: - TestSparkSubmit_StructuredCommand:断言 argv 首行为 --master,末行 仍为 script_path。 - TestSparkSubmit_EmptyMaster:空 master 返回错误。 - TestStore_Validate_RejectsSidecarPath:拒绝 .meta.json 路径。 - TestStore_Sweep_DeletesDataWithCorruptSidecar:损坏 sidecar 的数据文件 被清理。 未在本提交处理: - #9 cluster.DefaultSubmitArgs 弃用留作后续批次。 Co-Authored-By: tao.chen <93983997+taochen-ct@users.noreply.github.com>
spark-mcp-go
MCP (Model Context Protocol) server for Apache Spark on YARN. Lets an LLM agent
discover Spark/YARN endpoints, submit jobs, fetch logs, and analyze them via
11 typed Tools. Streamable HTTP transport, SQLite-backed configuration, and
log/slog structured logging with per-Tool call files.
What it does
- Exposes 11 MCP Tools over Streamable HTTP at
/mcp - Admin API at
/admin/*for cluster and audit configuration - HTTP Basic and YARN SimpleAuth for Spark/YARN endpoints
- SSRF protection with cluster URL allowlist plus DNS rebinding guard
- Per-Tool independent audit log (file mode
0600) spark-submitvia localexec.Command(no shell, no injection)
Quick start (3 steps)
-
Copy and edit the environment file:
cp .env.example .env # Edit ADMIN_TOKENS and AGENT_TOKEN -
Build and start the server:
go build ./... ./spark-mcp-go # Or use the helper: # ./scripts/dev.sh -
Initialize an MCP session, then call
list_clustersto discover endpoints.
11 Tools
| Tool | Type | Purpose |
|---|---|---|
list_clusters |
discovery | Return all active clusters (LLM entry point) |
spark_submit |
exec | Local spark-submit process, extracts app_id |
list_applications |
RM | GET /ws/v1/cluster/apps[?state=&user=] |
get_application_status |
RM | GET /ws/v1/cluster/apps/{id} |
get_application_logs |
RM | Fallback chain: amContainerLogs -> aggregated-logs -> logs |
kill_application |
RM | PUT state KILLED |
fetch_spark_metrics |
SHS | Executor metrics + summary mode triggers analyzer |
fetch_cluster_env |
RM | Aggregate /cluster/info and /cluster/metrics |
analyze_spark_log |
analyzer | RM logs + heuristic rules + LLM-ready prompt |
fetch_url |
HTTP | Generic primitive, uses cluster allowlist and auth |
upload_file |
FS | Write to ./data/uploads/, reference from spark_submit |
Environment variables
| Variable | Default | Required | Description |
|---|---|---|---|
LISTEN_ADDR |
:8080 |
no | HTTP listen address |
DATA_DIR |
./data |
no | SQLite, uploads, and log root |
ADMIN_TOKENS |
— | yes | Comma-separated tokens for /admin/* |
AGENT_TOKEN |
— | yes | Token for /mcp requests |
HTTP_CLIENT_TIMEOUT |
30s |
no | Timeout for fetch_url/RM Tool HTTP calls |
MAX_RESPONSE_BYTES |
1048576 (1 MiB) |
no | HTTP response truncation limit |
SPARK_SUBMIT_TIMEOUT |
60s |
no | spark-submit process timeout |
LOG_DIR |
./data/logs |
no | Log root directory |
LOG_LEVEL |
info |
no | debug, info, warn, or error |
LOG_FORMAT |
text |
no | text for terminal, json for log files |
ANALYZER_DATA_SKEW_RATIO |
3.0 |
no | Data-skew rule threshold (max/min ratio) |
ANALYZER_GC_PRESSURE_RATIO |
0.1 |
no | GC-pressure rule threshold (GC/CPU ratio) |
ANALYZER_BOTTLENECK_SHUFFLE_GB |
50 |
no | Bottleneck rule threshold (shuffle GB) |
Admin API
All endpoints require Authorization: Bearer <admin_token>.
| Method | Path | Description |
|---|---|---|
GET |
/admin/clusters |
List all clusters |
POST |
/admin/clusters |
Create a cluster |
GET |
/admin/clusters/:id |
Get one cluster |
PUT |
/admin/clusters/:id |
Update a cluster |
DELETE |
/admin/clusters/:id |
Delete a cluster |
GET |
/admin/audit?limit=100 |
Query audit log |
Create a cluster:
curl -sS -X POST \\
-H "Authorization: Bearer ${ADMIN_TOKEN}" \\
-H "Content-Type: application/json" \\
-d '{
"id": "prod",
"name": "prod",
"rm_url": "http://rm.example.com:8088",
"shs_url": "http://shs.example.com:18080",
"spark_submit_execute_bin": "/opt/spark/bin/spark-submit",
"is_active": true,
"auth_type": "simple",
"auth_username": "yarn",
"rate_limit_per_min": 10,
"url_allowlist": ["rm.example.com:8088", "shs.example.com:18080"]
}' \\
"http://127.0.0.1:${LISTEN_ADDR:-:8080}/admin/clusters"
Note: auth_password is not accepted via JSON (json:"-"). Set it through a
dedicated password endpoint or seed the database directly.
Security model
Three lines of defense:
- Token auth -
ADMIN_TOKENSprotects/admin/*;AGENT_TOKENprotects/mcp. - SSRF guard - Every outbound URL must match the cluster's
url_allowlist, with private-IP CIDR blacklist (incl.169.254.0.0/16AWS/GCP metadata). - Path allowlist -
upload_filewrites only to./data/uploads/, rejects absolute paths,.., and non-[a-zA-Z0-9._-]filenames.
Additional guarantees:
spark_submitruns the cluster's binary viaexec.Command(name, args...)— slice form, neversh -c. No shell metacharacter interpretation.auth_passwordis taggedjson:"-"; never serialized in responses, never accepted from JSON request bodies.- Cross-host redirects (RM → NM 307) preserve the
Authorizationheader throughDoWithRedirectonly when the destination host is inurl_allowlist. - Per-Tool log files are created with mode
0600and live under./data/logs/tools/.
Development
go build ./...
go test ./...
go vet ./...
More docs
- ARCHITECTURE.md - Design and package layout
- docs/runbook.md - Deployment, upgrade, and troubleshooting
Releases
2
Release v1.0.0
Latest
Languages
Go
86.5%
HTML
12.6%
Shell
0.9%