7920dc9597a7ccb0a4db830fdbed3c2eb7bd4217
将 spark_submit 的自由 args[] 替换为与 Python 服务一致的 schema:cluster_id、 master、deploy_mode、script_path、queue、executor_memory、executor_cores、 num_executors、spark_conf、extra_args。命令行顺序严格对齐 Python 实现: binary → --master → --queue → --executor-memory → --executor-cores → --num-executors → 按 key 排序的 --conf key=value → 按 key 排序的 --flag value → script_path 永远在最后。 deploy_mode 是 Python 参考函数的入参,但该函数实际并不输出 --deploy-mode; Go 端保留这个字段并校验其值,但同样不发射它,以保持行为一致。如果后续需要 --deploy-mode,再在此处添加并记录差异。 cluster.DefaultSubmitArgs 不再被拼接到 argv 中;该字段仍保留用于 admin API 向后兼容,运行时移除它的工作是后续独立的 follow-up。 删除了 translateArgs、pathLikeArg 以及 spark_submit.unminted_path 警告逻辑 (本次提交取代 2570713),因为自由参数已不存在,警告无从触发。新增 parseStringMap、sortedStringKeys 和 requireNonNegativeInt 辅助函数; requireNonNegativeInt 同时接受 int 与 float64,以兼容测试直接构造的整数 字面量和 JSON 解码后的 float64。 测试更新: - TestSparkSubmit_StructuredCommand:验证完整 argv 顺序,并断言 script_path 是最后一项。 - TestSparkSubmit_MissingRequiredField:校验缺少 master 时返回错误。 - TestSparkSubmit_BadSparkConfValue:校验 spark_conf 的值不是字符串时报错。 Co-Authored-By: Claude <noreply@anthropic.com>
spark-mcp-go
MCP (Model Context Protocol) server for Apache Spark on YARN. Lets an LLM agent
discover Spark/YARN endpoints, submit jobs, fetch logs, and analyze them via
11 typed Tools. Streamable HTTP transport, SQLite-backed configuration, and
log/slog structured logging with per-Tool call files.
What it does
- Exposes 11 MCP Tools over Streamable HTTP at
/mcp - Admin API at
/admin/*for cluster and audit configuration - HTTP Basic and YARN SimpleAuth for Spark/YARN endpoints
- SSRF protection with cluster URL allowlist plus DNS rebinding guard
- Per-Tool independent audit log (file mode
0600) spark-submitvia localexec.Command(no shell, no injection)
Quick start (3 steps)
-
Copy and edit the environment file:
cp .env.example .env # Edit ADMIN_TOKENS and AGENT_TOKEN -
Build and start the server:
go build ./... ./spark-mcp-go # Or use the helper: # ./scripts/dev.sh -
Initialize an MCP session, then call
list_clustersto discover endpoints.
11 Tools
| Tool | Type | Purpose |
|---|---|---|
list_clusters |
discovery | Return all active clusters (LLM entry point) |
spark_submit |
exec | Local spark-submit process, extracts app_id |
list_applications |
RM | GET /ws/v1/cluster/apps[?state=&user=] |
get_application_status |
RM | GET /ws/v1/cluster/apps/{id} |
get_application_logs |
RM | Fallback chain: amContainerLogs -> aggregated-logs -> logs |
kill_application |
RM | PUT state KILLED |
fetch_spark_metrics |
SHS | Executor metrics + summary mode triggers analyzer |
fetch_cluster_env |
RM | Aggregate /cluster/info and /cluster/metrics |
analyze_spark_log |
analyzer | RM logs + heuristic rules + LLM-ready prompt |
fetch_url |
HTTP | Generic primitive, uses cluster allowlist and auth |
upload_file |
FS | Write to ./data/uploads/, reference from spark_submit |
Environment variables
| Variable | Default | Required | Description |
|---|---|---|---|
LISTEN_ADDR |
:8080 |
no | HTTP listen address |
DATA_DIR |
./data |
no | SQLite, uploads, and log root |
ADMIN_TOKENS |
— | yes | Comma-separated tokens for /admin/* |
AGENT_TOKEN |
— | yes | Token for /mcp requests |
HTTP_CLIENT_TIMEOUT |
30s |
no | Timeout for fetch_url/RM Tool HTTP calls |
MAX_RESPONSE_BYTES |
1048576 (1 MiB) |
no | HTTP response truncation limit |
SPARK_SUBMIT_TIMEOUT |
60s |
no | spark-submit process timeout |
LOG_DIR |
./data/logs |
no | Log root directory |
LOG_LEVEL |
info |
no | debug, info, warn, or error |
LOG_FORMAT |
text |
no | text for terminal, json for log files |
ANALYZER_DATA_SKEW_RATIO |
3.0 |
no | Data-skew rule threshold (max/min ratio) |
ANALYZER_GC_PRESSURE_RATIO |
0.1 |
no | GC-pressure rule threshold (GC/CPU ratio) |
ANALYZER_BOTTLENECK_SHUFFLE_GB |
50 |
no | Bottleneck rule threshold (shuffle GB) |
Admin API
All endpoints require Authorization: Bearer <admin_token>.
| Method | Path | Description |
|---|---|---|
GET |
/admin/clusters |
List all clusters |
POST |
/admin/clusters |
Create a cluster |
GET |
/admin/clusters/:id |
Get one cluster |
PUT |
/admin/clusters/:id |
Update a cluster |
DELETE |
/admin/clusters/:id |
Delete a cluster |
GET |
/admin/audit?limit=100 |
Query audit log |
Create a cluster:
curl -sS -X POST \\
-H "Authorization: Bearer ${ADMIN_TOKEN}" \\
-H "Content-Type: application/json" \\
-d '{
"id": "prod",
"name": "prod",
"rm_url": "http://rm.example.com:8088",
"shs_url": "http://shs.example.com:18080",
"spark_submit_execute_bin": "/opt/spark/bin/spark-submit",
"is_active": true,
"auth_type": "simple",
"auth_username": "yarn",
"rate_limit_per_min": 10,
"url_allowlist": ["rm.example.com:8088", "shs.example.com:18080"]
}' \\
"http://127.0.0.1:${LISTEN_ADDR:-:8080}/admin/clusters"
Note: auth_password is not accepted via JSON (json:"-"). Set it through a
dedicated password endpoint or seed the database directly.
Security model
Three lines of defense:
- Token auth -
ADMIN_TOKENSprotects/admin/*;AGENT_TOKENprotects/mcp. - SSRF guard - Every outbound URL must match the cluster's
url_allowlist, with private-IP CIDR blacklist (incl.169.254.0.0/16AWS/GCP metadata). - Path allowlist -
upload_filewrites only to./data/uploads/, rejects absolute paths,.., and non-[a-zA-Z0-9._-]filenames.
Additional guarantees:
spark_submitruns the cluster's binary viaexec.Command(name, args...)— slice form, neversh -c. No shell metacharacter interpretation.auth_passwordis taggedjson:"-"; never serialized in responses, never accepted from JSON request bodies.- Cross-host redirects (RM → NM 307) preserve the
Authorizationheader throughDoWithRedirectonly when the destination host is inurl_allowlist. - Per-Tool log files are created with mode
0600and live under./data/logs/tools/.
Development
go build ./...
go test ./...
go vet ./...
More docs
- ARCHITECTURE.md - Design and package layout
- docs/runbook.md - Deployment, upgrade, and troubleshooting
Releases
2
Release v1.0.0
Latest
Languages
Go
86.5%
HTML
12.6%
Shell
0.9%