feat: add list_applications tool for YARN application enumeration

Add a new MCP tool that queries YARN's /ws/v1/cluster/apps endpoint
through a named Connection, returning a list of ApplicationSummary
records. Bypasses the local JobStore — useful for enumerating apps
that were not submitted through this service.

API:
  list_applications(
    connection_name: str,           # required, which YARN cluster
    state: str | None = None,       # YARN state filter: NEW/NEW_SAVING/
                                    # SUBMITTED/ACCEPTED/RUNNING/
                                    # FINISHED/FAILED/KILLED
    queue: str | None = None,       # YARN queue filter
    limit: int = 100,               # cap on returned apps (YARN has no
                                    # offset-based pagination; combine
                                    # state/queue filters for big clusters)
  ) -> list[ApplicationSummary]

Implementation:
  - yarn_client.list_applications(config, *, state, queue, limit) -> list[dict]
    Returns raw YARN app dicts; raises YarnError on 4xx/5xx; returns
    [] on 404 (no apps match). Uses the existing _request helper,
    which now accepts a "params" kwarg for query strings (one-line
    additive change).
  - external_jobs.list_applications(connection_name, state, queue, limit)
    -> list[ApplicationSummary]. Looks up the Connection, builds the
    YarnClientConfig, calls the yarn_client function, maps each raw
    YARN dict to ApplicationSummary (mirroring the manual field-mapping
    style of get_job_result). The yarn_client function is imported
    as "list_applications_yarn" to avoid name collision.
  - ApplicationSummary: 12-field Pydantic model with snake_case names
    (application_id, name, user, queue, state, final_status,
    application_type, application_tags, started_time, finished_time,
    tracking_url, progress). Unused YARN fields (memorySeconds,
    vcoreSeconds, preemptedResource*, etc.) are not exposed.
  - ListApplicationsRequest: Pydantic body model with Field(description=)
    for LLM-facing schema.
  - /list_applications route registered with operation_id=
    "list_applications", placed next to the other external YARN tools.

Tests:
  - 8 new unit tests in test_external_jobs.py (happy path, state/queue/
    limit pass-through, default limit, empty list, missing connection,
    full field mapping).
  - test_mcp_routes.py: assert 23 tool routes.
  - README: list_applications row added to the Spark Executor table.

Tests: 390 passed (was 382, +8 net).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Claude
2026-07-09 13:40:20 +08:00
co-authored by Claude Fable 5
parent a5b9539663
commit 7fbad97a87
8 changed files with 363 additions and 7 deletions
+57 -2
View File
@@ -10,8 +10,13 @@ YARN application_id and the name of a saved Connection.
import json
from common.logging import logger
from spark_executor.core.yarn_client import YarnClientConfig, get_application_status, get_application_logs
from spark_executor.models import JobStatus, JobResult
from spark_executor.core.yarn_client import (
YarnClientConfig,
get_application_logs,
get_application_status,
list_applications as list_applications_yarn, # alias to avoid collision
)
from spark_executor.models import ApplicationSummary, JobResult, JobStatus
from spark_executor.tools.connections import store as conn_store
@@ -63,3 +68,53 @@ def get_external_job_result(application_id: str, connection_name: str) -> JobRes
)
logger.info(f"get_external_job_result ok application_id={application_id} connection_name={connection_name} state={state}")
return result
def list_applications(
connection_name: str,
state: str | None = None,
queue: str | None = None,
limit: int = 100,
) -> list[ApplicationSummary]:
"""List YARN applications on the named cluster, optionally filtered.
Bypasses the local JobStore (this is for apps not submitted through
this service). The YARN ResourceManager REST endpoint
/ws/v1/cluster/apps is queried through the connection's auth/SSL
config.
Defaults: limit=100 (YARN has no offset-based pagination, so large
clusters should use state/queue filters to scope the result).
"""
logger.debug(
f"list_applications enter connection_name={connection_name} "
f"state={state} queue={queue} limit={limit}"
)
conn = conn_store.get(connection_name)
if conn is None:
raise KeyError(f"Connection not found: {connection_name}")
config = YarnClientConfig.from_connection(conn)
raw_apps = list_applications_yarn(
config, state=state, queue=queue, limit=limit
)
summaries = [
ApplicationSummary(
application_id=app.get("id", ""),
name=app.get("name", ""),
user=app.get("user", ""),
queue=app.get("queue", ""),
state=app.get("state", ""),
final_status=app.get("finalStatus"),
application_type=app.get("applicationType"),
application_tags=app.get("applicationTags", ""),
started_time=app.get("startedTime", 0),
finished_time=app.get("finishedTime", 0),
tracking_url=app.get("trackingUrl"),
progress=app.get("progress"),
)
for app in raw_apps
]
logger.info(
f"list_applications ok connection_name={connection_name} count={len(summaries)}"
)
return summaries
+36
View File
@@ -475,3 +475,39 @@ class UpdateConnectionRequest(BaseModel):
"Example: ['ccam*'] allows ccam1-ccam99."
),
)
class ListApplicationsRequest(BaseModel):
connection_name: str = Field(
...,
description=(
"Name of a saved Connection (see list_connections) pointing at "
"the YARN cluster to query."
),
)
state: str | None = Field(
default=None,
description=(
"Optional YARN application state filter. One of: 'NEW', "
"'NEW_SAVING', 'SUBMITTED', 'ACCEPTED', 'RUNNING', 'FINISHED', "
"'FAILED', 'KILLED'. 'FINISHED' is the umbrella state covering "
"SUCCEEDED/FAILED/KILLED. None = no state filter (returns all "
"states up to `limit`)."
),
)
queue: str | None = Field(
default=None,
description=(
"Optional YARN queue name filter (e.g. 'default', 'prod'). "
"None = no queue filter."
),
)
limit: int = Field(
default=100,
description=(
"Maximum number of applications to return. YARN has no "
"offset-based pagination, so for large clusters use state/queue "
"filters to scope the result. Max 10000 in practice (YARN's own "
"limit on the limit param)."
),
)