Files
mcp-server/.env.example
T
Claude 13e39b39f3 refactor: env-configurable loguru file paths
common/logging.py used hardcoded 'data/logs/{debug,info}/' paths,
which forced logs into the same dir as connections.json. Operators
couldn't put logs on a dedicated volume or a different disk for
retention policy reasons.

Now the base log dir comes from settings.log_dir (env var
SPARK_EXECUTOR_LOG_DIR, default <data_dir>/logs). Subdirs debug/ and
info/ are auto-created. Same behavior in dev (./data/logs/), but in
prod you can do:

  SPARK_EXECUTOR_LOG_DIR=/var/log/spark-executor   # dedicated volume
  SPARK_EXECUTOR_LOG_DIR=/mnt/slow-storage/logs   # cold storage for old logs

Other changes:
  - .env.example: new SPARK_EXECUTOR_LOG_DIR entry, split into its own
    'Loguru file sinks' section
  - docker-compose.yml: forwards SPARK_EXECUTOR_LOG_DIR with the
    container-side default /app/data/logs (the existing ./data:/app/data
    volume mount already covers this)
  - common/config.py: Settings gets a new log_dir field, defaults
    derived from data_dir; reload() also resets it

116/116 still pass. Live smoke verified: with
SPARK_EXECUTOR_LOG_DIR=/tmp/spark-logs, both
  /tmp/spark-logs/info/2026-06-25.log
  /tmp/spark-logs/debug/2026-06-25.log
are created on the first request.
2026-06-25 10:57:50 +08:00

65 lines
2.5 KiB
Bash

# Spark Executor MCP — environment template
# Copy to .env and edit. .env is gitignored.
#
# All env vars in this file are read by common/config.py (the single source
# of truth for application config). The only exception is the GUNICORN_*
# block at the bottom — those are read by gunicorn.conf.py.
# --- Data persistence ---
# Base directory for connections.json, pending_jobs.json, loguru logs/,
# and (by default) jobs/. Mount this from the host in production so
# state survives container restarts. The default ./data/ is fine in dev.
#
# SPARK_EXECUTOR_DATA_DIR=./data
# SPARK_EXECUTOR_DATA_DIR=/var/lib/spark-executor/data
# --- Job files (LLM-generated PySpark) ---
# Where generate_job_file writes PySpark source. Defaults to
# <SPARK_EXECUTOR_DATA_DIR>/jobs. Override to point at a larger disk
# (e.g. /var/spark-jobs) when the data volume is small.
#
# SPARK_EXECUTOR_JOBS_DIR=./data/jobs
# SPARK_EXECUTOR_JOBS_DIR=/var/spark-jobs
# --- YARN REST client ---
# Fallback URL when a Job's yarn_rm_url (snapshotted from its Connection
# at prepare_submit_job time) is unset. Set this OR per-Connection via
# save_connection.
#
# Examples:
# YARN_RESOURCE_MANAGER_URL=http://yarn-rm.prod.internal:8088
# YARN_RESOURCE_MANAGER_URL=https://yarn-rm.staging.example.com:8088
YARN_RESOURCE_MANAGER_URL=
# --- Loguru file sinks ---
# Base directory for loguru output. Subdirs debug/ and info/ are created
# automatically; rotated daily, gzipped, kept 30 days. Defaults to
# <SPARK_EXECUTOR_DATA_DIR>/logs. Override to point at a dedicated log
# volume (e.g. /var/log/spark-executor) or a network mount.
#
# SPARK_EXECUTOR_LOG_DIR=./data/logs
# SPARK_EXECUTOR_LOG_DIR=/var/log/spark-executor
# --- Loguru verbosity ---
# For stderr + the info-level file sink. The debug-level file sink
# always captures full DEBUG (audit trail regardless of level).
# DEBUG - default; full verbosity
# INFO - quieter; recommended for production
#
# SPARK_EXECUTOR_LOG_LEVEL=DEBUG
# SPARK_EXECUTOR_LOG_LEVEL=INFO
# --- Optional: JVM flags forwarded to spark-submit ---
# Useful for proxies, custom truststores, or driver memory caps.
# SPARK_SUBMIT_OPTS=-Dhttps.proxyHost=proxy.corp -Dhttps.proxyPort=3128
# --- Gunicorn process model (see gunicorn.conf.py; NOT read by common/config.py) ---
# Defaults shown. These are read by gunicorn directly, not by the app.
# GUNICORN_WORKERS=2
# GUNICORN_THREADS=1
# GUNICORN_TIMEOUT=120 # generous; yarn logs can be slow
# GUNICORN_GRACEFUL_TIMEOUT=30
# GUNICORN_KEEPALIVE=5
# GUNICORN_BIND=0.0.0.0:8000
# GUNICORN_LOGLEVEL=info