refactor: unify env-var config in common/config.py

Before: SPARK_EXECUTOR_DATA_DIR, SPARK_EXECUTOR_JOBS_DIR, and
YARN_RESOURCE_MANAGER_URL were each read directly via os.environ.get()
inside the module that used them. Log level was hardcoded DEBUG in
common/logging.py. There was no single file showing what the full set
of env vars the app reads is.

After: common/config.py defines a single Settings dataclass that
reads all env vars at import time and exposes them as fields on a
module-level singleton. App code uses "from common.config import
settings; settings.data_dir" etc. New SPARK_EXECUTOR_LOG_LEVEL env
var controls stderr + info file verbosity (debug file always gets full
DEBUG).

Improvements:
  - One file lists every env var the app reads (was: grep the codebase)
  - Tests can monkeypatch fields on the settings singleton directly
    instead of monkeypatching the env + reloading
  - Adding a new env var means adding one field in config.py, not
    editing 3+ call sites
  - settings.reload() method for tests that prefer env-var style

Out of scope (kept where they are):
  - GUNICORN_* env vars live in gunicorn.conf.py (gunicorn concept)
  - PYTHONUNBUFFERED in Dockerfile (Python runtime flag)
  - SPARK_SUBMIT_OPTS not in config (JVM flag, not Python)

Test changes:
  - test_job_writer.py: settings.jobs_dir instead of monkeypatching
    SPARK_EXECUTOR_JOBS_DIR
  - test_yarn_client.py: settings.yarn_resource_manager_url instead of
    monkeypatching YARN_RESOURCE_MANAGER_URL
  - test_generate_tool.py: same as job_writer
  - Each test file gets an autouse fixture that snapshots+restores
    settings so one test mutation does not leak into the next

116/116 still pass. Live verified: SPARK_EXECUTOR_LOG_LEVEL=INFO
suppresses DEBUG loguru output as expected.
This commit is contained in:
Claude
2026-06-25 10:52:53 +08:00
parent 3ec9ea37fd
commit f8e367d43c
9 changed files with 204 additions and 48 deletions
+7 -6
View File
@@ -16,20 +16,21 @@ import os
import secrets
from datetime import datetime
from common.config import settings
from common.logging import logger
DEFAULT_JOBS_DIR = "./data/jobs"
ENV_JOBS_DIR = "SPARK_EXECUTOR_JOBS_DIR"
def resolve_jobs_dir(jobs_dir: str | None = None) -> str:
"""Pick the effective jobs directory in priority order: arg > env > default."""
"""Pick the effective jobs directory in priority order: arg > settings.jobs_dir.
`settings.jobs_dir` itself is computed in `common/config.py` as:
SPARK_EXECUTOR_JOBS_DIR env var (if set) > <data_dir>/jobs
"""
if jobs_dir is not None:
return jobs_dir
env = os.environ.get(ENV_JOBS_DIR)
if env:
return env
return DEFAULT_JOBS_DIR
return settings.jobs_dir
def write_job_file(code: str, jobs_dir: str | None = None) -> str: