Before: SPARK_EXECUTOR_DATA_DIR, SPARK_EXECUTOR_JOBS_DIR, and
YARN_RESOURCE_MANAGER_URL were each read directly via os.environ.get()
inside the module that used them. Log level was hardcoded DEBUG in
common/logging.py. There was no single file showing what the full set
of env vars the app reads is.
After: common/config.py defines a single Settings dataclass that
reads all env vars at import time and exposes them as fields on a
module-level singleton. App code uses "from common.config import
settings; settings.data_dir" etc. New SPARK_EXECUTOR_LOG_LEVEL env
var controls stderr + info file verbosity (debug file always gets full
DEBUG).
Improvements:
- One file lists every env var the app reads (was: grep the codebase)
- Tests can monkeypatch fields on the settings singleton directly
instead of monkeypatching the env + reloading
- Adding a new env var means adding one field in config.py, not
editing 3+ call sites
- settings.reload() method for tests that prefer env-var style
Out of scope (kept where they are):
- GUNICORN_* env vars live in gunicorn.conf.py (gunicorn concept)
- PYTHONUNBUFFERED in Dockerfile (Python runtime flag)
- SPARK_SUBMIT_OPTS not in config (JVM flag, not Python)
Test changes:
- test_job_writer.py: settings.jobs_dir instead of monkeypatching
SPARK_EXECUTOR_JOBS_DIR
- test_yarn_client.py: settings.yarn_resource_manager_url instead of
monkeypatching YARN_RESOURCE_MANAGER_URL
- test_generate_tool.py: same as job_writer
- Each test file gets an autouse fixture that snapshots+restores
settings so one test mutation does not leak into the next
116/116 still pass. Live verified: SPARK_EXECUTOR_LOG_LEVEL=INFO
suppresses DEBUG loguru output as expected.
46 lines
1.5 KiB
Python
46 lines
1.5 KiB
Python
# coding=utf-8
|
|
import os
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
from common import config
|
|
from spark_executor.tools import generate
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _restore_settings():
|
|
snapshot = config.Settings(
|
|
data_dir=config.settings.data_dir,
|
|
jobs_dir=config.settings.jobs_dir,
|
|
yarn_resource_manager_url=config.settings.yarn_resource_manager_url,
|
|
log_level=config.settings.log_level,
|
|
)
|
|
yield
|
|
config.settings.data_dir = snapshot.data_dir
|
|
config.settings.jobs_dir = snapshot.jobs_dir
|
|
config.settings.yarn_resource_manager_url = snapshot.yarn_resource_manager_url
|
|
config.settings.log_level = snapshot.log_level
|
|
|
|
|
|
def test_generate_writes_code_and_returns_path(monkeypatch, tmp_path: Path):
|
|
config.settings.jobs_dir = str(tmp_path / "data" / "jobs")
|
|
out = generate.generate_job_file(
|
|
"from pyspark.sql import SparkSession\n"
|
|
"spark = SparkSession.builder.getOrCreate()\n"
|
|
)
|
|
assert "script_path" in out
|
|
p = out["script_path"]
|
|
assert os.path.isabs(p)
|
|
assert p.startswith(str(tmp_path / "data" / "jobs"))
|
|
assert p.endswith(".py")
|
|
with open(p) as f:
|
|
assert "SparkSession.builder.getOrCreate()" in f.read()
|
|
|
|
|
|
def test_generate_uses_settings_jobs_dir(tmp_path: Path):
|
|
config.settings.jobs_dir = str(tmp_path / "custom")
|
|
out = generate.generate_job_file("x = 1\n")
|
|
assert out["script_path"].startswith(str(tmp_path / "custom"))
|
|
assert os.path.isfile(out["script_path"])
|