Files
model-platform/DEVELOP.md
T
tao.chenandClaude b493907775 feat(scripts): 跨 owner 懒加载目录树 + 跨用户可见 workspace/public
修两个后端接口问题:
1) /api/v1/workspace-directories 返回为空,目录树结构消失
2) 同 workspace 内脚本/数据互相可见但默认排除 private

后端改动
--------
* list_scripts / list_resources / list_workspace_directories 新增
  owner_user_id 可选 query 参数;缺省 = 当前请求者本人(scope 到
  workspace/{me}/...),传值时 scope 到该 owner 的子树。前端根加载
  默认只见自己一级,其他成员以折叠分组呈现。
* visibility 过滤统一:非 admin 请求者只返回 owner==me 或
  visibility ∈ {workspace, public};admin 跳过。owner=me 含自己
  的 private,owner=other 只剩其 workspace/public,排除他人 private。
* create_workspace_directory 两个分支 visibility 默认 'public'
  (非 private),使跨 owner 目录树可见;响应新增 owner_user_id 字段。
* platform.list_members 鉴权从 system_admin_context 放宽为
  系统管理员或该 workspace 活跃成员(让普通用户也能渲染同
  workspace 成员名册,用于跨 owner 分组)。
* main.py 注册 platform 模块(随 list_members 改动补齐导入)。
* .env.example 同步 common/config.py 26 个字段。

前端改动
--------
* ScriptExplorer.memberScriptGroups 改由 members 列表播种分组,
  display_name 取 members.display_name;inferredDirectories 现在按
  owner_user_id 标记,统一跨 owner 目录渲染。删除脚本目录页头与
  树分组标题的工作副本数量角标。
* WorkspaceTree 新增 ownerUserId 透传到 store.toggleExpanded;
  仅"我"的分组 mount 时 auto-expand,他人分组默认折叠,展开才
  调 loadOwnerGroup / owner-scoped loadScripts / loadChildren。
* scriptWorkspaceStore 引入 namespaced cache key
  (ownerCacheKey = `${ownerUserId ?? me}:${path}`),loadedScriptPaths
  / loadedChildPaths / loadedOwnerGroups 全部按 owner 隔离;
  toggleExpanded 用 loadPath === undefined 区分 group 头与真实
  目录,修"他人子目录点击不触发接口"的 loadPath 前缀误判 bug。
* api.ts / AuthContext 透传 ownerUserId 给 listScripts /
  listResources / listWorkspaceDirectories。

文档
----
* API.md: §3.2 创建目录 visibility 默认 public + 响应加 owner_user_id;
  §3.3.1 GET directories 加 owner_user_id 参数 + 响应字段;
  §3.4 GET scripts 改写为 owner 作用域 + visibility 过滤语义;
  §五.1 GET data-resources 新增,同一套统一语义;
  §7 intro 例外 — GET members 对系统管理员或 workspace 活跃成员开放。
* DEVELOP.md: Code layout 重写以反映 backend api/services/clients/
  schemas 拆分 + schedule domain/scheduling/application/execution/
  infrastructure 拆分 + common 子包(auth/storage/backends);
  Configuration 系统补全 26 个 settings 字段;新增
  "Owner-scoping + visibility (cross-owner browsing)" 小节;
  Per-service dev 注释用 uv run 的源布局要求;Add a new DAG endpoint /
  storage bucket 路径改为 backend/src/backend/api/* 与 services/*。

测试
----
* test_list_scripts_parent_path.py /
  test_resources.py 补充 owner_user_id 参数化直接调用 + LIKE
  前缀断言(workspace/{owner}/... 前缀)。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-21 19:26:53 +08:00

22 KiB

DEVELOP.md — Developer Guide

This guide is for engineers working on the model platform codebase. For high-level design see ARCHITECTURE.md; for the current state of in-flight refactors see HANDOVER.md.

Code layout

All Python packages use the src/<pkg>/ layout; uv workspace glues them into one .venv. Always invoke via uv run [--package <pkg>] <cmd> (see "Local development" for the gotcha).

common/src/common/                Pure-Python shared library
  config.py                       Settings (pydantic-settings, lru_cache singleton)
  db/                             SQLAlchemy 2.0 async engine, session_scope, Base
  db/models/                      26 tables in 9 domain files (zero FK, zero relationship)
  auth/                           JWT / bcrypt / workspace membership helpers
  scheduler/                      APScheduler trigger helpers (delayed import)
  storage/                        AsyncStorageBackend abstraction + Pydantic schemas
    base.py                       Abstract interface
    factory.py                    create_storage + build_storage_config + PURPOSE_BUCKETS
    schemas.py                    CreateUploadRequest / ServerObjectRequest
    backends/local.py            Local filesystem impl
    backends/s3.py                S3-compatible impl (boto3)
    registry.py                   Bucket registry
  eventing.py                     add_outbox_event / utcnow / event_time
  service_app.py                  /health/ready TCP probe, /api/v1/health
  logging.py                      loguru config (LOG_LEVEL)
  schemas.py                      StrictModel base
  ids.py                          ULID generation helpers
  utils.py                        get_free_port, start_process

backend/src/backend/              Public FastAPI service
  main.py                         lifespan + route registration
  audit.py                        HTTP access log middleware (loguru sink)
  api/                            HTTP route handlers (one module per bounded context)
    auth.py                       /api/v1/auth/* (login / me / jupyter)
    jupyter.py                    /api/v1/auth/jupyter — the ONLY auth entry
    dependencies.py               request_context, database_session
    platform.py                   /api/v1/platform/* (system admin)
    admin.py                      /api/v1/admin/* (workspace-internal admin)
    scripts.py                    /api/v1/scripts/* + /api/v1/workspace-directories
    resources.py                  /api/v1/data-resources/*
    schedules/schedules.py        DAG CRUD
    schedules/runs.py             Run lifecycle
    storage.py                    /internal/v1/objects — single token-guarded endpoint (P0-1)
  services/                       Pure-Python business logic (no HTTP / no DI)
    scripts.py                    create_workspace_directory, visibility-filtered queries
    resources.py                  owner-scoped resource listing helpers
    jupyter.py                    jupyter_path / lock helpers
    storage.py                    object store helpers
    schedules.py                  DAG validation (cycle / orphan detection)
  schemas/                        Pydantic request / response models
    auth.py / common.py / jupyter.py / platform.py / resources.py / schedules.py / scripts.py
  clients/                        Outbound HTTP / RPC clients
    runtime.py                    Self-contained httpx wrapper for the runtime
    scheduler.py                  Backend → Schedule HTTP client (callback / dispatch)
    rclone.py                     rclone RC API client (FUSE cache invalidation)

schedule/src/schedule/            Schedule Executor (DAG worker)
  main.py                         Lifespan + FastAPI app
  notebook_runner.py              Subprocess entry point (nbclient) — DO NOT RENAME
  domain/                         Pure-Python domain types
    execution.py                  ExecutionResult (frozen dataclass) + state enums
    context.py                    Constants + naive_utc
  scheduling/                     Time-based trigger
    scheduler.py                  CronScheduler (APScheduler + 5s sync loop)
  application/                    Facades / orchestrators
    service.py                    SchedulerService (composes the three)
    orchestrator.py               DispatchOrchestrator (Outbox poll + DAG advance)
  execution/                      DAG node execution
    executor.py                   NodeExecutor (notebook / python dispatch)
    worker.py                     asyncio entry, schedule-spawned task boundary
    runners/notebook.py           nbclient subprocess path (6-line `notebook_runner` shim re-exports `main`)
  infrastructure/                 External-system adapters
    storage/client.py             SchedulerStorageClient — talks to backend /internal/v1/objects

runtime/src/runtime/              Jupyter Runtime
  main.py                         FastAPI entry: jupyter action endpoints
  process.py                      Per-workspace subprocess pool + asyncio locks
  mount.py                        rclone FUSE mount lifecycle

frontend/                         React Router SPA (vite build → nginx)
  app/                            features/ routes/ services/ components/

migrations/                       Alembic schema versions
docker-compose.yml                4 services (gateway / backend / schedule / runtime)
default.conf                      Nginx template
scripts/nginx-entrypoint.sh
.env.example                      All 26 config.py keys documented

Configuration system

All env vars go through one place: common/src/common/config.py.

from common.config import settings

# Auth / runtime
settings.database_url          # str — SQLAlchemy async URL (mysql+asyncmy, charset utf8mb4)
settings.jwt_secret             # HS256 secret for the auth_request handler
settings.cookie_force_secure    # bool — write Secure flag even on plain HTTP (TLS-terminating proxy)
settings.service_name          # surfaced in /health
settings.schedule_event_namespace  # APScheduler JobStore namespace + Outbox scope prefix
settings.readiness_targets     # CSV host:port list for /health/ready

# HTTP clients (intra-cluster URLs)
settings.runtime_api_url       # backend → runtime HTTP base
settings.public_base_url       # runtime public base URL (browser-facing /jupyter/)
settings.backend_api_url       # schedule → backend HTTP base
settings.rclone_rc_url         # backend → rclone RC control API
settings.internal_service_token  # Backend ↔ Schedule shared secret (X-Internal-Service-Token)

# Logging / audit
settings.log_level             # DEBUG / INFO / WARNING / ERROR / CRITICAL (lowercase → fallback INFO)
settings.audit_log_dir         # dir for daily audit logs (relative to cwd; "" disables file sink)
settings.audit_log_retention_days  # 0 disables cleanup
settings.audit_excluded_paths  # list[str] — paths skipped from audit (health probes, etc.)

# Storage
settings.storage_backend       # "s3" (default) or "local"
settings.local_storage_base_dir  # root dir for storage data (default "/data")
settings.s3_endpoint            # str (s3 mode only)
settings.s3_access_key          # str (s3 mode only)
settings.s3_secret_key          # str (s3 mode only)
settings.s3_workspace_bucket    # str (s3 mode only)
settings.s3_version_bucket      # str (s3 mode only)
settings.s3_run_log_bucket      # str (s3 mode only)
settings.s3_trash_bucket        # str (s3 mode only)
settings.s3_trash_retention_days  # int (s3 mode only)

# Schedule
settings.schedule_execution_concurrency  # int — max concurrent notebook subprocesses

Settings reads from process env first, then from a .env file at CWD if present. pydantic-settings auto-loads. case_sensitive=False so DATABASE_URL / database_url both work. The full list of 26 fields is in common/src/common/config.py.

Adding a new env var

  1. Add the field to Settings:
    new_var: str = Field(default="x", description="...")
    
  2. Add the line to .env.example with a comment (keep it synced — every field in config.py must have a matching .env.example entry).
  3. Use settings.new_var at the call site.

Do not call os.environ["NEW_VAR"] or os.getenv("NEW_VAR") in application code. The grep below should return zero hits:

grep -rnE 'os\.(environ\[?["\x27][A-Z_]+|getenv\(["\x27][A-Z_]+)' --include="*.py" \
  backend/src/ schedule/src/ runtime/src/ common/src/

Conventions

Database / SQLAlchemy

  • No foreign keys, no relationship — every join is explicit.
  • Every table has is_deleted TINYINT(1) NOT NULL DEFAULT 0 and deleted_at DATETIME(3) NULL; queries must filter is_deleted == 0 (or deleted_at.is_(None)) to avoid logical-deleted rows.
  • Domain files under common/src/common/db/models/ are split by bounded context: audit / events / experiments / identity / runtime / schedules / scripts / storage / workspaces.
  • All models are class X(Base) SQLAlchemy 2.0 declarative-mapped.
  • The full schema is in migrations/versions/. Apply with:
    uv run --frozen --package backend alembic upgrade head
    

Storage

  • All object bytes go through common.storage.AsyncStorageBackend, created by create_storage(config) from common.storage.factory.
  • Two backends are registered: local (filesystem, local mode) and s3 (S3-compatible service, s3 mode). Selection is per-deployment via settings.storage_backend ("s3" default, "local" for dev / single-node / air-gapped).
  • The factory helper build_storage_config(bucket_name) returns the right create_storage kwargs for each of the 4 purpose buckets (workspace, version, run_log, trash). Use it in lifespan code; route handlers don't see the difference.
  • Bucket resolution from usage_type is in one place (backend/storage_api.py:resolve_bucket); route handlers only know about app.state.object_stores[bucket_name].
  • The runtime's view of the workspace bucket on disk is exposed by common.storage.workspaces_root():
    • s3 mode: ${settings.local_storage_base_dir}/workspace (default /data/workspace, the rclone FUSE mount target).
    • local mode: ${settings.local_storage_base_dir}/workspace (default /data/workspace, a subdir of the shared local-storage volume). settings.local_storage_base_dir is the only path setting; the helper handles the per-mode suffix. Don't read settings.workspaces_root or any other path setting directly in runtime code — use this helper.
  • The pre-2026 abstraction (RustFSObjectStore / common.storage.client / StorageClient HTTP wrapper) is gone. Don't reintroduce it.

Auth

  • Browser → /jupyter/{workspace_id}/... → Nginx auth_requestGET /api/v1/auth/jupyter (Backend).
  • The handler:
    1. Parses cookie / Bearer JWT (HS256 + settings.jwt_secret).
    2. Verifies WorkspaceMembers for the workspace.
    3. Verifies Scripts.is_locked for the requested notebook path (owner / unlocked → allow; otherwise 403).
    4. Calls RuntimeClient.get_workspace / start_workspace.
    5. Returns x-upstream-addr + x-jupyter-internal-token response headers. Browser never holds the runtime token.

Service-to-service auth (P0-1 fix)

  • Schedule → Backend single endpoint POST /internal/v1/objects is guarded by require_internal_service in backend.api.storage.
  • The token header is X-Internal-Service-Token (case-insensitive on the wire because FastAPI Header lowercase-matches the name x-internal-service-token); the secret value comes from settings.internal_service_token / env INTERNAL_SERVICE_TOKEN.
  • Comparison uses secrets.compare_digest — never equality.
  • Backend and schedule must be configured with the same value; a mismatch fails fast at the first notebook run (401) which is intentional. .env.example ships a placeholder change-me-internal-service-token and the docker-compose ${INTERNAL_SERVICE_TOKEN:?...} reference forces production deployments to set a real value.
  • Removing the legacy backend / runtime host-port mappings (8891:8000 / 8892:8000) is part of the same fix — no service is reachable from the host except Nginx anymore.
  • Nginx captures the headers via auth_request_set and proxies to the upstream sub-process with Authorization: token $jupyter_token.

Permission gates

  • For write operations on a script/notebook, call require_script_modify_access(script, user_id=..., is_admin=...) from backend/src/backend/services/scripts.py. It enforces:
    • admin or owner → allow
    • non-owner, is_locked == 0 → allow
    • non-owner, is_locked == 1 → 403
  • Read endpoints (list_scripts, get_script) intentionally do not check is_locked — workspace members can see the script list.

Owner-scoping + visibility (cross-owner browsing)

GET /api/v1/scripts, GET /api/v1/data-resources, and GET /api/v1/workspace-directories all accept an optional owner_user_id query param and follow the same model:

  • owner_user_id 缺省 = 当前请求者本人(scope to workspace/{me}/...).
  • 传值时 scope 到 workspace/{owner_user_id}/...,用于前端"点开其他成员 分组"的懒加载(见 §3.4 of API.md).
  • visibility 过滤(非 admin):owner_user_id == me OR visibility IN (workspace, public) — 自己可见自己全部(含 private),他人只见其 workspace/public,排除他人 private.
  • 系统管理员跳过 visibility 过滤.
  • 目录行默认 visibility='public'(由 create_workspace_directory 写入), 不施加 visibility 过滤,使跨 owner 目录树可见.

实现 helper 在 backend/src/backend/services/scripts.py (_build_list_scripts_owner_descendant_prefix) 和 backend/src/backend/api/resources.py 中按相同模式分别构造 owner-scoped LIKE 前缀。

Outbox events

  • The platform's only async-messaging fabric is the MySQL OutboxEvents table. Producers (Backend) write rows in the same transaction as the business state. Consumers (Schedule Executor) poll every 250 ms and update event_status to published or failed (with retry).
  • Use common.eventing.add_outbox_event to write.
  • event_type values currently in use:
    • schedule.run.requested — produced by schedule_runs.py / cron post-back
    • job.node.execute — produced by orchestrator when a node is ready
    • job.node.finished — produced by worker after node execution
  • consumer_inbox provides exactly-once delivery per (consumer_name, event_id) (with process_status: processing → succeeded lifecycle).

Async / sync signatures

  • SchedulerService.start() and SchedulerService.close() are async def (so the FastAPI lifespan can await them).
  • CronScheduler.start(), DispatchOrchestrator.start() are def (sync) — they only create_task(...) and return. Don't await them.
  • cron.start(), orchestrator.start(), worker.handle_node_execute are wired together in SchedulerService.__init__; their lifetime is owned by the facade.

Local development

One-time setup

# Python workspace (monorepo via uv workspaces)
uv sync --all-packages

# Frontend deps
cd frontend && pnpm install && cd ..

Per-service dev

Always go through uv run so the workspace .venv is used — bare uvicorn / python resolves to system Python and from backend.X imports fail with ModuleNotFoundError.

# Backend (terminal 1)
export DATABASE_URL="mysql+asyncmy://model_platform:model_platform@127.0.0.1:3306/model_platform?charset=utf8mb4"
export STORAGE_BACKEND=s3
export S3_ACCESS_KEY=modelplatform
export S3_SECRET_KEY=modelplatformsecret
export S3_ENDPOINT=http://127.0.0.1:9000
# Or for local mode:
# export STORAGE_BACKEND=local
# export LOCAL_STORAGE_BASE_DIR=/data
uv run --package backend uvicorn backend.main:app --host 0.0.0.0 --port 8000 --reload

# Schedule Executor (terminal 2)
uv run --package schedule uvicorn schedule.main:app --host 0.0.0.0 --port 8001 --reload

# Runtime (terminal 3 — needs SYS_ADMIN, FUSE, devmode)
uv run --package runtime python -m runtime.main

Frontend dev

cd frontend
pnpm dev          # http://localhost:5173, proxies /api to backend
pnpm typecheck
pnpm build

Static checks

# Python compile
uv run --frozen --package backend python -m compileall -q backend/src common/src
uv run --frozen --package schedule python -m compileall -q schedule/src
uv run --frozen --package runtime python -m compileall -q runtime/src

# Type check (frontend)
cd frontend && pnpm typecheck && cd ..

# Docker compose config
docker compose config --quiet

Smoke test

# Run the full ORM import + Settings smoke test
PYTHONPATH="backend/src:common/src" uv run --frozen --package backend python -c "
from backend.main import app
from common.config import settings
print('backend:', len(app.routes), 'routes')
print('settings ok:', settings.s3_endpoint)
"

Common tasks

Add a new DAG endpoint

  1. Add the route handler in backend/schedules.py (DAG template) or backend/schedule_runs.py (run lifecycle).
  2. Validate request via backend/schedule_schemas.py.
  3. For mutations on nodes/edges/versions: route through create_script_record / get_script_row and apply require_script_modify_access if it touches a script.
  4. If it produces an outbox event, use add_outbox_event(session, event_type="...", producer="...", ...).

Add a new env var

See "Adding a new env var" above.

Add a new MySQL table

  1. Add a model class in common/src/common/db/models/<domain>.py. Include is_deleted TINYINT(1) NOT NULL DEFAULT 0 and deleted_at DATETIME(3) NULL.
  2. Export it from common/src/common/db/models/__init__.py.
  3. Generate the migration:
    uv run --frozen --package backend alembic revision --autogenerate -m "add <feature>"
    
  4. Review the generated migrations/versions/*.py — Alembic may miss comments / server defaults. Manually fix the migration.
  5. Apply locally:
    uv run --frozen --package backend alembic upgrade head
    

Wire a new storage bucket

The current 4 buckets are wired in backend/storage_api.py:resolve_bucket:

BUCKET_FOR_USAGE: dict[str, str] = {
    "working_copy":     settings.s3_workspace_bucket,
    "public_script":    settings.s3_workspace_bucket,
    "data_resource":    settings.s3_workspace_bucket,
    "snapshot":         settings.s3_workspace_bucket,
    "version_artifact": settings.s3_version_bucket,
    "run_log":          settings.s3_run_log_bucket,
    "run_result":       settings.s3_run_log_bucket,
}

The constant PURPOSE_BUCKETS = ("workspace", "version", "run_log", "trash") in common.storage.factory enumerates the four backends built in the backend lifespan. To add a fifth bucket:

  1. Add the env var to Settings (s3 mode only):
    s3_<feature>_bucket: str = Field(default="<feature>", description="...")
    
  2. Add to .env.example with a one-line comment.
  3. Append "<feature>" to the PURPOSE_BUCKETS tuple in common/storage/factory.py. build_storage_config("<feature>") will then automatically read settings.s3_<feature>_bucket (s3 mode) or use <local_storage_base_dir>/<feature> (local mode).
  4. Extend the Literal in common/storage/schemas.py (in CreateUploadRequest.usage_type, ServerObjectRequest.usage_type) to include the new value.
  5. Add an entry in BUCKET_FOR_USAGE mapping the new usage_type to the new bucket env var.
  6. Pre-create the bucket (s3 mode) or subdirectory (local mode) in the deployment. The backend no longer auto-creates buckets.

A workspace's artifact_bucket column (when non-null) overrides the default for that workspace, regardless of usage_type.

Add a new schedule node type

schedule/execution.py dispatches on script_type in execute_artifact. Add a new branch + a new _<type> function. worker.py does not need to change — the dispatch happens inside execute_artifact.

Tests

There is no formal test suite yet (see HANDOVER §Pending Tasks P1). A reasonable first test surface:

  • require_script_modify_access (admin / owner / non-owner-unlock / non-owner-lock): pure-function unit test, no DB.
  • validate_dag (cycle detection + orphan detection) in backend/schedules.py.
  • execute_artifact end-to-end with mocked content_hash and a real tempfile.TemporaryDirectory.

Test convention: pytest with pytest-asyncio for async def handlers. Use SQLite in-memory (or a MySQL test container) for DB integration. Use moto for S3.

Troubleshooting

"the greenlet library is required"

SQLAlchemy 2.0 needs greenlet for engine.dispose() in async contexts. Add greenlet>=3.0.0 to common/pyproject.toml and uv sync --all-packages. (Already present in this repo.)

"Can't connect to MySQL server"

Either MySQL isn't running, or the network namespace doesn't allow mysql:3306 resolution. Inside the Docker network, services reach each other by service name (mysql, backend, runtime, schedule, s3).

Jupyter routing 401s

Inspect docker compose logs backendjupyter.py:check_notebook_is_locked or load_active_membership will return an explicit reason. Then check the JWT (use JWT_SECRET from .env).

Schedule run never advances

schedule_runs.run_status is stuck at queued. Two likely causes:

  • outbox_events is empty (Backend's add_outbox_event failed — check add_outbox_event in schedule_runs.py).
  • The orchestrator's polling loop is dead. Check docker compose logs schedule and look for "database event loop failed" exceptions.

"AttributeError: 'SchedulerService' object has no attribute 'worker'"

worker must be constructed before orchestrator in SchedulerService.__init__, because orchestrator's dispatch table captures self.worker.handle_node_execute at construction time. See service.py — the order is load-bearing.

Style

  • Type hints everywhere (this repo uses from __future__ import annotations).
  • 4-space indent, double quotes, no trailing whitespace.
  • Comments are technical (explain why, not what).
  • Module docstrings document non-obvious invariants. Don't add docstrings to functions whose behavior is self-evident from the name.
  • 4 levels of indentation = "this function is doing too much; split it". (Project convention; see e.g. execute_artifact.)

See also

  • ARCHITECTURE.md — design diagrams
  • HANDOVER.md — current refactor state and pending work
  • CLAUDE.md — agent-facing conventions for the repo
  • models / __init__.py — exhaustive list of all 26 tables
  • common/config.py — all env vars in one place