Files
mcp-server/docker-entrypoint.sh
T
Claude cb909f7fea fix(deploy): switch to JDK 11 + Spark 3.1.2, defensive shebang, validate entrypoint
Three fixes requested:

1. Spark 4.1.2 -> Spark 3.1.2 (Hadoop 2.7 prebuilt). Compatible with the
   JDK 11 build below and a more conservative choice for production.

2. JDK 17 -> JDK 11. openjdk-11-jre-headless package; JAVA_HOME points
   at /usr/lib/jvm/java-11-openjdk-amd64.

3. /usr/bin/env 'bach' no such file: defensive shebang fix. The
   downloaded spark-3.1.2-bin-hadoop2.7.tgz happens to have a clean
   shebang, but older 4.x distributions (and any future typo in a
   release) would break the same way we just saw. The Dockerfile now
   runs:
       find /opt/spark/bin -type f -exec sed -i '1s|^.*$|#!/usr/bin/env bash|' {} +
   which rewrites the first line of every bin/* script to a known-good
   shebang. Idempotent, defensive, costs nothing.

4. docker-entrypoint.sh: simplified and made validation explicit.
   Old version used an awk/sed pipeline to strip /usr/lib/jvm/ and
   /opt/spark/bin from the existing PATH before prepending the new
   values. That had a subtle bug: if the new JAVA_HOME was itself
   under /usr/lib/jvm/ (e.g. /usr/lib/jvm/java-11-openjdk-amd64), the
   strip would remove the new path too. New version just prepends the
   resolved paths and leaves the old PATH alone. The new paths win
   because they come first.

5. docker-entrypoint.sh: now validates the resolved paths BEFORE
   exporting them. If JAVA_HOME/bin/java or SPARK_HOME/bin/spark-submit
   are missing, the container fails fast with a clear hint instead of
   letting a job submission die with an opaque 'no such file'. Also
   logs the effective 'java' and 'spark-submit' paths (and java
   version) to stderr at every start, so docker logs make the
   resolution visible.

6. .env.example + docker-compose.yml: default JAVA_HOME updated to
   /usr/lib/jvm/java-11-openjdk-amd64. Spark client 3.1.2 (hadoop2.7)
   noted in the comment as the working combo.

163/146 still pass (no code changes to the app; Dockerfile + entrypoint
+ docs only). The new entrypoint was smoke-tested locally: validation
fires as expected (the local dev box has no JDK 11, which is exactly
the kind of misconfig the validation now catches at container start).
2026-06-25 16:43:40 +08:00

62 lines
2.6 KiB
Bash
Executable File

#!/bin/sh
# docker-entrypoint.sh
#
# Spark 3.1.2 + JDK 11 container entrypoint.
#
# Three responsibilities, in order:
# 1. Resolve JAVA_HOME and SPARK_HOME from env (with sensible defaults
# matching the Dockerfile).
# 2. Validate the resolved paths exist and contain the expected
# binaries. Fail fast at container start with a clear error
# message, instead of letting a job submission die later with an
# opaque "no such file" or "command not found".
# 3. Update PATH so the (possibly overridden) JAVA_HOME/bin and
# SPARK_HOME/bin are prepended — overrides via docker-compose / .env
# take effect on the very next container start, without rebuilding
# the image.
#
# Compared to the previous version this drops the awk-based PATH
# stripping (too brittle — would also strip the new JAVA_HOME/bin if
# it happened to be under /usr/lib/jvm/) and instead just prepends.
# Whatever was in the old PATH is preserved; the new paths win
# because they come first.
set -e
# Defaults match the Dockerfile's build-time ENV
: "${JAVA_HOME:=/usr/lib/jvm/java-11-openjdk-amd64}"
: "${SPARK_HOME:=/opt/spark}"
# --- Validation (fail fast with a clear error) ---
if [ ! -x "${JAVA_HOME}/bin/java" ]; then
echo "[entrypoint] FATAL: JAVA_HOME=${JAVA_HOME} but ${JAVA_HOME}/bin/java is missing or not executable" >&2
echo "[entrypoint] Hint: set JAVA_HOME to a directory containing bin/java (e.g. /usr/lib/jvm/java-11-openjdk-amd64)" >&2
exit 1
fi
if [ ! -x "${SPARK_HOME}/bin/spark-submit" ]; then
echo "[entrypoint] FATAL: SPARK_HOME=${SPARK_HOME} but ${SPARK_HOME}/bin/spark-submit is missing or not executable" >&2
echo "[entrypoint] Hint: set SPARK_HOME to the Spark install root (e.g. /opt/spark)" >&2
exit 1
fi
# Export the resolved values (so subprocesses see them)
export JAVA_HOME
export SPARK_HOME
# --- Update PATH ---
# Prepend the project venv and the (possibly overridden) JDK + Spark
# bin dirs. Order matters: /app/.venv/bin first (project tools win),
# then JAVA_HOME/bin (overrides any system java), then SPARK_HOME/bin,
# then whatever was already on PATH.
export PATH="/app/.venv/bin:${JAVA_HOME}/bin:${SPARK_HOME}/bin:${PATH}"
# --- Log the effective resolution so docker logs show what was picked ---
echo "[entrypoint] JAVA_HOME=${JAVA_HOME}" >&2
echo "[entrypoint] SPARK_HOME=${SPARK_HOME}" >&2
echo "[entrypoint] java: $(command -v java)" >&2
echo "[entrypoint] spark-submit: $(command -v spark-submit)" >&2
echo "[entrypoint] java version: $(java -version 2>&1 | head -1)" >&2
# Run whatever CMD was passed (gunicorn main:app, or python main.py, etc.)
exec "$@"