Skip to content

Entrypoint: stop with Keycloak, persist setup marker with data, add KC_SKIP_DEFAULT_SETUP - #16

Merged
slominskir merged 1 commit into
mainfrom
entrypoint-signals-and-setup-marker
Sep 28, 2026
Merged

slominskir merged 1 commit into
mainfrom
entrypoint-signals-and-setup-marker

Conversation

@jlab-coding-agent

Copy link
Copy Markdown
Contributor

Fixes entrypoint problems found when running this image with restart: unless-stopped and a volume on /opt/keycloak/data (Sync Board). Existing users see no change unless they opt in, with the exceptions listed under Behavior changes.

Changes

1. The container now stops when Keycloak stops, and docker stop shuts Keycloak down cleanly (container-entrypoint.sh)

  • The entrypoint traps TERM and INT, passes TERM on to Keycloak, waits for it, and exits.
  • After setup it waits on Keycloak's PID instead of sleep infinity, then exits with Keycloak's status. If Keycloak crashes, the container exits and Docker's restart policy applies.
  • If Keycloak exits before it answers at KC_BACKEND_URL, the wait loop now exits too instead of polling forever.

2. The setup-complete marker is stored with the database (container-entrypoint.sh, container-healthcheck.sh)

  • With the default dev-file database (KC_DB unset or dev-file), the marker is now ${KC_HOME}/data/setup-complete, next to the H2 database. A container recreated on the same data volume keeps its realm and skips setup instead of re-running kcadm creates that fail.
  • Without a data volume, nothing changes: data/ is in the container layer, as the old marker was.
  • With an external database (e.g. KC_DB=oracle, as in this repo's compose.yaml), the marker stays at ${KC_HOME}/setup-complete, exactly as before. A marker in data/ would say nothing about whether an external database has been set up. It could even skip setup on a freshly wiped database if someone mounts a data volume. The README explains this and suggests KC_SKIP_DEFAULT_SETUP or idempotent scripts when the database outlives the container.
  • The healthcheck uses the same rule, duplicated in both scripts with a "keep in sync" comment. I didn't put it in kc-lib.sh because users who mount their own copy of kc-lib.sh would then break the healthcheck.

3. New opt-in KC_SKIP_DEFAULT_SETUP=true

  • It skips files at the top level of /container-entrypoint-initdb.d that are unmodified copies of the image's /defaults (00_config.env, 01_base.sh, 02_accounts.sh). Keycloak starts with only the master realm and bootstrap admin, and still becomes healthy.
  • Other scripts in that directory still run: extra mounted files, subdirectories, and replaced defaults such as this repo's own 02_accounts.sh mount. So the variable is a supported replacement for the tmpfs trick, and users can still add their own scripts.
  • It's documented in the README's environment variable table.

4. Small fixes

  • Empty KC_BACKEND_URL: the top-level return 0 (an error outside a function) is replaced by waiting on Keycloak without running setup. Before, the script carried on and polled an empty URL forever.
  • *.env branch: added the missing ; after . "${f}", so the echo no longer passes its words as arguments to the sourced file.

The README also gets a short Notes on Setup and Persistence section. VERSION and workflows are unchanged.

Behavior changes to note for the release

  • docker stop now exits with code 143 (Keycloak's status after a graceful SIGTERM shutdown), after about 1–2 s. Before, it took 10 s and exited with 137 (killed).
  • If Keycloak dies, the container now exits instead of staying up and unhealthy. With no restart policy it stays stopped. That's the intended fix, but anyone relying on the old "zombie" container should know.
  • Upgrading with an existing data volume (dev-file DB): the old marker was never in the volume. So on the first start with the new image, setup runs once more against the existing realm, logs the usual kcadm "Conflict" errors, and then writes the new marker. Later recreates skip setup. Users who hide the defaults (like Sync Board's tmpfs) see nothing.

Testing

Built locally with docker build -t jeffersonlab/keycloak:dev ..

Standalone image (scripted, dev-file DB unless noted):

# Scenario Result
T1 Defaults: test-realm, 4 users, marker in data/, 00_config.env sourced cleanly, healthy ✅
T2 docker stop: 1–2 s, exit 143, log shows Stopping Keycloak... and Keycloak stopped in 1.09s (the published :3 image: 10 s, exit 137) ✅
T3 kill -9 of the Java process with --restart unless-stopped: container exits (logs Keycloak exited with status 137), restarts (RestartCount=1), skips setup, healthy again ✅
T4 Volume on /opt/keycloak/data, add a user, stop and remove, recreate: setup skipped, no conflicts, realm and extra user kept ✅
T5 KC_SKIP_DEFAULT_SETUP=true: healthy, no test-realm, 3 defaults logged as skipped ✅
T5b KC_SKIP_DEFAULT_SETUP=true plus a mounted 99_extra.sh: the extra script runs ✅
T6 KC_DB=postgres: marker at ${KC_HOME}/setup-complete, none in data/, realm created, healthy ✅
T7 Empty KC_BACKEND_URL: Keycloak runs, no setup attempted, docker stop in 1–2 s ✅

Sync Board (the keycloak service pointed at :dev by a local, uncommitted override; existing kcdata volume; tmpfs over /container-entrypoint-initdb.d):

  • pnpm c:up: healthy on the existing volume. pnpm kc:setup: "Nothing was missing."
  • pnpm test:integration against the dev server: 20/20 pass.
  • docker restart: 1 s. --force-recreate: "Setup already run; skipping", still healthy.

Not tested: the Oracle path in this repo's compose.yaml. It's covered by the same non-dev-file branch as T6.

Note for local builds: on this Docker 29 / containerd 2 host, the dnf install step hangs (rpm looping over a huge nofile limit) unless you build with --ulimit nofile=1024:1024. CI is unaffected.

🤖 Generated with Claude Code

…SETUP

The entrypoint now waits on Keycloak instead of sleep infinity, so the
container exits (with Keycloak's status) when Keycloak does and Docker
restart policies apply. SIGTERM/SIGINT are passed on to Keycloak for a
clean shutdown instead of a kill after 10s.

With the default dev-file database the setup-complete marker moves to
${KC_HOME}/data, next to the database, so a recreated container with a
data volume skips setup. Other databases keep the old location.

KC_SKIP_DEFAULT_SETUP=true skips the unmodified default scripts while
still running user-provided ones.

Also fixes top-level `return 0` when KC_BACKEND_URL is empty and a missing
`;` after sourcing *.env files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@slominskir slominskir left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. This better aligns this container with app database containers

@slominskir
slominskir merged commit 66627ec into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant