Skip to content

[Bug] HStore Server image (hugegraph/server) keeps crash diagnostics only inside the container, so a restart on Kubernetes loses them #3256

Description

@bitflicker64

Bug Type (问题类型)

server status (启动/运行异常)

Before submit

  • I have confirmed and searched that there are no similar problems in the historical issue and documents

Searched open and closed issues and PRs for hs_err, ErrorFile, HeapDumpPath, additivity, STDOUT_MODE, Dockerfile-hstore, docker logs, kubectl logs and Starting HugeGraphServer failed. #2979 / #2980 wired the console appender and set STDOUT_MODE=true in hugegraph-server/Dockerfile (the standalone hugegraph/hugegraph image), but not in hugegraph-server/Dockerfile-hstore, which builds hugegraph/server (docker/README.md:240). #3203 hit the same wall during a cold-start failure but tracked the retry bug, not the logging.

Environment (环境信息)

Expected & Actual behavior (期望与实际表现)

Observed. In the 2026-09-10 run (#3203), kubectl logs --previous for a Server that exited during startup ended with these lines and held no stack trace:

Connecting to HugeGraphServer (http://0.0.0.0:8080/graphs)...........Starting HugeGraphServer failed
See /hugegraph-server/logs/hugegraph-server.log for HugeGraphServer log output.

(printed by util.sh:362 and start-hugegraph.sh:117). The second line is the non-STDOUT_MODE branch. The cause was in logs/hugegraph-server.log, which went away with the container; the trace in #3203 was only recovered by mounting a volume at /hugegraph-server/logs before the first boot. On 2026-09-22 one Server started during a helm upgrade exited once with Starting HugeGraphServer failed, and its cause was lost the same way.

Expected. After a Server container exits, its stdout (kubectl logs --previous) shows the error, and the JVM crash files sit in a documented directory an operator can mount.

Actual. Four things keep diagnostics inside the container filesystem.

  1. The HStore image does not set STDOUT_MODE. Dockerfile-hstore sets only JAVA_OPTS and HUGEGRAPH_HOME (:43-44), and the published hugegraph/server:latest (amd64 config, created 2026-10-01, read from the Docker Hub registry on 2026-10-02) has no STDOUT_MODE in its Env. Without it, hugegraph-server.sh sends the JVM's stdout and stderr, console appender included, to logs/hugegraph-server-stdout.log (:266-270), so nothing the JVM logs reaches kubectl logs.

  2. Even with STDOUT_MODE=true, five loggers write only to the file. log4j2.xml sets LOG_PATH to logs (:21) and the file appender writes ${LOG_PATH}/hugegraph-server.log (:32). Since fix(docker): enable docker logs for pd/store/server containers #2980 the root logger and org.apache.hugegraph also go to console (:106-109, :126-129). org.apache.hadoop, org.apache.zookeeper, com.alipay.sofa, io.netty and org.apache.commons are additivity="false" with only file (:110-124), so their errors never reach stdout. The file appender also has immediateFlush="false" (:34), so a killed JVM can lose its last lines from the file as well. The audit loggers (:130-135) and the slow-query log (:136-138) are file-only too. They are separate streams and out of scope here.

  3. JVM crash files land inside the install directory. hugegraph-server.sh sets LOGS="$TOP/logs" (:41), passes -XX:HeapDumpPath=${LOGS} (:124) and runs cd "${TOP}" before exec java (:91). Nothing in the repository sets -XX:ErrorFile (git grep ErrorFile on 176fb56dd returns nothing), so HotSpot writes hs_err_pid<pid>.log to the working directory, /hugegraph-server in the image. The image adds nothing here. JAVA_OPTS (Dockerfile-hstore:43) reaches the script through -j (docker-entrypoint.sh:243) and is appended on line 124. Line 124 sits inside the JAVA_OPTIONS default block (:118-129), so a user who sets JAVA_OPTIONS also loses -XX:+HeapDumpOnOutOfMemoryError.

  4. The only persistence is an image VOLUME. Dockerfile-hstore has WORKDIR /hugegraph-server/ (:46) and VOLUME /hugegraph-server (:78). Docker keeps that anonymous volume across docker restart. Kubernetes does not turn an image VOLUME into a pod volume. A restart starts a new container, and logs/, hs_err_pid*.log and heap dumps go with the old one.

Fix direction.

Related: #2979 / #2980, #3203, #3236 / #3253, #3211 (non-root image, which would chown the same log directory).

Vertex/Edge example (问题点 / 边数据举例)

Not applicable.

Schema [VertexLabel, EdgeLabel, IndexLabel] (元数据结构)

Not applicable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions