doc: add Helm chart deployment page for Kubernetes - #494
bitflicker64 wants to merge 12 commits into
Conversation
Documents the distributed Helm chart from apache/hugegraph#3218: install with the three values presets, the auth Secrets model, PD health vs readiness probing, upgrade and scale-down caveats, uninstall behavior, and current limitations. EN and CN pages under quickstart/hugegraph, weight 4, matching the sibling page skeleton.
The versioned build validates that every canonical Docs page has a logical ID in data/version_routes.json; the two new pages were missing from it, failing the Build latest job. Adds en: and cn: entries with latest populated and the four older versions null, the same shape as the other pages introduced after 1.7.
bitflicker64
left a comment
There was a problem hiding this comment.
Blocking: yes, until apache/hugegraph#3218 merges. Summary: The EN and CN pages follow the chart at 05f3d9e0b closely and render in CI, but the chart is not on apache/hugegraph master yet (#3132 was closed unmerged, #3218 is an open draft), so merging this first publishes install commands that fail and seven 404 chart links per language. Four smaller mismatches with the chart are inline; each applies to the CN page too. Evidence: compared both pages with helm/hugegraph at 05f3d9e0b (values*.yaml, values.schema.json, _helpers.tpl, NOTES.txt, README anchors); helm template with defaults, values-cluster.yaml, hubble.enabled=true and invalid secret values; gh api repos/apache/hugegraph/contents/helm?ref=master (404); gh pr checks 494 (all passing, publish skipped).
|
|
||
| The Helm chart deploys a distributed HugeGraph cluster on Kubernetes: PD, Store, and Server, plus the optional | ||
| Hubble UI. It lives in the main repository under | ||
| [`helm/hugegraph`](https://github.com/apache/hugegraph/tree/master/helm/hugegraph). |
There was a problem hiding this comment.
gh api repos/apache/hugegraph/contents/helm?ref=master returns 404, #3132 was closed without merging, and #3218 is still an open draft. If this PR merges first, the site publishes a guide where the helm install in 3.2 fails because helm/hugegraph does not exist on master, and all seven tree/master/helm/hugegraph links on the page (EN and CN) are 404s. Please convert this PR to draft until #3218 merges.
There was a problem hiding this comment.
Still applies at 795d5f8. Checked 2026-10-02: apache/hugegraph master (176fb56dd) has no helm/ directory, and apache/hugegraph#3218 is open (out of draft, head d05741d9c). Until it merges, the install in 3.2 and the tree/master/helm/hugegraph links fail. Leaving this thread open as the merge-order gate: merge #3218 first, or hold this PR until then.
There was a problem hiding this comment.
Still applies at a5fadcc. Checked 2026-10-03: apache/hugegraph master (02628ed50) has no helm/ directory, and apache/hugegraph#3218 is open at 2a6e2c404. Leaving this thread open as the merge-order gate: merge #3218 first, then this PR.
imbajin
left a comment
There was a problem hiding this comment.
Blocking: yes. Summary: The guide contains memory and probe claims that differ from the paired chart and a Store rollout check that can remove a second replica before recovery. Evidence: exact-head comparison with chart #3218 values, helpers, and README; the docs workflow checks passed.
Bring the deployment page level with the chart at 8603cdbb3: the values-cluster preset now ships NetworkPolicy and 5Gi/8Gi Store memory, PD startup and liveness derive from the replica count (pd.livenessPath, single-PD /v1/ready per apache/hugegraph#3222), the Store roll procedure no longer treats Up in PD as the between-pods check, and the Limitations follow the merged fixes, including in-place empty-PVC Store recovery on images carrying apache/hugegraph#3234. Add an operations page (EN and CN, registered in the version route map) rewritten for operators from the chart README: ports and health, scheduling, partition sharding, NetworkPolicy with the per-CNI NodePort client behavior, safe Store rolls, the disaster recovery runbook with the post-#3234 procedure, scaling including the Store drain steps, running Hubble outside the cluster, and the two causes of "Could not rebind". Verified: scripts/hugo.sh build passes and both new pages render with their cross-links in EN and CN.
Cut the README from 1,472 to 1,003 lines now that the docs site carries the operator walkthroughs. Each moved section keeps its load-bearing warning and commands and links its docs-site path: NetworkPolicy details, Cluster Health, Scheduling, Partition Sharding, the Disaster Recovery narrative (the keep-the-PVC rule and the retirement commands stay), the Scaling procedures, the outside-Hubble paths, and the Could-not-rebind measurements. The quickstart, presets, Kind flow, upgrade warnings with the OnDelete Store roll, the values tables, the validation list, every troubleshooting symptom and check command, and the Limitations stay. Deduplicate repeated passages to one home each: the ordinal-truncation explanation (Release Name Too Long), the anti-affinity trade (Installing), and the Hubble single-replica/H2 constraints (the Hubble section; the values.yaml comment now points there). README-only plus a values.yaml comment: helm template output is byte-identical, lint passes on the three presets, and all 139 unit tests pass. The docs-site links resolve once apache/hugegraph-doc#494 merges.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Several documentation corrections and navigation updates remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 6
Open (6)
Clarify that port-forward and curl require separate terminals · New Prevent port-forward from blocking installation verification · New Document leading-whitespace restrictions for admin passwords · New Clarify that port-forward and curl require separate terminals · New Prevent port-forward from blocking API verification · New Document leading-whitespace restrictions for admin passwords · New
What changed in this PR
Adds bilingual Helm deployment and operations documentation for HugeGraph on Kubernetes.
Changes:
- Adds English and Chinese deployment guides.
- Adds operations guides covering installation, networking, scaling, recovery, and Hubble.
- Registers the new pages in the version route map.
| File | Description |
|---|---|
data/version_routes.json |
Adds routes for the new pages. |
content/en/docs/quickstart/hugegraph/hugegraph-helm.md |
English deployment guide. |
content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md |
English operations guide. |
content/cn/docs/quickstart/hugegraph/hugegraph-helm.md |
Chinese deployment guide. |
content/cn/docs/quickstart/hugegraph/hugegraph-helm-operations.md |
Chinese operations guide. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
imbajin
left a comment
There was a problem hiding this comment.
Blocking: yes. Summary: The recovery guide compares a term that this Store endpoint does not return, can run balancing commands during a patrol-triggered cooldown, and leaves a re-used ordinal with a Tombstone Store ID on its retained PVC. Evidence: current Server master handlers and the chart head's default PVC-retention setting.
bitflicker64
left a comment
There was a problem hiding this comment.
Blocking: yes. Summary: The deployment page's Upgrade section says the first upgrade after install rolls PD and Server together but never warns that a Server started during that PD roll can lose its Gremlin binding for the life of the Pod, and its scaling paragraph says only PD and Store shrinks are rejected although the chart also rejects PD growth. Evidence: compared both EN and CN pages with helm/hugegraph at 3996bb12 (head of apache/hugegraph#3218): README.md Upgrading the Chart, templates/_helpers.tpl lines 649-681, values*.yaml, NOTES.txt, networkpolicy.yaml, README anchors; findings already raised in earlier reviews are not repeated.
There was a problem hiding this comment.
✅ Score: 8.5/10
Before merging:
- Coordinate with Helm #3218 and settle the operational wording.
- Resolve the missing
scripts/oink_module.pyprerequisite and rerun website checks. - Keep artifact/version-specific recovery limitations explicit.
The reported site failure occurs before it can establish a Helm documentation regression.
Describe the chart as of apache/hugegraph#3218 at 0e754b98c, in EN and CN: - helm test now queries Gremlin on every Ready Server Pod through the headless Service, with the replica floor and the 150 s retry window - values-cluster.yaml sets Store updateStrategy OnDelete; the other presets keep RollingUpdate, and OnDelete is not a safety check by itself - rollingUpdate.maxUnavailable accepts only the integer 1 - Server startup budget: guaranteed time is (failureThreshold - 1) * periodSeconds, default 91 x 5, HG_SERVER_STARTUP_TIMEOUT_S 150 s - PD secret and admin password whitespace and backslash constraints, and the Hubble egress render guard with server.advertiseUrl - Limitations and troubleshooting for apache/hugegraph#3226 and #3228, #3229 referenced from the Upgrade section, and the Scaling link moved from the chart README to the operations page where the procedure lives
Since the OINK migration the sidebar comes from data/docs_nav.json, and the two Helm pages were not in it, so no navigation led to them. Add both under the HugeGraph components entry after PD, regenerate the derived maps with materialize_docs_navigation, and update the pinned navigation stats: latest gains two pages, and each archived version, which lacks the pages, counts two more removed entries.
State each limitation as the behavior a user sees, with the issue link kept as a reference. The Store roll now says plainly that no endpoint reports restore completion. The uninitialized-PD entry separates what was observed outside Kubernetes from what the chart's probes do, and no longer offers an untested recovery step. helm test prints the failing Pod's IP, not its name.
kubectl port-forward stays in the foreground, so the curl commands that followed it in the same block never ran. Split the install check, the operations health check and the per-Pod Gremlin check into a forward block and a call block, in English and Chinese.
- values-cluster.yaml adds the Server PodDisruptionBudget; values.yaml already enables the PD and Store budgets. - Each JVM sets its heap to half of the free node memory, so the single-node preset needs memory sizing too. - Explain that --reuse-values keeps old values as the base and how to adopt new defaults; Prerequisites referred to this. - The chart rejects PD growth as well as PD and Store shrinks. - The Store per-group partition route never fills term. - patrolPartitions also sets the 180 s balance-shard flag when it reallocates a shard group. - PD exposes POST /v1/members/change behind its REST authentication. - A retained PVC of a retired Store ordinal cannot register again.
The admin password must now be printable ASCII with no spaces, colons or backslashes, and a bring-your-own Secret is checked at Pod start. A no-change upgrade rolls nothing, and a template-only pipeline rolls the Pods that read a generated credential on every sync instead of leaving the annotation inert. Document the Server and Hubble Service exposure acknowledgements, the install-time identity and bootstrap-only values, and why helm rollback bypasses the chart's guards. EN and CN updated together.

Purpose of the PR
apache/hugegraph#3218 adds a Helm chart for the distributed deployment (PD + Store + Server, optional Hubble), and the site has no pages for it yet. This adds two pages in EN and CN under
docs/quickstart/hugegraph/, following the same skeleton, heading numbering and mermaid usage as the sibling pages, both registered in the version route map:Deployment page (
hugegraph-helm, weight 4):values.yaml,values-single.yaml,values-cluster.yaml),--waitguidance, and the defaults to know before production use: no resources are set by default, image tags tracklatestuntil the next release, and the cluster preset ships NetworkPolicy plus 5Gi/8Gi Store memory (the RocksDB caches live off-heap).existingSecretbring-your-own paths, the value constraints (the admin password is printable ASCII with no spaces, colons or backslashes), the rotation caveats, and what a template-only pipeline does to chart-generated credentials./v1/ready, and the PD startup/liveness path derived from the replica count (/v1/healthabove one replica,/v1/readyat one, per [Bug] A single-node PD never recovers raft leadership after a failed periodic snapshot (disk full), even once the disk is freed; /v1/health stays 200 and /v1/ready hides the error state hugegraph#3222), withpd.livenessPathas the override; and the Server startup budget, counted as(failureThreshold - 1) * periodSeconds(91 x 5 by default) withHG_SERVER_STARTUP_TIMEOUT_Sat that time minus the 300 s storage wait (150 s by default).helm testchecks (a graph-bound Gremlin query on every Ready Server Pod through the headless Service, at least the replica floor, retried for 150 s), that a no-change upgrade rolls nothing, theOnDeleteStore strategy thatvalues-cluster.yamlnow sets and the between-pods check (Upin PD is not it), the restriction ofmaxUnavailableto1, the PVC resize constraint, the guard that reads the live StatefulSet and rejects any change to the PD count or a Store scale-down, the install-time identity and bootstrap-only values, and whyhelm rollbackbypasses the chart's guards.<component>.service.allowInsecureExposure) and the IngressallowPlainHttpopt-in are listed there too. Two more limitations: a Server that starts while PD is unreachable can lose its Gremlin binding while staying Ready ([Bug] A Server that starts while PD is briefly unreachable serves REST but loses its Gremlin binding for the life of the process hugegraph#3228, whichhelm testnow detects), and a PD that fails to open its RocksDB store stays uninitialized and, with more than one PD, is never restarted by its probes ([Bug] PD reports "Started" and keeps serving after failing to open its RocksDB (lock held by the previous instance): every request fails, /v1/health stays 200, no retry, no exit hugegraph#3226).Operations page (
hugegraph-helm-operations, weight 5), the chart README's deep-dive content rewritten for an operator: ports and health, scheduling, partition sharding, NetworkPolicy including how NodePort clients appear under kindnet/Calico/Cilium, the safe Store roll procedure (with theGET :8520/v1/partitionsfollower behavior before and after apache/hugegraph#3232), the disaster recovery runbook (leader targeting, patrol and balance semantics including the apache/hugegraph#3233 refusal body, telling a real run from a no-op, and the empty-PVC Store replacement on both pre- and post-#3234 images), scaling including the Store drain steps, running Hubble outside the cluster, and the two causes ofCould not rebind.The full values reference stays in the chart README; the pages link to its anchors instead of duplicating the tables, so they cannot drift apart.
Checked against the chart at commit
2a6e2c404of apache/hugegraph#3218 (2026-10-03). Merge order: apache/hugegraph#3218 first; until it lands, the install command and thetree/master/helm/hugegraphlinks on these pages do not resolve. Both pages are listed in the docs sidebar (data/docs_nav.json) under the HugeGraph components, after PD. Verified locally after mergingmaster(e6389aa), with Python 3.13 and Hugo 0.165.0 Extended:scripts/hugo.sh build,bash dist/validate-links.sh,python -m unittest discover -s scripts -p 'test_*.py'(227 tests, OK),scripts/versioning.py buildplusvalidateforlatest(305 pages), andtests/e2e/search-ranking.spec.jsagainst the aggregated site (24 passed). The pages render (mermaid diagram, tables, details blocks, cross-links) in EN and CN.Planned follow-ups
These are planned work in apache/hugegraph, not conditions for merging this PR. The pages here describe the chart as it is now and get updated when a follow-up changes user-facing behavior.
/versions. Switch it to the storage-awareGET /readinessendpoint once feat(server): add a storage-aware GET /readiness endpoint hugegraph#3221 merges and reaches the images the chart deploys, then update the health-check section here.helm upgradefrom that released version to the current chart.helm-chart-ci.ymlcovers the static layer only (lint, unit tests, render assertions, kubeconform). Add a job that installsvalues-single.yamlon Kind and runshelm test, with atimeout-minutesbudget sized for the Server and Store startup waits, as planned on 2026-08-27 (note on apache/hugegraph#3132).Rendered pages
Screenshots show the deployment page; the operations page uses the same skeleton.