Skip to content

[Bug] A Store rebuilt with an empty volume never rejoins its raft groups; retirement (Tombstone + patrol) cannot repair it and every health surface reads healthy #3227

Description

@bitflicker64

Bug Type (问题类型)

server status (启动/运行异常)

Before submit

  • I have confirmed and searched that there are no similar problems in the historical issue and documents

Environment (环境信息)

  • Server Version: master 83ef9f3 (Docker Hub :latest of 2026-09-21: pd sha256:69dc1d4e8625, store sha256:7bc0cbd81550, server sha256:9229e5a5f1bf)
  • Backend: HStore, 3 PD + 3 Store + 3 Server, shard count 3, 12 shard groups
  • OS: kind v0.33 (Kubernetes 1.37.0), 12 CPUs, 15G RAM, deployed with the Helm chart from feat(helm): add HStore deployment chart hugegraph/hugegraph#221
  • Data Size: about 9k vertices, 9k edges

Expected & Actual behavior (期望与实际表现)

Rebuilding one Store from scratch (delete its data volume, restart it at the same DNS/raft address, as a Kubernetes StatefulSet does) and retiring the old store id should end with three full replicas per shard group.

Measured on 2026-09-22:

  1. The rebuilt Store registers a new id (7279...) at the same raft address; the old id (1190...) stays with all 12 groups.
  2. POST /v1/store/{oldId} with {"storeState":"Tombstone"} returns 200. patrolPartitions on the PD leader logs shardOffline for all 12 partitions and reallocShards ShardGroup N, add shards from 2 to 3 with the NEW id in the computed list, and fires the conf change.
  3. The Store leader logs Raft N changePeers start, old peer is [... same address ...]: the new peer's endpoint is already in the raft configuration, so jraft has nothing to add and the rebuilt Store never receives a snapshot.
  4. PartitionEngine.onConfigurationCommitted then maps each endpoint back to a store id through the local shard group (PartitionManager.getStoreByRaftEndpoint, hg-store-core, PartitionManager.java:873-881), which still lists the OLD id, and pushes that group to PD (PartitionEngine.java:631-681), overwriting PD's corrected record.
  5. End state, stable for 20+ minutes across three patrols, DELETE /v1/store/{oldId} and balancePartitions (movedPartitions is empty ... newId=0): all 12 groups name a store id that no longer exists, the rebuilt Store answers 500 for every GET :8520/v1/partition/{id} (no partition engines), and every group runs on two live replicas.

The failure is invisible: /v1/stores counts 3 Up, cluster state stays Cluster_OK, Hubble shows 3 Stores UP, all Pods Ready. Writes kept working (2/3 replicas), 0 lost of 4498 acknowledged.

Related: #3150 (closed stale) reported Tombstone not migrating replicas in a 5-store bare-metal cluster; same code area, different trigger. This report is the address-reuse case, which is what any StatefulSet rebuild produces.

Possible directions (not verified): resolve endpoint-to-id against PD's current store list preferring Up over Tombstone; drive remove-then-add when the new id shares an endpoint so the rebuilt Store gets a snapshot; refuse or auto-retire a non-Tombstone store record at registration when a new id claims its raft address.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions