You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[Bug] A Store rebuilt with an empty volume never rejoins its raft groups; retirement (Tombstone + patrol) cannot repair it and every health surface reads healthy #3227
Rebuilding one Store from scratch (delete its data volume, restart it at the same DNS/raft address, as a Kubernetes StatefulSet does) and retiring the old store id should end with three full replicas per shard group.
Measured on 2026-09-22:
The rebuilt Store registers a new id (7279...) at the same raft address; the old id (1190...) stays with all 12 groups.
POST /v1/store/{oldId} with {"storeState":"Tombstone"} returns 200. patrolPartitions on the PD leader logs shardOffline for all 12 partitions and reallocShards ShardGroup N, add shards from 2 to 3 with the NEW id in the computed list, and fires the conf change.
The Store leader logs Raft N changePeers start, old peer is [... same address ...]: the new peer's endpoint is already in the raft configuration, so jraft has nothing to add and the rebuilt Store never receives a snapshot.
PartitionEngine.onConfigurationCommitted then maps each endpoint back to a store id through the local shard group (PartitionManager.getStoreByRaftEndpoint, hg-store-core, PartitionManager.java:873-881), which still lists the OLD id, and pushes that group to PD (PartitionEngine.java:631-681), overwriting PD's corrected record.
End state, stable for 20+ minutes across three patrols, DELETE /v1/store/{oldId} and balancePartitions (movedPartitions is empty ... newId=0): all 12 groups name a store id that no longer exists, the rebuilt Store answers 500 for every GET :8520/v1/partition/{id} (no partition engines), and every group runs on two live replicas.
The failure is invisible: /v1/stores counts 3 Up, cluster state stays Cluster_OK, Hubble shows 3 Stores UP, all Pods Ready. Writes kept working (2/3 replicas), 0 lost of 4498 acknowledged.
Related: #3150 (closed stale) reported Tombstone not migrating replicas in a 5-store bare-metal cluster; same code area, different trigger. This report is the address-reuse case, which is what any StatefulSet rebuild produces.
Possible directions (not verified): resolve endpoint-to-id against PD's current store list preferring Up over Tombstone; drive remove-then-add when the new id shares an endpoint so the rebuilt Store gets a snapshot; refuse or auto-retire a non-Tombstone store record at registration when a new id claims its raft address.
Bug Type (问题类型)
server status (启动/运行异常)
Before submit
Environment (环境信息)
83ef9f3(Docker Hub:latestof 2026-09-21: pdsha256:69dc1d4e8625, storesha256:7bc0cbd81550, serversha256:9229e5a5f1bf)Expected & Actual behavior (期望与实际表现)
Rebuilding one Store from scratch (delete its data volume, restart it at the same DNS/raft address, as a Kubernetes StatefulSet does) and retiring the old store id should end with three full replicas per shard group.
Measured on 2026-09-22:
7279...) at the same raft address; the old id (1190...) stays with all 12 groups.POST /v1/store/{oldId}with{"storeState":"Tombstone"}returns 200.patrolPartitionson the PD leader logsshardOfflinefor all 12 partitions andreallocShards ShardGroup N, add shards from 2 to 3with the NEW id in the computed list, and fires the conf change.Raft N changePeers start, old peer is [... same address ...]: the new peer's endpoint is already in the raft configuration, so jraft has nothing to add and the rebuilt Store never receives a snapshot.PartitionEngine.onConfigurationCommittedthen maps each endpoint back to a store id through the local shard group (PartitionManager.getStoreByRaftEndpoint, hg-store-core, PartitionManager.java:873-881), which still lists the OLD id, and pushes that group to PD (PartitionEngine.java:631-681), overwriting PD's corrected record.DELETE /v1/store/{oldId}andbalancePartitions(movedPartitions is empty ... newId=0): all 12 groups name a store id that no longer exists, the rebuilt Store answers 500 for everyGET :8520/v1/partition/{id}(no partition engines), and every group runs on two live replicas.The failure is invisible:
/v1/storescounts 3 Up, cluster state staysCluster_OK, Hubble shows 3 Stores UP, all Pods Ready. Writes kept working (2/3 replicas), 0 lost of 4498 acknowledged.Related: #3150 (closed stale) reported Tombstone not migrating replicas in a 5-store bare-metal cluster; same code area, different trigger. This report is the address-reuse case, which is what any StatefulSet rebuild produces.
Possible directions (not verified): resolve endpoint-to-id against PD's current store list preferring Up over Tombstone; drive remove-then-add when the new id shares an endpoint so the rebuilt Store gets a snapshot; refuse or auto-retire a non-Tombstone store record at registration when a new id claims its raft address.