Skip to content

Confirm a recovered cache entry's version before serving it (ADR-0049) - #89

Merged
aandraca-amazon merged 1 commit into
mainfrom
gh64-revalidate-after-restart
Oct 5, 2026
Merged

aandraca-amazon merged 1 commit into
mainfrom
gh64-revalidate-after-restart

Conversation

@aandraca-amazon

@aandraca-amazon aandraca-amazon commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

The cache directory is a hostPath, so a restarted daemon rebuilds both
disk tiers from the previous process's files: the chunk store from its
slot headers, foyer from its blocks. Neither knows whether what it
recovered is still current. An object overwritten while the node was
down came back as its old bytes, and so did one overwritten through
PACER just before a restart: ChunkStore::remove only updates the
in-memory index, and foyer keeps no tombstone for a removed entry.

A cached header is now used only once this process has confirmed its
object against the backend, so the first read of each object after a
restart costs one HeadObject. Every node-mode fill records the version
it filled, and a chunk is served only under the ETag the read
resolved: locally, from a peer, and by a holder's one-sided write into
client memory, whose response now carries the holder's witness. A
mismatched local copy is forgotten so its re-fill can land; a
mismatched peer copy is skipped and its holder sent an Invalidate.
ETags are compared without quotes, because the scatter path tags
chunks with CompleteMultipartUpload's quoted one.

tests/daemon/restart.rs restarts a daemon over its own cache directory
on both tiers: an overwrite while down, an invalidation before the
restart, and an unchanged object, which must still be served from the
recovered tier after one HeadObject and no re-fill.

Testing

  • The pacer-daemon, pacer-transport and pacer-cache test suites all pass. The new restart tests fail on main and pass here, on both disk tiers.
  • Clippy with -D warnings, including --features efa, is clean on Rust 1.96 and 1.99.
  • Not yet run: a daemon restart under load on a real cluster (Test daemon and node failures end to end under load #71). Everything above is in-process.

Fixes #64

🤖 Generated with Claude Code

The cache directory is a hostPath, so a restarted daemon rebuilds both
disk tiers from the previous process's files: the chunk store from its
slot headers, foyer from its blocks. Neither knows whether what it
recovered is still current. An object overwritten while the node was
down came back as its old bytes, and so did one overwritten through
PACER just before a restart: ChunkStore::remove only updates the
in-memory index, and foyer keeps no tombstone for a removed entry.

A cached header is now used only once this process has confirmed its
object against the backend, so the first read of each object after a
restart costs one HeadObject. Every node-mode fill records the version
it filled, and a chunk is served only under the ETag the read
resolved: locally, from a peer, and by a holder's one-sided write into
client memory, whose response now carries the holder's witness. A
mismatched local copy is forgotten so its re-fill can land; a
mismatched peer copy is skipped and its holder sent an Invalidate.
ETags are compared without quotes, because the scatter path tags
chunks with CompleteMultipartUpload's quoted one.

tests/daemon/restart.rs restarts a daemon over its own cache directory
on both tiers: an overwrite while down, an invalidation before the
restart, and an unchanged object, which must still be served from the
recovered tier after one HeadObject and no re-fill.

Fixes #64

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@aandraca-amazon
aandraca-amazon enabled auto-merge (rebase) October 5, 2026 10:56
@aandraca-amazon
aandraca-amazon merged commit 86aab2f into main Oct 5, 2026
8 checks passed
@aandraca-amazon
aandraca-amazon deleted the gh64-revalidate-after-restart branch October 5, 2026 11:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A restarted daemon serves the old bytes of an object overwritten while it was down

1 participant