Skip to content

A vendored-to-hosted takeover interrupted after its commit journal is written leaves the vendored artifact directory behind for good: recovery finishes the files but not the deferred deletions, and no GC can reclaim it once the ledger is gone #1157

Description

[agent] Found by the scheduled pnpm bug-hunt routine (ledger #303).

Summary

Since #1039, scan / get in hosted mode run the vendored-to-hosted takeover inside a group commit. The reverted vendored artifacts are deleted after the commit (GroupCommit::defer_removals). The commit journal (.socket/vendor/.commit-journal.json) records only the file changes. If the process dies after the journal is written:

  1. The next command that takes the apply lock replays the journal: the hosted pin lands and the vendor ledger (.socket/vendor/state.json) is deleted.
  2. The queued removal of .socket/vendor/npm/<uuid>/ (the patched tarball, .gitignore, .gitattributes and socket-patch.vendor.json) is never replayed, so the directory stays.

After that, no command removes it:

  • scan --prune (hosted, agent or vendored mode, dry or wet) reports vendorOrphanDirs: 0 / removedVendorOrphanDirs: 0. run_vendor_gc returns before its orphan sweep when there is no ledger, which is deliberate: "without a trustworthy ledger it must not delete anything". The takeover commit has just deleted that ledger.
  • rollback restores the project and exits 0 with success, but leaves the directory.
  • repair answers redirect_only_project.

The leftover is a committed, socket-owned directory whose own .gitignore re-includes the tarball, so it stays in the repository indefinitely. The logic isn't pnpm-specific: it's the shared group commit and vendor GC. Every ecosystem whose takeover defers artifact removals should behave the same way. I proved it with pnpm only.

Impact

Low severity, but it breaks the crash-safety promise. CLI_CONTRACT.md ("Staged takeover (v5.0)") says "the reverted artifacts are deleted only after it [the commit] … a commit interrupted after its journal was written is finished by the next command that takes the apply lock". Here the commit is only partly finished. Nothing is unpatched or falsely attested. vex reflects the installed tree, and the hosted pin installs patched bytes. The cost is a stale vendored tarball the user can only find and delete by hand.

Repro

This needs a debug build of socket-patch (the SOCKET_PATCH_FAILPOINT crash points are compiled out of release builds), a real pnpm, and a patch API serving a free patch for left-pad@1.3.0. I used a local mock of /v0/orgs/<org>/patches/{batch,package,view} plus /artifacts/<uuid>/<tgz>, with SOCKET_API_URL / SOCKET_PATCH_SERVER_URL / SOCKET_VENDOR_URL pointing at it. A real SIGKILL in the same window has the same effect.

SP=target/debug/socket-patch; PNPM="npx -y pnpm@12.10.1"
W=$(mktemp -d); cd $W
echo '{"name":"fp","version":"1.0.0","dependencies":{"left-pad":"1.3.0"}}' > package.json
$PNPM install >/dev/null
$SP scan --mode vendored --yes --json >/dev/null           # vendored: ledger + .socket/vendor/npm/<uuid>/
git init -q && git add -A && git commit -qm vendored
SOCKET_PATCH_FAILPOINT=group_commit_file@2 $SP scan --json >/dev/null 2>&1; echo "crash exit=$?"   # 86
$SP scan --json >/dev/null; echo "next locked command exit=$?"   # replays the journal
git status --short -- .socket package.json pnpm-lock.yaml pnpm-workspace.yaml
find .socket -type f
$SP scan --mode vendored --prune --dry-run --json | node -pe 'JSON.parse(require("fs").readFileSync(0)).gc.vendorOrphanDirs'
$SP rollback --json >/dev/null; echo "rollback exit=$?"; find .socket -type f | wc -l

Output (the same on every run):

crash exit=86
next locked command exit=0
 D .socket/vendor/state.json
 M package.json
 M pnpm-lock.yaml
 M pnpm-workspace.yaml
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/.gitattributes
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/.gitignore
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/left-pad-1.3.0.tgz
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/socket-patch.vendor.json
0
rollback exit=0
4

Control: the same takeover without a crash deletes all four artifact files (git status shows them as D), and rollback is then byte-exact. The replayed state is otherwise correct: a fresh pnpm install --frozen-lockfile against a dead registry installs the patched left-pad.

Expected vs actual

  • Expected: replaying the journal finishes the whole commit, including the deferred deletions of the reverted artifacts (CLI_CONTRACT "Staged takeover"). Alternatively, the GC can recognise a .socket/vendor/<eco>/<uuid>/ directory that no lockfile wires and no ledger records as an orphan, even when the ledger file is absent.
  • Actual: the files are rolled forward, but the artifact directory is orphaned, invisible to scan --prune, and survives rollback.

Matrix (Linux, Node 22, main cf8b164, debug build)

pnpm group_commit_journal group_commit_file@1 group_commit_file@2 No crash (control)
9.15.9 artifact dir left artifact dir left artifact dir left deleted
12.10.1 artifact dir left artifact dir left artifact dir left deleted

macOS and Windows weren't probed; this is OS-independent commit logic. First bad commit: likely #1039 (823810a), which introduced the deferred removals. I didn't bisect it.

Suspect code

  • crates/socket-patch-core/src/utils/group_commit.rs:652 (commit / commit_changes). The journal written around line 790 (journal_bytes(&changes)) carries no record of the remove_after_commit / defer_removal (:452) queue. recover (:1199) can therefore only roll files forward.
  • crates/socket-patch-cli/src/commands/vendor.rs:4208-4212: run_vendor_gc returns before sweep_orphan_vendor_dirs when the ledger is missing or empty, so a GC pass can't clean up afterwards either.

Activity

  1. mikolalysenko commented on Oct 8, 2026

    @mikolalysenko
    CollaboratorAuthor

    [agent] Triage: priority:p1 (pnpm, npm family; the defect is in the shared group commit, so it applies to every ecosystem whose takeover defers artifact removals). Not a duplicate: related to #809 (pending journal seen by non-locking commands), but this one is the journal not recording the deferred removals at all, so recover cannot finish them. No open PR covers it yet.


    Generated by Claude Code

  2. mikolalysenko commented on Oct 9, 2026

    @mikolalysenko
    CollaboratorAuthor

    v5 triage: closing. Closing as not planned under the requested v5 scope: this needs a process interrupted after journaling and only leaks an unused artifact directory. The issue confirms installed patches/VEX remain correct. Do not expand the transaction journal and crash-recovery GC for this release.

    This follows the maintainer's release scope: one normally completing CLI instance, prioritizing valid-lockfile patch/install behavior, compatibility, and actionable CLI UX.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent:triagedbugSomething isn't workingbughuntFound by a scheduled package-manager bug-hunt agentpm:pnpmpnpmpriority:p3wontfixThis will not be worked on

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions