[agent] Found by the scheduled pnpm bug-hunt routine (ledger #303).
Summary
Since #1039, scan / get in hosted mode run the vendored-to-hosted takeover inside a group commit. The reverted vendored artifacts are deleted after the commit (GroupCommit::defer_removals). The commit journal (.socket/vendor/.commit-journal.json) records only the file changes. If the process dies after the journal is written:
- The next command that takes the apply lock replays the journal: the hosted pin lands and the vendor ledger (
.socket/vendor/state.json) is deleted.
- The queued removal of
.socket/vendor/npm/<uuid>/ (the patched tarball, .gitignore, .gitattributes and socket-patch.vendor.json) is never replayed, so the directory stays.
After that, no command removes it:
scan --prune (hosted, agent or vendored mode, dry or wet) reports vendorOrphanDirs: 0 / removedVendorOrphanDirs: 0. run_vendor_gc returns before its orphan sweep when there is no ledger, which is deliberate: "without a trustworthy ledger it must not delete anything". The takeover commit has just deleted that ledger.
rollback restores the project and exits 0 with success, but leaves the directory.
repair answers redirect_only_project.
The leftover is a committed, socket-owned directory whose own .gitignore re-includes the tarball, so it stays in the repository indefinitely. The logic isn't pnpm-specific: it's the shared group commit and vendor GC. Every ecosystem whose takeover defers artifact removals should behave the same way. I proved it with pnpm only.
Impact
Low severity, but it breaks the crash-safety promise. CLI_CONTRACT.md ("Staged takeover (v5.0)") says "the reverted artifacts are deleted only after it [the commit] … a commit interrupted after its journal was written is finished by the next command that takes the apply lock". Here the commit is only partly finished. Nothing is unpatched or falsely attested. vex reflects the installed tree, and the hosted pin installs patched bytes. The cost is a stale vendored tarball the user can only find and delete by hand.
Repro
This needs a debug build of socket-patch (the SOCKET_PATCH_FAILPOINT crash points are compiled out of release builds), a real pnpm, and a patch API serving a free patch for left-pad@1.3.0. I used a local mock of /v0/orgs/<org>/patches/{batch,package,view} plus /artifacts/<uuid>/<tgz>, with SOCKET_API_URL / SOCKET_PATCH_SERVER_URL / SOCKET_VENDOR_URL pointing at it. A real SIGKILL in the same window has the same effect.
SP=target/debug/socket-patch; PNPM="npx -y pnpm@12.10.1"
W=$(mktemp -d); cd $W
echo '{"name":"fp","version":"1.0.0","dependencies":{"left-pad":"1.3.0"}}' > package.json
$PNPM install >/dev/null
$SP scan --mode vendored --yes --json >/dev/null # vendored: ledger + .socket/vendor/npm/<uuid>/
git init -q && git add -A && git commit -qm vendored
SOCKET_PATCH_FAILPOINT=group_commit_file@2 $SP scan --json >/dev/null 2>&1; echo "crash exit=$?" # 86
$SP scan --json >/dev/null; echo "next locked command exit=$?" # replays the journal
git status --short -- .socket package.json pnpm-lock.yaml pnpm-workspace.yaml
find .socket -type f
$SP scan --mode vendored --prune --dry-run --json | node -pe 'JSON.parse(require("fs").readFileSync(0)).gc.vendorOrphanDirs'
$SP rollback --json >/dev/null; echo "rollback exit=$?"; find .socket -type f | wc -l
Output (the same on every run):
crash exit=86
next locked command exit=0
D .socket/vendor/state.json
M package.json
M pnpm-lock.yaml
M pnpm-workspace.yaml
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/.gitattributes
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/.gitignore
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/left-pad-1.3.0.tgz
.socket/vendor/npm/11111111-2222-4333-8444-555555555555/socket-patch.vendor.json
0
rollback exit=0
4
Control: the same takeover without a crash deletes all four artifact files (git status shows them as D), and rollback is then byte-exact. The replayed state is otherwise correct: a fresh pnpm install --frozen-lockfile against a dead registry installs the patched left-pad.
Expected vs actual
- Expected: replaying the journal finishes the whole commit, including the deferred deletions of the reverted artifacts (CLI_CONTRACT "Staged takeover"). Alternatively, the GC can recognise a
.socket/vendor/<eco>/<uuid>/ directory that no lockfile wires and no ledger records as an orphan, even when the ledger file is absent.
- Actual: the files are rolled forward, but the artifact directory is orphaned, invisible to
scan --prune, and survives rollback.
Matrix (Linux, Node 22, main cf8b164, debug build)
| pnpm |
group_commit_journal |
group_commit_file@1 |
group_commit_file@2 |
No crash (control) |
| 9.15.9 |
artifact dir left |
artifact dir left |
artifact dir left |
deleted |
| 12.10.1 |
artifact dir left |
artifact dir left |
artifact dir left |
deleted |
macOS and Windows weren't probed; this is OS-independent commit logic. First bad commit: likely #1039 (823810a), which introduced the deferred removals. I didn't bisect it.
Suspect code
crates/socket-patch-core/src/utils/group_commit.rs:652 (commit / commit_changes). The journal written around line 790 (journal_bytes(&changes)) carries no record of the remove_after_commit / defer_removal (:452) queue. recover (:1199) can therefore only roll files forward.
crates/socket-patch-cli/src/commands/vendor.rs:4208-4212: run_vendor_gc returns before sweep_orphan_vendor_dirs when the ledger is missing or empty, so a GC pass can't clean up afterwards either.
[agent] Found by the scheduled pnpm bug-hunt routine (ledger #303).
Summary
Since #1039,
scan/getin hosted mode run the vendored-to-hosted takeover inside a group commit. The reverted vendored artifacts are deleted after the commit (GroupCommit::defer_removals). The commit journal (.socket/vendor/.commit-journal.json) records only the file changes. If the process dies after the journal is written:.socket/vendor/state.json) is deleted..socket/vendor/npm/<uuid>/(the patched tarball,.gitignore,.gitattributesandsocket-patch.vendor.json) is never replayed, so the directory stays.After that, no command removes it:
scan --prune(hosted, agent or vendored mode, dry or wet) reportsvendorOrphanDirs: 0/removedVendorOrphanDirs: 0.run_vendor_gcreturns before its orphan sweep when there is no ledger, which is deliberate: "without a trustworthy ledger it must not delete anything". The takeover commit has just deleted that ledger.rollbackrestores the project and exits 0 withsuccess, but leaves the directory.repairanswersredirect_only_project.The leftover is a committed, socket-owned directory whose own
.gitignorere-includes the tarball, so it stays in the repository indefinitely. The logic isn't pnpm-specific: it's the shared group commit and vendor GC. Every ecosystem whose takeover defers artifact removals should behave the same way. I proved it with pnpm only.Impact
Low severity, but it breaks the crash-safety promise. CLI_CONTRACT.md ("Staged takeover (v5.0)") says "the reverted artifacts are deleted only after it [the commit] … a commit interrupted after its journal was written is finished by the next command that takes the apply lock". Here the commit is only partly finished. Nothing is unpatched or falsely attested.
vexreflects the installed tree, and the hosted pin installs patched bytes. The cost is a stale vendored tarball the user can only find and delete by hand.Repro
This needs a debug build of socket-patch (the
SOCKET_PATCH_FAILPOINTcrash points are compiled out of release builds), a real pnpm, and a patch API serving a free patch forleft-pad@1.3.0. I used a local mock of/v0/orgs/<org>/patches/{batch,package,view}plus/artifacts/<uuid>/<tgz>, withSOCKET_API_URL/SOCKET_PATCH_SERVER_URL/SOCKET_VENDOR_URLpointing at it. A realSIGKILLin the same window has the same effect.Output (the same on every run):
Control: the same takeover without a crash deletes all four artifact files (
git statusshows them asD), androllbackis then byte-exact. The replayed state is otherwise correct: a freshpnpm install --frozen-lockfileagainst a dead registry installs the patchedleft-pad.Expected vs actual
.socket/vendor/<eco>/<uuid>/directory that no lockfile wires and no ledger records as an orphan, even when the ledger file is absent.scan --prune, and survivesrollback.Matrix (Linux, Node 22, main
cf8b164, debug build)group_commit_journalgroup_commit_file@1group_commit_file@2macOS and Windows weren't probed; this is OS-independent commit logic. First bad commit: likely #1039 (
823810a), which introduced the deferred removals. I didn't bisect it.Suspect code
crates/socket-patch-core/src/utils/group_commit.rs:652(commit/commit_changes). The journal written around line 790 (journal_bytes(&changes)) carries no record of theremove_after_commit/defer_removal(:452) queue.recover(:1199) can therefore only roll files forward.crates/socket-patch-cli/src/commands/vendor.rs:4208-4212:run_vendor_gcreturns beforesweep_orphan_vendor_dirswhen the ledger is missing or empty, so a GC pass can't clean up afterwards either.