Skip to content

docs: v5 waste review + repacking-to-depscan design - #286

Draft
Mikola Lysenko (mikolalysenko) wants to merge 3 commits into
release/v5-prereleasefrom
v5/waste-review
Draft

Mikola Lysenko (mikolalysenko) wants to merge 3 commits into
release/v5-prereleasefrom
v5/waste-review

Conversation

@mikolalysenko

Copy link
Copy Markdown
Collaborator

Docs only. No product code changes.

What's here

  • docs/design/v5-waste-review.md is a holistic waste review of socket-patch and depscan's patch system. 106 findings were checked on three criteria (re-measured facts, whether it is really waste under the owner decisions, and the savings estimate). A finding was kept when at least two held: 80 kept, 26 dropped.
  • docs/design/repacking-to-depscan.md is the chosen design for moving artifact repacking out of the CLI and into depscan. Four variants were scored on integrity, ops and delivery.
    • Current-state matrix per ecosystem and format.
    • An additive v2 contract on /v0/orgs/{slug}/patches/package.
    • Offline and airgap handling, outage idempotence, signing and provenance, and storage cost.
    • Quantified deletions.
    • A 30-step migration across both repos.
    • Open questions for the owner.

Headline numbers

  • Net-new savings, beyond what is already planned: about 32.8k lines actionable now (9.1k code, 14.9k test, 8.8k docs/YAML), plus 15.7k gated on a precondition or owner decision.
  • ci.yml: about 478 job-min saved per PR run and about 438 per main push, against a 945 job-min main-push run. Add about 95 job-min per PR push that triggers the compatibility workflows.
  • Biggest lever: building the e2e binaries once and fanning out (F51, about 280 job-min per run).
  • Repacking: server-primary with a single pypi local class. It deletes about 2,050 CLI src LOC and about 1,260 test LOC (net about −1,150 src). CI saving is about 0; this is a correctness and maintenance change.
  • Do first: the v5 base is red on crates/socket-patch-cli/tests/e2e_redirect_cargo_build.rs:829, so the e2e tier has not run on any v5 PR (F77).

All owner decisions in v5-plan.md are respected: vlt and every PM version stay (CI is tiered, not dropped), setup and the hosted ledger are removed, and maven/nuget stay frozen. Deleting the local rebuild is gated on the owner reversing the WS5 caveat after #283.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BgB7r9Ksfn5MJ6vWTFwRbL


Generated by Claude Code

Two design docs for the v5 train, docs only:

- v5-waste-review.md ranks 69 verified waste findings across
  socket-patch and depscan's patch system, with evidence, risk, and
  savings de-duplicated against WS1-WS8 and PRs #279-#283. Top-10
  cuts, totals, and the refuted findings so they are not re-raised.
- repacking-to-depscan.md is the chosen design for moving artifact
  repacking to depscan: server-primary with a single pypi local class,
  an additive v2 package contract, signed statements, offline bundles,
  and a phased migration across both repos gated on the WS5 caveat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BgB7r9Ksfn5MJ6vWTFwRbL
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] CI note: coverage is red, and test (*) / test-release will be too. None of it is caused by this PR, which only adds two docs under docs/design/.

  • The failures are cargo_get_uuid_hosted_fresh_checkout_fetch and cargo_hosted_fresh_checkout_fetch_pulls_patched_crate_and_vex_verifies in crates/socket-patch-cli/tests/e2e_redirect_cargo_build.rs. The assertion that fails is case (3), offline with no ledger, at around :829. It expects exit 1 and gets 0.
  • The base release/v5-prerelease (8ae7dc3) fails the same way in run 36352437716: the same two tests fail in coverage, test-release and test (ubuntu/macos/windows).
  • No fix I can port exists yet. v5 WS1+WS2: ledger-free hosted mode, upstream-restore rollback, vendor ejects hosted projects #280 (v5/ledger-free-hosted) rewrites this test as part of removing the ledger (823ce8a), and that rewrite only makes sense on top of that removal. The failure is deterministic, so I am not re-running it.
  • The waste review records this as F77. The fix is to settle the offline-VEX-without-ledger semantics in WS1 and then rebase the v5 PRs.

Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] follow-up instructions v1

Shared instructions for the cloud follow-up routines created from this waste review (2026-09-28). This PR itself is NOT merged; it is closed once every finding has an owner. Each routine's own prompt names its workstream, finding IDs, repo and target branch.

You implement ONE workstream of findings from the socket-patch v5 waste review, SocketDev/socket-patch PR #286 (read docs/design/v5-waste-review.md on branch v5/waste-review; for depscan items also docs/design/repacking-to-depscan.md). Each finding row has evidence (file:line, runs) and a recommendation. Work fully autonomously; the owner is traveling.

1. Re-verify each assigned finding against the CURRENT target branch first (the review was written against older heads; #280 has since landed and #283/#281/#279/#282 are landing now). Drop any finding that no longer holds and say why in the PR body. Respect the owner decisions in docs/design/v5-plan.md: vlt and every package-manager version stay (CI is tiered, not dropped), setup and the hosted ledger are removed, maven/nuget vendoring stay frozen.
2. Keep scope tight: only your assigned findings plus trivial same-file fallout. Do not touch files that an open v5 PR (#279 #281 #282 #283) is still rewriting unless your prompt says to wait for those to land; if you must, wait (poll every ~15 min) and merge the base after they land. Integrate with merges only - never rebase, never force-push.
3. Branch from the target branch named in your prompt. Tests: run the affected crates/workspaces' tests, clippy/lint, and for CI changes prove the new workflow is equivalent (same jobs covered, same failure modes caught) with a run link. Deletions must show nothing still references the removed code (grep + build).
4. Review your diff adversarially before opening the PR (parallel subagents: correctness/regression, coverage lost, CI equivalence, owner-decision conflicts); fix confirmed findings.
5. Commits signed/verified. socket-patch: subject <=50 chars, body wrap 72, end with `Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>`. depscan: follow AGENTS.md (subject <=50, NO Co-Authored-By, trailer `Assisted-by: Claude Code:claude-opus-5-5`; never export plain helpers from pipeline task modules).
6. Open the PR against the target branch in your prompt. Title as given in your prompt. Body: one table of finding ID -> done / dropped (reason) / deferred, measured savings (LOC, CI job-min with run links), and test evidence; end with `🤖 Generated with [Claude Code](https://claude.com/claude-code)`. Start as DRAFT.
7. Get CI green (fix real failures; rerun obvious flakes once; failures that also fail on the base's latest CI are base-inherited - say so). Then:
   - socket-patch PRs into release/v5-prerelease: mark ready for review (draft=false) and post a comment whose first line is `[agent] ready to land`. The hourly v5 landing coordinator squash-merges it into release/v5-prerelease. NEVER target or merge into main.
   - depscan PRs: leave them as DRAFT for owner review; do not merge.
8. Finally comment on socket-patch #286 with one line: workstream name, PR link, and per-finding outcome counts.

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] triage addendum

I checked all 69 ranked findings and the 11 measurement/checked-not-waste entries in docs/design/v5-waste-review.md against the workstream list in [agent] follow-up instructions v1. One ranked finding has no home. Two are assigned twice. None of the 11 measurement entries is assigned. Decisions are below. Each workstream routine should treat the rows marked for it as part of its scope.

Missing

  • F63 (five drafts cut in parallel from one base on the same hot files; recommendation: rebase serially and skip compatibility workflows on drafts) → W0 v5 fix: unblock the e2e tier on the v5 base (v5 fix: unblock the e2e tier on the v5 base #288). It is a subset of F78, which W0 already owns: the "skip compatibility workflows on drafts" half goes in the same change as F78's draft gating. The serial-rebase half is process only and needs no code. If W0 has already finished, it should mark F63 as done with F78, or as deferred, in its PR table.

Double-assigned

Measurement / checked-not-waste entries (no code unless noted)

No finding needs a new workstream, so there is no owner TODO for a new routine. The full map and the owner decisions follow in [agent] triage map.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] triage map

Every finding in docs/design/v5-waste-review.md has exactly one home (this includes the corrections in [agent] triage addendum). Key: W0–W6 are the follow-up routines listed in [agent] follow-up instructions v1. in-flight #N means a follow-up note routed to that PR's session. owner-gated means no agent will act until the owner decides.

Workstreams: W0 v5 fix: unblock the e2e tier on the v5 base (#288) · W1 v5 CI: build e2e binaries once and tier the PM matrix · W2 v5 tests: one suite per command, retire #257 oracles · W3 v5: remove dead code and v3 compatibility shims · W4 v5: fix partial-stage repair bug, cut redundant downloads · W5a Remove dead patch-system code and stale refactor docs (depscan) · W5b Cut redundant patch pipeline downloads and CI work (depscan) · W6 Prepare depscan for the socket-patch v5 CLI bump (depscan)

id finding home
F01 GitHub-app hosted PR flow keeps a ~6k-LOC TS port of the Rust rewriter owner-gated
F02 depscan patches/docs/refactor: 9.4k lines; 61% of cited paths no longer exist W5a (depscan)
F03 ~7.5k LOC of test-only "verbatim previous implementation" oracles from #257 W2
F04 covgap_/coverage_fix_/in_process_ categories no longer mean anything W2
F05 two builders per archive-shaped artifact produce different bytes owner-gated
F06 depscan legacy/ hosts the live admin package-detail engine; it blocks the 2026-11-04 table drop W5a (depscan)
F07 real-installer pdm/pipenv/poetry backtests duplicated in depscan W6 (gated: capture contract; F74 is the interim step, in W5b)
F09 Python lock family missing from WS3's PR order; 4–5 readers per format W3 (after #281; the #281 note sets ordering only)
F10 CRLF/corrupt/dry-run/round-trip invariants re-tested per writer W2
F11 depscan break list for the #279 bump: 30 loud setup e2e failures W6 (depscan)
F12 8 npm-family vendored backends repeat one vendor/revert skeleton W3
F13 shared redirect goldens cover ecosystems unevenly W6 (only if F01 is not done)
F14 depscan diff channel: on-demand bsdiff task, hidden route, proxy preflight W5b (interim fixes only; retiring it is an owner decision)
F15 setup leftovers: frozen hook wheel and Bundler plugin source, plus SetupConfig in-flight #279 note (owner confirms)
F16 WS1 upstream/* builds a second per-format grammar outside WS3 formats/ in-flight #281 note
F17 napi addon + hosted-bundle have 0 consumers; parity upkeep in 3 PRs owner-gated; F17(ci) path-gating → W1; the #282 note is sequencing only
F18 one-off cargo index-line backfill runs inside every converter cycle W5b (gated: backlog = 0)
F19 admin package-detail "lite" engine dead since #18370 W5a (depscan)
F20 purl-api-proxy keeps a full /patch/* forward that no client calls W5a (after a traffic check)
F21 hidden legacy scan spellings and the v3 SOCKET_PATCH_* env shim W3
F22 14 of 18 CLI telemetry events are never classified server-side W6 (owner decision: classify or delete)
F23 deploy-patches prod/staging workflows are 500-line twins W5a (depscan)
F24 patches-harness-contract carries a dead autopatch-backport config module W5a (depscan)
F25 parse_memo: process-global parse caches W3
F26 #282 moves the hosted ledger into core while #280 deletes it in-flight #282 note
F27 the GitHub app mirrors the redirect-state.json schema that WS1 deletes W6 (depscan)
F28 vendored path POSTs /patches/package once per uuid W4
F29 --one-off on get/rollback exists only to error W3
F30 .socket/packages/<uuid>.tar.gz is never written but still probed W3
F32 Auto/Service fallback policy hand-rolled in 7 backends W3
F33 npm .berry.zip sidecar is built, stored, served, never read W3 (cli) + W5a (server): intentional split
F34 every stored object is streamed back in full after upload W5b (depscan)
F35 Berry 10c0 checksum implemented twice with hand-copied fixtures owner-gated
F36 two npm-registry tarball fetchers inside depscan W5b (depscan)
F37 serve-route.ts re-inlines the stream/304/HEAD logic W5b (depscan)
F38 signature handling diverges; depscan signer seam and columns are dead W5a (depscan)
F39 3 regenerate/reset paths; admin routes skip the gopatch refusal W5b (depscan)
F40 setup-matrix (9 docker legs) is continue-on-error and tests setup in-flight #279 (WS7)
F41 e2e-docker is a strict subset of coverage-docker W1
F43 public core fns with zero callers W3
F44 #279 WS8 edits hosted text and tests that #280 deletes and #282 moves in-flight #279 note
F45 publish dry run re-extracts the file map and re-unpacks upstream W5b (depscan)
F46 e2e_maven/e2e_nuget/e2e_composer rows on 0.03 s hermetic tests W1
F47 cargo-vex 17-leg cross recompiles every leg W1
F48 module-wide #[allow(dead_code)] on vlt_lock_text W3
F49 default suite compiled twice per test leg W1
F50 next depscan bump fails golden.test.ts on the 68 new vlt cases W6 (depscan)
F51 e2e matrix recompiles the CLI and its test binary in each of 176 legs W1
F52 tier PM-version legs: boundaries on PRs, every version on main/nightly W1
F53 Dockerfile.base release build rebuilt with no cache 27× per run W1
F54 pdm-compat Windows native legs vacuous (47/47 SKIP, green) W1
F55 pdm capstone recompiles e2e_vex_build in 30 legs and repeats 7 ci.yml rows W1
F56 vlt capstones run in ci.yml (35 rows) and in vlt-compat install-proof (70 legs) W1
F57 vlt-serve-watchdog cannot alert; npm/pnpm compatibility have no path filter W1
F58 238 CLI integration-test binaries, 78 with ≤3 tests W2
F59 after WS1, vex refetches every record per run W4
F60 hosted scan downloads each patched wheel in full to read METADATA W4
F61 one publish downloads upstream up to 3× and repacks twice W5b (depscan)
F63 five drafts cut in parallel from one base on the same hot files W0 (#288), with F78 (see addendum)
F66 converter reads patched blobs one at a time W5b (depscan)
F67 disk stager treats diff-archive presence as full coverage W4
F69 pypi bz2/xz sdists are selected, downloaded, then refused W5b (depscan)
F70 same waste seen from the vendored download path (one fix with F71) W4
F71 qualifier-less pypi patches: CLI downloads the served sdist before rejecting it W4
F73 depscan api-v0 shards rebuild the pinned CLI every run W5b (depscan)
F74 depscan pdm/hatch workflows trigger on pypi task code they never run W5b (depscan)
F77 one red base test skips the e2e tier on all v5 PRs W0 (#288)
F78 in-flight v5 PRs: 34% of runs cancelled or failed W0 (#288)
F80 auto mode eagerly fetches pristine upstream for uninstalled packages W4
F08 measurement done with #280 (WS1 landed)
F31 measurement in-flight #282
F42 measurement W3 (v5 breaking-change notes)
F62 measurement none (not waste; context for #281)
F64 measurement reference only
F65 measurement context for #283
F68 measurement W5b guardrail for F14
F72 measurement owner-gated with F05
F75 measurement W5b guardrail (keep the repack tests)
F76 measurement reference only
F79 measurement W6 (bump sequencing)

Counts (ranked findings): W0 3 · W1 11 (+F17 ci) · W2 4 · W3 10 (F33 cli half) · W4 7 · W5a 8 (F33 server half) · W5b 12 · W6 6 · in-flight notes 5 (#279: F15 F40 F44; #281: F16; #282: F26) · owner-gated 4 (F01 F05 F17 F35). Total: 69, with F33 split between two workstreams.

Owner decisions needed

Gated, with no agent assigned:

  1. Repacking-to-depscan migration (docs/design/repacking-to-depscan.md): do we reverse the WS5 caveat and move all repacking to depscan, which drops the CLI local rebuild? This needs the service-hit vs local-fallback rate first, and that rate is still unmeasured. It also needs an answer for F72's fail-closed classes (pypi qualifier-less/wheel-only, locally pinned wheels/tarballs, bzip2 sdists).
  2. F01: should depscan's GitHub-app hosted PR flow call the Rust engine (socket-patch-node or the hosted-bundle subprocess) and delete the ~10.5k-line TS rewriter twin? If yes, which engine, and after which CLI bump? Until you decide, the TS port stays frozen. If the answer is no, F13 (shared goldens, W6) is the fallback.
  3. F05: while the local rebuild stays, should someone align the npm recipe (upstream entry order, mtime 0) so that both builders produce the same bytes and reuse.rs shrinks? Or do we leave it until decision 1?
  4. F35: keep both Berry 10c0 checksum implementations and only share the fixtures? And should vendored mode take the server's yarnBerry10c0 for service artifacts?
  5. F17: remove the napi addon and hosted-bundle, which have 0 consumers, or keep both engines as WS4 decided until depscan adopts one (F01)? The CI path-gating goes ahead in W1 either way.

Decisions that workstreams will surface in their PRs:
6. F14 (W5b): is the CLI diff mode staying? If the diff channel is retired, is a 404 the interim? W5b does only the interim cleanups (it removes the task throws and the doubled preflight).
7. F22 (W6): should the v5 telemetry events be classified and moved to /v1/orgs/{slug}/events (or the sunset extended)? Or should the unread track_* wrappers be deleted before the 2026-12-31 sunset?
8. F11 (W6): can 65_patch-telemetry.js's patch_setup case be retired?
9. F15 (#279): confirm deleting pypi/socket-patch-hook and gem/socket-patch-bundler.
10. F49 (W1): confirm that the #150 two-feature-set compile rationale no longer holds.
11. Production-data checks via Grafana/gateway logs (agents cannot run these): F20 (purl-api /patch/* traffic), F33 server (.berry.zip fetches), F18 (converter backfill backlog = 0), F71 (bare-purl pypi share).

Both docs are being moved to a docs-only reference PR against release/v5-prerelease so that closing this PR loses nothing.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] W6 (depscan v5 bump prep): https://github.com/SocketDev/depscan/pull/26900 (draft). Done 4 (F11, F27, F50, F22), partial 1 (F07; script deletion is still gated on socket-patch captures), deferred 1 (F13; waits for #281), dropped 0.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] CI note on c6c708e (merged the new base 686e5fb, with #280 landed): test (windows-latest) fails in -p socket-patch-cli --test covgap_commands_scan_mod.

  • Not caused by this PR, which only adds docs. scan_hosted_paths_run_once_per_project_directory compares against a forward-slash apps/a header, but Windows prints apps\a. The base has failed this way since 62f07c7.
  • The fix is bd632f5 "Expect native path separators in scan headers" in v5 fix: unblock the e2e tier on the v5 base #288 (v5/unblock-e2e-tier, W0). I didn't port it here, because this session isn't cleared to put another PR's commit on this branch. This PR goes green once v5 fix: unblock the e2e tier on the v5 base #288 lands and the base is merged in again.
  • The earlier e2e_redirect_cargo_build failure no longer appears on the new base.

Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] W5b (depscan pipeline/CI waste): drafts SocketDev/depscan#26898 (CI) and SocketDev/depscan#26903 (pipeline), both green. 5 done (F39 Go-refusal fix, F34, F66, F69, F74), 1 partly done (F36), 1 documented (F14), 5 deferred (F73 and F37 until #26860 and the annotation stack land, F61 with F45 needing an owner design call, F18 until its backlog drains).


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] W1 (v5 CI build-once + tiering): #291. Outcomes: 10 done (F51 F49 F52 F41 F53 F46 F47 F55 F54 F56), 1 partial (F57: the watchdog stays disarmed until the prod probe passes), 1 deferred (F40, left to #279). Measured: 945 → 399 job-min for the full tier; about 350 on a PR.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] W0 (fast lane): #288 is ready to land. F77 done: #280 fixed the cargo case, and #288 fixes the Windows covgap_commands_scan_mod test, so the e2e tier now runs for the first time on v5. It also fixes 3 stale e2e tests, and the tier surfaced #280 regressions (vlt/uv hosted rollback, yarn < 1.10) that are triaged on #288 for WS1. F78 done: concurrency cancels only superseded PR runs; draft-gating is deferred. 2 done, 0 dropped, 1 deferred.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] CI note on 50aed71 (merged the new base 06437d2, where #283 landed): coverage fails in socket-patch-core --lib at vendor::registry_fetch::tests::stage_local_artifact_caps_oversized_artifact_before_buffering, registry_fetch.rs:3334. The child process peaked at 610 MB RSS against a 512 MB bound.


Generated by Claude Code

@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

[agent] W4 (partial-stage repair bug + redundant downloads): #292 — done 5 (F67, F28, F70, F71, F80), dropped 0, deferred 2 (F59, F60)


Generated by Claude Code

Mikola Lysenko (mikolalysenko) added a commit that referenced this pull request Sep 28, 2026
Keep the v5 waste review and the repacking-to-depscan design from
PR #286 as reference docs. #286 is being closed without merging, and
its findings now belong to follow-up PRs. Each doc starts with a line
that links the triage map on #286, which gives every finding's owner.

Co-authored-by: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants