Skip to content

Add a scan benchmark suite and a CI performance gate - #485

Open
Mikola Lysenko (mikolalysenko) wants to merge 2 commits into
mainfrom
claude/great-heisenberg-buor2s
Open

Mikola Lysenko (mikolalysenko) wants to merge 2 commits into
mainfrom
claude/great-heisenberg-buor2s

Conversation

@mikolalysenko

@mikolalysenko Mikola Lysenko (mikolalysenko) commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds crates/socket-patch-bench (unpublished, no new third-party crates in Cargo.lock) and .github/workflows/bench.yml, a CI gate that compares every PR's socket-patch scan performance against its base.

What is measured

The harness generates a seed-deterministic project per package manager in its native lockfile format and install layout, then runs the real binary against a local, std-only mock of the patch API, the public proxy and the artifact host. Package managers: npm, pnpm (isolated store), yarn classic, yarn berry, bun, vlt, pip, uv, PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go, nuget, maven.

  • <pm>/hosted: the default scan on a freshly installed project (crawl, lockfile inventory, batch query, details, references, artifact checks, rewrite, views, writes).
  • <pm>/rescan: the same scan on an already-redirected project (pin discovery and update detection over rewritten locks).
  • npm also runs --dry-run, the public proxy (no token), and 40 ms simulated latency.

Why the numbers can be trusted

  • Every run is validated: scanned, lockfile-only and patched counts; every patch redirected; exactly the expected files rewritten, and those files actually changed on disk; only allowed warning codes; no unexpected requests; dry runs write nothing. A scan that skips work fails the scenario instead of looking fast.
  • Each run starts from the pristine tree (snapshot diff and restore) in an env -i environment. HOME and the caches point into the fixture, telemetry and the update check are off, and proxies point at a closed port. strace shows the CLI spawns nothing.

The gate

The workflow builds base and head with the new perf profile (release codegen, thin LTO) on one runner and interleaves their runs. A PR fails when any of these hold:

  • wall or CPU time slows past 10%, with a 95% sign-test interval above 1, confirmed by a second round of pairs;
  • peak RSS grows past 15%;
  • the head makes more API requests;
  • the head fails a scenario's validation.

The performance-regression-accepted label turns a regression into a report instead of a failure. Pushes to main are recorded as artifacts without gating.

Validation

  • An A/A run (identical binaries) across all 39 scenarios gave no false positives; the worst median drift was +7%.
  • A deliberately slowed binary was flagged at +29% and +77%.
  • A binary that skips work (--ecosystems pypi on an npm project) was rejected as invalid.
  • A full comparison takes about 11 minutes locally on 4 cores; the base build reuses the dependencies the head build compiled.
  • cargo clippy --workspace --all-features -D warnings passes, and the crate's 29 unit tests pass.

Notes

  • Not covered: scan --mode vendored/agent, which need real patched archives, and Deno, which has no hosted rewrite.
  • Benchmark observations: vlt is the most CPU-expensive per package. poetry and pdm are slow for their size. A Go rescan counts the Socket gopatch modules in go.sum as extra lockfile-only packages (412 vs 400 scanned). That looks like a double count; it is encoded as current behavior and not changed here.

🤖 Generated with Claude Code

https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ


Generated by Claude Code


Note

Low Risk
Adds CI tooling and an unpublished benchmark crate; it does not change shipped scan behavior, though flaky gates could block PRs until labeled or tuned.

Overview
Introduces a scan performance gate for Rust changes: a new unpublished workspace crate socket-patch-bench plus .github/workflows/bench.yml.

The harness builds deterministic synthetic projects per package manager (npm family, Python tools, Ruby, PHP, Rust, Go, NuGet, Maven), serves patches from a local HTTP mock, and runs the real socket-patch scan under an isolated environment. Each sample is validated (JSON counts, redirects, rewrites, request patterns) so skipped work cannot look like a win. The CLI supports list, run, compare, and serve (profiling), emitting markdown and results.json.

CI builds base and head with a new [profile.perf] (release semantics, thin LTO), runs interleaved paired comparisons, and fails PRs on statistically significant wall/CPU slowdown, RSS growth, extra API calls, or head validation failures. performance-regression-accepted suppresses timing/memory/request failures only; main pushes record results without gating. Docs now point performance work at this suite; record/replay remains for real API traffic.

Reviewed by Cursor Bugbot for commit 08936bd. Configure here.


Generated by Claude Code

New crate crates/socket-patch-bench (publish = false, no new third-party
dependencies) that benchmarks `socket-patch scan` end to end:

- Synthetic, seed-deterministic projects for every package manager in
  its native lockfile format and install layout: npm, pnpm (isolated
  store), yarn classic, yarn berry, bun, vlt, pip requirements, uv,
  PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go,
  nuget and maven. Per-user caches live under the fixture's own HOME.
- A std-only mock of the patch API, proxy and artifact host that answers
  from the scenario's catalog and counts requests per endpoint.
- Each PM runs a hosted scan of a fresh project and a rescan of an
  already-redirected one; npm also runs --dry-run, the public proxy and
  40 ms simulated latency (request concurrency).
- Every run is validated (scanned/lockfile-only/patched counts,
  redirected patches, exact rewritten files that really changed on disk,
  allowed warnings, no unexpected requests, dry runs write nothing), so
  a scan that skips work fails instead of looking fast. Runs start from
  the pristine tree via snapshot diff/restore, with a cleared env.
- `compare` interleaves base and head runs on one machine and fails on
  a wall/CPU slowdown (median paired ratio past 10% with a 95% sign-test
  interval above 1, confirmed by a second round), peak RSS growth, more
  API requests, or a head that fails validation.

.github/workflows/bench.yml builds the PR's base and head with the new
`perf` profile (release codegen, thin LTO) and runs the comparison on
every PR touching Rust code; pushes to main are recorded without gating.
The `performance-regression-accepted` label turns a regression into a
report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread crates/socket-patch-bench/src/fixtures/npm.rs
Comment thread crates/socket-patch-bench/src/engine.rs Outdated
Intra-store dependency symlinks live at {store}/{key}/node_modules/{dep},
so only the dependency name's scope adds a directory level. Counting the
parent's slashes pointed every link under a scoped parent one directory
too high, leaving those packages with a broken isolated layout.

Also replace the bare unwraps on the per-binary result maps with expect()
messages stating the invariant, per Bugbot review.

Co-Authored-By: Claude <noreply@anthropic.com>
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

bugbot run


Generated by Claude Code

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 08936bd. Configure here.

@mikolalysenko Mikola Lysenko (mikolalysenko) added the Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review label Oct 1, 2026
@mikolalysenko

Copy link
Copy Markdown
Collaborator Author

Burn-down agent: ready for review at 08936bd (08936bd7f3a7d0d3ba1dbffc71e9e59f0962e463).

  • CI: all check runs green on this head (401+ success, 5 skipped, 0 failed).
  • Bugbot: reviewed 08936bd, no new issues. Both findings on 06f2d40 were real and are fixed in 08936bd:
    • pnpm/vlt intra-store dep symlinks counted the parent's scope in the ../ depth, so links under scoped parents pointed one directory too high. Now only the dependency name's scope counts. pnpm/vlt hosted+rescan scenarios still validate locally.
    • Bare unwrap()s on the per-binary result maps are now expect() with the invariant.
  • Mergeable with no conflicts (20 commits behind main; only adds a new crate plus a workflow and profile).
  • Reviewer focus: the gate thresholds in .github/workflows/bench.yml (10% wall/CPU with sign test, 15% RSS, extra API requests) since this becomes a mandatory check, and the performance-regression-accepted label escape hatch.

Slack announcement not sent: no Slack send tool is available to this agent; the next run will retry.


Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Ready for review Agent-verified: mergeable, CI green, Bugbot clean — awaiting human review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants