Add a scan benchmark suite and a CI performance gate - #485
Open
Mikola Lysenko (mikolalysenko) wants to merge 2 commits into
Open
Mikola Lysenko (mikolalysenko) wants to merge 2 commits into
Mikola Lysenko (mikolalysenko) wants to merge 2 commits into
Conversation
New crate crates/socket-patch-bench (publish = false, no new third-party dependencies) that benchmarks `socket-patch scan` end to end: - Synthetic, seed-deterministic projects for every package manager in its native lockfile format and install layout: npm, pnpm (isolated store), yarn classic, yarn berry, bun, vlt, pip requirements, uv, PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go, nuget and maven. Per-user caches live under the fixture's own HOME. - A std-only mock of the patch API, proxy and artifact host that answers from the scenario's catalog and counts requests per endpoint. - Each PM runs a hosted scan of a fresh project and a rescan of an already-redirected one; npm also runs --dry-run, the public proxy and 40 ms simulated latency (request concurrency). - Every run is validated (scanned/lockfile-only/patched counts, redirected patches, exact rewritten files that really changed on disk, allowed warnings, no unexpected requests, dry runs write nothing), so a scan that skips work fails instead of looking fast. Runs start from the pristine tree via snapshot diff/restore, with a cleared env. - `compare` interleaves base and head runs on one machine and fails on a wall/CPU slowdown (median paired ratio past 10% with a 95% sign-test interval above 1, confirmed by a second round), peak RSS growth, more API requests, or a head that fails validation. .github/workflows/bench.yml builds the PR's base and head with the new `perf` profile (release codegen, thin LTO) and runs the comparison on every PR touching Rust code; pushes to main are recorded without gating. The `performance-regression-accepted` label turns a regression into a report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ
Intra-store dependency symlinks live at {store}/{key}/node_modules/{dep},
so only the dependency name's scope adds a directory level. Counting the
parent's slashes pointed every link under a scoped parent one directory
too high, leaving those packages with a broken isolated layout.
Also replace the bare unwraps on the per-binary result maps with expect()
messages stating the invariant, per Bugbot review.
Co-Authored-By: Claude <noreply@anthropic.com>
Collaborator
Author
|
bugbot run Generated by Claude Code |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 08936bd. Configure here.
Collaborator
Author
|
Burn-down agent: ready for review at
Slack announcement not sent: no Slack send tool is available to this agent; the next run will retry. Generated by Claude Code |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
crates/socket-patch-bench(unpublished, no new third-party crates inCargo.lock) and.github/workflows/bench.yml, a CI gate that compares every PR'ssocket-patch scanperformance against its base.What is measured
The harness generates a seed-deterministic project per package manager in its native lockfile format and install layout, then runs the real binary against a local, std-only mock of the patch API, the public proxy and the artifact host. Package managers: npm, pnpm (isolated store), yarn classic, yarn berry, bun, vlt, pip, uv, PEP 751 pylock, poetry, pipenv, pdm, bundler, composer, cargo, go, nuget, maven.
<pm>/hosted: the defaultscanon a freshly installed project (crawl, lockfile inventory, batch query, details, references, artifact checks, rewrite, views, writes).<pm>/rescan: the same scan on an already-redirected project (pin discovery and update detection over rewritten locks).--dry-run, the public proxy (no token), and 40 ms simulated latency.Why the numbers can be trusted
env -ienvironment.HOMEand the caches point into the fixture, telemetry and the update check are off, and proxies point at a closed port. strace shows the CLI spawns nothing.The gate
The workflow builds base and head with the new
perfprofile (release codegen, thin LTO) on one runner and interleaves their runs. A PR fails when any of these hold:The
performance-regression-acceptedlabel turns a regression into a report instead of a failure. Pushes tomainare recorded as artifacts without gating.Validation
--ecosystems pypion an npm project) was rejected as invalid.cargo clippy --workspace --all-features -D warningspasses, and the crate's 29 unit tests pass.Notes
scan --mode vendored/agent, which need real patched archives, and Deno, which has no hosted rewrite.gopatchmodules ingo.sumas extra lockfile-only packages (412 vs 400 scanned). That looks like a double count; it is encoded as current behavior and not changed here.🤖 Generated with Claude Code
https://claude.ai/code/session_017hKPGJkJfyhsQrzXYf7RJZ
Generated by Claude Code
Note
Low Risk
Adds CI tooling and an unpublished benchmark crate; it does not change shipped scan behavior, though flaky gates could block PRs until labeled or tuned.
Overview
Introduces a
scanperformance gate for Rust changes: a new unpublished workspace cratesocket-patch-benchplus.github/workflows/bench.yml.The harness builds deterministic synthetic projects per package manager (npm family, Python tools, Ruby, PHP, Rust, Go, NuGet, Maven), serves patches from a local HTTP mock, and runs the real
socket-patch scanunder an isolated environment. Each sample is validated (JSON counts, redirects, rewrites, request patterns) so skipped work cannot look like a win. The CLI supportslist,run,compare, andserve(profiling), emitting markdown andresults.json.CI builds base and head with a new
[profile.perf](release semantics, thin LTO), runs interleaved paired comparisons, and fails PRs on statistically significant wall/CPU slowdown, RSS growth, extra API calls, or head validation failures.performance-regression-acceptedsuppresses timing/memory/request failures only;mainpushes record results without gating. Docs now point performance work at this suite; record/replay remains for real API traffic.Reviewed by Cursor Bugbot for commit 08936bd. Configure here.
Generated by Claude Code