diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f52cefc..89a60b2 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -65,13 +65,13 @@ jobs: - run: cargo +stable test --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked -- --test-threads=1 msrv: - name: Rust 1.97.1 MSRV + name: Rust 1.96.0 MSRV runs-on: ubuntu-latest steps: - uses: actions/checkout@v5 - uses: dtolnay/rust-toolchain@master with: - toolchain: 1.97.1 + toolchain: 1.96.0 - uses: Swatinem/rust-cache@v2 with: shared-key: msrv diff --git a/CHANGELOG.md b/CHANGELOG.md index a768a7a..5fe7141 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -3,6 +3,155 @@ All notable public changes to Code System Graph are documented in this file. Code System Graph follows Semantic Versioning. +## [1.2.0] - 2026-10-03 + +### Changed + +- The SQLite store keeps one current graph per workspace and publishes each scan as a delta: + only inserted, changed, and removed nodes, edges, evidence, fingerprints, and extractor batches + are written. An unchanged republication writes no graph rows, and the database no longer grows + with the number of scans. +- The database records a schema identity derived from the embedded schema definition. `status`, + `doctor`, restore, and MCP status report `schema_id`. +- Communities are recomputed only when graph topology changes; the analyses of the current and the + previous snapshot are retained for comparison. +- FTS5 search rows are maintained by triggers keyed by node row. +- Every store read runs in one SQLite read transaction, so readers never observe a snapshot that a + concurrent publication replaced between statements. +- Scans fingerprint each physical file once and reuse its content hash from a per-file stat cache + (size, modification time, and, on Unix, device, inode, and change time) when unchanged; files + modified within two seconds of the cached observation are always read. +- Changed files are read once and shared by all of their extractors. Extraction runs on up to + `executionPolicy.maxExtractionWorkers` threads (default 8, bounded by available parallelism) and + merges results in artifact-key order, so the published graph is identical for every worker count. +- Event and literal-SQL source extractors skip files that contain none of their recognizer + markers, and Rust files are parsed for database calls only when they reference `sqlx` or + `mysql_async`. +- Checkpointed batches are written in one sidecar transaction and stored as raw payload BLOBs; + only artifacts that differ from the published graph are looked up in the checkpoint cache. +- Targeted `--repository` scans discover and fingerprint only the selected repository. +- The scan no longer stages a complete copy of the candidate graph in the operational sidecar + before publication. +- `ExecutionSummary` reports `statCacheHits`, `extractionWorkers`, `contentBytesRead`, + `publishedRows`, and per-phase wall time and resident memory in `phases`. +- Source files whose GraphQL, event, generated-protobuf, or literal-SQL scan finds no facts no + longer add a per-file artifact node and `contains` edge; their batches are persisted without + outputs, and reusing a batch without outputs decodes nothing. Every registered repository has a + repository node. +- A scan whose snapshot identity (a hash of the manifest, the extraction contract, the budgets, and + every artifact fingerprint) equals the current snapshot reuses it without loading the previous + fingerprints or any extractor batch. The snapshot identifier now covers artifact paths and sizes + in addition to content hashes. +- Incremental scans publish fingerprints and extractor batches as a delta planned against the + stored fingerprints; `--force` scans compare every stored artifact row. The incremental plan + lists only added, modified, and deleted artifacts, and `ArtifactChangeKind::Unchanged`, the + unused extractor-batch planner (`plan_extractor_batches`, `ExtractorBatchPlan`, `PlannedBatch`, + `BatchAction`), and `ArtifactKey` are removed from the public API. +- Community detection evaluates each Louvain move from incremental modularity gains and community + degree totals instead of recomputing modularity for every candidate. +- The scan worker reports progress at most every 50 ms. The supervisor's memory sampler reads only + process memory and parent links instead of listing every thread and reading CPU, disk usage, and + executable paths of every process, which cuts each sample by about four times on a desktop host + and keeps the no-progress watchdog on schedule under CPU contention. +- CodeGraph corroboration, extraction planning, and extractor-run accounting index artifacts and + graph elements by borrowed keys in `foldhash` maps instead of scanning or cloning them per + artifact, and reused extractor batches are moved rather than copied. +- Workspace registration inspects repositories in parallel and caches each `origin` remote by the + modification time and size of its Git config. `sync` resolves the workspace once per pass and runs + CodeGraph status and sync for up to four repositories concurrently. +- HTTP operations are identified by method and canonical route shape, so `{id}`, `:id`, + ``, `{id:int}`, `[id]`, and catch-all forms such as `{*rest}` or `[...slug]` declare the + same operation. One per-method route index resolves `calls_remote`, `validates`, and + `implemented_by`: concrete paths such as `/orders/42` match templates, the most specific template + wins, and equally specific providers are narrowed to an explicit repository restriction or the + caller's repository before being reported as ambiguous with their candidates. +- Route parameters embedded in a segment preserve its literal prefix and suffix. For example, + `/files/{name}.json` matches `/files/readme.json`, rejects `/files/readme.xml`, and has a distinct + identity from `/files/{id}`. Static segments take precedence over mixed segments, then whole + parameters and catch-alls; overlapping mixed templates are reported as ambiguous. +- Relationships are relinked from all current batches on every scan, so incremental and full scans + publish identical graphs. +- `sync --watch` passes after the initial one discover and synchronize only the repositories touched + by the coalesced filesystem events; manifest edits, ignore-rule changes, event overflow, and + watcher errors widen the pass to the whole workspace. +- Absolute consumer URLs are linked by path and scoped by their authority. Loopback hosts resolve + across the workspace, the new per-repository `authorities` manifest field restricts a host to one + repository, Compose and Kubernetes service names are inferred as authorities of the declaring + repository, and any other host is classified external instead of being reported as a call + without a provider. +- TypeScript, JavaScript, Go, and Java source extraction no longer treats `//` inside string + literals as a comment, so absolute URLs are recognized. +- Server routes are published with their full path. Router prefixes are composed per repository, + across files, for FastAPI `include_router`, Flask blueprints, Express `use`, NestJS controllers + and global prefixes, Spring and Feign class-level mappings, Gin and Chi groups and mounts, Axum + `nest`/`merge`, and Actix Web `scope`/`service`/`configure`. Go 1.22 `net/http` method patterns + such as `"GET /orders/{id}"` are recognized. +- Client URLs are evaluated instead of requiring one string literal: constants and variables bound + earlier in the same function or file, concatenation, Python f-strings, `%` and `.format`, Rust + `format!` and `concat!`, JavaScript template literals, Go `fmt.Sprintf`, Java `String.format`, + and conversions such as `encodeURIComponent` or `strconv.Itoa`. Runtime values in the scheme, + the authority, or a whole path segment keep the path exact, with the segment as a parameter. +- Client calls are composed per repository through wrapper functions: a call whose URL depends on a + parameter of its function is instantiated at every call site that binds it, also through + wrappers of wrappers and across modules. Tests are linked to the endpoints reached through the + helpers they call and the pytest fixtures they request, including fixtures in `conftest.py`. +- TypeScript and JavaScript clients are recognized per call instead of per statement, so every + `fetch` in a callback is reported; Axios instances created with `axios.create({ baseURL })`, + Axios config-object requests, Go `http.Client` receivers, `http.Head`, `http.PostForm`, and + `http.NewRequestWithContext` are recognized, and client calls carry their enclosing function. +- Tests are recognized in every supported language: Jest, Vitest, Mocha, and Playwright + `describe`/`it`/`test` blocks, identified by file, `describe` chain, and title and linked to the + `beforeEach`/`beforeAll` blocks that run before them; Go `TestX(t *testing.T)` functions; and + JUnit `@Test`, `@ParameterizedTest`, and `@RepeatedTest` methods. +- In-process test clients are recognized and resolve only against providers in their own + repository: Python `TestClient`, Flask `test_client()`, and HTTPX with `app=` or an ASGI/WSGI + transport, including pytest fixtures that return them; supertest and Playwright `request`; Go + `httptest.NewRequest` and `httptest.NewServer` URLs; Spring MockMvc, RestAssured, + `WebTestClient`, `RestTemplate`, and `TestRestTemplate`; Axum `Request` builders sent with + `oneshot`; and Actix `test::TestRequest`. +- Helm templates are recognized only as YAML or `.tpl` files under the `templates` directory of a + chart with `Chart.yaml`, so Jinja `*.j2` files and other `templates` directories are no longer + parsed as Helm. Path classification uses repository-relative paths only. +- Vendored, theme, bundled, and minified JavaScript is no longer scanned for literal SQL. +- Route handlers are resolved through middleware arguments, member references such as + `orders.list`, single-argument wrappers, FastAPI `add_api_route`, Flask `add_url_rule`, and Spring + mapping annotations followed by other annotations; inline handlers are identified by method and + path. +- The CodeGraph adapter is validated against CodeGraph 1.6.1. The structured CLI contract accepts + 1.6.1 and later 1.6.x patch releases; CodeGraph 1.5 and earlier are reported as incompatible. + +### Added + +- Every scan publishes an HTTP link report: counts of linked, provider-less, ambiguous, and + external consumer and test calls, and each unlinked call with its reason and candidate + providers. `status` reports it in `http_links`, query results list the unlinked calls among + their entities in `link_gaps`, and the MCP coverage resource and query rendering include both. + +### Fixed + +- `--database` accepts a bare file name such as `graph.db`; the database, its lock, and its work + sidecar are created in the current directory. + +### Release engineering + +- Added a release-mode 200-repository, 50,000-file synchronization acceptance test that gates + content reads of an unchanged scan, rows published by a one-file sync, and database growth + across 20 syncs. Results and a comparison with 1.1.0 are recorded in `docs/PERFORMANCE.md`. +- Added the `fixtures/cross-language-matrix` workspace: tests in Python, TypeScript, Go, Java, and + Rust call FastAPI, Flask, Express, NestJS, Next.js, Spring, Gin, Chi, `net/http`, Axum, and + Actix Web providers through base URLs and cross-file router prefixes, and every pair must + produce `validates` and `implemented_by` edges. +- Added differential tests: an incrementally maintained database, including its fingerprints and + extractor batches, equals a fresh scan after each step of a mutation sequence, and the published + graph, evidence, link report, and artifact rows are identical across repository order, file + creation order, and extraction worker count. +- Replaced the CodeGraph 1.5.0 fixtures with fixtures captured from CodeGraph 1.6.1; contract tests + parse every structured CLI output, and the live smoke test runs each structured operation. +- Lowered the workspace MSRV from 1.97.1 to 1.96.0, the lowest toolchain that builds every locked + dependency at its latest release. +- Updated all dependencies to their latest releases, including `jsonschema` 0.58, `sqlparser` + 0.63, `tree-sitter` 0.27, `rmcp` 3.5, and `serde-saphyr` 1.3. + ## [1.1.0] - 2026-10-02 ### Breaking changes diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index fce391c..a025315 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -16,7 +16,7 @@ have the right to provide it under the project's Apache-2.0 license. ## Development workflow Formatting and Clippy use the latest nightly Rust toolchain. Compilation, tests, documentation, -and release builds use the latest stable toolchain. The workspace MSRV is 1.97.1, declared by +and release builds use the latest stable toolchain. The workspace MSRV is 1.96.0, declared by every crate through the workspace package metadata and checked separately in CI. Install the additional quality toolchain once: @@ -46,7 +46,7 @@ cargo deny check Changes must also compile with the MSRV: ```text -cargo +1.97.1 check --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked +cargo +1.96.0 check --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked ``` ## Dependency updates diff --git a/Cargo.lock b/Cargo.lock index 16230b9..6f16f9b 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -18,9 +18,9 @@ dependencies = [ [[package]] name = "aho-corasick" -version = "1.1.4" +version = "1.1.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ddd31a130427c27518df266943a5308ed92d4b226cc639f5a8f1002816174301" +checksum = "c982642fa9e8606056828ee9a8505737230110bb1099153c79efe865c59d12ba" dependencies = [ "memchr", ] @@ -33,9 +33,9 @@ checksum = "683d7910e743518b0e34f1186f92494becacb047c7b6bf616c96772180fef923" [[package]] name = "android_system_properties" -version = "0.1.5" +version = "0.1.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "819e7219dbd41043ac279b19830f2efc897156490d7fd6ea916720117ee66311" +checksum = "ae221649c9976a6f6c56ae1facf410f3ddb33cc661c4b7b61020a912d4237fbc" dependencies = [ "libc", ] @@ -128,12 +128,6 @@ version = "0.5.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "7d902e3d592a523def97af8f317b08ce16b7ab854c1985a0c671e6f15cebc236" -[[package]] -name = "arrayref" -version = "0.3.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "76a2e8124351fda1ef8aaaa3bbd7ebbcb486bbcd4225aca0aa0d84bb2db8fecb" - [[package]] name = "arrayvec" version = "0.7.8" @@ -142,13 +136,13 @@ checksum = "d3fb67a6e08acf24fdeccbac2cb6ac4305825bd1f117462e0e6f2f193345ad56" [[package]] name = "async-trait" -version = "0.1.91" +version = "0.1.92" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ae36dc4177970ef04fde5178d3e2429882def40e57a451f919c098f72baa6cec" +checksum = "82f6aeea286b8eb4dd3431a1be1b59d290ace00f5bfd8e2a159bc2a05e2c1667" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] @@ -159,12 +153,12 @@ checksum = "1505bd5d3d116872e7271a6d4e16d81d0c8570876c8de68093a09ac269d8aac0" [[package]] name = "atomic-write-file" -version = "0.3.0" +version = "0.3.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "84790c55b5704b0d35130bf16a4ce22a8e70eb0ea773522557524d9a4852663d" +checksum = "ae67d5c03ee972c101a1d8db3ddbf8a9c819ad1ea9d6fd9cd385605547b60edb" dependencies = [ - "nix 0.30.1", - "rand 0.9.5", + "nix", + "rand", ] [[package]] @@ -175,9 +169,9 @@ checksum = "f2032f911046de80f0a198e0901378627c33f59ea0ac00e363d481118bd70a53" [[package]] name = "aws-lc-rs" -version = "1.18.0" +version = "1.18.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ce2b2dcc879c3bae0d371e77c99f2238400ef24ec001394befa67b6e543add9e" +checksum = "b281d307588d634de920874890732659e2e7672f72b5e10e81badc1a8a83621e" dependencies = [ "aws-lc-sys", "zeroize", @@ -185,9 +179,9 @@ dependencies = [ [[package]] name = "aws-lc-sys" -version = "0.44.0" +version = "0.45.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f09fae7be8bb3174e05c6afdb34199e6dc0c7c04ba9fa237b1967adfbde27483" +checksum = "9bff6c3b54fad79a2e60b8102caf565819711497c1f5f092f49508e2f5c31b27" dependencies = [ "cc", "cmake", @@ -250,15 +244,9 @@ dependencies = [ [[package]] name = "base64" -version = "0.22.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "72b3254f16251a8381aa12e40e3c4d2f0199f8c6508fbecb9d91f575e0fbb8c6" - -[[package]] -name = "base64" -version = "0.23.0" +version = "0.23.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b25655df2c3cdd83c5e5b293b88acd880332b2ddadd7c30ac43144fdc0033da9" +checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5" [[package]] name = "bit-set" @@ -277,17 +265,16 @@ checksum = "5e764a1d40d510daf35e07be9eb06e75770908c27d411ee6c92109c9840eaaf7" [[package]] name = "bitflags" -version = "2.13.1" +version = "2.13.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b588b76d00fde79687d7646a9b5bdf3cc0f655e0bbd080335a95d7e96f3587da" +checksum = "3ded4057c258ba199e2d26386d3af3780957ecaee6c4ef4041c6b4b8b97c0b06" [[package]] name = "blake3" -version = "1.8.5" +version = "1.8.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0aa83c34e62843d924f905e0f5c866eb1dd6545fc4d719e803d9ba6030371fce" +checksum = "6d9e454fc11f76977dc803893aff6304ed33d6a26efae8696573bea74baa27ae" dependencies = [ - "arrayref", "arrayvec", "cc", "cfg-if", @@ -303,9 +290,9 @@ checksum = "dc0b364ead1874514c8c2855ab558056ebfeb775653e7ae45ff72f28f8f3166c" [[package]] name = "bstr" -version = "1.13.0" +version = "1.13.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1f7dc094d718f2e1c1559ad110e27eeaae14a5465d3d56dd6dbd793079fbd530" +checksum = "6bb31b46c14244e20ee9984b11bf5c992b91fb6939fea616e3512c8baecdbe5f" dependencies = [ "memchr", "serde_core", @@ -329,20 +316,11 @@ version = "1.12.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "fc652a48c352aef3ea3aed32080501cf3ef6ed5da78602a020c991775b0aff04" -[[package]] -name = "camino" -version = "1.2.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb1307f12aa967b5a58416e87b3653360e0fd614a016b6e970db08fecbb1b80d" -dependencies = [ - "serde_core", -] - [[package]] name = "cc" -version = "1.4.0" +version = "1.5.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5add81bb678e6cb321aff7fa0dc7689ad82b112dbc032cea19f91d6b8e3582b9" +checksum = "f360145194ee8e21db5ee7f3fcd4fe52210864c75c985dae33218202c8bbe040" dependencies = [ "find-msvc-tools", "jobserver", @@ -352,9 +330,9 @@ dependencies = [ [[package]] name = "cfg-if" -version = "1.0.4" +version = "1.0.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" +checksum = "4e7648175b45a9a48536d676f68d918270699102aa8dab5496df06904c914600" [[package]] name = "cfg_aliases" @@ -370,7 +348,7 @@ checksum = "65c35e4b699c7e15ccbe7ee35c005e4fc0a278d22238a2857e6ce2dadeda1b06" dependencies = [ "cfg-if", "cpufeatures", - "rand_core 0.10.1", + "rand_core", ] [[package]] @@ -387,9 +365,9 @@ dependencies = [ [[package]] name = "clap" -version = "4.6.5" +version = "4.6.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "301b56658598e48f3648647ac6fc887be7e7108eddfa4e9b63fcf3ec58c0cadf" +checksum = "aa8876b300ab35ba921adea3dfd70157a46249b33f95c9084ae5709785478946" dependencies = [ "clap_builder", "clap_derive", @@ -397,9 +375,9 @@ dependencies = [ [[package]] name = "clap_builder" -version = "4.6.5" +version = "4.6.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "94a65403d1a1bd28f7dc68eb8506e8874808ee5eecb59298de588e2e1407a078" +checksum = "ec0797fb7aeb1406c84efac526901f7ec3ead2124f946b494e72879d4b54704d" dependencies = [ "anstream", "anstyle", @@ -409,30 +387,30 @@ dependencies = [ [[package]] name = "clap_complete" -version = "4.6.8" +version = "4.6.11" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b1f84a88507dbd05c695f2cb5e8558e747179134005e9893882dec964190ed89" +checksum = "037e2a1a92236d0aff7e845093f64661d6df4c02c9fcc61a60e9e1d736fa392f" dependencies = [ "clap", ] [[package]] name = "clap_derive" -version = "4.6.4" +version = "4.6.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d012d2b9d65aca7f18f4d9878a045bc17899bba951561ba5ec3c2ba1eed9a061" +checksum = "f9c751b79415d4e559e3d1fcf128e09e720eb673a06d26cf6f392d37d75b66e0" dependencies = [ "heck", "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] name = "clap_lex" -version = "1.1.0" +version = "1.1.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9" +checksum = "1c133bc6a41be0d194c306b5506d15e6feeea7b1d6604bd3f8310dfb2ca96486" [[package]] name = "cmake" @@ -445,7 +423,7 @@ dependencies = [ [[package]] name = "code-system-graph" -version = "1.1.0" +version = "1.2.0" dependencies = [ "anyhow", "atomic-write-file", @@ -457,8 +435,9 @@ dependencies = [ "code-system-graph-hooks", "code-system-graph-model", "code-system-graph-store-sqlite", + "foldhash", "jsonschema", - "nix 0.31.3", + "nix", "notify", "reqwest", "rmcp", @@ -471,28 +450,30 @@ dependencies = [ "subtle", "sysinfo", "tempfile", - "thiserror 2.0.19", + "thiserror 2.0.21", "tokio", "tokio-util", "tower", - "tower-http 0.7.0", + "tower-http 0.7.1", "windows-sys 0.61.2", ] [[package]] name = "code-system-graph-core" -version = "1.1.0" +version = "1.2.0" dependencies = [ + "aho-corasick", "async-trait", "atomic-write-file", "blake3", "code-system-graph-model", + "foldhash", "globset", "graphql-parser", "hcl-rs", "ignore", "libc", - "nix 0.31.3", + "nix", "proto-parser", "pulldown-cmark", "reqwest", @@ -504,7 +485,7 @@ dependencies = [ "serde_json", "sqlparser", "tempfile", - "thiserror 2.0.19", + "thiserror 2.0.21", "tokio", "tokio-util", "tree-sitter", @@ -529,34 +510,32 @@ dependencies = [ [[package]] name = "code-system-graph-hooks" -version = "1.1.0" +version = "1.2.0" dependencies = [ "atomic-write-file", "blake3", "libc", - "nix 0.31.3", + "nix", "schemars", "serde", "serde_json", "tempfile", - "thiserror 2.0.19", + "thiserror 2.0.21", ] [[package]] name = "code-system-graph-model" -version = "1.1.0" +version = "1.2.0" dependencies = [ "blake3", - "camino", "schemars", "semver", "serde", - "time", ] [[package]] name = "code-system-graph-store-sqlite" -version = "1.1.0" +version = "1.2.0" dependencies = [ "blake3", "code-system-graph-model", @@ -566,7 +545,7 @@ dependencies = [ "serde_json", "sysinfo", "tempfile", - "thiserror 2.0.19", + "thiserror 2.0.21", "windows-sys 0.61.2", ] @@ -578,9 +557,9 @@ checksum = "1d07550c9036bf2ae0c684c4297d503f838287c83c53686d05370d0e139ae570" [[package]] name = "combine" -version = "4.6.7" +version = "4.6.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ba5a308b75df32fe02788e748662718f03fde005016435c444eea572398219fd" +checksum = "cfc320937d09e6de266b31b9afb480f197d7a861be86be7cb2ea7e5d1bfffc5e" dependencies = [ "bytes", "memchr", @@ -608,20 +587,26 @@ version = "0.8.7" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "773648b94d0e5d620f64f280777445740e61fe701025087ec8b57f45c791888b" +[[package]] +name = "core_detect" +version = "1.0.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "7f8f80099a98041a3d1622845c271458a2d73e688351bf3cb999266764b81d48" + [[package]] name = "cpufeatures" -version = "0.3.0" +version = "0.3.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8b2a41393f66f16b0823bb79094d54ac5fbd34ab292ddafb9a0456ac9f87d201" +checksum = "5ca28b0ae3115b884660db4118d803791fd6756b6e88f39c0f3f7859060d7566" dependencies = [ "libc", ] [[package]] name = "crossbeam-deque" -version = "0.8.7" +version = "0.8.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5181e0de7b61eb03a81e347d6dd8797bae9da5146707b51077e2d71a54ec0ceb" +checksum = "622f3fc73690be383c7214310406f28a90e6edeadc3cea882f9d71e495b9711a" dependencies = [ "crossbeam-epoch", "crossbeam-utils", @@ -629,24 +614,24 @@ dependencies = [ [[package]] name = "crossbeam-epoch" -version = "0.9.20" +version = "0.9.21" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2d6914041f254d6e9176c01941b21115dcfb7089e55135a35411081bd106ef3f" +checksum = "dc74980687109a3b14c72fd458107bf0baa1da1a1a805e178d15501ba9b86d9d" dependencies = [ "crossbeam-utils", ] [[package]] name = "crossbeam-utils" -version = "0.8.22" +version = "0.8.23" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "61803da095bee82a81bb1a452ecc25d3b2f1416d1897eb86430c6159ef717c17" +checksum = "a31eee39dddec8330830986fcd7625edb5a24ec90ea038215273bbc3adb08ac6" [[package]] name = "darling" -version = "0.23.0" +version = "0.24.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "25ae13da2f202d56bd7f91c25fba009e7717a1e4a1cc98a76d844b65ae912e9d" +checksum = "ed17f5901b6630b993ca003def43f2f8ef4014fc13b047b57aad617ff32bc2ec" dependencies = [ "darling_core", "darling_macro", @@ -654,26 +639,26 @@ dependencies = [ [[package]] name = "darling_core" -version = "0.23.0" +version = "0.24.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9865a50f7c335f53564bb694ef660825eb8610e0a53d3e11bf1b0d3df31e03b0" +checksum = "6837e2cf7485aaae18f86181d2f0e9a7ed297a025e220aeabf63fdebd3a2ddff" dependencies = [ "ident_case", "proc-macro2", "quote", "strsim", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] name = "darling_macro" -version = "0.23.0" +version = "0.24.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ac3984ec7bd6cfa798e62b4a642426a5be0e68f9401cfc2a01e3fa9ea2fcdb8d" +checksum = "2ac7135c3ef02b2f7833bbeb1be5ba7f966dcde8a87c6b87f65a778d71a02785" dependencies = [ "darling_core", "quote", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] @@ -682,15 +667,6 @@ version = "2.11.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "4583a4551df46e2792f82ceeac45e850d2e2d5debba0b91f102385cda5b11f06" -[[package]] -name = "deranged" -version = "0.5.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7cd812cc2bc1d69d4764bd80df88b4317eaef9e773c75226407d9bc0876b211c" -dependencies = [ - "serde_core", -] - [[package]] name = "displaydoc" version = "0.2.7" @@ -699,7 +675,7 @@ checksum = "c6232dd377dcc64799954cbd3a9bb882e9cdc1308ccd87b1c098f1fb2eaf82a8" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] @@ -725,18 +701,23 @@ dependencies = [ [[package]] name = "encoding_rs" -version = "0.8.35" +version = "0.8.42" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "75030f3c4f45dafd7586dd6780965a8c7e8e285a5ecb86713e63a79c5b2766f3" +checksum = "8e985e0451871ad22fb8d2b6b076e2028a502a0d3950998c2c5c0a4f9b5d9679" dependencies = [ "cfg-if", + "core_detect", + "multiversion_no_op", + "rustversion", + "scopeguard", + "simdutf8", ] [[package]] name = "encoding_rs_io" -version = "0.1.7" +version = "0.1.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1cc3c5651fb62ab8aa3103998dade57efdd028544bd300516baa31840c252a83" +checksum = "fba3fe847045ecff794b9c138293a80db914678c453ad63fbf0c6a9eb6e00b22" dependencies = [ "encoding_rs", ] @@ -771,9 +752,9 @@ checksum = "7360491ce676a36bf9bb3c56c1aa791658183a54d2744120f27285738d90465a" [[package]] name = "fancy-regex" -version = "0.19.0" +version = "0.19.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "476de73bddf2ef8490aa4ee8f1cf40b430bf1d56c48c22080e5186952cd580e6" +checksum = "d301f5bf187b3c295fce6468d3875037a0bccc5f6b151c63cac2f85babf21912" dependencies = [ "bit-set", "regex-automata", @@ -788,9 +769,9 @@ checksum = "da7c62ceae207dd37ea5b845da6a0696c799f85e97da1ab5b7910be3c1c80223" [[package]] name = "find-msvc-tools" -version = "0.1.9" +version = "0.1.14" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5baebc0774151f905a1a2cc41989300b1e6fbb29aff0ceffa1064fdd3088d582" +checksum = "aedcfb3409746eddb02b9e19ebda1c3394f759a152e48ee875a0844d1b955484" [[package]] name = "fluent-uri" @@ -826,12 +807,12 @@ dependencies = [ [[package]] name = "fraction" -version = "0.15.4" +version = "0.17.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e076045bb43dac435333ed5f04caf35c7463631d0dae2deb2638d94dd0a5b872" +checksum = "e246562084dde8ebbcc943b261c406ce4f68e5032ec28029a251a47d6a295500" dependencies = [ - "lazy_static", "num", + "num-bigint", ] [[package]] @@ -851,9 +832,9 @@ dependencies = [ [[package]] name = "futures" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a88cf1f829d945f548cf8fec32c61b1f202b6d93b45848602fc02af4b12ad218" +checksum = "9a31d2a3fbaaeb2af2368bbdd904aa8e812d3c04a1ee10d3171f52d556e5d0a3" dependencies = [ "futures-channel", "futures-core", @@ -866,9 +847,9 @@ dependencies = [ [[package]] name = "futures-channel" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "262590f4fe6afeb0bc83be1daa64e52657fe185690a958af7f3ad0e92085c5ae" +checksum = "b1f9e3d69d39e4862ffed03ed071a76f9a13ba1d9109d355b0f0aa6b15e393c4" dependencies = [ "futures-core", "futures-sink", @@ -876,15 +857,15 @@ dependencies = [ [[package]] name = "futures-core" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2cd50c473c80f6d7c3670a752354b8e569b1a7cbfdc0419ec88e5edad85e0dc7" +checksum = "92d699e522242e69e3003b94ecc1f960f3a5e015aa7c5d7486e65ad01dd94f5e" [[package]] name = "futures-executor" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6754879cc9f2c66f88c6e5c35344bb0bdb0708b0352b1201815667c7eabc7458" +checksum = "031b47cf1a3c6cc8bc2fc76cd437f521619387907d469316e7c0bc278f1f5432" dependencies = [ "futures-core", "futures-task", @@ -893,38 +874,38 @@ dependencies = [ [[package]] name = "futures-io" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4577ecaa3c4f96589d473f679a71b596316f6641bc350038b962a5daf0085d7a" +checksum = "53c0fa8157de1303bfffdaa1cc2a673bfffb60102f76b0ef4441659124373fed" [[package]] name = "futures-macro" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2d6d3cde68c518367be28956066ddfef33813991b77a55005a69dae04bf3b10b" +checksum = "9fb9654ba8355388abeb8dcb4fc62f511300867002afc858860463bdd9fe0c44" dependencies = [ "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] name = "futures-sink" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e34418ac499d6305c2fb5ad0ed2f6ac998c5f8ca209b4510f7f94242c647e307" +checksum = "1944426bf7d03f1d14f708785e4b33efd750b36d48a157b836b3efc15ede8e1d" [[package]] name = "futures-task" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b231ed28831efb4a61a08580c4bc233ec56bc009f4cd8f52da2c3cb97df0c109" +checksum = "cd417de3d1d015fc3bfd2b1ea46dfc7bab72ef86f1cc7cc9c78e728b34a6d1fd" [[package]] name = "futures-util" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a77a90a256fce34da66415271e30f94ee91c57b04b8a2c042d9cf3220179deaa" +checksum = "0d50a92467f8ba5dd6e3ee5d4bd04d73ab2e4e1c44474a0674821dfce14b79bc" dependencies = [ "futures-channel", "futures-core", @@ -983,15 +964,15 @@ dependencies = [ "js-sys", "libc", "r-efi 6.0.0", - "rand_core 0.10.1", + "rand_core", "wasm-bindgen", ] [[package]] name = "globset" -version = "0.4.19" +version = "0.4.20" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e47d37d2ae4464254884b60ab7071be2b876a9c35b696bd018ddcc76847309cd" +checksum = "07c34a9410465b45bd9787443bc7370f37735bad04b0f0cd57ff1a3186c98988" dependencies = [ "aho-corasick", "bstr", @@ -1002,9 +983,9 @@ dependencies = [ [[package]] name = "granit-parser" -version = "1.0.0" +version = "1.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c60388b03522b86e24b6b45952255279936b08a871775156fc89c94e792d21d8" +checksum = "e20f99e46474f56bd905c56e817ebddcf377a611f94c53ac4649e4d3fa3c0cd0" dependencies = [ "arraydeque", "smallvec", @@ -1042,18 +1023,18 @@ dependencies = [ [[package]] name = "hashlink" -version = "0.12.1" +version = "0.12.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32069d97bb81e38fa67eab65e3393bf804bb85969f2bc06bf13f64aef5aba248" +checksum = "a596f1b20ed2cc5ecac41a164aaebc7258057060f06c0cf7a2ba3991ee7990fb" dependencies = [ "hashbrown 0.17.1", ] [[package]] name = "hcl-edit" -version = "0.9.6" +version = "0.9.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f641e65979f9c567246f13f9c64f7af41b3476e3171a63b10b36778af95b267b" +checksum = "8a67cc5751cb5996669b9780cd738d4541a76c2d5f2f10c17af4965c5a03a5f5" dependencies = [ "fnv", "hcl-primitives", @@ -1064,9 +1045,9 @@ dependencies = [ [[package]] name = "hcl-primitives" -version = "0.1.11" +version = "0.1.12" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "829a11d304c89e2cfe0dbb494a686bbe2b48ade17705c62cd1957b04aa4630f6" +checksum = "bd662a8afeca01b5b5318f35baed70017b9f854bfa38bdcdadb87de946a49071" dependencies = [ "itoa", "kstring", @@ -1077,9 +1058,9 @@ dependencies = [ [[package]] name = "hcl-rs" -version = "0.19.7" +version = "0.19.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "153fc72c1037f6933efe01c1c769235479e995fb82a799c5716dabad7375b671" +checksum = "212c6fce17a8b9e0eab43ccf61c8bc5d2c51fc24ad52d19bd1d2be2eeba22d02" dependencies = [ "hcl-edit", "hcl-primitives", @@ -1117,9 +1098,9 @@ dependencies = [ [[package]] name = "http-body-util" -version = "0.1.4" +version = "0.1.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e9f41fd6a08e4d4ec69df65976da761afd5ad5e58a9d4acb46bd1c953a9e3ff2" +checksum = "23169fe34a5fbcdd3f3862e78fb9b6fccd5f02a6dc6f732547005d45631ce71c" dependencies = [ "bytes", "futures-core", @@ -1142,9 +1123,9 @@ checksum = "df3b46402a9d5adb4c86a0cf463f42e19994e3ee891101b1841f30a545cb49a9" [[package]] name = "hyper" -version = "1.11.0" +version = "1.11.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d22053281f852e11534f5198498373cbb59295120a20771d90f7ed1897490a72" +checksum = "27b501faa50e7a26c3d3560ca625132f4078a17771f4810baf70475ae48cbe43" dependencies = [ "atomic-waker", "bytes", @@ -1163,9 +1144,9 @@ dependencies = [ [[package]] name = "hyper-rustls" -version = "0.27.9" +version = "0.27.10" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "33ca68d021ef39cf6463ab54c1d0f5daf03377b70561305bb89a8f83aab66e0f" +checksum = "dfa8e654703247911e29c23fbeaa261834bd9bb74efba2f9acddc37bfb127f53" dependencies = [ "http", "hyper", @@ -1178,16 +1159,17 @@ dependencies = [ [[package]] name = "hyper-util" -version = "0.1.20" +version = "0.1.21" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "96547c2556ec9d12fb1578c4eaf448b04993e7fb79cbaad930a656880a6bdfa0" +checksum = "ddc03d96684f9226b8a787cdb71488417b53ab5ea8fdb1dac946cb9431cc8bff" dependencies = [ - "base64 0.22.1", + "base64", "bytes", "futures-channel", "futures-util", "http", "http-body", + "httparse", "hyper", "ipnet", "libc", @@ -1225,9 +1207,9 @@ dependencies = [ [[package]] name = "icu_collections" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2984d1cd16c883d7935b9e07e44071dca8d917fd52ecc02c04d5fa0b5a3f191c" +checksum = "fa68d21081c4a05d5a901a1c62add574c77048b6a1c67be3b50ce0b60d4ca513" dependencies = [ "displaydoc", "potential_utf", @@ -1239,9 +1221,9 @@ dependencies = [ [[package]] name = "icu_locale_core" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92219b62b3e2b4d88ac5119f8904c10f8f61bf7e95b640d25ba3075e6cac2c29" +checksum = "d56e28588da92eee5c3201a6eff33fabdd49b62269c8938d4ff050ce4d900deb" dependencies = [ "displaydoc", "litemap", @@ -1252,9 +1234,9 @@ dependencies = [ [[package]] name = "icu_normalizer" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c56e5ee99d6e3d33bd91c5d85458b6005a22140021cc324cea84dd0e72cff3b4" +checksum = "12f9cf5f235641ed274641dd81c3f28d870e276763d0797aeeab72317b1c646f" dependencies = [ "icu_collections", "icu_normalizer_data", @@ -1266,16 +1248,17 @@ dependencies = [ [[package]] name = "icu_normalizer_data" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "da3be0ae77ea334f4da67c12f149704f19f81d1adf7c51cf482943e84a2bad38" +checksum = "1563da1ed3e0b3bf3d74c9b85917ac9c56464d2f57242270c09c9e752f8021a0" [[package]] name = "icu_properties" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bee3b67d0ea5c2cca5003417989af8996f8604e34fb9ddf96208a033901e70de" +checksum = "7e7ca276ad3145661a65914e6daf131ca5120cd3dcee8f8f3214b8875184a148" dependencies = [ + "displaydoc", "icu_collections", "icu_locale_core", "icu_properties_data", @@ -1286,15 +1269,15 @@ dependencies = [ [[package]] name = "icu_properties_data" -version = "2.2.0" +version = "2.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8e2bbb201e0c04f7b4b3e14382af113e17ba4f63e2c9d2ee626b720cbce54a14" +checksum = "e590f038c1464a96894fd6d10127e90a8be4509f56ff7ecef851b15cee0b7caa" [[package]] name = "icu_provider" -version = "2.2.0" +version = "2.3.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "139c4cf31c8b5f33d7e199446eff9c1e02decfc2f0eec2c8d71f65befa45b421" +checksum = "d27bbb9d3abbefac45d55f647c9de1d44aafcd1186eb91879afef17c396c3e73" dependencies = [ "displaydoc", "icu_locale_core", @@ -1350,9 +1333,9 @@ dependencies = [ [[package]] name = "indexmap" -version = "2.14.0" +version = "2.14.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d466e9454f08e4a911e14806c24e16fba1b4c121d1ea474396f396069cf949d9" +checksum = "cc4e190f5d26ca7051642629da2c52fc03bde85a03197c99408dcd291734c855" dependencies = [ "equivalent", "hashbrown 0.17.1", @@ -1362,9 +1345,9 @@ dependencies = [ [[package]] name = "inotify" -version = "0.11.4" +version = "0.11.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "153be1941a183ec9ccd095ddbe17a8b8d435ef6c76e9e02451b933c3999af2c8" +checksum = "4cc00ea907cab49550b7da656f80ebb97be1b997d931fbcd28d39734e17ce592" dependencies = [ "bitflags", "inotify-sys", @@ -1382,9 +1365,9 @@ dependencies = [ [[package]] name = "ipnet" -version = "2.12.0" +version = "2.12.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d98f6fed1fde3f8c21bc40a1abb88dd75e67924f9cffc3ef95607bad8017f8e2" +checksum = "791930b43c0d5973160d90a8f3894509f2b273430f5c5c73b668636d0287c5c0" [[package]] name = "is_terminal_polyfill" @@ -1410,7 +1393,7 @@ dependencies = [ "jni-sys", "log", "simd_cesu8", - "thiserror 2.0.19", + "thiserror 2.0.21", "walkdir", "windows-link", ] @@ -1459,9 +1442,9 @@ dependencies = [ [[package]] name = "js-sys" -version = "0.3.103" +version = "0.3.106" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "53b44bfcdb3f8d5837a46dae1ca9660a837176eee74a28b229bc626816589102" +checksum = "7883d941dae510fb2d978fc3fe018c71c9e2892fd38854de3e8b92c2e5ad9cc5" dependencies = [ "cfg-if", "futures-util", @@ -1470,9 +1453,9 @@ dependencies = [ [[package]] name = "jsonschema" -version = "0.49.8" +version = "0.58.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ce93912abc8220a3fdb768b2c4a826a7a9a4b1599cfb4d760d422eaa1faf88c7" +checksum = "ea18b8d5e1469b1bdd6b349169305141316a7c26da27f239773b65c830c0f5d6" dependencies = [ "ahash", "bytecount", @@ -1481,7 +1464,6 @@ dependencies = [ "fancy-regex", "fraction", "getrandom 0.3.4", - "idna", "itoa", "jsonschema-regex", "jsonschema-value", @@ -1499,32 +1481,34 @@ dependencies = [ [[package]] name = "jsonschema-regex" -version = "0.49.8" +version = "0.58.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8165657ebed4d32c50f3c250c1986d8428b16fbfeac355222e8fec50aa26eb1d" +checksum = "854e9e22c420535035e95eed240cbfb4a5f4b7da74d122d63144e0c31926dbfd" dependencies = [ "regex-syntax", ] [[package]] name = "jsonschema-value" -version = "0.49.8" +version = "0.58.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b069c5fdda3e9c2242bba49d811d5bbdb488abd8d234ed44a6312fd2891113e1" +checksum = "4e3bfac11ac357ec620880025757d38c67da8fd5b9ef5024d19f0f360002b4ca" dependencies = [ "ahash", "bytecount", "fraction", + "getrandom 0.3.4", "num-cmp", "num-traits", "serde_json", + "zmij", ] [[package]] name = "kqueue" -version = "1.2.0" +version = "1.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "273c0752728918e0ac4976f2b275b6fefb9ecd400585dec929419f3844cd87b5" +checksum = "8d763e5b24120b4ddf50de6c92308156765aabfbbccebf401da7cff2d70a41ea" dependencies = [ "kqueue-sys", "libc", @@ -1542,25 +1526,19 @@ dependencies = [ [[package]] name = "kstring" -version = "2.0.4" +version = "2.0.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b609e7ca5ea38f093c20a4a102335b247221c9643b7a6bc3510f196f99499a9e" +checksum = "7a09b82a7f771ed02dc0dd9b27130a0fa5499fa15ed3027116c1e5e4e591bd9e" dependencies = [ "serde", "static_assertions", ] -[[package]] -name = "lazy_static" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bbd2bcb4c963f2ddae06a2efc7e9f3591312473c50c6685e1f298068316e66fe" - [[package]] name = "libc" -version = "0.2.189" +version = "0.2.190" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3eaf3ede3fee6db1a4c2ee091bf8a8b4dccdc6d17f656fb07896ee72867612f2" +checksum = "ce5d3ddc6d3fa000eb1536d85e147bfe31aacaba692ed6a876f95cb7c855be78" [[package]] name = "libfuzzer-sys" @@ -1574,9 +1552,9 @@ dependencies = [ [[package]] name = "libsqlite3-sys" -version = "0.38.1" +version = "0.38.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f6c19a05435c21ac299d71b6a9c13db3e3f47c520517d58990a462a1397a61db" +checksum = "f1d20bef17f513b9b3004532233187769cd072d790971f4e4da0e346eb6401e8" dependencies = [ "cc", "pkg-config", @@ -1591,9 +1569,9 @@ checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53" [[package]] name = "litemap" -version = "0.8.2" +version = "0.8.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92daf443525c4cce67b150400bc2316076100ce0b3686209eb8cf3c31612e6f0" +checksum = "47d9d19d1d6efa0109d2f65ff4c85cddd50bd572e5a00127ab10987290bcefae" [[package]] name = "lock_api" @@ -1606,15 +1584,15 @@ dependencies = [ [[package]] name = "log" -version = "0.4.33" +version = "0.4.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad" +checksum = "f9f8bd3e56ce4dfc153cf470fffbfa98c7620958b312ca5c3a4b8d5181fd13c6" [[package]] name = "lru-slab" -version = "0.1.2" +version = "0.1.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "112b39cec0b298b6c1999fee3e31427f74f676e4cb9879ed1a121b43661a4154" +checksum = "4050469837a6ff301cd14c1f8f24f88549e6d548f24f64e2148eb0f72cebc51f" [[package]] name = "matchit" @@ -1642,9 +1620,9 @@ checksum = "6877bb514081ee2a7ff5ef9de3281f14a4dd4bceac4c09388074a6b5df8a139a" [[package]] name = "mio" -version = "1.2.2" +version = "1.2.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "30d65c71f1ce40ab09135ce117d742b9f8a19ff91a41a8b57ed50bc2de59c427" +checksum = "4b18443e9c262bfe8fa82f51666e2642c53393f7e5c27b3e1aeab922cff5b9d8" dependencies = [ "libc", "log", @@ -1653,16 +1631,10 @@ dependencies = [ ] [[package]] -name = "nix" -version = "0.30.1" +name = "multiversion_no_op" +version = "1.0.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "74523f3a35e05aba87a1d978330aef40f67b0304ac79c1c00b294c9830543db6" -dependencies = [ - "bitflags", - "cfg-if", - "cfg_aliases", - "libc", -] +checksum = "743fb55ba31b18fb1ecef6bdc9aa2743314978ac084044301a7eee33fb99a20d" [[package]] name = "nix" @@ -1757,17 +1729,11 @@ dependencies = [ "num-traits", ] -[[package]] -name = "num-conv" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "521739c6d2bac4aa25192232afe6841231376b2b26d4d9fae5ecf8ca5772e441" - [[package]] name = "num-integer" -version = "0.1.46" +version = "0.1.47" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7969661fd2958a5cb096e56c8e1ad0444ac2bbcd0061bd28660485a44879858f" +checksum = "7ce2d95d4b3734dc35aa2f45e1aa22cd416814592a4f9d9205e11affd5b8e10b" dependencies = [ "num-traits", ] @@ -1897,34 +1863,19 @@ checksum = "a89322df9ebe1c1578d689c92318e070967d1042b512afbe49518723f4e6d5cd" [[package]] name = "pkg-config" -version = "0.3.33" +version = "0.3.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "19f132c84eca552bf34cab8ec81f1c1dcc229b811638f9d283dceabe58c5569e" +checksum = "f6b464fbc74e149a392436b17d523f769e057cb6877f6a5c4618bc6f11800548" [[package]] name = "potential_utf" -version = "0.1.5" +version = "0.1.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0103b1cef7ec0cf76490e969665504990193874ea05c85ff9bab8b911d0a0564" +checksum = "d83eb9bc6d8e5cf568e7a1101d60ee05e81ed50ea106026f3d18deeb046d7661" dependencies = [ "zerovec", ] -[[package]] -name = "powerfmt" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391" - -[[package]] -name = "ppv-lite86" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "85eae3c4ed2f50dcfe72643da4befc30deadb458a9b590d720cde2f2b1e97da9" -dependencies = [ - "zerocopy", -] - [[package]] name = "pratt" version = "0.4.0" @@ -1942,13 +1893,13 @@ dependencies = [ [[package]] name = "process-wrap" -version = "9.1.0" +version = "10.0.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2e842efad9119158434d193c6682e2ebee4b44d6ad801d7b349623b3f57cdf55" +checksum = "1f21b97672d2dc848e7b25701ab4535618b92f4861c13cc3f7f7bed52ad3c8da" dependencies = [ "futures", "indexmap", - "nix 0.31.3", + "nix", "tokio", "tracing", "windows", @@ -1962,9 +1913,9 @@ checksum = "c10e0ce3022eceb3fd55e68577249e530fb2d2cfed8a585eb52fd90ab994e734" [[package]] name = "psm" -version = "0.1.31" +version = "0.1.32" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "645dbe486e346d9b5de3ef16ede18c26e6c70ad97418f4874b8b1889d6e761ea" +checksum = "4dcd034599e63b970727f70d79e02d62390a4a84f7c6b827c27c46d5ac3fa622" dependencies = [ "ar_archive_writer", "cc", @@ -1991,9 +1942,9 @@ checksum = "007d8adb5ddab6f8e3f491ac63566a7d5002cc7ed73901f72057943fa71ae1ae" [[package]] name = "quinn" -version = "0.11.11" +version = "0.11.12" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c1a41e437b6bbd489372cd4971de128e85c855f56c57f283d20ff016cf7c0a8" +checksum = "4051e23e9185c255a7e33ef59cdbca87a22d359052eecd22fc6b901fb37d9d11" dependencies = [ "bytes", "cfg_aliases", @@ -2003,7 +1954,7 @@ dependencies = [ "rustc-hash", "rustls", "socket2", - "thiserror 2.0.19", + "thiserror 2.0.21", "tokio", "tracing", "web-time", @@ -2011,22 +1962,22 @@ dependencies = [ [[package]] name = "quinn-proto" -version = "0.11.16" +version = "0.11.19" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2f4bfc015262b9df63c8845072ce59068853ff5872180c2ce2f13038b970e560" +checksum = "0e750cca55fe4f0439a15d0bb529da9651e79993e8e72c61a899a36d462befbe" dependencies = [ "aws-lc-rs", "bytes", "getrandom 0.4.3", "lru-slab", - "rand 0.10.2", + "rand", "rand_pcg", "ring", "rustc-hash", "rustls", "rustls-pki-types", "slab", - "thiserror 2.0.19", + "thiserror 2.0.21", "tinyvec", "tracing", "web-time", @@ -2034,9 +1985,9 @@ dependencies = [ [[package]] name = "quinn-udp" -version = "0.5.15" +version = "0.5.16" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "35a133f956daabe89a61a685c2649f13d82d5aa4bd5d12d1277e1072a21c0694" +checksum = "af66907df18639dcf4db56ca65490cabc4b27a97dbadd96f2926cca73298f016" dependencies = [ "cfg_aliases", "libc", @@ -2069,42 +2020,13 @@ checksum = "f8dcc9c7d52a811697d2151c701e0d08956f92b0e24136cf4cf27b57a6a0d9bf" [[package]] name = "rand" -version = "0.9.5" +version = "0.10.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b9ef1d0d795eb7d84685bca4f72f3649f064e6641543d3a8c415898726a57b41" -dependencies = [ - "rand_chacha", - "rand_core 0.9.5", -] - -[[package]] -name = "rand" -version = "0.10.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c7f5fa3a058cd35567ef9bfa5e75732bee0f9e4c55fa90477bef2dfcdbc4be80" +checksum = "65c9fb96cbc91e3478eaae79a69fcd3f1ae4ad052e471fe6732fff548984b4af" dependencies = [ "chacha20", "getrandom 0.4.3", - "rand_core 0.10.1", -] - -[[package]] -name = "rand_chacha" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3022b5f1df60f26e1ffddd6c66e8aa15de382ae63b3a0c1bfc0e4d3e3f325cb" -dependencies = [ - "ppv-lite86", - "rand_core 0.9.5", -] - -[[package]] -name = "rand_core" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "76afc826de14238e6e8c374ddcc1fa19e374fd8dd986b0d2af0d02377261d83c" -dependencies = [ - "getrandom 0.3.4", + "rand_core", ] [[package]] @@ -2119,7 +2041,7 @@ version = "0.10.2" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "caa0f4137e1c0a72f4c651489402276c8e8e1cf081f3b0ba156d2cbeef09e86a" dependencies = [ - "rand_core 0.10.1", + "rand_core", ] [[package]] @@ -2153,29 +2075,29 @@ dependencies = [ [[package]] name = "ref-cast" -version = "1.0.26" +version = "1.0.27" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "216e8f773d7923bcba9ceb86a86c93cabb3903a11872fc3f138c49630e50b96d" +checksum = "7e440fb4e4b4147295338efb76001ab9e4efc0e5839df2c47fc5ac2381d365c3" dependencies = [ "ref-cast-impl", ] [[package]] name = "ref-cast-impl" -version = "1.0.26" +version = "1.0.27" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2c9283685feec7d69af75fb0e858d5e7378f33fe4fc699383b2916ab9273e03c" +checksum = "92ecd8964f8453721699a1ed72037b0db49ce2f5a5138486ee89bed6f67cdf3a" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] name = "referencing" -version = "0.49.8" +version = "0.58.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f39f4c36ce0f50e96fb740d895f1cad34cfa76c4ab5c36934d6590bcd7029087" +checksum = "590efadb0a669f1712c1e3ab810a55d4c3d20d0127c455f7236dd0710aba0346" dependencies = [ "ahash", "fluent-uri", @@ -2219,11 +2141,11 @@ checksum = "d6f6ff9a378485b298a5286656da665ba74413d36db0979633275d2e708145d4" [[package]] name = "reqwest" -version = "0.13.4" +version = "0.13.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "219c5811de6525e5416c7d5d53bb656d3afdbc6c5af816e0802bcfa42dbdc1c3" +checksum = "16a1cfa75cc186dd73d5818e510e042e40927bccc9c236b061cea97e1eb08029" dependencies = [ - "base64 0.22.1", + "base64", "bytes", "futures-core", "http", @@ -2270,14 +2192,14 @@ dependencies = [ [[package]] name = "rmcp" -version = "3.1.0" +version = "3.5.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ad26b216c966e987e80e86daf784a455c039c43d98575ceed57b8faa259e5695" +checksum = "fae7019994ae0fe4ada40b732f798f3ff26f0f04facb1477f1bf37eb4f18a2d3" dependencies = [ - "async-trait", - "base64 0.23.0", + "base64", "chrono", "futures", + "indexmap", "pastey", "pin-project-lite", "process-wrap", @@ -2285,7 +2207,7 @@ dependencies = [ "schemars", "serde", "serde_json", - "thiserror 2.0.19", + "thiserror 2.0.21", "tokio", "tokio-stream", "tokio-util", @@ -2296,15 +2218,15 @@ dependencies = [ [[package]] name = "rmcp-macros" -version = "3.1.0" +version = "3.5.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41bc748630c2be2a71b614c2f40d27bc0df0060696d224e1692c72345b7e0b79" +checksum = "c7e66877136e969d5a2d561ee5841b17717a0619aebd64abb47a2a2f46e34683" dependencies = [ "darling", "proc-macro2", "quote", "serde_json", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] @@ -2314,14 +2236,14 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "c51c9ae4df8a7fba42103df5c621fa3c37eccf3a3c650879e90fc48b11cc192c" dependencies = [ "hashbrown 0.16.1", - "thiserror 2.0.19", + "thiserror 2.0.21", ] [[package]] name = "rusqlite" -version = "0.40.1" +version = "0.40.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "11438310b19e3109b6446c33d1ed5e889428cf2e278407bc7896bc4aaea43323" +checksum = "23f2a97da3e3873c73cb2a2e71b35c40ff95e0b1eefa8d72d8499a6928c3b5b3" dependencies = [ "bitflags", "fallible-iterator", @@ -2349,9 +2271,9 @@ dependencies = [ [[package]] name = "rustix" -version = "1.1.4" +version = "1.1.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190" +checksum = "891efababe418670775f199f0d233d84843c227a0949a883ce15b37c78d6629d" dependencies = [ "bitflags", "errno", @@ -2398,9 +2320,9 @@ dependencies = [ [[package]] name = "rustls-platform-verifier" -version = "0.7.0" +version = "0.7.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "26d1e2536ce4f35f4846aa13bff16bd0ff40157cdb14cc056c7b14ba41233ba0" +checksum = "1167586491e2b18b8bfbb293e8180ec17c201c4f076d7cb3070ca964e7598f98" dependencies = [ "core-foundation", "core-foundation-sys", @@ -2419,9 +2341,9 @@ dependencies = [ [[package]] name = "rustls-platform-verifier-android" -version = "0.1.1" +version = "0.2.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f87165f0995f63a9fbeea62b64d10b4d9d8e78ec6d7d51fb2125fda7bb36788f" +checksum = "eec689c0bc40ff2458a5977b6619cb718087084a18e02a131c599b62d05e1a5f" [[package]] name = "rustls-webpki" @@ -2488,7 +2410,7 @@ dependencies = [ "proc-macro2", "quote", "serde_derive_internals", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] @@ -2542,12 +2464,12 @@ dependencies = [ [[package]] name = "serde-saphyr" -version = "1.0.0" +version = "1.3.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5bb49c0aa8a5bb88c00b833c11bae50d1611fb89e01336f53b05f7c643b1c883" +checksum = "b8050abb251097357e24aff63ba2c52a6309ecb7d23a5474023df960a02694d8" dependencies = [ "annotate-snippets", - "base64 0.22.1", + "base64", "encoding_rs_io", "granit-parser", "nohash-hasher", @@ -2574,7 +2496,7 @@ checksum = "e7a5d71263a5a7d47b41f6b3f06ba276f10cc18b0931f1799f710578e2309348" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] @@ -2585,7 +2507,7 @@ checksum = "f852137cce035d6a4df67ccce505ff6b3e9fd3a10e3e52b24dc71e650bb1a9bd" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] @@ -2675,9 +2597,9 @@ checksum = "0c790de23124f9ab44544d7ac05d60440adc586479ce501c1d6d7da3cd8c9cf5" [[package]] name = "smallvec" -version = "1.15.2" +version = "1.16.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8ed6a63f02c8539c91a8685a86f4099661ba3da017932f6ebbea6de3f0fa7c90" +checksum = "f9395f0f0eee849a9b707b2f06bb92a6a422090e2123bb2ef8e87a0e61892a8e" [[package]] name = "socket2" @@ -2703,9 +2625,9 @@ dependencies = [ [[package]] name = "sqlparser" -version = "0.62.0" +version = "0.63.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "13c6d1b651dc4edf07eead2a0c6c78016ce971bc2c10da5266861b13f25e7cec" +checksum = "3679862809bd1f92e563cf6fd820e28f39135303668001c10331d5027183e8b4" dependencies = [ "log", "recursive", @@ -2719,9 +2641,9 @@ checksum = "6ce2be8dc25455e1f91df71bfa12ad37d7af1092ae736f3a6cd0e37bc7810596" [[package]] name = "stacker" -version = "0.1.24" +version = "0.1.25" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "640c8cdd92b6b12f5bcb1803ca3bbf5ab96e5e6b6b96b9ab77dabe9e880b3190" +checksum = "707f49d46706bacf8a2b00d51dace3f9de527c13eec3778f570c411f89e69967" dependencies = [ "cc", "cfg-if", @@ -2788,9 +2710,9 @@ dependencies = [ [[package]] name = "syn" -version = "3.0.3" +version = "3.0.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "53e9bae58849f64dfa4f5d5ae372c8341f7305f82a3868709269343628b659a3" +checksum = "8593e8e72159ed2257d083c7a454a85cbf854f37a0966d8d483aff8c8a3ebcee" dependencies = [ "proc-macro2", "quote", @@ -2808,13 +2730,13 @@ dependencies = [ [[package]] name = "synstructure" -version = "0.13.2" +version = "0.14.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "728a70f3dbaf5bab7f0c4b1ac8d7ae5ea60a4b5549c8a5914361c99147a709d2" +checksum = "901704edd0dfe137f1987838ee4f259e4e063c31371bdb423f7ae38ec6f77f02" dependencies = [ "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] @@ -2838,7 +2760,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" dependencies = [ "fastrand", - "getrandom 0.3.4", + "getrandom 0.4.3", "once_cell", "rustix", "windows-sys 0.61.2", @@ -2855,11 +2777,11 @@ dependencies = [ [[package]] name = "thiserror" -version = "2.0.19" +version = "2.0.21" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09a43598840e33d5b0331f38c5e30d13bb11c11210a4b58f0d9b18a5a5eefcd9" +checksum = "09e52cb86a36cede5cb101bf8908837b3e4c6e5e59fe7fd85c23fb56200d189e" dependencies = [ - "thiserror-impl 2.0.19", + "thiserror-impl 2.0.21", ] [[package]] @@ -2875,50 +2797,20 @@ dependencies = [ [[package]] name = "thiserror-impl" -version = "2.0.19" +version = "2.0.21" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "43cbfe0cf76104d42a574802844187e84a305e531ed54455f11fbde0f10541cd" +checksum = "fe5197923287db20a58125f0bc85c062f7f2c892de97b18c356f9efb14b28524" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", -] - -[[package]] -name = "time" -version = "0.3.55" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cdb87b95ec50ddfa440816d227a17b2ccbdda963a316a727fda0fc4334f7d134" -dependencies = [ - "deranged", - "num-conv", - "powerfmt", - "serde_core", - "time-core", - "time-macros", -] - -[[package]] -name = "time-core" -version = "0.1.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9e1c906769ad99c88eaa54e728060edef082f8e358ff32030cb7c7d315e81109" - -[[package]] -name = "time-macros" -version = "0.2.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7e689342a48d2ea927c87ea50cabf8594854bf940e9310208848d680d668ed85" -dependencies = [ - "num-conv", - "time-core", + "syn 3.0.6", ] [[package]] name = "tinystr" -version = "0.8.3" +version = "0.8.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8323304221c2a851516f22236c5722a72eaa19749016521d6dff0824447d96d" +checksum = "b1e27c91459209c2986af3dcf603a5a74a4368754ce37414f59acc971167f643" dependencies = [ "displaydoc", "zerovec", @@ -2926,18 +2818,9 @@ dependencies = [ [[package]] name = "tinyvec" -version = "1.12.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb4ebadaa0af04fab11ae01eb5f9fdb5f9c5b875506e210e71c07873528baa7f" -dependencies = [ - "tinyvec_macros", -] - -[[package]] -name = "tinyvec_macros" -version = "0.1.1" +version = "1.13.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1f3ccbac311fea05f86f61904b462b55fb3df8837a366dfc601a0161d0532f20" +checksum = "fd3ca314f692efd6c868f8408f53fe444634a845f96c028b97d35f6a1f79f0ee" [[package]] name = "tokio" @@ -2963,14 +2846,14 @@ checksum = "78773a2a397f451582ce068015985c33193cf6dea8b74d2a639fe457b2f07b0e" dependencies = [ "proc-macro2", "quote", - "syn 3.0.3", + "syn 3.0.6", ] [[package]] name = "tokio-rustls" -version = "0.26.4" +version = "0.26.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1729aa945f29d91ba541258c8df89027d5792d85a8841fb65e8bf0f4ede4ef61" +checksum = "c9cc2678c2cdd569ef8215e2afd7954ada2ae20b4fdd2c5fe6139a3b02d105db" dependencies = [ "rustls", "tokio", @@ -3039,9 +2922,9 @@ dependencies = [ [[package]] name = "tower-http" -version = "0.7.0" +version = "0.7.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b11f75e912b0c2be01b63d8cf8057b8c3f97cf34abb3d431a3a4c8675498e233" +checksum = "08a05a66a4fdd61cbbe0a1d755ffe0ca6aba159dd4820936a0ff8a8278245b9c" dependencies = [ "bitflags", "bytes", @@ -3103,13 +2986,12 @@ dependencies = [ [[package]] name = "tree-sitter" -version = "0.26.11" +version = "0.27.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "af1c71c1c4cc0920b20d6b0f6572e7682cd07a6a2faec71067a31fa394c586df" +checksum = "2038684e0058edba0d17302619f62eabce4a8e11c6ac59506996a8d79848851d" dependencies = [ "cc", "regex", - "regex-syntax", "serde_json", "streaming-iterator", "tree-sitter-language", @@ -3147,9 +3029,9 @@ dependencies = [ [[package]] name = "tree-sitter-language" -version = "0.1.7" +version = "0.1.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "009994f150cc0cd50ff54917d5bc8bffe8cad10ca10d81c34da2ec421ae61782" +checksum = "ca0d1bf6fdd806e43ae5198f82f527056d359def39e54e67a0f478ac09dac081" [[package]] name = "tree-sitter-python" @@ -3201,9 +3083,9 @@ checksum = "0b993bddc193ae5bd0d623b49ec06ac3e9312875fdae725a975c51db1cc1677f" [[package]] name = "unicode-ident" -version = "1.0.24" +version = "1.0.26" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" +checksum = "d245f478577f809a851594d02313b640fb437e0bb33866753cff937863096954" [[package]] name = "unicode-width" @@ -3243,9 +3125,9 @@ checksum = "06abde3611657adf66d383f00b093d7faecc7fa57071cce2578660c9f1010821" [[package]] name = "uuid" -version = "1.24.0" +version = "1.26.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bf3923a6f5c4c6382e0b653c4117f48d631ea17f38ed86e2a828e6f7412f5239" +checksum = "2ef6dac1e96601b4fb3acccccff2139741fcb757cb9a36089bf5be91cfb285ce" dependencies = [ "getrandom 0.4.3", "js-sys", @@ -3325,9 +3207,9 @@ dependencies = [ [[package]] name = "wasm-bindgen" -version = "0.2.126" +version = "0.2.129" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4b067c0c11094aef6b7a801c1e34a26affafdf3d051dba08456b868789aaf9a4" +checksum = "9bb54f33acc68fd454578d9820b0bde1a1a3d17aa17bb7b6595806d02886d409" dependencies = [ "cfg-if", "once_cell", @@ -3338,19 +3220,20 @@ dependencies = [ [[package]] name = "wasm-bindgen-futures" -version = "0.4.76" +version = "0.4.79" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c62df1340f32221cb9c54d6a27b030e3dba64361d4a95bed55f9aacb44da291d" +checksum = "3cbab34de2d982e9b48e18d216d04c4a6f641066ff19ffb699980f591ee3610e" dependencies = [ "js-sys", + "tokio", "wasm-bindgen", ] [[package]] name = "wasm-bindgen-macro" -version = "0.2.126" +version = "0.2.129" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "167ce5e579f6bcf889c4f7175a8a5a585de84e8ff93976ce393efa5f2837aab1" +checksum = "2e29d0c35b16e224a7eeb5cd2d25e3e1968fbd65604117b44d3b789d00ee8535" dependencies = [ "quote", "wasm-bindgen-macro-support", @@ -3358,31 +3241,31 @@ dependencies = [ [[package]] name = "wasm-bindgen-macro-support" -version = "0.2.126" +version = "0.2.129" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f3997c7839262f4ef12cf90b818d6340c18e80f263f1a94bf157d0ec4420380e" +checksum = "6f501a8bc3719dba86ef8ae4728879c08001bea749eb1333ac5b91e040e2a6b7" dependencies = [ "bumpalo", "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", "wasm-bindgen-shared", ] [[package]] name = "wasm-bindgen-shared" -version = "0.2.126" +version = "0.2.129" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dc1b4cb0cc549fcf58d7dfc081778139b3d283a081644e833e84682ad71cea24" +checksum = "23f0c9c52aa7cd7d77769a4cfe2a9adb1b331f489a41d912ce14513d5ab995c6" dependencies = [ "unicode-ident", ] [[package]] name = "web-sys" -version = "0.3.103" +version = "0.3.106" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8622dcb61c0bcc9fffa6938bed81210af2da9a7e4a1a834b2e37a59b6dfb6141" +checksum = "88261b9deccee56594c11a3460c462c41f58d148598fe70ad77070126a68aba4" dependencies = [ "js-sys", "wasm-bindgen", @@ -3409,9 +3292,9 @@ dependencies = [ [[package]] name = "which" -version = "8.0.5" +version = "8.0.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8f3ef584124b911bcc3875c2f1472e80f24361ceb789bd1c62b3e9a3df9ff43c" +checksum = "bae2f2b2b816647a1cab1acc91f5bd20812d53cb344382635ec2181940c8034f" dependencies = [ "libc", ] @@ -3730,9 +3613,9 @@ checksum = "1ebf944e87a7c253233ad6766e082e3cd714b5d03812acc24c318f549614536e" [[package]] name = "writeable" -version = "0.6.3" +version = "0.6.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1ffae5123b2d3fc086436f8834ae3ab053a283cfac8fe0a0b8eaae044768a4c4" +checksum = "3ad82d2a33cdc9674dc7465672f271e096168fcdbe0f799d9e6db8c5892679dc" [[package]] name = "yoke" @@ -3747,30 +3630,30 @@ dependencies = [ [[package]] name = "yoke-derive" -version = "0.8.2" +version = "0.8.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "de844c262c8848816172cef550288e7dc6c7b7814b4ee56b3e1553f275f1858e" +checksum = "ec8ebde2db3681e8c9980cc27822030e68752690ddfa9473e739aeb4dbde6d71" dependencies = [ "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", "synstructure", ] [[package]] name = "zerocopy" -version = "0.8.55" +version = "0.8.59" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b5a105cd7b140f6eeec8acff2ea38135d3cab283ada58540f629fe51e46696eb" +checksum = "6df92bf3d9227be3d53173901ddbffac2babc27ae50f397776ffd6dc33f800cb" dependencies = [ "zerocopy-derive", ] [[package]] name = "zerocopy-derive" -version = "0.8.55" +version = "0.8.59" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0fe976fb70c78cd64cccfe3a6fc142244e8a77b70959b30faf9d0ac37ee228eb" +checksum = "ac4f328cf2f05d084e496c3e9c3f33ed0a183656a16e1fcec4d464d8373aec82" dependencies = [ "proc-macro2", "quote", @@ -3788,13 +3671,13 @@ dependencies = [ [[package]] name = "zerofrom-derive" -version = "0.1.7" +version = "0.1.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "11532158c46691caf0f2593ea8358fed6bbf68a0315e80aae9bd41fbade684a1" +checksum = "f75b4683f6c7f45248d4d64056a24298c6281e0993356d7d1b4a1a962ef10d4a" dependencies = [ "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", "synstructure", ] @@ -3806,9 +3689,9 @@ checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e" [[package]] name = "zerotrie" -version = "0.2.4" +version = "0.2.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0f9152d31db0792fa83f70fb2f83148effb5c1f5b8c7686c3459e361d9bc20bf" +checksum = "4ea269c3bd32f0a32c321907a2ae912ba6f4649bb0fc764a15627e99a7095a3f" dependencies = [ "displaydoc", "yoke", @@ -3817,9 +3700,9 @@ dependencies = [ [[package]] name = "zerovec" -version = "0.11.6" +version = "0.11.8" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "90f911cbc359ab6af17377d242225f4d75119aec87ea711a880987b18cd7b239" +checksum = "bb0464e17806c1d976d5cba29399c7f08e516e279e2ba493f63123b5fca67dd8" dependencies = [ "yoke", "zerofrom", @@ -3828,13 +3711,13 @@ dependencies = [ [[package]] name = "zerovec-derive" -version = "0.11.3" +version = "0.11.6" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "625dc425cab0dca6dc3c3319506e6593dcb08a9f387ea3b284dbd52a92c40555" +checksum = "34df6fc39dbd26ddc9c10e6a2984476e13acce22e64e4487636ef494369225da" dependencies = [ "proc-macro2", "quote", - "syn 2.0.119", + "syn 3.0.6", ] [[package]] diff --git a/Cargo.toml b/Cargo.toml index aaac2ba..e758ff5 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -4,9 +4,9 @@ default-members = ["crates/*"] resolver = "3" [workspace.package] -version = "1.1.0" +version = "1.2.0" edition = "2024" -rust-version = "1.97.1" +rust-version = "1.96.0" description = "Local cross-repository code intelligence and dependency graph for impact analysis and AI coding agents." license = "Apache-2.0" repository = "https://github.com/dertin/code-system-graph" diff --git a/README.md b/README.md index c1d3e86..d36a4eb 100644 --- a/README.md +++ b/README.md @@ -9,7 +9,7 @@ Map APIs, events, schemas, packages, databases, and ownership across repositorie breaks another service. [![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE) -![Source version](https://img.shields.io/badge/source-v1.1.0-orange.svg) +![Source version](https://img.shields.io/badge/source-v1.2.0-orange.svg) [![crates.io](https://img.shields.io/crates/v/code-system-graph.svg)](https://crates.io/crates/code-system-graph) ![Platforms](https://img.shields.io/badge/validated-Linux%20%7C%20macOS%20%7C%20Windows-1793d1.svg) ![Privacy](https://img.shields.io/badge/privacy-local%20%7C%20no%20telemetry-2ea44f.svg) @@ -85,7 +85,8 @@ graph: | Literal SQL statement and one declared table | Code **reads from** or **writes to** the table | | Package, deployment, documentation, ownership, and config declarations | Dependency, deploys/provides, documents, owns, and config-key links | -It does not guess through dynamic routes, interpolated table names, duplicate providers, or prose. +It does not guess through dynamic routes, interpolated table names, indistinguishable providers, or +prose. Those cases remain incomplete or ambiguous so an agent cannot mistake missing evidence for safety. ## Supported languages and frameworks @@ -94,11 +95,16 @@ Source scanning currently recognizes these focused framework and language patter | Language | HTTP clients | HTTP servers | Database access | Recognized tests | | --- | --- | --- | --- | --- | -| TypeScript / JavaScript | Fetch, Axios | Express, Fastify, NestJS, Next.js App Router | Literal SQL | Not recognized | -| Python | requests, HTTPX, aiohttp, static method registries | FastAPI, Flask | psycopg/psycopg2, PyMySQL, SQLAlchemy, Alembic, literal SQL | pytest, unittest, Factory Boy model links | -| Go | `net/http` | `net/http`, Gin, Chi | Literal SQL | Not recognized | -| Java | WebClient, Feign | Spring MVC | Literal SQL | Not recognized | -| Rust | Reqwest | Axum, Actix Web; advisory Utoipa/OpenAPI operations | SQLx, `mysql_async`, Diesel, literal SQL | Built-in tests, Tokio tests, rstest | +| TypeScript / JavaScript | Fetch, Axios | Express, Fastify, NestJS, Next.js App Router | Literal SQL | Jest, Vitest, Mocha, Playwright; supertest | +| Python | requests, HTTPX, aiohttp, static method registries | FastAPI, Flask | psycopg/psycopg2, PyMySQL, SQLAlchemy, Alembic, literal SQL | pytest, unittest, Factory Boy model links; `TestClient`, Flask `test_client()` | +| Go | `net/http` | `net/http`, Gin, Chi | Literal SQL | `testing` functions; `httptest` | +| Java | WebClient, Feign, RestTemplate | Spring MVC | Literal SQL | JUnit; MockMvc, RestAssured, `WebTestClient` | +| Rust | Reqwest | Axum, Actix Web; advisory Utoipa/OpenAPI operations | SQLx, `mysql_async`, Diesel, literal SQL | Built-in tests, Tokio tests, rstest; Axum `oneshot`, Actix `test::TestRequest` | + +A test that calls an endpoint is linked to the contract it validates and to the handler that +implements it in any of these languages, through base URLs, router prefixes declared in other +files, client wrappers, test helpers, and fixtures. Every scan reports the calls it could not link +and why. Other boundary support is shared across these languages rather than tied to one web framework: @@ -141,7 +147,7 @@ installed to Cargo's binary directory, normally `$HOME/.cargo/bin`. The crates.io package is `code-system-graph`; the user-facing CLI command is `csgraph` (not `code-system-graph`). The hooks runtime installs as `code-system-graph-hooks`. -**From crates.io** (builds locally; requires Rust 1.97.1 or newer): +**From crates.io** (builds locally; requires Rust 1.96.0 or newer): ```bash cargo install code-system-graph code-system-graph-hooks @@ -328,7 +334,7 @@ plugin root. Re-run with `--replace-generated` after local paths change; only a Code System Graph's recognized ownership identity can be replaced. Until local binding creation runs, that MCP entry fails visibly while independent plugin skills and servers remain usable. -The generated MCP is read-only and requires `csgraph 1.1.0` in `PATH`; it does not bundle binaries. +The generated MCP is read-only and requires `csgraph 1.2.0` in `PATH`; it does not bundle binaries. Its plugin, server, and skill share a stable name derived from the declared workspace name, so clones produce the same versioned files. Install it project-locally for workspace-only activation; the generated skill also requires the nearest manifest and MCP `status` to report that workspace. diff --git a/SECURITY.md b/SECURITY.md index 667edf0..0dacbcb 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -2,8 +2,7 @@ ## Supported versions -- `1.0.x`: receives security fixes. -- Older versions: none have been published. +- `1.2.x`: receives security fixes. Platform support applies only to targets listed in the release notes with completed native validation. diff --git a/crates/code-system-graph-cli/Cargo.toml b/crates/code-system-graph-cli/Cargo.toml index 84af0f4..ca4e103 100644 --- a/crates/code-system-graph-cli/Cargo.toml +++ b/crates/code-system-graph-cli/Cargo.toml @@ -26,29 +26,30 @@ path = "src/main.rs" [dependencies] anyhow = "1.0.104" -atomic-write-file = "0.3.0" +atomic-write-file = "0.3.1" axum = "0.8.9" -blake3 = "1.8.5" -clap = { version = "4.6.4", features = ["derive"] } -clap_complete = "4.6.8" -code-system-graph-core = { version = "1.1.0", path = "../code-system-graph-core" } -code-system-graph-hooks = { version = "1.1.0", path = "../code-system-graph-hooks" } -code-system-graph-model = { version = "1.1.0", path = "../code-system-graph-model" } -code-system-graph-store-sqlite = { version = "1.1.0", path = "../code-system-graph-store-sqlite" } -rmcp = { version = "3.1.0", features = ["transport-io"] } +blake3 = "1.8.7" +foldhash = "0.2.0" +clap = { version = "4.6.7", features = ["derive"] } +clap_complete = "4.6.11" +code-system-graph-core = { version = "1.2.0", path = "../code-system-graph-core" } +code-system-graph-hooks = { version = "1.2.0", path = "../code-system-graph-hooks" } +code-system-graph-model = { version = "1.2.0", path = "../code-system-graph-model" } +code-system-graph-store-sqlite = { version = "1.2.0", path = "../code-system-graph-store-sqlite" } +rmcp = { version = "3.5.0", features = ["transport-io"] } notify = "8.2.0" -rusqlite = { version = "0.40.1", features = ["bundled"] } +rusqlite = { version = "0.40.2", features = ["bundled"] } schemars = "1.2.2" serde = { version = "1.0.229", features = ["derive"] } serde_json = "1.0.151" signal-hook = "0.4.4" subtle = "2.6.1" sysinfo = { version = "0.39.6", default-features = false, features = ["system"] } -thiserror = "2.0.19" +thiserror = "2.0.21" tokio = { version = "1.53.1", features = ["macros", "net", "rt-multi-thread", "signal", "sync", "time"] } tokio-util = { version = "0.7.19", features = ["rt"] } tower = { version = "0.5.3", features = ["util", "limit"] } -tower-http = { version = "0.7.0", features = ["catch-panic", "limit", "sensitive-headers", "set-header", "timeout"] } +tower-http = { version = "0.7.1", features = ["catch-panic", "limit", "sensitive-headers", "set-header", "timeout"] } [target.'cfg(unix)'.dependencies] nix = { version = "0.31.3", features = ["fs", "signal"] } @@ -64,8 +65,8 @@ windows-sys = { version = "0.61.2", features = [ workspace = true [dev-dependencies] -jsonschema = { version = "0.49.8", default-features = false } -reqwest = { version = "0.13.4", default-features = false, features = ["json", "rustls"] } -rusqlite = { version = "0.40.1", features = ["bundled"] } -serde-saphyr = "1.0.0" +jsonschema = { version = "0.58.5", default-features = false } +reqwest = { version = "0.13.5", default-features = false, features = ["json", "rustls"] } +rusqlite = { version = "0.40.2", features = ["bundled"] } +serde-saphyr = "1.3.0" tempfile = "3.27.0" diff --git a/crates/code-system-graph-cli/src/agent_plugin.rs b/crates/code-system-graph-cli/src/agent_plugin.rs index 5cc647a..e7d1b82 100644 --- a/crates/code-system-graph-cli/src/agent_plugin.rs +++ b/crates/code-system-graph-cli/src/agent_plugin.rs @@ -24,7 +24,6 @@ use thiserror::Error; use crate::{ApplicationError, application_exit_code, status_workspace}; const BINDING_GENERATOR: &str = "csgraph plugin binding"; -const LEGACY_COMPOSE_GENERATOR: &str = "csgraph plugin compose"; const BINDING_RELATIVE_PATH: &str = ".local/code-system-graph/mcp-binding.json"; const INTEGRATION_RECEIPT_RELATIVE_PATH: &str = ".local/code-system-graph/plugin-integration.json"; @@ -534,7 +533,7 @@ fn write_local_binding( } fn recognized_binding_generator(generator: &str) -> bool { - generator == BINDING_GENERATOR || generator == LEGACY_COMPOSE_GENERATOR + generator == BINDING_GENERATOR } fn read_json_file(path: &Path, file: &'static str) -> Result { diff --git a/crates/code-system-graph-cli/src/fingerprint.rs b/crates/code-system-graph-cli/src/fingerprint.rs new file mode 100644 index 0000000..42a3773 --- /dev/null +++ b/crates/code-system-graph-cli/src/fingerprint.rs @@ -0,0 +1,327 @@ +//! Artifact discovery and fingerprinting with one canonicalization, stat, and read per file. + +use std::collections::{BTreeMap, BTreeSet, HashMap, HashSet}; +use std::fs::Metadata; +#[cfg(unix)] +use std::os::unix::fs::MetadataExt; +use std::path::{Path, PathBuf}; +use std::time::{SystemTime, UNIX_EPOCH}; + +use code_system_graph_core::{ + ExtractionBudgets, ExtractionTracker, IgnorePolicy, JobPhase, discover_repository_files, encode_native_path, parallel +}; +use code_system_graph_model::{ArtifactFingerprint, RepositoryRecord, stable_id_bytes}; + +use super::{ApplicationError, WorkspaceContext, focused_extractors_for_path, read_bounded_bytes}; +use crate::telemetry::PhaseRecorder; +use crate::work_state::{FileStat, FileStatDelta, FileStatKey}; +use crate::{focused_extraction, worker}; + +/// A stat entry is trusted only when the file was last modified at least this long before the +/// entry was recorded, so a same-timestamp rewrite after recording always changes the metadata. +const RACY_WINDOW_NS: i64 = 2_000_000_000; + +/// Fingerprints and stat-cache maintenance produced by one discovery pass. +pub(crate) struct DiscoveredArtifacts { + pub(crate) fingerprints: Vec, + pub(crate) stat_delta: FileStatDelta, + /// Files whose content hash came from the stat cache without a read. + pub(crate) stat_hits: u64, +} + +struct FileRequest<'a> { + repository: &'a RepositoryRecord, + checkout: &'a Path, + relative: PathBuf, + extractors: BTreeSet, +} + +struct FileOutcome { + fingerprints: Vec, + stat: Option<(FileStatKey, FileStat)>, + stat_hit: bool, +} + +/// Discovers and fingerprints every artifact of the selected repositories. +/// +/// `selected` limits discovery to those manifest aliases; `None` scans every repository. +pub(crate) fn discover_artifact_fingerprints( + context: &WorkspaceContext, + selected: Option<&BTreeSet>, + stat_cache: &HashMap, + phases: &mut PhaseRecorder, +) -> Result { + let workers = context.execution_policy.effective_extraction_workers(); + let repository_records = context + .registry + .record + .repositories + .iter() + .map(|repository| (repository.alias.as_str(), repository)) + .collect::>(); + let aliases = context + .manifest + .repos + .keys() + .filter(|alias| selected.is_none_or(|selected| selected.contains(alias.as_str()))) + .collect::>(); + let per_repository = parallel::map_ordered(&aliases, workers, |alias| { + repository_requests(context, &repository_records, alias) + }); + let mut requests = Vec::new(); + for result in per_repository { + requests.extend(result?); + } + let scanned_checkouts = requests + .iter() + .map(|request| request.repository.checkout_id.as_str().to_owned()) + .collect::>(); + phases.complete(JobPhase::Discovery); + + let recorded_before_ns = unix_nanos(SystemTime::now()).saturating_sub(RACY_WINDOW_NS); + let outcomes = parallel::map_ordered(&requests, workers, |request| { + fingerprint_file( + request, + &context.extraction_budgets, + stat_cache, + recorded_before_ns, + ) + }); + + let mut fingerprints = Vec::new(); + let mut seen = HashSet::new(); + let mut stat_delta = FileStatDelta::default(); + let mut stat_hits = 0_u64; + for outcome in outcomes { + let outcome = outcome?; + stat_hits = stat_hits.saturating_add(u64::from(outcome.stat_hit)); + if let Some((key, stat)) = outcome.stat { + if stat_cache.get(&key) != Some(&stat) { + stat_delta.upserts.push((key.clone(), stat)); + } + seen.insert(key); + } + fingerprints.extend(outcome.fingerprints); + } + // Canonicalized symlinks can resolve two requests to one artifact; the last request wins. + fingerprints.sort_by(|left, right| { + focused_extraction::artifact_key(left).cmp(&focused_extraction::artifact_key(right)) + }); + fingerprints.dedup_by(|later, earlier| { + let duplicate = later.checkout_id == earlier.checkout_id + && later.path == earlier.path + && later.extractor == earlier.extractor; + if duplicate { + std::mem::swap(later, earlier); + } + duplicate + }); + stat_delta.removals = stat_cache + .keys() + .filter(|key| scanned_checkouts.contains(&key.0) && !seen.contains(*key)) + .cloned() + .collect(); + stat_delta.removals.sort(); + Ok(DiscoveredArtifacts { + fingerprints, + stat_delta, + stat_hits, + }) +} + +fn repository_requests<'a>( + context: &'a WorkspaceContext, + repository_records: &BTreeMap<&str, &'a RepositoryRecord>, + alias: &str, +) -> Result>, ApplicationError> { + let repository = *repository_records + .get(alias) + .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.to_owned()))?; + let checkout = context + .registry + .checkout_path(alias) + .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.to_owned()))?; + let effective = context + .repository_configs + .get(alias) + .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.to_owned()))?; + let mut files = BTreeMap::>::new(); + let mut declare = |path: &str, extractor: &str| { + files + .entry(PathBuf::from(path)) + .or_default() + .insert(extractor.to_owned()); + }; + for openapi in &effective.openapi { + declare(openapi, "code-system-graph.http.openapi"); + } + for consumer in &effective.http_consumers { + declare(&consumer.source, "code-system-graph.http.declared"); + } + for test in &effective.integration_tests { + declare(&test.path, "code-system-graph.tests.declared"); + } + for implementation in &effective.implementations { + declare( + &implementation.path, + "code-system-graph.implementations.declared", + ); + } + for (relative, extractors) in discover_focused_artifacts(checkout, &effective.ignore_policy)? { + files + .entry(relative) + .or_default() + .extend(extractors.into_iter().map(str::to_owned)); + } + Ok(files + .into_iter() + .map(|(relative, extractors)| FileRequest { + repository, + checkout, + relative, + extractors, + }) + .collect()) +} + +fn discover_focused_artifacts( + checkout: &Path, + ignore_policy: &IgnorePolicy, +) -> Result)>, ApplicationError> { + let mut discovered = Vec::new(); + for relative in discover_repository_files(checkout, ignore_policy, None)? { + worker::report_progress(JobPhase::Discovery, 1); + let extractors = focused_extractors_for_path(checkout, &relative); + if !extractors.is_empty() { + discovered.push((relative, extractors)); + } + } + Ok(discovered) +} + +fn fingerprint_file( + request: &FileRequest<'_>, + budgets: &ExtractionBudgets, + stat_cache: &HashMap, + recorded_before_ns: i64, +) -> Result { + let checkout = request.checkout; + let configured_path = checkout.join(&request.relative); + let canonical_path = + std::fs::canonicalize(&configured_path).map_err(|source| ApplicationError::ReadFile { + path: configured_path.clone(), + source, + })?; + let relative = canonical_path.strip_prefix(checkout).map_err(|_| { + ApplicationError::ArtifactOutsideCheckout { + path: configured_path.clone(), + checkout: checkout.to_path_buf(), + } + })?; + let metadata = + std::fs::metadata(&canonical_path).map_err(|source| ApplicationError::ReadFile { + path: canonical_path.clone(), + source, + })?; + let path = encode_native_path(relative); + code_system_graph_model::validate_safe_path_display(&path.display) + .map_err(|_| ApplicationError::UnsafeArtifactPath)?; + let first_extractor = request + .extractors + .first() + .map_or("code-system-graph.discovery", String::as_str); + let tracker = ExtractionTracker::new(&path.display, first_extractor, budgets); + tracker.check_input_bytes(metadata.len())?; + + let key = ( + request.repository.checkout_id.as_str().to_owned(), + path.bytes.clone(), + ); + let observed = observed_stat(&metadata); + let cached_hash = stat_cache.get(&key).and_then(|stat| { + observed + .as_ref() + .filter(|(size, modified, identity)| { + stat.size_bytes == *size + && stat.modified_unix_ns == *modified + && &stat.file_identity == identity + }) + .map(|_| stat.content_hash.clone()) + }); + let stat_hit = cached_hash.is_some(); + let content_hash = if let Some(hash) = cached_hash { + hash + } else { + let mut tracker = tracker; + let content = read_bounded_bytes(&canonical_path, &mut tracker)?; + stable_id_bytes("artifact-content", &content) + }; + let stat = observed + .filter(|(_, modified, _)| *modified < recorded_before_ns) + .map(|(size_bytes, modified_unix_ns, file_identity)| { + ( + key, + FileStat { + size_bytes, + modified_unix_ns, + file_identity, + content_hash: content_hash.clone(), + }, + ) + }); + let fingerprints = request + .extractors + .iter() + .map(|extractor| ArtifactFingerprint { + repo_id: request.repository.id.clone(), + checkout_id: request.repository.checkout_id.clone(), + path: path.clone(), + extractor: extractor.clone(), + content_hash: content_hash.clone(), + size_bytes: metadata.len(), + }) + .collect::>(); + worker::report_progress( + JobPhase::Fingerprinting, + u64::try_from(fingerprints.len()).unwrap_or(u64::MAX), + ); + Ok(FileOutcome { + fingerprints, + stat, + stat_hit, + }) +} + +/// Returns `(size, modified_unix_ns, file_identity)` when the platform exposes a modification +/// time; files without one are always read. +fn observed_stat(metadata: &Metadata) -> Option<(u64, i64, String)> { + let modified = metadata.modified().ok()?; + Some(( + metadata.len(), + unix_nanos(modified), + file_identity(metadata), + )) +} + +fn unix_nanos(time: SystemTime) -> i64 { + time.duration_since(UNIX_EPOCH) + .ok() + .and_then(|duration| i64::try_from(duration.as_nanos()).ok()) + .unwrap_or(0) +} + +#[cfg(unix)] +fn file_identity(metadata: &Metadata) -> String { + format!( + "{}:{}:{}.{}", + metadata.dev(), + metadata.ino(), + metadata.ctime(), + metadata.ctime_nsec() + ) +} + +#[cfg(not(unix))] +fn file_identity(_metadata: &Metadata) -> String { + String::new() +} diff --git a/crates/code-system-graph-cli/src/focused_extraction.rs b/crates/code-system-graph-cli/src/focused_extraction.rs new file mode 100644 index 0000000..90ee6e1 --- /dev/null +++ b/crates/code-system-graph-cli/src/focused_extraction.rs @@ -0,0 +1,783 @@ +//! Focused extractor batches: reuse of published batches, checkpoint persistence, and bounded +//! parallel extraction that reads each physical file once for all of its extractors. + +use std::collections::BTreeMap; +use std::path::Path; +use std::sync::atomic::{AtomicUsize, Ordering}; +use std::time::Instant; + +use code_system_graph_core::{ + DataDocument, DocumentationDocument, EXTRACTION_CONTRACT_VERSION, EventDocument, ExtractionBudgets, ExtractionTracker, ExtractorBatch, GeneratedClientMetadata, GraphqlDocument, IncrementalPlan, InfrastructureDocument, JobPhase, PackageManifest, ProtobufDocument, SafeConfigDocument, SourceEpistemicStatus, SourceLanguage, SourceObservation, SourceRole, SourceWarning, extract_asyncapi, extract_codeowners, extract_data_artifact, extract_docker_compose, extract_generated_client_metadata, extract_graphql_document_with_tracker, extract_graphql_persisted_operations_with_tracker, extract_helm, extract_kubernetes, extract_markdown, extract_package_manifest_with_tracker, extract_protobuf_with_tracker, extract_safe_config, extract_service_catalog, extract_terraform, inspect_source_syntax, load_extractor_batch_with_budgets, parallel, parse_event_source, parse_go_source_with_tracker, parse_graphql_source_with_tracker, parse_java_source_with_tracker, parse_javascript_source_at_path_with_tracker, parse_literal_sql_source_at_root, parse_protobuf_generated_source, parse_python_source_with_tracker, parse_rust_source_with_tracker, parse_typescript_source_at_path_with_tracker, precheck_focused_source_values, store_extractor_batch +}; +use code_system_graph_model::{ + ArtifactFingerprint, CheckoutId, NativePath, RepoId, StoredExtractorBatch +}; +use serde::Serialize; +use serde::de::DeserializeOwned; + +use super::{ + ApplicationError, FocusedBatchState, WorkspaceContext, cargo_crate_root, current_unix_millis, data_extractor, documentation_extractor, duration_millis, event_extractor, focused_extractor, graphql_extractor, infrastructure_extractor, native_relative_path, portable_path, protobuf_extractor, read_source_file, source_extractor, source_language_for_path, source_syntax_language +}; +use crate::{work_state, worker}; + +/// Typed outputs of one focused extractor invocation. +enum FocusedOutput { + Source(ExtractorBatch), + Package(ExtractorBatch), + GeneratedClient(ExtractorBatch), + Graphql(ExtractorBatch), + Event(ExtractorBatch), + Protobuf(ExtractorBatch), + Data(ExtractorBatch), + Infrastructure(ExtractorBatch), + Documentation(ExtractorBatch), + Config(ExtractorBatch), +} + +/// Owner of an outcome's persisted batch; reused batches stay in their input vectors until the +/// ordered merge moves them out, so no payload is copied. +enum StoredSlot { + Fresh(Box), + Previous(usize), + Checkpointed(usize), +} + +struct ArtifactOutcome { + stored: StoredSlot, + /// `None` when the document holds no graph facts. + output: Option>, + degradations: Vec, + duration_ms: u64, +} + +enum FocusedJob<'a> { + Reuse { + batch: &'a StoredExtractorBatch, + slot: fn(usize) -> StoredSlot, + index: usize, + }, + /// Every fingerprint shares one checkout and one relative path. + Extract { + checkout: &'a Path, + fingerprints: Vec<&'a ArtifactFingerprint>, + }, +} + +/// Artifact identity borrowed from a fingerprint, so lookups over every artifact clone nothing. +pub(super) type BorrowedArtifactKey<'a> = (&'a RepoId, &'a CheckoutId, &'a NativePath, &'a str); + +pub(super) fn artifact_key(fingerprint: &ArtifactFingerprint) -> BorrowedArtifactKey<'_> { + ( + &fingerprint.repo_id, + &fingerprint.checkout_id, + &fingerprint.path, + fingerprint.extractor.as_str(), + ) +} + +/// Lowest job index that failed, so later jobs can stop while every earlier job still runs and +/// the reported error stays independent of scheduling. +struct FailureFence(AtomicUsize); + +impl FailureFence { + fn new() -> Self { + Self(AtomicUsize::new(usize::MAX)) + } + + fn skips(&self, index: usize) -> bool { + index > self.0.load(Ordering::Acquire) + } + + fn record(&self, index: usize) { + self.0.fetch_min(index, Ordering::AcqRel); + } +} + +#[expect( + clippy::too_many_lines, + reason = "job planning, ordered merge, and checkpoint persistence share one fail-closed boundary" +)] +pub(super) fn assemble_focused_batches( + context: &WorkspaceContext, + fingerprints: &[ArtifactFingerprint], + previous: Vec, + plan: &IncrementalPlan, + checkpointed: Vec, + work_state: &mut work_state::WorkState, +) -> Result { + let budgets = &context.extraction_budgets; + let budget_fingerprint = budgets.fingerprint(); + let previous_by_key = previous + .iter() + .enumerate() + .map(|(index, batch)| (artifact_key(&batch.source), (index, batch))) + .collect::>(); + let checkpointed_by_key = checkpointed + .iter() + .enumerate() + .map(|(index, batch)| (artifact_key(&batch.source), (index, batch))) + .collect::>(); + let changed = plan + .changes + .iter() + .map(|change| { + ( + &change.repo_id, + &change.checkout_id, + &change.path, + change.extractor.as_str(), + ) + }) + .collect::>(); + let checkouts = context + .registry + .record + .repositories + .iter() + .filter_map(|repository| { + context + .registry + .checkout_path(&repository.alias) + .map(|path| (repository.id.clone(), path)) + }) + .collect::>(); + + let mut jobs = Vec::>::new(); + for fingerprint in fingerprints + .iter() + .filter(|fingerprint| focused_extractor(&fingerprint.extractor)) + { + let key = artifact_key(fingerprint); + let matches_fingerprint = |(_, batch): &(usize, &StoredExtractorBatch)| { + batch.source.content_hash == fingerprint.content_hash + && batch.extractor_version == EXTRACTION_CONTRACT_VERSION + && batch.budget_fingerprint == budget_fingerprint + }; + let reusable = checkpointed_by_key + .get(&key) + .copied() + .filter(matches_fingerprint) + .map(|(index, batch)| (index, batch, StoredSlot::Checkpointed as fn(usize) -> _)) + .or_else(|| { + previous_by_key + .get(&key) + .copied() + .filter(|_| !changed.contains(&key)) + .filter(matches_fingerprint) + .map(|(index, batch)| (index, batch, StoredSlot::Previous as fn(usize) -> _)) + }); + if let Some((index, batch, slot)) = reusable { + jobs.push(FocusedJob::Reuse { batch, slot, index }); + continue; + } + let checkout = *checkouts.get(&fingerprint.repo_id).ok_or_else(|| { + ApplicationError::RegistryAliasMissing(fingerprint.repo_id.as_str().to_owned()) + })?; + if let Some(FocusedJob::Extract { + fingerprints: group, + .. + }) = jobs.last_mut() + && group.last().is_some_and(|last| { + last.checkout_id == fingerprint.checkout_id && last.path == fingerprint.path + }) + { + group.push(fingerprint); + continue; + } + jobs.push(FocusedJob::Extract { + checkout, + fingerprints: vec![fingerprint], + }); + } + + let indexed_jobs = jobs.iter().enumerate().collect::>(); + let fence = FailureFence::new(); + let results = parallel::map_ordered( + &indexed_jobs, + context.execution_policy.effective_extraction_workers(), + |(index, job)| { + if fence.skips(*index) { + return None; + } + let result = run_job(context, job); + if result.is_err() { + fence.record(*index); + } + Some(result) + }, + ); + + drop(indexed_jobs); + drop(jobs); + drop(previous_by_key); + drop(checkpointed_by_key); + drop(changed); + let mut state = FocusedBatchState { + source_batches: Vec::new(), + package_batches: Vec::new(), + generated_client_batches: Vec::new(), + graphql_batches: Vec::new(), + event_batches: Vec::new(), + protobuf_batches: Vec::new(), + data_batches: Vec::new(), + infrastructure_batches: Vec::new(), + documentation_batches: Vec::new(), + config_batches: Vec::new(), + stored_batches: Vec::new(), + unpublished_batches: Vec::new(), + degradations: Vec::new(), + checkpoint_writes: 0, + artifact_durations_ms: Vec::new(), + }; + let mut previous = previous.into_iter().map(Some).collect::>(); + let mut checkpointed = checkpointed.into_iter().map(Some).collect::>(); + let mut stored_batches = Vec::<(StoredExtractorBatch, bool)>::new(); + let mut fresh_indices = Vec::new(); + let mut first_error = None; + for result in results.into_iter().flatten() { + let outcomes = match result { + Ok(outcomes) => outcomes, + Err(error) => { + first_error.get_or_insert(error); + continue; + } + }; + for outcome in outcomes { + let published = matches!(outcome.stored, StoredSlot::Previous(_)); + let stored = match outcome.stored { + StoredSlot::Fresh(stored) => { + fresh_indices.push(stored_batches.len()); + Some(*stored) + } + StoredSlot::Previous(index) => previous.get_mut(index).and_then(Option::take), + StoredSlot::Checkpointed(index) => { + checkpointed.get_mut(index).and_then(Option::take) + } + }; + let stored = stored.ok_or_else(|| { + ApplicationError::Initialization( + "reused extractor batch was claimed twice".to_owned(), + ) + })?; + stored_batches.push((stored, published)); + state.degradations.extend(outcome.degradations); + state.artifact_durations_ms.push(outcome.duration_ms); + let Some(output) = outcome.output else { + continue; + }; + match *output { + FocusedOutput::Source(batch) => state.source_batches.push(batch), + FocusedOutput::Package(batch) => state.package_batches.push(batch), + FocusedOutput::GeneratedClient(batch) => { + state.generated_client_batches.push(batch); + } + FocusedOutput::Graphql(batch) => state.graphql_batches.push(batch), + FocusedOutput::Event(batch) => state.event_batches.push(batch), + FocusedOutput::Protobuf(batch) => state.protobuf_batches.push(batch), + FocusedOutput::Data(batch) => state.data_batches.push(batch), + FocusedOutput::Infrastructure(batch) => state.infrastructure_batches.push(batch), + FocusedOutput::Documentation(batch) => state.documentation_batches.push(batch), + FocusedOutput::Config(batch) => state.config_batches.push(batch), + } + } + } + drop(previous); + drop(checkpointed); + let fresh = fresh_indices + .iter() + .map(|index| &stored_batches[*index].0) + .collect::>(); + state.checkpoint_writes = work_state + .put_batches( + &fresh, + context.execution_policy.max_checkpoint_cache_bytes, + current_unix_millis(), + ) + .map_err(ApplicationError::Initialization)?; + if let Some(error) = first_error { + return Err(error); + } + sort_by_source(&mut state.source_batches); + sort_by_source(&mut state.package_batches); + sort_by_source(&mut state.generated_client_batches); + sort_by_source(&mut state.graphql_batches); + sort_by_source(&mut state.event_batches); + sort_by_source(&mut state.protobuf_batches); + sort_by_source(&mut state.data_batches); + sort_by_source(&mut state.infrastructure_batches); + sort_by_source(&mut state.documentation_batches); + sort_by_source(&mut state.config_batches); + stored_batches + .sort_by(|left, right| artifact_key(&left.0.source).cmp(&artifact_key(&right.0.source))); + state.stored_batches.reserve_exact(stored_batches.len()); + for (index, (stored, published)) in stored_batches.into_iter().enumerate() { + if !published { + state.unpublished_batches.push(index); + } + state.stored_batches.push(stored); + } + Ok(state) +} + +fn sort_by_source(batches: &mut [ExtractorBatch]) { + batches.sort_by(|left, right| artifact_key(&left.source).cmp(&artifact_key(&right.source))); +} + +/// Payload of a batch without outputs; zero-output batches contribute nothing to the graph, so +/// reusing one needs no decoding. +const EMPTY_PAYLOAD: &[u8] = b"[]"; + +/// Whether a source-scan document holds no facts, so it is persisted without outputs: it only +/// records that the file was scanned and cannot add anything to the graph. +fn fact_free_source_output(output: &FocusedOutput) -> bool { + match output { + FocusedOutput::Graphql(batch) => { + batch.source.extractor == "code-system-graph.graphql.source" + && batch.outputs.iter().all(|document| { + document.types.is_empty() + && document.operations.is_empty() + && document.fragments.is_empty() + && document.persisted_operations.is_empty() + && document.resolvers.is_empty() + && document.federation.is_empty() + && document.warnings.is_empty() + && document.complete + }) + } + FocusedOutput::Event(batch) => { + batch.source.extractor == "code-system-graph.events.source" + && batch.outputs.iter().all(|document| { + document.observations.is_empty() + && document.warnings.is_empty() + && document.specification.is_none() + && !document.incomplete + }) + } + FocusedOutput::Protobuf(batch) => { + batch.source.extractor == "code-system-graph.protobuf.generated" + && batch.outputs.iter().all(|document| { + matches!(document, ProtobufDocument::Generated(markers) if markers.is_empty()) + }) + } + FocusedOutput::Data(batch) => { + batch.source.extractor == "code-system-graph.data.source" + && batch.outputs.iter().all(|document| { + document.tables.is_empty() + && document.migration.is_none() + && document.accesses.is_empty() + && document.frameworks.is_empty() + && document.references.is_empty() + && document.owners.is_empty() + && document.warnings.is_empty() + && !document.incomplete + }) + } + FocusedOutput::Source(_) + | FocusedOutput::Package(_) + | FocusedOutput::GeneratedClient(_) + | FocusedOutput::Infrastructure(_) + | FocusedOutput::Documentation(_) + | FocusedOutput::Config(_) => false, + } +} + +fn run_job( + context: &WorkspaceContext, + job: &FocusedJob<'_>, +) -> Result, ApplicationError> { + match job { + FocusedJob::Reuse { batch, slot, index } => { + let started = Instant::now(); + let output = if batch.output_count == 0 && batch.payload == EMPTY_PAYLOAD { + None + } else { + let output = decode_stored_batch(batch, &context.extraction_budgets)?; + (!fact_free_source_output(&output)).then(|| Box::new(output)) + }; + worker::report_progress(JobPhase::Extraction, 1); + Ok(vec![ArtifactOutcome { + stored: slot(*index), + output, + degradations: Vec::new(), + duration_ms: duration_millis(started.elapsed()), + }]) + } + FocusedJob::Extract { + checkout, + fingerprints, + } => extract_file(context, checkout, fingerprints), + } +} + +fn decode_stored_batch( + stored: &StoredExtractorBatch, + budgets: &ExtractionBudgets, +) -> Result { + fn decode( + stored: &StoredExtractorBatch, + budgets: &ExtractionBudgets, + ) -> Result, ApplicationError> { + Ok(load_extractor_batch_with_budgets(stored, budgets)?) + } + let extractor = stored.source.extractor.as_str(); + Ok(if source_extractor(extractor) { + FocusedOutput::Source(decode(stored, budgets)?) + } else if extractor == "code-system-graph.packages" { + FocusedOutput::Package(decode(stored, budgets)?) + } else if graphql_extractor(extractor) { + FocusedOutput::Graphql(decode(stored, budgets)?) + } else if event_extractor(extractor) { + FocusedOutput::Event(decode(stored, budgets)?) + } else if protobuf_extractor(extractor) { + FocusedOutput::Protobuf(decode(stored, budgets)?) + } else if data_extractor(extractor) { + FocusedOutput::Data(decode(stored, budgets)?) + } else if infrastructure_extractor(extractor) { + FocusedOutput::Infrastructure(decode(stored, budgets)?) + } else if documentation_extractor(extractor) { + FocusedOutput::Documentation(decode(stored, budgets)?) + } else if extractor == "code-system-graph.config.safe" { + FocusedOutput::Config(decode(stored, budgets)?) + } else { + FocusedOutput::GeneratedClient(decode(stored, budgets)?) + }) +} + +fn extract_file( + context: &WorkspaceContext, + checkout: &Path, + fingerprints: &[&ArtifactFingerprint], +) -> Result, ApplicationError> { + let Some(first) = fingerprints.first() else { + return Ok(Vec::new()); + }; + let budgets = &context.extraction_budgets; + let read_started = Instant::now(); + let relative_path = native_relative_path(&first.path); + let artifact_path = checkout.join(&relative_path); + let mut read_tracker = ExtractionTracker::new(&first.path.display, &first.extractor, budgets); + let (source, source_was_lossy) = read_source_file(&artifact_path, &mut read_tracker)?; + let mut read_ms = duration_millis(read_started.elapsed()); + let file = SharedSource { + checkout, + artifact_path: &artifact_path, + relative_path: &relative_path, + text: &source, + lossy: source_was_lossy, + }; + let mut outcomes = Vec::with_capacity(fingerprints.len()); + for fingerprint in fingerprints { + let started = Instant::now(); + let mut tracker = + ExtractionTracker::new(&fingerprint.path.display, &fingerprint.extractor, budgets); + let mut degradations = Vec::new(); + if source_was_lossy { + degradations.push(format!( + "{} contains invalid UTF-8 and was decoded lossily; extracted evidence is incomplete", + fingerprint.path.display + )); + } + let (mut stored, output) = + extract_artifact(&file, fingerprint, &mut tracker, &mut degradations)?; + let output = if fact_free_source_output(&output) { + stored.output_count = 0; + stored.payload = EMPTY_PAYLOAD.to_vec(); + None + } else { + Some(Box::new(output)) + }; + worker::report_progress(JobPhase::Extraction, 1); + outcomes.push(ArtifactOutcome { + stored: StoredSlot::Fresh(Box::new(stored)), + output, + degradations, + duration_ms: duration_millis(started.elapsed()).saturating_add(read_ms), + }); + read_ms = 0; + } + Ok(outcomes) +} + +/// One file read shared by every extractor of that file. +struct SharedSource<'a> { + checkout: &'a Path, + artifact_path: &'a Path, + relative_path: &'a Path, + text: &'a str, + lossy: bool, +} + +impl SharedSource<'_> { + fn language( + &self, + purpose: &str, + portable_path: &str, + ) -> Result { + source_language_for_path(self.relative_path).ok_or_else(|| { + ApplicationError::InvalidSourceObservation(format!( + "unsupported {purpose} language for `{portable_path}`" + )) + }) + } +} + +fn finish( + batch: ExtractorBatch, + tracker: &mut ExtractionTracker, + lossy: bool, + wrap: fn(ExtractorBatch) -> FocusedOutput, +) -> Result<(StoredExtractorBatch, FocusedOutput), ApplicationError> { + let stored = store_extractor_batch(&batch, tracker, lossy)?; + Ok((stored, wrap(batch))) +} + +#[expect( + clippy::too_many_lines, + reason = "one exhaustive dispatch keeps every focused extractor's typed contract visible" +)] +fn extract_artifact( + file: &SharedSource<'_>, + fingerprint: &ArtifactFingerprint, + tracker: &mut ExtractionTracker, + degradations: &mut Vec, +) -> Result<(StoredExtractorBatch, FocusedOutput), ApplicationError> { + let source = file.text; + let lossy = file.lossy; + let extractor = fingerprint.extractor.as_str(); + let portable_path = portable_path(&fingerprint.path.display); + if source_extractor(extractor) { + let observations = extract_source_observations(file, fingerprint, tracker)?; + return finish( + ExtractorBatch::new(fingerprint.clone(), observations), + tracker, + lossy, + FocusedOutput::Source, + ); + } + if extractor == "code-system-graph.packages" { + let manifest = extract_package_manifest_with_tracker(&portable_path, source, tracker)?; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![manifest]), + tracker, + lossy, + FocusedOutput::Package, + ); + } + if graphql_extractor(extractor) { + let document = match extractor { + "code-system-graph.graphql.document" => { + extract_graphql_document_with_tracker(&portable_path, source, tracker)? + } + "code-system-graph.graphql.persisted" => GraphqlDocument { + source_path: portable_path.clone(), + types: Vec::new(), + operations: Vec::new(), + fragments: Vec::new(), + persisted_operations: extract_graphql_persisted_operations_with_tracker( + &portable_path, + source, + tracker, + )?, + resolvers: Vec::new(), + federation: Vec::new(), + complete: true, + warnings: Vec::new(), + }, + _ => { + let language = file.language("GraphQL source", &portable_path)?; + let mut document = parse_graphql_source_with_tracker(language, source, tracker)?; + document.source_path = portable_path; + document + } + }; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Graphql, + ); + } + if event_extractor(extractor) { + let document = if extractor == "code-system-graph.events.asyncapi" { + extract_asyncapi(&portable_path, source)? + } else { + let language = file.language("event source", &portable_path)?; + let mut document = parse_event_source(language, source); + document.source_path = Some(portable_path); + document + }; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Event, + ); + } + if protobuf_extractor(extractor) { + let document = if extractor == "code-system-graph.protobuf" { + ProtobufDocument::File(Box::new(extract_protobuf_with_tracker( + &portable_path, + source, + tracker, + )?)) + } else { + let language = file.language("generated protobuf source", &portable_path)?; + ProtobufDocument::Generated(parse_protobuf_generated_source( + language, + &portable_path, + source, + )) + }; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Protobuf, + ); + } + if data_extractor(extractor) { + let document = if extractor == "code-system-graph.data.source" { + let language = file.language("data source", &portable_path)?; + // The crate root only resolves Rust database-crate migration paths, so the ancestor + // walk is skipped for every file that cannot reach that recognizer. + let crate_root = if language == SourceLanguage::Rust + && (source.contains("sqlx") || source.contains("mysql_async")) + { + cargo_crate_root(file.checkout, file.artifact_path) + } else { + String::new() + }; + let source = if language == SourceLanguage::JavaScript && is_minified_script(source) { + "" + } else { + source + }; + parse_literal_sql_source_at_root(language, &portable_path, &crate_root, source) + } else { + extract_data_artifact(&portable_path, source)? + }; + if extractor == "code-system-graph.data.artifact" && document.incomplete { + degradations.push(format!( + "{} data extraction is incomplete: {:?}", + fingerprint.path.display, document.warnings + )); + } + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Data, + ); + } + if infrastructure_extractor(extractor) { + let document = match extractor { + "code-system-graph.infrastructure.compose" => { + extract_docker_compose(&portable_path, source)? + } + "code-system-graph.infrastructure.kubernetes" => { + extract_kubernetes(&portable_path, source)? + } + "code-system-graph.infrastructure.helm" => extract_helm(&portable_path, source)?, + _ => extract_terraform(&portable_path, source)?, + }; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Infrastructure, + ); + } + if documentation_extractor(extractor) { + let document = match extractor { + "code-system-graph.documents.markdown" => extract_markdown(&portable_path, source)?, + "code-system-graph.documents.codeowners" => extract_codeowners(&portable_path, source)?, + _ => extract_service_catalog(&portable_path, source)?, + }; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Documentation, + ); + } + if extractor == "code-system-graph.config.safe" { + let document = extract_safe_config(&portable_path, source)?; + return finish( + ExtractorBatch::new(fingerprint.clone(), vec![document]), + tracker, + lossy, + FocusedOutput::Config, + ); + } + let metadata = extract_generated_client_metadata(&portable_path, source, tracker)?; + finish( + ExtractorBatch::new(fingerprint.clone(), metadata), + tracker, + lossy, + FocusedOutput::GeneratedClient, + ) +} + +fn extract_source_observations( + file: &SharedSource<'_>, + fingerprint: &ArtifactFingerprint, + tracker: &mut ExtractionTracker, +) -> Result, ApplicationError> { + let source = file.text; + let portable_path = portable_path(&fingerprint.path.display); + let syntax_language = source_syntax_language(&fingerprint.extractor); + let syntax = inspect_source_syntax(syntax_language, &portable_path, source, tracker)?; + let reserved_observations = u64::try_from(syntax.boundary_candidate_count).map_err(|_| { + ApplicationError::InvalidSourceObservation(format!( + "{} contains too many syntax candidates", + fingerprint.path.display + )) + })?; + tracker.charge_work(reserved_observations)?; + precheck_focused_source_values(source, syntax_language, tracker)?; + let mut observations = match fingerprint.extractor.as_str() { + "code-system-graph.source.javascript" => { + parse_javascript_source_at_path_with_tracker(&portable_path, source, tracker)? + } + "code-system-graph.source.typescript" => { + parse_typescript_source_at_path_with_tracker(&portable_path, source, tracker)? + } + "code-system-graph.source.rust" => parse_rust_source_with_tracker(source, tracker)?, + "code-system-graph.source.python" => parse_python_source_with_tracker(source, tracker)?, + "code-system-graph.source.go" => parse_go_source_with_tracker(source, tracker)?, + "code-system-graph.source.java" => parse_java_source_with_tracker(source, tracker)?, + _ => Vec::new(), + }; + if observations + .iter() + .any(|observation| observation.role != SourceRole::Test) + && syntax.boundary_candidate_count == 0 + { + return Err(ApplicationError::InvalidSourceObservation(format!( + "{} produced framework facts without a Tree-sitter boundary candidate", + fingerprint.path.display + ))); + } + if syntax.has_error { + for observation in &mut observations { + observation.status = SourceEpistemicStatus::Incomplete; + if !observation + .warnings + .contains(&SourceWarning::SyntaxErrorRecovery) + { + observation + .warnings + .push(SourceWarning::SyntaxErrorRecovery); + } + } + } + Ok(observations) +} + +/// Whether a script is minified: at least 4 KiB with lines averaging 500 bytes or more. +fn is_minified_script(source: &str) -> bool { + const MINIMUM_BYTES: usize = 4096; + const MINIMUM_AVERAGE_LINE: usize = 500; + source.len() >= MINIMUM_BYTES + && source.len() / source.lines().count().max(1) >= MINIMUM_AVERAGE_LINE +} diff --git a/crates/code-system-graph-cli/src/lib.rs b/crates/code-system-graph-cli/src/lib.rs index e537d2b..5eccc22 100644 --- a/crates/code-system-graph-cli/src/lib.rs +++ b/crates/code-system-graph-cli/src/lib.rs @@ -3,10 +3,13 @@ mod agent_markdown; mod agent_plugin; mod explore; +mod fingerprint; +mod focused_extraction; pub mod http_server; pub mod mcp; mod repository_ownership; mod sync; +mod telemetry; mod work_state; mod worker; @@ -16,23 +19,24 @@ use std::fs::{self, OpenOptions}; use std::io::{Read, Write}; use std::path::{Path, PathBuf}; use std::sync::Arc; -use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH}; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::time::{Duration, SystemTime, UNIX_EPOCH}; pub use agent_plugin::{ AgentPluginCreateMode, AgentPluginCreateReport, AgentPluginCreateRequest, AgentPluginCreateTarget, AgentPluginError, AgentPluginMcpBinding, AgentPluginUninstallReport, AgentPluginUninstallRequest, agent_plugin_exit_code, create_agent_plugin, load_agent_plugin_mcp_binding, uninstall_composed_integration }; use atomic_write_file::AtomicWriteFile; use code_system_graph_core::{ - AffectedTestsRequest, AgentNextAction, AnalyzerVersions, ArtifactKey, BatchAction, BatchPlanError, BitbucketProvider, ChangeAnalysisError, ChangeAnalysisOptions, ChangeError, ChangeImpactReport, ChangeProvider, ChangeRequest, ChangeScope, ChangeSet, CodeGraphConfig, CodeGraphProvider, CommunityError, ConfigDoctorInput, ConfigError, ConfigExtractionError, ContractReport, ContractRequest, CorroborationReport, DataDocument, DataExtractionError, DeclaredImplementation, DeclaredTestCase, DoctorReport, DoctorRequest, DocumentationDocument, DocumentationExtractionError, EXTRACTION_CONTRACT_VERSION, EffectiveRepositoryConfig, EventDocument, EventExtractionError, EventGraphFacts, ExecutionPolicy, ExitCode, ExportReport, ExportRequest, ExtractionBudgets, ExtractionGraphFacts, ExtractionLimitExceeded, ExtractionTracker, ExtractorBatch, ExtractorBatchPlan, FederatedGraph, FreshnessDoctorInput, GeneratedClientError, GeneratedClientMetadata, GitCliChangeProvider, GitHubProvider, GraphqlDocument, GraphqlExtractionError, GraphqlGraphFacts, HttpBoundary, HttpExtractionError, ImpactContext, ImpactError, ImpactReport, ImpactRequest, ImpactTarget, IncrementalPlan, InfrastructureDocument, InfrastructureExtractionError, IntegrityDoctorInput, InterfaceError, LinkError, LocalCodeIntelligenceProvider, LocalContextRequest, LocalEnrichmentInput, LocalEnrichmentStatus, LocalImpactItem, LocalImpactRequest, LocalNeighborDirection, LocalNeighborsRequest, ManifestEdit, ManifestEditError, ManifestError, ManualLinkConfig, ManualLinkError, PackageGraphFacts, PackageManifest, PackageManifestError, PrAuthToken, ProtobufDocument, ProtobufExtractionError, ProtobufGraphFacts, ProviderBudget, ProviderCapability, ProviderDoctorInput, ProviderDoctorStatus, ProviderRequest, ProviderStatus, PullRequestCoordinates, PullRequestError, PullRequestInspectRequest, PullRequestInspection, PullRequestListPage, PullRequestListRequest, PullRequestListState, PullRequestProvider, PullRequestProviderConfig, PullRequestProviderKind, QueryError, RecommendedCommand, RegisteredWorkspace, RegistryError, ReqwestPrHttpTransport, SafeConfigDocument, SchemaDoctorInput, SearchFilters, SearchReport, SearchRequest, SourceEpistemicStatus, SourceGraphFacts, SourceLanguage, SourceObservation, SourceRole, SourceSymbolIdentity, SourceSyntaxError, SourceSyntaxLanguage, SourceWarning, SymbolAnchor, SymbolCorroboration, TraceError, TraversalReport, TraversalRequest, WorkspaceManifest, affected_link_keys, analyze_changes, analyze_communities_with_progress, analyze_impact, apply_openapi_override, classify_interface_error, commit_manifest_edit, compare_community_snapshots, corroborate_repository, declared_implementation, declared_test_case, doctor, documents_to_graph, encode_native_path, event_documents_to_graph, export_graph, extract_asyncapi, extract_codeowners, extract_data_artifact, extract_docker_compose, extract_generated_client_metadata, extract_graphql_document_with_tracker, extract_graphql_persisted_operations_with_tracker, extract_helm, extract_kubernetes, extract_markdown, extract_openapi_with_tracker, extract_package_manifest_with_tracker, extract_protobuf_with_tracker, extract_safe_config, extract_service_catalog, extract_terraform, graphql_documents_to_graph, inspect_contracts, inspect_source_syntax, link_declared_implementations_with_ambiguities, link_declared_tests_with_ambiguities, link_http_boundaries_with_ambiguities, link_registered_package_owners, load_extractor_batch_with_budgets, merge_affected_link_neighborhoods, package_manifest_to_graph, parse_event_source, parse_go_source_with_tracker, parse_graphql_source_with_tracker, parse_java_source_with_tracker, parse_javascript_source_at_path_with_tracker, parse_literal_sql_source_at_root, parse_manifest, parse_manifest_with_extensions, parse_protobuf_generated_source, parse_python_source_with_tracker, parse_rust_source_with_tracker, parse_typescript_source_at_path_with_tracker, plan_extractor_batches, plan_incremental_scan, precheck_focused_source_values, preview_add_manual_link, preview_add_repository, preview_remove_repository, protobuf_documents_to_graph, register_workspace, resolve_manual_links, resolve_repository_config, resolve_repository_config_with_use_gitignore, search, source_observations_to_graph, store_extractor_batch, traverse + AffectedTestsRequest, AgentNextAction, AnalyzerVersions, AuthorityMap, BatchPlanError, BitbucketProvider, ChangeAnalysisError, ChangeAnalysisOptions, ChangeError, ChangeImpactReport, ChangeProvider, ChangeRequest, ChangeScope, ChangeSet, CodeGraphConfig, CodeGraphProvider, CommunityError, ConfigDoctorInput, ConfigError, ConfigExtractionError, ContractReport, ContractRequest, CorroborationReport, DataDocument, DataExtractionError, DeclaredImplementation, DeclaredTestCase, DoctorReport, DoctorRequest, DocumentationDocument, DocumentationExtractionError, EXTRACTION_CONTRACT_VERSION, EffectiveRepositoryConfig, EventDocument, EventExtractionError, EventGraphFacts, ExecutionPolicy, ExitCode, ExportReport, ExportRequest, ExtractionBudgets, ExtractionGraphFacts, ExtractionLimitExceeded, ExtractionTracker, ExtractorBatch, FederatedGraph, FreshnessDoctorInput, GeneratedClientError, GeneratedClientMetadata, GitCliChangeProvider, GitHubProvider, GraphqlDocument, GraphqlExtractionError, GraphqlGraphFacts, HttpBoundary, HttpExtractionError, ImpactContext, ImpactError, ImpactReport, ImpactRequest, ImpactTarget, IncrementalPlan, InfrastructureDocument, InfrastructureExtractionError, IntegrityDoctorInput, InterfaceError, LocalCodeIntelligenceProvider, LocalContextRequest, LocalEnrichmentInput, LocalEnrichmentStatus, LocalImpactItem, LocalImpactRequest, LocalNeighborDirection, LocalNeighborsRequest, ManifestEdit, ManifestEditError, ManifestError, ManualLinkConfig, ManualLinkError, PackageGraphFacts, PackageManifest, PackageManifestError, PrAuthToken, ProtobufDocument, ProtobufExtractionError, ProtobufGraphFacts, ProviderBudget, ProviderCapability, ProviderDoctorInput, ProviderDoctorStatus, ProviderRequest, ProviderStatus, PullRequestCoordinates, PullRequestError, PullRequestInspectRequest, PullRequestInspection, PullRequestListPage, PullRequestListRequest, PullRequestListState, PullRequestProvider, PullRequestProviderConfig, PullRequestProviderKind, QueryError, RecommendedCommand, RegisteredWorkspace, RegistryError, RepositorySourceFile, ReqwestPrHttpTransport, SafeConfigDocument, SchemaDoctorInput, SearchFilters, SearchReport, SearchRequest, SourceGraphFacts, SourceLanguage, SourceObservation, SourceRole, SourceSymbolIdentity, SourceSyntaxError, SourceSyntaxLanguage, SymbolAnchor, SymbolCorroboration, TraceError, TraversalReport, TraversalRequest, WorkspaceManifest, analyze_changes, analyze_communities_with_progress, analyze_impact, apply_openapi_override, classify_interface_error, commit_manifest_edit, compare_community_snapshots, compose_client_flows, compose_router_mounts, corroborate_repository, declared_implementation, declared_test_case, doctor, documents_to_graph, event_documents_to_graph, export_graph, extract_openapi_with_tracker, graphql_documents_to_graph, inspect_contracts, link_http_routes, link_registered_package_owners, load_extractor_batch_with_budgets, normalize_authority, package_manifest_to_graph, parse_manifest, parse_manifest_with_extensions, plan_incremental_scan, preview_add_manual_link, preview_add_repository, preview_remove_repository, protobuf_documents_to_graph, register_workspace, resolve_manual_links, resolve_repository_config, resolve_repository_config_with_use_gitignore, search, source_observations_to_graph, traverse }; pub use code_system_graph_core::{ ConfigSource, DEFAULT_EXCLUDES, IgnorePolicy, PROTECTED_EXCLUDES, discover_repository_files }; use code_system_graph_model::{ - ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityDelta, CommunityId, CommunityScope, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, ExtractorRunStatus, FreshnessSummary, LinkDecision, LinkStatus, Node, NodeId, NodeKind, OverallFreshness, Provenance, RepoFreshness, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, ToolEnvelope, ToolStatus, TraceReport, WorkspaceRecord, stable_id, stable_id_bytes + ArtifactChange, ArtifactChangeKind, ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityDelta, CommunityId, CommunityScope, CommunitySnapshot, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, ExtractorRunStatus, FreshnessSummary, HttpLinkCoverage, HttpLinkGap, HttpLinkReport, LinkDecision, LinkStatus, NativePathEncoding, Node, NodeId, NodeKind, OverallFreshness, Provenance, RepoFreshness, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, ToolEnvelope, ToolStatus, TraceReport, WorkspaceRecord, stable_id, stable_id_bytes }; use code_system_graph_store_sqlite::{ - ManualLinkDisposition, ManualLinkRecord, ProviderCapabilityRecord, QueryCacheRecord, SnapshotBatch, SqliteStore, StoreError, StoreLock, latest_schema_version + ArtifactDelta, ManualLinkDisposition, ManualLinkRecord, ProviderCapabilityRecord, QueryCacheRecord, SnapshotBatch, SqliteStore, StoreError, StoreLock, schema_identity }; pub use explore::{ ExploreCoverage, ExploreEvidenceLocation, ExploreExecution, ExploreFederatedHandoff, ExploreInput, ExploreLocalRelationship, ExploreReport, ExploreRepositoryContext, explore_repository @@ -51,6 +55,7 @@ pub use worker::{ const MAX_TRACE_DEPTH: usize = 32; const MAX_SCAN_DEGRADATIONS: usize = 25; +const DATA_ARTIFACT_EXTRACTOR: &str = "code-system-graph.data.artifact"; const QUERY_DELIVERY_REVISION: u32 = 3; const GENERATED_STATE_IGNORE_RULE: &[u8] = b".code-system-graph/"; pub(crate) const CODEGRAPH_DISABLED_CODE: &str = "codegraph_disabled"; @@ -134,9 +139,6 @@ pub enum ApplicationError { /// Incremental extractor batch state was inconsistent. #[error(transparent)] BatchPlan(#[from] BatchPlanError), - /// Deterministic linking failed. - #[error(transparent)] - Link(#[from] LinkError), /// Exact manual relationship resolution failed. #[error(transparent)] ManualLink(#[from] ManualLinkError), @@ -277,7 +279,6 @@ pub const fn application_exit_code(error: &ApplicationError) -> ExitCode { | ApplicationError::SourceSyntax(_) | ApplicationError::InvalidSourceObservation(_) | ApplicationError::BatchPlan(_) - | ApplicationError::Link(_) | ApplicationError::ManualLink(_) | ApplicationError::Registry(_) | ApplicationError::Config(_) @@ -388,10 +389,25 @@ pub struct ScanOverrides { pub codegraph_binary: Option, /// Restrict extractor work to one registered alias while reusing other repository batches. pub repository: Option, + /// Restrict discovery to these aliases while reusing other repository batches. An empty list + /// selects every repository; `repository` takes precedence when both are set. + pub touched_repositories: Vec, /// Recompute selected extractor batches even when fingerprints are unchanged. pub force: bool, } +impl ScanOverrides { + /// Returns the aliases selected for discovery, or `None` when every repository is scanned. + #[must_use] + pub fn selected_aliases(&self) -> Option> { + if let Some(alias) = &self.repository { + return Some(BTreeSet::from([alias.clone()])); + } + (!self.touched_repositories.is_empty()) + .then(|| self.touched_repositories.iter().cloned().collect()) + } +} + /// Action performed by one effective repository discovery rule. #[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] #[serde(rename_all = "snake_case")] @@ -523,8 +539,8 @@ pub struct ExtendedConfigReport { pub struct WorkspaceStatus { /// Workspace name. pub workspace: String, - /// Applied `SQLite` schema version. - pub schema_version: i64, + /// Applied `SQLite` schema identity. + pub schema_id: String, /// Result of the latest quick integrity check. pub integrity_ok: bool, /// Aggregate conservative freshness. @@ -533,8 +549,26 @@ pub struct WorkspaceStatus { pub repositories: Vec, /// Last persisted finite-watcher lifecycle state. pub watcher: WatcherStatus, + /// HTTP link coverage of the published graph. + pub http_links: HttpLinkStatus, +} + +/// HTTP link coverage of a published graph: counts by outcome and the calls left without a +/// provider edge. +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] +pub struct HttpLinkStatus { + /// Consumer and test calls by link outcome. + pub coverage: HttpLinkCoverage, + /// Total number of calls without a provider edge. + pub gap_count: usize, + /// Calls without a provider edge, bounded: calls without a provider first, then ambiguous + /// calls with their candidates, then external calls. + pub gaps: Vec, } +/// Most HTTP link gaps listed by `status`. +pub const STATUS_HTTP_LINK_GAP_LIMIT: usize = 50; + /// Persisted lifecycle of the foreground sync watcher. #[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] #[serde(rename_all = "snake_case")] @@ -578,8 +612,8 @@ pub struct RestoreSummary { pub source: String, /// Safety backup of the replaced database. pub safety_backup: Option, - /// Schema version after restore. - pub schema_version: i64, + /// Schema identity after restore. + pub schema_id: String, } /// Compact workspace registry item returned by CLI listing. @@ -949,14 +983,13 @@ struct GraphAssembly { edges: Vec, evidence: Vec, link_decisions: Vec, - link_node_keys: BTreeMap, degradations: Vec, coverage_gaps: Vec, + http_links: HttpLinkReport, } struct FocusedBatchState { source_batches: Vec>, - previous_source_batches: Vec>, package_batches: Vec>, generated_client_batches: Vec>, graphql_batches: Vec>, @@ -967,8 +1000,9 @@ struct FocusedBatchState { documentation_batches: Vec>, config_batches: Vec>, stored_batches: Vec, + /// Indices into `stored_batches` of batches the current graph does not store yet. + unpublished_batches: Vec, degradations: Vec, - force_relink: bool, checkpoint_writes: u64, artifact_durations_ms: Vec, } @@ -1255,7 +1289,7 @@ pub fn scan_workspace_with_overrides( /// Scans through an explicitly selected compatible worker executable. /// /// This entry point lets an embedding application re-execute its own binary after dispatching -/// `__worker-v1` to [`run_worker_from_stdio`], so using the library does not require Cargo to have +/// `__worker` to [`run_worker_from_stdio`], so using the library does not require Cargo to have /// built the `csgraph` binary target next to the host executable. /// /// # Errors @@ -1279,31 +1313,71 @@ pub(crate) fn scan_workspace_direct( scan_workspace_direct_with_mode(config_path, database_path, overrides, false) } -pub(crate) fn scan_workspace_direct_for_sync( +pub(crate) fn scan_loaded_workspace_for_sync( + context: WorkspaceContext, + phases: telemetry::PhaseRecorder, + database_path: &Path, + overrides: &ScanOverrides, + reuse_unchanged_with_codegraph: bool, +) -> Result { + scan_loaded_workspace( + context, + phases, + database_path, + overrides, + reuse_unchanged_with_codegraph, + ) +} + +fn scan_workspace_direct_with_mode( config_path: &Path, database_path: &Path, overrides: &ScanOverrides, reuse_unchanged_with_codegraph: bool, ) -> Result { - scan_workspace_direct_with_mode( - config_path, + let phases = telemetry::PhaseRecorder::start(); + let context = load_workspace_context(config_path, overrides)?; + scan_loaded_workspace( + context, + phases, database_path, overrides, reuse_unchanged_with_codegraph, ) } +fn scan_loaded_workspace( + context: WorkspaceContext, + phases: telemetry::PhaseRecorder, + database_path: &Path, + overrides: &ScanOverrides, + reuse_unchanged_with_codegraph: bool, +) -> Result { + let bytes_before = CONTENT_BYTES_READ.load(Ordering::Relaxed); + let mut summary = run_loaded_scan( + context, + phases, + database_path, + overrides, + reuse_unchanged_with_codegraph, + )?; + summary.execution.content_bytes_read = CONTENT_BYTES_READ + .load(Ordering::Relaxed) + .saturating_sub(bytes_before); + Ok(summary) +} + #[expect( clippy::too_many_lines, reason = "Atomic scan orchestration keeps lock, resume, publication, and summary sequencing visible" )] -fn scan_workspace_direct_with_mode( - config_path: &Path, +fn run_loaded_scan( + context: WorkspaceContext, + mut phases: telemetry::PhaseRecorder, database_path: &Path, overrides: &ScanOverrides, reuse_unchanged_with_codegraph: bool, ) -> Result { - let context = load_workspace_context(config_path, overrides)?; worker::report_progress(code_system_graph_core::JobPhase::Configuration, 1); if let Some(requested) = &overrides.workspace && requested != &context.manifest.name @@ -1313,215 +1387,109 @@ fn scan_workspace_direct_with_mode( manifest: context.manifest.name, }); } - let mut fingerprints = discover_artifact_fingerprints(&context)?; - worker::report_progress( - code_system_graph_core::JobPhase::Fingerprinting, - u64::try_from(fingerprints.len()).unwrap_or(u64::MAX), - ); let _writer_lock = StoreLock::acquire(database_path, Duration::from_mins(5))?; let mut store = SqliteStore::open(database_path)?; let database_instance_id = store.database_instance_id()?; let mut work_state = work_state::WorkState::open(database_path, &database_instance_id) .map_err(ApplicationError::Initialization)?; - let mut previous_fingerprints = - match store.load_current_artifact_fingerprints(&context.manifest.name) { - Ok(previous) => previous, - Err(StoreError::CurrentSnapshotMissing(_)) => Vec::new(), - Err(error) => return Err(error.into()), - }; - if let Some(alias) = &overrides.repository { - let selected = context - .registry - .record - .repositories - .iter() - .find(|repository| &repository.alias == alias) - .ok_or_else(|| ApplicationError::UnknownOverrideRepository(alias.clone()))? - .id - .clone(); - fingerprints.retain(|fingerprint| fingerprint.repo_id == selected); + phases.complete(code_system_graph_core::JobPhase::Configuration); + let selected_aliases = overrides.selected_aliases(); + let selected_repositories = selected_aliases + .as_ref() + .map(|aliases| { + aliases + .iter() + .map(|alias| { + context + .registry + .record + .repositories + .iter() + .find(|repository| &repository.alias == alias) + .map(|repository| repository.id.clone()) + .ok_or_else(|| ApplicationError::UnknownOverrideRepository(alias.clone())) + }) + .collect::, _>>() + }) + .transpose()?; + let stat_cache = if overrides.force { + std::collections::HashMap::new() + } else { + work_state + .load_file_stats() + .map_err(ApplicationError::Initialization)? + }; + let discovered = fingerprint::discover_artifact_fingerprints( + &context, + selected_aliases.as_ref(), + &stat_cache, + &mut phases, + )?; + drop(stat_cache); + let stat_cache_hits = discovered.stat_hits; + work_state + .apply_file_stat_delta(&discovered.stat_delta) + .map_err(ApplicationError::Initialization)?; + let mut fingerprints = discovered.fingerprints; + worker::report_progress( + code_system_graph_core::JobPhase::Fingerprinting, + u64::try_from(fingerprints.len()).unwrap_or(u64::MAX), + ); + phases.complete(code_system_graph_core::JobPhase::Fingerprinting); + let extraction_workers = + u64::try_from(context.execution_policy.effective_extraction_workers()).unwrap_or(1); + let budget_fingerprint = context.extraction_budgets.fingerprint(); + let maximum_payload_bytes = context + .extraction_budgets + .max_serialized_output_bytes_per_artifact; + let mut previous_fingerprints = None; + if let Some(selected) = &selected_repositories { + let previous = load_current_fingerprints(&store, &context.manifest.name)?; fingerprints.extend( - previous_fingerprints + previous .iter() - .filter(|fingerprint| fingerprint.repo_id != selected) + .filter(|fingerprint| !selected.contains(&fingerprint.repo_id)) .cloned(), ); - fingerprints.sort_by_key(|fingerprint| ArtifactKey::from(fingerprint)); - if overrides.force { - previous_fingerprints.retain(|fingerprint| fingerprint.repo_id != selected); - } - } else if overrides.force { - previous_fingerprints.clear(); + fingerprints.sort_by(|left, right| { + focused_extraction::artifact_key(left).cmp(&focused_extraction::artifact_key(right)) + }); + previous_fingerprints = Some(previous); } - let fingerprint_hashes = fingerprints - .iter() - .map(|fingerprint| fingerprint.content_hash.as_str()) - .collect::>() - .join(":"); - let snapshot_id = stable_id( - "snapshot", - &format!( - "{}:{fingerprint_hashes}:community-engine-v1", - context.registry.record.manifest_hash - ), - ); - let mut previous_extractor_batches = match store.load_current_extractor_batches_with_limit( - &context.manifest.name, - context - .extraction_budgets - .max_serialized_output_bytes_per_artifact, - ) { - Ok(previous) => previous, - Err(StoreError::CurrentSnapshotMissing(_)) => Vec::new(), - Err(error) => return Err(error.into()), - }; - let budget_fingerprint = context.extraction_budgets.fingerprint(); - let candidate_material = serde_json::to_vec(&( + let snapshot_id = snapshot_identity( &context.registry.record.manifest_hash, - &fingerprints, &budget_fingerprint, - EXTRACTION_CONTRACT_VERSION, - )) - .map_err(|error| ApplicationError::Initialization(error.to_string()))?; - let candidate_fingerprint = stable_id_bytes("work-candidate-v1", &candidate_material); - let _resumed_candidate = work_state - .begin_candidate( - &context.manifest.name, - &candidate_fingerprint, - current_unix_millis(), - ) - .map_err(ApplicationError::Initialization)?; - let cached_batches = if overrides.force { - Vec::new() - } else { - work_state - .load_batches( - &fingerprints, - &budget_fingerprint, - EXTRACTION_CONTRACT_VERSION, - context - .extraction_budgets - .max_serialized_output_bytes_per_artifact, - current_unix_millis(), - ) - .map_err(ApplicationError::Initialization)? - }; - let cached_keys = cached_batches - .iter() - .map(|batch| ArtifactKey::from(&batch.source)) - .collect::>(); - for cached in cached_batches { - let key = ArtifactKey::from(&cached.source); - if !previous_extractor_batches - .iter() - .any(|batch| ArtifactKey::from(&batch.source) == key && batch.source == cached.source) - { - previous_extractor_batches.push(cached); - } - } - let checkpoint_hits = u64::try_from(cached_keys.len()).unwrap_or(u64::MAX); - if overrides.repository.is_some() - && previous_extractor_batches - .iter() - .any(|batch| batch.budget_fingerprint != budget_fingerprint) - { - return Err(ApplicationError::PartialScanBudgetChanged); - } - let previous_graph = match store.load_current_graph(&context.manifest.name) { - Ok(graph) => graph, - Err(StoreError::CurrentSnapshotMissing(_)) => (Vec::new(), Vec::new()), - Err(error) => return Err(error.into()), - }; - let previous_communities = match store.load_current_community_snapshot(&context.manifest.name) { - Ok(snapshot) => Some(snapshot), - Err(StoreError::CurrentSnapshotMissing(_) | StoreError::CommunitySnapshotMissing(_)) => { - None - } - Err(error) => return Err(error.into()), - }; + &fingerprints, + ); let previous_manifest_matches = match store.load_workspace_registry(&context.manifest.name) { Ok(previous) => previous.manifest_hash == context.registry.record.manifest_hash, Err(StoreError::RegistryIncomplete(_)) => false, Err(error) => return Err(error.into()), }; - let plan = plan_incremental_scan(&previous_fingerprints, &fingerprints); - let staged_snapshot = if let Ok(candidate) = - work_state.load_candidate_snapshot(&context.manifest.name, &candidate_fingerprint) - { - candidate - } else { - work_state - .complete_candidate(&context.manifest.name) - .map_err(ApplicationError::Initialization)?; - work_state - .begin_candidate( - &context.manifest.name, - &candidate_fingerprint, - current_unix_millis(), - ) - .map_err(ApplicationError::Initialization)?; - None - }; - if let Some(candidate) = staged_snapshot { - publish_snapshot_candidate( - &mut store, - SnapshotBatch { - workspace: &context.registry.record, - snapshot_id: &candidate.snapshot_id, - nodes: &candidate.nodes, - edges: &candidate.edges, - evidence: &candidate.evidence, - fingerprints: &candidate.fingerprints, - extractor_batches: &candidate.extractor_batches, - extractor_runs: &candidate.extractor_runs, - manual_links: &candidate.manual_links, - community_snapshot: Some(&candidate.community_snapshot), - }, - &candidate.coverage_gaps, - )?; - work_state - .complete_candidate(&context.manifest.name) - .map_err(ApplicationError::Initialization)?; - let mut execution = candidate.execution; - execution.checkpoint_hits = checkpoint_hits; - return Ok(ScanSummary { - execution, - workspace: context.manifest.name, - snapshot_id: candidate.snapshot_id, - node_count: candidate.nodes.len(), - edge_count: candidate.edges.len(), - evidence_count: candidate.evidence.len(), - community_count: candidate.community_snapshot.communities.len(), - community_delta_count: candidate.community_delta_count, - discovered_input_count: candidate.fingerprints.len(), - changed_input_count: plan.changed_count(), - reused_snapshot: false, - corroborated_symbol_count: candidate.corroborated_symbol_count, - affected_test_count: candidate.affected_test_count, - degradation_count: candidate.degradations.len(), - degradations: candidate.degradations, - }); - } - if previous_manifest_matches - && !plan.has_changes() - && focused_batch_cache_complete( - &fingerprints, - &previous_extractor_batches, - &context.extraction_budgets, - ) - && previous_communities.is_some() + if !overrides.force + && previous_manifest_matches && (!overrides.codegraph || reuse_unchanged_with_codegraph) + && store + .find_current_snapshot_id(&context.manifest.name)? + .as_deref() + == Some(snapshot_id.as_str()) + && let Some(community_count) = store.current_community_count(&context.manifest.name)? { - work_state - .complete_candidate(&context.manifest.name) - .map_err(ApplicationError::Initialization)?; let current = store.current_snapshot_summary(&context.manifest.name)?; - let (degradation_count, degradations) = finalize_scan_degradations( - stored_batch_degradations(&previous_extractor_batches, &context.extraction_budgets)?, - ); + let (degradation_count, degradations) = + finalize_scan_degradations(stored_batch_degradations( + &store.load_current_degradation_batches( + &context.manifest.name, + maximum_payload_bytes, + DATA_ARTIFACT_EXTRACTOR, + )?, + &context.extraction_budgets, + )?); return Ok(ScanSummary { execution: code_system_graph_core::ExecutionSummary { - checkpoint_hits, + stat_cache_hits, + extraction_workers, + phases: phases.into_phases(), ..code_system_graph_core::ExecutionSummary::default() }, workspace: context.manifest.name, @@ -1529,9 +1497,7 @@ fn scan_workspace_direct_with_mode( node_count: current.node_count, edge_count: current.edge_count, evidence_count: current.evidence_count, - community_count: previous_communities - .as_ref() - .map_or(0, |snapshot| snapshot.communities.len()), + community_count, community_delta_count: 0, discovered_input_count: fingerprints.len(), changed_input_count: 0, @@ -1543,31 +1509,91 @@ fn scan_workspace_direct_with_mode( }); } - let batch_plan = plan_extractor_batches(&plan); - let focused_batches = assemble_focused_batches( + let previous_fingerprints = match (previous_fingerprints, &selected_repositories) { + (Some(mut previous), Some(selected)) if overrides.force => { + previous.retain(|fingerprint| !selected.contains(&fingerprint.repo_id)); + previous + } + (Some(previous), _) => previous, + (None, _) if overrides.force => Vec::new(), + (None, _) => load_current_fingerprints(&store, &context.manifest.name)?, + }; + let plan = plan_incremental_scan(&previous_fingerprints, &fingerprints); + drop(previous_fingerprints); + let previous_extractor_batches = match store + .load_current_extractor_batches_with_limit(&context.manifest.name, maximum_payload_bytes) + { + Ok(previous) => previous, + Err(StoreError::CurrentSnapshotMissing(_)) => Vec::new(), + Err(error) => return Err(error.into()), + }; + if selected_repositories.is_some() + && previous_extractor_batches + .iter() + .any(|batch| batch.budget_fingerprint != budget_fingerprint) + { + return Err(ApplicationError::PartialScanBudgetChanged); + } + let published_batches = previous_extractor_batches + .iter() + .map(|batch| (focused_extraction::artifact_key(&batch.source), batch)) + .collect::>(); + let checkpoint_candidates = fingerprints + .iter() + .filter(|fingerprint| focused_extractor(&fingerprint.extractor)) + .filter(|fingerprint| { + !published_batches + .get(&focused_extraction::artifact_key(fingerprint)) + .is_some_and(|batch| { + batch.source == **fingerprint + && batch.extractor_version == EXTRACTION_CONTRACT_VERSION + && batch.budget_fingerprint == budget_fingerprint + }) + }) + .collect::>(); + drop(published_batches); + let cached_batches = if overrides.force { + Vec::new() + } else { + work_state + .load_batches( + &checkpoint_candidates, + &budget_fingerprint, + EXTRACTION_CONTRACT_VERSION, + maximum_payload_bytes, + current_unix_millis(), + ) + .map_err(ApplicationError::Initialization)? + }; + drop(checkpoint_candidates); + let checkpoint_hits = u64::try_from(cached_batches.len()).unwrap_or(u64::MAX); + let previous_communities = match store.load_current_community_snapshot(&context.manifest.name) { + Ok(snapshot) => Some(snapshot), + Err(StoreError::CurrentSnapshotMissing(_) | StoreError::CommunitySnapshotMissing(_)) => { + None + } + Err(error) => return Err(error.into()), + }; + + let focused_batches = focused_extraction::assemble_focused_batches( &context, &fingerprints, - &previous_extractor_batches, - &batch_plan, - &cached_keys, + previous_extractor_batches, + &plan, + cached_batches, &mut work_state, )?; - work_state - .set_candidate_phase(&context.manifest.name, "extracted", current_unix_millis()) - .map_err(ApplicationError::Initialization)?; + phases.complete(code_system_graph_core::JobPhase::Extraction); + let previous_graph = match store.load_current_graph(&context.manifest.name) { + Ok(graph) => graph, + Err(StoreError::CurrentSnapshotMissing(_)) => (Vec::new(), Vec::new()), + Err(error) => return Err(error.into()), + }; worker::report_progress( code_system_graph_core::JobPhase::Extraction, u64::try_from(focused_batches.stored_batches.len()).unwrap_or(u64::MAX), ); let mut graph = assemble_graph(&context, &fingerprints, &focused_batches)?; - relink_affected_graph( - &mut graph, - &plan, - &batch_plan, - &focused_batches, - &previous_graph.0, - &previous_graph.1, - )?; let corroboration = if overrides.codegraph { run_codegraph_corroboration( &context, @@ -1591,6 +1617,7 @@ fn scan_workspace_direct_with_mode( code_system_graph_core::JobPhase::GraphAssembly, u64::try_from(graph.nodes.len().saturating_add(graph.edges.len())).unwrap_or(u64::MAX), ); + phases.complete(code_system_graph_core::JobPhase::GraphAssembly); let GraphAssembly { nodes, edges, @@ -1598,14 +1625,13 @@ fn scan_workspace_direct_with_mode( link_decisions, degradations: graph_degradations, coverage_gaps, - .. + http_links, } = graph; let community_config = default_community_config(); let community_snapshot = previous_communities .as_ref() .filter(|previous| { - previous.snapshot_id == snapshot_id - && previous.config == community_config + previous.config == community_config && community_topology_unchanged( &previous_graph.0, &previous_graph.1, @@ -1613,7 +1639,10 @@ fn scan_workspace_direct_with_mode( &edges, ) }) - .cloned() + .map(|previous| CommunitySnapshot { + snapshot_id: snapshot_id.clone(), + ..previous.clone() + }) .map_or_else( || { analyze_communities_with_progress( @@ -1641,6 +1670,7 @@ fn scan_workspace_direct_with_mode( code_system_graph_core::JobPhase::Communities, u64::try_from(community_snapshot.communities.len()).unwrap_or(u64::MAX), ); + phases.complete(code_system_graph_core::JobPhase::Communities); let extractor_runs = extractor_runs(&snapshot_id, &fingerprints, &plan); let manual_link_records = persisted_manual_link_records(&snapshot_id, &link_decisions)?; let mut staged_degradations = focused_batches.degradations.clone(); @@ -1650,52 +1680,35 @@ fn scan_workspace_direct_with_mode( &focused_batches.stored_batches, &context.extraction_budgets, )?); - let (_, staged_degradations) = finalize_scan_degradations(staged_degradations); let mut artifact_execution = artifact_execution_summary(&focused_batches.artifact_durations_ms); artifact_execution.checkpoint_hits = checkpoint_hits; artifact_execution.checkpoints_written = focused_batches.checkpoint_writes; - let candidate = work_state::StagedSnapshot { - snapshot_id, - nodes, - edges, - evidence, - fingerprints, - extractor_batches: focused_batches.stored_batches, - extractor_runs, - manual_links: manual_link_records, - community_snapshot, - community_delta_count, - corroborated_symbol_count: corroboration.confirmed_symbols, - affected_test_count: corroboration.affected_tests, - execution: artifact_execution, - degradations: staged_degradations, - coverage_gaps, - }; - work_state - .store_candidate_snapshot( - &context.manifest.name, - &candidate_fingerprint, - &candidate, - current_unix_millis(), - ) - .map_err(ApplicationError::Initialization)?; - publish_snapshot_candidate( + artifact_execution.stat_cache_hits = stat_cache_hits; + artifact_execution.extraction_workers = extraction_workers; + let artifact_delta = (!overrides.force) + .then(|| PlannedArtifactDelta::new(&fingerprints, &plan, &focused_batches)) + .transpose()?; + artifact_execution.published_rows = publish_snapshot_candidate( &mut store, + artifact_delta.as_ref().map(PlannedArtifactDelta::as_delta), SnapshotBatch { workspace: &context.registry.record, - snapshot_id: &candidate.snapshot_id, - nodes: &candidate.nodes, - edges: &candidate.edges, - evidence: &candidate.evidence, - fingerprints: &candidate.fingerprints, - extractor_batches: &candidate.extractor_batches, - extractor_runs: &candidate.extractor_runs, - manual_links: &candidate.manual_links, - community_snapshot: Some(&candidate.community_snapshot), + snapshot_id: &snapshot_id, + nodes: &nodes, + edges: &edges, + evidence: &evidence, + fingerprints: &fingerprints, + extractor_batches: &focused_batches.stored_batches, + extractor_runs: &extractor_runs, + manual_links: &manual_link_records, + community_snapshot: Some(&community_snapshot), }, - &candidate.coverage_gaps, + &coverage_gaps, + &http_links, )?; - let mut degradations = candidate.degradations.clone(); + phases.complete(code_system_graph_core::JobPhase::Publication); + artifact_execution.phases = phases.into_phases(); + let mut degradations = staged_degradations; for item in &corroboration.reports { let Some(capability) = &item.report.capability else { continue; @@ -1709,43 +1722,158 @@ fn scan_workspace_direct_with_mode( } } let (degradation_count, degradations) = finalize_scan_degradations(degradations); - let _ = work_state.complete_candidate(&context.manifest.name); Ok(ScanSummary { - execution: candidate.execution, + execution: artifact_execution, workspace: context.manifest.name, - snapshot_id: candidate.snapshot_id, - node_count: candidate.nodes.len(), - edge_count: candidate.edges.len(), - evidence_count: candidate.evidence.len(), - community_count: candidate.community_snapshot.communities.len(), - community_delta_count: candidate.community_delta_count, - discovered_input_count: candidate.fingerprints.len(), + snapshot_id, + node_count: nodes.len(), + edge_count: edges.len(), + evidence_count: evidence.len(), + community_count: community_snapshot.communities.len(), + community_delta_count, + discovered_input_count: fingerprints.len(), changed_input_count: plan.changed_count(), reused_snapshot: false, - corroborated_symbol_count: candidate.corroborated_symbol_count, - affected_test_count: candidate.affected_test_count, + corroborated_symbol_count: corroboration.confirmed_symbols, + affected_test_count: corroboration.affected_tests, degradation_count, degradations, }) } +fn load_current_fingerprints( + store: &SqliteStore, + workspace: &str, +) -> Result, ApplicationError> { + match store.load_current_artifact_fingerprints(workspace) { + Ok(previous) => Ok(previous), + Err(StoreError::CurrentSnapshotMissing(_)) => Ok(Vec::new()), + Err(error) => Err(error.into()), + } +} + +/// Identity of the manifest, the extraction contract, and every artifact fingerprint, hashed +/// field by field without materializing a canonical string. +/// +/// Every publication stores one batch per focused artifact, so an equal identity also proves +/// that the published batches are reusable under the active contract and budgets. +fn snapshot_identity( + manifest_hash: &str, + budget_fingerprint: &str, + fingerprints: &[ArtifactFingerprint], +) -> String { + fn update_field(hasher: &mut blake3::Hasher, bytes: &[u8]) { + hasher.update(&u64::try_from(bytes.len()).unwrap_or(u64::MAX).to_le_bytes()); + hasher.update(bytes); + } + let mut hasher = blake3::Hasher::new(); + update_field(&mut hasher, b"community-engine-v1"); + update_field(&mut hasher, EXTRACTION_CONTRACT_VERSION.as_bytes()); + update_field(&mut hasher, budget_fingerprint.as_bytes()); + update_field(&mut hasher, manifest_hash.as_bytes()); + for fingerprint in fingerprints { + let encoding: &[u8] = match fingerprint.path.encoding { + NativePathEncoding::UnixBytes => b"unix", + NativePathEncoding::WindowsWide => b"wide", + NativePathEncoding::Utf8 => b"utf8", + }; + update_field(&mut hasher, fingerprint.repo_id.as_str().as_bytes()); + update_field(&mut hasher, fingerprint.checkout_id.as_str().as_bytes()); + update_field(&mut hasher, encoding); + update_field(&mut hasher, &fingerprint.path.bytes); + update_field(&mut hasher, fingerprint.extractor.as_bytes()); + update_field(&mut hasher, fingerprint.content_hash.as_bytes()); + hasher.update(&fingerprint.size_bytes.to_le_bytes()); + } + format!("snapshot:{}", hasher.finalize().to_hex()) +} + +/// Artifact rows an incremental scan changed, planned against the stored fingerprints. +struct PlannedArtifactDelta<'a> { + upserted_fingerprints: Vec<&'a ArtifactFingerprint>, + upserted_batches: Vec<&'a StoredExtractorBatch>, + removed: Vec, +} + +impl<'a> PlannedArtifactDelta<'a> { + /// `fingerprints` must be sorted by artifact key, as discovery and partial scans leave them. + fn new( + fingerprints: &'a [ArtifactFingerprint], + plan: &IncrementalPlan, + focused: &'a FocusedBatchState, + ) -> Result { + let mut upserted_fingerprints = Vec::new(); + let mut removed = Vec::new(); + for change in &plan.changes { + if change.kind == ArtifactChangeKind::Deleted { + removed.push(change.clone()); + continue; + } + let key = ( + &change.repo_id, + &change.checkout_id, + &change.path, + change.extractor.as_str(), + ); + let index = fingerprints + .binary_search_by(|fingerprint| { + focused_extraction::artifact_key(fingerprint).cmp(&key) + }) + .map_err(|_| { + ApplicationError::Initialization(format!( + "changed artifact `{}` is missing from the current fingerprints", + change.path.display + )) + })?; + upserted_fingerprints.push(&fingerprints[index]); + } + Ok(Self { + upserted_fingerprints, + upserted_batches: focused + .unpublished_batches + .iter() + .map(|index| &focused.stored_batches[*index]) + .collect(), + removed, + }) + } + + fn as_delta(&self) -> ArtifactDelta<'_> { + ArtifactDelta { + upserted_fingerprints: &self.upserted_fingerprints, + upserted_batches: &self.upserted_batches, + removed: &self.removed, + } + } +} + fn publish_snapshot_candidate( store: &mut SqliteStore, + artifact_delta: Option>, batch: SnapshotBatch<'_>, coverage_gaps: &[RepositoryCoverageGap], -) -> Result<(), StoreError> { + http_links: &HttpLinkReport, +) -> Result { const REPORT_INTERVAL: u64 = 1_024; let mut pending = 0_u64; + let mut total = 0_u64; worker::report_progress(code_system_graph_core::JobPhase::Publication, 1); - store.publish_snapshot_with_progress_and_coverage(batch, coverage_gaps, |rows| { - worker::check_time(code_system_graph_core::JobPhase::Publication); - pending = pending.saturating_add(rows); - if pending >= REPORT_INTERVAL { - worker::report_progress(code_system_graph_core::JobPhase::Publication, pending); - pending = 0; - } - })?; - Ok(()) + store.publish_snapshot_with_progress_and_coverage( + batch, + artifact_delta, + coverage_gaps, + http_links, + |rows| { + worker::check_time(code_system_graph_core::JobPhase::Publication); + pending = pending.saturating_add(rows); + total = total.saturating_add(rows); + if pending >= REPORT_INTERVAL { + worker::report_progress(code_system_graph_core::JobPhase::Publication, pending); + pending = 0; + } + }, + )?; + Ok(total) } /// Traces the current stored graph and returns a versioned conservative envelope. @@ -1765,27 +1893,29 @@ pub fn trace_workspace( }); } let store = SqliteStore::open_read_only(database_path)?; - let snapshot = store.current_snapshot_summary(workspace)?; - let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; - let freshness = freshness_summary(&store.load_freshness_snapshot(&snapshot.snapshot_id)?); - let graph = FederatedGraph::new(nodes, edges)?; - let report = graph.trace( - &NodeId::new(&input.from), - &NodeId::new(&input.to), - input.max_depth, - )?; - let status = if report.coverage_gaps.is_empty() && freshness.overall == OverallFreshness::Fresh - { - ToolStatus::Ok - } else { - ToolStatus::Degraded - }; - Ok(ToolEnvelope { - schema_version: 2, - status, - data: Some(report), - freshness, - warnings: Vec::new(), + store.consistent_read(|| { + let snapshot = store.current_snapshot_summary(workspace)?; + let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let freshness = freshness_summary(&store.load_freshness_snapshot(&snapshot.snapshot_id)?); + let graph = FederatedGraph::new(nodes, edges)?; + let report = graph.trace( + &NodeId::new(&input.from), + &NodeId::new(&input.to), + input.max_depth, + )?; + let status = + if report.coverage_gaps.is_empty() && freshness.overall == OverallFreshness::Fresh { + ToolStatus::Ok + } else { + ToolStatus::Degraded + }; + Ok(ToolEnvelope { + schema_version: 2, + status, + data: Some(report), + freshness, + warnings: Vec::new(), + }) }) } @@ -1844,101 +1974,132 @@ pub(crate) fn search_workspace_for_delivery( action_capabilities: QueryActionCapabilities, ) -> Result, ApplicationError> { let store = SqliteStore::open_read_only(database_path)?; - let snapshot = store.current_snapshot_summary(workspace)?; - let input_fingerprint = query_cache_fingerprint(input, policy, action_capabilities)?; - let now_unix_ms = current_unix_millis(); - if let Some(cached) = store.load_query_cache( - workspace, - &snapshot.snapshot_id, - &input_fingerprint, - now_unix_ms, - )? && let Ok(envelope) = - serde_json::from_slice::>(&cached.result_summary_json) - { - return Ok(envelope); - } - let (mut nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; - let registry = store.load_workspace_registry(workspace)?; - apply_repository_alias_labels(&mut nodes, ®istry); - let repository_freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; - let freshness = freshness_summary(&repository_freshness); - let fts_hits = if input.query.trim().is_empty() { - Vec::new() - } else { - store.search_snapshot_nodes_ranked(&snapshot.snapshot_id, &input.query, 500)? - }; - let community_snapshot = store.load_community_snapshot(&snapshot.snapshot_id)?; - let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; - let request = SearchRequest { - query: input.query.clone(), - filters: SearchFilters { - node_kinds: input.node_kinds.clone(), - repo_ids: input.repo_ids.clone(), - workspace_nodes: Vec::new(), - service_ids: input.service_ids.clone(), - community_ids: input.community_ids.clone(), - }, - fts_scores: fts_search_scores(&fts_hits), - centrality_scores: community_centrality_scores(&community_snapshot.communities), - service_memberships: service_memberships(&nodes, &edges), - community_memberships: community_memberships(&community_snapshot.communities), - evidence: node_evidence(&edges, &evidence), - freshness: repository_freshness - .iter() - .map(|item| (item.repo_id.clone(), item.state)) - .collect(), - offset: input.offset, - limit: input.limit, - }; - let mut report = search(&nodes, &request)?; - report.next_actions = query_next_actions( - workspace, - input, - &report, - ®istry, - policy, - action_capabilities, - ); - if report.hits.is_empty() { - report.coverage.gaps.push(if action_capabilities.explore { - "Query searches persisted architecture entities and contracts, not source-code bodies; use Explore for implementation text." - .to_owned() + let (envelope, cache_record) = store.consistent_read(|| { + let snapshot = store.current_snapshot_summary(workspace)?; + let input_fingerprint = query_cache_fingerprint(input, policy, action_capabilities)?; + let now_unix_ms = current_unix_millis(); + if let Some(cached) = store.load_query_cache( + workspace, + &snapshot.snapshot_id, + &input_fingerprint, + now_unix_ms, + )? && let Ok(envelope) = + serde_json::from_slice::>(&cached.result_summary_json) + { + return Ok::<_, ApplicationError>((envelope, None)); + } + let (mut nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let registry = store.load_workspace_registry(workspace)?; + apply_repository_alias_labels(&mut nodes, ®istry); + let repository_freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; + let freshness = freshness_summary(&repository_freshness); + let fts_hits = if input.query.trim().is_empty() { + Vec::new() } else { - "Query searches persisted architecture entities and contracts, not source-code bodies; source exploration is unavailable in this delivery profile." - .to_owned() - }); + store.search_snapshot_nodes_ranked(&snapshot.snapshot_id, &input.query, 500)? + }; + let community_snapshot = store.load_community_snapshot(&snapshot.snapshot_id)?; + let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; + let request = SearchRequest { + query: input.query.clone(), + filters: SearchFilters { + node_kinds: input.node_kinds.clone(), + repo_ids: input.repo_ids.clone(), + workspace_nodes: Vec::new(), + service_ids: input.service_ids.clone(), + community_ids: input.community_ids.clone(), + }, + fts_scores: fts_search_scores(&fts_hits), + centrality_scores: community_centrality_scores(&community_snapshot.communities), + service_memberships: service_memberships(&nodes, &edges), + community_memberships: community_memberships(&community_snapshot.communities), + evidence: node_evidence(&edges, &evidence), + freshness: repository_freshness + .iter() + .map(|item| (item.repo_id.clone(), item.state)) + .collect(), + offset: input.offset, + limit: input.limit, + }; + let mut report = search(&nodes, &request)?; + let hit_ids = report + .hits + .iter() + .map(|hit| hit.node.id.clone()) + .collect::>(); + report.link_gaps = store.load_http_link_gaps_for_callers(workspace, &hit_ids)?; + report.next_actions = query_next_actions( + workspace, + input, + &report, + ®istry, + policy, + action_capabilities, + ); + if report.hits.is_empty() { + report + .coverage + .gaps + .push(zero_hit_gap(action_capabilities).to_owned()); + } + let status = + if report.coverage.gaps.is_empty() && freshness.overall == OverallFreshness::Fresh { + ToolStatus::Ok + } else { + ToolStatus::Degraded + }; + let envelope = ToolEnvelope { + schema_version: 2, + status, + data: Some(report), + freshness, + warnings: Vec::new(), + }; + let (envelope, encoded) = encode_query_cache(&envelope)?; + Ok(( + envelope, + Some(QueryCacheRecord { + workspace_name: workspace.to_owned(), + snapshot_id: snapshot.snapshot_id, + input_fingerprint, + result_summary_json: encoded, + stored_at_unix_ms: now_unix_ms, + expires_at_unix_ms: None, + }), + )) + })?; + drop(store); + if let Some(record) = cache_record { + store_query_cache(database_path, &record); } - let status = if report.coverage.gaps.is_empty() && freshness.overall == OverallFreshness::Fresh - { - ToolStatus::Ok - } else { - ToolStatus::Degraded - }; - let envelope = ToolEnvelope { - schema_version: 2, - status, - data: Some(report), - freshness, - warnings: Vec::new(), - }; - let encoded = serde_json::to_vec(&envelope) + Ok(envelope) +} + +/// Keeps freshly computed and cached query results identical after JSON normalization. +fn encode_query_cache( + envelope: &ToolEnvelope, +) -> Result<(ToolEnvelope, Vec), ApplicationError> { + let encoded = serde_json::to_vec(envelope) .map_err(|error| ApplicationError::Initialization(error.to_string()))?; let envelope = serde_json::from_slice(&encoded) .map_err(|error| ApplicationError::Initialization(error.to_string()))?; - drop(store); + Ok((envelope, encoded)) +} + +fn zero_hit_gap(capabilities: QueryActionCapabilities) -> &'static str { + if capabilities.explore { + "Query searches persisted architecture entities and contracts, not source-code bodies; use Explore for implementation text." + } else { + "Query searches persisted architecture entities and contracts, not source-code bodies; source exploration is unavailable in this delivery profile." + } +} + +fn store_query_cache(database_path: &Path, record: &QueryCacheRecord) { if let Ok(_lock) = StoreLock::acquire(database_path, Duration::from_secs(5)) && let Ok(mut writable) = SqliteStore::open(database_path) { - let _ = writable.put_query_cache(&QueryCacheRecord { - workspace_name: workspace.to_owned(), - snapshot_id: snapshot.snapshot_id, - input_fingerprint, - result_summary_json: encoded, - stored_at_unix_ms: now_unix_ms, - expires_at_unix_ms: None, - }); + let _ = writable.put_query_cache(record); } - Ok(envelope) } fn query_cache_fingerprint( @@ -2396,32 +2557,34 @@ pub async fn analyze_workspace_changes_with_cancellation( ) })?; let store = SqliteStore::open_read_only(database_path)?; - let snapshot = store.current_snapshot_summary(workspace)?; - let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; - let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; - let community = store.load_community_snapshot(&snapshot.snapshot_id)?; - let freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; - let report = analyze_changes( - &change_set, - &nodes, - &edges, - &evidence, - Some(&community), - &freshness, - &[], - options, - )?; - let status = if collected.status == ToolStatus::Ok && report.coverage.complete { - ToolStatus::Ok - } else { - ToolStatus::Degraded - }; - Ok(ToolEnvelope { - schema_version: 2, - status, - data: Some(report), - freshness: collected.freshness, - warnings: collected.warnings, + store.consistent_read(|| { + let snapshot = store.current_snapshot_summary(workspace)?; + let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; + let community = store.load_community_snapshot(&snapshot.snapshot_id)?; + let freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; + let report = analyze_changes( + &change_set, + &nodes, + &edges, + &evidence, + Some(&community), + &freshness, + &[], + options, + )?; + let status = if collected.status == ToolStatus::Ok && report.coverage.complete { + ToolStatus::Ok + } else { + ToolStatus::Degraded + }; + Ok(ToolEnvelope { + schema_version: 2, + status, + data: Some(report), + freshness: collected.freshness, + warnings: collected.warnings, + }) }) } @@ -2598,35 +2761,37 @@ fn load_impact_context( workspace: &str, ) -> Result { let store = SqliteStore::open_read_only(database_path)?; - let snapshot = store.current_snapshot_summary(workspace)?; - let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; - let persisted_freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; - let freshness = freshness_summary(&persisted_freshness); - let communities = store.load_community_snapshot(&snapshot.snapshot_id)?; - let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; - let registry = store.load_workspace_registry(workspace)?; - let node_evidence = node_evidence(&edges, &evidence); - let context = ImpactContext { - centrality: impact_centrality_scores(&communities.communities), - service_memberships: service_memberships(&nodes, &edges), - recommended_commands: recommended_test_commands(&nodes, &node_evidence), - nodes, - edges, - communities: Some(communities), - freshness: conservative_repo_freshness(&persisted_freshness), - compatibility: Vec::new(), - local_enrichment: Vec::new(), - public_contracts: Vec::new(), - criticality: Vec::new(), - environments: Vec::new(), - graph_complete: true, - coverage_gaps: Vec::new(), - }; - Ok(LoadedImpactContext { - context, - freshness, - registry, - node_evidence, + store.consistent_read(|| { + let snapshot = store.current_snapshot_summary(workspace)?; + let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let persisted_freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; + let freshness = freshness_summary(&persisted_freshness); + let communities = store.load_community_snapshot(&snapshot.snapshot_id)?; + let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; + let registry = store.load_workspace_registry(workspace)?; + let node_evidence = node_evidence(&edges, &evidence); + let context = ImpactContext { + centrality: impact_centrality_scores(&communities.communities), + service_memberships: service_memberships(&nodes, &edges), + recommended_commands: recommended_test_commands(&nodes, &node_evidence), + nodes, + edges, + communities: Some(communities), + freshness: conservative_repo_freshness(&persisted_freshness), + compatibility: Vec::new(), + local_enrichment: Vec::new(), + public_contracts: Vec::new(), + criticality: Vec::new(), + environments: Vec::new(), + graph_complete: true, + coverage_gaps: Vec::new(), + }; + Ok(LoadedImpactContext { + context, + freshness, + registry, + node_evidence, + }) }) } @@ -3103,13 +3268,20 @@ pub fn status_workspace( current.execution_policy.max_no_progress_time_ms, &store.database_instance_id()?, ); + let (http_report, gap_count) = + store.load_http_link_report(¤t.manifest.name, STATUS_HTTP_LINK_GAP_LIMIT)?; Ok(WorkspaceStatus { workspace: current.manifest.name, - schema_version: store.schema_version()?, + schema_id: store.schema_id()?, integrity_ok: store.integrity_check()?, freshness: freshness_summary(&repositories), repositories, watcher, + http_links: HttpLinkStatus { + coverage: http_report.coverage, + gap_count, + gaps: http_report.gaps, + }, }) } @@ -3303,10 +3475,8 @@ pub fn doctor_workspace( let store = SqliteStore::open_read_only(database_path)?; let diagnostics = store.diagnostics()?; let context = load_workspace_context(config_path, &ScanOverrides::default())?; - let expected_version = u32::try_from(latest_schema_version()).map_err(|error| { - ApplicationError::Initialization(format!("unsupported schema version: {error}")) - })?; - let actual_version = u32::try_from(status.schema_version).ok(); + let expected_schema = schema_identity().to_owned(); + let actual_schema = Some(status.schema_id.clone()); let freshness = status .repositories .iter() @@ -3337,9 +3507,9 @@ pub fn doctor_workspace( Ok(doctor(&DoctorRequest { schema: vec![SchemaDoctorInput { name: "sqlite".to_owned(), - expected_version, - actual_version, - metadata_consistent: Some(actual_version == Some(expected_version)), + metadata_consistent: Some(actual_schema.as_deref() == Some(expected_schema.as_str())), + expected_schema, + actual_schema, }], integrity: vec![ IntegrityDoctorInput { @@ -3519,7 +3689,7 @@ pub fn restore_database( safety_backup: report .safety_backup_path .map(|path| path.to_string_lossy().into_owned()), - schema_version: report.schema_version, + schema_id: report.schema_id, }) } @@ -3873,549 +4043,50 @@ fn validate_global_policy_source( if canonical_config.starts_with(&checkout) { return Err(ApplicationError::UntrustedGlobalPolicySource { config: canonical_config, - repository: alias.clone(), - }); - } - } - Ok(()) -} - -fn focused_batch_cache_complete( - fingerprints: &[ArtifactFingerprint], - stored: &[StoredExtractorBatch], - budgets: &ExtractionBudgets, -) -> bool { - let budget_fingerprint = budgets.fingerprint(); - fingerprints - .iter() - .filter(|fingerprint| focused_extractor(&fingerprint.extractor)) - .all(|fingerprint| { - stored.iter().any(|batch| { - batch.source == *fingerprint - && batch.extractor_version == EXTRACTION_CONTRACT_VERSION - && batch.budget_fingerprint == budget_fingerprint - }) - }) -} - -fn stored_batch_degradations( - stored: &[StoredExtractorBatch], - budgets: &ExtractionBudgets, -) -> Result, ApplicationError> { - let mut degradations = Vec::new(); - for batch in stored { - if batch.source_was_lossy { - degradations.push(format!("{} contains invalid UTF-8 and was decoded lossily; extracted evidence is incomplete", batch.source.path.display)); - } - if batch.source.extractor == "code-system-graph.data.artifact" { - let decoded: ExtractorBatch = - load_extractor_batch_with_budgets(batch, budgets)?; - for document in decoded.outputs { - if document.incomplete { - degradations.push(format!( - "{} data extraction is incomplete: {:?}", - batch.source.path.display, document.warnings - )); - } - } - } - } - Ok(degradations) -} - -fn finalize_scan_degradations(mut degradations: Vec) -> (usize, Vec) { - degradations.sort(); - degradations.dedup(); - let count = degradations.len(); - if count > MAX_SCAN_DEGRADATIONS { - degradations.truncate(MAX_SCAN_DEGRADATIONS - 1); - degradations.push(format!( - "{} additional degradations omitted", - count - (MAX_SCAN_DEGRADATIONS - 1) - )); - } - (count, degradations) -} - -#[expect( - clippy::too_many_lines, - reason = "Batch decode, reuse, extraction, and persistence share one fail-closed type boundary" -)] -fn assemble_focused_batches( - context: &WorkspaceContext, - fingerprints: &[ArtifactFingerprint], - previous: &[StoredExtractorBatch], - plan: &ExtractorBatchPlan, - cached_keys: &BTreeSet, - work_state: &mut work_state::WorkState, -) -> Result { - let previous_by_key = previous - .iter() - .map(|batch| (ArtifactKey::from(&batch.source), batch)) - .collect::>(); - let planned_actions = plan - .batches - .iter() - .map(|batch| (batch.key.clone(), batch.action)) - .collect::>(); - let mut previous_source_batches = previous - .iter() - .filter(|batch| source_extractor(&batch.source.extractor)) - .map(|batch| load_extractor_batch_with_budgets(batch, &context.extraction_budgets)) - .collect::>, _>>()?; - previous_source_batches.sort_by_key(ExtractorBatch::key); - - let checkouts = context - .registry - .record - .repositories - .iter() - .filter_map(|repository| { - context - .registry - .checkout_path(&repository.alias) - .map(|path| (repository.id.clone(), path)) - }) - .collect::>(); - let mut source_batches = Vec::new(); - let mut package_batches = Vec::new(); - let mut generated_client_batches = Vec::new(); - let mut graphql_batches = Vec::new(); - let mut event_batches = Vec::new(); - let mut protobuf_batches = Vec::new(); - let mut data_batches = Vec::new(); - let mut infrastructure_batches = Vec::new(); - let mut documentation_batches = Vec::new(); - let mut config_batches = Vec::new(); - let mut stored_batches = Vec::new(); - let mut degradations = Vec::new(); - let mut force_relink = false; - let mut checkpoint_writes = 0_u64; - let mut artifact_durations_ms = Vec::new(); - for fingerprint in fingerprints - .iter() - .filter(|fingerprint| focused_extractor(&fingerprint.extractor)) - { - let artifact_started = Instant::now(); - let key = ArtifactKey::from(fingerprint); - let reusable = previous_by_key.get(&key).copied().filter(|batch| { - (planned_actions.get(&key) == Some(&BatchAction::Reuse) || cached_keys.contains(&key)) - && batch.source.content_hash == fingerprint.content_hash - && batch.extractor_version == EXTRACTION_CONTRACT_VERSION - && batch.budget_fingerprint == context.extraction_budgets.fingerprint() - }); - if let Some(stored) = reusable { - stored_batches.push(stored.clone()); - if source_extractor(&fingerprint.extractor) { - source_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if fingerprint.extractor == "code-system-graph.packages" { - package_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if graphql_extractor(&fingerprint.extractor) { - graphql_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if event_extractor(&fingerprint.extractor) { - event_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if protobuf_extractor(&fingerprint.extractor) { - protobuf_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if data_extractor(&fingerprint.extractor) { - data_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if infrastructure_extractor(&fingerprint.extractor) { - infrastructure_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if documentation_extractor(&fingerprint.extractor) { - documentation_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else if fingerprint.extractor == "code-system-graph.config.safe" { - config_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } else { - generated_client_batches.push(load_extractor_batch_with_budgets( - stored, - &context.extraction_budgets, - )?); - } - worker::report_progress(code_system_graph_core::JobPhase::Extraction, 1); - artifact_durations_ms.push(duration_millis(artifact_started.elapsed())); - continue; - } - if previous_by_key.contains_key(&key) { - force_relink = true; - } - let checkout = checkouts.get(&fingerprint.repo_id).ok_or_else(|| { - ApplicationError::RegistryAliasMissing(fingerprint.repo_id.as_str().to_owned()) - })?; - let relative_path = native_relative_path(&fingerprint.path); - let artifact_path = checkout.join(&relative_path); - let mut tracker = ExtractionTracker::new( - &fingerprint.path.display, - &fingerprint.extractor, - &context.extraction_budgets, - ); - let (source, source_was_lossy) = read_source_file(&artifact_path, &mut tracker)?; - if source_was_lossy { - degradations.push(format!("{} contains invalid UTF-8 and was decoded lossily; extracted evidence is incomplete", fingerprint.path.display)); - } - let mut persist = |stored: StoredExtractorBatch| -> Result<(), ApplicationError> { - if work_state - .put_batch( - &stored, - context.execution_policy.max_checkpoint_cache_bytes, - current_unix_millis(), - ) - .map_err(ApplicationError::Initialization)? - { - checkpoint_writes = checkpoint_writes.checked_add(1).ok_or_else(|| { - ApplicationError::Initialization( - "checkpoint write counter overflowed".to_owned(), - ) - })?; - } - stored_batches.push(stored); - Ok(()) - }; - if source_extractor(&fingerprint.extractor) { - let syntax = inspect_source_syntax( - source_syntax_language(&fingerprint.extractor), - &portable_path(&fingerprint.path.display), - &source, - &mut tracker, - )?; - let reserved_observations = - u64::try_from(syntax.boundary_candidate_count).map_err(|_| { - ApplicationError::InvalidSourceObservation(format!( - "{} contains too many syntax candidates", - fingerprint.path.display - )) - })?; - tracker.charge_work(reserved_observations)?; - precheck_focused_source_values( - &source, - source_syntax_language(&fingerprint.extractor), - &mut tracker, - )?; - let mut observations = match fingerprint.extractor.as_str() { - "code-system-graph.source.javascript" => { - parse_javascript_source_at_path_with_tracker( - &portable_path(&fingerprint.path.display), - &source, - &mut tracker, - )? - } - "code-system-graph.source.typescript" => { - parse_typescript_source_at_path_with_tracker( - &portable_path(&fingerprint.path.display), - &source, - &mut tracker, - )? - } - "code-system-graph.source.rust" => { - parse_rust_source_with_tracker(&source, &mut tracker)? - } - "code-system-graph.source.python" => { - parse_python_source_with_tracker(&source, &mut tracker)? - } - "code-system-graph.source.go" => { - parse_go_source_with_tracker(&source, &mut tracker)? - } - "code-system-graph.source.java" => { - parse_java_source_with_tracker(&source, &mut tracker)? - } - _ => Vec::new(), - }; - if observations - .iter() - .any(|observation| observation.role != SourceRole::Test) - && syntax.boundary_candidate_count == 0 - { - return Err(ApplicationError::InvalidSourceObservation(format!( - "{} produced framework facts without a Tree-sitter boundary candidate", - fingerprint.path.display - ))); - } - if syntax.has_error { - for observation in &mut observations { - observation.status = SourceEpistemicStatus::Incomplete; - if !observation - .warnings - .contains(&SourceWarning::SyntaxErrorRecovery) - { - observation - .warnings - .push(SourceWarning::SyntaxErrorRecovery); - } - } - } - let unmapped_authorities = observations - .iter() - .filter(|observation| { - observation - .warnings - .contains(&SourceWarning::UnmappedAuthority) - }) - .count(); - if unmapped_authorities > 0 { - degradations.push(format!( - "{} contains {} absolute HTTP consumer URL{} without an explicit workspace authority mapping; those dependencies were not linked", - fingerprint.path.display, - unmapped_authorities, - if unmapped_authorities == 1 { "" } else { "s" } - )); - } - let batch = ExtractorBatch::new(fingerprint.clone(), observations); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - source_batches.push(batch); - } else if fingerprint.extractor == "code-system-graph.packages" { - let portable_path = portable_path(&fingerprint.path.display); - let manifest = - extract_package_manifest_with_tracker(&portable_path, &source, &mut tracker)?; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![manifest]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - package_batches.push(batch); - } else if graphql_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = match fingerprint.extractor.as_str() { - "code-system-graph.graphql.document" => { - extract_graphql_document_with_tracker(&portable_path, &source, &mut tracker)? - } - "code-system-graph.graphql.persisted" => GraphqlDocument { - source_path: portable_path.clone(), - types: Vec::new(), - operations: Vec::new(), - fragments: Vec::new(), - persisted_operations: extract_graphql_persisted_operations_with_tracker( - &portable_path, - &source, - &mut tracker, - )?, - resolvers: Vec::new(), - federation: Vec::new(), - complete: true, - warnings: Vec::new(), - }, - "code-system-graph.graphql.source" => { - let language = source_language_for_path(&relative_path).ok_or_else(|| { - ApplicationError::InvalidSourceObservation(format!( - "unsupported GraphQL source language for `{portable_path}`" - )) - })?; - let mut document = - parse_graphql_source_with_tracker(language, &source, &mut tracker)?; - document.source_path = portable_path; - document - } - _ => unreachable!("graphql extractor classification must be exhaustive"), - }; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - graphql_batches.push(batch); - } else if event_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = if fingerprint.extractor == "code-system-graph.events.asyncapi" { - extract_asyncapi(&portable_path, &source)? - } else { - let language = source_language_for_path(&relative_path).ok_or_else(|| { - ApplicationError::InvalidSourceObservation(format!( - "unsupported event source language for `{portable_path}`" - )) - })?; - let mut document = parse_event_source(language, &source); - document.source_path = Some(portable_path); - document - }; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - event_batches.push(batch); - } else if protobuf_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = if fingerprint.extractor == "code-system-graph.protobuf" { - ProtobufDocument::File(Box::new(extract_protobuf_with_tracker( - &portable_path, - &source, - &mut tracker, - )?)) - } else { - let language = source_language_for_path(&relative_path).ok_or_else(|| { - ApplicationError::InvalidSourceObservation(format!( - "unsupported generated protobuf source language for `{portable_path}`" - )) - })?; - ProtobufDocument::Generated(parse_protobuf_generated_source( - language, - &portable_path, - &source, - )) - }; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - protobuf_batches.push(batch); - } else if data_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = if fingerprint.extractor == "code-system-graph.data.source" { - let language = source_language_for_path(&relative_path).ok_or_else(|| { - ApplicationError::InvalidSourceObservation(format!( - "unsupported data source language for `{portable_path}`" - )) - })?; - let crate_root = cargo_crate_root(checkout, &artifact_path); - parse_literal_sql_source_at_root(language, &portable_path, &crate_root, &source) - } else { - extract_data_artifact(&portable_path, &source)? - }; - if fingerprint.extractor == "code-system-graph.data.artifact" && document.incomplete { - degradations.push(format!( - "{} data extraction is incomplete: {:?}", - fingerprint.path.display, document.warnings - )); - } - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - data_batches.push(batch); - } else if infrastructure_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = match fingerprint.extractor.as_str() { - "code-system-graph.infrastructure.compose" => { - extract_docker_compose(&portable_path, &source)? - } - "code-system-graph.infrastructure.kubernetes" => { - extract_kubernetes(&portable_path, &source)? - } - "code-system-graph.infrastructure.helm" => extract_helm(&portable_path, &source)?, - "code-system-graph.infrastructure.terraform" => { - extract_terraform(&portable_path, &source)? - } - _ => unreachable!("infrastructure extractor classification must be exhaustive"), - }; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - infrastructure_batches.push(batch); - } else if documentation_extractor(&fingerprint.extractor) { - let portable_path = portable_path(&fingerprint.path.display); - let document = match fingerprint.extractor.as_str() { - "code-system-graph.documents.markdown" => { - extract_markdown(&portable_path, &source)? - } - "code-system-graph.documents.codeowners" => { - extract_codeowners(&portable_path, &source)? - } - "code-system-graph.documents.catalog" => { - extract_service_catalog(&portable_path, &source)? + repository: alias.clone(), + }); + } + } + Ok(()) +} + +fn stored_batch_degradations( + stored: &[StoredExtractorBatch], + budgets: &ExtractionBudgets, +) -> Result, ApplicationError> { + let mut degradations = Vec::new(); + for batch in stored { + if batch.source_was_lossy { + degradations.push(format!("{} contains invalid UTF-8 and was decoded lossily; extracted evidence is incomplete", batch.source.path.display)); + } + if batch.source.extractor == DATA_ARTIFACT_EXTRACTOR { + let decoded: ExtractorBatch = + load_extractor_batch_with_budgets(batch, budgets)?; + for document in decoded.outputs { + if document.incomplete { + degradations.push(format!( + "{} data extraction is incomplete: {:?}", + batch.source.path.display, document.warnings + )); } - _ => unreachable!("documentation extractor classification must be exhaustive"), - }; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - documentation_batches.push(batch); - } else if fingerprint.extractor == "code-system-graph.config.safe" { - let portable_path = portable_path(&fingerprint.path.display); - let document = extract_safe_config(&portable_path, &source)?; - let batch = ExtractorBatch::new(fingerprint.clone(), vec![document]); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - config_batches.push(batch); - } else { - let portable_path = portable_path(&fingerprint.path.display); - let metadata = - extract_generated_client_metadata(&portable_path, &source, &mut tracker)?; - let batch = ExtractorBatch::new(fingerprint.clone(), metadata); - persist(store_extractor_batch( - &batch, - &mut tracker, - source_was_lossy, - )?)?; - generated_client_batches.push(batch); + } } - worker::report_progress(code_system_graph_core::JobPhase::Extraction, 1); - artifact_durations_ms.push(duration_millis(artifact_started.elapsed())); } - source_batches.sort_by_key(ExtractorBatch::key); - package_batches.sort_by_key(ExtractorBatch::key); - generated_client_batches.sort_by_key(ExtractorBatch::key); - graphql_batches.sort_by_key(ExtractorBatch::key); - event_batches.sort_by_key(ExtractorBatch::key); - protobuf_batches.sort_by_key(ExtractorBatch::key); - data_batches.sort_by_key(ExtractorBatch::key); - infrastructure_batches.sort_by_key(ExtractorBatch::key); - documentation_batches.sort_by_key(ExtractorBatch::key); - config_batches.sort_by_key(ExtractorBatch::key); - stored_batches.sort_by(|left, right| { - ArtifactKey::from(&left.source).cmp(&ArtifactKey::from(&right.source)) - }); - Ok(FocusedBatchState { - source_batches, - previous_source_batches, - package_batches, - generated_client_batches, - graphql_batches, - event_batches, - protobuf_batches, - data_batches, - infrastructure_batches, - documentation_batches, - config_batches, - stored_batches, - degradations, - force_relink, - checkpoint_writes, - artifact_durations_ms, - }) + Ok(degradations) +} + +fn finalize_scan_degradations(mut degradations: Vec) -> (usize, Vec) { + degradations.sort(); + degradations.dedup(); + let count = degradations.len(); + if count > MAX_SCAN_DEGRADATIONS { + degradations.truncate(MAX_SCAN_DEGRADATIONS - 1); + degradations.push(format!( + "{} additional degradations omitted", + count - (MAX_SCAN_DEGRADATIONS - 1) + )); + } + (count, degradations) } fn focused_extractor(extractor: &str) -> bool { @@ -4449,6 +4120,25 @@ fn event_extractor(extractor: &str) -> bool { ) } +/// Whether `path` is a YAML or `.tpl` file in the `templates` directory of a Helm chart. +fn is_helm_template(checkout: &Path, path: &Path, extension: Option<&str>) -> bool { + if !matches!(extension, Some("yaml" | "yml" | "tpl")) { + return false; + } + path.ancestors() + .skip(1) + .find(|ancestor| { + ancestor + .file_name() + .is_some_and(|name| name.eq_ignore_ascii_case("templates")) + }) + .and_then(Path::parent) + .is_some_and(|chart| { + let chart = checkout.join(chart); + chart.join("Chart.yaml").is_file() || chart.join("Chart.yml").is_file() + }) +} + fn protobuf_extractor(extractor: &str) -> bool { matches!( extractor, @@ -4623,6 +4313,31 @@ fn generated_client_graph( (nodes, edges, evidence) } +/// Maps call authorities to repositories: manifest `authorities` strictly, and service names from +/// Compose and Kubernetes declarations as inferred hints for the declaring repository. +fn workspace_authorities( + context: &WorkspaceContext, + infrastructure: &[(&RepoId, &str, &str, &InfrastructureDocument)], +) -> AuthorityMap { + let mut authorities = AuthorityMap::new(); + for (repo_id, _, _, document) in infrastructure { + authorities.infer_from_infrastructure(repo_id, document); + } + for repository in &context.registry.record.repositories { + let declared = context + .manifest + .repos + .get(&repository.alias) + .and_then(|config| config.authorities.as_ref()); + for authority in declared.into_iter().flatten() { + if let Some(authority) = normalize_authority(authority) { + authorities.declare(authority, repository.id.clone()); + } + } + } + authorities +} + #[expect( clippy::too_many_lines, reason = "Graph assembly keeps every source family in one deterministic merge boundary" @@ -4635,15 +4350,42 @@ fn assemble_graph( let mut boundaries = extract_boundaries(context)?; let mut tests = extract_declared_tests(context, fingerprints)?; let mut implementations = extract_declared_implementations(context, fingerprints)?; + let source_paths = focused + .source_batches + .iter() + .map(|batch| portable_path(&batch.source.path.display)) + .collect::>(); + let router_files = focused + .source_batches + .iter() + .zip(&source_paths) + .map(|(batch, path)| RepositorySourceFile { + repo_id: &batch.source.repo_id, + path, + observations: &batch.outputs, + }) + .collect::>(); + let routed_observations = compose_router_mounts(&router_files); + let flow_files = router_files + .iter() + .zip(&routed_observations) + .map(|(file, observations)| RepositorySourceFile { + observations, + ..*file + }) + .collect::>(); + let composed_observations = compose_client_flows(&flow_files); let source_facts = focused .source_batches .iter() - .map(|batch| { + .zip(&source_paths) + .zip(&composed_observations) + .map(|((batch, path), observations)| { source_observations_to_graph( &batch.source.repo_id, - &portable_path(&batch.source.path.display), + path, &batch.source.content_hash, - &batch.outputs, + observations, ) }) .collect::>(); @@ -4868,22 +4610,8 @@ fn assemble_graph( &repository_aliases, ); - let mut http_links = link_http_boundaries_with_ambiguities(&boundaries); - let test_links = link_declared_tests_with_ambiguities(&tests, &boundaries); - let implementation_links = - link_declared_implementations_with_ambiguities(&implementations, &boundaries); - http_links.ambiguities.extend(test_links.ambiguities); - http_links - .ambiguities - .extend(implementation_links.ambiguities); - http_links.ambiguities.sort_by(|left, right| { - (&left.method, &left.path, &left.candidates).cmp(&( - &right.method, - &right.path, - &right.candidates, - )) - }); - http_links.ambiguities.dedup(); + let authorities = workspace_authorities(context, &infrastructure_inputs); + let http_links = link_http_routes(&boundaries, &tests, &implementations, &authorities); let mut degradations = http_links .ambiguities .iter() @@ -4901,9 +4629,8 @@ fn assemble_graph( .iter() .map(|gap| format!("repository {}: {}", gap.repo_id.as_str(), gap.reason)), ); + let http_report = http_links.report; let mut edges = http_links.edges; - edges.extend(test_links.edges); - edges.extend(implementation_links.edges); edges.extend( source_facts .iter() @@ -4975,6 +4702,17 @@ fn assemble_graph( .into_iter() .map(|node| (node.id.clone(), node)), ); + for repository in &context.registry.record.repositories { + let stable_key = format!("repository:{}", repository.id.as_str()); + let id = NodeId::new(stable_id("node", &stable_key)); + nodes.entry(id.clone()).or_insert_with(|| Node { + id, + kind: NodeKind::Repository, + repo_id: Some(repository.id.clone()), + stable_key, + label: repository.id.as_str().to_owned(), + }); + } let mut evidence = boundaries .iter() .map(|boundary| (boundary.evidence.id.clone(), boundary.evidence.clone())) @@ -5033,155 +4771,17 @@ fn assemble_graph( .into_iter() .map(|item| (item.id.clone(), item)), ); - let mut link_node_keys = BTreeMap::new(); - link_node_keys.extend(boundaries.iter().map(|boundary| { - ( - boundary.node.id.clone(), - format!("{}:{}", boundary.method, boundary.path), - ) - })); - link_node_keys.extend(tests.iter().map(|test| { - ( - test.node.id.clone(), - format!("{}:{}", test.method, test.path), - ) - })); - link_node_keys.extend(implementations.iter().map(|implementation| { - ( - implementation.node.id.clone(), - format!("{}:{}", implementation.method, implementation.path), - ) - })); Ok(GraphAssembly { nodes: nodes.into_values().collect(), edges, evidence: evidence.into_values().collect(), link_decisions: Vec::new(), - link_node_keys, degradations, coverage_gaps, + http_links: http_report, }) } -#[expect( - clippy::too_many_lines, - reason = "Fail-closed relinking keeps old/new neighborhood handling in one audit boundary" -)] -fn relink_affected_graph( - graph: &mut GraphAssembly, - plan: &IncrementalPlan, - batch_plan: &ExtractorBatchPlan, - focused: &FocusedBatchState, - _previous_nodes: &[Node], - previous_edges: &[Edge], -) -> Result<(), ApplicationError> { - let full_relink = focused.force_relink - || plan.changes.iter().any(|change| { - change.kind != code_system_graph_model::ArtifactChangeKind::Unchanged - && change.extractor != "code-system-graph.packages" - && !source_extractor(&change.extractor) - }); - if full_relink || previous_edges.is_empty() { - return Ok(()); - } - let source_plan = ExtractorBatchPlan { - batches: batch_plan - .batches - .iter() - .filter(|batch| source_extractor(&batch.key.extractor)) - .cloned() - .collect(), - }; - let affected = affected_link_keys( - &source_plan, - &focused.previous_source_batches, - &focused.source_batches, - |observation| { - observation - .method - .as_ref() - .zip(observation.path.as_ref()) - .map(|(method, path)| format!("{method}:{path}")) - }, - )? - .into_iter() - .flatten() - .collect::>(); - if affected.is_empty() { - return Ok(()); - } - - let mut previous_node_keys = graph.link_node_keys.clone(); - for batch in &focused.previous_source_batches { - let facts = source_observations_to_graph( - &batch.source.repo_id, - &portable_path(&batch.source.path.display), - &batch.source.content_hash, - &batch.outputs, - ); - previous_node_keys.extend(facts.boundaries.iter().map(|boundary| { - ( - boundary.node.id.clone(), - format!("{}:{}", boundary.method, boundary.path), - ) - })); - previous_node_keys.extend(facts.tests.iter().map(|test| { - ( - test.node.id.clone(), - format!("{}:{}", test.method, test.path), - ) - })); - previous_node_keys.extend(facts.implementations.iter().map(|implementation| { - ( - implementation.node.id.clone(), - format!("{}:{}", implementation.method, implementation.path), - ) - })); - } - let current_node_ids = graph - .nodes - .iter() - .map(|node| node.id.clone()) - .collect::>(); - let is_http_link = |edge: &Edge| { - matches!( - edge.kind, - code_system_graph_model::EdgeKind::CallsRemote - | code_system_graph_model::EdgeKind::Validates - | code_system_graph_model::EdgeKind::ImplementedBy - ) - }; - let recomputed_http = graph - .edges - .iter() - .filter(|edge| is_http_link(edge)) - .cloned() - .collect::>(); - let previous_http = previous_edges - .iter() - .filter(|edge| is_http_link(edge)) - .cloned() - .collect::>(); - let mut edges = graph - .edges - .iter() - .filter(|edge| !is_http_link(edge)) - .cloned() - .collect::>(); - edges.extend(merge_affected_link_neighborhoods( - &previous_http, - &recomputed_http, - &affected, - &previous_node_keys, - &graph.link_node_keys, - ¤t_node_ids, - )); - edges.sort_by(|left, right| left.id.cmp(&right.id)); - edges.dedup_by(|left, right| left.id == right.id); - graph.edges = edges; - Ok(()) -} - fn run_codegraph_corroboration( context: &WorkspaceContext, focused: &FocusedBatchState, @@ -5285,11 +4885,7 @@ fn prepare_codegraph_jobs( })); } let mut changed_files_by_repository = BTreeMap::>::new(); - for change in plan - .changes - .iter() - .filter(|change| change.kind != code_system_graph_model::ArtifactChangeKind::Unchanged) - { + for change in &plan.changes { changed_files_by_repository .entry(change.repo_id.clone()) .or_default() @@ -5371,6 +4967,27 @@ fn limit_codegraph_anchors( } fn apply_codegraph_corroboration(graph: &mut GraphAssembly, reports: &[RepositoryCorroboration]) { + let symbol_refs = graph + .nodes + .iter() + .filter(|node| node.kind == NodeKind::SymbolRef) + .map(|node| node.id.clone()) + .collect::>(); + let mut source_locations = std::collections::HashMap::new(); + for evidence in &graph.evidence { + if let (Some(repo_id), Some(file_path)) = (&evidence.repo_id, &evidence.file_path) + && evidence.extractor.starts_with("code-system-graph.source.") + { + source_locations + .entry((repo_id.clone(), file_path.clone())) + .or_insert(( + evidence.start_line, + evidence.end_line, + evidence.content_hash.clone(), + )); + } + } + let mut implementation_evidence = BTreeMap::>::new(); for item in reports { for outcome in &item.report.symbols { let SymbolCorroboration::Confirmed { @@ -5387,29 +5004,15 @@ fn apply_codegraph_corroboration(graph: &mut GraphAssembly, reports: &[Repositor .node_id() }) .into_iter() - .filter(|candidate| { - graph - .nodes - .iter() - .any(|node| node.id == *candidate && node.kind == NodeKind::SymbolRef) - }) + .filter(|candidate| symbol_refs.contains(candidate)) .collect::>(); if implementation_ids.is_empty() { continue; } - let source_evidence = graph.evidence.iter().find(|evidence| { - evidence.repo_id.as_ref() == Some(&item.repo_id) - && evidence.file_path.as_deref() == Some(source_path) - && evidence.extractor.starts_with("code-system-graph.source.") - }); - let (start_line, end_line, content_hash) = - source_evidence.map_or((None, None, None), |evidence| { - ( - evidence.start_line, - evidence.end_line, - evidence.content_hash.clone(), - ) - }); + let (start_line, end_line, content_hash) = source_locations + .get(&(item.repo_id.clone(), source_path.clone())) + .cloned() + .unwrap_or((None, None, None)); let evidence_key = format!( "codegraph:{}:{source_path}:{symbol}:{}", item.repo_id.as_str(), @@ -5429,18 +5032,29 @@ fn apply_codegraph_corroboration(graph: &mut GraphAssembly, reports: &[Repositor content_hash, note: Some("exact provider symbol name, path, and line match".to_owned()), }; - for edge in graph.edges.iter_mut().filter(|edge| { - edge.kind == code_system_graph_model::EdgeKind::ImplementedBy - && implementation_ids.contains(&edge.target) - }) { - if !edge.evidence.contains(&evidence.id) { - edge.evidence.push(evidence.id.clone()); - edge.evidence.sort(); - } + for implementation_id in implementation_ids { + implementation_evidence + .entry(implementation_id) + .or_default() + .insert(evidence.id.clone()); } graph.evidence.push(evidence); } } + for edge in &mut graph.edges { + if edge.kind != code_system_graph_model::EdgeKind::ImplementedBy { + continue; + } + let Some(evidence_ids) = implementation_evidence.get(&edge.target) else { + continue; + }; + for evidence_id in evidence_ids { + if !edge.evidence.contains(evidence_id) { + edge.evidence.push(evidence_id.clone()); + } + } + edge.evidence.sort(); + } graph.evidence.sort_by(|left, right| left.id.cmp(&right.id)); graph.evidence.dedup_by(|left, right| left.id == right.id); } @@ -5710,102 +5324,8 @@ fn cargo_crate_root(checkout: &Path, source_path: &Path) -> String { .unwrap_or_default() } -fn discover_artifact_fingerprints( - context: &WorkspaceContext, -) -> Result, ApplicationError> { - let repository_records = context - .registry - .record - .repositories - .iter() - .map(|repository| (repository.alias.as_str(), repository)) - .collect::>(); - let mut fingerprints = BTreeMap::new(); - for alias in context.manifest.repos.keys() { - let repository = repository_records - .get(alias.as_str()) - .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.clone()))?; - let checkout_path = context - .registry - .checkout_path(alias) - .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.clone()))?; - let effective = context - .repository_configs - .get(alias) - .ok_or_else(|| ApplicationError::RegistryAliasMissing(alias.clone()))?; - for openapi in &effective.openapi { - let fingerprint = fingerprint_artifact( - repository, - checkout_path, - Path::new(openapi), - "code-system-graph.http.openapi", - &context.extraction_budgets, - )?; - fingerprints.insert(artifact_key(&fingerprint), fingerprint); - } - for consumer in &effective.http_consumers { - let fingerprint = fingerprint_artifact( - repository, - checkout_path, - Path::new(&consumer.source), - "code-system-graph.http.declared", - &context.extraction_budgets, - )?; - fingerprints.insert(artifact_key(&fingerprint), fingerprint); - } - for test in &effective.integration_tests { - let fingerprint = fingerprint_artifact( - repository, - checkout_path, - Path::new(&test.path), - "code-system-graph.tests.declared", - &context.extraction_budgets, - )?; - fingerprints.insert(artifact_key(&fingerprint), fingerprint); - } - for implementation in &effective.implementations { - let fingerprint = fingerprint_artifact( - repository, - checkout_path, - Path::new(&implementation.path), - "code-system-graph.implementations.declared", - &context.extraction_budgets, - )?; - fingerprints.insert(artifact_key(&fingerprint), fingerprint); - } - for (relative_path, extractor) in - discover_focused_artifacts(checkout_path, &effective.ignore_policy)? - { - let fingerprint = fingerprint_artifact( - repository, - checkout_path, - &relative_path, - extractor, - &context.extraction_budgets, - )?; - fingerprints.insert(artifact_key(&fingerprint), fingerprint); - } - } - Ok(fingerprints.into_values().collect()) -} - -fn discover_focused_artifacts( - checkout_path: &Path, - ignore_policy: &IgnorePolicy, -) -> Result, ApplicationError> { - let mut discovered = Vec::new(); - for relative in discover_repository_files(checkout_path, ignore_policy, None)? { - worker::report_progress(code_system_graph_core::JobPhase::Discovery, 1); - let path = checkout_path.join(&relative); - for extractor in focused_extractors_for_path(&path) { - discovered.push((relative.clone(), extractor)); - } - } - discovered.sort_by(|left, right| left.0.cmp(&right.0).then(left.1.cmp(right.1))); - Ok(discovered) -} - -fn focused_extractors_for_path(path: &Path) -> Vec<&'static str> { +/// Extractors selected for the repository-relative file `path` of the checkout `checkout`. +fn focused_extractors_for_path(checkout: &Path, path: &Path) -> Vec<&'static str> { let Some(name) = path.file_name().and_then(std::ffi::OsStr::to_str) else { return Vec::new(); }; @@ -5838,8 +5358,10 @@ fn focused_extractors_for_path(path: &Path) -> Vec<&'static str> { "code-system-graph.events.source", "code-system-graph.graphql.source", "code-system-graph.protobuf.generated", - "code-system-graph.data.source", ]); + if !is_third_party_script(path, name, extension.as_deref()) { + extractors.push("code-system-graph.data.source"); + } } if matches!(extension.as_deref(), Some("graphql" | "gql")) { extractors.push("code-system-graph.graphql.document"); @@ -5864,6 +5386,7 @@ fn focused_extractors_for_path(path: &Path) -> Vec<&'static str> { extractors.push("code-system-graph.graphql.persisted"); } extractors.extend(document_extractors_for_path( + checkout, path, name, extension.as_deref(), @@ -5897,7 +5420,37 @@ fn focused_extractors_for_path(path: &Path) -> Vec<&'static str> { extractors } +/// Whether `path` is a vendored, theme, or minified script rather than repository source. +fn is_third_party_script(path: &Path, name: &str, extension: Option<&str>) -> bool { + if !matches!(extension, Some("js" | "jsx")) { + return false; + } + let lower_name = name.to_ascii_lowercase(); + let minified = [".min.js", "-min.js", ".bundle.js", ".chunk.js"] + .iter() + .any(|suffix| lower_name.ends_with(suffix)); + minified + || path + .parent() + .into_iter() + .flat_map(Path::components) + .filter_map(|component| component.as_os_str().to_str()) + .any(|component| { + matches!( + component.to_ascii_lowercase().as_str(), + "theme" + | "themes" + | "vendor" + | "vendors" + | "third_party" + | "third-party" + | "bower_components" + ) + }) +} + fn document_extractors_for_path( + checkout: &Path, path: &Path, name: &str, extension: Option<&str>, @@ -5936,10 +5489,7 @@ fn document_extractors_for_path( if extension == Some("tf") { extractors.push("code-system-graph.infrastructure.terraform"); } - let in_helm_templates = lower_components - .iter() - .any(|component| component == "templates"); - let is_helm = in_helm_templates + let is_helm = is_helm_template(checkout, path, extension) || matches!( lower_name.as_str(), "chart.yaml" | "chart.yml" | "values.yaml" | "values.yml" @@ -6005,90 +5555,30 @@ fn document_extractors_for_path( extractors } -fn fingerprint_artifact( - repository: &RepositoryRecord, - checkout_path: &Path, - relative_path: &Path, - extractor: &str, - budgets: &ExtractionBudgets, -) -> Result { - let configured_path = checkout_path.join(relative_path); - let canonical_path = - std::fs::canonicalize(&configured_path).map_err(|source| ApplicationError::ReadFile { - path: configured_path.clone(), - source, - })?; - if !canonical_path.starts_with(checkout_path) { - return Err(ApplicationError::ArtifactOutsideCheckout { - path: configured_path, - checkout: checkout_path.to_path_buf(), - }); - } - let metadata = - std::fs::metadata(&canonical_path).map_err(|source| ApplicationError::ReadFile { - path: canonical_path.clone(), - source, - })?; - let relative = canonical_path.strip_prefix(checkout_path).map_err(|_| { - ApplicationError::ArtifactOutsideCheckout { - path: canonical_path.clone(), - checkout: checkout_path.to_path_buf(), - } - })?; - let path = encode_native_path(relative); - code_system_graph_model::validate_safe_path_display(&path.display) - .map_err(|_| ApplicationError::UnsafeArtifactPath)?; - let mut tracker = ExtractionTracker::new(&path.display, extractor, budgets); - let content = read_bounded_bytes(&canonical_path, &mut tracker)?; - let fingerprint = ArtifactFingerprint { - repo_id: repository.id.clone(), - checkout_id: repository.checkout_id.clone(), - path, - extractor: extractor.to_owned(), - content_hash: stable_id_bytes("artifact-content", &content), - size_bytes: metadata.len(), - }; - worker::report_progress(code_system_graph_core::JobPhase::Fingerprinting, 1); - Ok(fingerprint) -} - -fn artifact_key( - fingerprint: &ArtifactFingerprint, -) -> (CheckoutId, code_system_graph_model::NativePath, String) { - ( - fingerprint.checkout_id.clone(), - fingerprint.path.clone(), - fingerprint.extractor.clone(), - ) -} - fn extractor_runs( snapshot_id: &str, fingerprints: &[ArtifactFingerprint], plan: &IncrementalPlan, ) -> Vec { - let actions = plan + let changed_keys = plan .changes .iter() .map(|change| { ( - ArtifactKey { - repo_id: change.repo_id.clone(), - checkout_id: change.checkout_id.clone(), - path: change.path.clone(), - extractor: change.extractor.clone(), - }, - change.kind, + &change.repo_id, + &change.checkout_id, + &change.path, + change.extractor.as_str(), ) }) - .collect::>(); - let mut groups = BTreeMap::<(String, String, String), Vec<&ArtifactFingerprint>>::new(); + .collect::>(); + let mut groups = BTreeMap::<(&str, &str, &str), Vec<&ArtifactFingerprint>>::new(); for fingerprint in fingerprints { groups .entry(( - fingerprint.repo_id.as_str().to_owned(), - fingerprint.checkout_id.as_str().to_owned(), - fingerprint.extractor.clone(), + fingerprint.repo_id.as_str(), + fingerprint.checkout_id.as_str(), + fingerprint.extractor.as_str(), )) .or_default() .push(fingerprint); @@ -6099,8 +5589,7 @@ fn extractor_runs( let changed = inputs .iter() .filter(|fingerprint| { - actions.get(&ArtifactKey::from(**fingerprint)) - != Some(&code_system_graph_model::ArtifactChangeKind::Unchanged) + changed_keys.contains(&focused_extraction::artifact_key(fingerprint)) }) .count(); let discovered_files = u64::try_from(inputs.len()).unwrap_or(u64::MAX); @@ -6112,12 +5601,12 @@ fn extractor_runs( snapshot_id: snapshot_id.to_owned(), repo_id: RepoId::new(repo_id), checkout_id: CheckoutId::new(checkout_id), - extractor_version: if focused_extractor(&extractor) { + extractor_version: if focused_extractor(extractor) { EXTRACTION_CONTRACT_VERSION.to_owned() } else { env!("CARGO_PKG_VERSION").to_owned() }, - extractor, + extractor: extractor.to_owned(), status: if changed == 0 { ExtractorRunStatus::SkippedUnchanged } else { @@ -6150,6 +5639,9 @@ fn read_source_file( } } +/// Artifact content bytes read by this process, sampled before and after each scan. +static CONTENT_BYTES_READ: AtomicU64 = AtomicU64::new(0); + fn read_bounded_bytes( path: &Path, tracker: &mut ExtractionTracker, @@ -6171,12 +5663,16 @@ fn read_bounded_bytes( path: path.to_path_buf(), source, })?; - tracker.check_input_bytes(u64::try_from(bytes.len()).unwrap_or(u64::MAX))?; + let read = u64::try_from(bytes.len()).unwrap_or(u64::MAX); + CONTENT_BYTES_READ.fetch_add(read, Ordering::Relaxed); + tracker.check_input_bytes(read)?; Ok(bytes) } #[cfg(test)] mod budget_regression_tests { + use code_system_graph_core::encode_native_path; + use super::*; #[test] @@ -6248,6 +5744,45 @@ mod budget_regression_tests { } } +#[cfg(test)] +mod extractor_selection_tests { + use std::fs; + use std::path::Path; + + use super::focused_extractors_for_path; + + #[test] + fn helm_should_require_a_chart_and_yaml_templates() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let checkout = temporary.path().join("templates").join("deploy"); + fs::create_dir_all(checkout.join("charts/api/templates"))?; + fs::write(checkout.join("charts/api/Chart.yaml"), "name: api\n")?; + let helm = "code-system-graph.infrastructure.helm"; + let selects = |path: &str| focused_extractors_for_path(&checkout, Path::new(path)); + + assert!(selects("charts/api/templates/deployment.yaml").contains(&helm)); + assert!(selects("charts/api/templates/_helpers.tpl").contains(&helm)); + assert!(!selects("charts/api/templates/nginx.conf.j2").contains(&helm)); + assert!(!selects("roles/web/templates/site.yaml").contains(&helm)); + assert!(!selects("src/app.py").contains(&helm)); + Ok(()) + } + + #[test] + fn data_source_should_skip_vendored_theme_and_minified_scripts() { + let checkout = Path::new("/workspace/vendor/shop"); + let data = "code-system-graph.data.source"; + let selects = |path: &str| focused_extractors_for_path(checkout, Path::new(path)); + + assert!(selects("src/orders.js").contains(&data)); + assert!(selects("src/orders.ts").contains(&data)); + assert!(!selects("static/theme/app.js").contains(&data)); + assert!(!selects("public/js/jquery.min.js").contains(&data)); + assert!(!selects("vendor/lib/sql.js").contains(&data)); + assert!(selects("static/theme/app.js").contains(&"code-system-graph.source.javascript")); + } +} + #[cfg(test)] mod codegraph_job_tests { use super::{SymbolAnchor, limit_codegraph_anchors}; diff --git a/crates/code-system-graph-cli/src/main.rs b/crates/code-system-graph-cli/src/main.rs index a103ed8..e88701d 100644 --- a/crates/code-system-graph-cli/src/main.rs +++ b/crates/code-system-graph-cli/src/main.rs @@ -58,11 +58,11 @@ enum LogFormat { #[derive(Debug, Subcommand)] enum Command { /// Internal supervised worker protocol. - #[command(name = "__worker-v1", hide = true)] - WorkerV1, + #[command(name = "__worker", hide = true)] + Worker, /// Internal isolated filesystem-event worker protocol. - #[command(name = "__watch-events-v1", hide = true)] - WatchEventsV1 { + #[command(name = "__watch-events", hide = true)] + WatchEvents { #[arg(long)] config: PathBuf, #[arg(long)] @@ -280,10 +280,6 @@ enum Command { }, /// List, inspect, or compare deterministic graph communities. Communities { - /// `list`, `show`, `compare`, or `recompute`; omitted form preserves flag-based usage. - action: Option, - /// Community ID for `show` or historical snapshot ID for `compare`. - subject: Option, /// `SQLite` database path. #[arg(long)] database: PathBuf, @@ -293,7 +289,7 @@ enum Command { /// Optional exact community identity. #[arg(long)] community_id: Option, - /// Optional historical snapshot identity to compare. + /// Optional previous snapshot identity to compare. #[arg(long)] compare_snapshot: Option, /// Zero-based result offset. @@ -302,9 +298,6 @@ enum Command { /// Maximum communities. #[arg(long, default_value_t = 20)] limit: usize, - /// Workspace manifest used by `recompute`. - #[arg(long, default_value = "code-system-graph.yaml")] - config: PathBuf, }, /// Analyze conservative cross-repository impact and risk. Impact { @@ -351,7 +344,7 @@ enum Command { #[arg(long)] workspace: String, /// Registered repository alias. - #[arg(long = "repo", visible_alias = "repository")] + #[arg(long = "repo")] repository: String, /// `unstaged`, `staged`, `all`, `compare:`, `commit:`, or `range:..`. #[arg(long, default_value = "all", value_parser = parse_change_scope)] @@ -542,8 +535,8 @@ enum Command { const fn command_name(command: &Command) -> &'static str { match command { - Command::WorkerV1 => "__worker-v1", - Command::WatchEventsV1 { .. } => "__watch-events-v1", + Command::Worker => "__worker", + Command::WatchEvents { .. } => "__watch-events", Command::Init { .. } => "init", Command::Scan { .. } => "scan", Command::Sync { .. } => "sync", @@ -879,6 +872,7 @@ fn handle_scan( codegraph, codegraph_binary, repository, + touched_repositories: Vec::new(), force, }, )?; @@ -918,6 +912,7 @@ async fn handle_sync( codegraph: false, codegraph_binary, repository, + touched_repositories: Vec::new(), force, }; if watch { @@ -1376,69 +1371,6 @@ fn handle_communities( Ok(()) } -#[expect( - clippy::too_many_arguments, - reason = "Community compatibility syntax maps positional actions and existing bounded flags" -)] -fn handle_community_command( - database: &std::path::Path, - workspace: &str, - action: Option<&str>, - subject: Option, - community_id: Option, - compare_snapshot: Option, - offset: usize, - limit: usize, - config: &std::path::Path, -) -> anyhow::Result<()> { - if action == Some("recompute") { - if subject.is_some() || community_id.is_some() || compare_snapshot.is_some() { - anyhow::bail!("communities recompute does not accept selection arguments"); - } - let summary = scan_workspace_with_overrides(config, database, &ScanOverrides::default())?; - if summary.workspace != workspace { - anyhow::bail!( - "workspace name `{workspace}` does not match manifest name `{}`", - summary.workspace - ); - } - println!("{}", serde_json::to_string(&summary)?); - return Ok(()); - } - let (community_id, compare_snapshot_id) = match action { - None | Some("list") => (community_id, compare_snapshot), - Some("show") => ( - Some( - subject - .ok_or_else(|| anyhow::anyhow!("communities show requires a community ID"))?, - ), - None, - ), - Some("compare") => ( - None, - Some( - subject - .ok_or_else(|| anyhow::anyhow!("communities compare requires a snapshot ID"))?, - ), - ), - Some(other) => { - anyhow::bail!( - "unknown communities action `{other}`; expected list, show, compare, or recompute" - ) - } - }; - handle_communities( - database, - workspace, - &CommunityInput { - community_id: community_id.map(code_system_graph_model::CommunityId::new), - compare_snapshot_id, - offset, - limit, - }, - ) -} - #[expect( clippy::too_many_arguments, reason = "CLI impact controls remain explicit at the delivery boundary" @@ -1571,8 +1503,8 @@ fn log_command_event( )] async fn dispatch(cli: Cli) -> anyhow::Result<()> { match cli.command { - Command::WorkerV1 => run_worker_from_stdio().map_err(anyhow::Error::msg)?, - Command::WatchEventsV1 { + Command::Worker => run_worker_from_stdio().map_err(anyhow::Error::msg)?, + Command::WatchEvents { config, database, workspace, @@ -1767,25 +1699,21 @@ async fn dispatch(cli: Cli) -> anyhow::Result<()> { k, )?, Command::Communities { - action, - subject, database, workspace, community_id, compare_snapshot, offset, limit, - config, - } => handle_community_command( + } => handle_communities( &database, &workspace, - action.as_deref(), - subject, - community_id, - compare_snapshot, - offset, - limit, - &config, + &CommunityInput { + community_id: community_id.map(code_system_graph_model::CommunityId::new), + compare_snapshot_id: compare_snapshot, + offset, + limit, + }, )?, Command::Impact { database, diff --git a/crates/code-system-graph-cli/src/mcp.rs b/crates/code-system-graph-cli/src/mcp.rs index 618239b..021f896 100644 --- a/crates/code-system-graph-cli/src/mcp.rs +++ b/crates/code-system-graph-cli/src/mcp.rs @@ -13,7 +13,7 @@ use code_system_graph_store_sqlite::{SqliteStore, StoreLock}; use rmcp::handler::server::router::tool::ToolRouter; use rmcp::handler::server::wrapper::Parameters; use rmcp::model::{ - CallToolResult, ContentBlock, Implementation, ListResourceTemplatesResult, ListResourcesResult, PaginatedRequestParams, ReadResourceRequestParams, ReadResourceResponse, ReadResourceResult, Resource, ResourceContents, ResourceTemplate, ServerCapabilities, ServerInfo + CallToolResult, ContentBlock, Implementation, ListResourceTemplatesResult, ListResourcesResult, PaginatedRequestParams, ReadResourceRequestParams, ReadResourceResponse, ReadResourceResult, Resource, ResourceContents, ResourceTemplate, ServerCapabilities, ServerConfig }; use rmcp::service::{RequestContext, RoleServer}; use rmcp::{ErrorData as McpError, ServerHandler, tool, tool_handler, tool_router}; @@ -895,7 +895,7 @@ fn codegraph_disabled_error() -> McpError { )] #[tool_handler(router = self.tool_router)] impl ServerHandler for CodeSystemGraphServer { - fn get_info(&self) -> ServerInfo { + fn get_info(&self) -> ServerConfig { let capabilities = if self.codegraph.enabled { "Use query for persisted entity discovery, explore for ephemeral \ repository source, source_context for source-free evidence, impact for known targets, \ @@ -913,7 +913,7 @@ impl ServerHandler for CodeSystemGraphServer { only the schema catalog embeds fenced JSON. {} Administrative tools mutate state only when enabled and still require the exact workspace.", self.workspace, capabilities ); - ServerInfo::new( + ServerConfig::new( ServerCapabilities::builder() .enable_tools() .enable_resources() diff --git a/crates/code-system-graph-cli/src/mcp_support/agent_context.rs b/crates/code-system-graph-cli/src/mcp_support/agent_context.rs index 63972aa..020fad7 100644 --- a/crates/code-system-graph-cli/src/mcp_support/agent_context.rs +++ b/crates/code-system-graph-cli/src/mcp_support/agent_context.rs @@ -83,14 +83,15 @@ impl AgentPresentationContext { ) -> Result { let store = SqliteStore::open_read_only(database_path).map_err(|error| error.to_string())?; - let registry = store - .load_workspace_registry(workspace) - .map_err(|error| error.to_string())?; - let (nodes, edges) = store - .load_graph_snapshot(snapshot_id) - .map_err(|error| error.to_string())?; - let evidence = store - .load_evidence_snapshot(snapshot_id) + let (registry, nodes, edges, evidence) = store + .consistent_read(|| { + let registry = store.load_workspace_registry(workspace)?; + let (nodes, edges) = store.load_graph_snapshot(snapshot_id)?; + let evidence = store.load_evidence_snapshot(snapshot_id)?; + Ok::<_, code_system_graph_store_sqlite::StoreError>(( + registry, nodes, edges, evidence, + )) + }) .map_err(|error| error.to_string())?; Ok(Self::from_parts( registry diff --git a/crates/code-system-graph-cli/src/mcp_support/agent_views.rs b/crates/code-system-graph-cli/src/mcp_support/agent_views.rs index ca5fb6c..4fccfe7 100644 --- a/crates/code-system-graph-cli/src/mcp_support/agent_views.rs +++ b/crates/code-system-graph-cli/src/mcp_support/agent_views.rs @@ -6,7 +6,7 @@ use code_system_graph_core::{ AgentNextAction, ResolvedSymbol, SearchCoverage, SearchExplanation, SearchReport }; use code_system_graph_model::{ - EdgeKind, EpistemicStatus, EvidenceId, FreshnessSummary, Node, NodeId, NodeKind, Provenance, RepoFreshnessState, RepoId, ToolStatus, TraceReport + EdgeKind, EpistemicStatus, EvidenceId, FreshnessSummary, HttpLinkGap, Node, NodeId, NodeKind, Provenance, RepoFreshnessState, RepoId, ToolStatus, TraceReport }; use schemars::JsonSchema; use serde::{Deserialize, Serialize}; @@ -162,6 +162,8 @@ pub(super) struct AgentQueryReport { pub truncated: bool, pub coverage: SearchCoverage, pub next_actions: Vec, + #[serde(default, skip_serializing_if = "Vec::is_empty")] + pub link_gaps: Vec, } #[derive(Debug, Clone, PartialEq, Serialize, Deserialize, JsonSchema)] @@ -218,7 +220,7 @@ pub(super) struct AgentRepositoryFreshness { #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] pub(super) struct AgentStatusReport { pub workspace: String, - pub database_schema: i64, + pub database_schema: String, pub integrity_ok: bool, pub snapshot_id: String, pub node_count: usize, @@ -348,6 +350,7 @@ pub(super) fn query_view( .take(2) .cloned() .collect(), + link_gaps: report.link_gaps.clone(), } } @@ -495,7 +498,7 @@ pub(super) fn status_view( .count(); AgentStatusReport { workspace: report.workspace.clone(), - database_schema: report.schema_version, + database_schema: report.schema_id.clone(), integrity_ok: report.integrity_ok, snapshot_id: report.snapshot.snapshot_id.clone(), node_count: report.snapshot.node_count, diff --git a/crates/code-system-graph-cli/src/mcp_support/mod.rs b/crates/code-system-graph-cli/src/mcp_support/mod.rs index bf7eac8..2b3e33f 100644 --- a/crates/code-system-graph-cli/src/mcp_support/mod.rs +++ b/crates/code-system-graph-cli/src/mcp_support/mod.rs @@ -390,7 +390,7 @@ pub(super) struct SnapshotMetrics { #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] pub(super) struct GraphStatusReport { pub workspace: String, - pub schema_version: i64, + pub schema_id: String, pub integrity_ok: bool, pub snapshot: SnapshotMetrics, pub repositories: Vec, @@ -500,20 +500,18 @@ pub(super) fn contracts_envelope( report.truncated = total > report.contracts.len(); return Ok((report, freshness)); } - let snapshot = store - .current_snapshot_summary(configured_workspace) - .map_err(|error| error.to_string())?; - let (nodes, edges) = store - .load_graph_snapshot(&snapshot.snapshot_id) - .map_err(|error| error.to_string())?; - let evidence = store - .load_evidence_snapshot(&snapshot.snapshot_id) + let (nodes, edges, evidence, persisted_freshness) = store + .consistent_read(|| { + let snapshot = store.current_snapshot_summary(configured_workspace)?; + let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; + let freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; + Ok::<_, code_system_graph_store_sqlite::StoreError>(( + nodes, edges, evidence, freshness, + )) + }) .map_err(|error| error.to_string())?; - let freshness = freshness_summary( - &store - .load_freshness_snapshot(&snapshot.snapshot_id) - .map_err(|error| error.to_string())?, - ); + let freshness = freshness_summary(&persisted_freshness); let report = inspect_contracts(&nodes, &edges, &evidence, &[], &request) .map_err(|error| error.to_string())?; Ok::<_, String>((report, freshness)) @@ -560,11 +558,16 @@ pub(super) fn source_context_envelope( let result = (|| { let store = SqliteStore::open_read_only(database_path).map_err(|error| error.to_string())?; - let snapshot = store - .current_snapshot_summary(configured_workspace) - .map_err(|error| error.to_string())?; - let (nodes, edges) = store - .load_graph_snapshot(&snapshot.snapshot_id) + let (nodes, edges, persisted_evidence, persisted_freshness) = store + .consistent_read(|| { + let snapshot = store.current_snapshot_summary(configured_workspace)?; + let (nodes, edges) = store.load_graph_snapshot(&snapshot.snapshot_id)?; + let evidence = store.load_evidence_snapshot(&snapshot.snapshot_id)?; + let freshness = store.load_freshness_snapshot(&snapshot.snapshot_id)?; + Ok::<_, code_system_graph_store_sqlite::StoreError>(( + nodes, edges, evidence, freshness, + )) + }) .map_err(|error| error.to_string())?; let entity = nodes .iter() @@ -610,9 +613,7 @@ pub(super) fn source_context_envelope( }) .flat_map(|edge| edge.evidence.iter().cloned()) .collect::>(); - let all_evidence = store - .load_evidence_snapshot(&snapshot.snapshot_id) - .map_err(|error| error.to_string())? + let all_evidence = persisted_evidence .into_iter() .map(|evidence| (evidence.id.clone(), evidence)) .collect::>(); @@ -629,11 +630,7 @@ pub(super) fn source_context_envelope( .filter_map(|evidence_id| all_evidence.get(evidence_id).cloned()) .take(input.evidence_limit) .collect::>(); - let freshness = freshness_summary( - &store - .load_freshness_snapshot(&snapshot.snapshot_id) - .map_err(|error| error.to_string())?, - ); + let freshness = freshness_summary(&persisted_freshness); Ok::<_, String>(( SourceContextReport { workspace: configured_workspace.to_owned(), @@ -886,7 +883,7 @@ fn load_status( Ok(( GraphStatusReport { workspace: workspace.to_owned(), - schema_version: store.schema_version().map_err(|error| error.to_string())?, + schema_id: store.schema_id().map_err(|error| error.to_string())?, integrity_ok: store.integrity_check().map_err(|error| error.to_string())?, snapshot: SnapshotMetrics { snapshot_id: snapshot.snapshot_id, diff --git a/crates/code-system-graph-cli/src/mcp_support/presentation.rs b/crates/code-system-graph-cli/src/mcp_support/presentation.rs index b8d7589..0f70b58 100644 --- a/crates/code-system-graph-cli/src/mcp_support/presentation.rs +++ b/crates/code-system-graph-cli/src/mcp_support/presentation.rs @@ -6,7 +6,9 @@ use std::fmt::Write as _; use code_system_graph_core::{ AgentNextAction, ChangeImpactReport, ContractReport, ImpactReport, PullRequestInspection, SearchReport }; -use code_system_graph_model::{FreshnessSummary, ToolEnvelope, ToolStatus, TraceReport}; +use code_system_graph_model::{ + FreshnessSummary, HttpLinkGapReason, ToolEnvelope, ToolStatus, TraceReport +}; use serde::Serialize; use serde_json::{Value, json}; @@ -497,6 +499,7 @@ fn render_query( .filter(|gap| gap.as_str() != "Full-text scores were not provided.") .cloned() .collect::>(); + add_link_gaps(&mut document, view); add_gaps( &mut document, envelope, @@ -507,6 +510,56 @@ fn render_query( document } +fn add_link_gaps(document: &mut SemanticMarkdown, view: &AgentQueryReport) { + if view.link_gaps.is_empty() { + return; + } + let lines = view + .link_gaps + .iter() + .map(|gap| { + let caller = view + .results + .iter() + .find(|result| { + result.entity.node_id == gap.caller.as_str() + || result + .alternate_node_ids + .iter() + .any(|id| id == gap.caller.as_str()) + }) + .map_or(gap.caller.as_str(), |result| { + entity_display_label(&result.entity) + }); + let outcome = match gap.reason { + HttpLinkGapReason::NoProvider => { + "no provider in the workspace implements it".to_owned() + } + HttpLinkGapReason::Ambiguous => format!( + "{} providers match it equally: {}", + gap.candidates.len(), + gap.candidates + .iter() + .map(|candidate| format!("`{}`", candidate.as_str())) + .collect::>() + .join(", ") + ), + HttpLinkGapReason::External => "it targets a host outside the workspace".to_owned(), + }; + format!("- `{} {}` from {caller}: {outcome}.", gap.method, gap.path) + }) + .collect::>(); + document.add(format!( + "## Unlinked HTTP calls + +{}", + lines.join( + " +" + ) + )); +} + fn render_source_context( envelope: &ToolEnvelope, context: &AgentPresentationContext, diff --git a/crates/code-system-graph-cli/src/mcp_support/presentation/tests.rs b/crates/code-system-graph-cli/src/mcp_support/presentation/tests.rs index bf451de..27ef667 100644 --- a/crates/code-system-graph-cli/src/mcp_support/presentation/tests.rs +++ b/crates/code-system-graph-cli/src/mcp_support/presentation/tests.rs @@ -160,7 +160,7 @@ fn status_markdown_has_a_semantic_summary_heading() { status: ToolStatus::Ok, data: Some(GraphStatusReport { workspace: "example".to_owned(), - schema_version: 2, + schema_id: "schema:test".to_owned(), integrity_ok: true, snapshot: SnapshotMetrics { snapshot_id: "snapshot:example".to_owned(), @@ -416,6 +416,7 @@ fn query_markdown_has_semantic_summary_and_no_debug_syntax() { schema_version: 2, status: ToolStatus::Ok, data: Some(SearchReport { + link_gaps: Vec::new(), hits: Vec::new(), total_matches: 0, offset: 0, diff --git a/crates/code-system-graph-cli/src/mcp_support/presentation_goldens.rs b/crates/code-system-graph-cli/src/mcp_support/presentation_goldens.rs index 9fbebf9..c51a22e 100644 --- a/crates/code-system-graph-cli/src/mcp_support/presentation_goldens.rs +++ b/crates/code-system-graph-cli/src/mcp_support/presentation_goldens.rs @@ -163,6 +163,7 @@ fn search_envelope(fixture: &Fixture) -> ToolEnvelope { schema_version: 2, status: ToolStatus::Degraded, data: Some(SearchReport { + link_gaps: Vec::new(), hits: vec![ SearchHit { node: fixture.api.clone(), @@ -206,6 +207,7 @@ fn single_search_envelope(node: Node) -> ToolEnvelope { schema_version: 2, status: ToolStatus::Ok, data: Some(SearchReport { + link_gaps: Vec::new(), hits: vec![SearchHit { node, score: 10.0, @@ -286,6 +288,7 @@ fn repository_query_does_not_suggest_broad_source_exploration() { schema_version: 2, status: ToolStatus::Ok, data: Some(SearchReport { + link_gaps: Vec::new(), hits: vec![ SearchHit { node: repository.clone(), @@ -914,6 +917,7 @@ fn query_collapses_file_roles_and_omits_structural_contains_noise() { schema_version: 2, status: ToolStatus::Ok, data: Some(SearchReport { + link_gaps: Vec::new(), hits: vec![ SearchHit { node: artifact, diff --git a/crates/code-system-graph-cli/src/mcp_support/presentation_goldens/delivery.rs b/crates/code-system-graph-cli/src/mcp_support/presentation_goldens/delivery.rs index eb3ee2e..24f782d 100644 --- a/crates/code-system-graph-cli/src/mcp_support/presentation_goldens/delivery.rs +++ b/crates/code-system-graph-cli/src/mcp_support/presentation_goldens/delivery.rs @@ -238,7 +238,7 @@ fn healthy_status_is_under_750_bytes_and_uses_aliases() { status: ToolStatus::Ok, data: Some(GraphStatusReport { workspace: "hugint".to_owned(), - schema_version: 5, + schema_id: "schema:test".to_owned(), integrity_ok: true, snapshot: SnapshotMetrics { snapshot_id: "snapshot:current".to_owned(), diff --git a/crates/code-system-graph-cli/src/mcp_support/resources.rs b/crates/code-system-graph-cli/src/mcp_support/resources.rs index b2f00b3..8839d1e 100644 --- a/crates/code-system-graph-cli/src/mcp_support/resources.rs +++ b/crates/code-system-graph-cli/src/mcp_support/resources.rs @@ -6,7 +6,7 @@ use code_system_graph_core::CONTRACT_NODE_KINDS; #[cfg(test)] use code_system_graph_model::RepoFreshnessState; use code_system_graph_model::{ - Community, CommunityAlgorithm, CommunityConfig, CommunityEdgeWeight, CommunityId, CommunityLabelEvidence, CommunityMetrics, CommunityScope, Evidence, ExtractorRun, FreshnessSummary, Node, NodeId, NodeKind, OverallFreshness, RepoFreshness, RepoId, RepositoryRecord + Community, CommunityAlgorithm, CommunityConfig, CommunityEdgeWeight, CommunityId, CommunityLabelEvidence, CommunityMetrics, CommunityScope, Evidence, ExtractorRun, FreshnessSummary, HttpLinkCoverage, HttpLinkGap, HttpLinkReport, Node, NodeId, NodeKind, OverallFreshness, RepoFreshness, RepoId, RepositoryRecord }; use code_system_graph_store_sqlite::{SqliteStore, StoreError}; use serde_json::{Value, json}; @@ -155,6 +155,8 @@ struct CoverageResource { workspace: String, runs: BoundedCollection, freshness: FreshnessResource, + http_links: HttpLinkCoverage, + http_link_gaps: BoundedCollection, } #[derive(Debug)] @@ -226,7 +228,7 @@ fn render_overview(resource: &OverviewResource, maximum: usize) -> String { fn render_status(resource: &StatusResource, maximum: usize) -> String { let mut document = MarkdownDocument::resource("Workspace status", resource.schema_version); document.text("Workspace", &resource.status.workspace); - document.scalar("Database schema", resource.status.schema_version); + document.scalar("Database schema", &resource.status.schema_id); document.scalar("Integrity ok", resource.status.integrity_ok); document.debug("Snapshot", &resource.status.snapshot); render_collection(&mut document, "Repositories", &resource.repositories); @@ -333,6 +335,8 @@ fn render_coverage(resource: &CoverageResource, maximum: usize) -> String { document.text("Workspace", &resource.workspace); render_collection(&mut document, "Extractor runs", &resource.runs); render_freshness(&mut document, &resource.freshness); + document.debug("HTTP link coverage", &resource.http_links); + render_collection(&mut document, "HTTP link gaps", &resource.http_link_gaps); document.render(maximum) } @@ -444,7 +448,7 @@ pub(crate) fn resource_uris(workspace: &str) -> Vec<(String, String, String)> { ( format!("{prefix}/coverage"), "workspace-coverage".to_owned(), - "Bounded extractor and freshness coverage metadata.".to_owned(), + "Bounded extractor, freshness and HTTP link coverage metadata.".to_owned(), ), ] } @@ -639,10 +643,14 @@ fn workspace_coverage_resource( let freshness = store .load_current_freshness(workspace) .map_err(internal_store)?; + let http_links = store + .load_http_link_report(workspace, item_limit) + .map_err(internal_store)?; Ok(ResourceDocument::Coverage(coverage_resource_value( workspace, runs, freshness_summary(&freshness), + http_links, item_limit, ))) } @@ -651,6 +659,7 @@ fn coverage_resource_value( workspace: &str, runs: Vec, freshness: FreshnessSummary, + (http_links, http_link_gap_total): (HttpLinkReport, usize), item_limit: usize, ) -> CoverageResource { let runs = bounded_collection(runs, item_limit); @@ -659,6 +668,11 @@ fn coverage_resource_value( workspace: workspace.to_owned(), runs, freshness: bounded_freshness(freshness, item_limit), + http_links: http_links.coverage, + http_link_gaps: BoundedCollection { + total: http_link_gap_total, + items: http_links.gaps.into_iter().take(item_limit).collect(), + }, } } diff --git a/crates/code-system-graph-cli/src/mcp_support/resources/tests.rs b/crates/code-system-graph-cli/src/mcp_support/resources/tests.rs index a1effc3..e09daf9 100644 --- a/crates/code-system-graph-cli/src/mcp_support/resources/tests.rs +++ b/crates/code-system-graph-cli/src/mcp_support/resources/tests.rs @@ -1,3 +1,5 @@ +use code_system_graph_model::HttpLinkGapReason; + use super::*; fn repository_freshness(name: &str) -> RepoFreshness { RepoFreshness { @@ -79,7 +81,7 @@ fn workspace_resource_fixtures() -> Vec<(&'static str, ResourceDocument)> { schema_version: 2, status: GraphStatusReport { workspace: "workspace".to_owned(), - schema_version: 2, + schema_id: "schema:test".to_owned(), integrity_ok: true, snapshot: SnapshotMetrics { snapshot_id: "snapshot:one".to_owned(), @@ -137,6 +139,8 @@ fn graph_resource_fixtures() -> Vec<(&'static str, ResourceDocument)> { workspace: "workspace".to_owned(), runs: bounded_collection(Vec::new(), 1), freshness: empty_freshness_resource(), + http_links: HttpLinkCoverage::default(), + http_link_gaps: bounded_collection(Vec::new(), 1), }), ), ( @@ -174,12 +178,12 @@ fn every_typed_resource_variant_should_render_its_own_markdown_fixture() { let goldens = [ "mcp-resource-golden:4a0739c1a776d2951cc3f3da46f8818ffd62562b4c023a49103abc31f09ee5a5", "mcp-resource-golden:5bf552e3d0bdda5e42c1cf1ee9ae5de13e263870165cc5b1ed9f5850b801ada2", - "mcp-resource-golden:0bcdb56b1f90dc69cdb60ebd3816e410baa7e44e855fd97b668f2404fe6aa3bb", + "mcp-resource-golden:a7a632d875cbf35d5fed4932dc41d63f0860954dab54a8b907346283566a251e", "mcp-resource-golden:c21d183ace7081c78f80ecb6f3ae117c7fdde9219436dbc25cabe99fe794941a", "mcp-resource-golden:1fc4418c3ef31389ac48f43c3785de986be8621a1e06bac9361f7cec220b2c66", "mcp-resource-golden:cd051baf008bc31e8c87d86bf6a95d71ffb931eaddc9a4b7a90cc3a5d71caf0d", "mcp-resource-golden:d5f3b700d3f70bb59c0a03407f752f754a5233ce8751b56cefffafb9f023145b", - "mcp-resource-golden:0489a80e514c31d3f392dd464843579563baa7f8e3971e92d7a1fc466e457c68", + "mcp-resource-golden:59315d6855580169d482ff2bde2f2302c455a0f632c351aa73949087bf7fe30a", "mcp-resource-golden:e4a27ab1170b53c17cb04f9ba84dee61806da6b493fa258444221da65897fd22", "mcp-resource-golden:4869cf77e9b9e91b47a11f153130722dd46277844b4302f6265dad76f9e30cce", ]; @@ -225,7 +229,7 @@ fn status_resource_should_limit_repositories_with_exact_metadata() { let value = status_resource_value( GraphStatusReport { workspace: "workspace".to_owned(), - schema_version: 2, + schema_id: "schema:test".to_owned(), integrity_ok: true, snapshot: SnapshotMetrics { snapshot_id: "snapshot:one".to_owned(), @@ -258,15 +262,43 @@ fn coverage_resource_should_limit_runs_with_exact_metadata() { stale_repositories: Vec::new(), reasons: Vec::new(), }, + ( + HttpLinkReport { + coverage: HttpLinkCoverage { + linked: 4, + no_provider: 2, + ambiguous: 0, + external: 0, + }, + gaps: vec![ + http_link_gap("test:one", HttpLinkGapReason::NoProvider), + http_link_gap("test:two", HttpLinkGapReason::NoProvider), + ], + }, + 2, + ), 1, ); + assert_eq!(value.http_links.no_provider, 2); + assert_eq!(value.http_link_gaps.total, 2); + assert_eq!(value.http_link_gaps.retained(), 1); assert_eq!(value.runs.total, 2); assert_eq!(value.runs.retained(), 1); assert!(value.runs.truncated()); assert_eq!(value.runs.items.len(), 1); } +fn http_link_gap(caller: &str, reason: HttpLinkGapReason) -> HttpLinkGap { + HttpLinkGap { + caller: NodeId::new(caller), + method: "GET".to_owned(), + path: "/missing".to_owned(), + reason, + candidates: Vec::new(), + } +} + #[test] fn community_view_should_limit_every_nested_collection() { let community = Community { diff --git a/crates/code-system-graph-cli/src/sync.rs b/crates/code-system-graph-cli/src/sync.rs index 5666e53..9ffc567 100644 --- a/crates/code-system-graph-cli/src/sync.rs +++ b/crates/code-system-graph-cli/src/sync.rs @@ -5,6 +5,7 @@ use std::path::{Path, PathBuf}; use std::process::{Command, Stdio}; use std::time::{Duration, Instant}; +use code_system_graph_core::parallel::map_ordered; use code_system_graph_core::{ CodeGraphConfig, CodeGraphProvider, ConfigSource, EffectiveRepositoryConfig, IgnorePolicy, ProviderBudget, ProviderRequest, ProviderStatus }; @@ -13,11 +14,14 @@ use schemars::JsonSchema; use serde::{Deserialize, Serialize}; use tokio_util::sync::CancellationToken; +use super::telemetry::PhaseRecorder; use super::{ - ApplicationError, ScanOverrides, ScanSummary, load_workspace_context, work_database_instance_id + ApplicationError, ScanOverrides, ScanSummary, WorkspaceContext, load_workspace_context, work_database_instance_id }; const MAX_PERSISTED_WATCH_TARGET_BYTES: u64 = 8 * 1024 * 1024; +/// Upper bound on concurrent external `CodeGraph` status and sync processes. +const MAX_CONCURRENT_CODEGRAPH_SYNCS: usize = 4; /// One registered repository that participates in synchronization. #[derive(Debug, Clone, PartialEq, Eq)] @@ -161,25 +165,31 @@ pub fn workspace_sync_targets( overrides: &ScanOverrides, ) -> Result, ApplicationError> { let context = load_workspace_context(config_path, overrides)?; + sync_targets(&context, overrides) +} + +fn sync_targets( + context: &WorkspaceContext, + overrides: &ScanOverrides, +) -> Result, ApplicationError> { if let Some(requested) = &overrides.workspace && requested != &context.manifest.name { return Err(ApplicationError::WorkspaceNameMismatch { requested: requested.clone(), - manifest: context.manifest.name, + manifest: context.manifest.name.clone(), }); } - if let Some(selected) = &overrides.repository - && !context + let selected = overrides.selected_aliases(); + if let Some(unknown) = selected.iter().flatten().find(|alias| { + !context .registry .record .repositories .iter() - .any(|repository| &repository.alias == selected) - { - return Err(ApplicationError::UnknownOverrideRepository( - selected.clone(), - )); + .any(|repository| &repository.alias == *alias) + }) { + return Err(ApplicationError::UnknownOverrideRepository(unknown.clone())); } context @@ -188,10 +198,9 @@ pub fn workspace_sync_targets( .repositories .iter() .filter(|repository| { - overrides - .repository + selected .as_ref() - .is_none_or(|selected| selected == &repository.alias) + .is_none_or(|selected| selected.contains(&repository.alias)) }) .map(|repository| { let path = context @@ -258,7 +267,7 @@ pub fn sync_workspace_with_overrides( /// Synchronizes through an explicitly selected compatible worker executable. /// -/// Embedding applications can pass their own executable after dispatching `__worker-v1` to +/// Embedding applications can pass their own executable after dispatching `__worker` to /// [`crate::run_worker_from_stdio`]. /// /// # Errors @@ -303,24 +312,34 @@ pub(crate) fn sync_workspace_direct( overrides: &ScanOverrides, synchronize_codegraph: bool, ) -> Result { + let phases = PhaseRecorder::start(); let context = load_workspace_context(config_path, overrides)?; - let policy = context.execution_policy; - let workspace = context.manifest.name; - let targets = workspace_sync_targets(config_path, overrides)?; - persist_watch_targets(database_path, &workspace, &targets)?; + let targets = sync_targets(&context, overrides)?; + if overrides.touched_repositories.is_empty() { + persist_watch_targets(database_path, &context.manifest.name, &targets)?; + } let binary = codegraph_binary(overrides); - let codegraph_timeout = Duration::from_millis(policy.max_codegraph_sync_wall_time_ms_per_repo); + let codegraph_timeout = Duration::from_millis( + context + .execution_policy + .max_codegraph_sync_wall_time_ms_per_repo, + ); let codegraph = synchronize_codegraph_targets( &targets, synchronize_codegraph, codegraph_timeout, + context + .execution_policy + .effective_extraction_workers() + .min(MAX_CONCURRENT_CODEGRAPH_SYNCS), |target, deadline| run_codegraph_status(&binary, target, deadline), |path, deadline| run_codegraph_sync(&binary, path, deadline), ); let mut scan_overrides = overrides.clone(); scan_overrides.codegraph = codegraph.synchronized_count > 0; - let scan = super::scan_workspace_direct_for_sync( - config_path, + let scan = super::scan_loaded_workspace_for_sync( + context, + phases, database_path, &scan_overrides, codegraph.changed_count == 0, @@ -526,59 +545,32 @@ fn synchronize_codegraph_targets( targets: &[SyncTarget], enabled: bool, timeout: Duration, - mut inspect: I, - mut synchronize: S, + workers: usize, + inspect: I, + synchronize: S, ) -> CodeGraphSyncSummary where - I: FnMut(&SyncTarget, Instant) -> Result, - S: FnMut(&Path, Instant) -> Result<(), String>, + I: Fn(&SyncTarget, Instant) -> Result + Sync, + S: Fn(&Path, Instant) -> Result<(), String> + Sync, { - let mut observations = Vec::with_capacity(targets.len()); - let mut repositories = Vec::with_capacity(targets.len()); - if enabled { - for target in targets { - let deadline = Instant::now() + timeout; - let observation = if target.path.join(".codegraph").is_dir() { - match inspect(target, deadline) { - Ok(ProviderStatus::Available) => CodeGraphSyncObservation { - state: CodeGraphSyncState::Synchronized, - detail: None, - index_changed: false, - }, - Ok(ProviderStatus::Stale) => match synchronize(&target.path, deadline) { - Ok(()) => CodeGraphSyncObservation { - state: CodeGraphSyncState::Synchronized, - detail: None, - index_changed: true, - }, - Err(detail) => failed_sync_observation(&detail), - }, - Ok(ProviderStatus::IndexMissing) => CodeGraphSyncObservation { - state: CodeGraphSyncState::SkippedNotInitialized, - detail: Some("local CodeGraph index is not initialized".to_owned()), - index_changed: false, - }, - Ok(status) => failed_sync_observation(&format!( - "CodeGraph status is not usable for synchronization: {status:?}" - )), - Err(detail) => failed_sync_observation(&detail), - } - } else { - CodeGraphSyncObservation { - state: CodeGraphSyncState::SkippedNotInitialized, - detail: Some("local CodeGraph index is not initialized".to_owned()), - index_changed: false, - } - }; - repositories.push(CodeGraphRepositorySync { - repository: target.alias.clone(), - state: observation.state, - detail: observation.detail.clone(), - }); - observations.push(observation); + let observations = if enabled { + map_ordered(targets, workers, |target| { + let observation = synchronize_codegraph_target(target, timeout, &inspect, &synchronize); super::worker::report_progress(code_system_graph_core::JobPhase::CodeGraphSync, 1); - } - } + observation + }) + } else { + Vec::new() + }; + let mut repositories = targets + .iter() + .zip(&observations) + .map(|(target, observation)| CodeGraphRepositorySync { + repository: target.alias.clone(), + state: observation.state, + detail: observation.detail.clone(), + }) + .collect::>(); repositories.sort_by(|left, right| left.repository.cmp(&right.repository)); let synchronized_count = repositories .iter() @@ -612,6 +604,50 @@ where } } +fn synchronize_codegraph_target( + target: &SyncTarget, + timeout: Duration, + inspect: &I, + synchronize: &S, +) -> CodeGraphSyncObservation +where + I: Fn(&SyncTarget, Instant) -> Result, + S: Fn(&Path, Instant) -> Result<(), String>, +{ + let deadline = Instant::now() + timeout; + if !target.path.join(".codegraph").is_dir() { + return CodeGraphSyncObservation { + state: CodeGraphSyncState::SkippedNotInitialized, + detail: Some("local CodeGraph index is not initialized".to_owned()), + index_changed: false, + }; + } + match inspect(target, deadline) { + Ok(ProviderStatus::Available) => CodeGraphSyncObservation { + state: CodeGraphSyncState::Synchronized, + detail: None, + index_changed: false, + }, + Ok(ProviderStatus::Stale) => match synchronize(&target.path, deadline) { + Ok(()) => CodeGraphSyncObservation { + state: CodeGraphSyncState::Synchronized, + detail: None, + index_changed: true, + }, + Err(detail) => failed_sync_observation(&detail), + }, + Ok(ProviderStatus::IndexMissing) => CodeGraphSyncObservation { + state: CodeGraphSyncState::SkippedNotInitialized, + detail: Some("local CodeGraph index is not initialized".to_owned()), + index_changed: false, + }, + Ok(status) => failed_sync_observation(&format!( + "CodeGraph status is not usable for synchronization: {status:?}" + )), + Err(detail) => failed_sync_observation(&detail), + } +} + fn failed_sync_observation(detail: &str) -> CodeGraphSyncObservation { CodeGraphSyncObservation { state: CodeGraphSyncState::Failed, @@ -633,7 +669,7 @@ fn bounded_detail(detail: &str) -> String { #[cfg(test)] mod tests { - use std::cell::RefCell; + use std::sync::Mutex; use super::*; @@ -683,21 +719,25 @@ mod tests { let absent = temporary.path().join("absent"); std::fs::create_dir_all(initialized.join(".codegraph"))?; std::fs::create_dir(&absent)?; - let called = RefCell::new(Vec::new()); + let called = Mutex::new(Vec::new()); let targets = vec![target("zeta", initialized.clone()), target("alpha", absent)]; let report = synchronize_codegraph_targets( &targets, true, Duration::from_secs(1), + 2, |target, _| { - called.borrow_mut().push(target.path.clone()); + called + .lock() + .map_err(|_| "poisoned".to_owned())? + .push(target.path.clone()); Err("failure ".repeat(600)) }, |_, _| panic!("failed status inspection must not synchronize"), ); - assert_eq!(called.into_inner(), vec![initialized]); + assert_eq!(called.into_inner()?, vec![initialized]); assert_eq!(report.repository_count, 2); assert_eq!(report.synchronized_count, 0); assert_eq!(report.skipped_count, 1); @@ -723,6 +763,7 @@ mod tests { &[target], false, Duration::from_secs(1), + 2, |_, _| panic!("disabled synchronization must not inspect CodeGraph"), |_, _| panic!("disabled synchronization must not invoke CodeGraph"), ); @@ -738,13 +779,14 @@ mod tests { let stale = temporary.path().join("stale"); std::fs::create_dir_all(current.join(".codegraph"))?; std::fs::create_dir_all(stale.join(".codegraph"))?; - let synchronized = RefCell::new(Vec::new()); + let synchronized = Mutex::new(Vec::new()); let targets = vec![target("current", current), target("stale", stale.clone())]; let report = synchronize_codegraph_targets( &targets, true, Duration::from_secs(1), + 2, |target, _| { if target.alias == "stale" { Ok(ProviderStatus::Stale) @@ -753,12 +795,15 @@ mod tests { } }, |path, _| { - synchronized.borrow_mut().push(path.to_path_buf()); + synchronized + .lock() + .map_err(|_| "poisoned".to_owned())? + .push(path.to_path_buf()); Ok(()) }, ); - assert_eq!(synchronized.into_inner(), vec![stale]); + assert_eq!(synchronized.into_inner()?, vec![stale]); assert_eq!(report.synchronized_count, 2); assert_eq!(report.changed_count, 1); assert_eq!(report.unchanged_count, 1); @@ -770,27 +815,32 @@ mod tests { let temporary = tempfile::tempdir()?; let repository = temporary.path().join("stale"); std::fs::create_dir_all(repository.join(".codegraph"))?; - let inspected_deadline = RefCell::new(None); - let synchronized_deadline = RefCell::new(None); + let inspected_deadline = Mutex::new(None); + let synchronized_deadline = Mutex::new(None); let report = synchronize_codegraph_targets( &[target("stale", repository)], true, Duration::from_secs(1), + 2, |_, deadline| { - inspected_deadline.replace(Some(deadline)); + *inspected_deadline + .lock() + .map_err(|_| "poisoned".to_owned())? = Some(deadline); Ok(ProviderStatus::Stale) }, |_, deadline| { - synchronized_deadline.replace(Some(deadline)); + *synchronized_deadline + .lock() + .map_err(|_| "poisoned".to_owned())? = Some(deadline); Ok(()) }, ); assert_eq!(report.changed_count, 1); assert_eq!( - inspected_deadline.into_inner(), - synchronized_deadline.into_inner() + inspected_deadline.into_inner()?, + synchronized_deadline.into_inner()? ); Ok(()) } diff --git a/crates/code-system-graph-cli/src/sync_watch.rs b/crates/code-system-graph-cli/src/sync_watch.rs index b4e7f2c..3a47eaf 100644 --- a/crates/code-system-graph-cli/src/sync_watch.rs +++ b/crates/code-system-graph-cli/src/sync_watch.rs @@ -1,6 +1,5 @@ //! Portable filesystem watching for the `sync --watch` command. -#[cfg(target_os = "linux")] use std::collections::BTreeSet; use std::io::{BufRead, BufReader, Read, Write}; #[cfg(target_os = "linux")] @@ -52,7 +51,7 @@ enum WatchSignal { const DEFAULT_POLL_INTERVAL: Duration = Duration::from_secs(2); const MAX_SYNC_ATTEMPTS: usize = 3; -const WATCH_EVENT_PROTOCOL_MAX_BYTES: usize = 128; +const WATCH_EVENT_PROTOCOL_MAX_BYTES: usize = 16 * 1024; #[derive(Debug, serde::Deserialize, serde::Serialize)] #[serde(tag = "type", rename_all = "snake_case")] @@ -64,6 +63,9 @@ enum WatchEventMessage { schema_version: u8, #[serde(default)] refresh_scope: bool, + /// Aliases touched since the previous message; empty means the whole workspace. + #[serde(default)] + repositories: Vec, }, } @@ -72,9 +74,69 @@ struct WatchEventState { initial_ready: bool, refresh_started: Option, refresh_scope: bool, + touched: TouchedRepositories, failed: bool, } +/// Repositories changed since the last synchronization pass. +#[derive(Debug, Default)] +struct TouchedRepositories { + aliases: BTreeSet, + workspace: bool, +} + +impl TouchedRepositories { + fn record(&mut self, scope: EventScope) { + match scope { + EventScope::Irrelevant => {} + EventScope::Repositories(aliases) => self.aliases.extend(aliases), + EventScope::Workspace => self.workspace = true, + } + } + + /// Returns the touched aliases, or `None` when the next pass must cover the workspace. + fn take(&mut self) -> Option> { + let touched = std::mem::take(self); + (!touched.workspace && !touched.aliases.is_empty()) + .then(|| touched.aliases.into_iter().collect()) + } +} + +/// Part of the watched scope affected by one filesystem event. +#[derive(Debug, PartialEq, Eq)] +enum EventScope { + Irrelevant, + Repositories(BTreeSet), + Workspace, +} + +/// Change state accumulated by watcher callbacks between two `Dirty` messages. +#[derive(Debug, Default)] +struct PendingChanges { + refresh_directories: AtomicBool, + refresh_scope: AtomicBool, + touched: Mutex, +} + +impl PendingChanges { + fn record(&self, scope: EventScope) { + match self.touched.lock() { + Ok(mut touched) => touched.record(scope), + Err(poisoned) => poisoned.into_inner().record(EventScope::Workspace), + } + } + + fn take_touched(&self) -> Option> { + match self.touched.lock() { + Ok(mut touched) => touched.take(), + Err(poisoned) => { + poisoned.into_inner().take(); + None + } + } + } +} + struct WatchEventWorker { child: Child, group: code_system_graph::SupervisedProcessGroup, @@ -94,7 +156,7 @@ impl WatchEventWorker { let executable = std::env::current_exe().context("failed to locate watcher worker")?; let mut command = Command::new(executable); command - .arg("__watch-events-v1") + .arg("__watch-events") .arg("--config") .arg(config) .arg("--database") @@ -227,6 +289,14 @@ impl WatchEventWorker { .map_err(|_| anyhow::anyhow!("watcher protocol state was poisoned"))?; Ok(std::mem::take(&mut state.refresh_scope)) } + + fn take_touched_repositories(&self) -> anyhow::Result>> { + let mut state = self + .state + .lock() + .map_err(|_| anyhow::anyhow!("watcher protocol state was poisoned"))?; + Ok(state.touched.take()) + } } impl Drop for WatchEventWorker { @@ -268,9 +338,15 @@ fn spawn_watch_event_reader( WatchEventMessage::Dirty { schema_version: 2, refresh_scope, + repositories, } => { current.refresh_started.get_or_insert_with(Instant::now); current.refresh_scope |= refresh_scope; + current.touched.record(if repositories.is_empty() { + EventScope::Workspace + } else { + EventScope::Repositories(repositories.into_iter().collect()) + }); drop(current); match sender.try_send(()) { Ok(()) | Err(mpsc::error::TrySendError::Full(())) => {} @@ -312,6 +388,7 @@ struct WatchScope { #[derive(Debug, Clone)] struct WatchRepository { + alias: String, root: PathBuf, ignore_policy: IgnorePolicy, ignore_matcher: Arc>, @@ -319,9 +396,15 @@ struct WatchRepository { } impl WatchRepository { - fn new(root: PathBuf, ignore_policy: IgnorePolicy, explicit_paths: Vec) -> Self { + fn new( + alias: String, + root: PathBuf, + ignore_policy: IgnorePolicy, + explicit_paths: Vec, + ) -> Self { let ignore_matcher = RepositoryPathMatcher::new(&root, ignore_policy.clone()); Self { + alias, root, ignore_policy, ignore_matcher: Arc::new(Mutex::new(ignore_matcher)), @@ -376,7 +459,8 @@ impl WatchRepository { impl PartialEq for WatchRepository { fn eq(&self, other: &Self) -> bool { - self.root == other.root + self.alias == other.alias + && self.root == other.root && self.ignore_policy == other.ignore_policy && self.explicit_paths == other.explicit_paths } @@ -389,10 +473,16 @@ impl WatchScope { let mut repositories = load_persisted_watch_targets(database, workspace)? .into_iter() .map(|target| { - WatchRepository::new(target.path, target.ignore_policy, target.explicit_paths) + WatchRepository::new( + target.alias, + target.path, + target.ignore_policy, + target.explicit_paths, + ) }) .collect::>(); - repositories.sort_by(|left, right| left.root.cmp(&right.root)); + repositories + .sort_by(|left, right| (&left.root, &left.alias).cmp(&(&right.root, &right.alias))); Ok(Self { config: absolute_path(config)?, database: absolute_path(database)?, @@ -400,16 +490,23 @@ impl WatchScope { }) } - fn relevant_event(&self, event: &Event) -> anyhow::Result { - if event.paths.is_empty() { - return Ok(true); + fn event_scope(&self, event: &Event) -> anyhow::Result { + if event.paths.is_empty() || event.need_rescan() { + return Ok(EventScope::Workspace); } + let mut aliases = BTreeSet::new(); for path in &event.paths { - if self.relevant_path(path)? { - return Ok(true); + match self.path_scope(path)? { + EventScope::Irrelevant => {} + EventScope::Repositories(touched) => aliases.extend(touched), + EventScope::Workspace => return Ok(EventScope::Workspace), } } - Ok(false) + Ok(if aliases.is_empty() { + EventScope::Irrelevant + } else { + EventScope::Repositories(aliases) + }) } fn changes_enabled_gitignore(&self, event: &Event) -> bool { @@ -422,19 +519,29 @@ impl WatchScope { }) } - fn relevant_path(&self, path: &Path) -> anyhow::Result { + fn path_scope(&self, path: &Path) -> anyhow::Result { if path == self.config { - return Ok(true); + return Ok(EventScope::Workspace); } if database_artifact(path, &self.database) { - return Ok(false); + return Ok(EventScope::Irrelevant); } + let mut aliases = BTreeSet::new(); for repository in &self.repositories { if repository.relevant(path)? { - return Ok(true); + aliases.insert(repository.alias.clone()); } } - Ok(false) + Ok(if aliases.is_empty() { + EventScope::Irrelevant + } else { + EventScope::Repositories(aliases) + }) + } + + #[cfg(test)] + fn relevant_path(&self, path: &Path) -> anyhow::Result { + Ok(self.path_scope(path)? != EventScope::Irrelevant) } #[cfg(any(target_os = "linux", test))] @@ -623,24 +730,22 @@ pub(crate) async fn run_watch_event_worker( }; let scope = WatchScope::load(&config, &database, &workspace)?; let (sender, mut receiver) = mpsc::channel(1); - let refresh_required = Arc::new(AtomicBool::new(false)); - let scope_refresh_required = Arc::new(AtomicBool::new(false)); + let pending = Arc::new(PendingChanges::default()); let requested_poll = poll_interval_ms.map(Duration::from_millis); let mut watcher = build_watcher( &scope, sender.clone(), - Arc::clone(&refresh_required), - Arc::clone(&scope_refresh_required), + Arc::clone(&pending), requested_poll, &policy, )?; emit_watch_event(&WatchEventMessage::Ready { schema_version: 2 })?; while let Some(signal) = receiver.recv().await { - emit_watch_event(&WatchEventMessage::Dirty { - schema_version: 2, - refresh_scope: scope_refresh_required.swap(false, Ordering::AcqRel), - })?; - let refresh = refresh_required.swap(false, Ordering::AcqRel) + emit_dirty( + pending.refresh_scope.swap(false, Ordering::AcqRel), + pending.take_touched().unwrap_or_default(), + )?; + let refresh = pending.refresh_directories.swap(false, Ordering::AcqRel) || matches!(signal, WatchSignal::RefreshDirectories); if refresh { match watcher.refresh_native_directories(&scope, &policy) { @@ -650,8 +755,7 @@ pub(crate) async fn run_watch_event_worker( watcher = build_poll_watcher( &scope, sender.clone(), - Arc::clone(&refresh_required), - Arc::clone(&scope_refresh_required), + Arc::clone(&pending), requested_poll.unwrap_or(DEFAULT_POLL_INTERVAL), &policy, )?; @@ -663,6 +767,24 @@ pub(crate) async fn run_watch_event_worker( anyhow::bail!("filesystem event channel stopped") } +/// Emits a `Dirty` message, widening it to the whole workspace when the alias list would exceed +/// the protocol bound. +fn emit_dirty(refresh_scope: bool, repositories: Vec) -> anyhow::Result<()> { + let message = WatchEventMessage::Dirty { + schema_version: 2, + refresh_scope, + repositories, + }; + if serde_json::to_vec(&message)?.len() < WATCH_EVENT_PROTOCOL_MAX_BYTES { + return emit_watch_event(&message); + } + emit_watch_event(&WatchEventMessage::Dirty { + schema_version: 2, + refresh_scope, + repositories: Vec::new(), + }) +} + fn emit_watch_event(message: &WatchEventMessage) -> anyhow::Result<()> { let encoded = serde_json::to_vec(message)?; anyhow::ensure!( @@ -828,6 +950,14 @@ pub(crate) async fn watch_workspace( return Err(watcher_public_error(&error)); } }; + let touched = match watcher.take_touched_repositories() { + Ok(touched) => touched, + Err(error) => { + finish_after_error(&database, &workspace, &owner_token, &error)?; + return Err(watcher_public_error(&error)); + } + }; + let pass_overrides = targeted_pass_overrides(&overrides, refresh_scope, touched); let minimum_start = last_pass_started + Duration::from_millis(policy.min_watch_rescan_interval_ms); if Instant::now() < minimum_start { @@ -848,7 +978,7 @@ pub(crate) async fn watch_workspace( if let Err(error) = retry_sync( &config, &database, - &overrides, + &pass_overrides, synchronize_codegraph, session_started, policy.max_watch_session_wall_time_ms, @@ -930,6 +1060,19 @@ pub(crate) async fn watch_workspace( } } +/// Restricts a watch pass to the touched repositories unless the scope or workspace changed. +fn targeted_pass_overrides( + overrides: &ScanOverrides, + refresh_scope: bool, + touched: Option>, +) -> ScanOverrides { + let mut pass = overrides.clone(); + if !refresh_scope && overrides.repository.is_none() { + pass.touched_repositories = touched.unwrap_or_default(); + } + pass +} + fn emit_sync( config: &Path, database: &Path, @@ -1144,8 +1287,7 @@ fn finish_and_emit( fn build_watcher( scope: &WatchScope, sender: mpsc::Sender, - refresh_required: Arc, - scope_refresh_required: Arc, + pending: Arc, requested_poll_interval: Option, policy: &code_system_graph_core::ExecutionPolicy, ) -> anyhow::Result { @@ -1159,32 +1301,18 @@ fn build_watcher( return build_poll_watcher( scope, sender, - refresh_required, - scope_refresh_required, + pending, requested_poll_interval.unwrap_or(DEFAULT_POLL_INTERVAL), policy, ); } - match build_native_watcher( - scope, - sender.clone(), - Arc::clone(&refresh_required), - Arc::clone(&scope_refresh_required), - policy, - ) { + match build_native_watcher(scope, sender.clone(), Arc::clone(&pending), policy) { Ok(watcher) => Ok(watcher), Err(error) if watcher_limit_exceeded(&error) => Err(error), Err(error) => { eprintln!("csgraph sync native watcher unavailable ({error}); falling back to polling"); - build_poll_watcher( - scope, - sender, - refresh_required, - scope_refresh_required, - DEFAULT_POLL_INTERVAL, - policy, - ) + build_poll_watcher(scope, sender, pending, DEFAULT_POLL_INTERVAL, policy) } } } @@ -1198,20 +1326,13 @@ fn watcher_limit_exceeded(error: &anyhow::Error) -> bool { fn build_native_watcher( scope: &WatchScope, sender: mpsc::Sender, - refresh_required: Arc, - scope_refresh_required: Arc, + pending: Arc, policy: &code_system_graph_core::ExecutionPolicy, ) -> anyhow::Result { let callback_scope = scope.clone(); let mut watcher = RecommendedWatcher::new( move |result| { - forward_event( - result, - &callback_scope, - &sender, - &refresh_required, - &scope_refresh_required, - ); + forward_event(result, &callback_scope, &sender, &pending); }, Config::default().with_follow_symlinks(false), ) @@ -1233,8 +1354,7 @@ fn build_native_watcher( fn build_poll_watcher( scope: &WatchScope, sender: mpsc::Sender, - refresh_required: Arc, - scope_refresh_required: Arc, + pending: Arc, interval: Duration, policy: &code_system_graph_core::ExecutionPolicy, ) -> anyhow::Result { @@ -1245,13 +1365,7 @@ fn build_poll_watcher( .with_follow_symlinks(false); let mut watcher = PollWatcher::new( move |result| { - forward_event( - result, - &callback_scope, - &sender, - &refresh_required, - &scope_refresh_required, - ); + forward_event(result, &callback_scope, &sender, &pending); }, config, ) @@ -1336,27 +1450,28 @@ fn forward_event( result: notify::Result, scope: &WatchScope, sender: &mpsc::Sender, - refresh_required: &AtomicBool, - scope_refresh_required: &AtomicBool, + pending: &PendingChanges, ) { match result { - Ok(event) => match scope.relevant_event(&event) { - Ok(true) => { + Ok(event) => match scope.event_scope(&event) { + Ok(EventScope::Irrelevant) => {} + Ok(touched) => { + pending.record(touched); if scope.changes_enabled_gitignore(&event) { - scope_refresh_required.store(true, Ordering::Release); + pending.refresh_scope.store(true, Ordering::Release); } let signal = if event.kind.is_create() && event.paths.iter().any(|path| path.is_dir()) { - refresh_required.store(true, Ordering::Release); + pending.refresh_directories.store(true, Ordering::Release); WatchSignal::RefreshDirectories } else { WatchSignal::Dirty }; let _ignored = sender.try_send(signal); } - Ok(false) => {} Err(error) => { eprintln!("csgraph sync could not apply repository ignore policy: {error}"); + pending.record(EventScope::Workspace); let _ignored = sender.try_send(WatchSignal::Dirty); } }, @@ -1368,6 +1483,7 @@ fn forward_event( .all(|path| database_artifact(path, &scope.database)) => {} Err(error) => { eprintln!("csgraph sync filesystem watcher reported: {error}"); + pending.record(EventScope::Workspace); let _ignored = sender.try_send(WatchSignal::Dirty); } } @@ -1401,7 +1517,7 @@ fn database_artifact(path: &Path, database: &Path) -> bool { } let work_database = { let mut value = database.as_os_str().to_os_string(); - value.push(".work-v1.db"); + value.push(".work.db"); PathBuf::from(value) }; if path == work_database { @@ -1419,10 +1535,10 @@ fn database_artifact(path: &Path, database: &Path) -> bool { "-wal", "-shm", "-journal", - ".work-v1.db", - ".work-v1.db-wal", - ".work-v1.db-shm", - ".work-v1.db-journal", + ".work.db", + ".work.db-wal", + ".work.db-shm", + ".work.db-journal", ] .iter() .any(|suffix| name == format!("{database_name}{suffix}")) @@ -1479,6 +1595,7 @@ mod tests { initial_ready: true, refresh_started: Some(Instant::now()), refresh_scope: false, + touched: TouchedRepositories::default(), failed: false, })); let input = concat!( @@ -1516,6 +1633,7 @@ mod tests { initial_ready: true, refresh_started: Some(original), refresh_scope: false, + touched: TouchedRepositories::default(), failed: false, })); let input = concat!( @@ -1551,21 +1669,14 @@ mod tests { sender .try_send(WatchSignal::Dirty) .expect("channel should accept prefill"); - let refresh_required = AtomicBool::new(false); + let pending = PendingChanges::default(); let event = Event::new(notify::EventKind::Create(notify::event::CreateKind::Folder)) .add_path(directory); - let scope_refresh_required = AtomicBool::new(false); - forward_event( - Ok(event), - &scope, - &sender, - &refresh_required, - &scope_refresh_required, - ); + forward_event(Ok(event), &scope, &sender, &pending); assert!(matches!(receiver.try_recv(), Ok(WatchSignal::Dirty))); - assert!(refresh_required.swap(false, Ordering::AcqRel)); + assert!(pending.refresh_directories.swap(false, Ordering::AcqRel)); } #[test] @@ -1577,6 +1688,7 @@ mod tests { config: temporary.path().join("code-system-graph.yaml"), database: temporary.path().join("graph.db"), repositories: vec![WatchRepository::new( + "repo".to_owned(), root.clone(), gitignore_policy(), Vec::new(), @@ -1586,23 +1698,17 @@ mod tests { sender .try_send(WatchSignal::Dirty) .expect("channel should accept prefill"); - let refresh_required = AtomicBool::new(false); - let scope_refresh_required = AtomicBool::new(false); + let pending = PendingChanges::default(); let event = Event::new(notify::EventKind::Modify(notify::event::ModifyKind::Data( notify::event::DataChange::Any, ))) .add_path(root.join(".gitignore")); - forward_event( - Ok(event), - &scope, - &sender, - &refresh_required, - &scope_refresh_required, - ); + forward_event(Ok(event), &scope, &sender, &pending); assert!(matches!(receiver.try_recv(), Ok(WatchSignal::Dirty))); - assert!(scope_refresh_required.swap(false, Ordering::AcqRel)); + assert!(pending.refresh_scope.swap(false, Ordering::AcqRel)); + assert_eq!(pending.take_touched(), Some(vec!["repo".to_owned()])); } #[test] @@ -1671,6 +1777,7 @@ mod tests { config: temporary.path().join("code-system-graph.yaml"), database: temporary.path().join("graph.db"), repositories: vec![WatchRepository::new( + "repo".to_owned(), root.clone(), gitignore_policy(), Vec::new(), @@ -1688,6 +1795,7 @@ mod tests { config: temporary.path().join("code-system-graph.yaml"), database: temporary.path().join("graph.db"), repositories: vec![WatchRepository::new( + "repo".to_owned(), root.clone(), gitignore_policy(), Vec::new(), @@ -1704,6 +1812,7 @@ mod tests { config: PathBuf::from("/workspace/code-system-graph.yaml"), database: PathBuf::from("/workspace/.state/graph.db"), repositories: vec![WatchRepository::new( + "repo".to_owned(), PathBuf::from("/workspace/repo"), policy(&[], &[]), vec![PathBuf::from(".code-system-graph.yaml")], @@ -1713,8 +1822,8 @@ mod tests { assert!(scope.relevant_path(Path::new("/workspace/code-system-graph.yaml"))?); assert!(!scope.relevant_path(Path::new("/workspace/repo/.codegraph/codegraph.db-wal"))?); assert!(!scope.relevant_path(Path::new("/workspace/.state/graph.db-wal"))?); - assert!(!scope.relevant_path(Path::new("/workspace/.state/graph.db.work-v1.db-wal"))?); - assert!(!scope.relevant_path(Path::new("/workspace/.state/graph.db.work-v1.db-shm"))?); + assert!(!scope.relevant_path(Path::new("/workspace/.state/graph.db.work.db-wal"))?); + assert!(!scope.relevant_path(Path::new("/workspace/.state/graph.db.work.db-shm"))?); Ok(()) } @@ -1724,6 +1833,7 @@ mod tests { config: PathBuf::from("/workspace/code-system-graph.yaml"), database: PathBuf::from("/workspace/.state/graph.db"), repositories: vec![WatchRepository::new( + "repo".to_owned(), PathBuf::from("/workspace/repo"), policy(&["./generated//./**"], &["./vendor//internal-sdk/./**"]), vec![PathBuf::from("generated/explicit.yaml")], @@ -1754,17 +1864,90 @@ mod tests { config: temporary.path().join("code-system-graph.yaml"), database: temporary.path().join("graph.db"), repositories: vec![ - WatchRepository::new(root.clone(), policy(&["generated/*"], &[]), Vec::new()), - WatchRepository::new(root, policy(&[], &[]), Vec::new()), + WatchRepository::new( + "restrictive".to_owned(), + root.clone(), + policy(&["generated/*"], &[]), + Vec::new(), + ), + WatchRepository::new("permissive".to_owned(), root, policy(&[], &[]), Vec::new()), ], }; assert_eq!(scope.repositories.len(), 2); assert!(scope.relevant_path(&source)?); + assert!(matches!( + scope.path_scope(&source)?, + EventScope::Repositories(aliases) if aliases.contains("permissive") + )); assert!(scope.should_watch_directory(&repository)?); Ok(()) } + #[test] + fn manifest_and_rescan_events_should_widen_the_pass_to_the_workspace() -> anyhow::Result<()> { + let scope = WatchScope { + config: PathBuf::from("/workspace/code-system-graph.yaml"), + database: PathBuf::from("/workspace/.state/graph.db"), + repositories: vec![WatchRepository::new( + "repo".to_owned(), + PathBuf::from("/workspace/repo"), + policy(&[], &[]), + Vec::new(), + )], + }; + let modify = || { + Event::new(notify::EventKind::Modify(notify::event::ModifyKind::Data( + notify::event::DataChange::Any, + ))) + }; + + assert_eq!( + scope.event_scope(&modify().add_path(PathBuf::from("/workspace/repo/src/lib.rs")))?, + EventScope::Repositories(BTreeSet::from(["repo".to_owned()])) + ); + assert_eq!( + scope.event_scope( + &modify().add_path(PathBuf::from("/workspace/code-system-graph.yaml")) + )?, + EventScope::Workspace + ); + assert_eq!( + scope.event_scope(&modify().set_flag(notify::event::Flag::Rescan))?, + EventScope::Workspace + ); + assert_eq!( + scope.event_scope(&modify().add_path(PathBuf::from("/elsewhere/file.rs")))?, + EventScope::Irrelevant + ); + Ok(()) + } + + #[test] + fn touched_repositories_should_select_a_targeted_pass_only_when_bounded() { + let mut touched = TouchedRepositories::default(); + touched.record(EventScope::Repositories(BTreeSet::from(["b".to_owned()]))); + touched.record(EventScope::Repositories(BTreeSet::from(["a".to_owned()]))); + assert_eq!(touched.take(), Some(vec!["a".to_owned(), "b".to_owned()])); + assert_eq!(touched.take(), None); + + touched.record(EventScope::Repositories(BTreeSet::from(["a".to_owned()]))); + touched.record(EventScope::Workspace); + assert_eq!(touched.take(), None); + + let overrides = ScanOverrides::default(); + let targeted = targeted_pass_overrides(&overrides, false, Some(vec!["a".to_owned()])); + assert_eq!(targeted.touched_repositories, vec!["a".to_owned()]); + let refreshed = targeted_pass_overrides(&overrides, true, Some(vec!["a".to_owned()])); + assert_eq!(refreshed.touched_repositories, Vec::::new()); + let explicit = ScanOverrides { + repository: Some("b".to_owned()), + ..ScanOverrides::default() + }; + let explicit = targeted_pass_overrides(&explicit, false, Some(vec!["a".to_owned()])); + assert_eq!(explicit.touched_repositories, Vec::::new()); + } + #[test] fn recursive_parent_should_cover_nested_watch_roots() { let mut entries = vec![(PathBuf::from("/workspace"), RecursiveMode::NonRecursive)]; diff --git a/crates/code-system-graph-cli/src/telemetry.rs b/crates/code-system-graph-cli/src/telemetry.rs new file mode 100644 index 0000000..fe08eb8 --- /dev/null +++ b/crates/code-system-graph-cli/src/telemetry.rs @@ -0,0 +1,67 @@ +//! Per-phase wall time and resident memory for one scan pass. + +use std::time::Instant; + +use code_system_graph_core::{JobPhase, PhaseTelemetry}; +use sysinfo::{Pid, ProcessesToUpdate, System}; + +/// Records consecutive phases; each phase lasts from the previous mark to its own completion. +pub(crate) struct PhaseRecorder { + last_mark: Instant, + system: System, + phases: Vec, +} + +impl PhaseRecorder { + pub(crate) fn start() -> Self { + Self { + last_mark: Instant::now(), + system: System::new(), + phases: Vec::new(), + } + } + + pub(crate) fn complete(&mut self, phase: JobPhase) { + let now = Instant::now(); + let duration_ms = + u64::try_from(now.duration_since(self.last_mark).as_millis()).unwrap_or(u64::MAX); + self.last_mark = now; + let resident_memory_bytes = self.resident_memory_bytes(); + self.phases.push(PhaseTelemetry { + phase, + duration_ms, + resident_memory_bytes, + }); + } + + pub(crate) fn into_phases(self) -> Vec { + self.phases + } + + fn resident_memory_bytes(&mut self) -> u64 { + let pid = Pid::from_u32(std::process::id()); + self.system + .refresh_processes(ProcessesToUpdate::Some(&[pid]), true); + self.system.process(pid).map_or(0, sysinfo::Process::memory) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn phase_recorder_should_keep_completion_order() { + let mut recorder = PhaseRecorder::start(); + recorder.complete(JobPhase::Discovery); + recorder.complete(JobPhase::Extraction); + + let phases = recorder.into_phases(); + + assert_eq!( + phases.iter().map(|phase| phase.phase).collect::>(), + vec![JobPhase::Discovery, JobPhase::Extraction] + ); + assert!(phases.iter().any(|phase| phase.resident_memory_bytes > 0)); + } +} diff --git a/crates/code-system-graph-cli/src/work_state.rs b/crates/code-system-graph-cli/src/work_state.rs index 8028db9..e439e5b 100644 --- a/crates/code-system-graph-cli/src/work_state.rs +++ b/crates/code-system-graph-cli/src/work_state.rs @@ -1,20 +1,15 @@ //! Disposable operational state for resumable scans and finite watchers. +use std::collections::HashMap; use std::fs::{self, OpenOptions}; use std::path::{Path, PathBuf}; -use code_system_graph_core::ExecutionSummary; -use code_system_graph_model::{ - ArtifactFingerprint, CommunitySnapshot, Edge, Evidence, ExtractorRun, LinkDecision, Node, NodeId, RepositoryCoverageGap, StoredExtractorBatch, stable_id_bytes -}; -use code_system_graph_store_sqlite::{ - ManualLinkDisposition, ManualLinkRecord, set_owner_only_file -}; +use code_system_graph_model::{ArtifactFingerprint, StoredExtractorBatch, stable_id_bytes}; +use code_system_graph_store_sqlite::set_owner_only_file; use rusqlite::{Connection, ErrorCode, OpenFlags, OptionalExtension, TransactionBehavior, params}; -use serde::de::DeserializeOwned; use serde::{Deserialize, Serialize}; -const WORK_SCHEMA_VERSION: &str = "1.1.0"; +const WORK_SCHEMA_VERSION: &str = "1.2.0"; const SQLITE_ARTIFACT_SUFFIXES: [&str; 4] = ["", "-wal", "-shm", "-journal"]; enum WorkOpenError { @@ -35,48 +30,33 @@ pub(crate) struct WatcherLease { pub(crate) detail: Option, } -#[derive(Debug, Clone, Serialize, Deserialize)] -struct CandidateMetadata { - snapshot_id: String, - community_delta_count: usize, - corroborated_symbol_count: usize, - affected_test_count: usize, - execution: ExecutionSummary, - degradations: Vec, - #[serde(default)] - coverage_gaps: Vec, +/// Identity of one file in the stat cache: checkout id and canonical relative path bytes. +pub(crate) type FileStatKey = (String, Vec); + +/// Filesystem metadata that lets a scan reuse a content hash without reading the file. +#[derive(Debug, Clone, PartialEq, Eq)] +pub(crate) struct FileStat { + pub(crate) size_bytes: u64, + pub(crate) modified_unix_ns: i64, + pub(crate) file_identity: String, + pub(crate) content_hash: String, } -#[derive(Debug, Clone, Serialize, Deserialize)] -struct StagedManualLinkRecord { - id: String, - snapshot_id: String, - source_node_id: NodeId, - target_node_id: NodeId, - kind: String, - disposition: String, - reason: String, - decision: LinkDecision, - config_version: u32, +/// Stat-cache rows to write and remove after one fingerprinting pass. +#[derive(Debug, Default)] +pub(crate) struct FileStatDelta { + pub(crate) upserts: Vec<(FileStatKey, FileStat)>, + pub(crate) removals: Vec, } -#[derive(Debug, Clone)] -pub(crate) struct StagedSnapshot { - pub(crate) snapshot_id: String, - pub(crate) nodes: Vec, - pub(crate) edges: Vec, - pub(crate) evidence: Vec, - pub(crate) fingerprints: Vec, - pub(crate) extractor_batches: Vec, - pub(crate) extractor_runs: Vec, - pub(crate) manual_links: Vec, - pub(crate) community_snapshot: CommunitySnapshot, - pub(crate) community_delta_count: usize, - pub(crate) corroborated_symbol_count: usize, - pub(crate) affected_test_count: usize, - pub(crate) execution: ExecutionSummary, - pub(crate) degradations: Vec, - pub(crate) coverage_gaps: Vec, +/// Batch metadata stored beside the raw payload BLOB. +#[derive(Debug, Clone, Serialize, Deserialize)] +struct CachedBatchHeader { + source: ArtifactFingerprint, + extractor_version: String, + budget_fingerprint: String, + source_was_lossy: bool, + output_count: u64, } impl WorkState { @@ -118,26 +98,21 @@ impl WorkState { ); CREATE TABLE IF NOT EXISTS batch_cache ( cache_key TEXT PRIMARY KEY, - batch_json BLOB NOT NULL, + header_json BLOB NOT NULL, + payload BLOB NOT NULL, size_bytes INTEGER NOT NULL CHECK(size_bytes > 0), last_access_unix_ms INTEGER NOT NULL ); CREATE INDEX IF NOT EXISTS idx_batch_cache_lru ON batch_cache(last_access_unix_ms, cache_key); - CREATE TABLE IF NOT EXISTS active_candidate ( - workspace TEXT PRIMARY KEY, - compatibility_fingerprint TEXT NOT NULL, - phase TEXT NOT NULL, - started_at_unix_ms INTEGER NOT NULL, - updated_at_unix_ms INTEGER NOT NULL - ); - CREATE TABLE IF NOT EXISTS candidate_items ( - workspace TEXT NOT NULL, - compatibility_fingerprint TEXT NOT NULL, - item_kind TEXT NOT NULL, - item_key TEXT NOT NULL, - item_json BLOB NOT NULL, - PRIMARY KEY(workspace, item_kind, item_key) + CREATE TABLE IF NOT EXISTS file_stats ( + checkout_id TEXT NOT NULL, + path BLOB NOT NULL, + size_bytes INTEGER NOT NULL, + modified_unix_ns INTEGER NOT NULL, + file_identity TEXT NOT NULL, + content_hash TEXT NOT NULL, + PRIMARY KEY(checkout_id, path) ); CREATE TABLE IF NOT EXISTS watcher_lease ( workspace TEXT PRIMARY KEY, @@ -201,424 +176,217 @@ impl WorkState { Ok(Self { connection }) } - pub(crate) fn begin_candidate( - &mut self, - workspace: &str, - compatibility_fingerprint: &str, - now_unix_ms: u64, - ) -> Result { - let now_unix_ms = sqlite_integer(now_unix_ms, "candidate timestamp")?; - let previous = self - .connection - .query_row( - "SELECT compatibility_fingerprint FROM active_candidate WHERE workspace = ?1", - [workspace], - |row| row.get::<_, String>(0), - ) - .optional() - .map_err(|error| error.to_string())?; - let resumed = previous.as_deref() == Some(compatibility_fingerprint); - if !resumed { - self.connection - .execute( - "DELETE FROM candidate_items WHERE workspace = ?1", - [workspace], - ) - .map_err(|error| error.to_string())?; - } - self.connection - .execute( - "INSERT INTO active_candidate( - workspace, compatibility_fingerprint, phase, started_at_unix_ms, - updated_at_unix_ms - ) VALUES (?1, ?2, 'fingerprinted', ?3, ?3) - ON CONFLICT(workspace) DO UPDATE SET - compatibility_fingerprint = excluded.compatibility_fingerprint, - phase = excluded.phase, - started_at_unix_ms = CASE - WHEN active_candidate.compatibility_fingerprint = excluded.compatibility_fingerprint - THEN active_candidate.started_at_unix_ms - ELSE excluded.started_at_unix_ms - END, - updated_at_unix_ms = excluded.updated_at_unix_ms", - params![workspace, compatibility_fingerprint, now_unix_ms], - ) - .map_err(|error| error.to_string())?; - Ok(resumed) - } - - pub(crate) fn set_candidate_phase( - &self, - workspace: &str, - phase: &str, - now_unix_ms: u64, - ) -> Result<(), String> { - let now_unix_ms = sqlite_integer(now_unix_ms, "candidate timestamp")?; - self.connection - .execute( - "UPDATE active_candidate SET phase = ?2, updated_at_unix_ms = ?3 - WHERE workspace = ?1", - params![workspace, phase, now_unix_ms], - ) - .map_err(|error| error.to_string())?; - Ok(()) - } - - pub(crate) fn complete_candidate(&self, workspace: &str) -> Result<(), String> { - self.connection - .execute( - "DELETE FROM candidate_items WHERE workspace = ?1", - [workspace], - ) - .map_err(|error| error.to_string())?; - self.connection - .execute( - "DELETE FROM active_candidate WHERE workspace = ?1", - [workspace], - ) - .map_err(|error| error.to_string())?; - Ok(()) - } - - #[expect( - clippy::too_many_lines, - reason = "the candidate boundary records every immutable snapshot collection explicitly" - )] - pub(crate) fn store_candidate_snapshot( - &mut self, - workspace: &str, - compatibility_fingerprint: &str, - snapshot: &StagedSnapshot, - now_unix_ms: u64, - ) -> Result<(), String> { - let current = self - .connection - .query_row( - "SELECT compatibility_fingerprint FROM active_candidate WHERE workspace = ?1", - [workspace], - |row| row.get::<_, String>(0), - ) - .optional() - .map_err(|error| error.to_string())?; - if current.as_deref() != Some(compatibility_fingerprint) { - return Err("active candidate fingerprint changed before staging".to_owned()); - } - let transaction = self - .connection - .transaction() - .map_err(|error| error.to_string())?; - transaction - .execute( - "DELETE FROM candidate_items WHERE workspace = ?1", - [workspace], - ) - .map_err(|error| error.to_string())?; - let metadata = CandidateMetadata { - snapshot_id: snapshot.snapshot_id.clone(), - community_delta_count: snapshot.community_delta_count, - corroborated_symbol_count: snapshot.corroborated_symbol_count, - affected_test_count: snapshot.affected_test_count, - execution: snapshot.execution.clone(), - degradations: snapshot.degradations.clone(), - coverage_gaps: snapshot.coverage_gaps.clone(), - }; - insert_candidate_item( - &transaction, - workspace, - compatibility_fingerprint, - "metadata", - "snapshot-v1", - &metadata, - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "node", - &snapshot.nodes, - |node| node.id.as_str().to_owned(), - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "edge", - &snapshot.edges, - |edge| edge.id.as_str().to_owned(), - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "evidence", - &snapshot.evidence, - |evidence| evidence.id.as_str().to_owned(), - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "fingerprint", - &snapshot.fingerprints, - candidate_value_key, - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "extractor_batch", - &snapshot.extractor_batches, - candidate_value_key, - )?; - insert_candidate_values( - &transaction, - workspace, - compatibility_fingerprint, - "extractor_run", - &snapshot.extractor_runs, - |run| run.id.clone(), - )?; - for record in &snapshot.manual_links { - let staged = StagedManualLinkRecord::from(record); - insert_candidate_item( - &transaction, - workspace, - compatibility_fingerprint, - "manual_link", - &record.id, - &staged, - )?; - } - insert_candidate_item( - &transaction, - workspace, - compatibility_fingerprint, - "community", - "snapshot-v1", - &snapshot.community_snapshot, - )?; - let now = sqlite_integer(now_unix_ms, "candidate timestamp")?; - transaction - .execute( - "UPDATE active_candidate SET phase = 'ready', updated_at_unix_ms = ?2 - WHERE workspace = ?1 AND compatibility_fingerprint = ?3", - params![workspace, now, compatibility_fingerprint], - ) - .map_err(|error| error.to_string())?; - transaction.commit().map_err(|error| error.to_string()) - } - - pub(crate) fn load_candidate_snapshot( - &self, - workspace: &str, - compatibility_fingerprint: &str, - ) -> Result, String> { - let phase = self - .connection - .query_row( - "SELECT phase FROM active_candidate - WHERE workspace = ?1 AND compatibility_fingerprint = ?2", - params![workspace, compatibility_fingerprint], - |row| row.get::<_, String>(0), - ) - .optional() - .map_err(|error| error.to_string())?; - if phase.as_deref() != Some("ready") { - return Ok(None); - } - let metadata = load_single_candidate_value::( - &self.connection, - workspace, - compatibility_fingerprint, - "metadata", - )?; - let Some(metadata) = metadata else { - return Ok(None); - }; - let manual_links = load_candidate_values::( - &self.connection, - workspace, - compatibility_fingerprint, - "manual_link", - )? - .into_iter() - .map(ManualLinkRecord::try_from) - .collect::, _>>()?; - let Some(community_snapshot) = load_single_candidate_value( - &self.connection, - workspace, - compatibility_fingerprint, - "community", - )? - else { - return Ok(None); - }; - Ok(Some(StagedSnapshot { - snapshot_id: metadata.snapshot_id, - nodes: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "node", - )?, - edges: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "edge", - )?, - evidence: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "evidence", - )?, - fingerprints: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "fingerprint", - )?, - extractor_batches: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "extractor_batch", - )?, - extractor_runs: load_candidate_values( - &self.connection, - workspace, - compatibility_fingerprint, - "extractor_run", - )?, - manual_links, - community_snapshot, - community_delta_count: metadata.community_delta_count, - corroborated_symbol_count: metadata.corroborated_symbol_count, - affected_test_count: metadata.affected_test_count, - execution: metadata.execution, - degradations: metadata.degradations, - coverage_gaps: metadata.coverage_gaps, - })) - } - pub(crate) fn load_batches( &mut self, - fingerprints: &[ArtifactFingerprint], + fingerprints: &[&ArtifactFingerprint], budget_fingerprint: &str, extractor_version: &str, maximum_payload_bytes: u64, now_unix_ms: u64, ) -> Result, String> { let now_unix_ms = sqlite_integer(now_unix_ms, "cache timestamp")?; - // `Vec` is represented as comma-separated JSON integers. Four encoded bytes per - // payload byte plus bounded metadata is a conservative ceiling that lets SQLite reject - // an oversized cache row before copying its BLOB into this process. - let maximum_encoded_bytes = maximum_payload_bytes - .saturating_mul(4) - .saturating_add(1_048_576) - .min(i64::MAX as u64); - let maximum_encoded_bytes = sqlite_integer(maximum_encoded_bytes, "cache batch maximum")?; + let maximum_payload_bytes = sqlite_integer( + maximum_payload_bytes.min(i64::MAX as u64), + "cache batch maximum", + )?; let transaction = self .connection .transaction() .map_err(|error| error.to_string())?; let mut batches = Vec::new(); - for fingerprint in fingerprints { - let key = batch_cache_key(fingerprint, budget_fingerprint, extractor_version)?; - let encoded = transaction - .query_row( - "SELECT CASE WHEN length(batch_json) <= ?2 THEN batch_json END + { + let mut select = transaction + .prepare_cached( + "SELECT header_json, + CASE WHEN length(payload) <= ?2 THEN payload END FROM batch_cache WHERE cache_key = ?1", - params![key, maximum_encoded_bytes], - |row| row.get::<_, Option>>(0), ) - .optional() .map_err(|error| error.to_string())?; - let Some(Some(encoded)) = encoded else { - if encoded.is_some() { - transaction - .execute("DELETE FROM batch_cache WHERE cache_key = ?1", [&key]) + let mut touch = transaction + .prepare_cached( + "UPDATE batch_cache SET last_access_unix_ms = ?2 WHERE cache_key = ?1", + ) + .map_err(|error| error.to_string())?; + let mut remove = transaction + .prepare_cached("DELETE FROM batch_cache WHERE cache_key = ?1") + .map_err(|error| error.to_string())?; + for fingerprint in fingerprints { + let key = batch_cache_key(fingerprint, budget_fingerprint, extractor_version)?; + let row = select + .query_row(params![key, maximum_payload_bytes], |row| { + Ok((row.get::<_, Vec>(0)?, row.get::<_, Option>>(1)?)) + }) + .optional() + .map_err(|error| error.to_string())?; + let Some((header, payload)) = row else { + continue; + }; + let decoded = payload.and_then(|payload| { + serde_json::from_slice::(&header) + .ok() + .map(|header| (header, payload)) + }); + let Some((header, payload)) = decoded else { + remove.execute([&key]).map_err(|error| error.to_string())?; + continue; + }; + if header.source == **fingerprint + && header.budget_fingerprint == budget_fingerprint + && header.extractor_version == extractor_version + { + touch + .execute(params![key, now_unix_ms]) .map_err(|error| error.to_string())?; + batches.push(StoredExtractorBatch { + source: header.source, + extractor_version: header.extractor_version, + budget_fingerprint: header.budget_fingerprint, + source_was_lossy: header.source_was_lossy, + output_count: header.output_count, + payload, + }); } - continue; - }; - let batch: StoredExtractorBatch = if let Ok(batch) = serde_json::from_slice(&encoded) { - batch - } else { - transaction - .execute("DELETE FROM batch_cache WHERE cache_key = ?1", [&key]) - .map_err(|error| error.to_string())?; - continue; - }; - if batch.source == *fingerprint - && batch.budget_fingerprint == budget_fingerprint - && batch.extractor_version == extractor_version - { - transaction - .execute( - "UPDATE batch_cache SET last_access_unix_ms = ?2 WHERE cache_key = ?1", - params![key, now_unix_ms], - ) - .map_err(|error| error.to_string())?; - batches.push(batch); } } transaction.commit().map_err(|error| error.to_string())?; Ok(batches) } - pub(crate) fn put_batch( + /// Checkpoints completed batches in one transaction and returns how many were retained. + pub(crate) fn put_batches( &mut self, - batch: &StoredExtractorBatch, + batches: &[&StoredExtractorBatch], quota_bytes: u64, now_unix_ms: u64, - ) -> Result { - let encoded = serde_json::to_vec(batch).map_err(|error| error.to_string())?; - let size = u64::try_from(encoded.len()).unwrap_or(u64::MAX); - if size == 0 || size > quota_bytes || size > i64::MAX as u64 { - return Ok(false); + ) -> Result { + if batches.is_empty() { + return Ok(0); } - let size = sqlite_integer(size, "cache entry size")?; let now_unix_ms = sqlite_integer(now_unix_ms, "cache timestamp")?; - let key = batch_cache_key( - &batch.source, - &batch.budget_fingerprint, - &batch.extractor_version, - )?; let transaction = self .connection .transaction() .map_err(|error| error.to_string())?; - transaction - .execute( - "INSERT INTO batch_cache(cache_key, batch_json, size_bytes, last_access_unix_ms) - VALUES (?1, ?2, ?3, ?4) - ON CONFLICT(cache_key) DO UPDATE SET - batch_json = excluded.batch_json, - size_bytes = excluded.size_bytes, - last_access_unix_ms = excluded.last_access_unix_ms", - params![key, encoded, size, now_unix_ms], + let mut written = 0_u64; + { + let mut insert = transaction + .prepare_cached( + "INSERT INTO batch_cache( + cache_key, header_json, payload, size_bytes, last_access_unix_ms + ) VALUES (?1, ?2, ?3, ?4, ?5) + ON CONFLICT(cache_key) DO UPDATE SET + header_json = excluded.header_json, + payload = excluded.payload, + size_bytes = excluded.size_bytes, + last_access_unix_ms = excluded.last_access_unix_ms", + ) + .map_err(|error| error.to_string())?; + for batch in batches { + let header = serde_json::to_vec(&CachedBatchHeader { + source: batch.source.clone(), + extractor_version: batch.extractor_version.clone(), + budget_fingerprint: batch.budget_fingerprint.clone(), + source_was_lossy: batch.source_was_lossy, + output_count: batch.output_count, + }) + .map_err(|error| error.to_string())?; + let size = u64::try_from(header.len().saturating_add(batch.payload.len())) + .unwrap_or(u64::MAX); + if size > quota_bytes || size > i64::MAX as u64 { + continue; + } + let key = batch_cache_key( + &batch.source, + &batch.budget_fingerprint, + &batch.extractor_version, + )?; + insert + .execute(params![ + key, + header, + batch.payload, + sqlite_integer(size, "cache entry size")?, + now_unix_ms + ]) + .map_err(|error| error.to_string())?; + written = written.saturating_add(1); + } + } + evict_to_quota(&transaction, quota_bytes)?; + transaction.commit().map_err(|error| error.to_string())?; + Ok(written) + } + + pub(crate) fn load_file_stats(&self) -> Result, String> { + let mut statement = self + .connection + .prepare( + "SELECT checkout_id, path, size_bytes, modified_unix_ns, file_identity, + content_hash + FROM file_stats", ) .map_err(|error| error.to_string())?; - let mut total = cache_size(&transaction)?; - while total > quota_bytes { - let removed = transaction - .execute( - "DELETE FROM batch_cache WHERE cache_key = ( - SELECT cache_key FROM batch_cache - ORDER BY last_access_unix_ms, cache_key LIMIT 1 - )", - [], + let rows = statement + .query_map([], |row| { + Ok(( + (row.get::<_, String>(0)?, row.get::<_, Vec>(1)?), + FileStat { + size_bytes: u64::try_from(row.get::<_, i64>(2)?).unwrap_or(u64::MAX), + modified_unix_ns: row.get(3)?, + file_identity: row.get(4)?, + content_hash: row.get(5)?, + }, + )) + }) + .map_err(|error| error.to_string())?; + rows.collect::, _>>() + .map_err(|error| error.to_string()) + } + + pub(crate) fn apply_file_stat_delta(&mut self, delta: &FileStatDelta) -> Result<(), String> { + if delta.upserts.is_empty() && delta.removals.is_empty() { + return Ok(()); + } + let transaction = self + .connection + .transaction() + .map_err(|error| error.to_string())?; + { + let mut remove = transaction + .prepare_cached("DELETE FROM file_stats WHERE checkout_id = ?1 AND path = ?2") + .map_err(|error| error.to_string())?; + for (checkout_id, path) in &delta.removals { + remove + .execute(params![checkout_id, path]) + .map_err(|error| error.to_string())?; + } + let mut upsert = transaction + .prepare_cached( + "INSERT INTO file_stats( + checkout_id, path, size_bytes, modified_unix_ns, file_identity, + content_hash + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6) + ON CONFLICT(checkout_id, path) DO UPDATE SET + size_bytes = excluded.size_bytes, + modified_unix_ns = excluded.modified_unix_ns, + file_identity = excluded.file_identity, + content_hash = excluded.content_hash", ) .map_err(|error| error.to_string())?; - if removed == 0 { - break; + for ((checkout_id, path), stat) in &delta.upserts { + upsert + .execute(params![ + checkout_id, + path, + sqlite_integer(stat.size_bytes, "file size")?, + stat.modified_unix_ns, + stat.file_identity, + stat.content_hash + ]) + .map_err(|error| error.to_string())?; } - total = cache_size(&transaction)?; } - transaction.commit().map_err(|error| error.to_string())?; - Ok(true) + transaction.commit().map_err(|error| error.to_string()) } pub(crate) fn start_watcher( @@ -833,15 +601,42 @@ impl WorkState { } } -fn cache_size(transaction: &rusqlite::Transaction<'_>) -> Result { - let size = transaction +fn evict_to_quota(transaction: &rusqlite::Transaction<'_>, quota_bytes: u64) -> Result<(), String> { + let total = transaction .query_row( "SELECT COALESCE(SUM(size_bytes), 0) FROM batch_cache", [], |row| row.get::<_, i64>(0), ) .map_err(|error| error.to_string())?; - u64::try_from(size).map_err(|_| "negative work cache size".to_owned()) + let mut total = u64::try_from(total).map_err(|_| "negative work cache size".to_owned())?; + if total <= quota_bytes { + return Ok(()); + } + let mut oldest = transaction + .prepare( + "SELECT cache_key, size_bytes FROM batch_cache ORDER BY last_access_unix_ms, cache_key", + ) + .map_err(|error| error.to_string())?; + let mut victims = Vec::new(); + let mut rows = oldest.query([]).map_err(|error| error.to_string())?; + while total > quota_bytes { + let Some(row) = rows.next().map_err(|error| error.to_string())? else { + break; + }; + let key = row.get::<_, String>(0).map_err(|error| error.to_string())?; + let size = row.get::<_, i64>(1).map_err(|error| error.to_string())?; + total = total.saturating_sub(u64::try_from(size).unwrap_or(0)); + victims.push(key); + } + drop(rows); + let mut remove = transaction + .prepare_cached("DELETE FROM batch_cache WHERE cache_key = ?1") + .map_err(|error| error.to_string())?; + for key in victims { + remove.execute([key]).map_err(|error| error.to_string())?; + } + Ok(()) } fn sqlite_integer(value: u64, field: &str) -> Result { @@ -858,156 +653,9 @@ fn batch_cache_key( Ok(stable_id_bytes("work-batch-v1", &material)) } -fn candidate_value_key(value: &T) -> String { - serde_json::to_vec(value).map_or_else( - |_| "unencodable-candidate-value".to_owned(), - |encoded| stable_id_bytes("work-candidate-item-v1", &encoded), - ) -} - -fn insert_candidate_values( - transaction: &rusqlite::Transaction<'_>, - workspace: &str, - compatibility_fingerprint: &str, - item_kind: &str, - values: &[T], - key: F, -) -> Result<(), String> -where - T: Serialize, - F: Fn(&T) -> String, -{ - for value in values { - insert_candidate_item( - transaction, - workspace, - compatibility_fingerprint, - item_kind, - &key(value), - value, - )?; - } - Ok(()) -} - -fn insert_candidate_item( - transaction: &rusqlite::Transaction<'_>, - workspace: &str, - compatibility_fingerprint: &str, - item_kind: &str, - item_key: &str, - value: &T, -) -> Result<(), String> { - let encoded = serde_json::to_vec(value).map_err(|error| error.to_string())?; - transaction - .execute( - "INSERT INTO candidate_items( - workspace, compatibility_fingerprint, item_kind, item_key, item_json - ) VALUES (?1, ?2, ?3, ?4, ?5) - ON CONFLICT(workspace, item_kind, item_key) DO UPDATE SET - compatibility_fingerprint = excluded.compatibility_fingerprint, - item_json = excluded.item_json", - params![ - workspace, - compatibility_fingerprint, - item_kind, - item_key, - encoded - ], - ) - .map_err(|error| error.to_string())?; - Ok(()) -} - -fn load_candidate_values( - connection: &Connection, - workspace: &str, - compatibility_fingerprint: &str, - item_kind: &str, -) -> Result, String> { - let mut statement = connection - .prepare( - "SELECT item_json FROM candidate_items - WHERE workspace = ?1 AND compatibility_fingerprint = ?2 AND item_kind = ?3 - ORDER BY item_key", - ) - .map_err(|error| error.to_string())?; - let encoded = statement - .query_map( - params![workspace, compatibility_fingerprint, item_kind], - |row| row.get::<_, Vec>(0), - ) - .map_err(|error| error.to_string())? - .collect::, _>>() - .map_err(|error| error.to_string())?; - encoded - .into_iter() - .map(|value| serde_json::from_slice(&value).map_err(|error| error.to_string())) - .collect() -} - -fn load_single_candidate_value( - connection: &Connection, - workspace: &str, - compatibility_fingerprint: &str, - item_kind: &str, -) -> Result, String> { - let mut values = - load_candidate_values(connection, workspace, compatibility_fingerprint, item_kind)?; - match values.len() { - 0 => Ok(None), - 1 => Ok(values.pop()), - _ => Err(format!( - "candidate has multiple `{item_kind}` singleton rows" - )), - } -} - -impl From<&ManualLinkRecord> for StagedManualLinkRecord { - fn from(value: &ManualLinkRecord) -> Self { - Self { - id: value.id.clone(), - snapshot_id: value.snapshot_id.clone(), - source_node_id: value.source_node_id.clone(), - target_node_id: value.target_node_id.clone(), - kind: value.kind.clone(), - disposition: match value.disposition { - ManualLinkDisposition::Active => "active".to_owned(), - ManualLinkDisposition::Suppression => "suppression".to_owned(), - }, - reason: value.reason.clone(), - decision: value.decision.clone(), - config_version: value.config_version, - } - } -} - -impl TryFrom for ManualLinkRecord { - type Error = String; - - fn try_from(value: StagedManualLinkRecord) -> Result { - let disposition = match value.disposition.as_str() { - "active" => ManualLinkDisposition::Active, - "suppression" => ManualLinkDisposition::Suppression, - other => return Err(format!("invalid staged manual-link disposition `{other}`")), - }; - Ok(Self { - id: value.id, - snapshot_id: value.snapshot_id, - source_node_id: value.source_node_id, - target_node_id: value.target_node_id, - kind: value.kind, - disposition, - reason: value.reason, - decision: value.decision, - config_version: value.config_version, - }) - } -} - pub(crate) fn work_path(database: &Path) -> PathBuf { let mut value = database.as_os_str().to_os_string(); - value.push(".work-v1.db"); + value.push(".work.db"); PathBuf::from(value) } @@ -1015,6 +663,11 @@ fn canonicalize_parent(path: &Path) -> Result { let parent = path .parent() .ok_or_else(|| format!("work sidecar path `{}` has no parent", path.display()))?; + let parent = if parent.as_os_str().is_empty() { + Path::new(".") + } else { + parent + }; let file_name = path .file_name() .ok_or_else(|| format!("work sidecar path `{}` has no file name", path.display()))?; @@ -1044,7 +697,10 @@ fn classify_existing_sidecar_error(error: &rusqlite::Error) -> WorkOpenError { } fn ensure_private_file(path: &Path) -> Result<(), String> { - if let Some(parent) = path.parent() { + if let Some(parent) = path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + { fs::create_dir_all(parent).map_err(|error| error.to_string())?; } if path.exists() { @@ -1116,9 +772,7 @@ fn sqlite_artifact_path(path: &Path, suffix: &str) -> PathBuf { #[cfg(test)] mod tests { - use code_system_graph_model::{ - CheckoutId, CommunityAlgorithm, CommunityConfig, CommunityScope, NativePath, NativePathEncoding, RepoId - }; + use code_system_graph_model::{CheckoutId, NativePath, NativePathEncoding, RepoId}; use super::*; @@ -1137,7 +791,7 @@ mod tests { let state = WorkState::open(&database, "database-instance").expect("work state"); drop(state); - assert!(canonical_parent.join("graph.db.work-v1.db").is_file()); + assert!(canonical_parent.join("graph.db.work.db").is_file()); } fn fingerprint(hash: &str) -> ArtifactFingerprint { @@ -1168,21 +822,43 @@ mod tests { output_count: 1, payload: b"[{}]".to_vec(), }; - assert!(state.put_batch(&batch, 1_000_000, 1).expect("cache")); + assert_eq!( + state.put_batches(&[&batch], 1_000_000, 1).expect("cache"), + 1 + ); assert_eq!( state - .load_batches(&[fingerprint("one")], "budget", "1.0.0", 1_000_000, 2) + .load_batches(&[&fingerprint("one")], "budget", "1.0.0", 1_000_000, 2) .expect("load"), vec![batch.clone()] ); assert_eq!( state - .load_batches(&[fingerprint("two")], "budget", "1.0.0", 1_000_000, 3) + .load_batches(&[&fingerprint("two")], "budget", "1.0.0", 1_000_000, 3) .expect("load") .as_slice(), &[] ); - assert!(!state.put_batch(&batch, 1, 4).expect("oversized skip")); + assert_eq!( + state.put_batches(&[&batch], 1, 4).expect("oversized skip"), + 0 + ); + + let newer = StoredExtractorBatch { + source: fingerprint("two"), + ..batch.clone() + }; + let quota = 2 + * (serde_json::to_vec(&fingerprint("one")) + .expect("encode") + .len() as u64); + assert_eq!(state.put_batches(&[&newer], quota, 5).expect("evict"), 1); + assert_eq!( + state + .load_batches(&[&fingerprint("one")], "budget", "1.0.0", 1_000_000, 6) + .expect("evicted load"), + Vec::new() + ); } #[test] @@ -1204,7 +880,7 @@ mod tests { #[test] fn sidecar_cleanup_should_remove_every_sqlite_artifact() { let temporary = tempfile::tempdir().expect("temporary directory"); - let sidecar = temporary.path().join("graph.db.work-v1.db"); + let sidecar = temporary.path().join("graph.db.work.db"); for suffix in SQLITE_ARTIFACT_SUFFIXES { fs::write(sqlite_artifact_path(&sidecar, suffix), b"stale") .expect("create stale SQLite artifact"); @@ -1441,63 +1117,33 @@ mod tests { } #[test] - fn ready_candidate_should_round_trip_and_be_removed_after_publication() { + fn file_stat_delta_should_round_trip_and_remove_stale_rows() { let temporary = tempfile::tempdir().expect("temporary directory"); let mut state = WorkState::open(&temporary.path().join("graph.db"), "database-instance") .expect("work state"); - assert!( - !state - .begin_candidate("workspace", "compatible", 1) - .expect("begin") - ); - let snapshot = StagedSnapshot { - snapshot_id: "snapshot:one".to_owned(), - nodes: Vec::new(), - edges: Vec::new(), - evidence: Vec::new(), - fingerprints: vec![fingerprint("one")], - extractor_batches: Vec::new(), - extractor_runs: Vec::new(), - manual_links: Vec::new(), - community_snapshot: CommunitySnapshot { - snapshot_id: "snapshot:one".to_owned(), - engine_version: "1.0.0".to_owned(), - config: CommunityConfig { - algorithm: CommunityAlgorithm::Louvain, - scope: CommunityScope::Workspace, - seed: 0, - resolution: 1.0, - minimum_confidence: 0.0, - edge_weights: Vec::new(), - max_iterations: 1, - }, - communities: Vec::new(), - }, - community_delta_count: 0, - corroborated_symbol_count: 0, - affected_test_count: 0, - execution: ExecutionSummary { - checkpoints_written: 1, - ..ExecutionSummary::default() - }, - degradations: Vec::new(), - coverage_gaps: Vec::new(), + let stat = FileStat { + size_bytes: 4, + modified_unix_ns: 1_000, + file_identity: "1:2".to_owned(), + content_hash: "hash".to_owned(), }; + let key = ("checkout:api".to_owned(), b"src/main.rs".to_vec()); state - .store_candidate_snapshot("workspace", "compatible", &snapshot, 2) - .expect("stage"); - let loaded = state - .load_candidate_snapshot("workspace", "compatible") - .expect("load") - .expect("ready candidate"); - assert_eq!(loaded.snapshot_id, snapshot.snapshot_id); - assert_eq!(loaded.fingerprints, snapshot.fingerprints); - state.complete_candidate("workspace").expect("complete"); - assert!( - state - .load_candidate_snapshot("workspace", "compatible") - .expect("load after completion") - .is_none() + .apply_file_stat_delta(&FileStatDelta { + upserts: vec![(key.clone(), stat.clone())], + removals: Vec::new(), + }) + .expect("insert"); + assert_eq!( + state.load_file_stats().expect("load").get(&key), + Some(&stat) ); + state + .apply_file_stat_delta(&FileStatDelta { + upserts: Vec::new(), + removals: vec![key], + }) + .expect("remove"); + assert!(state.load_file_stats().expect("reload").is_empty()); } } diff --git a/crates/code-system-graph-cli/src/worker.rs b/crates/code-system-graph-cli/src/worker.rs index e0f8a14..4f421f5 100644 --- a/crates/code-system-graph-cli/src/worker.rs +++ b/crates/code-system-graph-cli/src/worker.rs @@ -13,7 +13,7 @@ use code_system_graph_core::{ ExecutionLimitExceeded, ExecutionPolicy, ExecutionResource, ExecutionSummary, ExtractionLimitExceeded, GraphqlExtractionError, JobPhase, ScanJobTracker }; use serde::{Deserialize, Serialize}; -use sysinfo::{Pid, ProcessesToUpdate, System}; +use sysinfo::{Pid, ProcessRefreshKind, ProcessesToUpdate, System}; use super::{ ApplicationError, ScanOverrides, ScanSummary, application_exit_code, load_execution_policy @@ -26,12 +26,17 @@ const MAX_PROTOCOL_LINE_BYTES: usize = 1_048_576; const PROTOCOL_QUEUE_CAPACITY: usize = 256; const MAX_PROTOCOL_MESSAGES_PER_POLL: usize = 1_024; const POLL_INTERVAL: Duration = Duration::from_millis(100); +/// Progress messages carry the cumulative total, so units completed within this interval of the +/// previous message are coalesced into the next one. +const PROGRESS_WRITE_INTERVAL_MS: u64 = 50; static WORKER_PROTOCOL_ACTIVE: AtomicBool = AtomicBool::new(false); static WORKER_LIMIT_REPORTED: AtomicBool = AtomicBool::new(false); static COMPLETED_UNITS: AtomicU64 = AtomicU64::new(0); static RUN_COUNTER: AtomicU64 = AtomicU64::new(0); static WORKER_TRACKER: OnceLock> = OnceLock::new(); +static PROGRESS_CLOCK: OnceLock = OnceLock::new(); +static LAST_PROGRESS_WRITE_MS: AtomicU64 = AtomicU64::new(0); #[derive(Debug, Clone, Serialize, Deserialize)] struct WorkerEnvelope { @@ -70,11 +75,11 @@ enum WorkerMessage { }, ScanResult { schema_version: u8, - summary: ScanSummary, + summary: Box, }, SyncResult { schema_version: u8, - summary: SyncSummary, + summary: Box, }, Failure { schema_version: u8, @@ -109,8 +114,8 @@ enum ExpectedResult { #[derive(Debug)] enum SupervisedResult { - Scan(ScanSummary), - Sync(SyncSummary), + Scan(Box), + Sync(Box), } struct ProtocolReader { @@ -174,7 +179,7 @@ pub fn run_worker_from_stdio() -> Result<(), String> { } => match super::scan_workspace_direct(&config, &database, &overrides) { Ok(summary) => WorkerMessage::ScanResult { schema_version: PROTOCOL_VERSION, - summary, + summary: Box::new(summary), }, Err(error) => failure_message(error), }, @@ -186,11 +191,12 @@ pub fn run_worker_from_stdio() -> Result<(), String> { } => match sync_workspace_direct(&config, &database, &overrides, synchronize_codegraph) { Ok(summary) => WorkerMessage::SyncResult { schema_version: PROTOCOL_VERSION, - summary, + summary: Box::new(summary), }, Err(error) => failure_message(error), }, }; + flush_progress()?; write_message(&message) } @@ -276,14 +282,6 @@ fn safe_application_diagnostic(error: &ApplicationError) -> String { source_path, }) => format!("unsupported persisted-operation manifest in `{source_path}`"), ApplicationError::HttpExtraction(error) => error.to_string(), - ApplicationError::Link(code_system_graph_core::LinkError::AmbiguousProvider { - method, - path, - candidates, - }) => format!( - "ambiguous HTTP provider for {method} {path}; stable candidates: {}", - candidates.join(", ") - ), ApplicationError::Config(error) => safe_config_diagnostic(error), ApplicationError::Manifest(_) => "workspace manifest is invalid".to_owned(), ApplicationError::ManualLink(error) => { @@ -465,11 +463,35 @@ pub(crate) fn report_progress(phase: JobPhase, completed: u64) { current.checked_add(completed) }) .map_or(u64::MAX, |previous| previous.saturating_add(completed)); - let _ = write_message(&WorkerMessage::Progress { + if progress_write_due() { + let _ = write_message(&WorkerMessage::Progress { + schema_version: PROTOCOL_VERSION, + phase, + completed_units: total, + }); + } +} + +fn progress_write_due() -> bool { + let now = duration_millis(PROGRESS_CLOCK.get_or_init(Instant::now).elapsed()); + let last = LAST_PROGRESS_WRITE_MS.load(Ordering::Acquire); + now.saturating_sub(last) >= PROGRESS_WRITE_INTERVAL_MS + && LAST_PROGRESS_WRITE_MS + .compare_exchange(last, now, Ordering::AcqRel, Ordering::Acquire) + .is_ok() +} + +/// Reports the cumulative total that coalescing may still be holding back. +fn flush_progress() -> Result<(), String> { + let phase = WORKER_TRACKER + .get() + .and_then(|tracker| tracker.lock().ok().map(|tracker| tracker.phase())) + .unwrap_or(JobPhase::Configuration); + write_message(&WorkerMessage::Progress { schema_version: PROTOCOL_VERSION, phase, - completed_units: total, - }); + completed_units: COMPLETED_UNITS.load(Ordering::Acquire), + }) } pub(crate) fn check_time(phase: JobPhase) { @@ -522,7 +544,7 @@ pub(crate) fn supervise_scan( overrides: overrides.clone(), }; match supervise(config, &request, ExpectedResult::Scan, None, None)? { - SupervisedResult::Scan(summary) => Ok(summary), + SupervisedResult::Scan(summary) => Ok(*summary), SupervisedResult::Sync(_) => unreachable!("worker result kind was validated"), } } @@ -545,7 +567,7 @@ pub(crate) fn supervise_scan_with_executable( None, Some(executable), )? { - SupervisedResult::Scan(summary) => Ok(summary), + SupervisedResult::Scan(summary) => Ok(*summary), SupervisedResult::Sync(_) => unreachable!("worker result kind was validated"), } } @@ -563,7 +585,7 @@ pub(crate) fn supervise_sync( synchronize_codegraph, }; match supervise(config, &request, ExpectedResult::Sync, None, None)? { - SupervisedResult::Sync(summary) => Ok(summary), + SupervisedResult::Sync(summary) => Ok(*summary), SupervisedResult::Scan(_) => unreachable!("worker result kind was validated"), } } @@ -588,7 +610,7 @@ pub(crate) fn supervise_sync_with_executable( None, Some(executable), )? { - SupervisedResult::Sync(summary) => Ok(summary), + SupervisedResult::Sync(summary) => Ok(*summary), SupervisedResult::Scan(_) => unreachable!("worker result kind was validated"), } } @@ -613,7 +635,7 @@ pub(crate) fn supervise_sync_with_wall_time_cap( Some(wall_time_cap_ms), None, )? { - SupervisedResult::Sync(summary) => Ok(summary), + SupervisedResult::Sync(summary) => Ok(*summary), SupervisedResult::Scan(_) => unreachable!("worker result kind was validated"), } } @@ -653,7 +675,7 @@ fn supervise( explicit_executable.map_or_else(worker_executable, validate_worker_executable)?; let mut command = Command::new(executable); command - .arg("__worker-v1") + .arg("__worker") .stdin(Stdio::piped()) .stdout(Stdio::piped()) .stderr(Stdio::inherit()); @@ -853,7 +875,7 @@ fn monitor_worker( )); } - system.refresh_processes(ProcessesToUpdate::All, true); + system.refresh_processes_specifics(ProcessesToUpdate::All, true, memory_refresh()); let memory = process_tree_memory(&system, root_pid); peak_memory = peak_memory.max(memory); if memory > policy.max_worker_memory_bytes { @@ -939,6 +961,11 @@ fn copy_artifact_timings(source: &ExecutionSummary, target: &mut ExecutionSummar target.artifact_duration_p50_ms = source.artifact_duration_p50_ms; target.artifact_duration_p95_ms = source.artifact_duration_p95_ms; target.artifact_duration_p99_ms = source.artifact_duration_p99_ms; + target.stat_cache_hits = source.stat_cache_hits; + target.content_bytes_read = source.content_bytes_read; + target.published_rows = source.published_rows; + target.extraction_workers = source.extraction_workers; + target.phases.clone_from(&source.phases); } fn application_failure(failure: WorkerFailure) -> ApplicationError { @@ -1139,10 +1166,19 @@ fn read_bounded_line( } } +/// Threads are listed as processes that report the memory of their whole process, so only +/// processes are summed. +/// The sampler reads only memory and parent links, which every refresh includes; the default +/// refresh also lists every thread of every process and reads CPU, disk usage, and executables. +fn memory_refresh() -> ProcessRefreshKind { + ProcessRefreshKind::nothing().with_memory().without_tasks() +} + fn process_tree_memory(system: &System, root: Pid) -> u64 { system .processes() .iter() + .filter(|(_, process)| process.thread_kind().is_none()) .filter(|(pid, _)| **pid == root || is_descendant(system, **pid, root)) .map(|(_, process)| process.memory()) .fold(0_u64, u64::saturating_add) @@ -1171,7 +1207,11 @@ pub fn terminate_process_tree(root_process_id: u32) { let root = Pid::from_u32(root_process_id); let mut system = System::new(); for _ in 0..3 { - system.refresh_processes(ProcessesToUpdate::All, true); + system.refresh_processes_specifics( + ProcessesToUpdate::All, + true, + ProcessRefreshKind::nothing().without_tasks(), + ); let descendants = system .processes() .iter() @@ -1638,7 +1678,7 @@ mod tests { .expect("policy message"); let result_message = serde_json::to_string(&WorkerMessage::ScanResult { schema_version: PROTOCOL_VERSION, - summary: ScanSummary { + summary: Box::new(ScanSummary { execution: ExecutionSummary::default(), workspace: "test".to_owned(), snapshot_id: "snapshot".to_owned(), @@ -1654,7 +1694,7 @@ mod tests { affected_test_count: 0, degradation_count: 0, degradations: Vec::new(), - }, + }), }) .expect("result message"); let script = format!( diff --git a/crates/code-system-graph-cli/tests/agent_plugin_e2e.rs b/crates/code-system-graph-cli/tests/agent_plugin_e2e.rs index 87304e7..c5f738c 100644 --- a/crates/code-system-graph-cli/tests/agent_plugin_e2e.rs +++ b/crates/code-system-graph-cli/tests/agent_plugin_e2e.rs @@ -8,7 +8,7 @@ use code_system_graph::{ AgentPluginCreateMode, AgentPluginCreateReport, AgentPluginUninstallReport, scan_workspace }; use rmcp::ServiceExt; -use rmcp::model::{ClientCapabilities, ClientInfo, Implementation}; +use rmcp::model::{ClientCapabilities, ClientConfig, Implementation}; const PLUGIN_SCHEMA: &str = include_str!( "../../code-system-graph-hooks/agent-integration-template/agent-plugin/schemas/1.0.0/plugin.schema.json" @@ -279,7 +279,7 @@ fn plugin_create_should_render_official_structure_and_be_idempotent() -> anyhow: "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json" ); assert_eq!(plugin["name"], first.plugin_name); - assert_eq!(plugin["version"], "1.1.0"); + assert_eq!(plugin["version"], env!("CARGO_PKG_VERSION")); let mcp: serde_json::Value = serde_json::from_slice(&std::fs::read(output.join("mcp.json"))?)?; let mcp_schema: serde_json::Value = serde_json::from_str(MCP_SCHEMA)?; assert!(jsonschema::validator_for(&mcp_schema)?.is_valid(&mcp)); @@ -497,7 +497,7 @@ async fn generated_mcp_should_handshake_with_read_only_profile() -> anyhow::Resu .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("agent-plugin-e2e", env!("CARGO_PKG_VERSION")), ); @@ -583,7 +583,7 @@ async fn generated_codegraph_profile_should_expose_explore_without_admin_tools() .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("agent-plugin-codegraph-e2e", env!("CARGO_PKG_VERSION")), ); @@ -929,7 +929,6 @@ fn plugin_create_should_replace_only_an_owned_existing_plugin_binding() -> anyho ); let binding_path = base.join(".local/code-system-graph/mcp-binding.json"); let mut binding: serde_json::Value = serde_json::from_slice(&std::fs::read(&binding_path)?)?; - binding["generator"] = serde_json::json!("csgraph plugin compose"); binding["workspace"] = serde_json::json!("locally-edited"); binding .as_object_mut() @@ -996,7 +995,7 @@ async fn existing_plugin_binding_should_start_the_read_only_mcp() -> anyhow::Res .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("agent-plugin-binding-e2e", env!("CARGO_PKG_VERSION")), ); diff --git a/crates/code-system-graph-cli/tests/authorities_e2e.rs b/crates/code-system-graph-cli/tests/authorities_e2e.rs new file mode 100644 index 0000000..a35a5d8 --- /dev/null +++ b/crates/code-system-graph-cli/tests/authorities_e2e.rs @@ -0,0 +1,113 @@ +//! Acceptance tests for absolute-URL authority resolution across repositories. + +use std::collections::BTreeMap; +use std::path::Path; + +use code_system_graph::scan_workspace; +use code_system_graph_model::{EdgeKind, Node, RepoId}; +use code_system_graph_store_sqlite::SqliteStore; + +const SOURCES: [(&str, &str); 6] = [ + ( + "api/src/server.ts", + "import express from 'express';\nconst app = express();\nfunction getOrder(req, res) {}\nfunction health(req, res) {}\napp.get('/v1/orders/:id', getOrder);\napp.get('/v1/health', health);\n", + ), + ( + "api/docker-compose.yml", + "services:\n orders-api:\n build: .\n ports:\n - \"8080:8080\"\n", + ), + ( + "billing/src/server.ts", + "import express from 'express';\nconst app = express();\nfunction readOrder(req, res) {}\napp.get('/v1/orders/:orderId', readOrder);\n", + ), + ( + "web/src/client.ts", + "export async function load() {\n await fetch('http://orders-api:8080/v1/orders/42');\n await fetch('https://payments.example.com/v1/orders/42');\n}\n", + ), + ( + "tests/tests/test_health.py", + "import requests\n\ndef test_health():\n requests.get(\"http://localhost:3000/v1/health\")\n", + ), + ( + "tests/tests/test_billing.py", + "import requests\n\ndef test_billing_order():\n requests.get(\"http://billing.internal/v1/orders/7\")\n", + ), +]; + +const MANIFEST: &str = "version: 1\nname: authorities\nrepos:\n api:\n path: api\n billing:\n path: billing\n authorities: [billing.internal]\n web:\n path: web\n tests:\n path: tests\n"; + +fn write_workspace(root: &Path) -> anyhow::Result { + for (path, contents) in SOURCES { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(&path, contents)?; + } + let manifest = root.join("code-system-graph.yaml"); + std::fs::write(&manifest, MANIFEST)?; + Ok(manifest) +} + +fn alias(aliases: &BTreeMap, node: &Node) -> String { + node.repo_id + .as_ref() + .and_then(|repo| aliases.get(repo)) + .cloned() + .unwrap_or_default() +} + +#[test] +fn absolute_urls_should_resolve_through_declared_inferred_and_loopback_authorities() +-> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path())?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, edges) = store.load_current_graph("authorities")?; + let aliases = store + .load_workspace_registry("authorities")? + .repositories + .into_iter() + .map(|repository| (repository.id, repository.alias)) + .collect::>(); + let nodes = nodes + .iter() + .map(|node| (&node.id, node)) + .collect::>(); + let mut links = edges + .iter() + .filter(|edge| matches!(edge.kind, EdgeKind::CallsRemote | EdgeKind::Validates)) + .map(|edge| { + let source = nodes[&edge.source]; + let target = nodes[&edge.target]; + ( + edge.kind, + alias(&aliases, source), + format!("{} -> {}", source.label, target.label), + alias(&aliases, target), + ) + }) + .collect::>(); + links.sort(); + + let mut summary = links + .iter() + .map(|(kind, source, _, target)| format!("{kind:?} {source} -> {target}")) + .collect::>(); + summary.sort(); + assert_eq!( + summary, + vec![ + "CallsRemote tests -> api".to_owned(), + "CallsRemote tests -> billing".to_owned(), + "CallsRemote web -> api".to_owned(), + "Validates tests -> api".to_owned(), + "Validates tests -> billing".to_owned(), + ], + "{links:#?}" + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/client_flows_e2e.rs b/crates/code-system-graph-cli/tests/client_flows_e2e.rs new file mode 100644 index 0000000..5221d0e --- /dev/null +++ b/crates/code-system-graph-cli/tests/client_flows_e2e.rs @@ -0,0 +1,130 @@ +//! Acceptance tests for tests linked to endpoints they reach through wrappers, helpers, and +//! fixtures in other files. + +use std::collections::BTreeMap; +use std::path::Path; + +use code_system_graph::scan_workspace; +use code_system_graph_model::{EdgeKind, RepoId}; +use code_system_graph_store_sqlite::SqliteStore; + +const SOURCES: [(&str, &str); 5] = [ + ( + "payments/src/main/java/com/shop/PaymentsController.java", + r#"package com.shop; + +import org.springframework.web.bind.annotation.*; + +@RestController +@RequestMapping("/payments") +public class PaymentsController { + @PostMapping("/{id}/refunds") + public Refund refund(@PathVariable String id) { + return null; + } +} +"#, + ), + ( + "catalog/src/routes.js", + r#"const express = require("express"); +const app = express(); +app.get("/products/:sku", getProduct); +function getProduct(req, res) {} +"#, + ), + ( + "tests/tests/conftest.py", + r#"import pytest +import requests + +CATALOG = "http://localhost:3000" +SKU = "sku-1" + +@pytest.fixture +def product(): + return requests.get(CATALOG + "/products/" + SKU).json() +"#, + ), + ( + "tests/tests/client.py", + r#"import requests + +PAYMENTS = "http://localhost:8080" + +def post_json(path, body): + return requests.post(PAYMENTS + path, json=body) + +def refund(payment_id): + return post_json(f"/payments/{payment_id}/refunds", {}) +"#, + ), + ( + "tests/tests/test_refunds.py", + r#"from tests.client import refund + +def test_refund(product): + refund(product["payment"]) +"#, + ), +]; + +const MANIFEST: &str = "version: 1\nname: flows\nrepos:\n payments:\n path: payments\n catalog:\n path: catalog\n tests:\n path: tests\n"; + +fn write_workspace(root: &Path) -> anyhow::Result { + for (path, contents) in SOURCES { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(&path, contents)?; + } + let manifest = root.join("code-system-graph.yaml"); + std::fs::write(&manifest, MANIFEST)?; + Ok(manifest) +} + +#[test] +fn tests_should_link_to_endpoints_reached_through_wrappers_and_fixtures() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path())?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, edges) = store.load_current_graph("flows")?; + let aliases = store + .load_workspace_registry("flows")? + .repositories + .into_iter() + .map(|repository| (repository.id, repository.alias)) + .collect::>(); + let nodes = nodes + .iter() + .map(|node| (&node.id, node)) + .collect::>(); + let mut links = edges + .iter() + .filter(|edge| edge.kind == EdgeKind::Validates) + .map(|edge| { + let target = nodes[&edge.target]; + let repo = target + .repo_id + .as_ref() + .and_then(|repo| aliases.get(repo)) + .cloned() + .unwrap_or_default(); + format!("{} -> {repo}:{}", nodes[&edge.source].label, target.label) + }) + .collect::>(); + links.sort(); + + assert_eq!( + links, + [ + "python/pytest::test_refund -> catalog:GET /products/:sku", + "python/pytest::test_refund -> payments:POST /payments/{id}/refunds", + ], + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/cross_language_matrix_e2e.rs b/crates/code-system-graph-cli/tests/cross_language_matrix_e2e.rs new file mode 100644 index 0000000..fe096cf --- /dev/null +++ b/crates/code-system-graph-cli/tests/cross_language_matrix_e2e.rs @@ -0,0 +1,229 @@ +//! Acceptance tests for the cross-language matrix: tests in five languages call providers in +//! eleven frameworks through base URLs, and every pair must be linked to the handler behind the +//! composed route. + +use std::collections::{BTreeMap, BTreeSet}; +use std::path::{Path, PathBuf}; + +use code_system_graph::{scan_workspace, status_workspace}; +use code_system_graph_model::{ + Edge, EdgeKind, HttpLinkCoverage, HttpLinkGapReason, Node, NodeId, RepoId +}; +use code_system_graph_store_sqlite::SqliteStore; + +const WORKSPACE: &str = "cross-language-matrix"; + +const PROVIDERS: [&str; 11] = [ + "actix", "axum", "chi", "express", "fastapi", "flask", "gin", "nest", "nethttp", "next", + "spring", +]; + +const TEST_REPOSITORIES: [&str; 5] = [ + "tests-go", + "tests-java", + "tests-python", + "tests-rust", + "tests-ts", +]; + +fn manifest() -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../../fixtures/cross-language-matrix/code-system-graph.yaml") +} + +struct Graph { + nodes: BTreeMap, + edges: Vec, + aliases: BTreeMap, +} + +impl Graph { + fn load(database: &Path) -> anyhow::Result { + let store = SqliteStore::open_read_only(database)?; + let (nodes, edges) = store.load_current_graph(WORKSPACE)?; + let aliases = store + .load_workspace_registry(WORKSPACE)? + .repositories + .into_iter() + .map(|repository| (repository.id, repository.alias)) + .collect(); + Ok(Self { + nodes: nodes + .into_iter() + .map(|node| (node.id.clone(), node)) + .collect(), + edges, + aliases, + }) + } + + fn alias(&self, id: &NodeId) -> String { + self.nodes[id] + .repo_id + .as_ref() + .and_then(|repo| self.aliases.get(repo)) + .cloned() + .unwrap_or_default() + } + + fn links(&self, kind: EdgeKind) -> Vec<(String, String, String, String)> { + let mut links = self + .edges + .iter() + .filter(|edge| edge.kind == kind) + .map(|edge| { + ( + self.alias(&edge.source), + self.nodes[&edge.source].label.clone(), + self.alias(&edge.target), + self.nodes[&edge.target].label.clone(), + ) + }) + .collect::>(); + links.sort(); + links + } +} + +#[test] +fn every_test_language_should_validate_every_provider_framework() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest(), &database)?; + let graph = Graph::load(&database)?; + let validates = graph.links(EdgeKind::Validates); + let implemented = graph.links(EdgeKind::ImplementedBy); + let pairs = validates + .iter() + .map(|(source, _, target, _)| (source.as_str(), target.as_str())) + .collect::>(); + let expected = TEST_REPOSITORIES + .iter() + .flat_map(|tests| PROVIDERS.iter().map(move |provider| (*tests, *provider))) + .collect::>(); + let handlers = implemented + .iter() + .map(|(provider, route, _, handler)| format!("{provider}:{route} -> {handler}")) + .collect::>(); + + assert_eq!(pairs, expected); + assert_eq!(validates.len(), expected.len()); + assert_eq!( + handlers, + [ + "actix:GET /actix/items/{id} -> rust::read_item", + "axum:GET /axum/items/:id -> rust::read_item", + "chi:GET /chi/items/{id} -> go::getItem", + "chi:GET /health -> go::health", + "express:GET /express/items/:id -> javascript::items.read", + "fastapi:GET /fastapi/items/{item_id} -> python::read_item", + "flask:GET /flask/items/ -> python::read_item", + "gin:GET /gin/items/:id -> go::getItem", + "gin:GET /health -> go::health", + "nest:GET /nest/items/:id -> typescript::findOne", + "nethttp:GET /nethttp/items/{id} -> go::getItem", + "next:GET /api/next/items/{id} -> typescript::GET", + "spring:GET /spring/items/{id} -> java::read", + ] + ); + Ok(()) +} + +#[test] +fn unlinked_calls_should_be_reported_with_their_reason() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = manifest(); + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let status = status_workspace(&manifest, &database)?; + let graph = Graph::load(&database)?; + let gaps = status + .http_links + .gaps + .iter() + .map(|gap| { + let mut candidates = gap + .candidates + .iter() + .map(|candidate| graph.alias(candidate)) + .collect::>(); + candidates.sort(); + ( + gap.reason, + format!( + "{}:{}", + graph.alias(&gap.caller), + graph.nodes[&gap.caller].label + ), + format!("{} {}", gap.method, gap.path), + candidates, + ) + }) + .collect::>(); + let gap = |reason, caller: &str, route: &str, candidates: &[&str]| { + ( + reason, + caller.to_owned(), + route.to_owned(), + candidates + .iter() + .map(|&candidate| candidate.to_owned()) + .collect(), + ) + }; + + assert_eq!( + status.http_links.coverage, + HttpLinkCoverage { + linked: 110, + no_provider: 2, + ambiguous: 2, + external: 2, + } + ); + assert_eq!(status.http_links.gap_count, 6); + assert_eq!( + gaps, + BTreeSet::from([ + gap( + HttpLinkGapReason::NoProvider, + "tests-python:python/pytest::test_missing_item", + "GET /missing/items/42", + &[], + ), + gap( + HttpLinkGapReason::NoProvider, + "tests-python:GET /missing/items/42", + "GET /missing/items/42", + &[], + ), + gap( + HttpLinkGapReason::Ambiguous, + "tests-python:python/pytest::test_health", + "GET /health", + &["chi", "gin"], + ), + gap( + HttpLinkGapReason::Ambiguous, + "tests-python:GET /health", + "GET /health", + &["chi", "gin"], + ), + gap( + HttpLinkGapReason::External, + "tests-python:python/pytest::test_public_status", + "GET /api/v2/status", + &[], + ), + gap( + HttpLinkGapReason::External, + "tests-python:GET /api/v2/status", + "GET /api/v2/status", + &[], + ), + ]) + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/differential_e2e.rs b/crates/code-system-graph-cli/tests/differential_e2e.rs new file mode 100644 index 0000000..a6afa71 --- /dev/null +++ b/crates/code-system-graph-cli/tests/differential_e2e.rs @@ -0,0 +1,236 @@ +//! Differential acceptance tests: an incrementally maintained graph equals a fresh scan after every +//! mutation, and the published graph does not depend on repository order, file creation order, or +//! extraction worker count. + +use std::path::{Path, PathBuf}; + +use code_system_graph::scan_workspace; +use code_system_graph_store_sqlite::SqliteStore; + +const WORKSPACE: &str = "cross-language-matrix"; + +fn fixture_root() -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")).join("../../fixtures/cross-language-matrix") +} + +fn fixture_files(root: &Path) -> anyhow::Result)>> { + let mut files = Vec::new(); + let mut pending = vec![root.to_path_buf()]; + while let Some(directory) = pending.pop() { + for entry in std::fs::read_dir(&directory)? { + let path = entry?.path(); + if path.is_dir() { + pending.push(path); + } else if path.file_name() != Some("code-system-graph.yaml".as_ref()) { + files.push(( + path.strip_prefix(root)?.to_path_buf(), + std::fs::read(&path)?, + )); + } + } + } + files.sort(); + Ok(files) +} + +fn write_files(root: &Path, files: &[(PathBuf, Vec)]) -> anyhow::Result<()> { + for (path, contents) in files { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(path, contents)?; + } + Ok(()) +} + +fn write_manifest(root: &Path, reversed: bool, workers: u64) -> anyhow::Result { + let template = std::fs::read_to_string(fixture_root().join("code-system-graph.yaml"))?; + let manifest = root.join("code-system-graph.yaml"); + std::fs::write(&manifest, manifest_contents(&template, reversed, workers)?)?; + Ok(manifest) +} + +fn manifest_contents(template: &str, reversed: bool, workers: u64) -> anyhow::Result { + let template = template.replace("\r\n", "\n"); + let (header, repositories) = template + .split_once("repos:\n") + .ok_or_else(|| anyhow::anyhow!("fixture manifest has no repositories"))?; + let mut entries = Vec::::new(); + for line in repositories.lines() { + if line.starts_with(" ") && !line.starts_with(" ") { + entries.push(String::new()); + } + if let Some(entry) = entries.last_mut() { + entry.push_str(line); + entry.push('\n'); + } + } + if reversed { + entries.reverse(); + } + Ok(format!( + "{header}executionPolicy:\n maxExtractionWorkers: {workers}\nrepos:\n{}", + entries.concat() + )) +} + +#[test] +fn fixture_manifest_should_support_crlf_checkouts() -> anyhow::Result<()> { + let template = std::fs::read_to_string(fixture_root().join("code-system-graph.yaml"))? + .replace("\r\n", "\n"); + for reversed in [false, true] { + assert_eq!( + manifest_contents(&template.replace('\n', "\r\n"), reversed, 4)?, + manifest_contents(&template, reversed, 4)?, + ); + } + Ok(()) +} + +/// Canonical serialization of everything a scan publishes for queries and for later reuse. +fn published(database: &Path) -> anyhow::Result { + let store = SqliteStore::open_read_only(database)?; + let (nodes, edges) = store.load_current_graph(WORKSPACE)?; + let evidence = store.load_current_evidence(WORKSPACE)?; + let links = store.load_http_link_report(WORKSPACE, usize::MAX)?; + let fingerprints = store.load_current_artifact_fingerprints(WORKSPACE)?; + let batches = store.load_current_extractor_batches(WORKSPACE)?; + Ok(serde_json::to_string(&( + nodes, + edges, + evidence, + links, + fingerprints, + batches, + ))?) +} + +fn edit(root: &Path, path: &str, from: &str, to: &str) -> anyhow::Result<()> { + let path = root.join(path); + let contents = std::fs::read_to_string(&path)?; + anyhow::ensure!( + contents.contains(from), + "`{from}` is not in {}", + path.display() + ); + std::fs::write(&path, contents.replace(from, to))?; + Ok(()) +} + +type Mutation = fn(&Path) -> anyhow::Result<()>; + +const MUTATIONS: [(&str, Mutation); 8] = [ + ("change a provider route", |root| { + edit( + root, + "gin/internal/items/items.go", + "\"/:id\"", + "\"/:id/detail\"", + ) + }), + ("change a mount prefix in another file", |root| { + edit( + root, + "fastapi/app/main.py", + "prefix=\"/fastapi\"", + "prefix=\"/fast\"", + ) + }), + ("follow the prefix in a test", |root| { + edit( + root, + "tests-python/tests/test_matrix.py", + "/fastapi/items/42", + "/fast/items/42", + ) + }), + ("add a provider file", |root| { + std::fs::write( + root.join("chi/stock.go"), + "package main\n\nimport \"github.com/go-chi/chi/v5\"\n\nfunc stock(r chi.Router) {\n\tr.Get(\"/stock/{sku}\", getStock)\n}\n", + )?; + Ok(()) + }), + ("add a test for the new route", |root| { + std::fs::write( + root.join("tests-go/stock_test.go"), + "package matrix\n\nimport (\n\t\"net/http\"\n\t\"testing\"\n)\n\nfunc TestStock(t *testing.T) {\n\thttp.Get(chiURL + \"/stock/A1\")\n}\n", + )?; + Ok(()) + }), + ("rename a test", |root| { + edit( + root, + "tests-rust/tests/matrix.rs", + "fn reads_axum_item", + "fn reads_one_axum_item", + ) + }), + ("delete a provider file", |root| { + std::fs::remove_file(root.join("express/src/routes/items.js"))?; + Ok(()) + }), + ("restore the original route", |root| { + edit( + root, + "gin/internal/items/items.go", + "\"/:id/detail\"", + "\"/:id\"", + ) + }), +]; + +#[test] +fn incremental_scans_should_equal_a_fresh_scan_after_every_mutation() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let root = temporary.path().join("workspace"); + write_files(&root, &fixture_files(&fixture_root())?)?; + let manifest = write_manifest(&root, false, 4)?; + let incremental = temporary.path().join("incremental.db"); + scan_workspace(&manifest, &incremental)?; + + for (step, (name, mutate)) in MUTATIONS.iter().enumerate() { + mutate(&root)?; + let summary = scan_workspace(&manifest, &incremental)?; + let fresh = temporary.path().join(format!("fresh-{step}.db")); + scan_workspace(&manifest, &fresh)?; + + assert!(!summary.reused_snapshot, "{name} was not observed"); + assert_eq!( + published(&incremental)?, + published(&fresh)?, + "incremental scan diverged after: {name}" + ); + } + Ok(()) +} + +#[test] +fn published_graph_should_not_depend_on_repository_order_file_order_or_workers() +-> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let root = temporary.path().join("workspace"); + let files = fixture_files(&fixture_root())?; + let mut outputs = Vec::new(); + for (reversed, workers) in [(false, 1), (true, 8), (true, 2)] { + if root.exists() { + std::fs::remove_dir_all(&root)?; + } + let mut ordered = files.clone(); + if reversed { + ordered.reverse(); + } + write_files(&root, &ordered)?; + let manifest = write_manifest(&root, reversed, workers)?; + let database = temporary + .path() + .join(format!("graph-{reversed}-{workers}.db")); + scan_workspace(&manifest, &database)?; + outputs.push(published(&database)?); + } + + assert_eq!(outputs[0], outputs[1]); + assert_eq!(outputs[0], outputs[2]); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/extraction_budgets_e2e.rs b/crates/code-system-graph-cli/tests/extraction_budgets_e2e.rs index 5c2d85c..19f424e 100644 --- a/crates/code-system-graph-cli/tests/extraction_budgets_e2e.rs +++ b/crates/code-system-graph-cli/tests/extraction_budgets_e2e.rs @@ -149,7 +149,7 @@ fn changed_budgets_should_reextract_and_fail_atomically_without_leaking_literals let persisted_text = String::from_utf8_lossy(&persisted_bytes); assert!(!persisted_text.contains("private-federation-value")); assert!(!persisted_text.contains("private-default-value")); - for suffix in [".work-v1.db", ".work-v1.db-wal", ".work-v1.db-shm"] { + for suffix in [".work.db", ".work.db-wal", ".work.db-shm"] { let path = temporary .path() .join(format!("code-system-graph.db{suffix}")); @@ -181,7 +181,7 @@ fn changed_budgets_should_reextract_and_fail_atomically_without_leaking_literals assert!( changed_batches .iter() - .all(|batch| batch.extractor_version == "1.1.0") + .all(|batch| batch.extractor_version == env!("CARGO_PKG_VERSION")) ); drop(changed_store); @@ -277,11 +277,10 @@ fn legacy_graphql_payload_should_fail_without_partial_publication() -> anyhow::R let initial = scan_workspace(&config, &database)?; let connection = rusqlite::Connection::open(&database)?; - let (snapshot_id, payload): (String, Vec) = connection.query_row( - "SELECT snapshot_id, payload + let (workspace_name, payload): (String, Vec) = connection.query_row( + "SELECT workspace_name, payload FROM extractor_batches WHERE extractor = 'code-system-graph.graphql.document' - ORDER BY rowid DESC LIMIT 1", [], |row| Ok((row.get(0)?, row.get(1)?)), @@ -294,8 +293,12 @@ fn legacy_graphql_payload_should_fail_without_partial_publication() -> anyhow::R connection.execute( "UPDATE extractor_batches SET payload = ?1 - WHERE snapshot_id = ?2 AND extractor = 'code-system-graph.graphql.document'", - rusqlite::params![tampered, snapshot_id], + WHERE workspace_name = ?2 AND extractor = 'code-system-graph.graphql.document'", + rusqlite::params![tampered, workspace_name], + )?; + std::fs::write( + repository.join("orders.graphql"), + "type Query { orders: String }\n", )?; let rejected = scan_workspace(&config, &database); diff --git a/crates/code-system-graph-cli/tests/http_trace_e2e.rs b/crates/code-system-graph-cli/tests/http_trace_e2e.rs index 3ebe872..5cc29d9 100644 --- a/crates/code-system-graph-cli/tests/http_trace_e2e.rs +++ b/crates/code-system-graph-cli/tests/http_trace_e2e.rs @@ -8,9 +8,9 @@ use code_system_graph::{ }; use code_system_graph_core::{RegisteredWorkspace, parse_manifest, register_workspace}; use code_system_graph_model::{ - EdgeKind, ExtractorRunStatus, OverallFreshness, RepoFreshnessState, ToolEnvelope, ToolStatus, TraceReport, stable_id + EdgeKind, ExtractorRunStatus, Node, NodeKind, OverallFreshness, RepoFreshnessState, ToolEnvelope, ToolStatus, TraceReport, stable_id }; -use code_system_graph_store_sqlite::{SqliteStore, latest_schema_version}; +use code_system_graph_store_sqlite::{SqliteStore, schema_identity}; #[derive(Debug, PartialEq, Eq)] struct TraceObservation { @@ -19,7 +19,7 @@ struct TraceObservation { evidence_count: usize, tool_status: ToolStatus, segment_count: usize, - schema_version: i64, + schema_id: String, freshness: OverallFreshness, discovered_inputs: usize, changed_inputs: usize, @@ -277,7 +277,7 @@ fn scan_and_trace_should_link_python_test_to_rust_implementation() -> anyhow::Re evidence_count: summary.evidence_count, tool_status: envelope.status, segment_count, - schema_version: status.schema_version, + schema_id: status.schema_id, freshness: status.freshness.overall, discovered_inputs: summary.discovered_input_count, changed_inputs: summary.changed_input_count, @@ -288,12 +288,12 @@ fn scan_and_trace_should_link_python_test_to_rust_implementation() -> anyhow::Re cross_language_edges, }, TraceObservation { - node_count: 101, - edge_count: 117, - evidence_count: 105, + node_count: 89, + edge_count: 105, + evidence_count: 93, tool_status: ToolStatus::Ok, segment_count: 1, - schema_version: latest_schema_version(), + schema_id: schema_identity().to_owned(), freshness: OverallFreshness::Fresh, discovered_inputs: 45, changed_inputs: 45, @@ -551,7 +551,7 @@ fn scan_should_apply_exact_codegraph_corroboration_without_source_payloads() -> let registry = register_workspace(&manifest, &source, &parsed)?; let repo_id = ®istry.record.repositories[0].id; let capability = - store.load_provider_capabilities("codegraph-scan", repo_id, "codegraph", "1.5.0")?; + store.load_provider_capabilities("codegraph-scan", repo_id, "codegraph", "1.6.1")?; let codegraph_evidence = store .search_current_nodes("codegraph-scan", "anchor", 10)? .len(); @@ -583,6 +583,14 @@ fn scan_should_apply_exact_codegraph_corroboration_without_source_payloads() -> Ok(()) } +fn http_operation_labels(nodes: &[Node]) -> Vec<&str> { + nodes + .iter() + .filter(|node| node.kind == NodeKind::HttpOperation) + .map(|node| node.label.as_str()) + .collect() +} + #[test] fn cli_openapi_override_should_take_precedence_and_invalidate_plain_status() -> anyhow::Result<()> { let temporary = tempfile::tempdir()?; @@ -618,10 +626,10 @@ fn cli_openapi_override_should_take_precedence_and_invalidate_plain_status() -> assert_eq!( ( summary.node_count, - nodes.first().map(|node| node.label.as_str()), + http_operation_labels(&nodes), status.freshness.overall, ), - (1, Some("GET /cli"), OverallFreshness::Stale) + (2, vec!["GET /cli"], OverallFreshness::Stale) ); Ok(()) } @@ -655,11 +663,8 @@ fn cli_repo_openapi_flag_should_override_manifest() -> anyhow::Result<()> { let (nodes, _) = SqliteStore::open_read_only(&database)?.load_current_graph("cli-override")?; assert_eq!( - ( - output.status.success(), - nodes.first().map(|node| node.label.as_str()), - ), - (true, Some("GET /cli")) + (output.status.success(), http_operation_labels(&nodes),), + (true, vec!["GET /cli"]) ); Ok(()) } diff --git a/crates/code-system-graph-cli/tests/in_process_tests_e2e.rs b/crates/code-system-graph-cli/tests/in_process_tests_e2e.rs new file mode 100644 index 0000000..7b41f2c --- /dev/null +++ b/crates/code-system-graph-cli/tests/in_process_tests_e2e.rs @@ -0,0 +1,238 @@ +//! Acceptance tests for tests that exercise their own application through in-process test +//! clients, linked to the provider of their repository even when another repository serves the +//! same route. + +use std::collections::BTreeMap; +use std::path::Path; + +use code_system_graph::scan_workspace; +use code_system_graph_model::{EdgeKind, RepoId}; +use code_system_graph_store_sqlite::SqliteStore; + +const SOURCES: [(&str, &str); 9] = [ + ( + "orders-py/app/main.py", + r#"from fastapi import FastAPI + +app = FastAPI() + +@app.get("/orders/{order_id}") +def read_order(order_id: str): + return {} +"#, + ), + ( + "orders-py/tests/conftest.py", + r"import pytest +from fastapi.testclient import TestClient +from app.main import app + +@pytest.fixture +def client(): + return TestClient(app) +", + ), + ( + "orders-py/tests/test_orders.py", + r#"def test_read_order(client): + assert client.get("/orders/42").status_code == 200 +"#, + ), + ( + "orders-js/src/app.js", + r#"const express = require("express"); +const orders = require("./orders"); +const app = express(); +app.get("/orders/:id", orders.read); +module.exports = app; +"#, + ), + ( + "orders-js/test/orders.test.js", + r#"const request = require("supertest"); +const app = require("../src/app"); + +describe("orders", () => { + it("reads one", async () => { + await request(app).get("/orders/7").expect(200); + }); +}); +"#, + ), + ( + "inventory/main.go", + r#"package main + +import "github.com/gin-gonic/gin" + +func router() *gin.Engine { + r := gin.Default() + r.GET("/stock/:sku", handlers.GetStock) + return r +} +"#, + ), + ( + "inventory/main_test.go", + r#"package main + +import ( + "net/http" + "net/http/httptest" + "testing" +) + +func TestGetStock(t *testing.T) { + req := httptest.NewRequest(http.MethodGet, "/stock/A1", nil) + rec := httptest.NewRecorder() + router().ServeHTTP(rec, req) +} +"#, + ), + ( + "billing/src/main/java/com/shop/InvoicesController.java", + r#"package com.shop; + +import org.springframework.web.bind.annotation.*; + +@RestController +@RequestMapping("/invoices") +public class InvoicesController { + @GetMapping("/{id}") + public Invoice read(@PathVariable String id) { + return null; + } +} +"#, + ), + ( + "billing/src/test/java/com/shop/InvoicesControllerTest.java", + r#"package com.shop; + +import org.junit.jupiter.api.Test; +import org.springframework.test.web.servlet.MockMvc; +import static org.springframework.test.web.servlet.request.MockMvcRequestBuilders.get; + +class InvoicesControllerTest { + private MockMvc mockMvc; + + @Test + void readsInvoice() throws Exception { + mockMvc.perform(get("/invoices/{id}", "inv-1")); + } +} +"#, + ), +]; + +const MANIFEST: &str = "version: 1\nname: inproc\nrepos:\n orders-py:\n path: orders-py\n orders-js:\n path: orders-js\n inventory:\n path: inventory\n billing:\n path: billing\n"; + +#[test] +fn absolute_urls_in_process_should_validate_only_the_local_provider() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path())?; + std::fs::write( + temporary.path().join("orders-py/tests/test_absolute.py"), + r#"import httpx +from fastapi.testclient import TestClient +from app.main import app + +client = TestClient(app) + +def test_absolute_url(): + client.get("http://testserver/orders/42") + +async def test_asgi_absolute_url(): + async with httpx.AsyncClient(transport=httpx.ASGITransport(app=app)) as api: + await api.get("http://testserver/orders/43") +"#, + )?; + let database = temporary.path().join("graph.db"); + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, edges) = store.load_current_graph("inproc")?; + let nodes = nodes + .iter() + .map(|node| (&node.id, node)) + .collect::>(); + for name in ["test_absolute_url", "test_asgi_absolute_url"] { + let links = edges + .iter() + .filter(|edge| { + edge.kind == EdgeKind::Validates && nodes[&edge.source].label.ends_with(name) + }) + .collect::>(); + assert_eq!(links.len(), 1, "missing local validation for {name}"); + let source = nodes[&links[0].source]; + let target = nodes[&links[0].target]; + assert_eq!(source.repo_id, target.repo_id); + assert_eq!(target.label, "GET /orders/{order_id}"); + } + Ok(()) +} + +fn write_workspace(root: &Path) -> anyhow::Result { + for (path, contents) in SOURCES { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(&path, contents)?; + } + let manifest = root.join("code-system-graph.yaml"); + std::fs::write(&manifest, MANIFEST)?; + Ok(manifest) +} + +#[test] +fn in_process_tests_should_validate_the_provider_of_their_own_repository() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path())?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, edges) = store.load_current_graph("inproc")?; + let aliases = store + .load_workspace_registry("inproc")? + .repositories + .into_iter() + .map(|repository| (repository.id, repository.alias)) + .collect::>(); + let nodes = nodes + .iter() + .map(|node| (&node.id, node)) + .collect::>(); + let alias = |repo: Option<&RepoId>| { + repo.and_then(|repo| aliases.get(repo)) + .cloned() + .unwrap_or_default() + }; + let mut links = edges + .iter() + .filter(|edge| edge.kind == EdgeKind::Validates) + .map(|edge| { + let source = nodes[&edge.source]; + let target = nodes[&edge.target]; + format!( + "{}:{} -> {}:{}", + alias(source.repo_id.as_ref()), + source.label, + alias(target.repo_id.as_ref()), + target.label + ) + }) + .collect::>(); + links.sort(); + + assert_eq!( + links, + [ + "billing:java/junit::readsInvoice -> billing:GET /invoices/{id}", + "inventory:go/go-test::TestGetStock -> inventory:GET /stock/:sku", + "orders-js:javascript/jest::orders > reads one -> orders-js:GET /orders/:id", + "orders-py:python/pytest::test_read_order -> orders-py:GET /orders/{order_id}", + ], + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/mcp_protocol_e2e.rs b/crates/code-system-graph-cli/tests/mcp_protocol_e2e.rs index 3af4a12..bcb429d 100644 --- a/crates/code-system-graph-cli/tests/mcp_protocol_e2e.rs +++ b/crates/code-system-graph-cli/tests/mcp_protocol_e2e.rs @@ -8,7 +8,7 @@ use code_system_graph::scan_workspace; use code_system_graph_store_sqlite::SqliteStore; use rmcp::ServiceExt; use rmcp::model::{ - CallToolRequestParams, CallToolResult, ClientCapabilities, ClientInfo, ContentBlock, Implementation, ReadResourceRequestParams, ResourceContents + CallToolRequestParams, CallToolResult, ClientCapabilities, ClientConfig, ContentBlock, Implementation, ReadResourceRequestParams, ResourceContents }; fn tool_text(result: &CallToolResult) -> &str { @@ -90,7 +90,7 @@ async fn stdio_should_initialize_without_noise_and_hide_admin_tools() -> anyhow: .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("code-system-graph-e2e", env!("CARGO_PKG_VERSION")), ); @@ -301,7 +301,7 @@ async fn stdio_explore_should_proxy_bounded_ephemeral_codegraph_context() -> any .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("code-system-graph-explore-e2e", env!("CARGO_PKG_VERSION")), ); @@ -422,7 +422,7 @@ async fn explicit_admin_stdio_should_apply_bounded_audited_mutations() -> anyhow .stdout .take() .ok_or_else(|| anyhow::anyhow!("stdout"))?; - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("code-system-graph-admin-e2e", env!("CARGO_PKG_VERSION")), ); diff --git a/crates/code-system-graph-cli/tests/query_communities_e2e.rs b/crates/code-system-graph-cli/tests/query_communities_e2e.rs index ae95d06..3d8f173 100644 --- a/crates/code-system-graph-cli/tests/query_communities_e2e.rs +++ b/crates/code-system-graph-cli/tests/query_communities_e2e.rs @@ -244,7 +244,7 @@ fn cli_should_expose_query_and_community_json() -> anyhow::Result<()> { .args(["--workspace", "commerce-platform", "--limit", "2"]) .output()?; let communities = std::process::Command::new(env!("CARGO_BIN_EXE_csgraph")) - .args(["communities", "list", "--database"]) + .args(["communities", "--database"]) .arg(&database) .args(["--workspace", "commerce-platform", "--limit", "2"]) .output()?; diff --git a/crates/code-system-graph-cli/tests/router_prefixes_e2e.rs b/crates/code-system-graph-cli/tests/router_prefixes_e2e.rs new file mode 100644 index 0000000..36fd425 --- /dev/null +++ b/crates/code-system-graph-cli/tests/router_prefixes_e2e.rs @@ -0,0 +1,95 @@ +//! Acceptance tests for router prefixes composed across files before cross-language linking. + +use std::collections::BTreeMap; +use std::path::Path; + +use code_system_graph::scan_workspace; +use code_system_graph_model::{EdgeKind, RepoId}; +use code_system_graph_store_sqlite::SqliteStore; + +const SOURCES: [(&str, &str); 6] = [ + ( + "users/app/routers/users.py", + "from fastapi import APIRouter\n\nrouter = APIRouter(prefix=\"/users\")\n\n@router.get(\"/{user_id}\")\ndef read_user(user_id: int):\n return {}\n", + ), + ( + "users/app/main.py", + "from fastapi import FastAPI\nfrom app.routers import users\n\napp = FastAPI()\napp.include_router(users.router, prefix=\"/v1\")\n", + ), + ( + "orders/internal/http/orders.go", + "package http\n\nimport \"github.com/gin-gonic/gin\"\n\nfunc RegisterOrders(rg *gin.RouterGroup) {\n\trg.GET(\"/:id\", getOrder)\n}\n\nfunc getOrder(c *gin.Context) {}\n", + ), + ( + "orders/cmd/server/main.go", + "package main\n\nimport (\n\t\"github.com/gin-gonic/gin\"\n\torders \"example.com/orders/internal/http\"\n)\n\nfunc main() {\n\tr := gin.Default()\n\tapi := r.Group(\"/api\")\n\torders.RegisterOrders(api.Group(\"/orders\"))\n\tr.Run()\n}\n", + ), + ( + "tests/tests/test_users.py", + "import httpx\n\ndef test_read_user():\n httpx.get(\"http://localhost:8000/v1/users/42\")\n", + ), + ( + "tests/tests/test_orders.py", + "import requests\n\ndef test_read_order():\n requests.get(\"http://localhost:8080/api/orders/7\")\n", + ), +]; + +const MANIFEST: &str = "version: 1\nname: routers\nrepos:\n users:\n path: users\n orders:\n path: orders\n tests:\n path: tests\n"; + +fn write_workspace(root: &Path) -> anyhow::Result { + for (path, contents) in SOURCES { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(&path, contents)?; + } + let manifest = root.join("code-system-graph.yaml"); + std::fs::write(&manifest, MANIFEST)?; + Ok(manifest) +} + +#[test] +fn tests_should_link_to_handlers_whose_paths_are_composed_across_files() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path())?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, edges) = store.load_current_graph("routers")?; + let aliases = store + .load_workspace_registry("routers")? + .repositories + .into_iter() + .map(|repository| (repository.id, repository.alias)) + .collect::>(); + let nodes = nodes + .iter() + .map(|node| (&node.id, node)) + .collect::>(); + let mut links = edges + .iter() + .filter(|edge| edge.kind == EdgeKind::Validates) + .map(|edge| { + let target = nodes[&edge.target]; + let repo = target + .repo_id + .as_ref() + .and_then(|repo| aliases.get(repo)) + .cloned() + .unwrap_or_default(); + format!("{} -> {repo}:{}", nodes[&edge.source].label, target.label) + }) + .collect::>(); + links.sort(); + + assert_eq!( + links, + [ + "python/pytest::test_read_order -> orders:GET /api/orders/:id", + "python/pytest::test_read_user -> users:GET /v1/users/{user_id}", + ], + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/scale_registry_e2e.rs b/crates/code-system-graph-cli/tests/scale_registry_e2e.rs index 2a43af6..4562119 100644 --- a/crates/code-system-graph-cli/tests/scale_registry_e2e.rs +++ b/crates/code-system-graph-cli/tests/scale_registry_e2e.rs @@ -170,7 +170,7 @@ fn interconnected_100k_file_workspace_should_complete_under_finite_defaults() -> assert!(incremental.changed_input_count > 0); let mut work_path = database.as_os_str().to_os_string(); - work_path.push(".work-v1.db"); + work_path.push(".work.db"); let work_path = std::path::PathBuf::from(work_path); let work_size = std::fs::metadata(&work_path)?.len(); let cache_bytes: u64 = rusqlite::Connection::open(&work_path)? diff --git a/crates/code-system-graph-cli/tests/scan_controls_e2e.rs b/crates/code-system-graph-cli/tests/scan_controls_e2e.rs index 0a2620d..e98fae1 100644 --- a/crates/code-system-graph-cli/tests/scan_controls_e2e.rs +++ b/crates/code-system-graph-cli/tests/scan_controls_e2e.rs @@ -6,7 +6,8 @@ use std::fmt::Write as _; use code_system_graph::{ ApplicationError, ScanOverrides, WatcherState, application_exit_code, finish_watcher_lease, scan_workspace, scan_workspace_with_overrides, scan_workspace_with_worker_executable, start_watcher_lease, status_workspace }; -use code_system_graph_core::ExitCode; +use code_system_graph_core::{ExitCode, encode_native_path}; +use code_system_graph_model::NodeKind; use code_system_graph_store_sqlite::SqliteStore; fn openapi(path: &str) -> String { @@ -72,10 +73,7 @@ fn targeted_scan_should_reuse_unselected_repository_batches() -> anyhow::Result< let corrupted = connection.execute( "UPDATE extractor_batches SET output_count = 0, payload = CAST('[]' AS BLOB) - WHERE snapshot_id = ( - SELECT id FROM repo_snapshots - WHERE workspace_name = 'targeted' AND is_current = 1 - ) + WHERE workspace_name = 'targeted' AND path_display = 'test_force.py' AND extractor = 'code-system-graph.source.python'", [], @@ -406,3 +404,129 @@ fn embedding_host_should_be_able_to_select_its_worker_executable() -> anyhow::Re assert_eq!(summary.workspace, "explicit-worker"); Ok(()) } + +#[test] +fn fact_free_source_files_should_not_add_artifact_nodes() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let repository = temporary.path().join("api"); + std::fs::create_dir_all(repository.join("src"))?; + std::fs::write( + repository.join("src/math.ts"), + "export function total(values: number[]): number {\n return values.reduce((sum, value) => sum + value, 0);\n}\n", + )?; + std::fs::write( + repository.join("src/events.ts"), + "import { Kafka } from 'kafkajs';\n\nconst producer = new Kafka({ brokers: ['kafka:9092'] }).producer();\n\nexport async function publish() {\n await producer.send({ topic: 'orders.created', messages: [] });\n}\n", + )?; + let manifest = temporary.path().join("code-system-graph.yaml"); + std::fs::write( + &manifest, + "version: 1\nname: fact-free\nrepos:\n api:\n path: api\n", + )?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let store = SqliteStore::open_read_only(&database)?; + let (nodes, _) = store.load_current_graph("fact-free")?; + let math_path = encode_native_path(&std::path::Path::new("src").join("math.ts")); + let math_events = store + .load_current_extractor_batches("fact-free")? + .into_iter() + .find(|batch| { + batch.source.path == math_path + && batch.source.extractor == "code-system-graph.events.source" + }) + .ok_or_else(|| anyhow::anyhow!("math.ts event batch is missing"))?; + let artifact_keys = nodes + .iter() + .filter(|node| node.kind == NodeKind::Artifact) + .map(|node| node.stable_key.as_str()) + .collect::>(); + + assert!( + artifact_keys + .iter() + .all(|key| key.ends_with(":src/events.ts")), + "{artifact_keys:?}" + ); + assert!( + artifact_keys + .iter() + .any(|key| key.starts_with("event-artifact:")) + ); + assert!(nodes.iter().any(|node| node.kind == NodeKind::Repository)); + assert_eq!( + (math_events.output_count, math_events.payload.as_slice()), + (0, b"[]".as_slice()) + ); + Ok(()) +} + +#[test] +fn unchanged_scan_should_reuse_the_snapshot_until_the_configuration_changes() -> anyhow::Result<()> +{ + let temporary = tempfile::tempdir()?; + let repository = temporary.path().join("api"); + std::fs::create_dir(&repository)?; + std::fs::write(repository.join("openapi.yaml"), openapi("/orders"))?; + std::fs::write( + repository.join("routes.ts"), + "import express from 'express';\nconst app = express();\napp.get('/orders', listOrders);\n", + )?; + let manifest = temporary.path().join("code-system-graph.yaml"); + let write_manifest = |budget: &str| { + std::fs::write( + &manifest, + format!( + "version: 1\nname: reuse\n{budget}repos:\n api:\n path: api\n openapi: openapi.yaml\n" + ), + ) + }; + write_manifest("")?; + let database = temporary.path().join("graph.db"); + + let first = scan_workspace(&manifest, &database)?; + let unchanged = scan_workspace(&manifest, &database)?; + write_manifest("extractionBudgets:\n maxWorkUnitsPerArtifact: 100000\n")?; + let rebudgeted = scan_workspace(&manifest, &database)?; + + assert!(!first.reused_snapshot); + assert!(unchanged.reused_snapshot); + assert_eq!(unchanged.snapshot_id, first.snapshot_id); + assert_eq!(unchanged.node_count, first.node_count); + assert!(!rebudgeted.reused_snapshot); + assert_ne!(rebudgeted.snapshot_id, first.snapshot_id); + Ok(()) +} + +#[test] +fn scan_should_accept_a_bare_relative_database_file_name() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let repository = temporary.path().join("api"); + std::fs::create_dir(&repository)?; + std::fs::write(repository.join("openapi.yaml"), openapi("/orders"))?; + std::fs::write( + temporary.path().join("code-system-graph.yaml"), + "version: 1\nname: relative-database\nrepos:\n api:\n path: api\n openapi: openapi.yaml\n", + )?; + + let output = std::process::Command::new(env!("CARGO_BIN_EXE_csgraph")) + .current_dir(temporary.path()) + .args([ + "scan", + "--config", + "code-system-graph.yaml", + "--database", + "graph.db", + ]) + .output()?; + + assert!( + output.status.success(), + "{}", + String::from_utf8_lossy(&output.stderr) + ); + assert!(temporary.path().join("graph.db").is_file()); + assert!(temporary.path().join("graph.db.work.db").is_file()); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/sync_e2e.rs b/crates/code-system-graph-cli/tests/sync_e2e.rs index 3e3ce04..41fb61d 100644 --- a/crates/code-system-graph-cli/tests/sync_e2e.rs +++ b/crates/code-system-graph-cli/tests/sync_e2e.rs @@ -158,7 +158,7 @@ fn stale_codegraph_index_should_corroborate_even_without_native_changes() -> any std::fs::write( &binary, format!( - "#!/bin/sh\nprintf '%s\\n' \"$*\" >> '{}'\ncase \"$1\" in\n status)\n if [ -f '{}' ]; then modified=0; else modified=1; fi\n printf '{{\"initialized\":true,\"version\":\"1.5.0\",\"pendingChanges\":{{\"added\":0,\"modified\":%s,\"removed\":0}},\"worktreeMismatch\":null,\"index\":{{\"reindexRecommended\":false,\"state\":\"complete\"}}}}\\n' \"$modified\"\n ;;\n sync) : > '{}' ;;\n *) exec python3 '{}' \"$@\" ;;\nesac\n", + "#!/bin/sh\nprintf '%s\\n' \"$*\" >> '{}'\ncase \"$1\" in\n status)\n if [ -f '{}' ]; then modified=0; else modified=1; fi\n printf '{{\"initialized\":true,\"version\":\"1.6.1\",\"pendingChanges\":{{\"added\":0,\"modified\":%s,\"removed\":0}},\"worktreeMismatch\":null,\"index\":{{\"reindexRecommended\":false,\"state\":\"complete\"}}}}\\n' \"$modified\"\n ;;\n sync) : > '{}' ;;\n *) exec python3 '{}' \"$@\" ;;\nesac\n", invocation_log.display(), synchronized_marker.display(), synchronized_marker.display(), @@ -444,7 +444,7 @@ fn watch_failure_should_not_persist_or_emit_parser_literals() -> anyhow::Result< .arg(&database) .output()?; anyhow::ensure!(status_output.status.success(), "status command failed"); - let sidecar = std::path::PathBuf::from(format!("{}.work-v1.db", database.display())); + let sidecar = std::path::PathBuf::from(format!("{}.work.db", database.display())); let persisted = std::fs::read(sidecar)?; for (surface, bytes) in [ ("termination JSONL", termination.as_bytes()), diff --git a/crates/code-system-graph-cli/tests/sync_pipeline_e2e.rs b/crates/code-system-graph-cli/tests/sync_pipeline_e2e.rs new file mode 100644 index 0000000..3000e67 --- /dev/null +++ b/crates/code-system-graph-cli/tests/sync_pipeline_e2e.rs @@ -0,0 +1,184 @@ +//! Acceptance tests for the single-read, stat-cached, parallel extraction pipeline. + +use std::path::Path; +use std::time::{Duration, SystemTime}; + +use code_system_graph::{ScanOverrides, scan_workspace, scan_workspace_with_overrides}; +use code_system_graph_store_sqlite::SqliteStore; + +const SOURCES: [(&str, &str); 6] = [ + ( + "api/src/server.ts", + "import express from 'express';\nconst app = express();\napp.get('/v1/orders/:id', (req, res) => res.json({}));\napp.post('/v1/orders', (req, res) => res.json({}));\n", + ), + ( + "api/src/publisher.ts", + "export async function announce(producer) {\n await producer.send({ topic: 'orders.created', messages: [] });\n}\n", + ), + ( + "api/src/db.py", + "def load(cursor):\n cursor.execute(\"SELECT id FROM orders WHERE id = %s\", (1,))\n", + ), + ( + "tests/tests/test_orders.py", + "import requests\n\ndef test_create_order():\n requests.post(\"http://localhost:8080/v1/orders\", json={})\n", + ), + ( + "tests/tests/test_lookup.py", + "import requests\n\ndef test_lookup_order():\n requests.get(\"http://localhost:8080/v1/orders/42\")\n", + ), + ("api/README.md", "# Orders API\n\nServes `/v1/orders`.\n"), +]; + +fn write_workspace(root: &Path, workers: u64) -> anyhow::Result { + for (path, contents) in SOURCES { + let path = root.join(path); + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(&path, contents)?; + } + let manifest = root.join("code-system-graph.yaml"); + std::fs::write( + &manifest, + format!( + "version: 1\nname: pipeline\nexecutionPolicy:\n maxExtractionWorkers: {workers}\nrepos:\n api:\n path: api\n tests:\n path: tests\n" + ), + )?; + Ok(manifest) +} + +fn age_sources(root: &Path, age: Duration) -> anyhow::Result<()> { + let modified = SystemTime::now() - age; + for (path, _) in SOURCES { + std::fs::File::options() + .write(true) + .open(root.join(path))? + .set_modified(modified)?; + } + Ok(()) +} + +#[test] +fn published_graph_should_not_depend_on_extraction_worker_count() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let mut graphs = Vec::new(); + for workers in [1, 8] { + let manifest = write_workspace(temporary.path(), workers)?; + let database = temporary.path().join(format!("graph-{workers}.db")); + let summary = scan_workspace(&manifest, &database)?; + assert_eq!( + summary.execution.extraction_workers, + workers.min( + std::thread::available_parallelism() + .map_or(1, std::num::NonZeroUsize::get) + .try_into()? + ) + ); + let (nodes, edges) = + SqliteStore::open_read_only(&database)?.load_current_graph("pipeline")?; + assert_ne!(nodes, Vec::new()); + graphs.push((nodes, edges)); + } + + assert_eq!(graphs[0], graphs[1]); + Ok(()) +} + +#[test] +fn unchanged_files_should_be_fingerprinted_from_the_stat_cache() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path(), 4)?; + let database = temporary.path().join("graph.db"); + age_sources(temporary.path(), Duration::from_mins(10))?; + + let initial = scan_workspace(&manifest, &database)?; + let unchanged = scan_workspace(&manifest, &database)?; + + assert_eq!(initial.execution.stat_cache_hits, 0); + assert!(unchanged.reused_snapshot); + assert_eq!(unchanged.snapshot_id, initial.snapshot_id); + assert!(unchanged.execution.stat_cache_hits >= u64::try_from(SOURCES.len())?); + assert!( + unchanged + .execution + .phases + .iter() + .any(|phase| phase.phase == code_system_graph_core::JobPhase::Fingerprinting) + ); + + let edited = temporary.path().join("api/src/server.ts"); + let original = std::fs::read_to_string(&edited)?; + std::fs::write( + &edited, + original.replace("/v1/orders/:id", "/v1/orders/:oid"), + )?; + std::fs::File::options() + .write(true) + .open(&edited)? + .set_modified(SystemTime::now() - Duration::from_mins(5))?; + let changed = scan_workspace(&manifest, &database)?; + + assert!(changed.changed_input_count > 0); + assert_ne!(changed.snapshot_id, initial.snapshot_id); + Ok(()) +} + +#[test] +fn recently_modified_files_should_always_be_read() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path(), 2)?; + let database = temporary.path().join("graph.db"); + + scan_workspace(&manifest, &database)?; + let second = scan_workspace(&manifest, &database)?; + + assert_eq!(second.execution.stat_cache_hits, 0); + assert!(second.reused_snapshot); + Ok(()) +} + +#[test] +fn touched_repository_scans_should_converge_with_a_full_scan() -> anyhow::Result<()> { + let temporary = tempfile::tempdir()?; + let manifest = write_workspace(temporary.path(), 2)?; + let database = temporary.path().join("graph.db"); + scan_workspace(&manifest, &database)?; + + let publisher = temporary.path().join("api/src/publisher.ts"); + std::fs::write( + &publisher, + std::fs::read_to_string(&publisher)?.replace("orders.created", "orders.updated"), + )?; + let test = temporary.path().join("tests/tests/test_orders.py"); + std::fs::write( + &test, + std::fs::read_to_string(&test)?.replace("test_create_order", "test_place_order"), + )?; + let targeted = scan_workspace_with_overrides( + &manifest, + &database, + &ScanOverrides { + touched_repositories: vec!["api".to_owned()], + ..ScanOverrides::default() + }, + )?; + let (nodes, _) = SqliteStore::open_read_only(&database)?.load_current_graph("pipeline")?; + let graph_json = serde_json::to_string(&nodes)?; + + assert!(!targeted.reused_snapshot); + assert!(graph_json.contains("orders.updated")); + assert!(graph_json.contains("test_create_order")); + assert!(!graph_json.contains("test_place_order")); + + let full = scan_workspace(&manifest, &database)?; + let fresh_database = temporary.path().join("fresh.db"); + scan_workspace(&manifest, &fresh_database)?; + + assert!(!full.reused_snapshot); + assert_eq!( + SqliteStore::open_read_only(&database)?.load_current_graph("pipeline")?, + SqliteStore::open_read_only(&fresh_database)?.load_current_graph("pipeline")? + ); + Ok(()) +} diff --git a/crates/code-system-graph-cli/tests/sync_scale_e2e.rs b/crates/code-system-graph-cli/tests/sync_scale_e2e.rs new file mode 100644 index 0000000..2204eff --- /dev/null +++ b/crates/code-system-graph-cli/tests/sync_scale_e2e.rs @@ -0,0 +1,229 @@ +//! Release-mode sync acceptance at 200 repositories and 50,000 files: content reads, publication +//! writes, and database growth across repeated syncs. + +use std::fmt::Write as _; +use std::path::{Path, PathBuf}; +use std::time::{Duration, Instant, SystemTime}; + +use code_system_graph::scan_workspace; +use code_system_graph_model::RepoId; +use code_system_graph_store_sqlite::SqliteStore; + +const WORKSPACE: &str = "sync-scale"; +const REPOSITORY_COUNT: usize = 200; +const FILES_PER_REPOSITORY: usize = 250; +const SYNC_ROUNDS: u64 = 20; +const CHANGED_REPOSITORY: &str = "repo-042"; +/// Directory where the workload is generated instead of a temporary one, so another build can +/// scan the same files. +const WORKLOAD_DIRECTORY: &str = "CODE_SYSTEM_GRAPH_SCALE_WORKLOAD"; + +#[test] +#[ignore = "run explicitly with --release for the 200-repository, 50,000-file sync workload"] +fn two_hundred_repository_sync_should_meet_read_write_and_growth_gates() -> anyhow::Result<()> { + if cfg!(debug_assertions) { + anyhow::bail!("run this acceptance test with --release"); + } + let temporary = tempfile::tempdir()?; + let root = std::env::var_os(WORKLOAD_DIRECTORY) + .map_or_else(|| temporary.path().join("workload"), PathBuf::from); + let manifest = write_workload(&root)?; + let database = temporary.path().join("graph.db"); + + let started = Instant::now(); + let cold = scan_workspace(&manifest, &database)?; + let cold_elapsed = started.elapsed(); + + let started = Instant::now(); + let unchanged = scan_workspace(&manifest, &database)?; + let unchanged_elapsed = started.elapsed(); + + let server = root.join(CHANGED_REPOSITORY).join("src/server.ts"); + let original = std::fs::read_to_string(&server)?; + let extended = original.replace( + "app.get('/svc-042/health', health);", + "app.get('/svc-042/health', health);\napp.get('/svc-042/audit/:id', getItem);", + ); + anyhow::ensure!(extended != original, "the changed route was not inserted"); + write_aged(&server, &extended, Duration::from_mins(5))?; + let started = Instant::now(); + let one_file = scan_workspace(&manifest, &database)?; + let one_file_elapsed = started.elapsed(); + let segment = segment_rows(&database, CHANGED_REPOSITORY)?; + let first_size = database_bytes(&database)?; + + for round in 0..SYNC_ROUNDS { + let contents = if round % 2 == 0 { &original } else { &extended }; + write_aged( + &server, + contents, + Duration::from_mins(4).saturating_sub(Duration::from_secs(round)), + )?; + scan_workspace(&manifest, &database)?; + } + let final_size = database_bytes(&database)?; + + eprintln!( + "sync_scale repos={REPOSITORY_COUNT} files={} cold_ms={} cold_peak_worker_rss_bytes={} \ + cold_bytes_read={} nodes={} edges={} unchanged_ms={} unchanged_bytes_read={} \ + unchanged_stat_hits={} one_file_ms={} one_file_bytes_read={} one_file_published_rows={} \ + segment_rows={segment} first_db_bytes={first_size} final_db_bytes={final_size}", + REPOSITORY_COUNT * FILES_PER_REPOSITORY, + cold_elapsed.as_millis(), + cold.execution.peak_worker_memory_bytes, + cold.execution.content_bytes_read, + cold.node_count, + cold.edge_count, + unchanged_elapsed.as_millis(), + unchanged.execution.content_bytes_read, + unchanged.execution.stat_cache_hits, + one_file_elapsed.as_millis(), + one_file.execution.content_bytes_read, + one_file.execution.published_rows, + ); + for phase in &cold.execution.phases { + eprintln!( + "sync_scale cold_phase={:?} ms={} rss_bytes={}", + phase.phase, phase.duration_ms, phase.resident_memory_bytes + ); + } + assert!(unchanged.reused_snapshot); + assert_eq!(unchanged.execution.content_bytes_read, 0); + assert!(!one_file.reused_snapshot); + assert!( + one_file.execution.published_rows <= segment, + "one changed file published {} rows; the repository segment has {segment}", + one_file.execution.published_rows + ); + assert!( + final_size.saturating_mul(100) <= first_size.saturating_mul(105), + "database grew from {first_size} to {final_size} bytes after {SYNC_ROUNDS} syncs" + ); + Ok(()) +} + +fn write_workload(root: &Path) -> anyhow::Result { + let mut manifest = format!("version: 1\nname: {WORKSPACE}\nrepos:\n"); + for index in 0..REPOSITORY_COUNT { + let alias = format!("repo-{index:03}"); + let repository = root.join(&alias); + let next = (index + 1) % REPOSITORY_COUNT; + let mut files = vec![ + ( + "src/server.ts".to_owned(), + format!( + "import express from 'express';\n\nconst app = express();\n\nfunction getItem(req, res) {{\n res.json({{ id: req.params.id }});\n}}\n\nfunction health(req, res) {{\n res.json({{ ok: true }});\n}}\n\napp.get('/svc-{index:03}/items/:id', getItem);\napp.get('/svc-{index:03}/health', health);\n\nexport default app;\n" + ), + ), + ( + "src/client.ts".to_owned(), + format!( + "const BASE_URL = 'http://localhost:8080';\n\nexport async function loadNext(id: string) {{\n return fetch(`${{BASE_URL}}/svc-{next:03}/items/${{id}}`);\n}}\n" + ), + ), + ( + "tests/test_api.py".to_owned(), + format!( + "import requests\n\nBASE_URL = \"http://localhost:8080\"\n\n\ndef test_item():\n requests.get(f\"{{BASE_URL}}/svc-{index:03}/items/42\")\n" + ), + ), + ]; + for module in 0..FILES_PER_REPOSITORY - files.len() { + files.push(module_file(module)); + } + for (path, contents) in files { + write_aged(&repository.join(path), &contents, Duration::from_mins(10))?; + } + writeln!(manifest, " {alias}:\n path: {alias}")?; + } + let manifest_path = root.join("code-system-graph.yaml"); + std::fs::write(&manifest_path, manifest)?; + Ok(manifest_path) +} + +fn module_file(module: usize) -> (String, String) { + match module % 3 { + 0 => ( + format!("src/modules/module_{module:03}.ts"), + format!( + "export interface Line{module} {{\n id: string;\n total: number;\n}}\n\nexport function total{module}(lines: Line{module}[]): number {{\n return lines.reduce((sum, line) => sum + line.total, 0);\n}}\n" + ), + ), + 1 => ( + format!("lib/module_{module:03}.py"), + format!( + "def total_{module}(lines):\n return sum(line[\"total\"] for line in lines)\n\n\ndef largest_{module}(lines):\n return max(lines, key=lambda line: line[\"total\"], default=None)\n" + ), + ), + _ => ( + format!("pkg/module_{module:03}.go"), + format!( + "package pkg\n\nfunc Total{module}(values []int) int {{\n\ttotal := 0\n\tfor _, value := range values {{\n\t\ttotal += value\n\t}}\n\treturn total\n}}\n" + ), + ), + } +} + +fn write_aged(path: &Path, contents: &str, age: Duration) -> anyhow::Result<()> { + if let Some(parent) = path.parent() { + std::fs::create_dir_all(parent)?; + } + std::fs::write(path, contents)?; + std::fs::File::options() + .write(true) + .open(path)? + .set_modified(SystemTime::now() - age)?; + Ok(()) +} + +fn database_bytes(database: &Path) -> anyhow::Result { + let mut total = std::fs::metadata(database)?.len(); + let mut wal = database.as_os_str().to_os_string(); + wal.push("-wal"); + if let Ok(metadata) = std::fs::metadata(PathBuf::from(wal)) { + total = total.saturating_add(metadata.len()); + } + Ok(total) +} + +/// Rows owned by one repository: its nodes, every edge touching them, its evidence, and its +/// fingerprints, extractor batches, and extractor runs. +fn segment_rows(database: &Path, alias: &str) -> anyhow::Result { + let store = SqliteStore::open_read_only(database)?; + let repository = store + .load_workspace_registry(WORKSPACE)? + .repositories + .into_iter() + .find(|repository| repository.alias == alias) + .map(|repository| repository.id) + .ok_or_else(|| anyhow::anyhow!("repository {alias} is not registered"))?; + let owned = |repo: Option<&RepoId>| repo == Some(&repository); + let (nodes, edges) = store.load_current_graph(WORKSPACE)?; + let node_ids = nodes + .iter() + .filter(|node| owned(node.repo_id.as_ref())) + .map(|node| &node.id) + .collect::>(); + let edge_count = edges + .iter() + .filter(|edge| node_ids.contains(&edge.source) || node_ids.contains(&edge.target)) + .count(); + let evidence_count = store + .load_current_evidence(WORKSPACE)? + .iter() + .filter(|evidence| owned(evidence.repo_id.as_ref())) + .count(); + let fingerprint_count = store + .load_current_artifact_fingerprints(WORKSPACE)? + .iter() + .filter(|fingerprint| fingerprint.repo_id == repository) + .count(); + let run_count = store + .load_current_extractor_runs(WORKSPACE)? + .iter() + .filter(|run| run.repo_id == repository) + .count(); + Ok(u64::try_from( + node_ids.len() + edge_count + evidence_count + 2 * fingerprint_count + run_count, + )?) +} diff --git a/crates/code-system-graph-core/Cargo.toml b/crates/code-system-graph-core/Cargo.toml index ab38034..0beec0c 100644 --- a/crates/code-system-graph-core/Cargo.toml +++ b/crates/code-system-graph-core/Cargo.toml @@ -12,20 +12,22 @@ keywords.workspace = true categories.workspace = true [dependencies] -async-trait = "0.1.91" -atomic-write-file = "0.3.0" -blake3 = "1.8.5" +aho-corasick = "1.1.5" +async-trait = "0.1.92" +atomic-write-file = "0.3.1" +blake3 = "1.8.7" +foldhash = "0.2.0" graphql-parser = "0.4.1" -globset = "0.4.19" -hcl-rs = "0.19.7" -ignore = "0.4.31" -libc = "0.2" +globset = "0.4.20" +hcl-rs = "0.19.8" +ignore = "0.4.33" +libc = "0.2.190" nix = { version = "0.31.3", features = ["fs"] } proto-parser = "1.14.3" pulldown-cmark = "0.13.4" -code-system-graph-model = { version = "1.1.0", path = "../code-system-graph-model" } -reqwest = { version = "0.13.4", default-features = false, features = ["json", "rustls"] } -rmcp = { version = "3.1.0", default-features = false, features = [ +code-system-graph-model = { version = "1.2.0", path = "../code-system-graph-model" } +reqwest = { version = "0.13.5", default-features = false, features = ["json", "rustls"] } +rmcp = { version = "3.5.0", default-features = false, features = [ "client", "transport-async-rw", "which-command", @@ -34,12 +36,12 @@ schemars = "1.2.2" semver = { version = "1.0.28", features = ["serde"] } serde = { version = "1.0.229", features = ["derive"] } serde_json = { version = "1.0.151", features = ["preserve_order"] } -serde-saphyr = "1.0.0" -sqlparser = "0.62.0" -thiserror = "2.0.19" +serde-saphyr = "1.3.0" +sqlparser = "0.63.0" +thiserror = "2.0.21" tokio = { version = "1.53.1", features = ["io-util", "macros", "process", "rt", "sync", "time"] } tokio-util = { version = "0.7.19", features = ["rt"] } -tree-sitter = "0.26.11" +tree-sitter = "0.27.0" tree-sitter-go = "0.25.0" tree-sitter-java = "0.23.5" tree-sitter-javascript = "0.25.0" diff --git a/crates/code-system-graph-core/src/batch.rs b/crates/code-system-graph-core/src/batch.rs index 9111324..6207994 100644 --- a/crates/code-system-graph-core/src/batch.rs +++ b/crates/code-system-graph-core/src/batch.rs @@ -1,40 +1,12 @@ -use std::collections::{BTreeMap, BTreeSet}; - -use code_system_graph_model::{ - ArtifactChangeKind, ArtifactFingerprint, CheckoutId, NativePath, RepoId -}; +use code_system_graph_model::ArtifactFingerprint; use serde::Serialize; use serde::de::DeserializeOwned; use thiserror::Error; use crate::{ - EXTRACTION_CONTRACT_VERSION, ExtractionBudgets, ExtractionLimitExceeded, ExtractionResource, ExtractionTracker, IncrementalPlan + EXTRACTION_CONTRACT_VERSION, ExtractionBudgets, ExtractionLimitExceeded, ExtractionResource, ExtractionTracker }; -/// Stable identity of one extractor input within a concrete checkout. -#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord)] -pub struct ArtifactKey { - /// Repository identity shared by linked worktrees. - pub repo_id: RepoId, - /// Concrete checkout identity. - pub checkout_id: CheckoutId, - /// Lossless repository-relative artifact path. - pub path: NativePath, - /// Extractor that owns this artifact. - pub extractor: String, -} - -impl From<&ArtifactFingerprint> for ArtifactKey { - fn from(fingerprint: &ArtifactFingerprint) -> Self { - Self { - repo_id: fingerprint.repo_id.clone(), - checkout_id: fingerprint.checkout_id.clone(), - path: fingerprint.path.clone(), - extractor: fingerprint.extractor.clone(), - } - } -} - /// Complete transient output owned by one extractor input. #[derive(Debug, Clone, PartialEq, Eq)] pub struct ExtractorBatch { @@ -50,99 +22,11 @@ impl ExtractorBatch { pub fn new(source: ArtifactFingerprint, outputs: Vec) -> Self { Self { source, outputs } } - - /// Returns the stable source key for planning and persistence. - #[must_use] - pub fn key(&self) -> ArtifactKey { - ArtifactKey::from(&self.source) - } -} - -/// Required action for one source-owned extractor batch. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum BatchAction { - /// Execute the extractor for a newly discovered source. - Add, - /// Execute the extractor and replace an existing source batch. - Replace, - /// Reuse the previous batch without extractor work. - Reuse, - /// Remove the previous batch because its source disappeared. - Delete, -} - -/// One deterministic source action derived from an incremental artifact plan. -#[derive(Debug, Clone, PartialEq, Eq)] -pub struct PlannedBatch { - /// Stable extractor input identity. - pub key: ArtifactKey, - /// Required action. - pub action: BatchAction, -} - -/// Source-level extraction and deletion work for one scan. -#[derive(Debug, Clone, PartialEq, Eq)] -pub struct ExtractorBatchPlan { - /// Actions sorted by repository, checkout, path, and extractor. - pub batches: Vec, -} - -impl ExtractorBatchPlan { - /// Returns only source batches requiring extraction or deletion. - pub fn changed(&self) -> impl Iterator { - self.batches - .iter() - .filter(|batch| batch.action != BatchAction::Reuse) - } -} - -/// Converts artifact changes into source-owned extractor batch actions. -#[must_use] -pub fn plan_extractor_batches(plan: &IncrementalPlan) -> ExtractorBatchPlan { - let batches = plan - .changes - .iter() - .map(|change| PlannedBatch { - key: ArtifactKey { - repo_id: change.repo_id.clone(), - checkout_id: change.checkout_id.clone(), - path: change.path.clone(), - extractor: change.extractor.clone(), - }, - action: match change.kind { - ArtifactChangeKind::Added => BatchAction::Add, - ArtifactChangeKind::Modified => BatchAction::Replace, - ArtifactChangeKind::Deleted => BatchAction::Delete, - ArtifactChangeKind::Unchanged => BatchAction::Reuse, - }, - }) - .collect(); - ExtractorBatchPlan { batches } } -/// Error returned when affected-neighborhood planning lacks a required batch. +/// Error returned while encoding or decoding a source-owned extractor batch. #[derive(Debug, Error, PartialEq, Eq)] pub enum BatchPlanError { - /// Multiple batches claim the same source key. - #[error("duplicate {side} extractor batch for `{extractor}` at `{path}`")] - DuplicateBatch { - /// Previous or current batch set. - side: &'static str, - /// Extractor owning the duplicate input. - extractor: String, - /// Diagnostic path display. - path: String, - }, - /// A changed source has no required previous or current output batch. - #[error("missing {side} extractor batch for `{extractor}` at `{path}`")] - MissingBatch { - /// Previous or current batch set. - side: &'static str, - /// Extractor owning the missing input. - extractor: String, - /// Diagnostic path display. - path: String, - }, /// A source-owned output payload could not be serialized or decoded. #[error("invalid extractor batch payload: {0}")] InvalidPayload(String), @@ -247,114 +131,16 @@ pub fn load_extractor_batch_with_budgets( Ok(ExtractorBatch::new(stored.source.clone(), outputs)) } -/// Computes exact link neighborhoods affected by add, modify, and delete actions. -/// -/// Modified sources contribute old and new keys, deleted sources contribute old keys, and added -/// sources contribute new keys. Unchanged sources do not trigger relinking. -/// -/// # Errors -/// -/// Returns [`BatchPlanError`] for duplicate source batches or when a changed action lacks the -/// required previous/current batch. -pub fn affected_link_keys( - plan: &ExtractorBatchPlan, - previous: &[ExtractorBatch], - current: &[ExtractorBatch], - link_key: impl Fn(&T) -> K, -) -> Result, BatchPlanError> -where - K: Ord, -{ - let previous = batch_map(previous, "previous")?; - let current = batch_map(current, "current")?; - let mut keys = BTreeSet::new(); - for batch in plan.changed() { - match batch.action { - BatchAction::Add => { - extend_link_keys( - &mut keys, - required_batch(¤t, batch, "current")?, - &link_key, - ); - } - BatchAction::Replace => { - extend_link_keys( - &mut keys, - required_batch(&previous, batch, "previous")?, - &link_key, - ); - extend_link_keys( - &mut keys, - required_batch(¤t, batch, "current")?, - &link_key, - ); - } - BatchAction::Delete => { - extend_link_keys( - &mut keys, - required_batch(&previous, batch, "previous")?, - &link_key, - ); - } - BatchAction::Reuse => {} - } - } - Ok(keys) -} - -fn batch_map<'a, T>( - batches: &'a [ExtractorBatch], - side: &'static str, -) -> Result>, BatchPlanError> { - let mut map = BTreeMap::new(); - for batch in batches { - let key = batch.key(); - if map.insert(key.clone(), batch).is_some() { - return Err(BatchPlanError::DuplicateBatch { - side, - extractor: key.extractor, - path: key.path.display, - }); - } - } - Ok(map) -} - -fn required_batch<'a, T>( - batches: &BTreeMap>, - planned: &PlannedBatch, - side: &'static str, -) -> Result<&'a ExtractorBatch, BatchPlanError> { - batches - .get(&planned.key) - .copied() - .ok_or_else(|| BatchPlanError::MissingBatch { - side, - extractor: planned.key.extractor.clone(), - path: planned.key.path.display.clone(), - }) -} - -fn extend_link_keys( - keys: &mut BTreeSet, - batch: &ExtractorBatch, - link_key: &impl Fn(&T) -> K, -) where - K: Ord, -{ - keys.extend(batch.outputs.iter().map(link_key)); -} - #[cfg(test)] mod tests { use code_system_graph_model::{ - ArtifactChange, ArtifactChangeKind, ArtifactFingerprint, CheckoutId, NativePath, NativePathEncoding, RepoId + ArtifactFingerprint, CheckoutId, NativePath, NativePathEncoding, RepoId }; use super::{ - BatchAction, BatchPlanError, ExtractorBatch, affected_link_keys, load_extractor_batch, load_extractor_batch_with_budgets, plan_extractor_batches, store_extractor_batch + BatchPlanError, ExtractorBatch, load_extractor_batch, load_extractor_batch_with_budgets, store_extractor_batch }; - use crate::{ExtractionBudgets, ExtractionLimitExceeded, ExtractionTracker, IncrementalPlan}; + use crate::{ExtractionBudgets, ExtractionLimitExceeded, ExtractionTracker}; fn tracker() -> ExtractionTracker { ExtractionTracker::new("src/routes.rs", "test", &ExtractionBudgets::default()) @@ -368,16 +154,6 @@ mod tests { } } - fn change(source: &str, kind: ArtifactChangeKind) -> ArtifactChange { - ArtifactChange { - repo_id: RepoId::new("repo:api"), - checkout_id: CheckoutId::new("checkout:api"), - path: path(source), - extractor: "code-system-graph.http.openapi".to_owned(), - kind, - } - } - fn batch(source: &str, hash: &str, outputs: &[&str]) -> ExtractorBatch { ExtractorBatch::new( ArtifactFingerprint { @@ -392,63 +168,6 @@ mod tests { ) } - #[test] - fn batch_plan_should_preserve_deterministic_source_actions() { - let plan = plan_extractor_batches(&IncrementalPlan { - changes: vec![ - change("added.yaml", ArtifactChangeKind::Added), - change("deleted.yaml", ArtifactChangeKind::Deleted), - change("same.yaml", ArtifactChangeKind::Unchanged), - ], - }); - - assert_eq!( - plan.batches - .iter() - .map(|batch| batch.action) - .collect::>(), - vec![BatchAction::Add, BatchAction::Delete, BatchAction::Reuse] - ); - } - - #[test] - fn affected_keys_should_include_old_and_new_modified_neighborhoods() { - let plan = plan_extractor_batches(&IncrementalPlan { - changes: vec![change("openapi.yaml", ArtifactChangeKind::Modified)], - }); - let result = affected_link_keys( - &plan, - &[batch("openapi.yaml", "old", &["POST:/v1/orders"])], - &[batch("openapi.yaml", "new", &["POST:/v2/orders"])], - Clone::clone, - ); - - assert_eq!( - result, - Ok(["POST:/v1/orders".to_owned(), "POST:/v2/orders".to_owned()] - .into_iter() - .collect()) - ); - } - - #[test] - fn affected_keys_should_require_deleted_previous_batch() { - let plan = plan_extractor_batches(&IncrementalPlan { - changes: vec![change("deleted.yaml", ArtifactChangeKind::Deleted)], - }); - let previous: Vec> = Vec::new(); - let current: Vec> = Vec::new(); - let result = affected_link_keys(&plan, &previous, ¤t, Clone::clone); - - assert!(matches!( - result, - Err(BatchPlanError::MissingBatch { - side: "previous", - .. - }) - )); - } - #[test] fn stored_batch_should_round_trip_without_source_text() { let original = batch("src/routes.rs", "hash", &["GET:/orders", "POST:/orders"]); diff --git a/crates/code-system-graph-core/src/client_flows.rs b/crates/code-system-graph-core/src/client_flows.rs new file mode 100644 index 0000000..45a06eb --- /dev/null +++ b/crates/code-system-graph-core/src/client_flows.rs @@ -0,0 +1,601 @@ +//! Repository-level composition of HTTP client calls through wrappers, helpers, and fixtures. +//! +//! Extractors record client calls whose URL depends on parameters of the enclosing function, and +//! the calls between functions with their string arguments. Composition first instantiates those +//! wrapper URLs at every call site that binds them, also through wrappers of wrappers. It then +//! attributes to each test every client call reached through the helpers and fixtures it calls, +//! so that a test is linked to the endpoints it exercises indirectly. +//! +//! A client call issued through a parameter of its function, such as a pytest `client` fixture, +//! is confirmed only when the parameter is a fixture returning an in-process test client, or is +//! passed such a fixture by every resolved caller path that reaches it. + +use std::collections::{BTreeMap, BTreeSet}; + +use crate::repository_symbols::{ + RepositorySourceFile, RepositorySymbols, SymbolFile, SymbolKey, files_by_repository +}; +use crate::source_http::instantiated_consumer; +use crate::url_template::BoundArguments; +use crate::{ + SourceEpistemicStatus, SourceFramework, SourceObservation, SourceRole, SymbolRef, UrlPart, UrlTemplate +}; + +/// Rounds of wrapper instantiation, which is also the longest chain of wrappers that pass a URL +/// parameter through to another wrapper. +const MAX_WRAPPER_HOPS: usize = 3; +/// Longest call chain followed from a test to the client calls of its helpers. +const MAX_TEST_CALL_DEPTH: usize = 3; +/// Most unresolved wrapper URLs retained for one function. +const MAX_OPEN_URLS: usize = 16; + +/// A named function of one file, identified by the file's index within its repository. +type Function<'a> = (usize, &'a str); + +/// Composes client calls through wrappers, helpers, and fixtures. +/// +/// The result has one observation list per input file, in input order. Call observations are +/// consumed; each function that binds a wrapper URL to an exact path gains a consumer +/// observation, and each test gains the client calls of the helpers and fixtures it reaches. +#[must_use] +pub fn compose_client_flows(files: &[RepositorySourceFile<'_>]) -> Vec> { + let mut output = files + .iter() + .map(|file| file.observations.to_vec()) + .collect::>(); + for members in files_by_repository(files).into_values() { + let composable = members.iter().any(|&index| { + files[index] + .observations + .iter() + .any(|observation| observation.role == SourceRole::Call || is_received(observation)) + }); + if !composable { + continue; + } + let flows = RepositoryFlows::new(files, &members); + for (local, additions) in flows.compose() { + output[members[local]].extend(additions); + } + } + for observations in &mut output { + observations.retain(|observation| { + !matches!(observation.role, SourceRole::Call | SourceRole::Client) + && !is_received(observation) + }); + } + output +} + +/// Whether `observation` is a client call issued through a parameter of its function. +fn is_received(observation: &SourceObservation) -> bool { + observation.role == SourceRole::Consumer + && matches!(observation.router, Some(SymbolRef::Parameter { .. })) +} + +/// A client call whose URL still depends on parameters of the function issuing it. +#[derive(Clone)] +struct OpenUrl<'a> { + origin: &'a SourceObservation, + url: UrlTemplate, +} + +struct RepositoryFlows<'a> { + symbols: RepositorySymbols<'a>, + functions: BTreeMap, Option>>, + locals: BTreeMap<(usize, String), Function<'a>>, + consumers: BTreeMap, Vec<&'a SourceObservation>>, + calls: BTreeMap, Vec<&'a SourceObservation>>, + callers: BTreeMap, Vec<(Function<'a>, &'a SourceObservation)>>, + tests: BTreeSet>, + clients: BTreeMap<(usize, String), SourceFramework>, + received: BTreeMap, Vec>, +} + +impl<'a> RepositoryFlows<'a> { + fn new(files: &'a [RepositorySourceFile<'a>], members: &[usize]) -> Self { + let mut symbols = RepositorySymbols::new( + members + .iter() + .map(|&index| SymbolFile { + path: files[index].path, + observations: files[index].observations, + }) + .collect(), + ); + let mut consumers = BTreeMap::, Vec<&'a SourceObservation>>::new(); + let mut calls = BTreeMap::, Vec<&'a SourceObservation>>::new(); + let mut tests = BTreeSet::new(); + let mut clients = BTreeMap::new(); + for (local, &index) in members.iter().enumerate() { + for observation in files[index].observations { + let Some(name) = observation.symbol_name.as_deref() else { + continue; + }; + match observation.role { + SourceRole::Consumer => consumers + .entry((local, name)) + .or_default() + .push(observation), + SourceRole::Call => calls.entry((local, name)).or_default().push(observation), + SourceRole::Test => { + tests.insert((local, name)); + } + SourceRole::Client => { + clients.insert((local, name.to_owned()), observation.framework); + } + SourceRole::Provider | SourceRole::Factory | SourceRole::Mount => {} + } + } + } + let mut functions = BTreeMap::, Option>>::new(); + let mut locals = BTreeMap::new(); + for &function in consumers.keys().chain(calls.keys()) { + locals.insert((function.0, function.1.to_owned()), function); + let key = symbols.declare_function(function.0, function.1); + functions + .entry(key) + .and_modify(|declared| { + if *declared != Some(function) { + *declared = None; + } + }) + .or_insert(Some(function)); + } + let mut flows = Self { + symbols, + functions, + locals, + consumers, + calls, + callers: BTreeMap::new(), + tests, + clients, + received: BTreeMap::new(), + }; + for (&caller, calls) in &flows.calls { + for &call in calls { + if let Some(callee) = flows.callee(caller.0, call) { + flows + .callers + .entry(callee) + .or_default() + .push((caller, call)); + } + } + } + flows.received = flows.resolve_received(); + flows + } + + /// Client calls issued through a parameter that receives an in-process test client. + fn resolve_received(&self) -> BTreeMap, Vec> { + let mut received = BTreeMap::, Vec>::new(); + for (&function, consumers) in &self.consumers { + for &consumer in consumers { + let Some(SymbolRef::Parameter { name, index }) = &consumer.router else { + continue; + }; + let Some(framework) = self.received_client(function, name, *index, 0) else { + continue; + }; + let mut confirmed = consumer.clone(); + confirmed.framework = framework; + confirmed.router = None; + if confirmed.method.is_some() && confirmed.path.is_some() { + confirmed.status = SourceEpistemicStatus::Confirmed; + confirmed.confidence = 1.0; + } + received.entry(function).or_default().push(confirmed); + } + } + received + } + + /// Test client framework passed to parameter `name` at position `index` of `function`. + fn received_client( + &self, + function: Function<'a>, + name: &str, + index: usize, + depth: usize, + ) -> Option { + let fixture = SymbolRef::Fixture(name.to_owned()); + let requests_fixture = self.calls.get(&function).is_some_and(|calls| { + calls.iter().any(|call| { + call.call + .as_ref() + .is_some_and(|call| call.callee == fixture) + }) + }); + if requests_fixture { + let Some(SymbolKey::Local { file, name }) = self.symbols.resolve(function.0, &fixture) + else { + return None; + }; + return self.clients.get(&(file, name)).copied(); + } + if depth >= MAX_TEST_CALL_DEPTH { + return None; + } + let mut frameworks = self.callers.get(&function)?.iter().map(|(caller, call)| { + let arguments = bound_arguments(call); + let argument = arguments + .by_name + .get(name) + .or_else(|| arguments.by_index.get(&index))?; + match argument.parts.as_slice() { + [UrlPart::Parameter { name, index }] => { + self.received_client(*caller, name, *index, depth + 1) + } + _ => None, + } + }); + let first = frameworks.next()??; + frameworks + .all(|framework| framework == Some(first)) + .then_some(first) + } + + /// Function called by `call` from file `local`, when it is known in this repository. + fn callee(&self, local: usize, call: &SourceObservation) -> Option> { + let reference = &call.call.as_ref()?.callee; + match self.symbols.resolve(local, reference)? { + SymbolKey::Local { file, name } => self.locals.get(&(file, name)).copied(), + SymbolKey::Function(key) => self.functions.get(&key).copied().flatten(), + } + } + + /// Observations added to each file, keyed by the file's index within the repository. + fn compose(&self) -> BTreeMap> { + let resolved = self.instantiate_wrappers(); + let mut additions = BTreeMap::>::new(); + for (function, observations) in self.received.iter().chain(&resolved) { + additions + .entry(function.0) + .or_default() + .extend(observations.iter().cloned()); + } + for &test in &self.tests { + let mut seen = self + .exact_consumers(test, &resolved) + .into_iter() + .map(consumer_identity) + .collect::>(); + for observation in self.reached_consumers(test, &resolved) { + if seen.insert(consumer_identity(&observation)) { + additions.entry(test.0).or_default().push(observation); + } + } + } + additions + } + + /// Consumers obtained by binding wrapper URLs at call sites, keyed by the calling function. + fn instantiate_wrappers(&self) -> BTreeMap, Vec> { + let mut open = BTreeMap::, Vec>>::new(); + for (&function, consumers) in &self.consumers { + for &origin in consumers { + if let (Some(url), None) = (&origin.url, &origin.path) { + open.entry(function).or_default().push(OpenUrl { + origin, + url: url.clone(), + }); + } + } + } + let mut resolved = BTreeMap::, Vec>::new(); + for _ in 0..MAX_WRAPPER_HOPS { + let snapshot = open.clone(); + let mut changed = false; + for (&caller, calls) in &self.calls { + for &call in calls { + let Some(callee) = self + .callee(caller.0, call) + .filter(|callee| *callee != caller) + else { + continue; + }; + let Some(wrappers) = snapshot.get(&callee) else { + continue; + }; + let arguments = bound_arguments(call); + for wrapper in wrappers { + let url = wrapper.url.bind(&arguments); + if let Some(consumer) = + instantiated_consumer(wrapper.origin, &url, caller.1, call.lines) + { + let consumers = resolved.entry(caller).or_default(); + let identity = consumer_identity(&consumer); + if !consumers + .iter() + .any(|known| consumer_identity(known) == identity) + { + consumers.push(consumer); + changed = true; + } + } else if url.has_parameters() { + let urls = open.entry(caller).or_default(); + if urls.len() < MAX_OPEN_URLS + && !urls.iter().any(|known| known.url == url) + { + urls.push(OpenUrl { + origin: wrapper.origin, + url, + }); + changed = true; + } + } + } + } + } + if !changed { + break; + } + } + resolved + } + + /// Exact consumers issued by `function`, directly or through an instantiated wrapper. + fn exact_consumers<'r>( + &'r self, + function: Function<'a>, + resolved: &'r BTreeMap, Vec>, + ) -> Vec<&'r SourceObservation> + where + 'a: 'r, + { + self.consumers + .get(&function) + .into_iter() + .flatten() + .copied() + .filter(|observation| { + observation.status == SourceEpistemicStatus::Confirmed + && observation.method.is_some() + && observation.path.is_some() + }) + .chain(self.received.get(&function).into_iter().flatten()) + .chain(resolved.get(&function).into_iter().flatten()) + .collect() + } + + /// Exact consumers of the functions reached from `test`, attributed to `test` at the call + /// site in the test through which each one is reached. + fn reached_consumers( + &self, + test: Function<'a>, + resolved: &BTreeMap, Vec>, + ) -> Vec { + let mut reached = Vec::new(); + let mut visited = BTreeSet::from([test]); + let mut frontier = self + .calls + .get(&test) + .into_iter() + .flatten() + .filter_map(|call| Some((self.callee(test.0, call)?, call.lines))) + .collect::>(); + for _ in 0..MAX_TEST_CALL_DEPTH { + let mut next = Vec::new(); + for (function, lines) in frontier { + if !visited.insert(function) { + continue; + } + for consumer in self.exact_consumers(function, resolved) { + let mut attributed = consumer.clone(); + attributed.symbol_name = Some(test.1.to_owned()); + attributed.lines = lines; + attributed.url = None; + reached.push(attributed); + } + next.extend( + self.calls + .get(&function) + .into_iter() + .flatten() + .filter_map(|call| Some((self.callee(function.0, call)?, lines))), + ); + } + frontier = next; + } + reached + } +} + +/// Arguments of `call` keyed by keyword, or by position among the positional arguments. +fn bound_arguments(call: &SourceObservation) -> BoundArguments { + let mut bound = BoundArguments::default(); + let mut position = 0; + for argument in call.call.iter().flat_map(|call| &call.arguments) { + match (&argument.keyword, &argument.value) { + (Some(keyword), Some(value)) => { + bound.by_name.insert(keyword.clone(), value.clone()); + } + (Some(_), None) => {} + (None, value) => { + if let Some(value) = value { + bound.by_index.insert(position, value.clone()); + } + position += 1; + } + } + } + bound +} + +/// Method, path, and authority of a consumer. +type ConsumerIdentity = (Option, Option, Option); + +fn consumer_identity(observation: &SourceObservation) -> ConsumerIdentity { + ( + observation.method.clone(), + observation.path.clone(), + observation.authority.clone(), + ) +} + +#[cfg(test)] +mod tests { + use code_system_graph_model::RepoId; + + use super::compose_client_flows; + use crate::{ + RepositorySourceFile, SourceObservation, SourceRole, parse_python_source, parse_typescript_source + }; + + /// Consumers after composition as `path: symbol METHOD authority path`. + fn composed_consumers(files: &[(&str, Vec)]) -> Vec { + let repo = RepoId::new("repo:shop"); + let inputs = files + .iter() + .map(|(path, observations)| RepositorySourceFile { + repo_id: &repo, + path, + observations, + }) + .collect::>(); + let composed = compose_client_flows(&inputs); + assert!( + composed + .iter() + .flatten() + .all(|item| item.role != SourceRole::Call) + ); + files + .iter() + .zip(composed) + .flat_map(|((path, _), observations)| { + observations + .into_iter() + .filter(|item| item.role == SourceRole::Consumer && item.path.is_some()) + .map(move |item| { + format!( + "{path}: {} {} {} {}", + item.symbol_name.unwrap_or_default(), + item.method.unwrap_or_default(), + item.authority.unwrap_or_default(), + item.path.unwrap_or_default() + ) + }) + }) + .collect() + } + + #[test] + fn python_tests_should_reach_fixture_and_wrapped_helper_calls_across_files() { + let conftest = r#" +import pytest +import requests +BASE = "http://orders:8000" + +@pytest.fixture +def order(): + return requests.post(BASE + "/orders", json={}).json() +"#; + let helpers = r#" +import requests +API = "http://orders:8000/api" + +def get_json(path): + return requests.get(API + path).json() + +def read_order(order_id): + return get_json(f"/orders/{order_id}") +"#; + let test = r#" +from tests.helpers import read_order + +def test_read(order): + read_order(order["id"]) +"#; + let consumers = composed_consumers(&[ + ("tests/conftest.py", parse_python_source(conftest)), + ("tests/helpers.py", parse_python_source(helpers)), + ("tests/test_orders.py", parse_python_source(test)), + ]); + + assert_eq!( + consumers, + [ + "tests/conftest.py: order POST orders:8000 /orders", + "tests/helpers.py: read_order GET orders:8000 /api/orders/{order_id}", + "tests/test_orders.py: test_read POST orders:8000 /orders", + "tests/test_orders.py: test_read GET orders:8000 /api/orders/{order_id}", + ] + ); + } + + #[test] + fn python_parameter_clients_should_resolve_through_test_client_fixtures() { + let conftest = r" +import pytest +from fastapi.testclient import TestClient +from app.main import app + +@pytest.fixture +def client(): + with TestClient(app) as test_client: + yield test_client +"; + let helpers = r#" +def create_order(client, sku): + return client.post("/orders", json={"sku": sku}) + +def health(session): + return session.get("/health") +"#; + let test = r#" +from tests.helpers import create_order, health + +def test_create(client): + create_order(client, "A-1") + +def test_read(client): + client.get("/orders/42") + +def test_health(): + health(object()) +"#; + let files = [ + ("tests/conftest.py", parse_python_source(conftest)), + ("tests/helpers.py", parse_python_source(helpers)), + ("tests/test_orders.py", parse_python_source(test)), + ]; + let consumers = composed_consumers(&files); + + assert_eq!( + consumers, + [ + "tests/helpers.py: create_order POST /orders", + "tests/test_orders.py: test_read GET /orders/42", + "tests/test_orders.py: test_create POST /orders", + ] + ); + } + + #[test] + fn script_wrappers_should_be_instantiated_at_call_sites_in_other_modules() { + let api = r#" +const API = "/api"; +export function request(path: string) { + return fetch(API + path); +} +"#; + let orders = r#" +import { request } from "./api"; +export function getOrder(id: string) { + return request(`/orders/${id}`); +} +"#; + let consumers = composed_consumers(&[ + ("web/src/api.ts", parse_typescript_source(api)), + ("web/src/orders.ts", parse_typescript_source(orders)), + ]); + + assert_eq!( + consumers, + ["web/src/orders.ts: getOrder GET /api/orders/{id}"] + ); + } +} diff --git a/crates/code-system-graph-core/src/codegraph/contract.rs b/crates/code-system-graph-core/src/codegraph/contract.rs index 9c6d328..a75baf3 100644 --- a/crates/code-system-graph-core/src/codegraph/contract.rs +++ b/crates/code-system-graph-core/src/codegraph/contract.rs @@ -243,9 +243,16 @@ pub(crate) fn cli_operations(version: &str) -> Vec .collect() } +/// The structured CLI adapter is validated against `CodeGraph` 1.6.1; pre-releases and other minor +/// versions are rejected until their fixtures are captured. pub(crate) fn supports_cli_contract(version: &str) -> bool { - let version = version.trim().trim_start_matches('v'); - version == "1.5" || version.starts_with("1.5.") + let mut parts = version.trim().trim_start_matches('v').split('.'); + let (Some(major), Some(minor), Some(patch), None) = + (parts.next(), parts.next(), parts.next(), parts.next()) + else { + return false; + }; + major == "1" && minor == "6" && patch.parse::().is_ok_and(|patch| patch >= 1) } fn schema_parameters(schema: &Value) -> Vec { @@ -270,13 +277,17 @@ fn contains_any(value: &str, candidates: &[&str]) -> bool { #[cfg(test)] mod tests { - use super::{DiscoveredTool, StatusContract, map_mcp_tools, supports_cli_contract}; + use serde_json::Value; + + use super::{ + AffectedTestsContract, DiscoveredTool, ImpactContract, NeighborsContract, StatusContract, SymbolQueryContract, map_mcp_tools, supports_cli_contract + }; use crate::{ProviderOperation, ProviderStatus}; #[test] - fn codegraph_1_5_tools_should_map_context_by_name_and_schema() { + fn codegraph_1_6_tools_should_map_context_by_name_and_schema() { let tools: Vec = serde_json::from_str(include_str!( - "../../../../fixtures/codegraph/1.5.0/tools-list.json" + "../../../../fixtures/codegraph/1.6.1/tools-list.json" )) .expect("fixture should be valid"); let operations = map_mcp_tools(&tools); @@ -286,6 +297,61 @@ mod tests { assert!(operations[0].parameters.contains(&"projectPath".to_owned())); } + #[test] + fn codegraph_1_6_initialize_should_negotiate_the_tested_protocol() { + let initialize: Value = serde_json::from_str(include_str!( + "../../../../fixtures/codegraph/1.6.1/initialize.json" + )) + .expect("fixture should be valid"); + + assert_eq!(initialize["protocolVersion"], "2024-11-05"); + assert_eq!(initialize["serverInfo"]["version"], "1.6.1"); + assert!(initialize["capabilities"]["tools"].is_object()); + } + + #[test] + fn codegraph_1_6_cli_json_should_match_the_structured_contracts() { + let symbols: Vec = serde_json::from_str(include_str!( + "../../../../fixtures/codegraph/1.6.1/query.json" + )) + .expect("query fixture should parse"); + let neighbors: NeighborsContract = serde_json::from_str(include_str!( + "../../../../fixtures/codegraph/1.6.1/neighbors.json" + )) + .expect("neighbors fixture should parse"); + let impact: ImpactContract = serde_json::from_str(include_str!( + "../../../../fixtures/codegraph/1.6.1/impact.json" + )) + .expect("impact fixture should parse"); + let affected: AffectedTestsContract = serde_json::from_str(include_str!( + "../../../../fixtures/codegraph/1.6.1/affected.json" + )) + .expect("affected fixture should parse"); + + let symbol = symbols + .into_iter() + .next() + .expect("query fixture should resolve one symbol") + .into_symbol(); + assert_eq!( + ( + symbol.name.as_str(), + symbol.file_path.as_str(), + symbol.start_line + ), + ("scan_workspace", "src/lib.rs", 3) + ); + assert_eq!(neighbors.symbol, "discover"); + assert_eq!(neighbors.callers.len(), 1); + assert_eq!(neighbors.callers[0].name, "scan_workspace"); + assert!(neighbors.callees.is_empty()); + assert_eq!((impact.symbol.as_str(), impact.depth), ("discover", 2)); + assert_eq!(impact.node_count, impact.affected.len()); + assert_eq!(affected.changed_files, ["src/provider.ts"]); + assert_eq!(affected.affected_tests, ["tests/provider.test.ts"]); + assert_eq!(affected.total_dependents_traversed, 1); + } + #[test] fn missing_optional_tool_should_not_create_false_capability() { let tools: Vec = serde_json::from_str(include_str!( @@ -312,11 +378,11 @@ mod tests { } #[test] - fn status_contract_should_accept_null_index_state_from_codegraph_1_5() { + fn status_contract_should_treat_a_null_index_state_as_stale() { let status: StatusContract = serde_json::from_str( r#"{ "initialized": true, - "version": "1.5.0", + "version": "1.6.1", "worktreeMismatch": null, "pendingChanges": {"added": 0, "modified": 0, "removed": 0}, "index": {"reindexRecommended": true, "state": null} @@ -330,16 +396,36 @@ mod tests { #[test] fn status_contract_should_recognize_current_index() { let status: StatusContract = serde_json::from_str(include_str!( - "../../../../fixtures/codegraph/1.5.0/status.json" + "../../../../fixtures/codegraph/1.6.1/status.json" )) .expect("fixture should be valid"); + assert_eq!(status.version.as_deref(), Some("1.6.1")); assert_eq!(status.status(), ProviderStatus::Available); } #[test] - fn cli_contract_should_reject_unvalidated_versions() { - assert!(supports_cli_contract("1.5.0")); - assert!(!supports_cli_contract("1.6.0")); + fn cli_contract_should_accept_only_validated_versions() { + for version in ["1.6.1", "v1.6.1", "1.6.2", " 1.6.10\n"] { + assert!( + supports_cli_contract(version), + "{version} should be accepted" + ); + } + for version in [ + "1.5.0", + "1.6", + "1.6.0", + "1.6.1-rc.1", + "1.6.1.0", + "1.7.0", + "2.6.1", + "", + ] { + assert!( + !supports_cli_contract(version), + "{version} should be rejected" + ); + } } } diff --git a/crates/code-system-graph-core/src/codegraph/mcp.rs b/crates/code-system-graph-core/src/codegraph/mcp.rs index af17174..2ff47b0 100644 --- a/crates/code-system-graph-core/src/codegraph/mcp.rs +++ b/crates/code-system-graph-core/src/codegraph/mcp.rs @@ -8,7 +8,7 @@ use std::task::{Context, Poll}; use std::time::Duration; use rmcp::model::{ - CallToolRequestParams, ClientCapabilities, ClientInfo, Implementation, JsonObject + CallToolRequestParams, ClientCapabilities, ClientConfig, Implementation, JsonObject }; use rmcp::service::RunningService; use rmcp::{RoleClient, ServiceExt}; @@ -38,7 +38,7 @@ pub(crate) struct CodeGraphMcp { binary: OsString, } -type McpService = RunningService; +type McpService = RunningService; struct McpConnection { service: McpService, @@ -227,7 +227,7 @@ impl CodeGraphMcp { .max_output_bytes .saturating_add(PROTOCOL_OVERHEAD_BYTES); let reader = LimitedAsyncRead::new(stdout, protocol_limit, protocol_exceeded.clone()); - let client = ClientInfo::new( + let client = ClientConfig::new( ClientCapabilities::default(), Implementation::new("code-system-graph", env!("CARGO_PKG_VERSION")), ); diff --git a/crates/code-system-graph-core/src/communities.rs b/crates/code-system-graph-core/src/communities.rs index 92b5117..9ed2d9e 100644 --- a/crates/code-system-graph-core/src/communities.rs +++ b/crates/code-system-graph-core/src/communities.rs @@ -66,7 +66,6 @@ struct AcceptedEdge { struct WeightedGraph { nodes: Vec, adjacency: Vec>, - undirected_edges: Vec<(usize, usize, f64)>, accepted_edges: Vec, weighted_degree: Vec, neighbor_count: Vec, @@ -496,7 +495,6 @@ fn build_graph( WeightedGraph { nodes: graph_nodes, adjacency, - undirected_edges, accepted_edges, weighted_degree, neighbor_count, @@ -578,6 +576,10 @@ where assignments } +/// Local-moving Louvain phase. Each move is scored by its exact modularity gain: for node `i` with +/// weighted degree `k`, links `k_c` to community `c`, and community degree totals `d_c`, moving from +/// `a` to `b` changes modularity by +/// `(k_b - k_a) / m - resolution * k * (d_b - (d_a - k)) / (2 * m^2)`. fn louvain( graph: &WeightedGraph, seed: u64, @@ -588,43 +590,62 @@ fn louvain( where F: FnMut(u64), { - let mut assignments: Vec = (0..graph.nodes.len()).collect(); - if graph.total_undirected_weight <= EPSILON { + let count = graph.nodes.len(); + let mut assignments: Vec = (0..count).collect(); + let total = graph.total_undirected_weight; + if total <= EPSILON { return assignments; } + let label_keys = (0..count) + .map(|label| seeded_label_key(graph, seed, label)) + .collect::>(); + let mut community_degree = graph.weighted_degree.clone(); + let mut members = vec![1_usize; count]; + let mut empty = BTreeSet::new(); let order = seeded_node_order(graph, seed); for _ in 0..max_iterations { let mut changed = false; for &node in &order { let current = assignments[node]; - let baseline = modularity(graph, &assignments, resolution); - let mut candidates = BTreeSet::from([current]); - for &neighbor in graph.adjacency[node].keys() { - candidates.insert(assignments[neighbor]); + let degree = graph.weighted_degree[node]; + let mut links = BTreeMap::from([(current, 0.0)]); + for (&neighbor, &weight) in &graph.adjacency[node] { + if neighbor != node { + *links.entry(assignments[neighbor]).or_insert(0.0) += weight; + } } - if let Some(empty) = first_empty_label(&assignments) { - candidates.insert(empty); + if let Some(&label) = empty.first() { + links.entry(label).or_insert(0.0); } + let current_links = links[¤t]; + let source_degree = community_degree[current] - degree; let mut best = current; - let mut best_modularity = baseline; - for candidate in candidates { + let mut best_gain = 0.0; + for (&candidate, &candidate_links) in &links { if candidate == current { continue; } - assignments[node] = candidate; - let candidate_modularity = modularity(graph, &assignments, resolution); - assignments[node] = current; - if candidate_modularity > best_modularity + EPSILON - || ((candidate_modularity - best_modularity).abs() <= EPSILON - && candidate_modularity > baseline + EPSILON - && seeded_label_key(graph, seed, candidate) - < seeded_label_key(graph, seed, best)) + let gain = (candidate_links - current_links) / total + - resolution * degree * (community_degree[candidate] - source_degree) + / (2.0 * total * total); + if gain > best_gain + EPSILON + || ((gain - best_gain).abs() <= EPSILON + && gain > EPSILON + && label_keys[candidate] < label_keys[best]) { best = candidate; - best_modularity = candidate_modularity; + best_gain = gain; } } - if best != current && best_modularity > baseline + EPSILON { + if best != current && best_gain > EPSILON { + community_degree[current] -= degree; + community_degree[best] += degree; + members[current] -= 1; + members[best] += 1; + if members[current] == 0 { + empty.insert(current); + } + empty.remove(&best); assignments[node] = best; changed = true; } @@ -637,42 +658,6 @@ where assignments } -/// Computes generalized undirected modularity: -/// `sum_c(internal_weight_c / m - resolution * (degree_c / (2m))^2)`. -fn modularity(graph: &WeightedGraph, assignments: &[usize], resolution: f64) -> f64 { - let total = graph.total_undirected_weight; - if total <= EPSILON { - return 0.0; - } - let mut internal = BTreeMap::::new(); - let mut degree = BTreeMap::::new(); - for (node, &community) in assignments.iter().enumerate() { - *degree.entry(community).or_default() += graph.weighted_degree[node]; - } - for &(source, target, weight) in &graph.undirected_edges { - if assignments[source] == assignments[target] { - *internal.entry(assignments[source]).or_default() += weight; - } - } - degree - .into_iter() - .map(|(community, community_degree)| { - let inside = internal.get(&community).copied().unwrap_or_default(); - inside / total - resolution * (community_degree / (2.0 * total)).powi(2) - }) - .sum() -} - -fn first_empty_label(assignments: &[usize]) -> Option { - let mut used = vec![false; assignments.len()]; - for &assignment in assignments { - if let Some(slot) = used.get_mut(assignment) { - *slot = true; - } - } - used.iter().position(|value| !value) -} - fn seeded_node_order(graph: &WeightedGraph, seed: u64) -> Vec { let mut order: Vec = (0..graph.nodes.len()).collect(); order.sort_by(|&left, &right| { @@ -1532,6 +1517,149 @@ mod tests { )); } + /// Computes generalized undirected modularity: + /// `sum_c(internal_weight_c / m - resolution * (degree_c / (2m))^2)`. + fn modularity(graph: &WeightedGraph, assignments: &[usize], resolution: f64) -> f64 { + let total = graph.total_undirected_weight; + if total <= EPSILON { + return 0.0; + } + let mut internal = BTreeMap::::new(); + let mut degree = BTreeMap::::new(); + for (node, &community) in assignments.iter().enumerate() { + *degree.entry(community).or_default() += graph.weighted_degree[node]; + } + for (source, neighbors) in graph.adjacency.iter().enumerate() { + for (&target, &weight) in neighbors.range(source..) { + if assignments[source] == assignments[target] { + *internal.entry(assignments[source]).or_default() += weight; + } + } + } + degree + .into_iter() + .map(|(community, community_degree)| { + let inside = internal.get(&community).copied().unwrap_or_default(); + inside / total - resolution * (community_degree / (2.0 * total)).powi(2) + }) + .sum() + } + + fn first_empty_label(assignments: &[usize]) -> Option { + let mut used = vec![false; assignments.len()]; + for &assignment in assignments { + if let Some(slot) = used.get_mut(assignment) { + *slot = true; + } + } + used.iter().position(|value| !value) + } + + /// Louvain local moving that recomputes full modularity for every candidate move. + fn reference_louvain( + graph: &WeightedGraph, + seed: u64, + resolution: f64, + max_iterations: u32, + ) -> Vec { + let mut assignments: Vec = (0..graph.nodes.len()).collect(); + if graph.total_undirected_weight <= EPSILON { + return assignments; + } + let order = seeded_node_order(graph, seed); + for _ in 0..max_iterations { + let mut changed = false; + for &node in &order { + let current = assignments[node]; + let baseline = modularity(graph, &assignments, resolution); + let mut candidates = BTreeSet::from([current]); + for &neighbor in graph.adjacency[node].keys() { + candidates.insert(assignments[neighbor]); + } + if let Some(empty) = first_empty_label(&assignments) { + candidates.insert(empty); + } + let mut best = current; + let mut best_modularity = baseline; + for candidate in candidates { + if candidate == current { + continue; + } + assignments[node] = candidate; + let candidate_modularity = modularity(graph, &assignments, resolution); + assignments[node] = current; + if candidate_modularity > best_modularity + EPSILON + || ((candidate_modularity - best_modularity).abs() <= EPSILON + && candidate_modularity > baseline + EPSILON + && seeded_label_key(graph, seed, candidate) + < seeded_label_key(graph, seed, best)) + { + best = candidate; + best_modularity = candidate_modularity; + } + } + if best != current && best_modularity > baseline + EPSILON { + assignments[node] = best; + changed = true; + } + } + if !changed { + break; + } + } + assignments + } + + fn pseudo_random_graph(seed: u64) -> (Vec, Vec) { + let mut state = seed; + let mut next = move |bound: u64| { + state = state + .wrapping_mul(6_364_136_223_846_793_005) + .wrapping_add(1_442_695_040_888_963_407); + (state >> 33) % bound + }; + let nodes = (0..60) + .map(|index| { + let id = format!("n{index:02}"); + node(&id, NodeKind::Service, "repo:one", &id) + }) + .collect::>(); + let kinds = [ + EdgeKind::CallsRemote, + EdgeKind::ManualLink, + EdgeKind::Consumes, + EdgeKind::Validates, + ]; + let edges = (0..180) + .map(|number| { + let source = format!("n{:02}", next(60)); + let target = format!("n{:02}", next(60)); + let kind = kinds[usize::try_from(next(4)).expect("small index")]; + let confidence = 0.5 + f32::from(u8::try_from(next(50)).expect("small")) / 100.0; + edge(&format!("e{number}"), &source, &target, kind, confidence) + }) + .collect(); + (nodes, edges) + } + + #[test] + fn incremental_louvain_should_match_full_modularity_recomputation() { + for graph_seed in [1, 2, 3, 4] { + let (nodes, edges) = pseudo_random_graph(graph_seed); + let config = config(CommunityAlgorithm::Louvain); + let weights = validate_inputs(&nodes, &edges, &config).expect("weights"); + let selected = scoped_node_ids(&nodes, &edges, &config, &weights).expect("scope"); + let graph = build_graph(&nodes, &edges, &config, &weights, &selected); + for (seed, resolution) in [(17, 1.0), (3, 0.5), (99, 2.0)] { + assert_eq!( + louvain(&graph, seed, resolution, 100, &mut |_| {}), + reference_louvain(&graph, seed, resolution, 100), + "graph {graph_seed}, seed {seed}, resolution {resolution}" + ); + } + } + } + #[test] fn completed_community_units_should_report_progress() { let (nodes, edges) = two_cluster_graph(); diff --git a/crates/code-system-graph-core/src/config.rs b/crates/code-system-graph-core/src/config.rs index 8300355..f0824bf 100644 --- a/crates/code-system-graph-core/src/config.rs +++ b/crates/code-system-graph-core/src/config.rs @@ -563,6 +563,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let resolved = resolve_repository_config(repository.path(), &workspace)?; @@ -600,6 +601,7 @@ mod tests { implementations: None, excludes: Some(vec!["workspace/**".to_owned()]), include_defaults: Some(Vec::new()), + authorities: None, }; let resolved = resolve_repository_config(repository.path(), &workspace)?; @@ -637,6 +639,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let resolved = resolve_repository_config_with_use_gitignore( @@ -669,6 +672,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let resolved = resolve_repository_config(repository.path(), &workspace)?; @@ -697,6 +701,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let resolved = resolve_repository_config_with_use_gitignore( @@ -721,10 +726,12 @@ mod tests { implementations: None, excludes: Some(vec!["coverage/**".to_owned()]), include_defaults: Some(vec!["vendor/internal-sdk/**".to_owned()]), + authorities: None, }; let redundant = RepositoryConfig { excludes: Some(vec!["./coverage//./**".to_owned()]), include_defaults: Some(vec!["./vendor//internal-sdk/./**".to_owned()]), + authorities: None, ..canonical.clone() }; @@ -754,6 +761,7 @@ mod tests { implementations: None, excludes: None, include_defaults: Some(vec!["vendor/internal-sdk/**".to_owned()]), + authorities: None, }; let resolved = resolve_repository_config(repository.path(), &workspace)?; @@ -786,6 +794,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let resolved = resolve_repository_config(repository.path(), &workspace)?; @@ -815,6 +824,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let result = resolve_repository_config(repository.path(), &workspace); @@ -839,6 +849,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let result = resolve_repository_config(repository.path(), &workspace); @@ -866,6 +877,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let result = resolve_repository_config(&repository, &workspace); @@ -895,6 +907,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let result = resolve_repository_config(repository.path(), &workspace); @@ -916,6 +929,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let result = resolve_repository_config(repository.path(), &workspace); @@ -935,6 +949,7 @@ mod tests { implementations: None, excludes: None, include_defaults: None, + authorities: None, }; let mut resolved = resolve_repository_config(repository.path(), &workspace)?; diff --git a/crates/code-system-graph-core/src/data_contracts.rs b/crates/code-system-graph-core/src/data_contracts.rs index c28d93f..d7ca209 100644 --- a/crates/code-system-graph-core/src/data_contracts.rs +++ b/crates/code-system-graph-core/src/data_contracts.rs @@ -5,6 +5,7 @@ use std::collections::{BTreeMap, BTreeSet}; use std::path::Path; +use std::sync::LazyLock; use serde::{Deserialize, Serialize}; use sqlparser::ast::{ @@ -16,6 +17,7 @@ use thiserror::Error; use tree_sitter::Node as SyntaxNode; use crate::SourceLanguage; +use crate::markers::MarkerSet; /// Maximum source size accepted by the database artifact extractor. pub const MAX_DATA_INPUT_BYTES: usize = 1_048_576; @@ -397,15 +399,31 @@ pub fn parse_literal_sql_source_at_root( source_path: &str, crate_root: &str, input: &str, +) -> DataDocument { + parse_literal_sql_source_with_prefilter(language, source_path, crate_root, input, true) +} + +fn parse_literal_sql_source_with_prefilter( + language: SourceLanguage, + source_path: &str, + crate_root: &str, + input: &str, + prefilter: bool, ) -> DataDocument { let mut document = empty_document(source_path, DataArtifactKind::LiteralQuerySource); if validate_input(input).is_err() { mark_incomplete(&mut document, DataWarning::LimitExceeded); return document; } + if prefilter && !DATA_SOURCE_MARKERS.any_in(input) { + finish_document(&mut document); + return document; + } if language == SourceLanguage::Rust { - extract_rust_database_source(input, crate_root, &mut document); + if !prefilter || input.contains("sqlx") || input.contains("mysql_async") { + extract_rust_database_source(input, crate_root, &mut document); + } } else if language == SourceLanguage::Python { if input.contains("pymysql") { document.frameworks.push(DataFramework::PyMysql); @@ -3245,16 +3263,27 @@ fn adjacent_dynamic_operator(input: &str, start: usize, end: usize) -> bool { || after.starts_with(".format(") } +const QUERY_CONTEXT_MARKERS: [&str; 10] = [ + "query", "execute", "fetch", "select", ".sql(", "sql =", "sql!", "sqlx", "prepare", "raw", +]; + +// Every recognizer in `parse_literal_sql_source_at_root` requires one of these substrings: +// query-context markers for literals, dynamic-call markers, framework imports, and the Rust +// database crates. +static DATA_SOURCE_MARKERS: LazyLock = LazyLock::new(|| { + let mut markers = QUERY_CONTEXT_MARKERS.to_vec(); + markers.extend(["exec(", "pymysql", "psycopg", "sqlalchemy", "mysql_async"]); + MarkerSet::new(&markers) +}); + fn has_query_context(input: &str, literal: &SourceLiteral) -> bool { let context_start = floor_char_boundary(input, literal.start.saturating_sub(160)); let before = input[context_start..literal.start].to_ascii_lowercase(); let after_end = floor_char_boundary(input, (literal.end + 80).min(input.len())); let after = input[literal.end..after_end].to_ascii_lowercase(); - [ - "query", "execute", "fetch", "select", ".sql(", "sql =", "sql!", "sqlx", "prepare", "raw", - ] - .iter() - .any(|marker| before.contains(marker) || after.contains(marker)) + QUERY_CONTEXT_MARKERS + .iter() + .any(|marker| before.contains(marker) || after.contains(marker)) } fn floor_char_boundary(input: &str, mut index: usize) -> usize { @@ -3349,6 +3378,22 @@ fn java_method_name(line: &str) -> Option { #[cfg(test)] mod tests { + #[test] + fn source_marker_prefilter_should_not_change_any_corpus_document() { + let corpus = crate::markers::differential_corpus(); + assert!( + corpus.len() > 100, + "differential corpus is unexpectedly small" + ); + for (language, path, input) in corpus { + assert_eq!( + super::parse_literal_sql_source_with_prefilter(language, &path, ".", &input, true), + super::parse_literal_sql_source_with_prefilter(language, &path, ".", &input, false), + "data prefilter changed {path}" + ); + } + } + use super::*; #[test] diff --git a/crates/code-system-graph-core/src/events.rs b/crates/code-system-graph-core/src/events.rs index b890476..6514e3d 100644 --- a/crates/code-system-graph-core/src/events.rs +++ b/crates/code-system-graph-core/src/events.rs @@ -1,6 +1,7 @@ //! Conservative extraction of event contracts from `AsyncAPI` and source boundaries. use std::collections::BTreeSet; +use std::sync::LazyLock; use serde::{Deserialize, Serialize}; use serde_json::Value; @@ -9,6 +10,7 @@ type Mapping = serde_json::Map; use thiserror::Error; use crate::SourceLanguage; +use crate::markers::MarkerSet; const MAX_ASYNCAPI_BYTES: usize = 4 * 1024 * 1024; const MAX_SOURCE_BYTES: usize = 1024 * 1024; @@ -268,8 +270,15 @@ fn parse_asyncapi_value(source_path: &str, input: &str) -> Result EventDocument { + parse_event_source_with_prefilter(language, input, true) +} + +fn parse_event_source_with_prefilter( + language: SourceLanguage, + input: &str, + prefilter: bool, +) -> EventDocument { let (source, truncated) = bounded_source(input); - let sanitized = sanitize_source(language, source); let mut document = EventDocument::default(); if truncated { document.incomplete = true; @@ -283,6 +292,13 @@ pub fn parse_event_source(language: SourceLanguage, input: &str) -> EventDocumen .warnings .push("source input contains a NUL byte".to_owned()); } + // Comment removal can only join two tokens, which never forms a real call, so the raw text + // contains every marker that a sanitized candidate line can contain. + if prefilter && !EVENT_MARKERS.any_in(source) { + finish_document(&mut document); + return document; + } + let sanitized = sanitize_source(language, source); for index in 0..sanitized.len() { if document.observations.len() >= MAX_OBSERVATIONS { @@ -1797,25 +1813,26 @@ fn source_candidate(lines: &[String], start: usize) -> Option { Some(candidate) } +const EVENT_LINE_MARKERS: [&str; 13] = [ + "publish", + "subscribe", + "send", + "receive", + "consume", + "produce", + "queue", + "topic", + "writemessages", + "basic_", + "createstream", + "create_stream", + "pull(", +]; + +static EVENT_MARKERS: LazyLock = LazyLock::new(|| MarkerSet::new(&EVENT_LINE_MARKERS)); + fn possible_event_line(line: &str) -> bool { - let lower = line.to_ascii_lowercase(); - [ - "publish", - "subscribe", - "send", - "receive", - "consume", - "produce", - "queue", - "topic", - "writemessages", - "basic_", - "createstream", - "create_stream", - "pull(", - ] - .iter() - .any(|marker| lower.contains(marker)) + EVENT_MARKERS.any_in(line) } fn delimiter_balance(line: &str) -> i32 { @@ -2021,6 +2038,22 @@ fn observation_sort_key(observation: &EventObservation) -> String { #[cfg(test)] mod tests { + #[test] + fn source_marker_prefilter_should_not_change_any_corpus_document() { + let corpus = crate::markers::differential_corpus(); + assert!( + corpus.len() > 100, + "differential corpus is unexpectedly small" + ); + for (language, path, input) in corpus { + assert_eq!( + super::parse_event_source_with_prefilter(language, &input, true), + super::parse_event_source_with_prefilter(language, &input, false), + "event prefilter changed {path}" + ); + } + } + use super::*; fn source_observation(language: SourceLanguage, source: &str) -> EventObservation { diff --git a/crates/code-system-graph-core/src/execution_policy.rs b/crates/code-system-graph-core/src/execution_policy.rs index 5155dd8..32e2ad8 100644 --- a/crates/code-system-graph-core/src/execution_policy.rs +++ b/crates/code-system-graph-core/src/execution_policy.rs @@ -29,8 +29,13 @@ pub const DEFAULT_WATCH_IDLE_TIMEOUT_MS: u64 = 28_800_000; pub const DEFAULT_MAX_WATCH_SESSION_WALL_TIME_MS: u64 = 86_400_000; /// Default minimum delay between watched sync pass starts. pub const DEFAULT_MIN_WATCH_RESCAN_INTERVAL_MS: u64 = 10_000; -/// Default maximum retained historical checkpoint-cache bytes. +/// Default maximum retained checkpoint-cache bytes. pub const DEFAULT_MAX_CHECKPOINT_CACHE_BYTES: u64 = 10_737_418_240; +/// Default maximum concurrent extraction workers; the effective value never exceeds the host's +/// available parallelism. +pub const DEFAULT_MAX_EXTRACTION_WORKERS: u64 = 8; +/// Largest accepted extraction-worker count. +pub const MAX_EXTRACTION_WORKERS_CEILING: u64 = 256; pub const DEFAULT_MAX_EXPLORE_WALL_TIME_MS: u64 = 8_000; pub const DEFAULT_MAX_EXPLORE_CODEGRAPH_OPERATIONS: u64 = 8; pub const DEFAULT_MAX_EXPLORE_CONCURRENT_CODEGRAPH_PROCESSES: u64 = 2; @@ -53,10 +58,9 @@ pub const DEFAULT_MAX_MCP_SCHEMA_CATALOG_BYTES: u64 = 2_097_152; /// Smallest Markdown response budget that can retain the mandatory MCP control block. pub const MIN_MCP_MARKDOWN_BYTES: u64 = 256; -// Keep the agent-facing policy inventory in one declarative list. The two public serde structs -// intentionally remain flat for schema-v2 compatibility; defaults, override resolution, and the -// delivery fingerprint are generated from this list so a new field cannot silently omit one of -// those behaviors. +// Keep the agent-facing policy inventory in one declarative list. Defaults, override resolution, +// and the delivery fingerprint are generated from this list so a new field cannot silently omit +// one of those behaviors. macro_rules! with_agent_policy_fields { ($consumer:ident) => { $consumer! { @@ -182,6 +186,8 @@ pub struct ExecutionPolicyOverrides { pub min_watch_rescan_interval_ms: Option, /// Optional maximum retained checkpoint-cache bytes. pub max_checkpoint_cache_bytes: Option, + /// Optional maximum concurrent extraction workers. + pub max_extraction_workers: Option, /// Optional corroboration-anchor count, or `-1` for unlimited. #[serde(rename = "maxCodeGraphCorroborationAnchorsPerRepo")] pub max_codegraph_corroboration_anchors_per_repo: Option, @@ -248,8 +254,10 @@ pub struct ExecutionPolicy { pub max_watch_session_wall_time_ms: u64, /// Minimum delay between watched sync pass starts. pub min_watch_rescan_interval_ms: u64, - /// Maximum retained historical checkpoint-cache bytes. + /// Maximum retained checkpoint-cache bytes. pub max_checkpoint_cache_bytes: u64, + /// Maximum concurrent extraction workers. Output is identical for every value. + pub max_extraction_workers: u64, /// Effective corroboration-anchor count, or unlimited. #[serde(rename = "maxCodeGraphCorroborationAnchorsPerRepo")] #[schemars(with = "i64")] @@ -311,6 +319,7 @@ impl Default for ExecutionPolicy { max_watch_session_wall_time_ms: DEFAULT_MAX_WATCH_SESSION_WALL_TIME_MS, min_watch_rescan_interval_ms: DEFAULT_MIN_WATCH_RESCAN_INTERVAL_MS, max_checkpoint_cache_bytes: DEFAULT_MAX_CHECKPOINT_CACHE_BYTES, + max_extraction_workers: DEFAULT_MAX_EXTRACTION_WORKERS, max_codegraph_corroboration_anchors_per_repo: CodeGraphCorroborationAnchorLimit::try_from( DEFAULT_MAX_CODEGRAPH_CORROBORATION_ANCHORS_PER_REPO, @@ -351,6 +360,7 @@ impl ExecutionPolicy { apply!(max_watch_session_wall_time_ms); apply!(min_watch_rescan_interval_ms); apply!(max_checkpoint_cache_bytes); + apply!(max_extraction_workers); if let Some(value) = values.max_codegraph_corroboration_anchors_per_repo { policy.max_codegraph_corroboration_anchors_per_repo = value.try_into()?; } @@ -415,9 +425,26 @@ impl ExecutionPolicy { return Err(InvalidExecutionPolicy::InvalidValue { field, value }); } } + if self.max_extraction_workers == 0 + || self.max_extraction_workers > MAX_EXTRACTION_WORKERS_CEILING + { + return Err(InvalidExecutionPolicy::InvalidValue { + field: "maxExtractionWorkers", + value: self.max_extraction_workers, + }); + } Ok(()) } + /// Returns the extraction worker count bounded by the host's available parallelism. + #[must_use] + pub fn effective_extraction_workers(&self) -> usize { + let available = std::thread::available_parallelism().map_or(1, NonZeroUsize::get); + usize::try_from(self.max_extraction_workers) + .unwrap_or(usize::MAX) + .clamp(1, available.max(1)) + } + fn validate_scan_relationships(&self) -> Result<(), InvalidExecutionPolicy> { Self::require_not_greater( "maxNoProgressTimeMs", @@ -570,12 +597,6 @@ impl ExecutionPolicy { stable_id("execution-policy", &canonical) } - #[must_use] - /// Returns the historical scan-only fingerprint alias. - pub fn fingerprint(&self) -> String { - self.scan_fingerprint() - } - /// Returns a fingerprint that also includes the additive `CodeGraph` corroboration bound. #[must_use] pub fn fingerprint_with_codegraph_limit( @@ -717,6 +738,28 @@ pub struct ExecutionSummary { pub artifact_duration_p95_ms: u64, /// 99th-percentile artifact-extractor duration in milliseconds. pub artifact_duration_p99_ms: u64, + /// Files whose content hash was reused from filesystem metadata without a read. + pub stat_cache_hits: u64, + /// Bytes of artifact content read from disk for fingerprinting and extraction. + pub content_bytes_read: u64, + /// Rows inserted, updated, or deleted by snapshot publication. + pub published_rows: u64, + /// Effective concurrent extraction workers. + pub extraction_workers: u64, + /// Wall time and resident memory at the end of each completed worker phase. + pub phases: Vec, +} + +/// Duration and memory observation for one completed worker phase. +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] +#[serde(rename_all = "camelCase")] +pub struct PhaseTelemetry { + /// Completed phase. + pub phase: JobPhase, + /// Wall time spent in the phase. + pub duration_ms: u64, + /// Resident memory of the worker process when the phase completed. + pub resident_memory_bytes: u64, } /// Typed failure produced when one supervised execution resource is exhausted. @@ -902,7 +945,7 @@ mod tests { let second = ExecutionPolicy::resolve(Some(&ExecutionPolicyOverrides::default())) .expect("defaults valid"); - assert_eq!(first.fingerprint(), second.fingerprint()); + assert_eq!(first.scan_fingerprint(), second.scan_fingerprint()); } #[test] diff --git a/crates/code-system-graph-core/src/execution_policy/tracker.rs b/crates/code-system-graph-core/src/execution_policy/tracker.rs index b736bd9..0e3fef5 100644 --- a/crates/code-system-graph-core/src/execution_policy/tracker.rs +++ b/crates/code-system-graph-core/src/execution_policy/tracker.rs @@ -66,6 +66,12 @@ impl ScanJobTracker { } } + /// Phase of the most recent progress or phase change. + #[must_use] + pub const fn phase(&self) -> JobPhase { + self.phase + } + /// Changes phase after checking the active deadlines. /// /// # Errors diff --git a/crates/code-system-graph-core/src/extraction_budget.rs b/crates/code-system-graph-core/src/extraction_budget.rs index 60cb550..04223de 100644 --- a/crates/code-system-graph-core/src/extraction_budget.rs +++ b/crates/code-system-graph-core/src/extraction_budget.rs @@ -8,7 +8,7 @@ use serde::{Deserialize, Serialize}; use thiserror::Error; /// Extraction payload/cache contract for the 1.1 source and relationship semantics. -pub const EXTRACTION_CONTRACT_VERSION: &str = "1.1.0"; +pub const EXTRACTION_CONTRACT_VERSION: &str = "1.2.0"; /// Default maximum input bytes accepted for one artifact-extractor invocation. pub const DEFAULT_MAX_INPUT_BYTES_PER_ARTIFACT: u64 = 8_388_608; diff --git a/crates/code-system-graph-core/src/http.rs b/crates/code-system-graph-core/src/http.rs index 218139d..2d7b286 100644 --- a/crates/code-system-graph-core/src/http.rs +++ b/crates/code-system-graph-core/src/http.rs @@ -4,7 +4,9 @@ use code_system_graph_model::{ use serde::{Deserialize, Serialize}; use thiserror::Error; -use crate::{ExtractionBudgets, ExtractionLimitExceeded, ExtractionTracker, HttpConsumerConfig}; +use crate::{ + CallScope, ExtractionBudgets, ExtractionLimitExceeded, ExtractionTracker, HttpConsumerConfig, canonical_route +}; const HTTP_METHODS: [&str; 8] = [ "DELETE", "GET", "HEAD", "OPTIONS", "PATCH", "POST", "PUT", "TRACE", @@ -31,6 +33,8 @@ pub struct HttpBoundary { pub path: String, /// Consumer or provider direction. pub role: BoundaryRole, + /// Repositories allowed to provide a consumer's operation. + pub scope: CallScope, /// Direct evidence for the boundary. pub evidence: Evidence, } @@ -498,7 +502,11 @@ fn boundary( Provenance::Declared, ), }; - let stable_key = format!("http:{}:{role_key}:{method}:{path}", repo_id.as_str()); + let stable_key = format!( + "http:{}:{role_key}:{method}:{}", + repo_id.as_str(), + canonical_route(path) + ); let evidence_key = format!("{stable_key}:{source_path}"); HttpBoundary { node: Node { @@ -511,6 +519,7 @@ fn boundary( method: method.to_owned(), path: path.to_owned(), role, + scope: CallScope::Workspace, evidence: Evidence { id: EvidenceId::new(stable_id("evidence", &evidence_key)), repo_id: Some(repo_id), diff --git a/crates/code-system-graph-core/src/incremental.rs b/crates/code-system-graph-core/src/incremental.rs index 94bb72d..1a01bfa 100644 --- a/crates/code-system-graph-core/src/incremental.rs +++ b/crates/code-system-graph-core/src/incremental.rs @@ -1,12 +1,13 @@ -use std::collections::{BTreeMap, BTreeSet}; - use code_system_graph_model::{ ArtifactChange, ArtifactChangeKind, ArtifactFingerprint, CheckoutId, NativePath, RepoId }; -type ArtifactKey = (RepoId, CheckoutId, NativePath, String); +type ArtifactKey<'a> = (&'a RepoId, &'a CheckoutId, &'a NativePath, &'a str); /// Deterministic incremental plan for extractor-relevant artifacts. +/// +/// Only artifacts that require extraction or deletion are listed; every current artifact absent +/// from [`Self::changes`] is unchanged. #[derive(Debug, Clone, PartialEq, Eq)] pub struct IncrementalPlan { /// Changes sorted by repository, checkout, path, and extractor. @@ -17,18 +18,13 @@ impl IncrementalPlan { /// Returns whether any artifact requires add, modify, or delete processing. #[must_use] pub fn has_changes(&self) -> bool { - self.changes - .iter() - .any(|change| change.kind != ArtifactChangeKind::Unchanged) + !self.changes.is_empty() } /// Returns the number of artifacts requiring extractor or deletion work. #[must_use] pub fn changed_count(&self) -> usize { - self.changes - .iter() - .filter(|change| change.kind != ArtifactChangeKind::Unchanged) - .count() + self.changes.len() } } @@ -40,47 +36,59 @@ pub fn plan_incremental_scan( ) -> IncrementalPlan { let previous = fingerprint_map(previous); let current = fingerprint_map(current); - let keys = previous + let added_or_modified = current.iter().filter_map(|(key, after)| { + let kind = match previous.get(key) { + None => ArtifactChangeKind::Added, + Some(before) if before.content_hash != after.content_hash => { + ArtifactChangeKind::Modified + } + Some(_) => return None, + }; + Some((*key, kind)) + }); + let deleted = previous .keys() - .chain(current.keys()) - .cloned() - .collect::>(); - let changes = keys - .into_iter() - .filter_map(|key| { - let kind = match (previous.get(&key), current.get(&key)) { - (None, Some(_)) => ArtifactChangeKind::Added, - (Some(_), None) => ArtifactChangeKind::Deleted, - (Some(before), Some(after)) if before.content_hash != after.content_hash => { - ArtifactChangeKind::Modified - } - (Some(_), Some(_)) => ArtifactChangeKind::Unchanged, - (None, None) => return None, - }; - Some(ArtifactChange { - repo_id: key.0, - checkout_id: key.1, - path: key.2, - extractor: key.3, - kind, - }) + .filter(|key| !current.contains_key(*key)) + .map(|key| (*key, ArtifactChangeKind::Deleted)); + let mut changes = added_or_modified + .chain(deleted) + .map(|(key, kind)| ArtifactChange { + repo_id: key.0.clone(), + checkout_id: key.1.clone(), + path: key.2.clone(), + extractor: key.3.to_owned(), + kind, }) - .collect(); + .collect::>(); + changes.sort_by(|left, right| { + ( + &left.repo_id, + &left.checkout_id, + &left.path, + &left.extractor, + ) + .cmp(&( + &right.repo_id, + &right.checkout_id, + &right.path, + &right.extractor, + )) + }); IncrementalPlan { changes } } fn fingerprint_map( fingerprints: &[ArtifactFingerprint], -) -> BTreeMap { +) -> foldhash::HashMap, &ArtifactFingerprint> { fingerprints .iter() .map(|fingerprint| { ( ( - fingerprint.repo_id.clone(), - fingerprint.checkout_id.clone(), - fingerprint.path.clone(), - fingerprint.extractor.clone(), + &fingerprint.repo_id, + &fingerprint.checkout_id, + &fingerprint.path, + fingerprint.extractor.as_str(), ), fingerprint, ) @@ -112,7 +120,7 @@ mod tests { } #[test] - fn plan_should_classify_add_modify_delete_and_unchanged() { + fn plan_should_list_only_added_deleted_and_modified_artifacts_in_key_order() { let previous = vec![ fingerprint("deleted.yaml", "a"), fingerprint("modified.yaml", "a"), @@ -125,20 +133,29 @@ mod tests { ]; let plan = plan_incremental_scan(&previous, ¤t); - let kinds = plan + let changes = plan .changes .iter() - .map(|change| change.kind) + .map(|change| (change.path.display.as_str(), change.kind)) .collect::>(); assert_eq!( - kinds, + changes, vec![ - ArtifactChangeKind::Added, - ArtifactChangeKind::Deleted, - ArtifactChangeKind::Modified, - ArtifactChangeKind::Unchanged, + ("added.yaml", ArtifactChangeKind::Added), + ("deleted.yaml", ArtifactChangeKind::Deleted), + ("modified.yaml", ArtifactChangeKind::Modified), ] ); + assert_eq!(plan.changed_count(), 3); + } + + #[test] + fn plan_should_report_no_changes_for_identical_artifact_sets() { + let fingerprints = vec![fingerprint("same.yaml", "a")]; + + let plan = plan_incremental_scan(&fingerprints, &fingerprints); + + assert!(!plan.has_changes()); } } diff --git a/crates/code-system-graph-core/src/interfaces.rs b/crates/code-system-graph-core/src/interfaces.rs index b877930..ab59671 100644 --- a/crates/code-system-graph-core/src/interfaces.rs +++ b/crates/code-system-graph-core/src/interfaces.rs @@ -177,10 +177,10 @@ pub enum DoctorStatus { pub struct SchemaDoctorInput { /// Stable store or component name. pub name: String, - /// Schema version expected by this binary. - pub expected_version: u32, - /// Observed schema version, or `None` when unavailable. - pub actual_version: Option, + /// Schema identity expected by this binary. + pub expected_schema: String, + /// Observed schema identity, or `None` when unavailable. + pub actual_schema: Option, /// Whether the exact schema metadata is internally consistent. pub metadata_consistent: Option, } @@ -991,32 +991,32 @@ fn append_schema_checks(inputs: &[SchemaDoctorInput], checks: &mut Vec ( - DoctorStatus::Healthy, - format!("Schema version {actual} and metadata are consistent."), - None, - ), - (Some(actual), _) if actual != input.expected_version => ( - DoctorStatus::Failed, - format!( - "Schema version {actual} does not match expected version {}.", - input.expected_version + let (status, summary, remediation) = + match (input.actual_schema.as_deref(), input.metadata_consistent) { + (Some(actual), Some(true)) if actual == input.expected_schema => ( + DoctorStatus::Healthy, + "Schema identity matches this binary and metadata is consistent.".to_owned(), + None, ), - Some("Remove the incompatible local database and run a full scan.".to_owned()), - ), - (Some(_), Some(false)) => ( - DoctorStatus::Failed, - "Schema metadata is inconsistent.".to_owned(), - Some("Restore an exact 1.0.0 backup or rebuild the local database.".to_owned()), - ), - _ => ( - DoctorStatus::Unknown, - "Schema state was not fully observed.".to_owned(), - Some("Open the store and validate its exact schema metadata.".to_owned()), - ), - }; + (Some(actual), _) if actual != input.expected_schema => ( + DoctorStatus::Failed, + format!( + "Schema identity `{actual}` does not match `{}`.", + input.expected_schema + ), + Some("Remove the local database and run a full scan.".to_owned()), + ), + (Some(_), Some(false)) => ( + DoctorStatus::Failed, + "Schema metadata is inconsistent.".to_owned(), + Some("Remove the local database and run a full scan.".to_owned()), + ), + _ => ( + DoctorStatus::Unknown, + "Schema state was not fully observed.".to_owned(), + Some("Open the store and validate its exact schema metadata.".to_owned()), + ), + }; checks.push(DoctorCheck { category: DoctorCategory::Schema, name: input.name.clone(), @@ -1285,8 +1285,8 @@ mod tests { DoctorRequest { schema: vec![SchemaDoctorInput { name: "store".to_owned(), - expected_version: 8, - actual_version: Some(8), + expected_schema: "schema:test".to_owned(), + actual_schema: Some("schema:test".to_owned()), metadata_consistent: Some(true), }], integrity: vec![IntegrityDoctorInput { diff --git a/crates/code-system-graph-core/src/lib.rs b/crates/code-system-graph-core/src/lib.rs index d6db45e..66102e9 100644 --- a/crates/code-system-graph-core/src/lib.rs +++ b/crates/code-system-graph-core/src/lib.rs @@ -5,6 +5,7 @@ mod builtin_extractors; mod capability_dir; mod change_analysis; mod changes; +mod client_flows; mod codegraph; mod communities; mod config; @@ -30,26 +31,33 @@ mod interfaces; mod linker; mod manifest; mod manifest_edit; +mod markers; mod package_graph; mod packages; +pub mod parallel; mod protobuf_contracts; mod protobuf_graph; mod provider; mod pull_requests; mod query; mod registry; +mod repository_symbols; +mod router_mounts; +mod routes; mod secret_safety; mod source_graph; mod source_http; mod source_polyglot; +mod source_routers; mod source_symbol; mod source_syntax; mod test_links; mod trace; +mod url_template; mod yaml; pub use batch::{ - ArtifactKey, BatchAction, BatchPlanError, ExtractorBatch, ExtractorBatchPlan, PlannedBatch, affected_link_keys, load_extractor_batch, load_extractor_batch_with_budgets, plan_extractor_batches, store_extractor_batch + BatchPlanError, ExtractorBatch, load_extractor_batch, load_extractor_batch_with_budgets, store_extractor_batch }; pub use builtin_extractors::{ FocusedSourceExtractor, FocusedSourceLanguage, GeneratedClientMetadataExtractor, PackageManifestExtractor, charge_source_observation, precheck_focused_source_values @@ -63,6 +71,7 @@ pub use change_analysis::{ pub use changes::{ AnalyzerVersions, ChangeError, ChangeHunk, ChangeProvider, ChangeRequest, ChangeScope, ChangeSet, ChangeSourceLayer, ChangeValidity, ChangeValidityInput, ChangedFile, ChangedFileStatus, ChangedLine, ChangedLineKind, CommitFileSelection, CommitGate, CommitIntent, CommitSelection, GitCliChangeProvider, StaleReason, evaluate_commit_gate, validate_change_set }; +pub use client_flows::compose_client_flows; pub use codegraph::{CodeGraphConfig, CodeGraphProvider}; pub use communities::{ CommunityError, analyze_communities, analyze_communities_with_progress, compare_community_snapshots @@ -89,7 +98,7 @@ pub use events::{ DeliverySemantics, EventBroker, EventDocument, EventEvidenceLine, EventExtractionError, EventObservation, EventRole, EventSchemaDefinition, EventSchemaField, extract_asyncapi, parse_event_source }; pub use execution_policy::{ - CodeGraphCorroborationAnchorLimit, DEFAULT_MAX_CODEGRAPH_CORROBORATION_ANCHORS_PER_REPO, ExecutionLimitExceeded, ExecutionPolicy, ExecutionPolicyOverrides, ExecutionResource, ExecutionSummary, InvalidExecutionPolicy, JobPhase, MIN_MCP_MARKDOWN_BYTES, MonotonicClock, ScanJobTracker + CodeGraphCorroborationAnchorLimit, DEFAULT_MAX_CODEGRAPH_CORROBORATION_ANCHORS_PER_REPO, ExecutionLimitExceeded, ExecutionPolicy, ExecutionPolicyOverrides, ExecutionResource, ExecutionSummary, InvalidExecutionPolicy, JobPhase, MIN_MCP_MARKDOWN_BYTES, MonotonicClock, PhaseTelemetry, ScanJobTracker }; pub use extraction_budget::{ BoundedJsonWriter, EXTRACTION_CONTRACT_VERSION, ExtractionBudgetOverrides, ExtractionBudgets, ExtractionClock, ExtractionLimitExceeded, ExtractionResource, ExtractionTracker, InvalidExtractionBudget @@ -122,7 +131,7 @@ pub use interfaces::{ Ambiguity, CONTRACT_NODE_KINDS, ConfigDoctorInput, ContractAction, ContractCompatibility, ContractCompatibilitySummary, ContractDifference, ContractFinding, ContractIssue, ContractIssueSeverity, ContractLink, ContractReport, ContractRequest, ContractView, DELIVERY_METADATA_VERSION, DoctorCategory, DoctorCheck, DoctorReport, DoctorRequest, DoctorStatus, DomainErrorKind, EvidenceMetadata, ExitCode, ExportFormat, ExportReport, ExportRequest, FreshnessDoctorInput, INTERFACE_RESULT_VERSION, INTERFACE_SCHEMA_VERSION, IntegrityDoctorInput, InterfaceError, MAX_EXPORT_EDGES, MAX_EXPORT_NODES, NextAction, Page, Pagination, ProviderDoctorInput, ProviderDoctorStatus, PublicSchema, PublicSchemaCatalog, SchemaDoctorInput, Summary, Warning, classify_exit_code, classify_interface_error, doctor, export_graph, inspect_contracts, paginate, public_schema_catalog }; pub use linker::{ - HttpLinkAmbiguity, HttpLinkResolution, LinkError, ManualLinkEndpoint, ManualLinkError, ManualLinkResolution, link_http_boundaries, link_http_boundaries_with_ambiguities, merge_affected_link_neighborhoods, resolve_manual_links + HttpLinkAmbiguity, HttpLinkResolution, ManualLinkEndpoint, ManualLinkError, ManualLinkResolution, resolve_manual_links }; pub use manifest::{ ContractImplementationConfig, HttpConsumerConfig, HttpContractConfig, IntegrationTestConfig, ManifestError, ManifestExtensions, ManualLinkConfig, RepositoryConfig, WorkspaceManifest, parse_manifest, parse_manifest_with_extensions, validate_manual_links @@ -150,12 +159,17 @@ pub use query::{ AgentNextAction, EdgeKindCost, PathSegment, PathSegmentScope, QueryError, SearchCoverage, SearchExplanation, SearchFilters, SearchHit, SearchReport, SearchRequest, TraversalAlgorithm, TraversalDirection, TraversalFilters, TraversalLimits, TraversalOptions, TraversalPath, TraversalReport, TraversalRequest, search, traverse }; pub use registry::{RegisteredWorkspace, RegistryError, encode_native_path, register_workspace}; +pub use repository_symbols::RepositorySourceFile; +pub use router_mounts::compose_router_mounts; +pub use routes::{ + AuthorityMap, CallScope, RouteSegment, RouteShape, canonical_route, link_http_routes, normalize_authority +}; pub use secret_safety::{ ConfigArtifactKind, ConfigExtractionError, SafeConfigDocument, SafeConfigKey, SensitiveKeyKind, classify_sensitive_key, extract_safe_config, is_safe_literal_reference }; pub use source_graph::{SourceGraphFacts, source_observations_to_graph}; pub use source_http::{ - SourceEpistemicStatus, SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SourceRole, SourceWarning, normalize_source_http_path, parse_python_source, parse_python_source_with_tracker, parse_rust_source, parse_rust_source_with_tracker + CallArgument, CallSite, SourceEpistemicStatus, SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SourceRole, SourceWarning, SymbolRef, UrlPart, UrlTemplate, normalize_source_http_path, parse_python_source, parse_python_source_with_tracker, parse_rust_source, parse_rust_source_with_tracker }; pub use source_polyglot::{ parse_go_source, parse_go_source_with_tracker, parse_java_source, parse_java_source_with_tracker, parse_javascript_source, parse_javascript_source_at_path, parse_javascript_source_at_path_with_tracker, parse_typescript_source, parse_typescript_source_at_path, parse_typescript_source_at_path_with_tracker @@ -165,6 +179,6 @@ pub use source_syntax::{ SourceSyntaxError, SourceSyntaxInspection, SourceSyntaxLanguage, inspect_source_syntax }; pub use test_links::{ - DeclaredImplementation, DeclaredTestCase, declared_implementation, declared_test_case, link_declared_implementations, link_declared_implementations_with_ambiguities, link_declared_tests, link_declared_tests_with_ambiguities + DeclaredImplementation, DeclaredTestCase, declared_implementation, declared_test_case }; pub use trace::{FederatedGraph, TraceError}; diff --git a/crates/code-system-graph-core/src/linker.rs b/crates/code-system-graph-core/src/linker.rs index 856df73..f1e1546 100644 --- a/crates/code-system-graph-core/src/linker.rs +++ b/crates/code-system-graph-core/src/linker.rs @@ -1,31 +1,16 @@ use std::collections::{BTreeMap, BTreeSet}; use code_system_graph_model::{ - Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, EvidenceRef, LinkDecision, LinkStatus, Node, NodeId, Provenance, stable_id + Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, EvidenceRef, HttpLinkReport, LinkDecision, LinkStatus, Node, NodeId, Provenance, stable_id }; use semver::Version; use thiserror::Error; -use crate::{BoundaryRole, HttpBoundary, ManifestError, ManualLinkConfig, validate_manual_links}; +use crate::{ManifestError, ManualLinkConfig, validate_manual_links}; const MANUAL_LINK_MATCHER: &str = "manual_exact"; const MANUAL_LINK_EXTRACTOR: &str = "code-system-graph.manual-link"; -/// Error returned by deterministic boundary linking. -#[derive(Debug, Error, PartialEq, Eq)] -pub enum LinkError { - /// More than one provider has the same exact contract identity. - #[error("ambiguous HTTP provider for {method} {path}: {candidates:?}")] - AmbiguousProvider { - /// Canonical HTTP method. - method: String, - /// Canonical path template. - path: String, - /// Stable candidate node identifiers in deterministic order. - candidates: Vec, - }, -} - /// One exact HTTP contract that could not be linked because several providers matched. #[derive(Debug, Clone, PartialEq, Eq)] pub struct HttpLinkAmbiguity { @@ -44,20 +29,8 @@ pub struct HttpLinkResolution { pub edges: Vec, /// Source-free ambiguous contract identities omitted from the edge set. pub ambiguities: Vec, -} - -impl HttpLinkResolution { - pub(crate) fn into_legacy_result(self) -> Result, LinkError> { - let Self { edges, ambiguities } = self; - let Some(ambiguity) = ambiguities.into_iter().next() else { - return Ok(edges); - }; - Err(LinkError::AmbiguousProvider { - method: ambiguity.method, - path: ambiguity.path, - candidates: ambiguity.candidates, - }) - } + /// Consumer and test calls left without a provider edge, with outcome counts. + pub report: HttpLinkReport, } /// Endpoint field being resolved for a manual relationship. @@ -161,91 +134,6 @@ pub struct ManualLinkResolution { pub decisions: Vec, } -/// Links HTTP consumers to providers by exact canonical method and path. -/// -/// Consumers without an observed provider remain unlinked. Callers must represent that as a -/// coverage gap rather than concluding that no dependency exists. -/// -/// Duplicate providers remain fail-closed for compatibility. Use -/// [`link_http_boundaries_with_ambiguities`] to preserve ambiguities as data while continuing with -/// unrelated contracts. -/// -/// # Errors -/// -/// Returns [`LinkError::AmbiguousProvider`] instead of silently omitting an ambiguous relationship. -pub fn link_http_boundaries(boundaries: &[HttpBoundary]) -> Result, LinkError> { - link_http_boundaries_with_ambiguities(boundaries).into_legacy_result() -} - -/// Links exact HTTP boundaries while preserving duplicate-provider decisions. -#[must_use] -pub fn link_http_boundaries_with_ambiguities(boundaries: &[HttpBoundary]) -> HttpLinkResolution { - let mut providers: BTreeMap<(&str, &str), Vec<&HttpBoundary>> = BTreeMap::new(); - for boundary in boundaries { - if boundary.role == BoundaryRole::Provider { - let candidates = providers - .entry((&boundary.method, &boundary.path)) - .or_default(); - if let Some(existing) = candidates - .iter() - .position(|candidate| candidate.node.id == boundary.node.id) - { - if boundary.evidence.confidence > candidates[existing].evidence.confidence { - candidates[existing] = boundary; - } - } else { - candidates.push(boundary); - } - } - } - - let mut edges = Vec::new(); - let mut ambiguities = Vec::new(); - for consumer in boundaries - .iter() - .filter(|boundary| boundary.role == BoundaryRole::Consumer) - { - let Some(candidates) = providers.get(&(consumer.method.as_str(), consumer.path.as_str())) - else { - continue; - }; - if candidates.len() > 1 { - ambiguities.push(http_link_ambiguity( - &consumer.method, - &consumer.path, - candidates.iter().map(|candidate| &candidate.node.id), - )); - continue; - } - let provider = candidates[0]; - let edge_key = format!( - "{}:calls_remote:{}", - consumer.node.id.as_str(), - provider.node.id.as_str() - ); - edges.push(Edge { - id: EdgeId::new(stable_id("edge", &edge_key)), - source: consumer.node.id.clone(), - target: provider.node.id.clone(), - kind: EdgeKind::CallsRemote, - confidence: consumer - .evidence - .confidence - .min(provider.evidence.confidence), - status: consensus_status( - consumer - .evidence - .confidence - .min(provider.evidence.confidence), - ), - evidence: vec![consumer.evidence.id.clone(), provider.evidence.id.clone()], - }); - } - edges.sort_by(|left, right| left.id.cmp(&right.id)); - sort_http_ambiguities(&mut ambiguities); - HttpLinkResolution { edges, ambiguities } -} - pub(crate) fn http_link_ambiguity<'a>( method: &str, path: &str, @@ -275,14 +163,6 @@ pub(crate) fn sort_http_ambiguities(ambiguities: &mut Vec) { ambiguities.dedup(); } -fn consensus_status(confidence: f32) -> EpistemicStatus { - if confidence >= 1.0 { - EpistemicStatus::Confirmed - } else { - EpistemicStatus::Inferred - } -} - /// Resolves and applies exact manual relationships to an automatic edge set. /// /// Endpoint values match only [`NodeId`] text or [`Node::stable_key`] text. A creation replaces an @@ -517,64 +397,17 @@ fn manual_link_decision( } } -/// Reuses unaffected edges and atomically replaces every affected link neighborhood. -/// -/// Both the previous and current node-to-key maps are required because replacements may change -/// their canonical key. Edges referencing removed nodes are never reused. -#[must_use] -pub fn merge_affected_link_neighborhoods( - previous_edges: &[Edge], - recomputed_edges: &[Edge], - affected_keys: &BTreeSet, - previous_node_keys: &BTreeMap, - current_node_keys: &BTreeMap, - current_node_ids: &BTreeSet, -) -> Vec -where - K: Ord, -{ - let mut merged = BTreeMap::new(); - for edge in previous_edges { - let endpoints_exist = - current_node_ids.contains(&edge.source) && current_node_ids.contains(&edge.target); - if endpoints_exist && !edge_touches_keys(edge, previous_node_keys, affected_keys) { - merged.insert(edge.id.clone(), edge.clone()); - } - } - for edge in recomputed_edges { - if edge_touches_keys(edge, current_node_keys, affected_keys) { - merged.insert(edge.id.clone(), edge.clone()); - } - } - merged.into_values().collect() -} - -fn edge_touches_keys( - edge: &Edge, - node_keys: &BTreeMap, - affected_keys: &BTreeSet, -) -> bool -where - K: Ord, -{ - [&edge.source, &edge.target] - .into_iter() - .filter_map(|node| node_keys.get(node)) - .any(|key| affected_keys.contains(key)) -} - #[cfg(test)] mod tests { - use std::collections::{BTreeMap, BTreeSet}; use code_system_graph_model::{ Edge, EdgeId, EdgeKind, EpistemicStatus, LinkStatus, Node, NodeId, NodeKind, Provenance, RepoId }; - use super::{ - LinkError, ManualLinkEndpoint, ManualLinkError, link_http_boundaries, link_http_boundaries_with_ambiguities, merge_affected_link_neighborhoods, resolve_manual_links + use super::{ManualLinkEndpoint, ManualLinkError, resolve_manual_links}; + use crate::{ + AuthorityMap, HttpConsumerConfig, ManualLinkConfig, extract_openapi, link_http_routes }; - use crate::{HttpConsumerConfig, ManualLinkConfig, extract_openapi}; fn consumer() -> crate::HttpBoundary { crate::HttpBoundary::consumer( @@ -637,19 +470,24 @@ paths: #[test] fn link_http_boundaries_should_require_bilateral_evidence() { - let result = link_http_boundaries(&[consumer(), provider("repo:api")]); - let evidence_count = result.map(|edges| edges[0].evidence.len()); + let result = link_http_routes( + &[consumer(), provider("repo:api")], + &[], + &[], + &AuthorityMap::new(), + ); - assert_eq!(evidence_count, Ok(2)); + assert_eq!(result.edges[0].evidence.len(), 2); } #[test] fn link_http_boundaries_should_preserve_duplicate_providers_as_ambiguity() { - let result = link_http_boundaries_with_ambiguities(&[ - consumer(), - provider("repo:api-a"), - provider("repo:api-b"), - ]); + let result = link_http_routes( + &[consumer(), provider("repo:api-a"), provider("repo:api-b")], + &[], + &[], + &AuthorityMap::new(), + ); assert_eq!(result.edges, Vec::new()); assert_eq!(result.ambiguities.len(), 1); @@ -658,14 +496,6 @@ paths: assert_eq!(result.ambiguities[0].candidates.len(), 2); } - #[test] - fn link_http_boundaries_compatibility_wrapper_should_reject_duplicate_providers() { - let result = - link_http_boundaries(&[consumer(), provider("repo:api-a"), provider("repo:api-b")]); - - assert!(matches!(result, Err(LinkError::AmbiguousProvider { .. }))); - } - #[test] fn resolve_manual_links_should_create_confirmed_exact_edge_with_manual_evidence() { let nodes = [ @@ -811,40 +641,4 @@ paths: second.unwrap_or_else(|error| panic!("second resolution should succeed: {error}")) ); } - - #[test] - fn relinking_should_replace_only_affected_neighborhoods() { - let orders = link_http_boundaries(&[consumer(), provider("repo:api")]) - .unwrap_or_else(|error| panic!("fixture must link: {error}")); - let mut health_consumer = consumer(); - health_consumer.method = "GET".to_owned(); - health_consumer.path = "/health".to_owned(); - let mut health_provider = provider("repo:health"); - health_provider.method = "GET".to_owned(); - health_provider.path = "/health".to_owned(); - let health = link_http_boundaries(&[health_consumer, health_provider]) - .unwrap_or_else(|error| panic!("fixture must link: {error}")); - let mut previous = orders.clone(); - previous.extend(health.clone()); - let affected = BTreeSet::from(["POST:/api/orders".to_owned()]); - let previous_keys = BTreeMap::from([ - (orders[0].source.clone(), "POST:/api/orders".to_owned()), - (orders[0].target.clone(), "POST:/api/orders".to_owned()), - (health[0].source.clone(), "GET:/health".to_owned()), - (health[0].target.clone(), "GET:/health".to_owned()), - ]); - let current_ids = previous_keys.keys().cloned().collect::>(); - let result = merge_affected_link_neighborhoods( - &previous, - &orders, - &affected, - &previous_keys, - &previous_keys, - ¤t_ids, - ); - - assert_eq!(result.len(), 2); - assert!(result.contains(&health[0])); - assert!(result.contains(&orders[0])); - } } diff --git a/crates/code-system-graph-core/src/manifest.rs b/crates/code-system-graph-core/src/manifest.rs index 0a2b90e..7a6126e 100644 --- a/crates/code-system-graph-core/src/manifest.rs +++ b/crates/code-system-graph-core/src/manifest.rs @@ -10,6 +10,7 @@ use crate::extraction_budget::{ ExtractionBudgetOverrides, ExtractionBudgets, InvalidExtractionBudget }; use crate::ignore_policy::{IgnorePatternError, validate_excludes, validate_include_defaults}; +use crate::routes::normalize_authority; const MANUAL_ENDPOINT_MAX_BYTES: usize = 2_048; const MANUAL_CONTRACT_MAX_BYTES: usize = 1_024; @@ -57,6 +58,9 @@ pub struct RepositoryConfig { pub excludes: Option>, /// Repository-relative exceptions to reactivable built-in exclusions. pub include_defaults: Option>, + /// `host` or `host:port` authorities served by this repository. An entry without a port + /// matches every port of that host. + pub authorities: Option>, } /// Additive manifest settings introduced without changing exhaustively constructible public @@ -101,6 +105,7 @@ struct RepositoryConfigWire { excludes: Option>, include_defaults: Option>, use_gitignore: Option, + authorities: Option>, } /// Exact manual relationship or automatic-link suppression. @@ -254,6 +259,24 @@ pub enum ManifestError { /// Repeated relationship. relation: EdgeKind, }, + /// A repository authority is not a `host` or `host:port` value. + #[error( + "manifest field `{field}` must be a `host` or `host:port` authority without scheme or path" + )] + InvalidAuthority { + /// Dot-style field location. + field: String, + }, + /// Two repositories declare the same authority. + #[error("authority `{authority}` is declared by both `repos.{first}` and `repos.{second}`")] + DuplicateAuthority { + /// Normalized authority. + authority: String, + /// Alias of the first declaring repository. + first: String, + /// Alias of the repeating repository. + second: String, + }, } /// Parses and semantically validates a strict workspace manifest. @@ -293,6 +316,7 @@ pub fn parse_manifest_with_extensions( implementations: repository.implementations, excludes: repository.excludes, include_defaults: repository.include_defaults, + authorities: repository.authorities, }, ) }) @@ -386,9 +410,32 @@ fn validate_manifest(manifest: WorkspaceManifest) -> Result) -> Result<(), ManifestError> { + let mut owners = BTreeMap::::new(); + for (alias, repository) in repos { + for (index, authority) in repository.authorities.iter().flatten().enumerate() { + let field = format!("repos.{alias}.authorities[{index}]"); + validate_not_empty(&field, authority)?; + let normalized = + normalize_authority(authority).ok_or(ManifestError::InvalidAuthority { field })?; + if let Some(first) = owners.insert(normalized.clone(), alias) + && first != alias + { + return Err(ManifestError::DuplicateAuthority { + authority: normalized, + first: first.to_owned(), + second: alias.clone(), + }); + } + } + } + Ok(()) +} + fn validate_repository_ignore_patterns( alias: &str, repository: &RepositoryConfig, @@ -878,4 +925,59 @@ repos: Err(ManifestError::ManualSelfLink { index: 0, .. }) )); } + + const AUTHORITIES: &str = r" +version: 1 +name: commerce +repos: + orders: + path: ../orders + authorities: [orders-api:8080, orders.internal] + billing: + path: ../billing + authorities: [BILLING] +"; + + #[test] + fn parse_manifest_should_accept_repository_authorities() { + let manifest = parse_manifest(AUTHORITIES).expect("valid authorities"); + + assert_eq!( + manifest.repos["orders"].authorities, + Some(vec![ + "orders-api:8080".to_owned(), + "orders.internal".to_owned() + ]) + ); + } + + #[test] + fn parse_manifest_should_reject_authorities_with_scheme_or_path() { + for authority in ["http://orders", "orders/v1", "orders:http"] { + let input = AUTHORITIES.replace("orders.internal", authority); + + let result = parse_manifest(&input); + + assert!( + matches!( + result, + Err(ManifestError::InvalidAuthority { ref field }) + if field == "repos.orders.authorities[1]" + ), + "{authority}: {result:?}" + ); + } + } + + #[test] + fn parse_manifest_should_reject_an_authority_declared_by_two_repositories() { + let input = AUTHORITIES.replace("[BILLING]", "[Orders-API:8080]"); + + let result = parse_manifest(&input); + + assert!(matches!( + result, + Err(ManifestError::DuplicateAuthority { authority, .. }) if authority == "orders-api:8080" + )); + } } diff --git a/crates/code-system-graph-core/src/markers.rs b/crates/code-system-graph-core/src/markers.rs new file mode 100644 index 0000000..3c7d026 --- /dev/null +++ b/crates/code-system-graph-core/src/markers.rs @@ -0,0 +1,99 @@ +//! Exact ASCII case-insensitive marker prefilters for lexical extractors. + +#[cfg(test)] +use std::path::{Path, PathBuf}; + +use aho_corasick::AhoCorasick; + +#[cfg(test)] +use crate::SourceLanguage; + +/// Multi-pattern substring matcher used to skip inputs that cannot produce observations. +pub(crate) struct MarkerSet { + matcher: Option, +} + +impl MarkerSet { + pub(crate) fn new(markers: &[&str]) -> Self { + Self { + matcher: AhoCorasick::builder() + .ascii_case_insensitive(true) + .build(markers) + .ok(), + } + } + + /// Returns whether any marker occurs in `haystack`. + /// + /// A matcher that failed to build never filters, so a construction error cannot drop facts. + pub(crate) fn any_in(&self, haystack: &str) -> bool { + self.matcher + .as_ref() + .is_none_or(|matcher| matcher.is_match(haystack)) + } +} + +/// Source files used by prefilter differential tests: repository fixtures plus this workspace's +/// own crates, whose test modules embed snippets for every supported language. +#[cfg(test)] +pub(crate) fn differential_corpus() -> Vec<(SourceLanguage, String, String)> { + let root = Path::new(env!("CARGO_MANIFEST_DIR")).join("../.."); + let mut files = Vec::new(); + visit(&root.join("fixtures"), &mut files); + visit(&root.join("crates"), &mut files); + files + .into_iter() + .filter_map(|path| { + let language = match path.extension().and_then(|extension| extension.to_str())? { + "py" => SourceLanguage::Python, + "js" | "mjs" | "cjs" | "jsx" => SourceLanguage::JavaScript, + "ts" | "tsx" | "mts" | "cts" => SourceLanguage::TypeScript, + "go" => SourceLanguage::Go, + "java" => SourceLanguage::Java, + "rs" => SourceLanguage::Rust, + _ => return None, + }; + let input = std::fs::read_to_string(&path).ok()?; + Some((language, path.display().to_string(), input)) + }) + .collect() +} + +#[cfg(test)] +fn visit(directory: &Path, files: &mut Vec) { + let Ok(entries) = std::fs::read_dir(directory) else { + return; + }; + let mut entries = entries + .flatten() + .map(|entry| entry.path()) + .collect::>(); + entries.sort(); + for path in entries { + let name = path + .file_name() + .and_then(|name| name.to_str()) + .unwrap_or(""); + if path.is_dir() { + if !matches!(name, "target" | "node_modules" | ".git") { + visit(&path, files); + } + } else { + files.push(path); + } + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn marker_set_should_match_ascii_case_insensitively() { + let markers = MarkerSet::new(&["publish", "pull("]); + + assert!(markers.any_in("client.PUBLISH(topic)")); + assert!(markers.any_in("sub.Pull(ctx)")); + assert!(!markers.any_in("fn handler() {}")); + } +} diff --git a/crates/code-system-graph-core/src/parallel.rs b/crates/code-system-graph-core/src/parallel.rs new file mode 100644 index 0000000..0708612 --- /dev/null +++ b/crates/code-system-graph-core/src/parallel.rs @@ -0,0 +1,66 @@ +//! Bounded scoped-thread fan-out whose results never depend on scheduling. + +use std::sync::atomic::{AtomicUsize, Ordering}; + +/// Applies `operation` to every item with at most `workers` threads and returns results in item +/// order, so callers observe the same sequence for every worker count. +pub fn map_ordered(items: &[T], workers: usize, operation: F) -> Vec +where + T: Sync, + R: Send, + F: Fn(&T) -> R + Sync, +{ + let workers = workers.clamp(1, items.len().max(1)); + if workers == 1 { + return items.iter().map(operation).collect(); + } + let next = AtomicUsize::new(0); + let mut indexed = std::thread::scope(|scope| { + let handles = (0..workers) + .map(|_| { + scope.spawn(|| { + let mut produced = Vec::new(); + loop { + let index = next.fetch_add(1, Ordering::Relaxed); + let Some(item) = items.get(index) else { + break; + }; + produced.push((index, operation(item))); + } + produced + }) + }) + .collect::>(); + handles + .into_iter() + .flat_map(|handle| match handle.join() { + Ok(produced) => produced, + Err(panic) => std::panic::resume_unwind(panic), + }) + .collect::>() + }); + indexed.sort_unstable_by_key(|(index, _)| *index); + indexed.into_iter().map(|(_, result)| result).collect() +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn map_ordered_should_return_item_order_for_every_worker_count() { + let items = (0..257_u64).collect::>(); + let expected = items.iter().map(|value| value * 3).collect::>(); + + for workers in [1, 2, 7, 64, 1_000] { + assert_eq!(map_ordered(&items, workers, |value| value * 3), expected); + } + } + + #[test] + fn map_ordered_should_accept_empty_input() { + let items: [u8; 0] = []; + + assert_eq!(map_ordered(&items, 8, |value| *value), Vec::::new()); + } +} diff --git a/crates/code-system-graph-core/src/query.rs b/crates/code-system-graph-core/src/query.rs index 830022f..e40bff0 100644 --- a/crates/code-system-graph-core/src/query.rs +++ b/crates/code-system-graph-core/src/query.rs @@ -5,7 +5,7 @@ use std::collections::{BTreeMap, BTreeSet, VecDeque}; use std::time::{Duration, Instant}; use code_system_graph_model::{ - CommunityId, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, Node, NodeId, NodeKind, RepoFreshnessState, RepoId, TraceSegment + CommunityId, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, HttpLinkGap, Node, NodeId, NodeKind, RepoFreshnessState, RepoId, TraceSegment }; use schemars::JsonSchema; use serde::{Deserialize, Serialize}; @@ -156,6 +156,10 @@ pub struct SearchReport { /// Navigation based only on observed repositories and entity identifiers. #[serde(default)] pub next_actions: Vec, + /// HTTP calls among the hits that have no provider edge: calls without a provider, + /// ambiguous calls with their candidates, and calls to external hosts. + #[serde(default)] + pub link_gaps: Vec, } /// Traversal strategy used to find confirmed paths. @@ -452,6 +456,7 @@ pub fn search(nodes: &[Node], request: &SearchRequest) -> Result)>; + +/// Process-wide remote cache shared by every registration in this process. +static REMOTE_CACHE: LazyLock> = LazyLock::new(|| Mutex::new(HashMap::new())); /// Validated workspace record plus native paths used by scanners. #[derive(Debug, Clone)] @@ -82,75 +101,22 @@ pub fn register_workspace( allowed_roots.sort(); allowed_roots.dedup(); - let mut records = Vec::with_capacity(manifest.repos.len()); + let configured = manifest + .repos + .iter() + .map(|(alias, repository)| (alias.as_str(), repository.path.as_str())) + .collect::>(); + let workers = std::thread::available_parallelism() + .map_or(1, std::num::NonZeroUsize::get) + .min(MAX_REGISTRATION_WORKERS); + let mut records = Vec::with_capacity(configured.len()); let mut checkout_paths = BTreeMap::new(); - for (alias, repository) in &manifest.repos { - let configured_path = resolve_path(&config_directory, Path::new(&repository.path)); - let configured_path = canonicalize(&configured_path)?; - let git_root = configured_path - .join(".git") - .exists() - .then(|| git_path(&configured_path, &["rev-parse", "--show-toplevel"])) - .flatten() - .and_then(|path| canonicalize(&path).ok()) - .filter(|path| path == &configured_path); - let is_git_repository = git_root.is_some(); - let checkout_path = git_root.unwrap_or(configured_path); - ensure_allowed(alias, &checkout_path, &allowed_roots)?; - - let git_common_dir = is_git_repository - .then(|| git_common_directory(&checkout_path)) - .flatten(); - if let Some(common_dir) = &git_common_dir { - ensure_allowed(alias, common_dir, &allowed_roots)?; - } - let normalized_remote = is_git_repository - .then(|| git_text(&checkout_path, &["remote", "get-url", "origin"])) - .flatten() - .map(|remote| normalize_remote(&remote)); - let head_commit = is_git_repository - .then(|| git_text(&checkout_path, &["rev-parse", "HEAD"])) - .flatten(); - let working_tree_dirty = is_git_repository - && git_output( - &checkout_path, - &["status", "--porcelain", "--untracked-files=normal"], - ) - .is_some_and(|output| !output.is_empty()); - let native_checkout = encode_native_path(&checkout_path); - let native_common = git_common_dir.as_deref().map(encode_native_path); - let repository_key = normalized_remote.as_ref().map_or_else( - || { - native_common.as_ref().map_or_else( - || format!("path:{}", native_path_fingerprint(&native_checkout)), - |common| format!("git:{}", native_path_fingerprint(common)), - ) - }, - |remote| format!("remote:{remote}"), - ); - let repo_id = RepoId::new(stable_id("repo", &repository_key)); - let checkout_id = CheckoutId::new(stable_id( - "checkout", - &format!( - "{}:{}", - repo_id.as_str(), - native_path_fingerprint(&native_checkout) - ), - )); - let is_linked_worktree = checkout_path.join(".git").is_file(); - - records.push(RepositoryRecord { - id: repo_id, - checkout_id, - alias: alias.clone(), - canonical_path: native_checkout, - git_common_dir: native_common, - normalized_remote, - head_commit, - is_linked_worktree, - working_tree_dirty, - }); - checkout_paths.insert(alias.clone(), checkout_path); + for registered in map_ordered(&configured, workers, |(alias, path)| { + register_repository(alias, path, &config_directory, &allowed_roots) + }) { + let (record, checkout_path) = registered?; + checkout_paths.insert(record.alias.clone(), checkout_path); + records.push(record); } records.sort_by(|left, right| left.alias.cmp(&right.alias)); @@ -175,6 +141,109 @@ pub fn register_workspace( }) } +fn register_repository( + alias: &str, + configured_path: &str, + config_directory: &Path, + allowed_roots: &[PathBuf], +) -> Result<(RepositoryRecord, PathBuf), RegistryError> { + let configured_path = resolve_path(config_directory, Path::new(configured_path)); + let configured_path = canonicalize(&configured_path)?; + let git_root = configured_path + .join(".git") + .exists() + .then(|| git_path(&configured_path, &["rev-parse", "--show-toplevel"])) + .flatten() + .and_then(|path| canonicalize(&path).ok()) + .filter(|path| path == &configured_path); + let is_git_repository = git_root.is_some(); + let checkout_path = git_root.unwrap_or(configured_path); + ensure_allowed(alias, &checkout_path, allowed_roots)?; + + let git_common_dir = is_git_repository + .then(|| git_common_directory(&checkout_path)) + .flatten(); + if let Some(common_dir) = &git_common_dir { + ensure_allowed(alias, common_dir, allowed_roots)?; + } + let normalized_remote = is_git_repository + .then(|| origin_remote(&checkout_path, git_common_dir.as_deref())) + .flatten() + .map(|remote| normalize_remote(&remote)); + let head_commit = is_git_repository + .then(|| git_text(&checkout_path, &["rev-parse", "HEAD"])) + .flatten(); + let working_tree_dirty = is_git_repository + && git_output( + &checkout_path, + &["status", "--porcelain", "--untracked-files=normal"], + ) + .is_some_and(|output| !output.is_empty()); + let native_checkout = encode_native_path(&checkout_path); + let native_common = git_common_dir.as_deref().map(encode_native_path); + let repository_key = normalized_remote.as_ref().map_or_else( + || { + native_common.as_ref().map_or_else( + || format!("path:{}", native_path_fingerprint(&native_checkout)), + |common| format!("git:{}", native_path_fingerprint(common)), + ) + }, + |remote| format!("remote:{remote}"), + ); + let repo_id = RepoId::new(stable_id("repo", &repository_key)); + let checkout_id = CheckoutId::new(stable_id( + "checkout", + &format!( + "{}:{}", + repo_id.as_str(), + native_path_fingerprint(&native_checkout) + ), + )); + let is_linked_worktree = checkout_path.join(".git").is_file(); + let record = RepositoryRecord { + id: repo_id, + checkout_id, + alias: alias.to_owned(), + canonical_path: native_checkout, + git_common_dir: native_common, + normalized_remote, + head_commit, + is_linked_worktree, + working_tree_dirty, + }; + Ok((record, checkout_path)) +} + +/// Returns the raw `origin` URL, reusing the cached value while the Git config is unchanged. +fn origin_remote(checkout_path: &Path, common_dir: Option<&Path>) -> Option { + let config_path = common_dir.map_or_else( + || checkout_path.join(".git").join("config"), + |common| common.join("config"), + ); + let stamp = std::fs::metadata(&config_path) + .ok() + .and_then(|metadata| Some((metadata.modified().ok()?, metadata.len()))) + .filter(|(modified, _)| { + SystemTime::now() + .duration_since(*modified) + .is_ok_and(|age| age >= REMOTE_CACHE_RACY_WINDOW) + }); + if let Some(stamp) = stamp + && let Ok(cache) = REMOTE_CACHE.lock() + && let Some((cached_stamp, remote)) = cache.get(&config_path) + && *cached_stamp == stamp + { + return remote.clone(); + } + let remote = git_text(checkout_path, &["remote", "get-url", "origin"]); + if let Some(stamp) = stamp + && let Ok(mut cache) = REMOTE_CACHE.lock() + { + cache.insert(config_path, (stamp, remote.clone())); + } + remote +} + /// Encodes a native path without requiring UTF-8. #[must_use] pub fn encode_native_path(path: &Path) -> NativePath { diff --git a/crates/code-system-graph-core/src/repository_symbols.rs b/crates/code-system-graph-core/src/repository_symbols.rs new file mode 100644 index 0000000..44e5128 --- /dev/null +++ b/crates/code-system-graph-core/src/repository_symbols.rs @@ -0,0 +1,333 @@ +//! Per-repository resolution of source-level symbol references across files. +//! +//! References are resolved through file-local names, relative script imports, Python package +//! imports, and module-qualified function paths. A reference that does not resolve to exactly one +//! declaration stays unresolved. + +use std::collections::{BTreeMap, BTreeSet}; +use std::path::Path; + +use code_system_graph_model::RepoId; + +use crate::{SourceLanguage, SourceObservation, SymbolRef}; + +const SCRIPT_EXTENSIONS: [&str; 8] = ["ts", "tsx", "js", "jsx", "mjs", "cjs", "mts", "cts"]; + +/// Source observations of one repository-relative file, as input to repository-level +/// composition. +#[derive(Debug, Clone, Copy)] +pub struct RepositorySourceFile<'a> { + /// Repository owning the file. + pub repo_id: &'a RepoId, + /// Portable repository-relative path. + pub path: &'a str, + /// Extracted observations of the file. + pub observations: &'a [SourceObservation], +} + +/// Indices of `files` grouped by repository, in input order within each repository. +pub(crate) fn files_by_repository<'a>( + files: &[RepositorySourceFile<'a>], +) -> BTreeMap<&'a RepoId, Vec> { + let mut by_repository = BTreeMap::<&RepoId, Vec>::new(); + for (index, file) in files.iter().enumerate() { + by_repository.entry(file.repo_id).or_default().push(index); + } + by_repository +} + +/// Repository-relative source file with its extracted observations. +#[derive(Debug, Clone, Copy)] +pub(crate) struct SymbolFile<'a> { + pub(crate) path: &'a str, + pub(crate) observations: &'a [SourceObservation], +} + +/// Resolved identity of a referenced symbol. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash)] +pub(crate) enum SymbolKey { + /// A name bound in one file, identified by its index. + Local { file: usize, name: String }, + /// A function qualified by its module path. + Function(Vec), +} + +/// Files of one repository and the module-qualified functions declared in them. +pub(crate) struct RepositorySymbols<'a> { + files: Vec>, + paths: BTreeMap<&'a str, usize>, + functions: BTreeSet>, +} + +impl<'a> RepositorySymbols<'a> { + pub(crate) fn new(files: Vec>) -> Self { + let paths = files + .iter() + .enumerate() + .map(|(index, file)| (file.path, index)) + .collect(); + Self { + files, + paths, + functions: BTreeSet::new(), + } + } + + pub(crate) fn files(&self) -> &[SymbolFile<'a>] { + &self.files + } + + /// Declares a function of file `index`, returning its module-qualified key. + pub(crate) fn declare_function(&mut self, index: usize, name: &str) -> Vec { + let key = self.function_key(index, name); + self.functions.insert(key.clone()); + key + } + + pub(crate) fn resolve(&self, index: usize, reference: &SymbolRef) -> Option { + match reference { + SymbolRef::Local(name) => Some(SymbolKey::Local { + file: index, + name: name.clone(), + }), + SymbolRef::Function(name) => Some(SymbolKey::Function(self.function_key(index, name))), + SymbolRef::Call(path) => self.resolve_call(index, path).map(SymbolKey::Function), + SymbolRef::Import { module, name } => self.resolve_import(index, module, name), + SymbolRef::Fixture(name) => self.resolve_fixture(index, name), + SymbolRef::Parameter { .. } => None, + } + } + + pub(crate) fn language(&self, index: usize) -> Option { + self.files[index] + .observations + .first() + .map(|observation| observation.language) + } + + /// Qualifies a function defined in file `index` with its module path. + /// + /// Names starting with `@` identify repository-wide functions. + pub(crate) fn function_key(&self, index: usize, name: &str) -> Vec { + if name.starts_with('@') { + return vec![name.to_owned()]; + } + let mut segments = path_segments(self.files[index].path); + match self.language(index) { + Some(SourceLanguage::Go) => { + segments.pop(); + } + Some(SourceLanguage::Rust) => { + if let Some(source_root) = segments.iter().rposition(|segment| segment == "src") { + segments.drain(..=source_root); + } + if segments + .last() + .is_some_and(|stem| matches!(stem.as_str(), "main" | "lib" | "mod")) + { + segments.pop(); + } + } + _ => { + if segments + .last() + .is_some_and(|stem| matches!(stem.as_str(), "index" | "__init__")) + { + segments.pop(); + } + } + } + segments.push(name.to_owned()); + segments + } + + fn resolve_call(&self, index: usize, path: &str) -> Option> { + let segments = path + .split("::") + .flat_map(|segment| segment.split('.')) + .filter(|segment| { + !segment.is_empty() && !matches!(*segment, "crate" | "self" | "super" | "this") + }) + .map(str::to_owned) + .collect::>(); + if segments.len() == 1 { + let local = self.function_key(index, &segments[0]); + if self.functions.contains(&local) { + return Some(local); + } + } + let last = segments.last()?; + let suffix_matches = self + .functions + .iter() + .filter(|candidate| candidate.ends_with(&segments)) + .collect::>(); + if !suffix_matches.is_empty() { + return unique(suffix_matches.into_iter()).cloned(); + } + // Import aliases, such as Go `orders "example.com/internal/http"`, rename the qualifier. + unique( + self.functions + .iter() + .filter(|candidate| candidate.last() == Some(last)), + ) + .cloned() + } + + fn resolve_import(&self, index: usize, module: &str, name: &str) -> Option { + let file = |target: usize, name: &str| SymbolKey::Local { + file: target, + name: name.to_owned(), + }; + if self.language(index) == Some(SourceLanguage::Python) { + let separator = if module.ends_with('.') { "" } else { "." }; + if let Some((head, tail)) = name.split_once('.') + && let Some(target) = + self.python_module(index, &format!("{module}{separator}{head}")) + { + return Some(file(target, tail)); + } + return self + .python_module(index, module) + .map(|target| file(target, name)); + } + self.script_module(index, module) + .map(|target| file(target, name)) + } + + /// A fixture declared in the same file, or else in the nearest enclosing `conftest.py`. + fn resolve_fixture(&self, index: usize, name: &str) -> Option { + let declares = |file: usize| { + self.files[file].observations.iter().any(|observation| { + observation.symbol_name.as_deref() == Some(name) + && observation.role != crate::SourceRole::Test + }) + }; + if declares(index) { + return Some(SymbolKey::Local { + file: index, + name: name.to_owned(), + }); + } + let mut directory = path_components(self.files[index].path); + while directory.pop().is_some() { + let candidate = directory + .iter() + .copied() + .chain(["conftest.py"]) + .collect::>() + .join("/"); + if let Some(&file) = self.paths.get(candidate.as_str()) + && declares(file) + { + return Some(SymbolKey::Local { + file, + name: name.to_owned(), + }); + } + } + None + } + + fn script_module(&self, index: usize, module: &str) -> Option { + if !module.starts_with('.') { + return None; + } + let mut base = path_components(self.files[index].path); + base.pop(); + let joined = join_components(base, module.split('/'))?.join("/"); + let stem = SCRIPT_EXTENSIONS + .iter() + .find_map(|extension| joined.strip_suffix(&format!(".{extension}"))) + .unwrap_or(&joined); + std::iter::once(joined.clone()) + .chain( + SCRIPT_EXTENSIONS + .iter() + .map(|extension| format!("{stem}.{extension}")), + ) + .chain( + SCRIPT_EXTENSIONS + .iter() + .map(|extension| format!("{joined}/index.{extension}")), + ) + .find_map(|candidate| self.paths.get(candidate.as_str()).copied()) + } + + fn python_module(&self, index: usize, module: &str) -> Option { + let dots = module + .chars() + .take_while(|character| *character == '.') + .count(); + let rest = module[dots..] + .split('.') + .filter(|segment| !segment.is_empty()); + let target = if dots == 0 { + rest.map(str::to_owned).collect::>() + } else { + let mut base = path_components(self.files[index].path); + base.pop(); + for _ in 1..dots { + base.pop()?; + } + base.into_iter() + .chain(rest) + .map(str::to_owned) + .collect::>() + }; + if target.is_empty() { + return None; + } + unique(self.paths.iter().filter_map(|(path, candidate)| { + let mut segments = path_segments(path); + if segments.last().is_some_and(|stem| stem == "__init__") { + segments.pop(); + } + let is_python = Path::new(path) + .extension() + .is_some_and(|extension| extension.eq_ignore_ascii_case("py")); + (is_python && segments.ends_with(&target)).then_some(*candidate) + })) + } +} + +pub(crate) fn unique(mut candidates: impl Iterator) -> Option { + let first = candidates.next()?; + candidates.next().is_none().then_some(first) +} + +/// Path components with the file extension removed from the last one. +fn path_segments(path: &str) -> Vec { + let mut segments = path_components(path) + .into_iter() + .map(str::to_owned) + .collect::>(); + if let Some(last) = segments.last_mut() + && let Some((stem, _)) = last.rsplit_once('.') + { + *last = stem.to_owned(); + } + segments +} + +fn path_components(path: &str) -> Vec<&str> { + path.split('/') + .filter(|component| !component.is_empty() && *component != ".") + .collect() +} + +fn join_components<'p>( + mut base: Vec<&'p str>, + relative: impl Iterator, +) -> Option> { + for component in relative { + match component { + "" | "." => {} + ".." => { + base.pop()?; + } + component => base.push(component), + } + } + Some(base) +} diff --git a/crates/code-system-graph-core/src/router_mounts.rs b/crates/code-system-graph-core/src/router_mounts.rs new file mode 100644 index 0000000..72b00e0 --- /dev/null +++ b/crates/code-system-graph-core/src/router_mounts.rs @@ -0,0 +1,654 @@ +//! Repository-level composition of router prefixes declared across source files. +//! +//! Extractors record the router each provider route is registered on and every mount of one +//! router under a prefix on another. Composition resolves those references per repository, through +//! file-local names, relative and package imports, and function paths, then rewrites each route +//! path with every prefix chain that reaches it. + +use std::collections::{BTreeMap, BTreeSet}; + +use crate::repository_symbols::{ + RepositorySourceFile, RepositorySymbols, SymbolFile, SymbolKey, files_by_repository +}; +use crate::{ + SourceEpistemicStatus, SourceLanguage, SourceObservation, SourceRole, SymbolRef, normalize_source_http_path +}; + +/// Longest mount chain followed from a route to the application root. +const MAX_MOUNT_DEPTH: usize = 8; +/// Most composed paths retained for one router. +const MAX_COMPOSED_PREFIXES: usize = 16; + +/// Rewrites provider paths with the router prefixes mounted above them. +/// +/// The result has one observation list per input file, in input order. Mount observations are +/// consumed, a route on a router mounted at several places is repeated once per composed path, +/// and routes whose router is not mounted anywhere keep their path. +#[must_use] +pub fn compose_router_mounts(files: &[RepositorySourceFile<'_>]) -> Vec> { + let by_repository = files_by_repository(files); + let mut output = files + .iter() + .map(|file| file.observations.to_vec()) + .collect::>(); + for members in by_repository.into_values() { + let repository = RepositoryRouters::new(files, &members); + if repository.mounts.is_empty() { + for &index in &members { + output[index].retain(|observation| observation.role != SourceRole::Mount); + } + continue; + } + let mut cache = BTreeMap::new(); + for (local, &index) in members.iter().enumerate() { + output[index] = repository.rewrite(local, &mut cache); + } + } + output +} + +struct RepositoryRouters<'a> { + symbols: RepositorySymbols<'a>, + mounts: BTreeMap, String)>>, +} + +impl<'a> RepositoryRouters<'a> { + fn new(files: &'a [RepositorySourceFile<'a>], members: &[usize]) -> Self { + let mut symbols = RepositorySymbols::new( + members + .iter() + .map(|&index| SymbolFile { + path: files[index].path, + observations: files[index].observations, + }) + .collect(), + ); + for local in 0..members.len() { + for observation in symbols.files()[local].observations { + for reference in [&observation.router, &observation.mount_parent] + .into_iter() + .flatten() + { + if let SymbolRef::Function(name) = reference { + symbols.declare_function(local, name); + } + } + } + } + let mut mounts = BTreeMap::, String)>>::new(); + for local in 0..members.len() { + for observation in symbols.files()[local] + .observations + .iter() + .filter(|observation| observation.role == SourceRole::Mount) + { + let Some(child) = observation + .router + .as_ref() + .and_then(|reference| symbols.resolve(local, reference)) + else { + continue; + }; + let parent = observation + .mount_parent + .as_ref() + .and_then(|reference| symbols.resolve(local, reference)); + if parent.as_ref() == Some(&child) { + continue; + } + let prefix = observation + .path + .as_deref() + .map(prefix_text) + .unwrap_or_default(); + mounts.entry(child).or_default().push((parent, prefix)); + } + } + for edges in mounts.values_mut() { + edges.sort(); + edges.dedup(); + } + Self { symbols, mounts } + } + + fn rewrite( + &self, + local: usize, + cache: &mut BTreeMap>, + ) -> Vec { + let observations = self.symbols.files()[local].observations; + let mut output = Vec::with_capacity(observations.len()); + for observation in observations { + match observation.role { + SourceRole::Mount => {} + SourceRole::Provider => { + let prefixes = observation + .router + .as_ref() + .and_then(|reference| self.symbols.resolve(local, reference)) + .map(|key| self.prefixes(&key, cache)) + .unwrap_or_default(); + match (&observation.path, prefixes.as_slice()) { + (Some(path), [_, ..]) if prefixes != [""] => { + output.extend(prefixes.iter().map(|prefix| { + let mut composed = observation.clone(); + composed.path = + Some(normalize_source_http_path(&format!("{prefix}{path}"))); + composed + })); + } + _ => output.push(observation.clone()), + } + } + SourceRole::Consumer + | SourceRole::Test + | SourceRole::Factory + | SourceRole::Call + | SourceRole::Client => { + output.push(observation.clone()); + } + } + } + output + } + + /// Returns every prefix chain from `key` to a root, outermost prefix first. + fn prefixes( + &self, + key: &SymbolKey, + cache: &mut BTreeMap>, + ) -> Vec { + if let Some(cached) = cache.get(key) { + return cached.clone(); + } + let mut visiting = Vec::new(); + let mut prefixes = self.walk(key, &mut visiting); + if prefixes.is_empty() { + prefixes.push(String::new()); + } + cache.insert(key.clone(), prefixes.clone()); + prefixes + } + + /// Chains through a cycle or deeper than [`MAX_MOUNT_DEPTH`] are dropped. + fn walk(&self, key: &SymbolKey, visiting: &mut Vec) -> Vec { + let Some(edges) = self.mounts.get(key) else { + return vec![String::new()]; + }; + if visiting.len() >= MAX_MOUNT_DEPTH || visiting.contains(key) { + return Vec::new(); + } + visiting.push(key.clone()); + let mut prefixes = BTreeSet::new(); + for (parent, prefix) in edges { + let outer = parent + .as_ref() + .map_or_else(|| vec![String::new()], |parent| self.walk(parent, visiting)); + prefixes.extend(outer.into_iter().map(|outer| format!("{outer}{prefix}"))); + } + visiting.pop(); + prefixes.into_iter().take(MAX_COMPOSED_PREFIXES).collect() + } +} + +/// Normalized prefix text without a trailing slash; the root prefix is empty. +fn prefix_text(path: &str) -> String { + let normalized = normalize_source_http_path(path); + if normalized == "/" { + String::new() + } else { + normalized + } +} + +/// Builds a mount observation recorded by a source extractor. +#[must_use] +pub(crate) fn mount_observation( + language: SourceLanguage, + framework: crate::SourceFramework, + child: SymbolRef, + parent: Option, + prefix: Option<&str>, + lines: crate::SourceLineRange, +) -> SourceObservation { + SourceObservation { + language, + framework, + role: SourceRole::Mount, + method: None, + path: Some(normalize_source_http_path(prefix.unwrap_or("/"))), + symbol_name: None, + related_symbol: None, + related_path: None, + authority: None, + router: Some(child), + mount_parent: parent, + url: None, + call: None, + lines, + status: SourceEpistemicStatus::Confirmed, + confidence: 1.0, + warnings: Vec::new(), + } +} + +#[cfg(test)] +mod tests { + use code_system_graph_model::RepoId; + + use super::{compose_router_mounts, mount_observation}; + use crate::{ + RepositorySourceFile, SourceEpistemicStatus, SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SourceRole, SymbolRef + }; + + const LINES: SourceLineRange = SourceLineRange { start: 1, end: 1 }; + + fn route(language: SourceLanguage, path: &str, router: SymbolRef) -> SourceObservation { + SourceObservation { + language, + framework: SourceFramework::Express, + role: SourceRole::Provider, + method: Some("GET".to_owned()), + path: Some(path.to_owned()), + symbol_name: Some("handler".to_owned()), + related_symbol: None, + related_path: None, + authority: None, + router: Some(router), + mount_parent: None, + url: None, + call: None, + lines: LINES, + status: SourceEpistemicStatus::Confirmed, + confidence: 1.0, + warnings: Vec::new(), + } + } + + fn mount( + language: SourceLanguage, + child: SymbolRef, + parent: Option, + prefix: &str, + ) -> SourceObservation { + mount_observation( + language, + SourceFramework::Express, + child, + parent, + Some(prefix), + LINES, + ) + } + + fn composed_paths(files: &[(&str, Vec)]) -> Vec { + let repo = RepoId::new("repo:api"); + let inputs = files + .iter() + .map(|(path, observations)| RepositorySourceFile { + repo_id: &repo, + path, + observations, + }) + .collect::>(); + let mut paths = compose_router_mounts(&inputs) + .into_iter() + .flatten() + .filter_map(|observation| observation.path) + .collect::>(); + paths.sort(); + paths + } + + fn local(name: &str) -> SymbolRef { + SymbolRef::Local(name.to_owned()) + } + + #[test] + fn script_imports_and_default_exports_should_compose_nested_prefixes() { + let js = SourceLanguage::TypeScript; + let files = [ + ( + "src/routes/orders.ts", + vec![ + route(js, "/:id", local("router")), + mount(js, local("router"), Some(local("default")), "/"), + ], + ), + ( + "src/routes/index.ts", + vec![mount( + js, + SymbolRef::Import { + module: "./orders.js".to_owned(), + name: "default".to_owned(), + }, + Some(local("api")), + "/orders", + )], + ), + ( + "src/app.ts", + vec![ + mount( + js, + SymbolRef::Import { + module: "./routes".to_owned(), + name: "api".to_owned(), + }, + Some(local("app")), + "/v1", + ), + mount( + js, + SymbolRef::Import { + module: "./routes/index".to_owned(), + name: "api".to_owned(), + }, + Some(local("app")), + "/v2/", + ), + ], + ), + ]; + + assert_eq!( + composed_paths(&files), + ["/v1/orders/:id".to_owned(), "/v2/orders/:id".to_owned()] + ); + } + + #[test] + fn python_modules_should_resolve_absolute_and_relative_imports() { + let py = SourceLanguage::Python; + let files = [ + ( + "src/app/api/users.py", + vec![route(py, "/users/{id}", local("router"))], + ), + ( + "src/app/api/__init__.py", + vec![mount( + py, + SymbolRef::Import { + module: ".users".to_owned(), + name: "router".to_owned(), + }, + Some(local("api")), + "/internal", + )], + ), + ( + "src/app/main.py", + vec![mount( + py, + SymbolRef::Import { + module: "app".to_owned(), + name: "api.api".to_owned(), + }, + Some(local("app")), + "/v1", + )], + ), + ]; + + assert_eq!( + composed_paths(&files), + ["/v1/internal/users/{id}".to_owned()] + ); + } + + #[test] + fn function_routers_should_resolve_by_module_path() { + let rust = SourceLanguage::Rust; + let files = [ + ( + "crates/api/src/orders/mod.rs", + vec![route( + rust, + "/{id}", + SymbolRef::Function("routes".to_owned()), + )], + ), + ( + "crates/api/src/users.rs", + vec![route( + rust, + "/{id}", + SymbolRef::Function("routes".to_owned()), + )], + ), + ( + "crates/api/src/main.rs", + vec![ + mount( + rust, + SymbolRef::Call("crate::orders::routes".to_owned()), + Some(SymbolRef::Function("app".to_owned())), + "/orders", + ), + mount( + rust, + SymbolRef::Call("routes".to_owned()), + Some(SymbolRef::Function("app".to_owned())), + "/ambiguous", + ), + ], + ), + ]; + + assert_eq!( + composed_paths(&files), + ["/orders/{id}".to_owned(), "/{id}".to_owned()] + ); + } + + fn extracted_routes(files: &[(&str, &str)]) -> Vec { + let repo = RepoId::new("repo:api"); + let observations = files + .iter() + .map(|(path, source)| { + let extension = path.rsplit_once('.').map(|(_, extension)| extension); + match extension { + Some("py") => crate::parse_python_source(source), + Some("rs") => crate::parse_rust_source(source), + Some("go") => crate::parse_go_source(source), + Some("java") => crate::parse_java_source(source), + Some("js") => crate::parse_javascript_source_at_path(path, source), + _ => crate::parse_typescript_source_at_path(path, source), + } + }) + .collect::>(); + let inputs = files + .iter() + .zip(&observations) + .map(|((path, _), observations)| RepositorySourceFile { + repo_id: &repo, + path, + observations, + }) + .collect::>(); + let mut routes = compose_router_mounts(&inputs) + .into_iter() + .flatten() + .filter(|observation| { + observation.role == SourceRole::Provider + && observation.status == SourceEpistemicStatus::Confirmed + }) + .filter_map(|observation| { + Some(format!( + "{} {} {}", + observation.method?, observation.path?, observation.symbol_name? + )) + }) + .collect::>(); + routes.sort(); + routes + } + + #[test] + fn fastapi_and_flask_mounts_should_compose_across_modules() { + let routes = extracted_routes(&[ + ( + "app/api/users.py", + "from fastapi import APIRouter\nrouter = APIRouter(prefix=\"/users\")\n\n@router.get(\"/{user_id}\")\ndef read_user(user_id: int):\n return {}\n", + ), + ( + "app/main.py", + "from fastapi import FastAPI\nfrom app.api import users\nfrom app.api.users import router as users_router\napp = FastAPI()\napp.include_router(users.router, prefix=\"/v1\")\napp.include_router(users_router, prefix=\"/v2\")\n", + ), + ( + "shop/views.py", + "from flask import Blueprint\nbp = Blueprint(\"orders\", __name__, url_prefix=\"/orders\")\n\n@bp.route(\"/\", methods=[\"GET\"])\ndef show(order_id):\n return ''\n", + ), + ( + "shop/__init__.py", + "from flask import Flask\nfrom .views import bp\napp = Flask(__name__)\napp.register_blueprint(bp, url_prefix=\"/shop\")\n", + ), + ]); + + assert_eq!( + routes, + [ + "GET /shop/orders/ show", + "GET /v1/users/{user_id} read_user", + "GET /v2/users/{user_id} read_user", + ] + ); + } + + #[test] + fn express_and_nest_mounts_should_compose_across_modules() { + let routes = extracted_routes(&[ + ( + "src/routes/orders.ts", + "import { Router } from 'express';\nconst router = Router();\nrouter.get('/:id', getOrder);\nexport default router;\n", + ), + ( + "src/routes/users.js", + "const express = require('express');\nfunction buildUsers() {\n const users = express.Router();\n users.get('/:id', getUser);\n return users;\n}\nmodule.exports = { buildUsers };\n", + ), + ( + "src/app.ts", + "import express from 'express';\nimport orders from './routes/orders';\nconst { buildUsers } = require('./routes/users');\nconst app = express();\nconst api = express.Router();\napi.use('/orders', orders);\napi.use('/users', authenticate, buildUsers());\napp.use('/api', api);\n", + ), + ( + "src/catalog/catalog.controller.ts", + "import { Controller, Get } from '@nestjs/common';\n@Controller('catalog')\nexport class CatalogController {\n @Get(':sku')\n findOne() {}\n @Get()\n findAll() {}\n}\n", + ), + ( + "src/main.ts", + "import { NestFactory } from '@nestjs/core';\nasync function bootstrap() {\n const app = await NestFactory.create(AppModule);\n app.setGlobalPrefix('v2');\n}\n", + ), + ]); + + assert_eq!( + routes, + [ + "GET /api/orders/:id getOrder", + "GET /api/users/:id getUser", + "GET /v2/catalog findAll", + "GET /v2/catalog/:sku findOne", + ] + ); + } + + #[test] + fn spring_class_mappings_should_prefix_method_mappings() { + let routes = extracted_routes(&[( + "src/main/java/shop/OrderController.java", + "import org.springframework.web.bind.annotation.*;\n@RestController\n@RequestMapping(\"/api/orders\")\npublic class OrderController {\n @GetMapping(\"/{id}\")\n public Order get(@PathVariable long id) { return null; }\n @PostMapping\n public Order create() { return null; }\n}\n", + )]); + + assert_eq!( + routes, + ["GET /api/orders/{id} get", "POST /api/orders create"] + ); + } + + #[test] + fn gin_chi_and_net_http_mounts_should_compose_across_packages() { + let routes = extracted_routes(&[ + ( + "internal/routes/orders.go", + "package routes\n\nimport \"github.com/gin-gonic/gin\"\n\nfunc RegisterOrders(rg *gin.RouterGroup) {\n\trg.GET(\"/:id\", getOrder)\n}\n", + ), + ( + "cmd/api/main.go", + "package main\n\nimport (\n\t\"github.com/gin-gonic/gin\"\n\t\"example.com/internal/routes\"\n)\n\nfunc main() {\n\tr := gin.Default()\n\tv1 := r.Group(\"/v1\")\n\t{\n\t\tv1.GET(\"/health\", health)\n\t}\n\troutes.RegisterOrders(v1.Group(\"/orders\"))\n}\n", + ), + ( + "billing/server.go", + "package billing\n\nimport (\n\t\"net/http\"\n\t\"github.com/go-chi/chi/v5\"\n)\n\nfunc invoices() http.Handler {\n\tr := chi.NewRouter()\n\tr.Get(\"/{id}\", getInvoice)\n\treturn r\n}\n\nfunc Router() http.Handler {\n\tr := chi.NewRouter()\n\tr.Route(\"/billing\", func(r chi.Router) {\n\t\tr.Get(\"/status\", status)\n\t\tr.Mount(\"/invoices\", invoices())\n\t})\n\tmux := http.NewServeMux()\n\tmux.HandleFunc(\"GET /payments/{id}\", getPayment)\n\treturn r\n}\n", + ), + ]); + + assert_eq!( + routes, + [ + "GET /billing/invoices/{id} getInvoice", + "GET /billing/status status", + "GET /payments/{id} getPayment", + "GET /v1/health health", + "GET /v1/orders/:id getOrder", + ] + ); + } + + #[test] + fn axum_and_actix_mounts_should_compose_across_modules() { + let routes = extracted_routes(&[ + ( + "orders/src/orders.rs", + "use axum::{routing::get, Router};\npub fn routes() -> Router {\n Router::new().route(\"/{id}\", get(get_order))\n}\n", + ), + ( + "orders/src/main.rs", + "use axum::{routing::get, Router};\nfn app() -> Router {\n let api = Router::new().route(\"/health\", get(health));\n Router::new()\n .nest(\"/orders\", crate::orders::routes())\n .nest(\"/api\", api)\n}\n", + ), + ( + "billing/src/handlers.rs", + "use actix_web::{get, web, HttpResponse};\n#[get(\"/invoices/{id}\")]\nasync fn get_invoice() -> HttpResponse { HttpResponse::Ok().finish() }\npub fn config(cfg: &mut web::ServiceConfig) {\n cfg.service(web::scope(\"/admin\").route(\"/audit\", web::get().to(audit)));\n}\n", + ), + ( + "billing/src/main.rs", + "use actix_web::{web, App, HttpServer};\nasync fn main() {\n HttpServer::new(|| {\n App::new()\n .service(web::scope(\"/api\").service(handlers::get_invoice))\n .configure(handlers::config)\n });\n}\n", + ), + ]); + + assert_eq!( + routes, + [ + "GET /admin/audit audit", + "GET /api/health health", + "GET /api/invoices/{id} get_invoice", + "GET /orders/{id} get_order", + ] + ); + } + + #[test] + fn mount_cycles_and_unresolved_routers_should_keep_route_paths() { + let go = SourceLanguage::Go; + let files = [( + "internal/http/routes.go", + vec![ + route(go, "/orders", local("a")), + route(go, "/users", local("lonely")), + mount(go, local("a"), Some(local("b")), "/a"), + mount(go, local("b"), Some(local("a")), "/b"), + mount( + go, + SymbolRef::Call("missing.Register".to_owned()), + None, + "/x", + ), + ], + )]; + + assert_eq!( + composed_paths(&files), + ["/orders".to_owned(), "/users".to_owned()] + ); + } +} diff --git a/crates/code-system-graph-core/src/routes.rs b/crates/code-system-graph-core/src/routes.rs new file mode 100644 index 0000000..5a1c0a5 --- /dev/null +++ b/crates/code-system-graph-core/src/routes.rs @@ -0,0 +1,1218 @@ +//! Canonical HTTP route shapes and the single provider index shared by every HTTP link kind. +//! +//! Providers are indexed by method and canonical shape. Consumers, tests, and implementation +//! anchors resolve through the same index, so a concrete request path such as `/orders/42` reaches +//! the `/orders/{id}` template regardless of the framework syntax that declared it. + +use std::collections::BTreeMap; + +use code_system_graph_model::{ + Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, HttpLinkGap, HttpLinkGapReason, HttpLinkReport, Node, NodeId, RepoId, stable_id +}; + +use crate::linker::{http_link_ambiguity, sort_http_ambiguities}; +use crate::{ + BoundaryRole, DeclaredImplementation, DeclaredTestCase, HttpBoundary, HttpLinkAmbiguity, HttpLinkResolution, InfrastructureDocument, InfrastructureResourceKind +}; + +/// One segment of a canonical route template. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash)] +pub enum RouteSegment { + /// Literal path segment. + Static(String), + /// Single-segment parameter such as `{id}`, `:id`, ``, `{id:int}`, or `[id]`. + Param, + /// Literal pieces separated by single-segment parameters, such as `{name}.json`. + /// The first and last pieces may be empty; parameter names are discarded. + Mixed(Vec), + /// Parameter that consumes one or more remaining segments, such as `{*rest}`, `{path...}`, + /// ``, `[...slug]`, or `*filepath`. + CatchAll, +} + +/// Framework-independent route template. Parameter names are not part of the shape. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash)] +pub struct RouteShape { + segments: Vec, +} + +impl RouteShape { + /// Parses a normalized path or route template. Query strings and fragments are ignored. + #[must_use] + pub fn parse(path: &str) -> Self { + let path = path.split(['?', '#']).next().unwrap_or_default(); + let mut segments = Vec::new(); + for segment in path.split('/').filter(|segment| !segment.is_empty()) { + let classified = classify_segment(segment); + let terminal = classified == RouteSegment::CatchAll; + segments.push(classified); + if terminal { + break; + } + } + Self { segments } + } + + /// Returns the canonical segments in path order. + #[must_use] + pub fn segments(&self) -> &[RouteSegment] { + &self.segments + } + + /// Returns `true` when the shape has no parameter segment. + #[must_use] + pub fn is_concrete(&self) -> bool { + self.segments + .iter() + .all(|segment| matches!(segment, RouteSegment::Static(_))) + } + + /// Renders the identity form: parameters become `{}` and catch-alls become `{*}`. + #[must_use] + pub fn canonical(&self) -> String { + if self.segments.is_empty() { + return "/".to_owned(); + } + let mut rendered = String::new(); + for segment in &self.segments { + rendered.push('/'); + match segment { + RouteSegment::Static(value) => rendered.push_str(value), + RouteSegment::Param => rendered.push_str("{}"), + RouteSegment::Mixed(literals) => rendered.push_str(&literals.join("{}")), + RouteSegment::CatchAll => rendered.push_str("{*}"), + } + } + rendered + } +} + +/// Returns the canonical identity form of a path or route template. +#[must_use] +pub fn canonical_route(path: &str) -> String { + RouteShape::parse(path).canonical() +} + +fn classify_segment(segment: &str) -> RouteSegment { + if segment == "*" || segment == "**" { + return RouteSegment::CatchAll; + } + if let Some(inner) = segment + .strip_prefix('{') + .and_then(|value| value.strip_suffix('}')) + .filter(|inner| !inner.contains(['{', '}'])) + { + return if inner.starts_with('*') + || inner.ends_with("...") + || inner.ends_with(":path") + || inner.ends_with(":.*") + || inner.ends_with(":.+") + { + RouteSegment::CatchAll + } else { + RouteSegment::Param + }; + } + if (segment.starts_with("[...") && segment.ends_with(']')) + || (segment.starts_with("[[...") && segment.ends_with("]]")) + { + return RouteSegment::CatchAll; + } + if segment.len() > 2 && segment.starts_with('[') && segment.ends_with(']') { + return RouteSegment::Param; + } + if let Some(inner) = segment + .strip_prefix('<') + .and_then(|value| value.strip_suffix('>')) + { + return if inner.starts_with("path:") { + RouteSegment::CatchAll + } else { + RouteSegment::Param + }; + } + if let Some(name) = segment.strip_prefix(':').filter(|name| !name.is_empty()) { + return if name.ends_with('*') || name.ends_with('+') { + RouteSegment::CatchAll + } else { + RouteSegment::Param + }; + } + if segment.len() > 1 && segment.starts_with('*') { + return RouteSegment::CatchAll; + } + if let Some(literals) = mixed_segment_literals(segment) { + if literals.len() == 2 && literals.iter().all(String::is_empty) { + return RouteSegment::Param; + } + return RouteSegment::Mixed(literals); + } + RouteSegment::Static(segment.to_owned()) +} + +fn mixed_segment_literals(segment: &str) -> Option> { + let mut remaining = segment; + let mut literals = Vec::new(); + while let Some((prefix, parameter)) = remaining.split_once('{') { + let (name, suffix) = parameter.split_once('}')?; + if name.contains('{') || prefix.contains('}') { + return None; + } + // JavaScript interpolation uses `${name}` for the same path parameter. + literals.push(prefix.strip_suffix('$').unwrap_or(prefix).to_owned()); + remaining = suffix; + } + if literals.is_empty() || remaining.contains('}') { + return None; + } + literals.push(remaining.to_owned()); + Some(literals) +} + +fn mixed_segment_matches(literals: &[String], value: &str) -> bool { + let [prefix, middle @ .., suffix] = literals else { + return false; + }; + let Some(remaining) = value.strip_prefix(prefix.as_str()) else { + return false; + }; + let Some(mut remaining) = remaining.strip_suffix(suffix.as_str()) else { + return false; + }; + for literal in middle { + // Each parameter consumes at least one character, including adjacent parameters. + let Some(first) = remaining.chars().next() else { + return false; + }; + remaining = &remaining[first.len_utf8()..]; + let Some((_, rest)) = remaining.split_once(literal.as_str()) else { + return false; + }; + remaining = rest; + } + !remaining.is_empty() +} + +/// Returns the normalized authority of an absolute `http`/`https` URL. +pub(crate) fn url_authority(url: &str) -> Option { + let rest = url + .strip_prefix("http://") + .or_else(|| url.strip_prefix("https://"))?; + let authority = rest.split(['/', '?', '#']).next().unwrap_or_default(); + normalize_authority(authority) +} + +/// Normalizes a `host[:port]` authority: lower-case, without user information or a trailing +/// root dot. Returns `None` for an empty host or a value containing a scheme, path, or whitespace. +#[must_use] +pub fn normalize_authority(value: &str) -> Option { + let value = value.trim(); + if value.is_empty() || value.contains(['/', '?', '#', ' ', '\t']) || value.contains("://") { + return None; + } + let host_and_port = value.rsplit_once('@').map_or(value, |(_, host)| host); + let (host, port) = split_port(host_and_port); + let host = host.trim_end_matches('.').to_ascii_lowercase(); + if host.is_empty() || port.is_some_and(|port| port.parse::().is_err()) { + return None; + } + Some(match port { + Some(port) if host.contains(':') => format!("[{host}]:{port}"), + Some(port) => format!("{host}:{port}"), + None => host, + }) +} + +fn split_port(authority: &str) -> (&str, Option<&str>) { + if let Some(rest) = authority.strip_prefix('[') { + return match rest.split_once(']') { + Some((host, tail)) => (host, tail.strip_prefix(':')), + None => (authority, None), + }; + } + match authority.rsplit_once(':') { + Some((host, port)) if !host.contains(':') => (host, Some(port)), + _ => (authority, None), + } +} + +fn authority_host(authority: &str) -> &str { + split_port(authority).0 +} + +/// Returns `true` for loopback and wildcard hosts used by locally started services and tests. +fn is_loopback(host: &str) -> bool { + matches!( + host, + "localhost" | "::1" | "0.0.0.0" | "host.docker.internal" + ) || host.ends_with(".localhost") + || host.starts_with("127.") +} + +/// Reduces cluster-internal service DNS names such as `orders.shop.svc.cluster.local` to the +/// service name. +fn service_host(host: &str) -> &str { + let trimmed = host + .strip_suffix(".svc.cluster.local") + .or_else(|| host.strip_suffix(".svc")); + trimmed.map_or(host, |name| name.split('.').next().unwrap_or(name)) +} + +/// Repository that serves each known authority. +/// +/// Declared authorities come from the manifest and strictly restrict resolution to their +/// repository. Inferred authorities come from deployment declarations; when the inferred repository +/// has no matching route, resolution falls back to workspace rules. A host inferred for several +/// repositories is not used. +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub struct AuthorityMap { + declared: BTreeMap, + inferred: BTreeMap>, +} + +/// Resolution target of one call authority. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum AuthorityTarget<'a> { + Workspace, + Declared(&'a RepoId), + Inferred(&'a RepoId), + External, +} + +impl AuthorityMap { + /// Creates an empty map. + #[must_use] + pub fn new() -> Self { + Self::default() + } + + /// Declares the repository serving a normalized `host` or `host:port` authority. + pub fn declare(&mut self, authority: String, repository: RepoId) { + self.declared.insert(authority, repository); + } + + /// Records a host inferred from a deployment declaration in `repository`. + pub fn infer(&mut self, host: &str, repository: &RepoId) { + let Some(host) = normalize_authority(host) else { + return; + }; + let host = service_host(authority_host(&host)).to_owned(); + self.inferred + .entry(host) + .and_modify(|existing| { + if existing.as_ref() != Some(repository) { + *existing = None; + } + }) + .or_insert_with(|| Some(repository.clone())); + } + + /// Records every service, deployment, and host alias declared by an infrastructure document. + pub fn infer_from_infrastructure( + &mut self, + repository: &RepoId, + document: &InfrastructureDocument, + ) { + for unit in &document.deployment_units { + for host in std::iter::once(&unit.name) + .chain(&unit.service_names) + .chain(&unit.host_aliases) + { + self.infer(host, repository); + } + } + for resource in document + .resources + .iter() + .filter(|resource| resource.kind == InfrastructureResourceKind::Service) + { + self.infer(&resource.name, repository); + } + } + + fn lookup(&self, authority: &str) -> AuthorityTarget<'_> { + let host = authority_host(authority); + if is_loopback(host) { + return AuthorityTarget::Workspace; + } + if let Some(repository) = self + .declared + .get(authority) + .or_else(|| self.declared.get(host)) + { + return AuthorityTarget::Declared(repository); + } + match self.inferred.get(service_host(host)) { + Some(Some(repository)) => AuthorityTarget::Inferred(repository), + Some(None) => AuthorityTarget::Workspace, + None => AuthorityTarget::External, + } + } +} + +/// How a caller constrains which repositories may provide the operation it invokes. +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub enum CallScope { + /// Any repository may provide the operation. + #[default] + Workspace, + /// The call runs in-process against the application in its own repository, such as a + /// `TestClient(app)`, `httptest`, or `MockMvc` request. + InProcess, + /// The call targets a normalized `host:port` authority. + Authority(String), + /// The call targets a host outside the workspace and is never linked. + External, +} + +/// Providers sharing one canonical shape, de-duplicated by node identity. +#[derive(Debug, Default)] +struct ProviderSet<'a> { + providers: Vec<&'a HttpBoundary>, +} + +impl<'a> ProviderSet<'a> { + fn insert(&mut self, provider: &'a HttpBoundary) { + if let Some(existing) = self + .providers + .iter_mut() + .find(|candidate| candidate.node.id == provider.node.id) + { + if provider.evidence.confidence > existing.evidence.confidence { + *existing = provider; + } + } else { + self.providers.push(provider); + } + } + + fn accepted(&self, accept: &impl Fn(&HttpBoundary) -> bool) -> Vec<&'a HttpBoundary> { + self.providers + .iter() + .copied() + .filter(|provider| accept(provider)) + .collect() + } +} + +#[derive(Debug, Default)] +struct RouteNode<'a> { + statics: BTreeMap>, + mixed: BTreeMap, RouteNode<'a>>, + param: Option>>, + catch_all: ProviderSet<'a>, + terminal: ProviderSet<'a>, +} + +impl<'a> RouteNode<'a> { + fn insert(&mut self, segments: &[RouteSegment], provider: &'a HttpBoundary) { + let Some((first, rest)) = segments.split_first() else { + self.terminal.insert(provider); + return; + }; + match first { + RouteSegment::Static(value) => self + .statics + .entry(value.clone()) + .or_default() + .insert(rest, provider), + RouteSegment::Param => self + .param + .get_or_insert_with(Box::default) + .insert(rest, provider), + RouteSegment::Mixed(literals) => self + .mixed + .entry(literals.clone()) + .or_default() + .insert(rest, provider), + RouteSegment::CatchAll => self.catch_all.insert(provider), + } + } + + /// Finds the most specific accepted providers: at every segment a static match is preferred + /// over a mixed segment, then a whole-segment parameter, then a catch-all. + fn find( + &self, + query: &[RouteSegment], + accept: &impl Fn(&HttpBoundary) -> bool, + ) -> Vec<&'a HttpBoundary> { + let Some((first, rest)) = query.split_first() else { + return self.terminal.accepted(accept); + }; + if let RouteSegment::Static(value) = first + && let Some(child) = self.statics.get(value) + { + let found = child.find(rest, accept); + if !found.is_empty() { + return found; + } + } + let mut mixed_matches = Vec::new(); + for (literals, child) in &self.mixed { + let matches = match first { + RouteSegment::Static(value) => mixed_segment_matches(literals, value), + RouteSegment::Mixed(query_literals) => literals == query_literals, + RouteSegment::Param | RouteSegment::CatchAll => false, + }; + if matches { + mixed_matches.extend(child.find(rest, accept)); + } + } + if !mixed_matches.is_empty() { + return mixed_matches; + } + if !matches!(first, RouteSegment::CatchAll) + && let Some(child) = &self.param + { + let found = child.find(rest, accept); + if !found.is_empty() { + return found; + } + } + self.catch_all.accepted(accept) + } +} + +/// Outcome of resolving one HTTP call against the provider index. +#[derive(Debug, PartialEq)] +enum RouteResolution<'a> { + Unique(&'a HttpBoundary), + Ambiguous(Vec<&'a HttpBoundary>), + Unmatched, +} + +/// Per-method tries of provider shapes. +#[derive(Debug, Default)] +struct RouteIndex<'a> { + methods: BTreeMap<&'a str, RouteNode<'a>>, +} + +impl<'a> RouteIndex<'a> { + fn new(boundaries: &'a [HttpBoundary]) -> Self { + let mut index = Self::default(); + for provider in boundaries + .iter() + .filter(|boundary| boundary.role == BoundaryRole::Provider) + { + index + .methods + .entry(provider.method.as_str()) + .or_default() + .insert(RouteShape::parse(&provider.path).segments(), provider); + } + index + } + + /// Applies the provider scope rules in order: an explicit repository restriction (mapped + /// authority or in-process client), then the caller's own repository, then a unique provider + /// in the workspace. Anything else is ambiguous. + fn resolve( + &self, + method: &str, + path: &str, + caller_repository: Option<&RepoId>, + restriction: Option<&RepoId>, + ) -> RouteResolution<'a> { + let Some(root) = self.methods.get(method) else { + return RouteResolution::Unmatched; + }; + let shape = RouteShape::parse(path); + let mut found = match restriction { + Some(repository) => root.find(shape.segments(), &|provider: &HttpBoundary| { + provider.node.repo_id.as_ref() == Some(repository) + }), + None => root.find(shape.segments(), &|_: &HttpBoundary| true), + }; + if found.len() > 1 + && let Some(caller) = caller_repository + { + let local = found + .iter() + .copied() + .filter(|provider| provider.node.repo_id.as_ref() == Some(caller)) + .collect::>(); + if local.len() == 1 { + found = local; + } + } + match found.as_slice() { + [] => RouteResolution::Unmatched, + [provider] => RouteResolution::Unique(provider), + _ => RouteResolution::Ambiguous(found), + } + } +} + +/// Repository restriction implied by a call scope. +#[derive(Debug, Clone, Copy)] +enum Restriction<'a> { + /// Resolve across the workspace. + Workspace, + /// Resolve only inside one repository. + Strict(&'a RepoId), + /// Prefer one repository and fall back to the workspace when it has no match. + Preferred(&'a RepoId), + /// The call leaves the workspace and is never linked. + External, +} + +fn scope_restriction<'a>( + scope: &'a CallScope, + caller_repository: Option<&'a RepoId>, + authorities: &'a AuthorityMap, +) -> Restriction<'a> { + match scope { + CallScope::Workspace => Restriction::Workspace, + CallScope::InProcess => { + caller_repository.map_or(Restriction::Workspace, Restriction::Strict) + } + CallScope::Authority(authority) => match authorities.lookup(authority) { + AuthorityTarget::Workspace => Restriction::Workspace, + AuthorityTarget::Declared(repository) => Restriction::Strict(repository), + AuthorityTarget::Inferred(repository) => Restriction::Preferred(repository), + AuthorityTarget::External => Restriction::External, + }, + CallScope::External => Restriction::External, + } +} + +/// Outcome of resolving one scoped call. +enum CallResolution<'a> { + Route(RouteResolution<'a>), + External, +} + +fn resolve_call<'a>( + index: &RouteIndex<'a>, + method: &str, + path: &str, + caller_repository: Option<&'a RepoId>, + scope: &'a CallScope, + authorities: &'a AuthorityMap, +) -> CallResolution<'a> { + let resolution = match scope_restriction(scope, caller_repository, authorities) { + Restriction::Workspace => index.resolve(method, path, caller_repository, None), + Restriction::Strict(repository) => { + index.resolve(method, path, caller_repository, Some(repository)) + } + Restriction::Preferred(repository) => { + match index.resolve(method, path, caller_repository, Some(repository)) { + RouteResolution::Unmatched => index.resolve(method, path, caller_repository, None), + resolution => resolution, + } + } + Restriction::External => return CallResolution::External, + }; + CallResolution::Route(resolution) +} + +/// Links HTTP consumers (`calls_remote`), tests (`validates`), and implementation anchors +/// (`implemented_by`) to providers through one route index. +/// +/// Concrete paths match provider templates; the most specific template wins and equal-shape +/// providers are narrowed by scope. Remaining ties are reported as ambiguities with their +/// candidates instead of edges. `authorities` maps the hosts named by absolute URLs to the +/// repository that serves them; calls to unknown hosts leave the workspace and are reported as +/// external rather than unmatched. +#[must_use] +pub fn link_http_routes( + boundaries: &[HttpBoundary], + tests: &[DeclaredTestCase], + implementations: &[DeclaredImplementation], + authorities: &AuthorityMap, +) -> HttpLinkResolution { + let index = RouteIndex::new(boundaries); + let mut sink = LinkSink::default(); + + let consumers = boundaries + .iter() + .filter(|boundary| boundary.role == BoundaryRole::Consumer) + .map(|consumer| RemoteCall { + node: &consumer.node, + method: &consumer.method, + path: &consumer.path, + scope: &consumer.scope, + evidence: &consumer.evidence, + kind: EdgeKind::CallsRemote, + relation: "calls_remote", + }); + let test_calls = tests.iter().map(|test| RemoteCall { + node: &test.node, + method: &test.method, + path: &test.path, + scope: &test.scope, + evidence: &test.evidence, + kind: EdgeKind::Validates, + relation: "validates", + }); + for call in consumers.chain(test_calls) { + link_remote_call(&mut sink, &index, authorities, &call); + } + + for implementation in implementations { + let Some(repository) = implementation.node.repo_id.as_ref() else { + continue; + }; + match index.resolve( + &implementation.method, + &implementation.path, + Some(repository), + Some(repository), + ) { + RouteResolution::Unique(provider) => sink.link( + &provider.node.id, + &implementation.node.id, + EdgeKind::ImplementedBy, + "implemented_by", + provider + .evidence + .confidence + .min(implementation.evidence.confidence), + vec![ + provider.evidence.id.clone(), + implementation.evidence.id.clone(), + ], + ), + RouteResolution::Ambiguous(candidates) => { + sink.ambiguous(&implementation.method, &implementation.path, &candidates); + } + RouteResolution::Unmatched => {} + } + } + + sink.finish() +} + +/// Consumer or test call resolved against provider routes. +struct RemoteCall<'a> { + node: &'a Node, + method: &'a str, + path: &'a str, + scope: &'a CallScope, + evidence: &'a Evidence, + kind: EdgeKind, + relation: &'static str, +} + +fn link_remote_call<'a>( + sink: &mut LinkSink, + index: &RouteIndex<'a>, + authorities: &'a AuthorityMap, + call: &RemoteCall<'a>, +) { + let caller = call.node.repo_id.as_ref(); + match resolve_call( + index, + call.method, + call.path, + caller, + call.scope, + authorities, + ) { + CallResolution::Route(RouteResolution::Unique(provider)) => { + sink.report.coverage.linked += 1; + sink.link( + &call.node.id, + &provider.node.id, + call.kind, + call.relation, + call.evidence.confidence.min(provider.evidence.confidence), + vec![call.evidence.id.clone(), provider.evidence.id.clone()], + ); + } + CallResolution::Route(RouteResolution::Ambiguous(candidates)) => { + sink.ambiguous(call.method, call.path, &candidates); + sink.gap(call, HttpLinkGapReason::Ambiguous, &candidates); + } + CallResolution::Route(RouteResolution::Unmatched) => { + sink.gap(call, HttpLinkGapReason::NoProvider, &[]); + } + CallResolution::External => sink.gap(call, HttpLinkGapReason::External, &[]), + } +} + +/// Accumulates de-duplicated route edges and unresolved ambiguities. +#[derive(Debug, Default)] +struct LinkSink { + edges: BTreeMap, + ambiguities: Vec, + report: HttpLinkReport, +} + +impl LinkSink { + fn link( + &mut self, + source: &NodeId, + target: &NodeId, + kind: EdgeKind, + relation: &str, + confidence: f32, + evidence: Vec, + ) { + let edge = Edge { + id: EdgeId::new(stable_id( + "edge", + &format!("{}:{relation}:{}", source.as_str(), target.as_str()), + )), + source: source.clone(), + target: target.clone(), + kind, + confidence, + status: consensus_status(confidence), + evidence, + }; + self.edges + .entry(edge.id.clone()) + .and_modify(|existing| merge_equivalent_edge(existing, &edge)) + .or_insert(edge); + } + + fn ambiguous(&mut self, method: &str, path: &str, candidates: &[&HttpBoundary]) { + self.ambiguities.push(http_link_ambiguity( + method, + path, + candidates.iter().map(|candidate| &candidate.node.id), + )); + } + + fn gap( + &mut self, + call: &RemoteCall<'_>, + reason: HttpLinkGapReason, + candidates: &[&HttpBoundary], + ) { + let coverage = &mut self.report.coverage; + match reason { + HttpLinkGapReason::NoProvider => coverage.no_provider += 1, + HttpLinkGapReason::Ambiguous => coverage.ambiguous += 1, + HttpLinkGapReason::External => coverage.external += 1, + } + let mut candidates = candidates + .iter() + .map(|candidate| candidate.node.id.clone()) + .collect::>(); + candidates.sort(); + candidates.dedup(); + self.report.gaps.push(HttpLinkGap { + caller: call.node.id.clone(), + method: call.method.to_owned(), + path: call.path.to_owned(), + reason, + candidates, + }); + } + + fn finish(mut self) -> HttpLinkResolution { + sort_http_ambiguities(&mut self.ambiguities); + self.report.gaps.sort(); + self.report.gaps.dedup(); + HttpLinkResolution { + edges: self.edges.into_values().collect(), + ambiguities: self.ambiguities, + report: self.report, + } + } +} + +fn merge_equivalent_edge(existing: &mut Edge, candidate: &Edge) { + existing.confidence = existing.confidence.max(candidate.confidence); + existing.status = consensus_status(existing.confidence); + existing.evidence.extend(candidate.evidence.iter().cloned()); + existing.evidence.sort(); + existing.evidence.dedup(); +} + +fn consensus_status(confidence: f32) -> EpistemicStatus { + if confidence >= 1.0 { + EpistemicStatus::Confirmed + } else { + EpistemicStatus::Inferred + } +} + +#[cfg(test)] +mod tests { + use code_system_graph_model::{EdgeKind, HttpLinkGapReason, RepoId}; + + use super::{ + AuthorityMap, CallScope, RouteSegment, RouteShape, canonical_route, link_http_routes, normalize_authority, url_authority + }; + use crate::{HttpBoundary, HttpConsumerConfig, extract_openapi}; + + fn provider(repo: &str, method: &str, path: &str) -> HttpBoundary { + let document = format!( + "openapi: 3.0.3\npaths:\n {path}:\n {}: {{}}\n", + method.to_ascii_lowercase() + ); + match extract_openapi(&RepoId::new(repo), "openapi.yaml", &document) { + Ok(mut boundaries) => boundaries.remove(0), + Err(error) => panic!("provider fixture must be valid: {error}"), + } + } + + fn consumer(repo: &str, method: &str, path: &str) -> HttpBoundary { + HttpBoundary::consumer( + RepoId::new(repo), + &HttpConsumerConfig { + method: method.to_owned(), + path: path.to_owned(), + source: "src/client.ts".to_owned(), + }, + ) + } + + fn linked_targets(boundaries: &[HttpBoundary]) -> Vec { + let resolution = link_http_routes(boundaries, &[], &[], &AuthorityMap::new()); + resolution + .edges + .iter() + .filter(|edge| edge.kind == EdgeKind::CallsRemote) + .filter_map(|edge| { + boundaries + .iter() + .find(|boundary| boundary.node.id == edge.target) + .map(|boundary| boundary.path.clone()) + }) + .collect() + } + + #[test] + fn route_shape_should_normalize_every_framework_parameter_syntax() { + let parameter = "/orders/{}"; + for template in [ + "/orders/{id}", + "/orders/:id", + "/orders/", + "/orders/", + "/orders/{id:int}", + "/orders/{id:[0-9]+}", + "/orders/[id]", + "/orders/${orderId}", + ] { + assert_eq!(canonical_route(template), parameter, "{template}"); + } + for template in [ + "/files/[...slug]", + "/files/[[...slug]]", + "/files/{*rest}", + "/files/{path...}", + "/files/{rest:path}", + "/files/", + "/files/*filepath", + "/files/:rest*", + "/files/*", + "/files/**", + ] { + assert_eq!(canonical_route(template), "/files/{*}", "{template}"); + } + assert_eq!(canonical_route("/"), "/"); + assert_eq!(canonical_route("/orders/42?expand=true"), "/orders/42"); + assert!(RouteShape::parse("/orders/42").is_concrete()); + assert_eq!( + RouteShape::parse("/v1/{tenant}/orders").segments(), + &[ + RouteSegment::Static("v1".to_owned()), + RouteSegment::Param, + RouteSegment::Static("orders".to_owned()), + ] + ); + } + + #[test] + fn concrete_path_should_link_to_the_most_specific_template() { + let boundaries = [ + provider("repo:api", "GET", "/orders/{id}"), + provider("repo:api", "GET", "/orders/export"), + provider("repo:files", "GET", "/{tenant}/export"), + provider("repo:files", "GET", "/orders/{rest:path}"), + consumer("repo:web", "GET", "/orders/42"), + consumer("repo:web", "GET", "/orders/export"), + consumer("repo:web", "GET", "/orders/42/lines/7"), + ]; + + let mut targets = linked_targets(&boundaries); + targets.sort(); + + assert_eq!( + targets, + vec![ + "/orders/export".to_owned(), + "/orders/{id}".to_owned(), + "/orders/{rest:path}".to_owned(), + ] + ); + } + + #[test] + fn equal_shapes_should_prefer_the_callers_repository_then_report_ambiguity() { + let local = [ + provider("repo:web", "POST", "/orders"), + provider("repo:api", "POST", "/orders"), + consumer("repo:web", "POST", "/orders"), + ]; + assert_eq!(linked_targets(&local), vec!["/orders".to_owned()]); + + let remote = [ + provider("repo:api-a", "POST", "/orders/{id}"), + provider("repo:api-b", "POST", "/orders/:orderId"), + consumer("repo:web", "POST", "/orders/9"), + ]; + let resolution = link_http_routes(&remote, &[], &[], &AuthorityMap::new()); + assert_eq!(resolution.edges, Vec::new()); + assert_eq!(resolution.ambiguities.len(), 1); + assert_eq!(resolution.ambiguities[0].path, "/orders/9"); + assert_eq!(resolution.ambiguities[0].candidates.len(), 2); + } + + #[test] + fn scoped_calls_should_resolve_only_inside_their_restriction() { + let mut in_process = consumer("repo:api-b", "GET", "/health"); + in_process.scope = CallScope::InProcess; + let mut mapped = consumer("repo:web", "GET", "/health"); + mapped.scope = CallScope::Authority("orders-api:8080".to_owned()); + let mut external = consumer("repo:web", "GET", "/health"); + external.scope = CallScope::External; + let boundaries = [ + provider("repo:api-a", "GET", "/health"), + provider("repo:api-b", "GET", "/health"), + in_process, + mapped, + external, + ]; + let mut authorities = AuthorityMap::new(); + authorities.declare("orders-api:8080".to_owned(), RepoId::new("repo:api-a")); + + let resolution = link_http_routes(&boundaries, &[], &[], &authorities); + let mut pairs = resolution + .edges + .iter() + .filter_map(|edge| { + let source = boundaries.iter().find(|item| item.node.id == edge.source)?; + let target = boundaries.iter().find(|item| item.node.id == edge.target)?; + Some(( + source.node.repo_id.clone()?.as_str().to_owned(), + target.node.repo_id.clone()?.as_str().to_owned(), + )) + }) + .collect::>(); + pairs.sort(); + + assert_eq!( + pairs, + vec![ + ("repo:api-b".to_owned(), "repo:api-b".to_owned()), + ("repo:web".to_owned(), "repo:api-a".to_owned()), + ] + ); + assert_eq!(resolution.ambiguities, Vec::new()); + assert!( + resolution + .report + .gaps + .iter() + .any(|call| call.reason == HttpLinkGapReason::External) + ); + } + + fn linked_repositories(boundaries: &[HttpBoundary], authorities: &AuthorityMap) -> Vec { + link_http_routes(boundaries, &[], &[], authorities) + .edges + .iter() + .filter_map(|edge| { + let target = boundaries.iter().find(|item| item.node.id == edge.target)?; + Some(target.node.repo_id.clone()?.as_str().to_owned()) + }) + .collect() + } + + fn authority_consumer(authority: &str) -> HttpBoundary { + let mut call = consumer("repo:web", "GET", "/orders/42"); + call.scope = CallScope::Authority(authority.to_owned()); + call + } + + #[test] + fn authorities_should_resolve_loopback_declared_inferred_and_external_hosts() { + let providers = [ + provider("repo:api-a", "GET", "/orders/{id}"), + provider("repo:api-b", "GET", "/orders/{id}"), + ]; + let mut authorities = AuthorityMap::new(); + authorities.declare("orders-api".to_owned(), RepoId::new("repo:api-a")); + authorities.infer("billing", &RepoId::new("repo:api-b")); + authorities.infer("shared", &RepoId::new("repo:api-a")); + authorities.infer("shared", &RepoId::new("repo:api-b")); + let resolve = |authority: &str| { + let mut boundaries = providers.to_vec(); + boundaries.push(authority_consumer(authority)); + linked_repositories(&boundaries, &authorities) + }; + + assert_eq!(resolve("orders-api:8080"), vec!["repo:api-a".to_owned()]); + assert_eq!( + resolve("billing.payments.svc.cluster.local"), + vec!["repo:api-b".to_owned()] + ); + assert_eq!(resolve("localhost:8080"), Vec::::new()); + assert_eq!(resolve("shared:80"), Vec::::new()); + assert_eq!(resolve("api.example.com"), Vec::::new()); + + let mut single = vec![provider("repo:api-a", "GET", "/orders/{id}")]; + single.push(authority_consumer("127.0.0.1:3000")); + assert_eq!( + linked_repositories(&single, &authorities), + vec!["repo:api-a".to_owned()] + ); + } + + #[test] + fn inferred_authority_should_fall_back_to_the_workspace_when_its_repository_has_no_route() { + let mut boundaries = vec![provider("repo:api-a", "GET", "/orders/{id}")]; + boundaries.push(authority_consumer("gateway:8080")); + let mut authorities = AuthorityMap::new(); + authorities.infer("gateway", &RepoId::new("repo:edge")); + assert_eq!( + linked_repositories(&boundaries, &authorities), + vec!["repo:api-a".to_owned()] + ); + + let mut strict = AuthorityMap::new(); + strict.declare("gateway".to_owned(), RepoId::new("repo:edge")); + assert_eq!( + linked_repositories(&boundaries, &strict), + Vec::::new() + ); + } + + #[test] + fn authority_normalization_should_reject_schemes_and_paths() { + assert_eq!( + url_authority("https://User@Orders-API.:8080/v1?x=1"), + Some("orders-api:8080".to_owned()) + ); + assert_eq!( + url_authority("http://[::1]:9000/x"), + Some("[::1]:9000".to_owned()) + ); + assert_eq!(normalize_authority("http://orders"), None); + assert_eq!(normalize_authority("orders/v1"), None); + assert_eq!(normalize_authority("orders:http"), None); + assert_eq!(normalize_authority("Orders"), Some("orders".to_owned())); + } + + #[test] + fn provider_identity_should_ignore_parameter_names() { + let openapi = provider("repo:api", "GET", "/orders/{orderId}"); + let express = provider("repo:api", "GET", "/orders/:id"); + + assert_eq!(openapi.node.id, express.node.id); + } + + #[test] + fn mixed_route_segments_should_preserve_literals_in_their_identity() { + for (template, canonical) in [ + ("/files/{name}.json", "/files/{}.json"), + ("/files/report-{name}", "/files/report-{}"), + ("/files/{name}.{format}", "/files/{}.{}"), + ("/files/${name}.json", "/files/{}.json"), + ] { + assert_eq!(canonical_route(template), canonical, "{template}"); + assert!(!RouteShape::parse(template).is_concrete(), "{template}"); + } + let json = provider("repo:api", "GET", "/files/{name}.json"); + let renamed = provider("repo:api", "GET", "/files/{id}.json"); + let plain = provider("repo:api", "GET", "/files/{id}"); + let xml = provider("repo:api", "GET", "/files/{id}.xml"); + assert_eq!(json.node.id, renamed.node.id); + assert_ne!(json.node.id, plain.node.id); + assert_ne!(json.node.id, xml.node.id); + } + + #[test] + fn mixed_route_segments_should_link_only_calls_satisfying_every_literal() { + for (template, matching, mismatching) in [ + ( + "/files/{name}.json", + "/files/readme.json", + "/files/readme.xml", + ), + ( + "/files/report-{name}", + "/files/report-readme", + "/files/other-readme", + ), + ( + "/files/{name}.{format}", + "/files/readme.json", + "/files/readme", + ), + ( + "/files/${name}.json", + "/files/readme.json", + "/files/readme.xml", + ), + ("/files/{a}{b}.json", "/files/ab.json", "/files/a.json"), + ] { + let boundaries = [ + provider("repo:api", "GET", template), + consumer("repo:web", "GET", matching), + consumer("repo:web", "GET", mismatching), + ]; + assert_eq!( + linked_targets(&boundaries), + vec![template.to_owned()], + "{template}" + ); + } + } + + #[test] + fn mixed_route_segments_should_rank_between_static_and_whole_segment_parameters() { + for (call, target) in [ + ("/files/readme.json", "/files/readme.json"), + ("/files/other.json", "/files/{name}.json"), + ("/files/readme.xml", "/files/{id}"), + ("/files/{id}.json", "/files/{name}.json"), + ] { + let boundaries = [ + provider("repo:api", "GET", "/files/{id}"), + provider("repo:api", "GET", "/files/{name}.json"), + provider("repo:api", "GET", "/files/readme.json"), + consumer("repo:web", "GET", call), + ]; + assert_eq!( + linked_targets(&boundaries), + vec![target.to_owned()], + "{call}" + ); + } + } + + #[test] + fn overlapping_mixed_route_segments_should_report_ambiguity() { + let boundaries = [ + provider("repo:api", "GET", "/files/{name}.json"), + provider("repo:api", "GET", "/files/report-{name}"), + consumer("repo:web", "GET", "/files/report-readme.json"), + ]; + let resolution = link_http_routes(&boundaries, &[], &[], &AuthorityMap::new()); + assert_eq!(resolution.edges, Vec::new()); + assert_eq!(resolution.ambiguities.len(), 1); + assert_eq!(resolution.ambiguities[0].candidates.len(), 2); + } + + #[test] + fn mixed_route_parameters_should_consume_nonempty_unicode_values() { + for (call, expected_links) in [ + ("/files/éñ.json", 1), + ("/files/é.json", 0), + ("/files/.json", 0), + ] { + let boundaries = [ + provider("repo:api", "GET", "/files/{a}{b}.json"), + consumer("repo:web", "GET", call), + ]; + assert_eq!(linked_targets(&boundaries).len(), expected_links, "{call}"); + } + } +} diff --git a/crates/code-system-graph-core/src/source_graph.rs b/crates/code-system-graph-core/src/source_graph.rs index 71a4b28..a2ac090 100644 --- a/crates/code-system-graph-core/src/source_graph.rs +++ b/crates/code-system-graph-core/src/source_graph.rs @@ -5,7 +5,7 @@ use code_system_graph_model::{ }; use crate::{ - BoundaryRole, DeclaredImplementation, DeclaredTestCase, HttpBoundary, SourceFramework, SourceLanguage, SourceObservation, SourceRole, SourceSymbolIdentity + BoundaryRole, CallScope, DeclaredImplementation, DeclaredTestCase, HttpBoundary, SourceFramework, SourceLanguage, SourceObservation, SourceRole, SourceSymbolIdentity }; /// Graph-ready facts derived from one focused Rust or Python source file. @@ -99,7 +99,7 @@ pub fn source_observations_to_graph( result.relation_edges.push(edge); result.relation_evidence.push(evidence); } - SourceRole::Test => {} + SourceRole::Test | SourceRole::Mount | SourceRole::Call | SourceRole::Client => {} } } @@ -111,12 +111,17 @@ pub fn source_observations_to_graph( observation.symbol_name.as_deref()?, observation.method.as_deref()?, observation.path.as_deref()?, + observation.authority.as_deref(), + observation.framework.is_in_process_client(), )) }) .fold( - BTreeMap::<&str, BTreeSet<(&str, &str)>>::new(), - |mut grouped, (symbol, method, path)| { - grouped.entry(symbol).or_default().insert((method, path)); + BTreeMap::<&str, BTreeSet<(&str, &str, Option<&str>, bool)>>::new(), + |mut grouped, (symbol, method, path, authority, in_process)| { + grouped + .entry(symbol) + .or_default() + .insert((method, path, authority, in_process)); grouped }, ); @@ -128,17 +133,19 @@ pub fn source_observations_to_graph( continue; }; if let Some(targets) = consumers_by_symbol.get(symbol) { - result.tests.extend(targets.iter().map(|(method, path)| { - source_test_case( - repo_id, - source_path, - content_hash, - observation, - symbol, - method, - path, - ) - })); + result + .tests + .extend(targets.iter().map(|(method, path, authority, in_process)| { + source_test_case( + repo_id, + source_path, + content_hash, + observation, + symbol, + (method, path), + call_scope(*authority, *in_process), + ) + })); } else { let (node, evidence) = standalone_test(repo_id, source_path, content_hash, observation, symbol); @@ -202,7 +209,11 @@ fn source_boundary( let role = match observation.role { SourceRole::Provider => BoundaryRole::Provider, SourceRole::Consumer => BoundaryRole::Consumer, - SourceRole::Test | SourceRole::Factory => { + SourceRole::Test + | SourceRole::Factory + | SourceRole::Mount + | SourceRole::Call + | SourceRole::Client => { unreachable!("test and factory observations are not HTTP boundaries") } }; @@ -210,7 +221,11 @@ fn source_boundary( BoundaryRole::Provider => "provider", BoundaryRole::Consumer => "consumer", }; - let stable_key = format!("http:{}:{role_key}:{method}:{path}", repo_id.as_str()); + let stable_key = format!( + "http:{}:{role_key}:{method}:{}", + repo_id.as_str(), + crate::canonical_route(path) + ); let evidence = source_evidence( repo_id, source_path, @@ -229,10 +244,24 @@ fn source_boundary( method: method.to_owned(), path: path.to_owned(), role, + scope: call_scope( + observation.authority.as_deref(), + observation.framework.is_in_process_client(), + ), evidence, } } +fn call_scope(authority: Option<&str>, in_process: bool) -> CallScope { + if in_process { + return CallScope::InProcess; + } + match authority { + Some(authority) => CallScope::Authority(authority.to_owned()), + None => CallScope::Workspace, + } +} + fn source_implementation( repo_id: &RepoId, source_path: &str, @@ -274,8 +303,8 @@ fn source_test_case( content_hash: &str, observation: &SourceObservation, symbol: &str, - method: &str, - path: &str, + (method, path): (&str, &str), + scope: CallScope, ) -> DeclaredTestCase { let stable_key = test_stable_key(repo_id, source_path, observation, symbol); DeclaredTestCase { @@ -292,6 +321,7 @@ fn source_test_case( }, method: method.to_owned(), path: path.to_owned(), + scope, evidence: source_evidence( repo_id, source_path, @@ -451,6 +481,23 @@ fn framework_name(framework: SourceFramework) -> &'static str { SourceFramework::SpringMvc => "spring-mvc", SourceFramework::WebClient => "webclient", SourceFramework::Feign => "feign", + SourceFramework::RestTemplate => "rest-template", + SourceFramework::TestClient => "test-client", + SourceFramework::FlaskTestClient => "flask-test-client", + SourceFramework::Jest => "jest", + SourceFramework::Vitest => "vitest", + SourceFramework::Mocha => "mocha", + SourceFramework::Playwright => "playwright", + SourceFramework::Supertest => "supertest", + SourceFramework::GoTest => "go-test", + SourceFramework::Httptest => "httptest", + SourceFramework::JUnit => "junit", + SourceFramework::MockMvc => "mockmvc", + SourceFramework::RestAssured => "rest-assured", + SourceFramework::WebTestClient => "webtestclient", + SourceFramework::TestRestTemplate => "test-rest-template", + SourceFramework::AxumOneshot => "axum-oneshot", + SourceFramework::ActixTest => "actix-test", } } @@ -460,17 +507,19 @@ fn role_name(role: SourceRole) -> &'static str { SourceRole::Consumer => "consumer", SourceRole::Test => "test", SourceRole::Factory => "factory", + SourceRole::Mount => "mount", + SourceRole::Call => "call", + SourceRole::Client => "client", } } #[cfg(test)] mod tests { + use code_system_graph_model::{EdgeKind, RepoId}; use super::source_observations_to_graph; - use crate::{ - link_declared_implementations, link_declared_tests, parse_python_source, parse_rust_source - }; + use crate::{AuthorityMap, link_http_routes, parse_python_source, parse_rust_source}; #[test] fn source_graph_should_link_python_test_to_rust_handler() { @@ -489,26 +538,32 @@ mod tests { &parse_python_source( r#"import requests def test_create_order(): - requests.post("https://api.test/orders") + requests.post("/orders") "#, ), ); - let test_edges = link_declared_tests(&python.tests, &rust.boundaries) - .unwrap_or_else(|error| panic!("fixtures must link: {error}")); - let implementation_edges = - link_declared_implementations(&rust.implementations, &rust.boundaries) - .unwrap_or_else(|error| panic!("fixtures must link: {error}")); + let edges = link_http_routes( + &rust.boundaries, + &python.tests, + &rust.implementations, + &AuthorityMap::new(), + ) + .edges; + let kinds = edges.iter().map(|edge| edge.kind).collect::>(); - assert!( - test_edges - .iter() - .all(|edge| edge.kind == EdgeKind::Validates) - ); - assert!( - implementation_edges - .iter() - .all(|edge| edge.kind == EdgeKind::ImplementedBy) + assert_eq!( + ( + kinds + .iter() + .filter(|kind| **kind == EdgeKind::Validates) + .count(), + kinds + .iter() + .filter(|kind| **kind == EdgeKind::ImplementedBy) + .count(), + ), + (1, 1) ); } diff --git a/crates/code-system-graph-core/src/source_http.rs b/crates/code-system-graph-core/src/source_http.rs index 8f769bc..b29efd8 100644 --- a/crates/code-system-graph-core/src/source_http.rs +++ b/crates/code-system-graph-core/src/source_http.rs @@ -10,9 +10,23 @@ use std::collections::{BTreeMap, BTreeSet}; use serde::{Deserialize, Serialize}; +use crate::router_mounts::mount_observation; use crate::{ExtractionLimitExceeded, ExtractionTracker}; +mod brace_clients; +mod brace_flows; +mod brace_lexer; +mod brace_tests; +mod python_flows; mod rust_bindings; +mod rust_flows; +mod rust_routers; +mod rust_test_clients; + +pub(crate) use brace_clients::{BraceClients, collect_brace_clients, java_method_lines}; +use python_flows::{PythonScopes, keyword_argument, record_fixture_requests}; +use rust_flows::RustScopes; +use rust_routers::{RustRouters, join_prefixes}; const HTTP_METHODS: [&str; 8] = [ "DELETE", "GET", "HEAD", "OPTIONS", "PATCH", "POST", "PUT", "TRACE", @@ -98,6 +112,61 @@ pub enum SourceFramework { WebClient, /// Feign client declarations. Feign, + /// Spring `RestTemplate` calls. + RestTemplate, + /// Starlette and `FastAPI` `TestClient`, and HTTPX clients bound to an ASGI or WSGI app. + TestClient, + /// Flask `app.test_client()` requests. + FlaskTestClient, + /// Jest tests. + Jest, + /// Vitest tests. + Vitest, + /// Mocha tests. + Mocha, + /// Playwright tests and `APIRequestContext` requests. + Playwright, + /// Supertest `request(app)` requests. + Supertest, + /// Go `testing` test functions. + GoTest, + /// Go `net/http/httptest` requests and servers. + Httptest, + /// `JUnit` `@Test` methods. + JUnit, + /// Spring `MockMvc` requests. + MockMvc, + /// REST Assured requests. + RestAssured, + /// Spring `WebTestClient` requests. + WebTestClient, + /// Spring Boot `TestRestTemplate` requests. + TestRestTemplate, + /// Axum `Router` requests sent with `tower::ServiceExt::oneshot`. + AxumOneshot, + /// Actix Web `test::TestRequest` requests. + ActixTest, +} + +impl SourceFramework { + /// Whether requests of this client run in-process against the application of their own + /// repository. + #[must_use] + pub fn is_in_process_client(self) -> bool { + matches!( + self, + Self::TestClient + | Self::FlaskTestClient + | Self::Supertest + | Self::Httptest + | Self::MockMvc + | Self::RestAssured + | Self::WebTestClient + | Self::TestRestTemplate + | Self::AxumOneshot + | Self::ActixTest + ) + } } /// Repository-boundary role represented by an observation. @@ -112,6 +181,86 @@ pub enum SourceRole { Test, /// A data factory with one statically declared model target. Factory, + /// A router mounted under a path prefix on another router or on the application root. + Mount, + /// A call from the enclosing function to another function of the repository. + Call, + /// A function, such as a pytest fixture, that returns an in-process test client. + Client, +} + +/// Source-level reference to a router or function, resolved per repository. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] +#[serde(rename_all = "snake_case")] +pub enum SymbolRef { + /// A router bound to a name in the declaring file; `default` names a default export. + Local(String), + /// A router imported from another module, with the module specifier as written. + Import { + /// Module specifier, such as `./routes/users` or `app.api.users`. + module: String, + /// Imported name; `default` for default exports. + name: String, + }, + /// The router built, returned, or configured by a function defined in the declaring file. + Function(String), + /// A function referenced by its path as written, such as `users::routes` or `routes.Register`. + Call(String), + /// A pytest fixture injected by parameter name, declared in the same file or in the nearest + /// enclosing `conftest.py`. + Fixture(String), + /// A parameter of the enclosing function, by name and zero-based position after receivers. + Parameter { + /// Declared parameter name. + name: String, + /// Position among the parameters a caller passes. + index: usize, + }, +} + +/// Client URL assembled from literal text and runtime values. +#[derive(Debug, Clone, Default, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] +pub struct UrlTemplate { + /// Parts in source order; adjacent text parts are merged. + pub parts: Vec, +} + +/// One part of a [`UrlTemplate`]. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] +#[serde(rename_all = "snake_case")] +pub enum UrlPart { + /// Literal text. + Text(String), + /// A parameter of the enclosing function, by name and zero-based position after receivers. + Parameter { + /// Declared parameter name. + name: String, + /// Position among the parameters a caller passes. + index: usize, + }, + /// Any other runtime value, named when it is a plain identifier. + Value(Option), +} + +/// Call recorded for repository-level client-wrapper and test-helper correlation. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] +pub struct CallSite { + /// Called function. + pub callee: SymbolRef, + /// Arguments in source order, up to the last string expression. + #[serde(default, skip_serializing_if = "Vec::is_empty")] + pub arguments: Vec, +} + +/// One argument of a [`CallSite`]. +#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)] +pub struct CallArgument { + /// Keyword under which the argument is passed. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub keyword: Option, + /// String value; `None` when the argument is not a string expression. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub value: Option, } /// Epistemic state of a source observation. @@ -145,8 +294,6 @@ pub enum SourceWarning { DynamicPath, /// A literal URL is not an absolute URL or root-relative path. UnsupportedLiteralPath, - /// An absolute consumer URL names an authority with no explicit workspace mapping. - UnmappedAuthority, /// The framework declaration did not expose an implementation symbol. MissingSymbol, /// Tree-sitter recovered from at least one syntax error in the source artifact. @@ -179,6 +326,22 @@ pub struct SourceObservation { pub related_symbol: Option, /// Repository-relative module path that declares `related_symbol`, when statically imported. pub related_path: Option, + /// Lower-case `host[:port]` named by an absolute consumer URL. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub authority: Option, + /// Router a provider route is registered on, the router mounted by a mount observation, or + /// the parameter through which a consumer receives its client. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub router: Option, + /// Router receiving a mount observation; `None` mounts at the application root. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub mount_parent: Option, + /// Consumer URL that depends on parameters of the enclosing function. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub url: Option, + /// Called function and arguments of a call observation. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub call: Option, /// Inclusive source range containing the direct evidence. pub lines: SourceLineRange, /// Epistemic state of this observation. @@ -237,6 +400,12 @@ impl<'a> SourceObservationCollector<'a> { pub(crate) fn into_unbounded(self) -> Vec { self.observations } + + pub(crate) fn has_role(&self, role: SourceRole) -> bool { + self.observations + .iter() + .any(|observation| observation.role == role) + } } pub(crate) fn charge_source_observation_values( @@ -261,6 +430,22 @@ pub(crate) fn charge_source_observation_values( if let Some(path) = &observation.related_path { tracker.charge_portable_path(path)?; } + let call_values = observation + .call + .iter() + .flat_map(|call| &call.arguments) + .filter_map(|argument| argument.value.as_ref()); + for template in observation.url.iter().chain(call_values) { + for part in &template.parts { + match part { + UrlPart::Text(text) => tracker.charge_string(text)?, + UrlPart::Parameter { name, .. } | UrlPart::Value(Some(name)) => { + tracker.charge_identifier(name)?; + } + UrlPart::Value(None) => {} + } + } + } Ok(()) } @@ -281,16 +466,31 @@ fn collect_rust_source<'a>( ) -> SourceObservationCollector<'a> { let tokens = lex_rust(source); let functions = rust_functions(&tokens); + let routers = RustRouters::new(&tokens, &functions); parse_rust_attributes(&tokens, &mut observations); if has_ident(&tokens, "axum") { - parse_axum_routes(&tokens, &mut observations); + parse_axum_routes(&tokens, &routers, &mut observations); + for mount in routers.mounts(false) { + observations.push(mount); + } } if has_ident(&tokens, "actix_web") { - parse_actix_builder_routes(&tokens, &mut observations); - } - if has_ident(&tokens, "reqwest") { - parse_reqwest_calls(&tokens, &functions, &mut observations); + parse_actix_builder_routes(&tokens, &routers, &mut observations); + for mount in routers.mounts(true) { + observations.push(mount); + } } + let scopes = RustScopes::new(&tokens, &functions); + rust_test_clients::parse_rust_test_requests(&tokens, &functions, &scopes, &mut observations); + let clients = if has_ident(&tokens, "reqwest") { + let receivers = rust_bindings::reqwest_receiver_tokens(&tokens, &functions); + parse_reqwest_calls(&tokens, &functions, &receivers, &scopes, &mut observations); + receivers + } else { + BTreeSet::new() + }; + let declares_tests = observations.has_role(SourceRole::Test) || has_ident(&tokens, "test"); + scopes.record_calls(&clients, declares_tests, &mut observations); observations } @@ -328,11 +528,39 @@ fn collect_python_source<'a>( let tokens = lex_python(source); let functions = python_functions(source, &tokens); let contexts = PythonContexts::discover(&tokens); + let scopes = PythonScopes::new(&tokens, &functions); parse_python_routes(&tokens, &contexts, &mut observations); + parse_python_router_mounts(&tokens, &contexts, &mut observations); parse_python_http_registries(&tokens, &mut observations); - parse_python_http_calls(&tokens, &functions, &contexts, &mut observations); + parse_python_route_calls(&tokens, &contexts, &mut observations); + parse_python_http_calls(&tokens, &functions, &contexts, &scopes, &mut observations); + parse_python_client_fixtures(&tokens, &functions, &contexts, &mut observations); parse_python_factories(source, &tokens, &mut observations); parse_python_tests(source, &tokens, &mut observations); + let clients = contexts + .request_modules + .iter() + .chain(&contexts.request_clients) + .chain(&contexts.httpx_modules) + .chain(&contexts.httpx_clients) + .chain(&contexts.aiohttp_modules) + .chain(&contexts.aiohttp_clients) + .chain(contexts.direct_calls.keys()) + .chain(contexts.test_clients.keys()) + .map(String::as_str) + .collect::>(); + let declares_tests = has_ident(&tokens, "pytest") + || has_ident(&tokens, "unittest") + || functions + .iter() + .any(|function| function.name.starts_with("test_")); + scopes.record_calls( + &python_router_imports(&tokens), + &clients, + declares_tests, + &mut observations, + ); + record_fixture_requests(&scopes, &mut observations); observations } @@ -385,6 +613,8 @@ pub fn normalize_source_http_path(path: &str) -> String { enum TokenKind { Ident(String), Literal(Option), + /// Raw body of a JavaScript template literal, including `${...}` substitutions. + Template(String), Punct(char), } @@ -743,6 +973,41 @@ fn rust_functions(tokens: &[Token]) -> Vec { functions } +/// Position of the smallest function span containing token `index`. +fn innermost_function(functions: &[FunctionSpan], index: usize) -> Option { + functions + .iter() + .enumerate() + .filter(|(_, function)| function.start_token <= index && index <= function.end_token) + .min_by_key(|(_, function)| function.end_token - function.start_token) + .map(|(position, _)| position) +} + +/// Splits `start..end` at top-level `separator` punctuation. +fn split_operands( + tokens: &[Token], + start: usize, + end: usize, + separator: char, +) -> Vec<(usize, usize)> { + let mut operands = Vec::new(); + let mut depth = 0_i32; + let mut operand_start = start; + for (index, token) in tokens.iter().enumerate().take(end).skip(start) { + match token.kind { + TokenKind::Punct('(' | '[' | '{') => depth += 1, + TokenKind::Punct(')' | ']' | '}') => depth -= 1, + TokenKind::Punct(character) if character == separator && depth == 0 => { + operands.push((operand_start, index)); + operand_start = index + 1; + } + _ => {} + } + } + operands.push((operand_start, end)); + operands +} + fn enclosing_symbol(functions: &[FunctionSpan], token_index: usize) -> Option { functions .iter() @@ -771,6 +1036,11 @@ fn parse_rust_attributes(tokens: &[Token], observations: &mut SourceObservationC [first, second] if first == "tokio" && second == "test" => { Some(SourceFramework::TokioTest) } + [first, second] + if (first == "actix_web" || first == "actix_rt") && second == "test" => + { + Some(SourceFramework::RustTest) + } [name] if name == "rstest" => Some(SourceFramework::Rstest), _ => None, }; @@ -988,7 +1258,7 @@ fn parse_actix_attribute( return; } for method in methods { - observations.push(http_from_literal( + let mut observation = http_from_literal( SourceLanguage::Rust, SourceFramework::ActixWeb, SourceRole::Provider, @@ -1000,11 +1270,17 @@ fn parse_actix_attribute( end: tokens[signature_end].end_line, }, true, - )); + ); + observation.router = Some(SymbolRef::Function(symbol.clone())); + observations.push(observation); } } -fn parse_axum_routes(tokens: &[Token], observations: &mut SourceObservationCollector<'_>) { +fn parse_axum_routes( + tokens: &[Token], + routers: &RustRouters<'_>, + observations: &mut SourceObservationCollector<'_>, +) { for (index, token) in tokens.iter().enumerate() { if !token.is_ident("route") || !tokens @@ -1024,9 +1300,10 @@ fn parse_axum_routes(tokens: &[Token], observations: &mut SourceObservationColle continue; } let literal = tokens.get(arguments[0]).and_then(Token::literal); + let router = routers.chain_owner(index - 1).router; let methods = method_calls(tokens, arguments[1], close); for (method, handler) in methods { - observations.push(http_from_literal( + let mut observation = http_from_literal( SourceLanguage::Rust, SourceFramework::Axum, SourceRole::Provider, @@ -1038,12 +1315,18 @@ fn parse_axum_routes(tokens: &[Token], observations: &mut SourceObservationColle end: tokens[close].end_line, }, true, - )); + ); + observation.router = Some(router.clone()); + observations.push(observation); } } } -fn parse_actix_builder_routes(tokens: &[Token], observations: &mut SourceObservationCollector<'_>) { +fn parse_actix_builder_routes( + tokens: &[Token], + routers: &RustRouters<'_>, + observations: &mut SourceObservationCollector<'_>, +) { for (index, token) in tokens.iter().enumerate() { if !token.is_ident("route") || !tokens @@ -1059,21 +1342,29 @@ fn parse_actix_builder_routes(tokens: &[Token], observations: &mut SourceObserva if arguments.len() < 2 || !contains_ident(tokens, arguments[1], close, "web") { continue; } + let owner = + (index > 0 && tokens[index - 1].is_punct('.')).then(|| routers.chain_owner(index - 1)); let literal = tokens.get(arguments[0]).and_then(Token::literal); + let literal = match owner.as_ref().and_then(|owner| owner.prefix.as_deref()) { + Some(prefix) => join_prefixes(Some(prefix), literal), + None => literal.map(str::to_owned), + }; for (method, handler) in method_calls(tokens, arguments[1], close) { - observations.push(http_from_literal( + let mut observation = http_from_literal( SourceLanguage::Rust, SourceFramework::ActixWeb, SourceRole::Provider, Some(method), - literal, + literal.as_deref(), handler, SourceLineRange { start: token.line, end: tokens[close].end_line, }, true, - )); + ); + observation.router = owner.as_ref().map(|owner| owner.router.clone()); + observations.push(observation); } } for (index, token) in tokens.iter().enumerate() { @@ -1097,8 +1388,12 @@ fn parse_actix_builder_routes(tokens: &[Token], observations: &mut SourceObserva continue; } let literal = tokens.get(index + 2).and_then(Token::literal); + let router = tokens + .get(resource_close + 1) + .is_some_and(|token| token.is_punct('.')) + .then(|| routers.chain_owner(resource_close + 1).router); for (method, handler) in method_calls(tokens, resource_close + 1, end) { - observations.push(http_from_literal( + let mut observation = http_from_literal( SourceLanguage::Rust, SourceFramework::ActixWeb, SourceRole::Provider, @@ -1110,7 +1405,9 @@ fn parse_actix_builder_routes(tokens: &[Token], observations: &mut SourceObserva end: tokens[end].end_line, }, true, - )); + ); + observation.router.clone_from(&router); + observations.push(observation); } } } @@ -1118,45 +1415,12 @@ fn parse_actix_builder_routes(tokens: &[Token], observations: &mut SourceObserva fn parse_reqwest_calls( tokens: &[Token], functions: &[FunctionSpan], + reqwest_receivers: &BTreeSet, + scopes: &RustScopes<'_>, observations: &mut SourceObservationCollector<'_>, ) { - let reqwest_receivers = rust_bindings::reqwest_receiver_tokens(tokens, functions); for index in 0..tokens.len() { - let explicit = tokens[index].is_ident("reqwest") - && tokens - .get(index + 1) - .is_some_and(|token| token.is_punct(':')) - && tokens - .get(index + 2) - .is_some_and(|token| token.is_punct(':')); - let (receiver, method_index) = if explicit { - let direct = index + 3; - let method_index = if tokens - .get(direct) - .is_some_and(|token| token.is_ident("blocking")) - && tokens - .get(direct + 1) - .is_some_and(|token| token.is_punct(':')) - && tokens - .get(direct + 2) - .is_some_and(|token| token.is_punct(':')) - { - direct + 3 - } else { - direct - }; - ("reqwest", method_index) - } else if let Some(receiver) = tokens[index].ident() { - if reqwest_receivers.contains(&index) - && tokens - .get(index + 1) - .is_some_and(|token| token.is_punct('.')) - { - (receiver, index + 2) - } else { - continue; - } - } else { + let Some(method_index) = reqwest_method_index(tokens, index, reqwest_receivers) else { continue; }; let Some(method_name) = tokens.get(method_index).and_then(Token::ident) else { @@ -1191,15 +1455,20 @@ fn parse_reqwest_calls( if method.is_none() && method_name != "request" { continue; } - let literal = path_argument - .and_then(|argument| tokens.get(argument)) - .and_then(Token::literal); + let template = path_argument.map(|argument| { + let end = arguments + .iter() + .find(|start| **start > argument) + .map_or(close, |next| next - 1); + scopes.template(argument, end, index) + }); + let literal = template.as_ref().and_then(UrlTemplate::client_literal); let mut observation = http_from_literal( SourceLanguage::Rust, SourceFramework::Reqwest, SourceRole::Consumer, method, - literal, + literal.as_deref(), enclosing_symbol(functions, index), SourceLineRange { start: tokens[index].line, @@ -1214,11 +1483,37 @@ fn parse_reqwest_calls( observation.warnings.sort(); observation.warnings.dedup(); } - let _ = receiver; + observation.url = template.filter(UrlTemplate::has_parameters); observations.push(observation); } } +/// Index of the method name called at token `index` on the `reqwest` or `reqwest::blocking` +/// module, or on a Reqwest client receiver. +fn reqwest_method_index( + tokens: &[Token], + index: usize, + reqwest_receivers: &BTreeSet, +) -> Option { + let path_separator = |at: usize| { + tokens.get(at).is_some_and(|token| token.is_punct(':')) + && tokens.get(at + 1).is_some_and(|token| token.is_punct(':')) + }; + if tokens[index].is_ident("reqwest") && path_separator(index + 1) { + let direct = index + 3; + let blocking = tokens + .get(direct) + .is_some_and(|token| token.is_ident("blocking")) + && path_separator(direct + 1); + return Some(if blocking { direct + 3 } else { direct }); + } + (reqwest_receivers.contains(&index) + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('.'))) + .then_some(index + 2) +} + fn rust_method_expression(tokens: &[Token], start: usize, end: usize) -> Option { if let Some(literal) = tokens.get(start).and_then(Token::literal) { return canonical_method(literal).map(str::to_owned); @@ -1341,6 +1636,64 @@ fn canonical_method(name: &str) -> Option<&'static str> { .find(|method| method.eq_ignore_ascii_case(name)) } +/// Call from function `function` recorded for repository-level flow composition. +pub(crate) fn call_observation( + language: SourceLanguage, + callee: SymbolRef, + arguments: Vec, + function: String, + lines: SourceLineRange, +) -> SourceObservation { + let framework = match language { + SourceLanguage::Python => SourceFramework::PythonTest, + SourceLanguage::Rust => SourceFramework::RustTest, + SourceLanguage::TypeScript | SourceLanguage::JavaScript => SourceFramework::Fetch, + SourceLanguage::Go => SourceFramework::GoNetHttp, + SourceLanguage::Java => SourceFramework::SpringMvc, + }; + SourceObservation { + language, + framework, + role: SourceRole::Call, + method: None, + path: None, + symbol_name: Some(function), + related_symbol: None, + related_path: None, + authority: None, + router: None, + mount_parent: None, + url: None, + call: Some(CallSite { callee, arguments }), + lines, + status: SourceEpistemicStatus::Confirmed, + confidence: 1.0, + warnings: Vec::new(), + } +} + +/// Confirmed consumer of `template` issued from `caller`, with the client coordinates of +/// `origin`; `None` when the template does not yield an exact method and path. +pub(crate) fn instantiated_consumer( + origin: &SourceObservation, + template: &UrlTemplate, + caller: &str, + lines: SourceLineRange, +) -> Option { + let literal = template.client_literal()?; + let observation = http_from_literal( + origin.language, + origin.framework, + SourceRole::Consumer, + origin.method.clone(), + Some(&literal), + Some(caller.to_owned()), + lines, + false, + ); + (observation.status == SourceEpistemicStatus::Confirmed).then_some(observation) +} + fn confirmed_test( language: SourceLanguage, framework: SourceFramework, @@ -1357,6 +1710,11 @@ fn confirmed_test( symbol_name: Some(name), related_symbol: None, related_path: None, + authority: None, + router: None, + mount_parent: None, + url: None, + call: None, lines: SourceLineRange { start, end }, status: SourceEpistemicStatus::Confirmed, confidence: 1.0, @@ -1385,16 +1743,14 @@ fn http_from_literal( client_literal_identity(value) } }); - let warning = if authority.is_some() { - Some(SourceWarning::UnmappedAuthority) - } else if literal.is_none() { + let warning = if literal.is_none() { Some(SourceWarning::DynamicPath) } else if path.is_none() { Some(SourceWarning::UnsupportedLiteralPath) } else { None }; - let status = if path.is_some() && method.is_some() && authority.is_none() { + let status = if path.is_some() && method.is_some() { SourceEpistemicStatus::Confirmed } else { SourceEpistemicStatus::Ambiguous @@ -1408,6 +1764,11 @@ fn http_from_literal( symbol_name, related_symbol: None, related_path: None, + authority, + router: None, + mount_parent: None, + url: None, + call: None, lines, status, confidence: if status == SourceEpistemicStatus::Confirmed { @@ -1448,6 +1809,11 @@ fn inexact_http( symbol_name, related_symbol: None, related_path: None, + authority: None, + router: None, + mount_parent: None, + url: None, + call: None, lines, status: SourceEpistemicStatus::Incomplete, confidence: 0.0, @@ -1479,13 +1845,12 @@ fn client_literal_identity(value: &str) -> (Option, Option) { return (None, None); } let separator = after_authority.find('/'); - let authority = separator.map_or(after_authority, |index| &after_authority[..index]); let path = separator.map_or("/", |index| &after_authority[index..]); ( Some(normalize_source_http_path( path.split(['?', '#']).next().unwrap_or(path), )), - Some(authority.to_ascii_lowercase()), + crate::routes::url_authority(value), ) } @@ -1502,6 +1867,7 @@ struct PythonContexts { aiohttp_clients: BTreeSet, direct_calls: BTreeMap, string_constants: BTreeMap, + test_clients: BTreeMap, } impl PythonContexts { @@ -1545,27 +1911,31 @@ impl PythonContexts { { contexts.string_constants.insert(name.to_owned(), value); } + if let Some(framework) = python_test_client(tokens, index + 2, end) { + contexts.test_clients.insert(name.to_owned(), framework); + continue; + } if contains_ident(tokens, index + 2, end, "FastAPI") || contains_ident(tokens, index + 2, end, "APIRouter") { contexts.fastapi_apps.insert(name.to_owned()); if contains_ident(tokens, index + 2, end, "APIRouter") - && let Some(open) = - (index + 2..end).find(|candidate| tokens[*candidate].is_punct('(')) - && let Some(close) = matching(tokens, open, '(', ')') - && let Some(prefix) = parse_keyword_string_values(tokens, open, close, "prefix") - .into_iter() - .next() + && let Some(prefix) = + python_constructor_prefix(tokens, index + 2, end, "prefix") { - contexts - .route_prefixes - .insert(name.to_owned(), normalize_source_http_path(&prefix)); + contexts.route_prefixes.insert(name.to_owned(), prefix); } } if contains_ident(tokens, index + 2, end, "Flask") || contains_ident(tokens, index + 2, end, "Blueprint") { contexts.flask_apps.insert(name.to_owned()); + if contains_ident(tokens, index + 2, end, "Blueprint") + && let Some(prefix) = + python_constructor_prefix(tokens, index + 2, end, "url_prefix") + { + contexts.route_prefixes.insert(name.to_owned(), prefix); + } } if contains_ident(tokens, index + 2, end, "Session") && (contains_any(tokens, index + 2, end, &contexts.request_modules) @@ -1587,27 +1957,141 @@ impl PythonContexts { contexts.aiohttp_clients.insert(name.to_owned()); } } - if aiohttp_client_imported { - for index in 0..tokens.len() { - if !tokens[index].is_ident("as") { - continue; - } - let Some(name) = tokens.get(index + 1).and_then(Token::ident) else { - continue; - }; - let start = index.saturating_sub(32); - if tokens[start..index] + contexts.discover_with_targets(tokens, aiohttp_client_imported); + contexts + } + + /// Clients bound by `with ... as NAME`. + fn discover_with_targets(&mut self, tokens: &[Token], aiohttp_client_imported: bool) { + for index in 0..tokens.len() { + if !tokens[index].is_ident("as") { + continue; + } + let Some(name) = tokens.get(index + 1).and_then(Token::ident) else { + continue; + }; + let start = index.saturating_sub(32); + let with = (start..index) + .rev() + .find(|candidate| tokens[*candidate].is_ident("with")) + .unwrap_or(start); + if let Some(framework) = python_test_client(tokens, with, index) { + self.test_clients.insert(name.to_owned(), framework); + } else if aiohttp_client_imported + && tokens[start..index] .iter() .any(|token| token.is_ident("ClientSession")) - { - contexts.aiohttp_clients.insert(name.to_owned()); - } + { + self.aiohttp_clients.insert(name.to_owned()); } } - contexts + } + + /// Client receivers and the framework of each. + fn http_receivers(&self) -> BTreeMap<&str, SourceFramework> { + self.request_modules + .iter() + .chain(&self.request_clients) + .map(|name| (name.as_str(), SourceFramework::Requests)) + .chain( + self.httpx_modules + .iter() + .chain(&self.httpx_clients) + .map(|name| (name.as_str(), SourceFramework::Httpx)), + ) + .chain( + self.aiohttp_clients + .iter() + .map(|name| (name.as_str(), SourceFramework::AioHttp)), + ) + .chain( + self.test_clients + .iter() + .map(|(name, framework)| (name.as_str(), *framework)), + ) + .collect() } } +/// In-process test client constructed by the expression in `start..end`. +fn python_test_client(tokens: &[Token], start: usize, end: usize) -> Option { + if contains_ident(tokens, start, end, "TestClient") { + return Some(SourceFramework::TestClient); + } + let flask = (start.max(1)..end).any(|index| { + tokens[index].is_ident("test_client") + && tokens[index - 1].is_punct('.') + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + }); + if flask { + return Some(SourceFramework::FlaskTestClient); + } + let httpx = contains_ident(tokens, start, end, "AsyncClient") + || contains_ident(tokens, start, end, "Client"); + let in_process = contains_ident(tokens, start, end, "ASGITransport") + || contains_ident(tokens, start, end, "WSGITransport") + || (start.max(1)..end).any(|index| { + tokens[index].is_ident("app") + && (tokens[index - 1].is_punct('(') || tokens[index - 1].is_punct(',')) + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('=')) + }); + (httpx && in_process).then_some(SourceFramework::TestClient) +} + +/// Functions, such as pytest fixtures, that return or yield an in-process test client. +fn parse_python_client_fixtures( + tokens: &[Token], + functions: &[FunctionSpan], + contexts: &PythonContexts, + observations: &mut SourceObservationCollector<'_>, +) { + for index in 0..tokens.len() { + if !tokens[index].is_ident("return") && !tokens[index].is_ident("yield") { + continue; + } + let Some(function) = innermost_function(functions, index) else { + continue; + }; + let end = python_expression_end(tokens, index + 1); + let named = tokens + .get(index + 1) + .and_then(Token::ident) + .filter(|_| end == index + 2) + .and_then(|name| contexts.test_clients.get(name).copied()); + let Some(framework) = named.or_else(|| python_test_client(tokens, index + 1, end)) else { + continue; + }; + let mut observation = confirmed_test( + SourceLanguage::Python, + framework, + functions[function].name.clone(), + tokens[index].line, + tokens[index].end_line, + ); + observation.role = SourceRole::Client; + observations.push(observation); + } +} + +/// Normalized `keyword=` string of the first call between `start` and `end`. +fn python_constructor_prefix( + tokens: &[Token], + start: usize, + end: usize, + keyword: &str, +) -> Option { + let open = (start..end).find(|candidate| tokens[*candidate].is_punct('('))?; + let close = matching(tokens, open, '(', ')')?; + let prefix = parse_keyword_string_values(tokens, open, close, keyword) + .into_iter() + .next()?; + Some(normalize_source_http_path(&prefix)) +} + fn python_imports_item(tokens: &[Token], module: &str, item: &str) -> bool { tokens.iter().enumerate().any(|(index, token)| { token.is_ident("from") @@ -1849,8 +2333,9 @@ fn parse_python_routes( }); let literal = prefixed_path.as_deref(); let methods = python_route_methods(tokens, open, close, decorator, framework); + let router = Some(SymbolRef::Local(receiver.to_owned())); if methods.is_empty() { - observations.push(inexact_http( + let mut observation = inexact_http( SourceLanguage::Python, framework, SourceRole::Provider, @@ -1862,10 +2347,12 @@ fn parse_python_routes( end: tokens[signature_end].end_line, }, SourceWarning::DynamicMethod, - )); + ); + observation.router = router; + observations.push(observation); } else { for method in methods { - observations.push(http_from_literal( + let mut observation = http_from_literal( SourceLanguage::Python, framework, SourceRole::Provider, @@ -1877,10 +2364,257 @@ fn parse_python_routes( end: tokens[signature_end].end_line, }, true, - )); + ); + observation.router.clone_from(&router); + observations.push(observation); + } + } + } +} + +/// Records `include_router` and `register_blueprint` mounts on known applications and routers. +fn parse_python_router_mounts( + tokens: &[Token], + contexts: &PythonContexts, + observations: &mut SourceObservationCollector<'_>, +) { + let imports = python_router_imports(tokens); + for index in 0..tokens.len().saturating_sub(3) { + let Some(receiver) = tokens[index].ident() else { + continue; + }; + let (framework, keyword) = match tokens[index + 2].ident() { + Some("include_router") if contexts.fastapi_apps.contains(receiver) => { + (SourceFramework::FastApi, "prefix") + } + Some("register_blueprint") if contexts.flask_apps.contains(receiver) => { + (SourceFramework::Flask, "url_prefix") + } + _ => continue, + }; + if !tokens[index + 1].is_punct('.') || !tokens[index + 3].is_punct('(') { + continue; + } + let open = index + 3; + let Some(close) = matching(tokens, open, '(', ')') else { + continue; + }; + let Some(&first) = top_level_arguments(tokens, open, close).first() else { + continue; + }; + let mut dotted = Vec::new(); + let mut cursor = first; + while let Some(name) = tokens.get(cursor).and_then(Token::ident) { + dotted.push(name); + if !tokens + .get(cursor + 1) + .is_some_and(|token| token.is_punct('.')) + { + break; + } + cursor += 2; + } + let Some(child) = python_router_reference(&dotted, contexts, &imports) else { + continue; + }; + let prefix = parse_keyword_string_values(tokens, open, close, keyword) + .into_iter() + .next(); + observations.push(mount_observation( + SourceLanguage::Python, + framework, + child, + Some(SymbolRef::Local(receiver.to_owned())), + prefix.as_deref(), + SourceLineRange { + start: tokens[index].line, + end: tokens[close].end_line, + }, + )); + } +} + +/// Routes registered by `FastAPI` `add_api_route(path, endpoint)` and Flask +/// `add_url_rule(rule, endpoint, view_func)` calls. +fn parse_python_route_calls( + tokens: &[Token], + contexts: &PythonContexts, + observations: &mut SourceObservationCollector<'_>, +) { + for index in 0..tokens.len().saturating_sub(3) { + let Some(receiver) = tokens[index].ident() else { + continue; + }; + if !tokens[index + 1].is_punct('.') || !tokens[index + 3].is_punct('(') { + continue; + } + let (framework, path_keyword, handler_keyword, handler_position) = + match tokens[index + 2].ident() { + Some("add_api_route") if contexts.fastapi_apps.contains(receiver) => { + (SourceFramework::FastApi, "path", "endpoint", 1) + } + Some("add_url_rule") if contexts.flask_apps.contains(receiver) => { + (SourceFramework::Flask, "rule", "view_func", 2) + } + _ => continue, + }; + let open = index + 3; + let Some(close) = matching(tokens, open, '(', ')') else { + continue; + }; + let arguments = top_level_arguments(tokens, open, close) + .into_iter() + .map(|start| { + let end = python_argument_end(tokens, start, close); + match keyword_argument(tokens, start, end) { + Some((keyword, value)) => (Some(keyword), value, end), + None => (None, start, end), + } + }) + .collect::>(); + let positional = arguments + .iter() + .filter(|(keyword, ..)| keyword.is_none()) + .collect::>(); + let path = arguments + .iter() + .find(|(keyword, ..)| keyword.as_deref() == Some(path_keyword)) + .or_else(|| positional.first().copied()) + .and_then(|(_, start, end)| { + python_static_string_expression(tokens, *start, *end, &contexts.string_constants) + }); + let handler = arguments + .iter() + .find(|(keyword, ..)| keyword.as_deref() == Some(handler_keyword)) + .or_else(|| positional.get(handler_position).copied()) + .and_then(|(_, start, end)| { + (*start..*end) + .map(|token| { + tokens[token] + .ident() + .or_else(|| tokens[token].is_punct('.').then_some(".")) + }) + .collect::>() + }); + let mut methods = + python_route_methods(tokens, open, close, "route", SourceFramework::Flask); + if methods.is_empty() { + methods.insert("GET".to_owned()); + } + for method in methods { + let mut observation = http_from_literal( + SourceLanguage::Python, + framework, + SourceRole::Provider, + Some(method), + path.as_deref(), + handler.clone(), + SourceLineRange { + start: tokens[index].line, + end: tokens[close].end_line, + }, + true, + ); + observation.router = Some(SymbolRef::Local(receiver.to_owned())); + observations.push(observation); + } + } +} + +fn python_router_reference( + dotted: &[&str], + contexts: &PythonContexts, + imports: &BTreeMap, +) -> Option { + let (head, rest) = dotted.split_first()?; + if let Some((module, name)) = imports.get(*head) { + let name = std::iter::once(name.as_str()) + .filter(|name| !name.is_empty()) + .chain(rest.iter().copied()) + .collect::>() + .join("."); + return (!name.is_empty()).then(|| SymbolRef::Import { + module: module.clone(), + name, + }); + } + (rest.is_empty() + && (contexts.fastapi_apps.contains(*head) || contexts.flask_apps.contains(*head))) + .then(|| SymbolRef::Local((*head).to_owned())) +} + +/// Maps local names to `(module, imported name)`; `import a.b as c` binds `c` to `(a.b, "")`. +fn python_router_imports(tokens: &[Token]) -> BTreeMap { + let mut imports = BTreeMap::new(); + for index in 0..tokens.len() { + if tokens[index].is_ident("import") + && (index == 0 || tokens[index - 1].line != tokens[index].line) + { + let mut cursor = index + 1; + let mut module = String::new(); + while let Some(token) = tokens + .get(cursor) + .filter(|token| token.line == tokens[index].line) + { + match (token.ident(), token.is_punct('.')) { + (Some(name), _) if name != "as" => module.push_str(name), + (None, true) => module.push('.'), + _ => break, + } + cursor += 1; + } + if tokens.get(cursor).is_some_and(|token| token.is_ident("as")) + && let Some(alias) = tokens.get(cursor + 1).and_then(Token::ident) + { + imports.insert(alias.to_owned(), (module, String::new())); } + continue; + } + if !tokens[index].is_ident("from") { + continue; + } + let Some(import_index) = (index + 1..tokens.len()) + .take_while(|candidate| tokens[*candidate].line == tokens[index].line) + .find(|candidate| tokens[*candidate].is_ident("import")) + else { + continue; + }; + let module = tokens[index + 1..import_index] + .iter() + .map(|token| token.ident().unwrap_or(".")) + .collect::(); + let parenthesized = tokens + .get(import_index + 1) + .is_some_and(|token| token.is_punct('(')); + let end = if parenthesized { + matching(tokens, import_index + 1, '(', ')').unwrap_or(import_index + 1) + } else { + (import_index + 1..tokens.len()) + .find(|candidate| tokens[*candidate].line != tokens[index].line) + .unwrap_or(tokens.len()) + }; + let mut cursor = import_index + 1; + while cursor < end { + let Some(imported) = tokens[cursor].ident() else { + cursor += 1; + continue; + }; + let aliased = tokens + .get(cursor + 1) + .is_some_and(|token| token.is_ident("as")); + let local = if aliased { + tokens + .get(cursor + 2) + .and_then(Token::ident) + .unwrap_or(imported) + } else { + imported + }; + imports.insert(local.to_owned(), (module.clone(), imported.to_owned())); + cursor += if aliased { 3 } else { 1 }; } } + imports } fn python_decorator_call(tokens: &[Token], at: usize) -> Option<(&str, &str, usize)> { @@ -2185,6 +2919,11 @@ fn parse_python_factories( symbol_name: Some(factory_name.to_owned()), related_symbol: Some(model_name.to_owned()), related_path: imported_paths.get(model_name).cloned(), + authority: None, + router: None, + mount_parent: None, + url: None, + call: None, lines: SourceLineRange { start: start_line, end: model_line, @@ -2248,44 +2987,34 @@ fn parse_python_http_calls( tokens: &[Token], functions: &[FunctionSpan], contexts: &PythonContexts, + scopes: &PythonScopes<'_>, observations: &mut SourceObservationCollector<'_>, ) { - let receivers = contexts - .request_modules - .iter() - .chain(&contexts.request_clients) - .map(|name| (name.as_str(), SourceFramework::Requests)) - .chain( - contexts - .httpx_modules - .iter() - .chain(&contexts.httpx_clients) - .map(|name| (name.as_str(), SourceFramework::Httpx)), - ) - .chain( - contexts - .aiohttp_clients - .iter() - .map(|name| (name.as_str(), SourceFramework::AioHttp)), - ) - .collect::>(); + let receivers = contexts.http_receivers(); for index in 0..tokens.len() { let Some(receiver) = tokens[index].ident() else { continue; }; - let (framework, call, open) = if let Some(framework) = receivers.get(receiver).copied() { - if !tokens + let member = || { + tokens .get(index + 1) .is_some_and(|token| token.is_punct('.')) - { - continue; - } - let Some(call) = tokens.get(index + 2).and_then(Token::ident) else { + .then(|| tokens.get(index + 2).and_then(Token::ident)) + .flatten() + }; + let mut parameter = None; + let (framework, call, open) = if let Some(framework) = receivers.get(receiver).copied() { + let Some(call) = member() else { continue; }; (framework, call, index + 3) } else if let Some((framework, call)) = contexts.direct_calls.get(receiver) { (*framework, call.as_str(), index + 1) + } else if let Some(call) = member().filter(|call| canonical_method(call).is_some()) + && let Some(received) = python_received_client(tokens, index, scopes) + { + parameter = Some(received); + (SourceFramework::TestClient, call, index + 3) } else { continue; }; @@ -2315,25 +3044,20 @@ fn parse_python_http_calls( if method.is_none() && call != "request" { continue; } - let resolved_path = path_argument.and_then(|argument| { - python_static_string_expression( - tokens, - argument, - python_argument_end(tokens, argument, close), - &contexts.string_constants, - ) - }); - let literal = resolved_path.as_deref().or_else(|| { - path_argument - .and_then(|argument| tokens.get(argument)) - .and_then(Token::literal) + let template = path_argument.map(|argument| { + let end = python_argument_end(tokens, argument, close); + let start = keyword_argument(tokens, argument, end) + .filter(|(keyword, _)| keyword == "url") + .map_or(argument, |(_, value)| value); + scopes.template(start, end, index) }); + let literal = template.as_ref().and_then(UrlTemplate::client_literal); let mut observation = http_from_literal( SourceLanguage::Python, framework, SourceRole::Consumer, method, - literal, + literal.as_deref(), enclosing_symbol(functions, index), SourceLineRange { start: tokens[index].line, @@ -2341,6 +3065,15 @@ fn parse_python_http_calls( }, false, ); + observation.url = template.filter(UrlTemplate::has_parameters); + if let Some(received) = parameter { + if !literal.as_deref().is_some_and(|path| path.starts_with('/')) { + continue; + } + observation.router = Some(received); + observation.status = SourceEpistemicStatus::Ambiguous; + observation.confidence = 0.0; + } if call == "request" && observation.method.is_none() { observation.status = SourceEpistemicStatus::Incomplete; observation.confidence = 0.0; @@ -2352,6 +3085,27 @@ fn parse_python_http_calls( } } +/// Parameter of the enclosing function through which the call at `index` receives its client. +fn python_received_client( + tokens: &[Token], + index: usize, + scopes: &PythonScopes<'_>, +) -> Option { + if index > 0 && tokens[index - 1].is_punct('.') { + return None; + } + let name = tokens[index].ident()?; + let function = scopes.innermost(index)?; + let position = scopes + .parameters_of(function) + .iter() + .position(|parameter| parameter == name)?; + Some(SymbolRef::Parameter { + name: name.to_owned(), + index: position, + }) +} + fn parse_python_tests( source: &str, tokens: &[Token], @@ -2496,14 +3250,81 @@ fn compare_observations(left: &SourceObservation, right: &SourceObservation) -> right.status, &right.warnings, )) + .then_with(|| { + ( + &left.authority, + &left.router, + &left.mount_parent, + &left.url, + &left.call, + ) + .cmp(&( + &right.authority, + &right.router, + &right.mount_parent, + &right.url, + &right.call, + )) + }) } #[cfg(test)] mod tests { use super::{ - SourceEpistemicStatus, SourceFramework, SourceRole, SourceWarning, parse_python_source, parse_rust_source + SourceEpistemicStatus, SourceFramework, SourceObservation, SourceRole, SourceWarning, SymbolRef, UrlPart, parse_python_source }; + /// HTTP, route, and test facts without call observations. + fn parse_rust_source(source: &str) -> Vec { + super::parse_rust_source(source) + .into_iter() + .filter(|item| item.role != SourceRole::Call) + .collect() + } + + #[test] + fn rust_reqwest_urls_should_resolve_constants_format_strings_and_wrapper_parameters() { + let source = r#" +use reqwest::Client; +const BASE: &str = "http://orders:8080"; +async fn get_order(client: &Client, id: u64) { + client.get(format!("{BASE}/orders/{id}")).send().await; +} +async fn post_json(client: &Client, path: &str) { + client.post(format!("{}{}", BASE, path)).send().await; +} +async fn refund(client: &Client, id: &str) { + post_json(client, &format!("/orders/{id}/refunds")).await; +} +"#; + let result = super::parse_rust_source(source); + let consumers = result + .iter() + .filter(|item| item.role == SourceRole::Consumer) + .map(|item| { + ( + item.symbol_name.as_deref(), + item.path.as_deref(), + item.url.is_some(), + ) + }) + .collect::>(); + let calls = result + .iter() + .filter_map(|item| item.call.as_ref()) + .map(|call| (call.callee.clone(), call.arguments.len())) + .collect::>(); + + assert_eq!( + consumers, + [ + (Some("get_order"), Some("/orders/{id}"), true), + (Some("post_json"), None, true), + ] + ); + assert_eq!(calls, [(SymbolRef::Call("post_json".to_owned()), 2)]); + } + #[test] fn axum_should_extract_multiline_route_and_handler() { let source = r#" @@ -2572,6 +3393,7 @@ fn app() { assert_eq!( result .iter() + .filter(|item| item.role == SourceRole::Provider) .map(|item| ( item.framework, item.method.as_deref(), @@ -2702,11 +3524,11 @@ async fn synchronize<'request>(client: &'request Client) { } #[test] - fn absolute_consumer_url_should_require_an_explicit_authority_mapping() { + fn absolute_consumer_url_should_record_its_authority() { let source = r#" use reqwest::Client; async fn synchronize(client: &Client) { - client.get("https://third-party.example/health").send().await; + client.get("https://User@Third-Party.example:8443/health?probe=1").send().await; client.get("/internal-health").send().await; } "#; @@ -2720,9 +3542,14 @@ async fn synchronize(client: &Client) { .find(|item| item.path.as_deref() == Some("/internal-health")) .expect("relative call"); - assert_eq!(external.status, SourceEpistemicStatus::Ambiguous); - assert_eq!(external.warnings, [SourceWarning::UnmappedAuthority]); + assert_eq!(external.status, SourceEpistemicStatus::Confirmed); + assert_eq!( + external.authority.as_deref(), + Some("third-party.example:8443") + ); + assert_eq!(external.warnings, []); assert_eq!(internal.status, SourceEpistemicStatus::Confirmed); + assert_eq!(internal.authority, None); } #[test] @@ -3407,22 +4234,324 @@ def send(method, base, path): )); } + fn consumer_paths(source: &str) -> Vec<(Option, Option, bool)> { + parse_python_source(source) + .into_iter() + .filter(|item| item.role == SourceRole::Consumer) + .map(|item| (item.symbol_name, item.path, item.url.is_some())) + .collect() + } + #[test] - fn python_f_strings_should_be_dynamic_not_exact() { + fn python_url_expressions_should_resolve_whole_segment_values_and_scoped_names() { let source = r#" import requests +BASE_URL = "http://orders:8080" +API = BASE_URL + "/api" def fetch(user_id): requests.get(f"https://example.test/users/{user_id}") +def report(name): + requests.get(f"/reports/{name}.csv") +def scoped(order): + url = f"{API}/orders/{order.id}" + requests.get(url) +def formatted(sku): + requests.get("{}/items/{}".format(API, sku)) + requests.delete("/items/%s" % sku) +"#; + + assert_eq!( + consumer_paths(source), + [ + ( + Some("fetch".to_owned()), + Some("/users/{user_id}".to_owned()), + true + ), + (Some("report".to_owned()), None, true), + ( + Some("scoped".to_owned()), + Some("/api/orders/{id}".to_owned()), + false + ), + ( + Some("formatted".to_owned()), + Some("/api/items/{sku}".to_owned()), + true + ), + ( + Some("formatted".to_owned()), + Some("/items/{sku}".to_owned()), + true + ), + ] + ); + } + + #[test] + fn python_test_clients_should_be_recognized_and_parameter_clients_left_for_composition() { + let source = r#" +import httpx +import pytest +from fastapi.testclient import TestClient +from app import app, flask_app + +client = TestClient(app) + +def test_direct(): + client.get("/items") + +class OrdersTest(unittest.TestCase): + def setUp(self): + self.web = flask_app.test_client() + + def test_flask(self): + self.web.post("/orders") + +async def test_async(): + async with httpx.AsyncClient(app=app, base_url="http://test") as api: + await api.delete("/orders/1") + +@pytest.fixture +def api_client(): + return TestClient(app) + +def test_injected(api_client): + api_client.put("/orders/2") "#; let result = parse_python_source(source); + let consumers = result + .iter() + .filter(|item| item.role == SourceRole::Consumer) + .map(|item| { + ( + item.framework, + item.method.clone().unwrap_or_default(), + item.path.clone().unwrap_or_default(), + item.status, + item.router.clone(), + ) + }) + .collect::>(); + let clients = result + .iter() + .filter(|item| item.role == SourceRole::Client) + .map(|item| (item.framework, item.symbol_name.clone().unwrap_or_default())) + .collect::>(); - assert!(matches!( - result.as_slice(), - [item] - if item.path.is_none() - && item.status == SourceEpistemicStatus::Ambiguous - && item.warnings == [SourceWarning::DynamicPath] - )); + assert_eq!( + consumers, + [ + ( + SourceFramework::TestClient, + "GET".to_owned(), + "/items".to_owned(), + SourceEpistemicStatus::Confirmed, + None + ), + ( + SourceFramework::FlaskTestClient, + "POST".to_owned(), + "/orders".to_owned(), + SourceEpistemicStatus::Confirmed, + None + ), + ( + SourceFramework::TestClient, + "DELETE".to_owned(), + "/orders/1".to_owned(), + SourceEpistemicStatus::Confirmed, + None + ), + ( + SourceFramework::TestClient, + "PUT".to_owned(), + "/orders/2".to_owned(), + SourceEpistemicStatus::Ambiguous, + Some(SymbolRef::Parameter { + name: "api_client".to_owned(), + index: 0 + }) + ), + ] + ); + assert_eq!( + clients, + [(SourceFramework::TestClient, "api_client".to_owned())] + ); + } + + #[test] + fn python_imperative_routes_should_name_their_handlers() { + let source = r#" +from fastapi import APIRouter +from flask import Flask +from orders import views + +router = APIRouter() +router.add_api_route("/orders", views.list_orders, methods=["GET", "POST"]) +router.add_api_route(path="/health", endpoint=health) +app = Flask(__name__) +app.add_url_rule("/legacy/", "legacy", views.legacy) +"#; + let providers = parse_python_source(source) + .into_iter() + .filter(|item| item.role == SourceRole::Provider) + .map(|item| { + format!( + "{} {} -> {}", + item.method.unwrap_or_default(), + item.path.unwrap_or_default(), + item.symbol_name.unwrap_or_default() + ) + }) + .collect::>(); + + assert_eq!( + providers, + [ + "GET /orders -> views.list_orders", + "POST /orders -> views.list_orders", + "GET /health -> health", + "GET /legacy/ -> views.legacy", + ] + ); + } + + #[test] + fn rust_test_requests_should_be_recognized_for_axum_oneshot_and_actix() { + let source = r#" +use axum::http::{Method, Request}; +use tower::ServiceExt; + +#[tokio::test] +async fn creates_order() { + let request = Request::builder() + .method(Method::POST) + .uri("/orders") + .body(Body::empty()) + .unwrap(); + app().oneshot(request).await.unwrap(); +} + +#[actix_web::test] +async fn reads_order() { + let id = 42; + let request = test::TestRequest::get().uri(&format!("/orders/{}", id)).to_request(); +} +"#; + let facts = parse_rust_source(source) + .into_iter() + .map(|item| { + format!( + "{:?} {:?} {} {} {}", + item.role, + item.framework, + item.symbol_name.unwrap_or_default(), + item.method.unwrap_or_default(), + item.path.unwrap_or_default() + ) + }) + .filter(|fact| !fact.starts_with("Call")) + .collect::>(); + + assert_eq!( + facts, + [ + "Test TokioTest creates_order ", + "Consumer AxumOneshot creates_order POST /orders", + "Test RustTest reads_order ", + "Consumer ActixTest reads_order GET /orders/{value}", + ] + ); + } + + #[test] + fn python_wrappers_calls_and_fixture_requests_should_be_recorded() { + let source = r#" +import pytest +import requests +from tests.helpers import create_order + +def get_json(path): + return requests.get(BASE + path) + +@pytest.fixture +def order(client): + return create_order(client, "/orders") + +def test_reads_order(order, api): + api.read_order(order) + get_json("/orders/7") +"#; + let result = parse_python_source(source); + let calls = result + .iter() + .filter_map(|item| { + let call = item.call.as_ref()?; + Some(( + item.symbol_name.clone()?, + call.callee.clone(), + call.arguments + .iter() + .filter_map(|argument| argument.value.as_ref()?.client_literal()) + .collect::>(), + )) + }) + .collect::>(); + let wrapper = result + .iter() + .find(|item| item.role == SourceRole::Consumer) + .and_then(|item| item.url.clone()); + + assert_eq!( + wrapper.map(|url| url.parts), + Some(vec![ + UrlPart::Value(Some("BASE".to_owned())), + UrlPart::Parameter { + name: "path".to_owned(), + index: 0 + }, + ]) + ); + assert_eq!( + calls, + [ + ( + "order".to_owned(), + SymbolRef::Fixture("client".to_owned()), + vec![] + ), + ( + "order".to_owned(), + SymbolRef::Import { + module: "tests.helpers".to_owned(), + name: "create_order".to_owned() + }, + vec!["/orders".to_owned()] + ), + ( + "test_reads_order".to_owned(), + SymbolRef::Fixture("api".to_owned()), + vec![] + ), + ( + "test_reads_order".to_owned(), + SymbolRef::Fixture("order".to_owned()), + vec![] + ), + ( + "test_reads_order".to_owned(), + SymbolRef::Call("api.read_order".to_owned()), + vec![] + ), + ( + "test_reads_order".to_owned(), + SymbolRef::Local("get_json".to_owned()), + vec!["/orders/7".to_owned()] + ), + ] + ); } #[test] diff --git a/crates/code-system-graph-core/src/source_http/brace_clients.rs b/crates/code-system-graph-core/src/source_http/brace_clients.rs new file mode 100644 index 0000000..d504e7f --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_clients.rs @@ -0,0 +1,627 @@ +//! HTTP client calls and call sites in JavaScript and TypeScript, Go, and Java. +//! +//! Client URLs are evaluated as [`UrlTemplate`]s, so constants, format strings, and wrapper +//! parameters reach repository-level composition instead of being dropped as dynamic. + +use std::collections::BTreeMap; + +use self::test_clients::{httptest_servers, targets_httptest}; +use super::brace_flows::{BraceScopes, expression_end}; +use super::brace_lexer::{BraceDialect, lex_brace}; +use super::brace_tests::{ + ScriptBlock, record_go_tests, record_java_tests, record_script_tests, script_blocks +}; +use super::{SourceObservationCollector, Token, http_from_literal, matching, split_operands}; +use crate::{ + SourceFramework, SourceLanguage, SourceLineRange, SourceRole, SourceWarning, UrlPart, UrlTemplate +}; + +const METHODS: [&str; 8] = [ + "GET", "POST", "PUT", "PATCH", "DELETE", "HEAD", "OPTIONS", "TRACE", +]; + +/// Names whose calls are runtime or assertion utilities rather than repository functions. +const SCRIPT_UTILITIES: [&str; 13] = [ + "console", "JSON", "Object", "Array", "Math", "Promise", "Number", "String", "Date", "expect", + "fetch", "axios", "require", +]; +const GO_UTILITIES: [&str; 8] = [ + "http", "fmt", "strings", "strconv", "errors", "json", "log", "t", +]; +const JAVA_UTILITIES: [&str; 4] = ["System", "String", "Objects", "Arrays"]; + +mod test_clients; +#[cfg(test)] +mod tests; + +/// Client libraries that a file imports. +#[derive(Debug, Clone, Copy, Default)] +pub(crate) struct BraceClients { + pub(crate) axios: bool, + pub(crate) go_http: bool, + pub(crate) web_client: bool, +} + +/// Records HTTP client calls and call sites of one JavaScript, TypeScript, Go, or Java file. +pub(crate) fn collect_brace_clients( + source: &str, + language: SourceLanguage, + clients: BraceClients, + observations: &mut SourceObservationCollector<'_>, +) { + let dialect = match language { + SourceLanguage::JavaScript | SourceLanguage::TypeScript => BraceDialect::Script, + SourceLanguage::Go => BraceDialect::Go, + SourceLanguage::Java => BraceDialect::Java, + SourceLanguage::Python | SourceLanguage::Rust => return, + }; + let tokens = lex_brace(source, dialect); + let blocks = if dialect == BraceDialect::Script { + script_blocks(&tokens) + } else { + Vec::new() + }; + let scopes = BraceScopes::new( + &tokens, + dialect, + blocks.iter().filter_map(ScriptBlock::function).collect(), + ); + let consumers = Consumers { + scopes: &scopes, + language, + }; + let all_calls = declares_tests(&tokens, dialect); + match dialect { + BraceDialect::Script => { + consumers.fetch(observations); + let instances = if clients.axios { + axios_instances(&scopes) + } else { + BTreeMap::new() + }; + if clients.axios { + consumers.axios(&instances, observations); + } + consumers.supertest(observations); + consumers.playwright(observations); + let is_client = + |head: &str| SCRIPT_UTILITIES.contains(&head) || instances.contains_key(head); + scopes.record_calls(language, &is_client, all_calls, observations); + record_script_tests(&tokens, &blocks, language, observations); + } + BraceDialect::Go => { + if clients.go_http { + consumers.go_http(&go_client_names(&tokens), observations); + } + consumers.httptest(observations); + let is_client = |head: &str| GO_UTILITIES.contains(&head); + scopes.record_calls(language, &is_client, all_calls, observations); + record_go_tests(&scopes, observations); + } + BraceDialect::Java => { + let mentions = |name: &str| tokens.iter().any(|token| token.is_ident(name)); + if mentions("WebTestClient") { + consumers.web_client(SourceFramework::WebTestClient, observations); + } else if clients.web_client { + consumers.web_client(SourceFramework::WebClient, observations); + } + if mentions("MockMvc") || mentions("MockMvcRequestBuilders") { + consumers.mock_mvc(observations); + } + if mentions("restassured") || mentions("RestAssured") { + consumers.rest_assured(observations); + } + consumers.rest_template(observations); + let is_client = |head: &str| JAVA_UTILITIES.contains(&head); + scopes.record_calls(language, &is_client, all_calls, observations); + record_java_tests(&scopes, observations); + } + } +} + +/// Start line and name of each Java method, in declaration order. +pub(crate) fn java_method_lines(source: &str) -> Vec<(u32, String)> { + let tokens = lex_brace(source, BraceDialect::Java); + let scopes = BraceScopes::new(&tokens, BraceDialect::Java, Vec::new()); + let mut methods = scopes + .functions() + .iter() + .map(|function| (tokens[function.start_token].line, function.name.clone())) + .collect::>(); + methods.sort(); + methods +} + +/// Whether the file declares tests, so that every call from its functions is recorded. +fn declares_tests(tokens: &[Token], dialect: BraceDialect) -> bool { + tokens + .iter() + .enumerate() + .any(|(index, token)| match dialect { + BraceDialect::Script => { + ["describe", "it", "test"] + .iter() + .any(|name| token.is_ident(name)) + && tokens.get(index + 1).is_some_and(|next| next.is_punct('(')) + && index + .checked_sub(1) + .is_none_or(|previous| !tokens[previous].is_punct('.')) + } + BraceDialect::Go => token.is_ident("testing"), + BraceDialect::Java => { + token.is_ident("Test") && index > 0 && tokens[index - 1].is_punct('@') + } + }) +} + +struct Consumers<'s, 'a> { + scopes: &'s BraceScopes<'a>, + language: SourceLanguage, +} + +impl Consumers<'_, '_> { + fn tokens(&self) -> &[Token] { + self.scopes.tokens + } + + fn argument_template( + &self, + argument: Option<&(usize, usize)>, + at: usize, + ) -> Option { + argument + .filter(|(start, end)| start < end) + .map(|(start, end)| self.scopes.template(*start, *end, at)) + } + + /// Records a consumer call spanning tokens `start..=close`; a missing method is dynamic. + fn push( + &self, + framework: SourceFramework, + method: Option, + template: Option, + (start, close): (usize, usize), + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens(); + let literal = template.as_ref().and_then(UrlTemplate::client_literal); + let mut observation = http_from_literal( + self.language, + framework, + SourceRole::Consumer, + method, + literal.as_deref(), + self.scopes.function_name(start), + SourceLineRange { + start: tokens[start].line, + end: tokens[close].end_line, + }, + false, + ); + observation.url = template.filter(UrlTemplate::has_parameters); + if observation.method.is_none() { + observation.warnings.push(SourceWarning::DynamicMethod); + observation.warnings.sort(); + } + observations.push(observation); + } + + fn arguments(&self, open: usize) -> Option<(usize, Vec<(usize, usize)>)> { + let tokens = self.tokens(); + let close = matching(tokens, open, '(', ')')?; + let arguments = split_operands(tokens, open + 1, close, ','); + Some((close, arguments)) + } + + fn fetch(&self, observations: &mut SourceObservationCollector<'_>) { + let tokens = self.tokens(); + for index in 0..tokens.len() { + if !tokens[index].is_ident("fetch") + || !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let global = index < 2 + || !tokens[index - 1].is_punct('.') + || ["window", "globalThis", "self"] + .iter() + .any(|name| tokens[index - 2].is_ident(name)); + if !global { + continue; + } + let Some((close, arguments)) = self.arguments(index + 1) else { + continue; + }; + let method = options_method(tokens, arguments.get(1).copied()); + let template = self.argument_template(arguments.first(), index); + self.push( + SourceFramework::Fetch, + method, + template, + (index, close), + observations, + ); + } + } + + fn axios( + &self, + instances: &BTreeMap, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens(); + for index in 0..tokens.len() { + let Some(name) = tokens[index].ident() else { + continue; + }; + let Some((receiver, call, open)) = receiver_call(tokens, index) else { + if name == "axios" + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + && index + .checked_sub(1) + .is_none_or(|previous| !tokens[previous].is_punct('.')) + { + self.axios_request(None, index, index + 1, observations); + } + continue; + }; + let base = match instances.get(&receiver) { + Some(base) => Some(base), + None if receiver == "axios" => None, + None => continue, + }; + if call == "request" { + self.axios_request(base, index, open, observations); + continue; + } + let Some(method) = canonical(&call) else { + continue; + }; + let Some((close, arguments)) = self.arguments(open) else { + continue; + }; + let template = self + .argument_template(arguments.first(), index) + .map(|path| join_base(base, path)); + self.push( + SourceFramework::Axios, + Some(method), + template, + (index, close), + observations, + ); + } + } + + /// `axios(config)`, `axios(url, config)`, or `instance.request(config)`. + fn axios_request( + &self, + base: Option<&UrlTemplate>, + start: usize, + open: usize, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens(); + let Some((close, arguments)) = self.arguments(open) else { + return; + }; + let (url, config) = match arguments.as_slice() { + [(first, end), rest @ ..] if !tokens[*first].is_punct('{') => ( + Some(self.scopes.template(*first, *end, start)), + rest.first().copied(), + ), + [config, ..] => (None, Some(*config)), + [] => return, + }; + let url = url.or_else(|| { + let (config_start, config_end) = config?; + let value = object_value(tokens, config_start, config_end, "url")?; + Some( + self.scopes + .template(value, expression_end(tokens, value).min(config_end), start), + ) + }); + let method = options_method(tokens, config); + let template = url.map(|path| join_base(base, path)); + self.push( + SourceFramework::Axios, + method, + template, + (start, close), + observations, + ); + } + + fn go_http(&self, clients: &[String], observations: &mut SourceObservationCollector<'_>) { + let tokens = self.tokens(); + let servers = httptest_servers(tokens); + let uses_httptest = tokens.iter().any(|token| token.is_ident("httptest")); + for index in 0..tokens.len() { + let Some((receiver, call, open)) = receiver_call(tokens, index) else { + continue; + }; + let last = receiver.rsplit('.').next().unwrap_or_default(); + let package = receiver == "http"; + let client = + receiver == "http.DefaultClient" || clients.iter().any(|name| name == last); + if !package && !client { + continue; + } + let Some((close, arguments)) = self.arguments(open) else { + continue; + }; + let (method, url) = match call.as_str() { + "Get" => (Some("GET".to_owned()), arguments.first()), + "Head" => (Some("HEAD".to_owned()), arguments.first()), + "Post" | "PostForm" => (Some("POST".to_owned()), arguments.first()), + "NewRequest" if package => (go_method(tokens, arguments.first()), arguments.get(1)), + "NewRequestWithContext" if package => { + (go_method(tokens, arguments.get(1)), arguments.get(2)) + } + _ => continue, + }; + let template = self.argument_template(url, index); + let framework = if template + .as_ref() + .is_some_and(|template| targets_httptest(template, &servers, uses_httptest)) + { + SourceFramework::Httptest + } else { + SourceFramework::GoNetHttp + }; + self.push(framework, method, template, (index, close), observations); + } + } + + fn web_client( + &self, + framework: SourceFramework, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens(); + for index in 1..tokens.len() { + if !tokens[index].is_ident("uri") + || !tokens[index - 1].is_punct('.') + || !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let Some((close, arguments)) = self.arguments(index + 1) else { + continue; + }; + let method = chain_method(tokens, index); + let template = self.argument_template(arguments.first(), index); + self.push(framework, method, template, (index, close), observations); + } + } +} + +fn canonical(name: &str) -> Option { + METHODS + .iter() + .find(|method| method.eq_ignore_ascii_case(name)) + .map(|method| (*method).to_owned()) +} + +/// `receiver.call(` at identifier `index`, where `index` starts the receiver path. +fn receiver_call(tokens: &[Token], index: usize) -> Option<(String, String, usize)> { + if index > 0 && tokens[index - 1].is_punct('.') { + return None; + } + let mut segments = vec![tokens[index].ident()?]; + let mut cursor = index + 1; + while tokens.get(cursor)?.is_punct('.') { + segments.push(tokens.get(cursor + 1)?.ident()?); + cursor += 2; + } + if !tokens.get(cursor)?.is_punct('(') || segments.len() < 2 { + return None; + } + let call = segments.pop()?.to_owned(); + Some((segments.join("."), call, cursor)) +} + +/// Span of the `key` entry at the top level of the object literal in `start..end`. +fn object_entry(tokens: &[Token], start: usize, end: usize, key: &str) -> Option<(usize, usize)> { + if !tokens.get(start)?.is_punct('{') { + return None; + } + let close = matching(tokens, start, '{', '}')?.min(end); + split_operands(tokens, start + 1, close, ',') + .into_iter() + .find(|(entry, entry_end)| { + entry < entry_end + && (tokens[*entry].ident() == Some(key) || tokens[*entry].literal() == Some(key)) + }) +} + +/// Value start of `key:` at the top level of the object literal in `start..end`. +fn object_value(tokens: &[Token], start: usize, end: usize, key: &str) -> Option { + let (entry, entry_end) = object_entry(tokens, start, end, key)?; + (tokens.get(entry + 1)?.is_punct(':') && entry + 2 < entry_end).then_some(entry + 2) +} + +/// Method of a request whose options object spans `options`: the literal `method` entry, `GET` +/// without options or without that entry, and `None` when the method is computed. +fn options_method(tokens: &[Token], options: Option<(usize, usize)>) -> Option { + let Some((start, end)) = options.filter(|(start, end)| start < end) else { + return Some("GET".to_owned()); + }; + if !tokens[start].is_punct('{') { + return None; + } + let Some((entry, entry_end)) = object_entry(tokens, start, end, "method") else { + return Some("GET".to_owned()); + }; + (tokens.get(entry + 1)?.is_punct(':') && entry + 3 == entry_end) + .then(|| tokens[entry + 2].literal()) + .flatten() + .and_then(canonical) +} + +fn join_base(base: Option<&UrlTemplate>, path: UrlTemplate) -> UrlTemplate { + let Some(base) = base else { + return path; + }; + let absolute = path + .as_literal() + .is_some_and(|path| path.starts_with("http://") || path.starts_with("https://")); + if absolute || base.parts.is_empty() { + return path; + } + let mut joined = base.clone(); + let base_slash = base + .client_literal() + .is_some_and(|base| base.ends_with('/')); + let mut path = path; + if base_slash + && let Some(UrlPart::Text(text)) = path.parts.first_mut() + && let Some(stripped) = text.strip_prefix('/') + { + *text = stripped.to_owned(); + } + joined.extend(path); + joined +} + +/// Instances created by `axios.create({ baseURL })`, keyed by their binding path. +fn axios_instances(scopes: &BraceScopes<'_>) -> BTreeMap { + let tokens = scopes.tokens; + let mut instances = BTreeMap::new(); + for index in 2..tokens.len() { + if !(tokens[index].is_ident("create") + && tokens[index - 1].is_punct('.') + && tokens[index - 2].is_ident("axios") + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('('))) + { + continue; + } + let Some(assign) = index + .checked_sub(3) + .filter(|assign| tokens[*assign].is_punct('=')) + else { + continue; + }; + let Some(name) = assign.checked_sub(1).and_then(|name| tokens[name].ident()) else { + continue; + }; + let this_member = + assign >= 3 && tokens[assign - 2].is_punct('.') && tokens[assign - 3].is_ident("this"); + let binding = if this_member { + format!("this.{name}") + } else { + name.to_owned() + }; + let base = matching(tokens, index + 1, '(', ')') + .and_then(|close| { + let value = object_value(tokens, index + 2, close, "baseURL")?; + Some(scopes.template(value, expression_end(tokens, value).min(close), index)) + }) + .unwrap_or_default(); + instances.insert(binding, base); + } + instances +} + +/// Names of variables, fields, and parameters that hold a Go `http.Client`. +fn go_client_names(tokens: &[Token]) -> Vec { + let mut names = Vec::new(); + for index in 0..tokens.len() { + if !(tokens[index].is_ident("http") + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('.')) + && tokens + .get(index + 2) + .is_some_and(|token| token.is_ident("Client"))) + { + continue; + } + let mut cursor = index; + while cursor > 0 && (tokens[cursor - 1].is_punct('*') || tokens[cursor - 1].is_punct('&')) { + cursor -= 1; + } + if cursor >= 2 && tokens[cursor - 1].is_punct('=') { + let assign = if tokens[cursor - 2].is_punct(':') { + cursor - 2 + } else { + cursor - 1 + }; + if let Some(name) = assign.checked_sub(1).and_then(|name| tokens[name].ident()) { + names.push(name.to_owned()); + } + } else if let Some(name) = cursor.checked_sub(1).and_then(|name| tokens[name].ident()) { + names.push(name.to_owned()); + } + } + names.sort(); + names.dedup(); + names +} + +/// Method argument of `http.NewRequest`: a literal or an `http.MethodX` constant. +fn go_method(tokens: &[Token], argument: Option<&(usize, usize)>) -> Option { + let (start, end) = *argument?; + match end - start { + 1 => tokens[start].literal().and_then(canonical), + 3 if tokens[start].is_ident("http") => tokens[start + 2] + .ident() + .and_then(|name| name.strip_prefix("Method")) + .and_then(canonical), + _ => None, + } +} + +/// HTTP method selected earlier in a `WebClient` chain ending at the `uri` identifier `index`. +fn chain_method(tokens: &[Token], index: usize) -> Option { + let mut depth = 0_i32; + let mut cursor = index; + while cursor > 0 { + cursor -= 1; + let token = &tokens[cursor]; + if token.is_punct(')') { + depth += 1; + } else if token.is_punct('(') { + depth -= 1; + if depth < 0 { + return None; + } + } else if depth == 0 + && (token.is_punct(';') + || token.is_punct('{') + || token.is_punct('}') + || token.is_punct('=')) + { + return None; + } + if depth != 0 || cursor == 0 || !tokens[cursor - 1].is_punct('.') { + continue; + } + let Some(name) = token.ident() else { + continue; + }; + if name == "method" { + let argument = tokens.get(cursor + 4).and_then(Token::ident); + return argument.and_then(canonical); + } + if tokens + .get(cursor + 1) + .is_some_and(|next| next.is_punct('(')) + && tokens + .get(cursor + 2) + .is_some_and(|next| next.is_punct(')')) + && let Some(method) = canonical(name) + { + return Some(method); + } + } + None +} diff --git a/crates/code-system-graph-core/src/source_http/brace_clients/test_clients.rs b/crates/code-system-graph-core/src/source_http/brace_clients/test_clients.rs new file mode 100644 index 0000000..2784b42 --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_clients/test_clients.rs @@ -0,0 +1,448 @@ +//! Test clients: supertest and Playwright in scripts, `httptest` in Go, and `MockMvc`, +//! `RestAssured`, `WebTestClient`, and `RestTemplate` in Java. + +use std::collections::BTreeSet; + +use super::{Consumers, canonical, join_base, receiver_call}; +use crate::source_http::brace_flows::matching_open; +use crate::source_http::{Token, matching}; +use crate::{SourceFramework, UrlPart, UrlTemplate}; + +const SUPERTEST_METHODS: [&str; 7] = ["get", "post", "put", "patch", "delete", "head", "options"]; + +impl Consumers<'_, '_> { + /// `request(app).get(url)` and agents bound to `request(app)` or `request.agent(app)`. + pub(super) fn supertest(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + let locals = module_locals(tokens, "supertest"); + if locals.is_empty() { + return; + } + let agents = bound_calls(tokens, &locals); + for index in 0..tokens.len() { + let Some(name) = tokens[index].ident() else { + continue; + }; + if index > 0 && tokens[index - 1].is_punct('.') { + continue; + } + let (base, verb) = if locals.contains(name) + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + let Some(close) = matching(tokens, index + 1, '(', ')') else { + continue; + }; + let base = tokens + .get(index + 2) + .filter(|token| token.literal().is_some()) + .map(|_| self.scopes.template(index + 2, close, index)); + (base, close + 1) + } else if agents.contains(name) { + (None, index + 1) + } else { + continue; + }; + self.chained_verb( + SourceFramework::Supertest, + base.as_ref(), + index, + verb, + observations, + ); + } + } + + /// Playwright `request.get(url)` through the `request` fixture, `page.request`, or a + /// context created by `request.newContext()`. + pub(super) fn playwright(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + if !tokens + .iter() + .any(|token| token.literal() == Some("@playwright/test")) + { + return; + } + let mut receivers = BTreeSet::from(["request".to_owned()]); + for index in 2..tokens.len() { + if tokens[index].is_ident("newContext") && tokens[index - 1].is_punct('.') { + let assign = (index.saturating_sub(8)..index) + .rev() + .find(|candidate| tokens[*candidate].is_punct('=')); + if let Some(name) = assign + .and_then(|assign| assign.checked_sub(1)) + .and_then(|name| tokens[name].ident()) + { + receivers.insert(name.to_owned()); + } + } + } + for index in 0..tokens.len() { + let Some(name) = tokens[index].ident() else { + continue; + }; + let qualified = + index < 2 || !tokens[index - 1].is_punct('.') || tokens[index - 2].is_ident("page"); + if receivers.contains(name) && qualified { + self.chained_verb( + SourceFramework::Playwright, + None, + index, + index + 1, + observations, + ); + } + } + } + + /// `.verb(url, ...)` starting at token `dot`, attributed to the call starting at `start`. + fn chained_verb( + &self, + framework: SourceFramework, + base: Option<&UrlTemplate>, + start: usize, + dot: usize, + observations: &mut super::SourceObservationCollector<'_>, + ) { + let tokens = self.tokens(); + if !tokens.get(dot).is_some_and(|token| token.is_punct('.')) { + return; + } + let Some(verb) = tokens.get(dot + 1).and_then(Token::ident) else { + return; + }; + if !SUPERTEST_METHODS.contains(&verb) { + return; + } + let Some((close, arguments)) = self.arguments(dot + 2) else { + return; + }; + let template = self + .argument_template(arguments.first(), start) + .map(|path| join_base(base, path)); + self.push( + framework, + canonical(verb), + template, + (start, close), + observations, + ); + } + + /// `httptest.NewRequest(method, url, body)`. + pub(super) fn httptest(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + for index in 0..tokens.len() { + let Some((receiver, call, open)) = receiver_call(tokens, index) else { + continue; + }; + if receiver != "httptest" || call != "NewRequest" { + continue; + } + let Some((close, arguments)) = self.arguments(open) else { + continue; + }; + let method = super::go_method(tokens, arguments.first()); + let template = self.argument_template(arguments.get(1), index); + self.push( + SourceFramework::Httptest, + method, + template, + (index, close), + observations, + ); + } + } + + /// `mockMvc.perform(get(url))`, also through `MockMvcRequestBuilders` and `request`. + pub(super) fn mock_mvc(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + for index in 1..tokens.len() { + if !tokens[index].is_ident("perform") + || !tokens[index - 1].is_punct('.') + || !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let mut builder = index + 2; + if tokens + .get(builder) + .is_some_and(|token| token.is_ident("MockMvcRequestBuilders")) + && tokens + .get(builder + 1) + .is_some_and(|token| token.is_punct('.')) + { + builder += 2; + } + let Some(verb) = tokens.get(builder).and_then(Token::ident) else { + continue; + }; + let Some((close, arguments)) = self.arguments(builder + 1) else { + continue; + }; + let (method, url) = match verb { + "request" => ( + java_http_method(tokens, arguments.first()), + arguments.get(1), + ), + "multipart" => (Some("POST".to_owned()), arguments.first()), + verb => match canonical(verb) { + Some(method) => (Some(method), arguments.first()), + None => continue, + }, + }; + let template = self.argument_template(url, index); + self.push( + SourceFramework::MockMvc, + method, + template, + (index, close), + observations, + ); + } + } + + /// `given()...when().get(url)` and `RestAssured.get(url)` chains. + pub(super) fn rest_assured(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + for index in 1..tokens.len() { + let Some(verb) = tokens[index].ident() else { + continue; + }; + if !SUPERTEST_METHODS.contains(&verb) + || !tokens[index - 1].is_punct('.') + || !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let heads = chain_heads(tokens, index - 1); + if !heads + .iter() + .any(|head| matches!(*head, "when" | "given" | "RestAssured")) + { + continue; + } + let Some((close, arguments)) = self.arguments(index + 1) else { + continue; + }; + let Some(template) = self.argument_template(arguments.first(), index) else { + continue; + }; + self.push( + SourceFramework::RestAssured, + canonical(verb), + Some(template), + (index, close), + observations, + ); + } + } + + /// Calls on fields, parameters, and variables declared as `RestTemplate` or + /// `TestRestTemplate`. + pub(super) fn rest_template(&self, observations: &mut super::SourceObservationCollector<'_>) { + let tokens = self.tokens(); + let mut receivers = Vec::new(); + for index in 0..tokens.len().saturating_sub(2) { + let framework = match tokens[index].ident() { + Some("RestTemplate") => SourceFramework::RestTemplate, + Some("TestRestTemplate") => SourceFramework::TestRestTemplate, + _ => continue, + }; + let declares = tokens[index + 2].is_punct(';') + || tokens[index + 2].is_punct('=') + || tokens[index + 2].is_punct(',') + || tokens[index + 2].is_punct(')'); + if let Some(name) = tokens[index + 1].ident().filter(|_| declares) { + receivers.push((name.to_owned(), framework)); + } + } + if receivers.is_empty() { + return; + } + for index in 0..tokens.len() { + let Some((receiver, call, open)) = receiver_call(tokens, index) else { + continue; + }; + let receiver = receiver.strip_prefix("this.").unwrap_or(&receiver); + let Some((_, framework)) = receivers.iter().find(|(name, _)| name == receiver) else { + continue; + }; + let Some((close, arguments)) = self.arguments(open) else { + continue; + }; + let method = match call.as_str() { + "getForObject" | "getForEntity" => Some("GET".to_owned()), + "postForObject" | "postForEntity" | "postForLocation" => Some("POST".to_owned()), + "put" => Some("PUT".to_owned()), + "delete" => Some("DELETE".to_owned()), + "patchForObject" => Some("PATCH".to_owned()), + "headForHeaders" => Some("HEAD".to_owned()), + "optionsForAllow" => Some("OPTIONS".to_owned()), + "exchange" => java_http_method(tokens, arguments.get(1)), + _ => continue, + }; + let template = self.argument_template(arguments.first(), index); + self.push(*framework, method, template, (index, close), observations); + } + } +} + +/// Whether `template` is addressed to a Go `httptest` server: it starts with the `URL` field of +/// one of `servers`, or it is a relative path in a file that uses `httptest`. +pub(super) fn targets_httptest( + template: &UrlTemplate, + servers: &[String], + uses_httptest: bool, +) -> bool { + match template.parts.first() { + Some(UrlPart::Value(Some(name))) => name + .strip_suffix(".URL") + .is_some_and(|server| servers.iter().any(|candidate| candidate == server)), + Some(UrlPart::Text(text)) => uses_httptest && text.starts_with('/'), + _ => false, + } +} + +/// Names bound to `httptest.NewServer`, `NewTLSServer`, or `NewUnstartedServer`. +pub(super) fn httptest_servers(tokens: &[Token]) -> Vec { + let mut servers = Vec::new(); + for index in 0..tokens.len() { + let Some((receiver, call, _)) = receiver_call(tokens, index) else { + continue; + }; + if receiver != "httptest" + || !matches!( + call.as_str(), + "NewServer" | "NewTLSServer" | "NewUnstartedServer" + ) + { + continue; + } + let Some(assign) = index + .checked_sub(1) + .filter(|assign| tokens[*assign].is_punct('=')) + else { + continue; + }; + let target = if assign > 0 && tokens[assign - 1].is_punct(':') { + assign - 1 + } else { + assign + }; + if let Some(name) = target.checked_sub(1).and_then(|name| tokens[name].ident()) { + servers.push(name.to_owned()); + } + } + servers +} + +/// `HttpMethod.X` argument of a Spring call. +fn java_http_method(tokens: &[Token], argument: Option<&(usize, usize)>) -> Option { + let (start, end) = *argument?; + (end == start + 3 && tokens[start].is_ident("HttpMethod") && tokens[start + 1].is_punct('.')) + .then(|| tokens[start + 2].ident()) + .flatten() + .and_then(canonical) +} + +/// Identifiers of the calls and names chained before the `.` at token `dot`. +fn chain_heads(tokens: &[Token], dot: usize) -> Vec<&str> { + let mut heads = Vec::new(); + let mut cursor = dot; + while cursor > 0 && tokens[cursor].is_punct('.') { + let mut previous = cursor - 1; + if tokens[previous].is_punct(')') { + let Some(open) = matching_open(tokens, previous) else { + break; + }; + let Some(callee) = open.checked_sub(1) else { + break; + }; + previous = callee; + } + let Some(name) = tokens[previous].ident() else { + break; + }; + heads.push(name); + let Some(next) = previous.checked_sub(1) else { + break; + }; + cursor = next; + } + heads +} + +/// Local names bound to the default or namespace export of the script module `module`. +fn module_locals(tokens: &[Token], module: &str) -> BTreeSet { + let mut locals = BTreeSet::new(); + for index in 0..tokens.len() { + if tokens[index].literal() != Some(module) || index < 2 { + continue; + } + if tokens[index - 1].is_ident("from") { + let Some(import) = (index.saturating_sub(16)..index) + .rev() + .find(|candidate| tokens[*candidate].is_ident("import")) + else { + continue; + }; + let clause = &tokens[import + 1..index - 1]; + if let Some(name) = clause + .first() + .and_then(Token::ident) + .filter(|name| *name != "type") + { + locals.insert(name.to_owned()); + } else if let [star, alias, name, ..] = clause + && star.is_punct('*') + && alias.is_ident("as") + && let Some(name) = name.ident() + { + locals.insert(name.to_owned()); + } + } else if tokens[index - 1].is_punct('(') + && tokens[index - 2].is_ident("require") + && index >= 4 + && tokens[index - 3].is_punct('=') + && let Some(name) = tokens[index - 4].ident() + { + locals.insert(name.to_owned()); + } + } + locals +} + +/// Names bound to a call of one of `callees`, directly or through one member such as `agent`. +fn bound_calls(tokens: &[Token], callees: &BTreeSet) -> BTreeSet { + let mut names = BTreeSet::new(); + for index in 2..tokens.len() { + let Some(callee) = tokens[index].ident() else { + continue; + }; + if !callees.contains(callee) || !tokens[index - 1].is_punct('=') { + continue; + } + let direct = tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')); + let member = tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('.')) + && tokens + .get(index + 3) + .is_some_and(|token| token.is_punct('(')); + if !direct && !member { + continue; + } + if let Some(name) = tokens[index - 2].ident() { + names.insert(name.to_owned()); + } + } + names +} diff --git a/crates/code-system-graph-core/src/source_http/brace_clients/tests.rs b/crates/code-system-graph-core/src/source_http/brace_clients/tests.rs new file mode 100644 index 0000000..164426d --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_clients/tests.rs @@ -0,0 +1,377 @@ +use crate::{ + SourceObservation, SourceRole, SymbolRef, UrlPart, parse_go_source, parse_java_source, parse_javascript_source, parse_typescript_source +}; + +type Consumer<'a> = ( + Option<&'a str>, + Option<&'a str>, + Option<&'a str>, + Option<&'a str>, +); + +fn consumers(observations: &[SourceObservation]) -> Vec> { + observations + .iter() + .filter(|item| item.role == SourceRole::Consumer) + .map(|item| { + ( + item.method.as_deref(), + item.path.as_deref(), + item.authority.as_deref(), + item.symbol_name.as_deref(), + ) + }) + .collect() +} + +#[test] +fn typescript_clients_should_resolve_constants_templates_instances_and_wrappers() { + let source = r#" +import axios from "axios"; +const API = "/api/v1"; +const api = axios.create({ baseURL: "http://orders:8080/api/" }); +export async function getOrder(id: string): Promise { + return fetch(`${API}/orders/${id}`).then((response) => response.json()); +} +export const deleteOrder = async (id: string) => api.delete(`/orders/${encodeURIComponent(id)}`); +export function request(path: string, method: string) { + return fetch(API + path, { method }); +} +document.querySelector("a").addEventListener("click", () => { + fetch("/first"); + fetch("/second", { method: "POST" }); +}); +"#; + let result = parse_typescript_source(source); + + assert_eq!( + consumers(&result), + [ + ( + Some("GET"), + Some("/api/v1/orders/{id}"), + None, + Some("getOrder") + ), + ( + Some("DELETE"), + Some("/api/orders/{id}"), + Some("orders:8080"), + Some("deleteOrder") + ), + (None, None, None, Some("request")), + (Some("GET"), Some("/first"), None, None), + (Some("POST"), Some("/second"), None, None), + ] + ); + let wrapper = result + .iter() + .find(|item| item.symbol_name.as_deref() == Some("request")) + .and_then(|item| item.url.as_ref()) + .map(|url| url.parts.clone()); + assert_eq!( + wrapper, + Some(vec![ + UrlPart::Text("/api/v1".to_owned()), + UrlPart::Parameter { + name: "path".to_owned(), + index: 0 + }, + ]) + ); +} + +#[test] +fn typescript_test_files_should_record_relative_import_calls() { + let source = r#" +import { createOrder } from "./helpers"; +import * as api from "../api"; +export async function seed() { + await createOrder("/orders/1"); + await api.listOrders(); + localHelper(); +} +describe("orders", () => {}); +"#; + let calls = parse_typescript_source(source) + .into_iter() + .filter(|item| item.role == SourceRole::Call) + .filter_map(|item| item.call.map(|call| call.callee)) + .collect::>(); + + assert_eq!( + calls, + [ + SymbolRef::Import { + module: "./helpers".to_owned(), + name: "createOrder".to_owned() + }, + SymbolRef::Import { + module: "../api".to_owned(), + name: "listOrders".to_owned() + }, + SymbolRef::Local("localHelper".to_owned()), + ] + ); +} + +#[test] +fn go_clients_should_resolve_sprintf_constants_and_client_receivers() { + let source = r#" +package users + +import ( + "fmt" + "net/http" +) + +const base = "http://users:9000" + +type Client struct { + client *http.Client +} + +func (c *Client) GetUser(id string) (*User, error) { + url := fmt.Sprintf("%s/v1/users/%s", base, id) + req, _ := http.NewRequest(http.MethodGet, url, nil) + return c.do(req) +} + +func (c *Client) Ping() error { + _, err := c.client.Head(base + "/healthz") + return err +} + +func Post(body io.Reader) { + http.Post("/v1/users", "application/json", body) +} +"#; + let result = parse_go_source(source); + + assert_eq!( + consumers(&result), + [ + ( + Some("GET"), + Some("/v1/users/{id}"), + Some("users:9000"), + Some("GetUser") + ), + ( + Some("HEAD"), + Some("/healthz"), + Some("users:9000"), + Some("Ping") + ), + (Some("POST"), Some("/v1/users"), None, Some("Post")), + ] + ); +} + +#[test] +fn java_web_client_should_resolve_chain_methods_and_constants() { + let source = r#" +import org.springframework.web.reactive.function.client.WebClient; + +class OrdersClient { + private static final String ORDERS = "/orders"; + + Mono find(String id) { + return client.get().uri(ORDERS + "/" + id).retrieve().bodyToMono(Order.class); + } + + void create() { + client.method(HttpMethod.POST).uri("/orders/{id}", 7).retrieve(); + } +} +"#; + let result = parse_java_source(source); + + assert_eq!( + consumers(&result), + [ + (Some("GET"), Some("/orders/{id}"), None, Some("find")), + (Some("POST"), Some("/orders/{id}"), None, Some("create")), + ] + ); +} + +/// Sorted tests, consumers, and local calls as `role framework symbol: detail`. +fn test_facts(observations: &[SourceObservation]) -> Vec { + let mut facts = observations + .iter() + .filter_map(|item| { + let detail = match item.role { + SourceRole::Test => String::new(), + SourceRole::Consumer => format!( + "{} {}", + item.method.as_deref().unwrap_or("?"), + item.path.as_deref().unwrap_or("?") + ), + SourceRole::Call => match &item.call.as_ref()?.callee { + SymbolRef::Local(name) => format!("-> {name}"), + _ => return None, + }, + _ => return None, + }; + Some(format!( + "{:?} {:?} {}: {detail}", + item.role, + item.framework, + item.symbol_name.as_deref().unwrap_or("-") + )) + }) + .collect::>(); + facts.sort(); + facts +} + +#[test] +fn script_tests_should_be_named_by_describe_chain_and_reach_hooks_and_supertest() { + let source = r#" +const request = require("supertest"); +const app = require("../app"); + +describe("orders", () => { + let agent; + beforeEach(async () => { + agent = request.agent(app); + await fetch("/seed", { method: "POST" }); + }); + + it("creates an order", async () => { + await request(app).post("/orders").send({ sku: "A" }); + }); + + describe("by id", function () { + test.only(`reads one`, async () => { + await agent.get("/orders/42"); + }); + }); +}); +"#; + let facts = test_facts(&parse_javascript_source(source)); + + assert_eq!( + facts, + [ + "Call Fetch orders > by id > reads one: -> orders > beforeEach", + "Call Fetch orders > creates an order: -> orders > beforeEach", + "Call Fetch orders > creates an order: -> request", + "Consumer Fetch orders > beforeEach: POST /seed", + "Consumer Supertest orders > by id > reads one: GET /orders/42", + "Consumer Supertest orders > creates an order: POST /orders", + "Test Jest orders > by id > reads one: ", + "Test Jest orders > creates an order: ", + ] + ); +} + +#[test] +fn playwright_request_fixture_calls_should_belong_to_their_test() { + let source = r#" +import { test, expect } from "@playwright/test"; + +test.describe("catalog", () => { + test("lists products", async ({ request }) => { + const response = await request.get("/api/products"); + expect(response.ok()).toBeTruthy(); + }); +}); +"#; + let facts = test_facts(&parse_typescript_source(source)); + + assert_eq!( + facts, + [ + "Consumer Playwright catalog > lists products: GET /api/products", + "Test Playwright catalog > lists products: ", + ] + ); +} + +#[test] +fn go_tests_should_recognize_httptest_requests_and_servers() { + let source = r#" +package orders + +import ( + "net/http" + "net/http/httptest" + "testing" +) + +func TestGetOrder(t *testing.T) { + req := httptest.NewRequest(http.MethodGet, "/orders/42", nil) + rec := httptest.NewRecorder() + router().ServeHTTP(rec, req) +} + +func TestCreateOrder(t *testing.T) { + srv := httptest.NewServer(router()) + defer srv.Close() + http.Post(srv.URL+"/orders", "application/json", nil) +} + +func helper(t *testing.T) {} +"#; + let facts = test_facts(&parse_go_source(source)); + + assert_eq!( + facts, + [ + "Consumer Httptest TestCreateOrder: POST /orders", + "Consumer Httptest TestGetOrder: GET /orders/42", + "Test GoTest TestCreateOrder: ", + "Test GoTest TestGetOrder: ", + ] + ); +} + +#[test] +fn java_tests_should_recognize_mockmvc_restassured_and_rest_templates() { + let source = r#" +import org.junit.jupiter.api.Test; +import org.springframework.test.web.servlet.MockMvc; +import static org.springframework.test.web.servlet.request.MockMvcRequestBuilders.get; +import static io.restassured.RestAssured.given; + +class OrdersTest { + @Autowired + private MockMvc mockMvc; + @Autowired + private TestRestTemplate rest; + + @Test + void readsOrder() throws Exception { + mockMvc.perform(get("/orders/{id}", 42)).andExpect(status().isOk()); + } + + @ParameterizedTest + @ValueSource(strings = {"A", "B"}) + void createsOrder(String sku) { + given().body(sku).when().post("/orders").then().statusCode(201); + } + + @Test + public void deletesOrder() { + rest.exchange("/orders/7", HttpMethod.DELETE, null, Void.class); + } + + private void helper() {} +} +"#; + let facts = test_facts(&parse_java_source(source)); + + assert_eq!( + facts, + [ + "Consumer MockMvc readsOrder: GET /orders/{id}", + "Consumer RestAssured createsOrder: POST /orders", + "Consumer TestRestTemplate deletesOrder: DELETE /orders/7", + "Test JUnit createsOrder: ", + "Test JUnit deletesOrder: ", + "Test JUnit readsOrder: ", + ] + ); +} diff --git a/crates/code-system-graph-core/src/source_http/brace_flows.rs b/crates/code-system-graph-core/src/source_http/brace_flows.rs new file mode 100644 index 0000000..9ec085e --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_flows.rs @@ -0,0 +1,962 @@ +//! Function scopes, URL expressions, and call sites for JavaScript and TypeScript, Go, and Java. +//! +//! Named functions are discovered per dialect: declarations, functions and arrow functions bound +//! to a name, class and object methods, Go functions and methods, and Java methods. URL +//! expressions are evaluated through literals, template literals, `+` concatenation, +//! `fmt.Sprintf`, `String.format`, conversions, and names bound earlier in the same function or +//! at file level. + +use std::cmp::Reverse; +use std::collections::BTreeMap; + +use super::brace_lexer::BraceDialect; +use super::{ + FunctionSpan, SourceObservationCollector, Token, TokenKind, call_observation, matching, split_operands +}; +use crate::url_template::{CONVERSIONS, printf_format, template_literal}; +use crate::{CallArgument, SourceLanguage, SourceLineRange, SymbolRef, UrlPart, UrlTemplate}; + +const SCRIPT_KEYWORDS: [&str; 22] = [ + "if", "for", "while", "switch", "catch", "function", "return", "with", "typeof", "new", + "await", "yield", "delete", "void", "super", "import", "require", "else", "do", "try", "in", + "of", +]; + +const GO_KEYWORDS: [&str; 14] = [ + "if", "for", "switch", "func", "return", "go", "defer", "select", "range", "make", "new", + "append", "len", "panic", +]; + +const JAVA_KEYWORDS: [&str; 16] = [ + "if", + "for", + "while", + "switch", + "catch", + "synchronized", + "return", + "new", + "try", + "else", + "do", + "throw", + "super", + "this", + "assert", + "case", +]; + +const DECLARATION_MODIFIERS: [&str; 10] = [ + "async", + "static", + "get", + "set", + "public", + "private", + "protected", + "readonly", + "override", + "abstract", +]; + +struct Binding { + function: Option, + index: usize, + name: String, + value: UrlTemplate, +} + +/// Named functions, parameters, bindings, and relative imports of one brace-language file. +pub(super) struct BraceScopes<'a> { + pub(super) tokens: &'a [Token], + pub(super) dialect: BraceDialect, + functions: Vec, + /// Innermost named function of each token. + owners: Vec>, + /// Whether each token belongs to a function signature. + signatures: Vec, + parameters: Vec>, + bindings: Vec, + imports: BTreeMap, +} + +impl<'a> BraceScopes<'a> { + /// Scopes of `tokens`, with `extra` functions found by the caller, such as test callbacks. + pub(super) fn new( + tokens: &'a [Token], + dialect: BraceDialect, + extra: Vec<(FunctionSpan, Vec)>, + ) -> Self { + let mut declared = match dialect { + BraceDialect::Script => script_functions(tokens), + BraceDialect::Go => go_functions(tokens), + BraceDialect::Java => java_functions(tokens), + }; + declared.extend(extra); + let (functions, parameters): (Vec, _) = declared.into_iter().unzip(); + let mut owners = vec![None; tokens.len()]; + let mut signatures = vec![false; tokens.len()]; + let mut by_size = (0..functions.len()).collect::>(); + by_size.sort_by_key(|&function| { + Reverse(functions[function].end_token - functions[function].start_token) + }); + for function in by_size { + let span = &functions[function]; + let end = span.end_token.min(tokens.len().saturating_sub(1)); + for owner in owners.iter_mut().take(end + 1).skip(span.start_token) { + *owner = Some(function); + } + for signature in signatures + .iter_mut() + .take(span.body_start_token) + .skip(span.start_token) + { + *signature = true; + } + } + let mut scopes = Self { + tokens, + dialect, + functions, + owners, + signatures, + parameters, + bindings: Vec::new(), + imports: if dialect == BraceDialect::Script { + script_imports(tokens) + } else { + BTreeMap::new() + }, + }; + let mut index = 0; + while index < tokens.len() { + if let Some((name, value_start)) = scopes.binding_target(index) { + let end = expression_end(tokens, value_start); + let value = scopes.template(value_start, end, index); + scopes.bindings.push(Binding { + function: scopes.innermost(index), + index, + name, + value, + }); + } + index += 1; + } + scopes + } + + pub(super) fn functions(&self) -> &[FunctionSpan] { + &self.functions + } + + pub(super) fn innermost(&self, index: usize) -> Option { + self.owners.get(index).copied().flatten() + } + + pub(super) fn function_name(&self, index: usize) -> Option { + self.innermost(index) + .map(|function| self.functions[function].name.clone()) + } + + /// Bound name and value start of a binding at token `index`. + fn binding_target(&self, index: usize) -> Option<(String, usize)> { + let tokens = self.tokens; + let name = tokens[index].ident()?; + let previous = index.checked_sub(1).map(|previous| &tokens[previous]); + if tokens.get(index + 1)?.is_punct(':') && tokens.get(index + 2)?.is_punct('=') { + return Some((name.to_owned(), index + 3)); + } + let this_member = previous.is_some_and(|token| token.is_punct('.')) + && index >= 2 + && tokens[index - 2].is_ident("this"); + if previous.is_some_and(|token| token.is_punct('.')) && !this_member { + return None; + } + let mut equals = index + 1; + if self.dialect == BraceDialect::Script + && tokens.get(equals)?.is_punct(':') + && previous.is_some_and(|token| { + token.is_ident("const") || token.is_ident("let") || token.is_ident("var") + }) + { + equals = (equals..tokens.len()) + .take_while(|candidate| !tokens[*candidate].is_punct(';')) + .find(|candidate| is_assignment(tokens, *candidate))?; + } + if !is_assignment(tokens, equals) { + return None; + } + let name = if this_member { + format!("this.{name}") + } else { + name.to_owned() + }; + Some((name, equals + 1)) + } + + /// Evaluates the expression in `start..end`, resolving names as seen at token `at`. + pub(super) fn template(&self, start: usize, end: usize, at: usize) -> UrlTemplate { + let mut template = UrlTemplate::default(); + for (operand_start, operand_end) in split_operands(self.tokens, start, end, '+') { + template.extend(self.operand(operand_start, operand_end, at)); + } + template + } + + fn operand(&self, start: usize, mut end: usize, at: usize) -> UrlTemplate { + let tokens = self.tokens; + let anonymous = || UrlTemplate::part(UrlPart::Value(None)); + while end > start + 3 + && tokens[end - 1].is_punct(')') + && tokens[end - 2].is_punct('(') + && tokens[end - 3].is_ident("toString") + && tokens[end - 4].is_punct('.') + { + end -= 4; + } + if start >= end { + return anonymous(); + } + if end == start + 1 { + return match &tokens[start].kind { + TokenKind::Literal(Some(value)) => UrlTemplate::text(value), + TokenKind::Template(body) => { + template_literal(body, &|name| self.resolve(name, at)).unwrap_or_else(anonymous) + } + TokenKind::Ident(_) => self.resolve(tokens[start].ident().unwrap_or_default(), at), + TokenKind::Literal(None) | TokenKind::Punct(_) => anonymous(), + }; + } + if tokens[start].is_punct('(') && matching(tokens, start, '(', ')') == Some(end - 1) { + return self.template(start + 1, end - 1, at); + } + if let Some(open) = (start..end).find(|index| tokens[*index].is_punct('(')) + && matching(tokens, open, '(', ')') == Some(end - 1) + { + return self.call_value(start, open, end - 1, at); + } + match dotted_name(tokens, start, end) { + Some(name) => self.resolve(&name, at), + None => anonymous(), + } + } + + /// Value of a call expression whose callee spans `start..open`. + fn call_value(&self, start: usize, open: usize, close: usize, at: usize) -> UrlTemplate { + let tokens = self.tokens; + let callee = dotted_name(tokens, start, open).unwrap_or_default(); + let callee = callee.strip_prefix("new ").unwrap_or(&callee); + let arguments = split_operands(tokens, open + 1, close, ',') + .into_iter() + .filter(|(argument_start, argument_end)| argument_start < argument_end) + .map(|(argument_start, argument_end)| self.template(argument_start, argument_end, at)) + .collect::>(); + let last = callee.rsplit('.').next().unwrap_or_default(); + match (callee, last) { + ("fmt.Sprintf" | "String.format", _) => { + let Some(format) = arguments.first().and_then(UrlTemplate::as_literal) else { + return UrlTemplate::part(UrlPart::Value(None)); + }; + printf_format(format, &arguments[1..]) + } + ("URL", _) if start > 0 && tokens[start - 1].is_ident("new") => arguments + .first() + .filter(|path| path.as_literal().is_some_and(|path| path.starts_with('/'))) + .cloned() + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(None))), + (_, conversion) if CONVERSIONS.contains(&conversion) => match arguments.as_slice() { + [value] if value.as_literal().is_none() => value.clone(), + _ => UrlTemplate::part(UrlPart::Value(None)), + }, + _ => UrlTemplate::part(UrlPart::Value(None)), + } + } + + /// Resolves a plain or dotted name at token `at`. + pub(super) fn resolve(&self, name: &str, at: usize) -> UrlTemplate { + let function = self.innermost(at); + if let Some(function) = function + && let Some(index) = self.parameters[function] + .iter() + .position(|parameter| parameter == name) + { + return UrlTemplate::part(UrlPart::Parameter { + name: name.to_owned(), + index, + }); + } + let visible = |binding: &&Binding| { + binding.name == name + && (binding.function.is_none() + || (binding.function == function && binding.index < at)) + }; + let local = self + .bindings + .iter() + .rev() + .filter(visible) + .find(|binding| binding.function.is_some()); + let global = || { + self.bindings + .iter() + .rev() + .filter(visible) + .find(|binding| binding.function.is_none()) + }; + let member = || { + name.strip_prefix("this.").and_then(|field| { + let mut values = self + .bindings + .iter() + .filter(|binding| { + binding.name == name || binding.function.is_none() && binding.name == field + }) + .map(|binding| &binding.value); + let first = values.next()?; + values.all(|value| value == first).then_some(first) + }) + }; + local + .or_else(global) + .map(|binding| &binding.value) + .or_else(member) + .filter(|value| !value.has_parameters() || local.is_some()) + .cloned() + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(Some(name.to_owned())))) + } + + /// Records calls from named functions; with `all_calls` every call, otherwise only calls + /// that pass a path literal or forward a parameter of the caller. + pub(super) fn record_calls( + &self, + language: SourceLanguage, + is_client: &dyn Fn(&str) -> bool, + all_calls: bool, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens; + let keywords: &[&str] = match self.dialect { + BraceDialect::Script => &SCRIPT_KEYWORDS, + BraceDialect::Go => &GO_KEYWORDS, + BraceDialect::Java => &JAVA_KEYWORDS, + }; + for index in 0..tokens.len() { + if !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let Some(name) = tokens[index].ident() else { + continue; + }; + if keywords.contains(&name) { + continue; + } + let Some((start, path)) = dotted_before(tokens, index) else { + continue; + }; + let head = path.split('.').next().unwrap_or_default(); + if self.signatures[index] + || is_client(head) + || start > 0 + && (tokens[start - 1].is_ident("new") || tokens[start - 1].is_ident("func")) + { + continue; + } + let Some(caller) = self.function_name(index) else { + continue; + }; + let Some(close) = matching(tokens, index + 1, '(', ')') else { + continue; + }; + let arguments = self.call_arguments(index + 1, close, index); + let forwards_url = arguments.iter().any(|argument| { + argument.value.as_ref().is_some_and(|value| { + value.has_parameters() + || value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(text) if text.contains('/'))) + }) + }); + if !all_calls && !forwards_url { + continue; + } + observations.push(call_observation( + language, + self.callee(&path), + arguments, + caller, + SourceLineRange { + start: tokens[start].line, + end: tokens[close].end_line, + }, + )); + } + } + + fn callee(&self, path: &str) -> SymbolRef { + let (head, rest) = path + .split_once('.') + .map_or((path, None), |(head, rest)| (head, Some(rest))); + if let Some((module, name)) = self.imports.get(head) { + let name = match (name.as_str(), rest) { + ("*", Some(rest)) => rest.to_owned(), + (name, None) => name.to_owned(), + (_, Some(_)) => return SymbolRef::Call(path.to_owned()), + }; + return SymbolRef::Import { + module: module.clone(), + name, + }; + } + match (self.dialect, rest) { + (BraceDialect::Script, None) => SymbolRef::Local(head.to_owned()), + (BraceDialect::Script, Some(method)) if head == "this" && !method.contains('.') => { + SymbolRef::Local(method.to_owned()) + } + _ => SymbolRef::Call(path.to_owned()), + } + } + + fn call_arguments(&self, open: usize, close: usize, at: usize) -> Vec { + let mut arguments = split_operands(self.tokens, open + 1, close, ',') + .into_iter() + .filter(|(start, end)| start < end) + .map(|(start, end)| { + let value = self.template(start, end, at); + let is_string = value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(_) | UrlPart::Parameter { .. })); + CallArgument { + keyword: None, + value: is_string.then_some(value), + } + }) + .collect::>(); + while arguments + .last() + .is_some_and(|argument| argument.value.is_none()) + { + arguments.pop(); + } + arguments + } +} + +fn is_assignment(tokens: &[Token], index: usize) -> bool { + tokens.get(index).is_some_and(|token| token.is_punct('=')) + && !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('=') || token.is_punct('>')) + && !index.checked_sub(1).is_some_and(|previous| { + matches!( + tokens[previous].kind, + TokenKind::Punct( + '=' | '!' | '<' | '>' | '+' | '-' | '*' | '/' | '%' | '&' | '|' | '^' | '?' + ) + ) + }) +} + +/// End of the expression starting at `start`: a top-level `;` or `,`, a closing bracket of an +/// enclosing group, or a new line once all groups are closed. +pub(super) fn expression_end(tokens: &[Token], start: usize) -> usize { + let mut depth = 0_i32; + let mut previous_line = tokens.get(start).map_or(0, |token| token.end_line); + for (index, token) in tokens.iter().enumerate().skip(start) { + if index > start && depth == 0 && token.line > previous_line { + let continues = matches!(token.kind, TokenKind::Punct('+' | '.' | '?' | ':')) + || matches!( + tokens[index - 1].kind, + TokenKind::Punct('+' | '=' | '(' | ',') + ); + if !continues { + return index; + } + } + match token.kind { + TokenKind::Punct('(' | '[' | '{') => depth += 1, + TokenKind::Punct(')' | ']' | '}') => { + depth -= 1; + if depth < 0 { + return index; + } + } + TokenKind::Punct(';' | ',') if depth == 0 => return index, + _ => {} + } + previous_line = token.end_line; + } + tokens.len() +} + +fn dotted_name(tokens: &[Token], start: usize, end: usize) -> Option { + let mut name = String::new(); + let mut cursor = start; + if tokens.get(cursor)?.is_ident("new") { + name.push_str("new "); + cursor += 1; + } + let mut expect_word = true; + for token in tokens.get(cursor..end)? { + match (expect_word, token.ident()) { + (true, Some(word)) => name.push_str(word), + (false, None) if token.is_punct('.') => name.push('.'), + (false, None) if token.is_punct('?') || token.is_punct('!') => { + expect_word = true; + continue; + } + _ => return None, + } + expect_word = !expect_word; + } + (!expect_word).then_some(name) +} + +/// Dotted callee ending at identifier `index`, when it does not continue another expression. +fn dotted_before(tokens: &[Token], index: usize) -> Option<(usize, String)> { + let mut start = index; + let mut segments = vec![tokens[index].ident()?]; + while start >= 2 && tokens[start - 1].is_punct('.') { + segments.insert(0, tokens[start - 2].ident()?); + start -= 2; + } + if start > 0 && tokens[start - 1].is_punct('.') { + return None; + } + Some((start, segments.join("."))) +} + +/// Body span of a function whose signature ends at `close`, skipping a return type. +fn body_after(tokens: &[Token], close: usize, end: usize) -> Option<(usize, usize)> { + let mut depth = 0_i32; + let mut cursor = close + 1; + while cursor < end.min(tokens.len()) { + let token = &tokens[cursor]; + match token.kind { + TokenKind::Punct('(' | '<' | '[') => depth += 1, + TokenKind::Punct(')' | '>' | ']') => depth -= 1, + TokenKind::Punct('{') if depth == 0 => { + let is_empty_type = tokens + .get(cursor + 1) + .is_some_and(|next| next.is_punct('}')) + && cursor > 0 + && (tokens[cursor - 1].is_ident("interface") + || tokens[cursor - 1].is_ident("struct")); + if !is_empty_type { + return Some((cursor, matching(tokens, cursor, '{', '}')?)); + } + cursor += 1; + } + TokenKind::Punct(';' | '=') if depth == 0 => return None, + _ => {} + } + cursor += 1; + } + None +} + +fn parameter_list( + tokens: &[Token], + open: usize, + close: usize, + dialect: BraceDialect, +) -> Vec { + let parts = split_operands(tokens, open + 1, close, ',') + .into_iter() + .filter(|(start, end)| start < end) + .map(|(start, end)| strip_annotations(&tokens[start..end])) + .collect::>(); + match dialect { + BraceDialect::Script => parts + .iter() + .filter_map(|part| { + part.iter() + .map(|token| token.ident()) + .take_while(Option::is_some) + .flatten() + .find(|word| !DECLARATION_MODIFIERS.contains(word) && *word != "this") + .map(str::to_owned) + }) + .collect(), + BraceDialect::Java => parts + .iter() + .filter_map(|part| { + part.iter() + .rev() + .find_map(|token| token.ident()) + .map(str::to_owned) + }) + .collect(), + BraceDialect::Go => { + let mut names = Vec::new(); + let mut pending = Vec::new(); + for part in &parts { + let Some(name) = part.first().and_then(|token| token.ident()) else { + pending.clear(); + continue; + }; + pending.push(name.to_owned()); + if part.len() > 1 { + names.append(&mut pending); + } + } + names + } + } +} + +/// Tokens of one parameter without `@Annotation(...)` prefixes. +fn strip_annotations(tokens: &[Token]) -> Vec<&Token> { + let mut output = Vec::new(); + let mut index = 0; + while index < tokens.len() { + if tokens[index].is_punct('@') { + index += 2; + if tokens.get(index).is_some_and(|token| token.is_punct('(')) { + let mut depth = 0_i32; + while index < tokens.len() { + if tokens[index].is_punct('(') { + depth += 1; + } else if tokens[index].is_punct(')') { + depth -= 1; + if depth == 0 { + index += 1; + break; + } + } + index += 1; + } + } + continue; + } + output.push(&tokens[index]); + index += 1; + } + output +} + +fn span(name: &str, start: usize, body: (usize, usize)) -> FunctionSpan { + FunctionSpan { + name: name.to_owned(), + start_token: start, + body_start_token: body.0, + end_token: body.1, + } +} + +fn script_functions(tokens: &[Token]) -> Vec<(FunctionSpan, Vec)> { + let mut functions = Vec::new(); + for index in 0..tokens.len() { + if tokens[index].is_ident("function") { + let named = tokens.get(index + 1).and_then(Token::ident); + let open = if named.is_some() { + index + 2 + } else { + index + 1 + }; + let open = if tokens.get(open).is_some_and(|token| token.is_punct('*')) { + open + 1 + } else { + open + }; + let Some(close) = matching(tokens, open, '(', ')') else { + continue; + }; + let Some(name) = named + .map(str::to_owned) + .or_else(|| bound_name(tokens, index)) + else { + continue; + }; + if let Some(body) = body_after(tokens, close, tokens.len()) { + functions.push(( + span(&name, index, body), + parameter_list(tokens, open, close, BraceDialect::Script), + )); + } + } else if tokens[index].is_punct('=') + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('>')) + { + if let Some(function) = script_arrow(tokens, index) { + functions.push(function); + } + } else if let Some(function) = script_method(tokens, index) { + functions.push(function); + } + } + functions +} + +/// Name bound to the function or arrow expression starting at `start`. +fn bound_name(tokens: &[Token], start: usize) -> Option { + let mut cursor = start; + if cursor > 0 && tokens[cursor - 1].is_ident("async") { + cursor -= 1; + } + let previous = tokens.get(cursor.checked_sub(1)?)?; + if previous.is_punct(':') { + return tokens + .get(cursor.checked_sub(2)?)? + .ident() + .map(str::to_owned); + } + if !previous.is_punct('=') { + return None; + } + let declaration = (cursor.saturating_sub(16)..cursor - 1) + .rev() + .take_while(|index| !matches!(tokens[*index].kind, TokenKind::Punct(';' | '{' | '}'))) + .find(|index| { + tokens[*index].is_ident("const") + || tokens[*index].is_ident("let") + || tokens[*index].is_ident("var") + }); + match declaration { + Some(declaration) => tokens.get(declaration + 1)?.ident().map(str::to_owned), + None => tokens + .get(cursor.checked_sub(2)?)? + .ident() + .map(str::to_owned), + } +} + +/// Opening parenthesis of the group closed at `close`. +pub(super) fn matching_open(tokens: &[Token], close: usize) -> Option { + let mut depth = 0_u32; + for index in (0..=close).rev() { + if tokens[index].is_punct(')') { + depth += 1; + } else if tokens[index].is_punct('(') { + depth = depth.checked_sub(1)?; + if depth == 0 { + return Some(index); + } + } + } + None +} + +fn script_arrow(tokens: &[Token], arrow: usize) -> Option<(FunctionSpan, Vec)> { + let mut before = arrow.checked_sub(1)?; + let typed_return = matches!(tokens[before].kind, TokenKind::Punct('>' | ']')) + || before > 0 && tokens[before - 1].is_punct(':'); + if typed_return { + before = (before.saturating_sub(32)..before) + .rev() + .find(|index| tokens[*index].is_punct(')') && tokens[index + 1].is_punct(':'))?; + } + let (start, parameters) = if tokens[before].is_punct(')') { + let open = matching_open(tokens, before)?; + ( + open, + parameter_list(tokens, open, before, BraceDialect::Script), + ) + } else { + (before, vec![tokens[before].ident()?.to_owned()]) + }; + let name = bound_name(tokens, start)?; + let body_start = arrow + 2; + let end = if tokens + .get(body_start) + .is_some_and(|token| token.is_punct('{')) + { + matching(tokens, body_start, '{', '}')? + } else { + expression_end(tokens, body_start).saturating_sub(1) + }; + Some((span(&name, start, (body_start, end)), parameters)) +} + +fn script_method(tokens: &[Token], index: usize) -> Option<(FunctionSpan, Vec)> { + let name = tokens[index].ident()?; + if SCRIPT_KEYWORDS.contains(&name) || !tokens.get(index + 1)?.is_punct('(') { + return None; + } + let previous = index.checked_sub(1).map(|previous| &tokens[previous]); + let declares = previous.is_none_or(|token| { + matches!( + token.kind, + TokenKind::Punct('{' | '}' | ';' | ',' | '*' | ')') + ) || token + .ident() + .is_some_and(|word| DECLARATION_MODIFIERS.contains(&word)) + }); + if !declares { + return None; + } + let close = matching(tokens, index + 1, '(', ')')?; + if !tokens.get(close + 1)?.is_punct('{') && !tokens.get(close + 1)?.is_punct(':') { + return None; + } + let body = body_after(tokens, close, close + 24)?; + Some(( + span(name, index, body), + parameter_list(tokens, index + 1, close, BraceDialect::Script), + )) +} + +fn go_functions(tokens: &[Token]) -> Vec<(FunctionSpan, Vec)> { + let mut functions = Vec::new(); + for index in 0..tokens.len() { + if !tokens[index].is_ident("func") { + continue; + } + let mut cursor = index + 1; + if tokens.get(cursor).is_some_and(|token| token.is_punct('(')) { + let Some(receiver_close) = matching(tokens, cursor, '(', ')') else { + continue; + }; + cursor = receiver_close + 1; + } + let Some(name) = tokens.get(cursor).and_then(Token::ident) else { + continue; + }; + let open = cursor + 1; + let Some(close) = matching(tokens, open, '(', ')') else { + continue; + }; + if let Some(body) = body_after(tokens, close, tokens.len()) { + functions.push(( + span(name, index, body), + parameter_list(tokens, open, close, BraceDialect::Go), + )); + } + } + functions +} + +fn java_functions(tokens: &[Token]) -> Vec<(FunctionSpan, Vec)> { + let mut functions = Vec::new(); + for index in 1..tokens.len() { + let Some(name) = tokens[index].ident() else { + continue; + }; + if JAVA_KEYWORDS.contains(&name) + || !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let previous = &tokens[index - 1]; + if !(previous + .ident() + .is_some_and(|word| !JAVA_KEYWORDS.contains(&word)) + || previous.is_punct('>') + || previous.is_punct(']')) + { + continue; + } + let Some(close) = matching(tokens, index + 1, '(', ')') else { + continue; + }; + let mut cursor = close + 1; + if tokens + .get(cursor) + .is_some_and(|token| token.is_ident("throws")) + { + while tokens + .get(cursor) + .is_some_and(|token| !token.is_punct('{') && !token.is_punct(';')) + { + cursor += 1; + } + } + if !tokens.get(cursor).is_some_and(|token| token.is_punct('{')) { + continue; + } + let Some(end) = matching(tokens, cursor, '{', '}') else { + continue; + }; + functions.push(( + span(name, index, (cursor, end)), + parameter_list(tokens, index + 1, close, BraceDialect::Java), + )); + } + functions +} + +/// Relative ES-module and `CommonJS` imports: local name to `(module, imported name)`; namespace +/// imports use `*` and default imports `default`. +fn script_imports(tokens: &[Token]) -> BTreeMap { + let mut imports = BTreeMap::new(); + for index in 0..tokens.len() { + if tokens[index].is_ident("import") { + let Some(from) = (index + 1..tokens.len().min(index + 64)) + .find(|candidate| tokens[*candidate].is_ident("from")) + else { + continue; + }; + let Some(module) = tokens.get(from + 1).and_then(Token::literal) else { + continue; + }; + if !module.starts_with('.') { + continue; + } + bind_imported_names(&tokens[index + 1..from], module, &mut imports); + } else if tokens[index].is_ident("require") + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + && let Some(module) = tokens.get(index + 2).and_then(Token::literal) + && module.starts_with('.') + && index >= 2 + && tokens[index - 1].is_punct('=') + { + let target = &tokens[..index - 1]; + if let Some(name) = target.last().and_then(Token::ident) { + imports.insert(name.to_owned(), (module.to_owned(), "*".to_owned())); + } else if target.last().is_some_and(|token| token.is_punct('}')) + && let Some(open) = target.iter().rposition(|token| token.is_punct('{')) + { + bind_imported_names(&target[open..], module, &mut imports); + } + } + } + imports +} + +fn bind_imported_names( + clause: &[Token], + module: &str, + imports: &mut BTreeMap, +) { + let mut braced = false; + let mut index = 0; + while index < clause.len() { + let token = &clause[index]; + if token.is_punct('{') { + braced = true; + } else if token.is_punct('}') { + braced = false; + } else if token.is_punct('*') + && clause + .get(index + 1) + .is_some_and(|next| next.is_ident("as")) + && let Some(alias) = clause.get(index + 2).and_then(Token::ident) + { + imports.insert(alias.to_owned(), (module.to_owned(), "*".to_owned())); + index += 2; + } else if let Some(name) = token.ident().filter(|name| *name != "type") { + let (local, skip) = if clause + .get(index + 1) + .is_some_and(|next| next.is_ident("as") || next.is_punct(':')) + && let Some(alias) = clause.get(index + 2).and_then(Token::ident) + { + (alias, 2) + } else { + (name, 0) + }; + let imported = if braced { name } else { "default" }; + imports.insert(local.to_owned(), (module.to_owned(), imported.to_owned())); + index += skip; + } + index += 1; + } +} diff --git a/crates/code-system-graph-core/src/source_http/brace_lexer.rs b/crates/code-system-graph-core/src/source_http/brace_lexer.rs new file mode 100644 index 0000000..4a978f8 --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_lexer.rs @@ -0,0 +1,255 @@ +//! Lexer for the brace languages recognized by client-flow analysis: JavaScript and TypeScript, +//! Go, and Java. +//! +//! Comments are dropped, string literals become literal tokens, JavaScript template literals keep +//! their raw body, and JavaScript regular expressions become opaque literals so that quotes inside +//! them cannot desynchronize the token stream. Single- and double-quoted strings end at a newline, +//! which confines the damage of stray apostrophes such as those in JSX text to one line. + +use super::{Token, TokenKind}; + +/// Source language whose lexical rules apply. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(super) enum BraceDialect { + /// JavaScript and TypeScript. + Script, + /// Go. + Go, + /// Java. + Java, +} + +/// Keywords after which a JavaScript `/` starts a regular expression. +const REGEX_KEYWORDS: [&str; 13] = [ + "return", "typeof", "case", "do", "else", "in", "of", "new", "delete", "void", "throw", + "yield", "await", +]; + +pub(super) fn lex_brace(source: &str, dialect: BraceDialect) -> Vec { + let characters = source.chars().collect::>(); + let mut tokens = Vec::::new(); + let mut index = 0; + let mut line = 1_u32; + let push = |tokens: &mut Vec, kind: TokenKind, start: u32, end: u32| { + tokens.push(Token { + kind, + line: start, + end_line: end, + }); + }; + while index < characters.len() { + let character = characters[index]; + let next = characters.get(index + 1).copied(); + match character { + '\n' => { + line += 1; + index += 1; + } + character if character.is_whitespace() => index += 1, + '/' if next == Some('/') => { + while index < characters.len() && characters[index] != '\n' { + index += 1; + } + } + '/' if next == Some('*') => { + index += 2; + while index < characters.len() + && !(characters[index] == '*' && characters.get(index + 1) == Some(&'/')) + { + if characters[index] == '\n' { + line += 1; + } + index += 1; + } + index += 2; + } + '/' if dialect == BraceDialect::Script && regex_allowed(tokens.last()) => { + index = skip_regex(&characters, index); + push(&mut tokens, TokenKind::Literal(None), line, line); + } + '"' if dialect == BraceDialect::Java + && next == Some('"') + && characters.get(index + 2) == Some(&'"') => + { + let start = line; + let (end, value) = text_block(&characters, index + 3, &mut line); + index = end; + push(&mut tokens, TokenKind::Literal(value), start, line); + } + '"' | '\'' => { + let (end, value) = quoted(&characters, index, character); + index = end; + push(&mut tokens, TokenKind::Literal(value), line, line); + } + '`' if dialect == BraceDialect::Script => { + let start = line; + let (end, body) = template(&characters, index + 1, &mut line); + index = end; + push(&mut tokens, TokenKind::Template(body), start, line); + } + '`' if dialect == BraceDialect::Go => { + let start = line; + let (end, value) = raw_string(&characters, index + 1, &mut line); + index = end; + push(&mut tokens, TokenKind::Literal(Some(value)), start, line); + } + character if character.is_ascii_digit() => { + while index < characters.len() + && (characters[index].is_ascii_alphanumeric() + || matches!(characters[index], '.' | '_')) + { + index += 1; + } + push(&mut tokens, TokenKind::Literal(None), line, line); + } + character if is_word_start(character, dialect) => { + let start = index; + index += 1; + while index < characters.len() && is_word_continue(characters[index], dialect) { + index += 1; + } + let word = characters[start..index].iter().collect(); + push(&mut tokens, TokenKind::Ident(word), line, line); + } + character => { + push(&mut tokens, TokenKind::Punct(character), line, line); + index += 1; + } + } + } + tokens +} + +fn is_word_start(character: char, dialect: BraceDialect) -> bool { + character == '_' + || character.is_alphabetic() + || (character == '$' && dialect != BraceDialect::Go) +} + +fn is_word_continue(character: char, dialect: BraceDialect) -> bool { + character == '_' + || character.is_alphanumeric() + || (character == '$' && dialect != BraceDialect::Go) +} + +fn regex_allowed(previous: Option<&Token>) -> bool { + match previous.map(|token| &token.kind) { + None => true, + Some(TokenKind::Punct(character)) => !matches!(character, ')' | ']' | '}'), + Some(TokenKind::Ident(word)) => REGEX_KEYWORDS.contains(&word.as_str()), + Some(TokenKind::Literal(_) | TokenKind::Template(_)) => false, + } +} + +/// Skips a regular expression literal and its flags; an unterminated one ends at the newline. +fn skip_regex(characters: &[char], start: usize) -> usize { + let mut index = start + 1; + let mut class = false; + while index < characters.len() && characters[index] != '\n' { + match characters[index] { + '\\' => index += 1, + '[' => class = true, + ']' => class = false, + '/' if !class => { + index += 1; + while index < characters.len() && characters[index].is_ascii_alphabetic() { + index += 1; + } + return index; + } + _ => {} + } + index += 1; + } + index +} + +/// Single-line quoted string; `None` when it is unterminated on its line. +fn quoted(characters: &[char], start: usize, quote: char) -> (usize, Option) { + let mut value = String::new(); + let mut index = start + 1; + while index < characters.len() { + match characters[index] { + '\n' => return (index, None), + '\\' => { + if let Some(escaped) = characters.get(index + 1) { + value.push(match escaped { + 'n' => '\n', + 't' => '\t', + 'r' => '\r', + other => *other, + }); + } + index += 2; + } + character if character == quote => return (index + 1, Some(value)), + character => { + value.push(character); + index += 1; + } + } + } + (index, None) +} + +/// Go raw string body up to the closing backtick. +fn raw_string(characters: &[char], start: usize, line: &mut u32) -> (usize, String) { + let mut index = start; + while index < characters.len() && characters[index] != '`' { + if characters[index] == '\n' { + *line += 1; + } + index += 1; + } + (index + 1, characters[start..index].iter().collect()) +} + +fn text_block(characters: &[char], start: usize, line: &mut u32) -> (usize, Option) { + let mut index = start; + while index + 2 < characters.len() { + if characters[index] == '"' && characters[index + 1] == '"' && characters[index + 2] == '"' + { + let value = characters[start..index].iter().collect::(); + return (index + 3, Some(value.trim().to_owned())); + } + if characters[index] == '\n' { + *line += 1; + } + index += 1; + } + (characters.len(), None) +} + +/// Raw template body up to the closing backtick, skipping nested `${...}` expressions. +fn template(characters: &[char], start: usize, line: &mut u32) -> (usize, String) { + let mut index = start; + let mut depth = 0_u32; + while index < characters.len() { + match characters[index] { + '\n' => *line += 1, + '\\' => index += 1, + '$' if depth == 0 && characters.get(index + 1) == Some(&'{') => { + depth = 1; + index += 1; + } + '{' if depth > 0 => depth += 1, + '}' if depth > 0 => depth -= 1, + '`' if depth == 0 => { + return (index + 1, characters[start..index].iter().collect()); + } + '`' => { + let (end, _) = template(characters, index + 1, line); + index = end; + continue; + } + quote @ ('"' | '\'') if depth > 0 => { + let (end, _) = quoted(characters, index, quote); + index = end; + continue; + } + _ => {} + } + index += 1; + } + (index, characters[start..].iter().collect()) +} diff --git a/crates/code-system-graph-core/src/source_http/brace_tests.rs b/crates/code-system-graph-core/src/source_http/brace_tests.rs new file mode 100644 index 0000000..e47b84d --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/brace_tests.rs @@ -0,0 +1,324 @@ +//! Tests of JavaScript and TypeScript, Go, and Java test frameworks. +//! +//! Script tests have no named function: `it` and `test` callbacks are named by their `describe` +//! chain and title, and `beforeEach`-style callbacks by their `describe` chain, so each test keeps +//! a stable identity and calls the setup blocks that run before it. Go tests are `TestX` +//! functions taking `*testing.T`, and Java tests are methods annotated with `@Test`. + +use super::brace_flows::{BraceScopes, matching_open}; +use super::{ + FunctionSpan, SourceObservationCollector, Token, TokenKind, call_observation, confirmed_test, matching, split_operands +}; +use crate::{SourceFramework, SourceLanguage, SourceLineRange, SymbolRef}; + +const SUITES: [&str; 3] = ["describe", "context", "suite"]; +const TESTS: [&str; 3] = ["it", "test", "specify"]; +const HOOKS: [&str; 3] = ["beforeEach", "beforeAll", "before"]; +const SUITE_MODIFIERS: [&str; 5] = ["only", "skip", "serial", "parallel", "concurrent"]; +const TEST_MODIFIERS: [&str; 7] = [ + "only", + "skip", + "concurrent", + "fails", + "fixme", + "fail", + "slow", +]; +const JAVA_TEST_ANNOTATIONS: [&str; 4] = + ["Test", "ParameterizedTest", "RepeatedTest", "TestFactory"]; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum BlockKind { + Suite, + Test, + Hook, +} + +/// A `describe`, `it`/`test`, or setup hook call with its callback. +pub(super) struct ScriptBlock { + kind: BlockKind, + title: String, + name: String, + start: usize, + body: (usize, usize), + parameters: Vec, + /// Tokens of the innermost enclosing suite body, or of the whole file. + scope: (usize, usize), +} + +impl ScriptBlock { + /// Function span of a test or hook callback. + pub(super) fn function(&self) -> Option<(FunctionSpan, Vec)> { + (self.kind != BlockKind::Suite).then(|| { + ( + FunctionSpan { + name: self.name.clone(), + start_token: self.start, + body_start_token: self.body.0, + end_token: self.body.1, + }, + self.parameters.clone(), + ) + }) + } +} + +/// Test blocks of one script file, named by their `describe` chain. +pub(super) fn script_blocks(tokens: &[Token]) -> Vec { + let mut blocks = (0..tokens.len()) + .filter_map(|index| script_block(tokens, index)) + .collect::>(); + let suites = blocks + .iter() + .filter(|block| block.kind == BlockKind::Suite) + .map(|block| (block.body, block.title.clone())) + .collect::>(); + for block in &mut blocks { + let enclosing = suites + .iter() + .filter(|((start, end), _)| *start < block.start && block.start < *end) + .collect::>(); + if let Some(((start, end), _)) = enclosing.last() { + block.scope = (*start, *end); + } + block.name = enclosing + .iter() + .map(|(_, title)| title.as_str()) + .chain([block.title.as_str()]) + .collect::>() + .join(" > "); + } + blocks +} + +fn script_block(tokens: &[Token], index: usize) -> Option { + let head = tokens[index].ident()?; + if index > 0 && tokens[index - 1].is_punct('.') { + return None; + } + let mut open = index + 1; + let mut modifier = None; + if tokens.get(open)?.is_punct('.') { + modifier = Some(tokens.get(open + 1)?.ident()?); + open += 2; + } + if !tokens.get(open)?.is_punct('(') { + return None; + } + let kind = match (head, modifier) { + (head, None) if SUITES.contains(&head) => BlockKind::Suite, + (head, Some(modifier)) if SUITES.contains(&head) && SUITE_MODIFIERS.contains(&modifier) => { + BlockKind::Suite + } + ("test", Some("describe")) => BlockKind::Suite, + (head, None) if HOOKS.contains(&head) => BlockKind::Hook, + ("test", Some(modifier)) if HOOKS.contains(&modifier) => BlockKind::Hook, + (head, None) if TESTS.contains(&head) => BlockKind::Test, + (head, Some(modifier)) if TESTS.contains(&head) && TEST_MODIFIERS.contains(&modifier) => { + BlockKind::Test + } + _ => return None, + }; + let close = matching(tokens, open, '(', ')')?; + let arguments = split_operands(tokens, open + 1, close, ','); + let title = match kind { + BlockKind::Hook => modifier.unwrap_or(head).to_owned(), + BlockKind::Suite | BlockKind::Test => { + let (start, end) = *arguments.first()?; + static_title(tokens, start, end)? + } + }; + let (callback_start, callback_end) = *arguments.last()?; + let (body, parameters) = callback(tokens, callback_start, callback_end)?; + Some(ScriptBlock { + kind, + title: title.clone(), + name: title, + start: index, + body, + parameters, + scope: (0, tokens.len()), + }) +} + +fn static_title(tokens: &[Token], start: usize, end: usize) -> Option { + if end != start + 1 { + return None; + } + match &tokens[start].kind { + TokenKind::Literal(Some(value)) => Some(value.clone()), + TokenKind::Template(body) if !body.contains("${") => Some(body.clone()), + _ => None, + } +} + +/// Body and parameters of the function or arrow expression in `start..end`. +fn callback(tokens: &[Token], start: usize, end: usize) -> Option<((usize, usize), Vec)> { + let mut cursor = start; + if tokens.get(cursor)?.is_ident("async") { + cursor += 1; + } + if tokens.get(cursor)?.is_ident("function") { + cursor += 1; + if tokens.get(cursor)?.ident().is_some() { + cursor += 1; + } + let close = matching(tokens, cursor, '(', ')')?; + let body = (close + 1..end).find(|index| tokens[*index].is_punct('{'))?; + return Some(( + (body, matching(tokens, body, '{', '}')?), + parameter_names(tokens, cursor, close), + )); + } + let (parameters, arrow) = if tokens.get(cursor)?.is_punct('(') { + let close = matching(tokens, cursor, '(', ')')?; + let arrow = (close + 1..end).find(|index| { + tokens[*index].is_punct('=') + && tokens.get(index + 1).is_some_and(|next| next.is_punct('>')) + })?; + (parameter_names(tokens, cursor, close), arrow) + } else { + let name = tokens.get(cursor)?.ident()?; + if !tokens.get(cursor + 1)?.is_punct('=') || !tokens.get(cursor + 2)?.is_punct('>') { + return None; + } + (vec![name.to_owned()], cursor + 1) + }; + let body_start = arrow + 2; + if body_start >= end { + return None; + } + let body_end = if tokens[body_start].is_punct('{') { + matching(tokens, body_start, '{', '}')? + } else { + end - 1 + }; + Some(((body_start, body_end), parameters)) +} + +fn parameter_names(tokens: &[Token], open: usize, close: usize) -> Vec { + split_operands(tokens, open + 1, close, ',') + .into_iter() + .filter(|(start, end)| start < end) + .filter_map(|(start, _)| tokens[start].ident().map(str::to_owned)) + .collect() +} + +/// Framework of a script test file, from the test runner it imports; Jest otherwise. +fn script_framework(tokens: &[Token]) -> SourceFramework { + let imports = |module: &str| tokens.iter().any(|token| token.literal() == Some(module)); + if imports("vitest") { + SourceFramework::Vitest + } else if imports("@playwright/test") { + SourceFramework::Playwright + } else if imports("mocha") || imports("chai") { + SourceFramework::Mocha + } else { + SourceFramework::Jest + } +} + +/// Records script tests and a call from each test to every setup hook that runs before it. +pub(super) fn record_script_tests( + tokens: &[Token], + blocks: &[ScriptBlock], + language: SourceLanguage, + observations: &mut SourceObservationCollector<'_>, +) { + let framework = script_framework(tokens); + for test in blocks.iter().filter(|block| block.kind == BlockKind::Test) { + let lines = SourceLineRange { + start: tokens[test.start].line, + end: tokens[test.body.0].end_line, + }; + observations.push(confirmed_test( + language, + framework, + test.name.clone(), + lines.start, + lines.end, + )); + for hook in blocks.iter().filter(|block| { + block.kind == BlockKind::Hook + && block.scope.0 <= test.start + && test.start < block.scope.1 + }) { + observations.push(call_observation( + language, + SymbolRef::Local(hook.name.clone()), + Vec::new(), + test.name.clone(), + lines, + )); + } + } +} + +/// Records Go `TestX(t *testing.T)` functions. +pub(super) fn record_go_tests( + scopes: &BraceScopes<'_>, + observations: &mut SourceObservationCollector<'_>, +) { + let tokens = scopes.tokens; + for function in scopes.functions() { + let signature = &tokens[function.start_token..function.body_start_token]; + let takes_testing = signature.windows(3).any(|window| { + window[0].is_ident("testing") && window[1].is_punct('.') && window[2].is_ident("T") + }); + let named = function.name.strip_prefix("Test").is_some_and(|rest| { + rest.chars() + .next() + .is_none_or(|first| !first.is_lowercase()) + }); + if !named || !takes_testing { + continue; + } + observations.push(confirmed_test( + SourceLanguage::Go, + SourceFramework::GoTest, + function.name.clone(), + tokens[function.start_token].line, + tokens[function.body_start_token].end_line, + )); + } +} + +/// Records Java methods annotated with `@Test` or another `JUnit` test annotation. +pub(super) fn record_java_tests( + scopes: &BraceScopes<'_>, + observations: &mut SourceObservationCollector<'_>, +) { + let tokens = scopes.tokens; + for function in scopes.functions() { + let mut cursor = function.start_token; + let mut annotated = false; + while cursor > 1 && !annotated { + cursor -= 1; + let token = &tokens[cursor]; + if token.is_punct(')') { + let Some(open) = matching_open(tokens, cursor) else { + break; + }; + cursor = open; + continue; + } + if matches!(token.kind, TokenKind::Punct(';' | '{' | '}')) { + break; + } + annotated = tokens[cursor - 1].is_punct('@') + && token + .ident() + .is_some_and(|name| JAVA_TEST_ANNOTATIONS.contains(&name)); + } + if !annotated { + continue; + } + observations.push(confirmed_test( + SourceLanguage::Java, + SourceFramework::JUnit, + function.name.clone(), + tokens[function.start_token].line, + tokens[function.body_start_token].end_line, + )); + } +} diff --git a/crates/code-system-graph-core/src/source_http/python_flows.rs b/crates/code-system-graph-core/src/source_http/python_flows.rs new file mode 100644 index 0000000..e1417bb --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/python_flows.rs @@ -0,0 +1,551 @@ +//! Python URL expressions, scoped string constants, client wrappers, and call sites. +//! +//! URL arguments are evaluated through literals, implicit and `+` concatenation, f-strings, +//! `str.format`, `%` formatting, and names assigned earlier in the same function or at module +//! level. Calls are recorded so repository composition can instantiate client wrappers at call +//! sites and attribute helper and fixture requests to the tests that use them. + +use std::collections::{BTreeMap, BTreeSet}; + +use super::{ + FunctionSpan, SourceObservationCollector, Token, TokenKind, call_observation, innermost_function, matching, split_operands, top_level_arguments +}; +use crate::url_template::{brace_format, printf_format}; +use crate::{CallArgument, SourceLanguage, SourceLineRange, SymbolRef, UrlPart, UrlTemplate}; + +const KEYWORDS: [&str; 30] = [ + "and", "as", "assert", "async", "await", "class", "def", "del", "elif", "else", "except", + "for", "from", "global", "if", "import", "in", "is", "lambda", "nonlocal", "not", "or", "pass", + "raise", "return", "while", "with", "yield", "print", "super", +]; + +const BUILTINS: [&str; 34] = [ + "abs", + "all", + "any", + "bool", + "bytes", + "dict", + "enumerate", + "filter", + "float", + "format", + "getattr", + "hasattr", + "hash", + "id", + "int", + "isinstance", + "issubclass", + "iter", + "len", + "list", + "map", + "max", + "min", + "next", + "open", + "range", + "repr", + "round", + "set", + "setattr", + "sorted", + "str", + "sum", + "tuple", +]; + +const STRING_PREFIXES: [&str; 12] = [ + "f", "F", "rf", "fr", "Rf", "fR", "RF", "FR", "rF", "Fr", "b", "u", +]; + +struct Assignment { + function: Option, + index: usize, + name: String, + value: UrlTemplate, +} + +/// Function parameters and string assignments of one Python file. +pub(super) struct PythonScopes<'a> { + tokens: &'a [Token], + functions: &'a [FunctionSpan], + parameters: Vec>, + assignments: Vec, +} + +impl<'a> PythonScopes<'a> { + pub(super) fn new(tokens: &'a [Token], functions: &'a [FunctionSpan]) -> Self { + let parameters = functions + .iter() + .map(|function| python_parameters(tokens, function)) + .collect(); + let mut scopes = Self { + tokens, + functions, + parameters, + assignments: Vec::new(), + }; + for index in 0..tokens.len() { + if index > 0 && tokens[index - 1].line == tokens[index].line { + continue; + } + let Some((name, value_start)) = assignment_target(tokens, index) else { + continue; + }; + let end = statement_end(tokens, value_start); + let value = scopes.template(value_start, end, index); + scopes.assignments.push(Assignment { + function: scopes.innermost(index), + index, + name, + value, + }); + } + scopes + } + + /// Caller-visible parameters of the function named `name`, receivers excluded. + pub(super) fn parameters_of(&self, function: usize) -> &[String] { + &self.parameters[function] + } + + pub(super) fn innermost(&self, index: usize) -> Option { + innermost_function(self.functions, index) + } + + /// Evaluates the expression in `start..end`, resolving names as seen at token `at`. + pub(super) fn template(&self, start: usize, end: usize, at: usize) -> UrlTemplate { + let mut template = UrlTemplate::default(); + for (operand_start, operand_end) in split_operands(self.tokens, start, end, '+') { + template.extend(self.operand(operand_start, operand_end, at)); + } + template + } + + fn operand(&self, start: usize, end: usize, at: usize) -> UrlTemplate { + let tokens = self.tokens; + let anonymous = || UrlTemplate::part(UrlPart::Value(None)); + if start >= end { + return anonymous(); + } + if let [(left_start, left_end), (right_start, right_end)] = + split_operands(tokens, start, end, '%').as_slice() + && let Some(format) = self.literal_run(*left_start, *left_end, at) + && let Some(format) = format.as_literal() + { + let arguments = if tokens[*right_start].is_punct('(') + && matching(tokens, *right_start, '(', ')') == Some(right_end - 1) + { + split_operands(tokens, right_start + 1, right_end - 1, ',') + .into_iter() + .map(|(start, end)| self.template(start, end, at)) + .collect() + } else { + vec![self.template(*right_start, *right_end, at)] + }; + return printf_format(format, &arguments); + } + if let Some(template) = self.literal_run(start, end, at) { + return template; + } + if let Some(format_end) = self.format_call(start, end, at) { + return format_end; + } + if tokens[start].is_punct('(') && matching(tokens, start, '(', ')') == Some(end - 1) { + return self.template(start + 1, end - 1, at); + } + match dotted_name(tokens, start, end) { + Some(name) => self.resolve(&name, at), + None => anonymous(), + } + } + + /// Adjacent string literals, including f-strings, that span all of `start..end`. + fn literal_run(&self, start: usize, end: usize, at: usize) -> Option { + let tokens = self.tokens; + let mut template = UrlTemplate::default(); + let mut cursor = start; + while cursor < end { + let formatted = tokens[cursor] + .ident() + .filter(|prefix| STRING_PREFIXES.contains(prefix)) + .filter(|_| { + tokens + .get(cursor + 1) + .is_some_and(|next| next.literal().is_some()) + }); + if let Some(prefix) = formatted { + let literal = tokens[cursor + 1].literal().unwrap_or_default(); + if prefix.contains(['f', 'F']) { + template.extend(brace_format(literal, &[], &BTreeMap::new(), &|name| { + self.inline(name, at) + })); + } else { + template.push_text(literal); + } + cursor += 2; + } else { + template.push_text(tokens[cursor].literal()?); + cursor += 1; + } + } + Some(template) + } + + /// `"...".format(...)` spanning all of `start..end`. + fn format_call(&self, start: usize, end: usize, at: usize) -> Option { + let tokens = self.tokens; + let dot = (start..end).find(|index| tokens[*index].is_punct('.'))?; + let format = self.literal_run(start, dot, at)?; + let format = format.as_literal()?; + if !tokens + .get(dot + 1) + .is_some_and(|token| token.is_ident("format")) + { + return None; + } + let open = dot + 2; + let close = matching(tokens, open, '(', ')')?; + if close + 1 != end { + return None; + } + let mut positional = Vec::new(); + let mut named = BTreeMap::new(); + for (argument_start, argument_end) in split_operands(tokens, open + 1, close, ',') { + match keyword_argument(tokens, argument_start, argument_end) { + Some((keyword, value_start)) => { + named.insert(keyword, self.template(value_start, argument_end, at)); + } + None => positional.push(self.template(argument_start, argument_end, at)), + } + } + Some(brace_format(format, &positional, &named, &|name| { + self.inline(name, at) + })) + } + + fn inline(&self, expression: &str, at: usize) -> UrlTemplate { + let expression = expression.trim(); + let is_dotted = expression.split('.').all(|segment| { + !segment.is_empty() + && segment + .chars() + .all(|character| character.is_ascii_alphanumeric() || character == '_') + }); + if is_dotted { + self.resolve(expression, at) + } else { + UrlTemplate::part(UrlPart::Value(None)) + } + } + + /// Resolves a plain or dotted name at token `at`. + pub(super) fn resolve(&self, name: &str, at: usize) -> UrlTemplate { + let function = self.innermost(at); + if let Some(function) = function + && let Some(index) = self.parameters[function] + .iter() + .position(|parameter| parameter == name) + { + return UrlTemplate::part(UrlPart::Parameter { + name: name.to_owned(), + index, + }); + } + let local = self + .assignments + .iter() + .rev() + .filter(|assignment| assignment.name == name && assignment.index < at) + .find(|assignment| function.is_some() && assignment.function == function); + let module = || { + self.assignments + .iter() + .rev() + .find(|assignment| assignment.name == name && assignment.function.is_none()) + }; + let attribute = || { + name.starts_with("self.") + .then(|| { + let mut values = self + .assignments + .iter() + .filter(|assignment| assignment.name == name) + .map(|assignment| &assignment.value); + let first = values.next()?; + values.all(|value| value == first).then_some(first) + }) + .flatten() + }; + local + .or_else(module) + .map(|assignment| &assignment.value) + .or_else(attribute) + .filter(|value| !value.has_parameters() || local.is_some()) + .cloned() + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(Some(name.to_owned())))) + } + + /// Records calls from functions, with their URL-like arguments. + /// + /// With `all_calls`, every call is recorded; otherwise only calls that pass a path literal or + /// forward a parameter of the caller. + pub(super) fn record_calls( + &self, + imports: &BTreeMap, + clients: &BTreeSet<&str>, + all_calls: bool, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens; + for index in 0..tokens.len() { + if !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let Some(name) = tokens[index].ident() else { + continue; + }; + if KEYWORDS.contains(&name) || BUILTINS.contains(&name) { + continue; + } + let Some((start, dotted)) = dotted_before(tokens, index) else { + continue; + }; + if start > 0 + && (tokens[start - 1].is_ident("def") + || tokens[start - 1].is_ident("class") + || tokens[start - 1].is_punct('@')) + { + continue; + } + if clients.contains(dotted[0].as_str()) { + continue; + } + let Some(function) = self.innermost(index) else { + continue; + }; + let Some(close) = matching(tokens, index + 1, '(', ')') else { + continue; + }; + let arguments = self.call_arguments(index + 1, close, index); + let forwards_url = arguments.iter().any(|argument| { + argument.value.as_ref().is_some_and(|value| { + value.has_parameters() + || value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(text) if text.contains('/'))) + }) + }); + if !all_calls && !forwards_url { + continue; + } + observations.push(call_observation( + SourceLanguage::Python, + python_callee(&dotted, imports), + arguments, + self.functions[function].name.clone(), + SourceLineRange { + start: tokens[start].line, + end: tokens[close].end_line, + }, + )); + } + } + + fn call_arguments(&self, open: usize, close: usize, at: usize) -> Vec { + let tokens = self.tokens; + let mut arguments = split_operands(tokens, open + 1, close, ',') + .into_iter() + .filter(|(start, end)| start < end) + .map(|(start, end)| { + let (keyword, value_start) = keyword_argument(tokens, start, end) + .map_or((None, start), |(keyword, value)| (Some(keyword), value)); + let value = self.template(value_start, end, at); + let is_string = value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(_) | UrlPart::Parameter { .. })); + CallArgument { + keyword, + value: is_string.then_some(value), + } + }) + .collect::>(); + while arguments + .last() + .is_some_and(|argument| argument.value.is_none()) + { + arguments.pop(); + } + arguments + } +} + +/// Records pytest fixture requests of tests and fixtures as calls to the requested fixtures. +pub(super) fn record_fixture_requests( + scopes: &PythonScopes<'_>, + observations: &mut SourceObservationCollector<'_>, +) { + let tokens = scopes.tokens; + for (position, function) in scopes.functions.iter().enumerate() { + let is_test = function.name.starts_with("test_"); + let is_fixture = + (function.start_token.saturating_sub(12)..function.start_token).any(|index| { + tokens[index].is_ident("fixture") + && tokens[..index] + .iter() + .rev() + .take(3) + .any(|token| token.is_punct('@')) + }); + if !is_test && !is_fixture { + continue; + } + for parameter in scopes.parameters_of(position) { + observations.push(call_observation( + SourceLanguage::Python, + SymbolRef::Fixture(parameter.clone()), + Vec::new(), + function.name.clone(), + SourceLineRange { + start: tokens[function.start_token].line, + end: tokens[function.start_token].end_line, + }, + )); + } + } +} + +fn python_callee(dotted: &[String], imports: &BTreeMap) -> SymbolRef { + let head = &dotted[0]; + let rest = &dotted[1..]; + if let Some((module, name)) = imports.get(head) { + let name = std::iter::once(name.as_str()) + .filter(|name| !name.is_empty()) + .chain(rest.iter().map(String::as_str)) + .collect::>() + .join("."); + if !name.is_empty() { + return SymbolRef::Import { + module: module.clone(), + name, + }; + } + } + match (head.as_str(), rest) { + (_, []) => SymbolRef::Local(head.clone()), + ("self" | "cls", [method]) => SymbolRef::Local(method.clone()), + _ => SymbolRef::Call(dotted.join(".")), + } +} + +/// Parameter names a caller binds, without `self`/`cls` receivers or `*`/`**` markers. +fn python_parameters(tokens: &[Token], function: &FunctionSpan) -> Vec { + let open = function.start_token + 2; + let Some(close) = matching(tokens, open, '(', ')') else { + return Vec::new(); + }; + let mut parameters = Vec::new(); + for (position, start) in top_level_arguments(tokens, open, close) + .into_iter() + .enumerate() + { + let Some(name) = tokens[start..close] + .iter() + .take_while(|token| !token.is_punct(':') && !token.is_punct('=')) + .find_map(Token::ident) + else { + continue; + }; + if position == 0 && matches!(name, "self" | "cls") { + continue; + } + parameters.push(name.to_owned()); + } + parameters +} + +/// Target and value start of a statement such as `NAME = ...` or `self.NAME = ...`. +fn assignment_target(tokens: &[Token], index: usize) -> Option<(String, usize)> { + let head = tokens[index].ident()?; + let (name, equals) = if head == "self" + && tokens.get(index + 1)?.is_punct('.') + && let Some(attribute) = tokens.get(index + 2)?.ident() + { + (format!("self.{attribute}"), index + 3) + } else { + (head.to_owned(), index + 1) + }; + let is_assignment = tokens.get(equals)?.is_punct('=') + && !tokens + .get(equals + 1) + .is_some_and(|token| token.is_punct('=')); + (is_assignment && !KEYWORDS.contains(&head)).then_some((name, equals + 1)) +} + +fn statement_end(tokens: &[Token], start: usize) -> usize { + let mut depth = 0_i32; + let mut previous_line = tokens.get(start).map_or(0, |token| token.end_line); + for (index, token) in tokens.iter().enumerate().skip(start) { + if index > start && depth == 0 && token.line > previous_line { + return index; + } + match token.kind { + TokenKind::Punct('(' | '[' | '{') => depth += 1, + TokenKind::Punct(')' | ']' | '}') => depth -= 1, + TokenKind::Punct(';') if depth == 0 => return index, + _ => {} + } + previous_line = token.end_line; + } + tokens.len() +} + +pub(super) fn keyword_argument( + tokens: &[Token], + start: usize, + end: usize, +) -> Option<(String, usize)> { + let keyword = tokens.get(start)?.ident()?; + (start + 2 <= end + && tokens.get(start + 1)?.is_punct('=') + && !tokens + .get(start + 2) + .is_some_and(|token| token.is_punct('='))) + .then(|| (keyword.to_owned(), start + 2)) +} + +fn dotted_name(tokens: &[Token], start: usize, end: usize) -> Option { + let mut name = String::new(); + for (offset, token) in tokens.get(start..end)?.iter().enumerate() { + if offset % 2 == 0 { + name.push_str(token.ident()?); + } else if token.is_punct('.') { + name.push('.'); + } else { + return None; + } + } + (!name.is_empty() && !name.ends_with('.')).then_some(name) +} + +/// Dotted callee ending at identifier `index`, when it does not continue another expression. +fn dotted_before(tokens: &[Token], index: usize) -> Option<(usize, Vec)> { + let mut start = index; + let mut segments = vec![tokens[index].ident()?.to_owned()]; + while start >= 2 && tokens[start - 1].is_punct('.') { + let previous = tokens[start - 2].ident()?; + segments.insert(0, previous.to_owned()); + start -= 2; + } + if start > 0 && tokens[start - 1].is_punct('.') { + return None; + } + Some((start, segments)) +} diff --git a/crates/code-system-graph-core/src/source_http/rust_flows.rs b/crates/code-system-graph-core/src/source_http/rust_flows.rs new file mode 100644 index 0000000..0418514 --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/rust_flows.rs @@ -0,0 +1,399 @@ +//! Rust URL expressions, string constants, client wrappers, and call sites. +//! +//! URL arguments are evaluated through literals, `format!`, `concat!`, `+` concatenation, +//! `const`/`static` items, and `let` bindings earlier in the same function. Calls are recorded so +//! repository composition can instantiate client wrappers and attribute helper requests to tests. + +use std::collections::{BTreeMap, BTreeSet}; + +use super::{ + FunctionSpan, SourceObservationCollector, Token, TokenKind, call_observation, innermost_function, matching, split_operands +}; +use crate::url_template::brace_format; +use crate::{CallArgument, SourceLanguage, SourceLineRange, SymbolRef, UrlPart, UrlTemplate}; + +const KEYWORDS: [&str; 24] = [ + "if", "while", "match", "for", "loop", "return", "fn", "let", "in", "as", "move", "async", + "await", "Some", "Ok", "Err", "Box", "Vec", "String", "Self", "self", "unsafe", "where", + "impl", +]; + +/// Conversions that leave a string value unchanged. +const IDENTITY_METHODS: [&str; 6] = ["to_string", "to_owned", "as_str", "into", "as_ref", "clone"]; + +struct Binding { + function: usize, + index: usize, + name: String, + value: UrlTemplate, +} + +/// Function parameters, string constants, and `let` bindings of one Rust file. +pub(super) struct RustScopes<'a> { + tokens: &'a [Token], + functions: &'a [FunctionSpan], + parameters: Vec>, + constants: BTreeMap, + bindings: Vec, +} + +impl<'a> RustScopes<'a> { + pub(super) fn new(tokens: &'a [Token], functions: &'a [FunctionSpan]) -> Self { + let mut scopes = Self { + tokens, + functions, + parameters: functions + .iter() + .map(|function| rust_parameters(tokens, function)) + .collect(), + constants: BTreeMap::new(), + bindings: Vec::new(), + }; + for index in 0..tokens.len() { + let is_item = tokens[index].is_ident("const") || tokens[index].is_ident("static"); + let is_let = tokens[index].is_ident("let"); + if !is_item && !is_let { + continue; + } + let mut name_index = index + 1; + if tokens + .get(name_index) + .is_some_and(|token| token.is_ident("mut")) + { + name_index += 1; + } + let Some(name) = tokens.get(name_index).and_then(Token::ident) else { + continue; + }; + let Some(equals) = (name_index + 1..tokens.len()) + .take_while(|candidate| !tokens[*candidate].is_punct(';')) + .find(|candidate| { + tokens[*candidate].is_punct('=') + && !tokens + .get(candidate + 1) + .is_some_and(|token| token.is_punct('=') || token.is_punct('>')) + }) + else { + continue; + }; + let end = statement_end(tokens, equals + 1); + let value = scopes.template(equals + 1, end, index); + match (is_item, scopes.innermost(index)) { + (true, _) => { + scopes.constants.insert(name.to_owned(), value); + } + (false, Some(function)) => scopes.bindings.push(Binding { + function, + index, + name: name.to_owned(), + value, + }), + (false, None) => {} + } + } + scopes + } + + pub(super) fn innermost(&self, index: usize) -> Option { + innermost_function(self.functions, index) + } + + /// Evaluates the expression in `start..end`, resolving names as seen at token `at`. + pub(super) fn template(&self, start: usize, end: usize, at: usize) -> UrlTemplate { + let mut template = UrlTemplate::default(); + for (operand_start, operand_end) in split_operands(self.tokens, start, end, '+') { + template.extend(self.operand(operand_start, operand_end, at)); + } + template + } + + fn operand(&self, mut start: usize, mut end: usize, at: usize) -> UrlTemplate { + let tokens = self.tokens; + while start < end && tokens[start].is_punct('&') { + start += 1; + } + while end >= start + 4 + && tokens[end - 1].is_punct(')') + && tokens[end - 2].is_punct('(') + && tokens[end - 3] + .ident() + .is_some_and(|method| IDENTITY_METHODS.contains(&method)) + && tokens[end - 4].is_punct('.') + { + end -= 4; + } + if start >= end { + return UrlTemplate::part(UrlPart::Value(None)); + } + if end == start + 1 + && let Some(literal) = tokens[start].literal() + { + return UrlTemplate::text(literal); + } + if let Some(macro_name) = tokens[start].ident() + && tokens + .get(start + 1) + .is_some_and(|token| token.is_punct('!')) + && matching(tokens, start + 2, '(', ')') == Some(end - 1) + { + return self.macro_call(macro_name, start + 2, end - 1, at); + } + if tokens[start].is_ident("String") + && end > start + 5 + && tokens[start + 3].is_ident("from") + && matching(tokens, start + 4, '(', ')') == Some(end - 1) + { + return self.template(start + 5, end - 1, at); + } + if tokens[start].is_punct('(') && matching(tokens, start, '(', ')') == Some(end - 1) { + return self.template(start + 1, end - 1, at); + } + match path_name(tokens, start, end) { + Some(name) => self.resolve(&name, at), + None => UrlTemplate::part(UrlPart::Value(None)), + } + } + + fn macro_call(&self, name: &str, open: usize, close: usize, at: usize) -> UrlTemplate { + let arguments = split_operands(self.tokens, open + 1, close, ','); + match name { + "format" => { + let Some(format) = arguments + .first() + .and_then(|(start, _)| self.tokens[*start].literal()) + else { + return UrlTemplate::part(UrlPart::Value(None)); + }; + let mut positional = Vec::new(); + let mut named = BTreeMap::new(); + for &(start, end) in &arguments[1..] { + if let Some(keyword) = self.tokens.get(start).and_then(Token::ident) + && self + .tokens + .get(start + 1) + .is_some_and(|token| token.is_punct('=')) + { + named.insert(keyword.to_owned(), self.template(start + 2, end, at)); + } else { + positional.push(self.template(start, end, at)); + } + } + brace_format(format, &positional, &named, &|name| self.resolve(name, at)) + } + "concat" => { + let mut joined = UrlTemplate::default(); + for (start, end) in arguments { + joined.extend(self.template(start, end, at)); + } + joined + } + _ => UrlTemplate::part(UrlPart::Value(None)), + } + } + + fn resolve(&self, name: &str, at: usize) -> UrlTemplate { + let function = self.innermost(at); + if let Some(function) = function + && let Some(index) = self.parameters[function] + .iter() + .position(|parameter| parameter == name) + { + return UrlTemplate::part(UrlPart::Parameter { + name: name.to_owned(), + index, + }); + } + let binding = self.bindings.iter().rev().find(|binding| { + Some(binding.function) == function && binding.index < at && binding.name == name + }); + let constant = || { + let last = name.rsplit(':').next().unwrap_or(name); + self.constants + .get(name) + .or_else(|| self.constants.get(last)) + }; + binding + .map(|binding| &binding.value) + .or_else(constant) + .cloned() + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(Some(name.to_owned())))) + } + + /// Records calls from functions; see the Python counterpart for the selection rule. + pub(super) fn record_calls( + &self, + clients: &BTreeSet, + all_calls: bool, + observations: &mut SourceObservationCollector<'_>, + ) { + let tokens = self.tokens; + for index in 0..tokens.len() { + if !tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + { + continue; + } + let Some(name) = tokens[index].ident() else { + continue; + }; + if KEYWORDS.contains(&name) { + continue; + } + let Some((start, path)) = path_before(tokens, index) else { + continue; + }; + if start > 0 && tokens[start - 1].is_ident("fn") + || clients.contains(&start) + || path.starts_with("reqwest::") + { + continue; + } + let Some(function) = self.innermost(index) else { + continue; + }; + let Some(close) = matching(tokens, index + 1, '(', ')') else { + continue; + }; + let arguments = self.call_arguments(index + 1, close, index); + let forwards_url = arguments.iter().any(|argument| { + argument.value.as_ref().is_some_and(|value| { + value.has_parameters() + || value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(text) if text.contains('/'))) + }) + }); + if !all_calls && !forwards_url { + continue; + } + observations.push(call_observation( + SourceLanguage::Rust, + SymbolRef::Call(path), + arguments, + self.functions[function].name.clone(), + SourceLineRange { + start: tokens[start].line, + end: tokens[close].end_line, + }, + )); + } + } + + fn call_arguments(&self, open: usize, close: usize, at: usize) -> Vec { + let mut arguments = split_operands(self.tokens, open + 1, close, ',') + .into_iter() + .filter(|(start, end)| start < end) + .map(|(start, end)| { + let value = self.template(start, end, at); + let is_string = value + .parts + .iter() + .any(|part| matches!(part, UrlPart::Text(_) | UrlPart::Parameter { .. })); + CallArgument { + keyword: None, + value: is_string.then_some(value), + } + }) + .collect::>(); + while arguments + .last() + .is_some_and(|argument| argument.value.is_none()) + { + arguments.pop(); + } + arguments + } +} + +/// Parameter names a caller binds, without `self` receivers. +fn rust_parameters(tokens: &[Token], function: &FunctionSpan) -> Vec { + let Some(open) = (function.start_token..function.body_start_token) + .find(|index| tokens[*index].is_punct('(')) + else { + return Vec::new(); + }; + let Some(close) = matching(tokens, open, '(', ')') else { + return Vec::new(); + }; + split_operands(tokens, open + 1, close, ',') + .into_iter() + .filter_map(|(start, end)| { + let pattern = tokens[start..end] + .iter() + .take_while(|token| !token.is_punct(':')) + .filter_map(Token::ident) + .filter(|name| *name != "mut") + .collect::>(); + match pattern.as_slice() { + [name] if *name != "self" => Some((*name).to_owned()), + _ => None, + } + }) + .collect() +} + +fn statement_end(tokens: &[Token], start: usize) -> usize { + let mut depth = 0_i32; + for (index, token) in tokens.iter().enumerate().skip(start) { + match token.kind { + TokenKind::Punct('(' | '[' | '{') => depth += 1, + TokenKind::Punct(')' | ']' | '}') => { + depth -= 1; + if depth < 0 { + return index; + } + } + TokenKind::Punct(';') if depth == 0 => return index, + _ => {} + } + } + tokens.len() +} + +/// Path such as `BASE`, `config::BASE`, or `self.base_url` spanning all of `start..end`. +fn path_name(tokens: &[Token], start: usize, end: usize) -> Option { + let mut name = String::new(); + let mut cursor = start; + while cursor < end { + if let Some(ident) = tokens[cursor].ident() { + name.push_str(ident); + cursor += 1; + } else if tokens[cursor].is_punct('.') { + name.push('.'); + cursor += 1; + } else if tokens[cursor].is_punct(':') + && tokens + .get(cursor + 1) + .is_some_and(|token| token.is_punct(':')) + { + name.push_str("::"); + cursor += 2; + } else { + return None; + } + } + (!name.is_empty() && !name.ends_with(['.', ':'])).then_some(name) +} + +/// Callee path ending at identifier `index`, such as `crate::client::get` or `self.fetch`. +fn path_before(tokens: &[Token], index: usize) -> Option<(usize, String)> { + let mut start = index; + let mut segments = vec![tokens[index].ident()?]; + loop { + if start >= 2 && tokens[start - 1].is_punct('.') { + segments.insert(0, tokens[start - 2].ident()?); + start -= 2; + } else if start >= 3 && tokens[start - 1].is_punct(':') && tokens[start - 2].is_punct(':') { + segments.insert(0, tokens[start - 3].ident()?); + start -= 3; + } else { + break; + } + } + if start > 0 && (tokens[start - 1].is_punct('.') || tokens[start - 1].is_punct('!')) { + return None; + } + Some((start, segments.join("::"))) +} diff --git a/crates/code-system-graph-core/src/source_http/rust_routers.rs b/crates/code-system-graph-core/src/source_http/rust_routers.rs new file mode 100644 index 0000000..3b70200 --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/rust_routers.rs @@ -0,0 +1,254 @@ +//! Router ownership and mounts for Axum and Actix Web builder chains. +//! +//! A builder chain belongs to the variable it is bound to, to the function it is returned from, +//! or, when it is passed as an argument, to a synthetic router keyed by its first token. Actix +//! `web::scope` chains prefix every route and service registered on them. + +use super::{FunctionSpan, Token, enclosing_symbol, matching, top_level_arguments}; +use crate::router_mounts::mount_observation; +use crate::{SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SymbolRef}; + +/// Router facts of one Rust file. +pub(super) struct RustRouters<'a> { + tokens: &'a [Token], + functions: &'a [FunctionSpan], +} + +/// Router that owns a builder chain plus the prefix its `web::scope` base applies. +pub(super) struct ChainOwner { + pub(super) router: SymbolRef, + pub(super) prefix: Option, +} + +impl<'a> RustRouters<'a> { + pub(super) fn new(tokens: &'a [Token], functions: &'a [FunctionSpan]) -> Self { + Self { tokens, functions } + } + + /// Owner of the chain whose method call is introduced by the `.` at `dot`. + pub(super) fn chain_owner(&self, dot: usize) -> ChainOwner { + let base = self.chain_start(dot); + ChainOwner { + router: self.base_router(base), + prefix: self.scope_prefix(base), + } + } + + /// Records `nest`, `merge`, `service`, and `configure` mounts. + pub(super) fn mounts(&self, actix: bool) -> Vec { + let mut output = Vec::new(); + for (index, token) in self.tokens.iter().enumerate() { + let Some(name) = token.ident() else { + continue; + }; + let framework = match name { + "nest" | "merge" if !actix => SourceFramework::Axum, + "service" | "configure" if actix => SourceFramework::ActixWeb, + _ => continue, + }; + if index == 0 + || !self.tokens[index - 1].is_punct('.') + || !self + .tokens + .get(index + 1) + .is_some_and(|next| next.is_punct('(')) + { + continue; + } + let open = index + 1; + let Some(close) = matching(self.tokens, open, '(', ')') else { + continue; + }; + let arguments = top_level_arguments(self.tokens, open, close); + let (prefix, argument) = match (name, arguments.as_slice()) { + ("nest", [path, argument]) => (self.tokens[*path].literal(), *argument), + ("merge" | "service" | "configure", [argument]) => (None, *argument), + _ => continue, + }; + let end = arguments + .iter() + .find(|start| **start > argument) + .map_or(close, |next| next - 1); + let Some(child) = self.argument_router(name, argument, end) else { + continue; + }; + let owner = self.chain_owner(index - 1); + let prefix = join_prefixes(owner.prefix.as_deref(), prefix); + output.push(mount_observation( + SourceLanguage::Rust, + framework, + child, + Some(owner.router), + prefix.as_deref(), + SourceLineRange { + start: token.line, + end: self.tokens[close].end_line, + }, + )); + } + output + } + + /// Router passed as an argument spanning `start..end`. + fn argument_router(&self, method: &str, start: usize, end: usize) -> Option { + let span = self.tokens.get(start..end)?; + if span.iter().any(|token| token.is_punct('.')) { + return Some(SymbolRef::Local(format!("chain@{start}"))); + } + let path = span + .iter() + .take_while(|token| !token.is_punct('(')) + .filter_map(Token::ident) + .collect::>(); + if path.is_empty() { + return None; + } + let called = span.iter().any(|token| token.is_punct('(')); + match (method, called, path.as_slice()) { + ("nest" | "merge", false, [variable]) => Some(self.local(start, variable)), + _ => Some(SymbolRef::Call(path.join("::"))), + } + } + + fn local(&self, index: usize, variable: &str) -> SymbolRef { + match enclosing_symbol(self.functions, index) { + Some(function) => SymbolRef::Local(format!("{function}.{variable}")), + None => SymbolRef::Local(variable.to_owned()), + } + } + + /// Walks back from a method-call `.` over `ident`, `::`, `.`, and call groups. + fn chain_start(&self, dot: usize) -> usize { + let tokens = self.tokens; + let mut cursor = dot; + loop { + let Some(previous) = cursor.checked_sub(1) else { + return cursor; + }; + let segment = if tokens[previous].is_punct(')') { + match matching_backward(tokens, previous) { + Some(open) if open > 0 && tokens[open - 1].ident().is_some() => open - 1, + _ => return cursor, + } + } else if tokens[previous].ident().is_some() { + previous + } else { + return cursor; + }; + match segment.checked_sub(1).map(|index| &tokens[index]) { + Some(token) if token.is_punct('.') => cursor = segment - 1, + Some(token) + if token.is_punct(':') && segment >= 2 && tokens[segment - 2].is_punct(':') => + { + cursor = segment - 2; + } + _ => return segment, + } + } + } + + fn base_router(&self, base: usize) -> SymbolRef { + let tokens = self.tokens; + let previous = base.checked_sub(1).map(|index| &tokens[index]); + if previous.is_some_and(|token| token.is_punct('(') || token.is_punct(',')) { + return SymbolRef::Local(format!("chain@{base}")); + } + let is_variable = tokens[base].ident().is_some() + && tokens + .get(base + 1) + .is_some_and(|token| token.is_punct('.')); + if is_variable && let Some(variable) = tokens[base].ident() { + if self.is_parameter(base, variable) + && let Some(function) = enclosing_symbol(self.functions, base) + { + return SymbolRef::Function(function); + } + return self.local(base, variable); + } + if previous.is_some_and(|token| token.is_punct('=')) + && let Some(binding) = (base.saturating_sub(8)..base.saturating_sub(1)) + .rev() + .find(|index| tokens[*index].is_ident("let")) + .and_then(|index| { + let name = tokens.get(index + 1)?; + let name = if name.is_ident("mut") { + tokens.get(index + 2)? + } else { + name + }; + name.ident() + }) + { + return self.local(base, binding); + } + enclosing_symbol(self.functions, base).map_or_else( + || SymbolRef::Local(format!("chain@{base}")), + SymbolRef::Function, + ) + } + + fn is_parameter(&self, index: usize, variable: &str) -> bool { + let Some(function) = self + .functions + .iter() + .filter(|function| function.start_token <= index && index <= function.end_token) + .min_by_key(|function| function.end_token - function.start_token) + else { + return false; + }; + let tokens = self.tokens; + let Some(open) = (function.start_token..function.body_start_token) + .find(|candidate| tokens[*candidate].is_punct('(')) + else { + return false; + }; + let close = matching(tokens, open, '(', ')').unwrap_or(function.body_start_token); + (open + 1..close).any(|candidate| { + tokens[candidate].is_ident(variable) + && tokens + .get(candidate + 1) + .is_some_and(|token| token.is_punct(':')) + && !tokens + .get(candidate + 2) + .is_some_and(|token| token.is_punct(':')) + }) + } + + /// Literal of a `web::scope("/p")` chain base. + fn scope_prefix(&self, base: usize) -> Option { + let tokens = self.tokens; + let scope = (base..base + 4).find(|index| { + tokens + .get(*index) + .is_some_and(|token| token.is_ident("scope")) + && tokens + .get(index + 1) + .is_some_and(|token| token.is_punct('(')) + })?; + tokens.get(scope + 2)?.literal().map(str::to_owned) + } +} + +fn matching_backward(tokens: &[Token], close: usize) -> Option { + let mut depth = 0_u32; + for index in (0..=close).rev() { + if tokens[index].is_punct(')') { + depth += 1; + } else if tokens[index].is_punct('(') { + depth -= 1; + if depth == 0 { + return Some(index); + } + } + } + None +} + +/// Joins an outer chain prefix with an inner literal prefix. +pub(super) fn join_prefixes(outer: Option<&str>, inner: Option<&str>) -> Option { + match (outer, inner) { + (None, None) => None, + (Some(prefix), None) | (None, Some(prefix)) => Some(prefix.to_owned()), + (Some(outer), Some(inner)) => Some(format!("{outer}/{inner}")), + } +} diff --git a/crates/code-system-graph-core/src/source_http/rust_test_clients.rs b/crates/code-system-graph-core/src/source_http/rust_test_clients.rs new file mode 100644 index 0000000..40a6a21 --- /dev/null +++ b/crates/code-system-graph-core/src/source_http/rust_test_clients.rs @@ -0,0 +1,115 @@ +//! In-process Rust test requests: axum `Request` builders sent through `oneshot`, and actix +//! `test::TestRequest` builders. + +use super::rust_flows::RustScopes; +use super::{ + FunctionSpan, SourceObservationCollector, Token, canonical_method, enclosing_symbol, http_from_literal, matching, rust_method_expression +}; +use crate::{SourceFramework, SourceLanguage, SourceLineRange, SourceRole, UrlTemplate}; + +/// A request builder chain: its method, URI argument span, and last token. +struct RequestChain { + method: Option, + uri: Option<(usize, usize)>, + end: usize, +} + +/// Records axum and actix test requests of one Rust file. +pub(super) fn parse_rust_test_requests( + tokens: &[Token], + functions: &[FunctionSpan], + scopes: &RustScopes<'_>, + observations: &mut SourceObservationCollector<'_>, +) { + let oneshot = tokens.iter().any(|token| token.is_ident("oneshot")); + for index in 0..tokens.len() { + let framework = match tokens[index].ident() { + Some("Request") if oneshot => SourceFramework::AxumOneshot, + Some("TestRequest") => SourceFramework::ActixTest, + _ => continue, + }; + if !path_separator(tokens, index + 1) { + continue; + } + let Some(constructor) = tokens.get(index + 3).and_then(Token::ident) else { + continue; + }; + let Some(chain) = request_chain(tokens, index + 3, constructor) else { + continue; + }; + let template = chain + .uri + .map(|(start, end)| scopes.template(start, end, index)); + let literal = template.as_ref().and_then(UrlTemplate::client_literal); + let mut observation = http_from_literal( + SourceLanguage::Rust, + framework, + SourceRole::Consumer, + chain.method, + literal.as_deref(), + enclosing_symbol(functions, index), + SourceLineRange { + start: tokens[index].line, + end: tokens[chain.end].end_line, + }, + false, + ); + observation.url = template.filter(UrlTemplate::has_parameters); + observations.push(observation); + } +} + +/// The builder chain whose constructor `Type::constructor(...)` is at token `constructor_index`. +fn request_chain( + tokens: &[Token], + constructor_index: usize, + constructor: &str, +) -> Option { + let open = constructor_index + 1; + if !tokens.get(open)?.is_punct('(') { + return None; + } + let close = matching(tokens, open, '(', ')')?; + let mut chain = RequestChain { + method: None, + uri: None, + end: close, + }; + match constructor { + "builder" | "default" => chain.method = Some("GET".to_owned()), + constructor => { + chain.method = Some(canonical_method(constructor)?.to_owned()); + if close > open + 1 { + chain.uri = Some((open + 1, close)); + } + } + } + let mut cursor = close + 1; + while tokens.get(cursor).is_some_and(|token| token.is_punct('.')) { + let Some(name) = tokens.get(cursor + 1).and_then(Token::ident) else { + break; + }; + if !tokens + .get(cursor + 2) + .is_some_and(|token| token.is_punct('(')) + { + break; + } + let Some(call_close) = matching(tokens, cursor + 2, '(', ')') else { + break; + }; + match name { + "uri" => chain.uri = Some((cursor + 3, call_close)), + "method" => chain.method = rust_method_expression(tokens, cursor + 3, call_close), + _ => {} + } + chain.end = call_close; + cursor = call_close + 1; + } + Some(chain) +} + +fn path_separator(tokens: &[Token], at: usize) -> bool { + tokens.get(at).is_some_and(|token| token.is_punct(':')) + && tokens.get(at + 1).is_some_and(|token| token.is_punct(':')) +} diff --git a/crates/code-system-graph-core/src/source_polyglot.rs b/crates/code-system-graph-core/src/source_polyglot.rs index 6d58e76..a67e14c 100644 --- a/crates/code-system-graph-core/src/source_polyglot.rs +++ b/crates/code-system-graph-core/src/source_polyglot.rs @@ -1,8 +1,14 @@ use std::collections::BTreeSet; -use crate::source_http::SourceObservationCollector; +use crate::router_mounts::mount_observation; +use crate::source_http::{ + BraceClients, SourceObservationCollector, collect_brace_clients, java_method_lines +}; +use crate::source_routers::{ + GoRouterScope, receiver_before, script_imports, script_router_observations +}; use crate::{ - ExtractionLimitExceeded, ExtractionTracker, SourceEpistemicStatus, SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SourceRole, SourceWarning, normalize_source_http_path + ExtractionLimitExceeded, ExtractionTracker, SourceEpistemicStatus, SourceFramework, SourceLanguage, SourceLineRange, SourceObservation, SourceRole, SourceWarning, SymbolRef, normalize_source_http_path }; const METHODS: [(&str, &str); 8] = [ @@ -120,66 +126,65 @@ fn collect_go_source<'a>( let has_http = source.contains("\"net/http\""); let has_gin = source.contains("github.com/gin-gonic/gin"); let has_chi = source.contains("github.com/go-chi/chi"); - for statement in statements(source) { - if has_http { - if let Some(method) = method_call(&statement.text, "http.") - && matches!(method.as_str(), "GET" | "POST") - { - observations.push(http_observation( - SourceLanguage::Go, - SourceFramework::GoNetHttp, - SourceRole::Consumer, - Some(method), - first_literal_after_call(&statement.text), - None, - statement.lines, - )); - } - if statement.text.contains("http.NewRequest(") { - observations.push(http_observation( - SourceLanguage::Go, - SourceFramework::GoNetHttp, - SourceRole::Consumer, - literal_method_argument(&statement.text), - nth_literal(&statement.text, 2), - None, - statement.lines, - )); - } - if statement.text.contains("http.HandleFunc(") { - observations.push(incomplete_observation( - SourceLanguage::Go, - SourceFramework::GoNetHttp, - SourceRole::Provider, - None, - first_literal_after_call(&statement.text) - .as_deref() - .and_then(literal_path), - argument_identifier(&statement.text, 2), - statement.lines, - SourceWarning::DynamicMethod, - )); + let router_framework = if has_gin { + SourceFramework::Gin + } else if has_chi { + SourceFramework::Chi + } else { + SourceFramework::GoNetHttp + }; + let mut scope = GoRouterScope::default(); + for statement in go_statements(source) { + if has_http || has_gin || has_chi { + for mount in scope.observe(&statement, router_framework) { + observations.push(mount); } } + if has_http { + append_go_pattern_route(&mut observations, &statement, &scope); + } + if has_http + && statement.text.contains("http.HandleFunc(") + && !first_literal_after_call(&statement.text) + .is_some_and(|pattern| pattern.contains(' ')) + { + observations.push(incomplete_observation( + SourceLanguage::Go, + SourceFramework::GoNetHttp, + SourceRole::Provider, + None, + first_literal_after_call(&statement.text) + .as_deref() + .and_then(literal_path), + route_symbol(&statement.text, None), + statement.lines, + SourceWarning::DynamicMethod, + )); + } if has_gin { append_receiver_route( &mut observations, - SourceLanguage::Go, SourceFramework::Gin, &statement, true, + &scope, ); } if has_chi { append_receiver_route( &mut observations, - SourceLanguage::Go, SourceFramework::Chi, &statement, false, + &scope, ); } } + let clients = BraceClients { + go_http: has_http, + ..BraceClients::default() + }; + collect_brace_clients(source, SourceLanguage::Go, clients, &mut observations); observations } @@ -212,17 +217,43 @@ fn collect_java_source<'a>( let web_client = source.contains("org.springframework.web.reactive.function.client.WebClient"); let feign = source.contains("@FeignClient") || source.contains("openfeign.FeignClient"); let lines = source.lines().collect::>(); + let methods = if spring && !feign { + java_method_lines(source) + } else { + Vec::new() + }; + let mut class_prefix = None::; for (index, line) in lines.iter().enumerate() { let trimmed = line.trim(); + if java_annotates_type(&lines, index) { + if trimmed.starts_with("@RequestMapping") { + class_prefix = quoted_value(trimmed); + } else if trimmed.starts_with("@FeignClient") { + class_prefix = feign_path(trimmed); + } + continue; + } if spring && trimmed.starts_with('@') && let Some((method, path)) = java_mapping(trimmed) { - let symbol = lines - .iter() - .skip(index + 1) - .take(5) - .find_map(|candidate| java_method_name(candidate)); + let path = match (&class_prefix, path) { + (Some(prefix), path) => Some(format!("/{prefix}/{}", path.unwrap_or_default())), + (None, path) => path, + }; + let line = line_number(index); + let symbol = if feign { + lines + .iter() + .skip(index + 1) + .take(5) + .find_map(|candidate| java_method_name(candidate)) + } else { + methods + .iter() + .find(|(start, _)| *start >= line) + .map(|(_, name)| name.clone()) + }; observations.push(http_observation( SourceLanguage::Java, if feign { @@ -245,28 +276,11 @@ fn collect_java_source<'a>( )); } } - if web_client { - for statement in statements(source) { - if !statement.text.contains(".uri(") { - continue; - } - let method = METHODS.iter().find_map(|(name, method)| { - statement - .text - .contains(&format!(".{name}()")) - .then_some((*method).to_owned()) - }); - observations.push(http_observation( - SourceLanguage::Java, - SourceFramework::WebClient, - SourceRole::Consumer, - method, - literal_after(&statement.text, ".uri("), - None, - statement.lines, - )); - } - } + let clients = BraceClients { + web_client, + ..BraceClients::default() + }; + collect_brace_clients(source, SourceLanguage::Java, clients, &mut observations); observations } @@ -299,35 +313,20 @@ fn collect_ecmascript_at_path<'a>( || source.contains("require(\"axios\")") || source.contains("require('axios')"); let nest = source.contains("@nestjs/common"); - let express_receivers = assigned_receivers(source, "express"); + let mut express_receivers = assigned_receivers(source, "express"); + if express && source.contains("Router") { + express_receivers.extend(assigned_receivers(source, "Router")); + } let fastify_receivers = assigned_receivers(source, "fastify"); - for statement in statements(source) { - if statement.text.contains("fetch(") { - let method = object_method(&statement.text).or_else(|| Some("GET".to_owned())); - observations.push(http_observation( - language, - SourceFramework::Fetch, - SourceRole::Consumer, - method, - literal_after(&statement.text, "fetch("), - None, - statement.lines, - )); - } - if axios && statement.text.contains("axios.") { - let method = method_call(&statement.text, "axios."); - if method.is_some() { - observations.push(http_observation( - language, - SourceFramework::Axios, - SourceRole::Consumer, - method, - first_literal_after_call(&statement.text), - None, - statement.lines, - )); - } + let statements = statements(source); + if express { + let imports = script_imports(&statements); + for mount in script_router_observations(&statements, language, &express_receivers, &imports) + { + observations.push(mount); } + } + for statement in statements { if express { append_js_route( &mut observations, @@ -348,34 +347,81 @@ fn collect_ecmascript_at_path<'a>( } } if nest { - let lines = source.lines().collect::>(); - for (index, line) in lines.iter().enumerate() { - let Some((method, path)) = nest_mapping(line.trim()) else { - continue; - }; - let symbol = lines - .iter() - .skip(index + 1) - .take(5) - .find_map(|candidate| ecmascript_method_name(candidate)); - observations.push(http_observation( - language, - SourceFramework::NestJs, - SourceRole::Provider, - Some(method), - path, - symbol, - SourceLineRange { - start: line_number(index), - end: line_number(index), - }, - )); - } + append_nest_routes(&mut observations, source, language); + } + if let Some(prefix) = nest_global_prefix(source) { + observations.push(mount_observation( + language, + SourceFramework::NestJs, + SymbolRef::Function(NEST_APPLICATION.to_owned()), + None, + Some(&prefix), + SourceLineRange { start: 1, end: 1 }, + )); } append_next_app_routes(&mut observations, source_path, source, language); + let clients = BraceClients { + axios, + ..BraceClients::default() + }; + collect_brace_clients(source, language, clients, &mut observations); observations } +/// Repository-wide router of every `NestJS` controller, mounted by `setGlobalPrefix`. +const NEST_APPLICATION: &str = "@nestjs"; + +fn append_nest_routes( + observations: &mut SourceObservationCollector<'_>, + source: &str, + language: SourceLanguage, +) { + let lines = source.lines().collect::>(); + let mut controller = None::; + for (index, line) in lines.iter().enumerate() { + let trimmed = line.trim(); + if let Some(arguments) = trimmed.strip_prefix("@Controller(") { + let prefix = quoted_value(arguments.split(')').next().unwrap_or_default()); + controller = Some(prefix.unwrap_or_default()); + continue; + } + let Some((method, path)) = nest_mapping(trimmed) else { + continue; + }; + let symbol = lines + .iter() + .skip(index + 1) + .take(5) + .find_map(|candidate| ecmascript_method_name(candidate)); + let path = match (&controller, path) { + (Some(prefix), path) => Some(format!("/{prefix}/{}", path.unwrap_or_default())), + (None, Some(path)) if !path.starts_with('/') => Some(format!("/{path}")), + (None, path) => path, + }; + let mut observation = http_observation( + language, + SourceFramework::NestJs, + SourceRole::Provider, + Some(method), + path, + symbol, + SourceLineRange { + start: line_number(index), + end: line_number(index), + }, + ); + observation.router = Some(SymbolRef::Function(NEST_APPLICATION.to_owned())); + observations.push(observation); + } +} + +fn nest_global_prefix(source: &str) -> Option { + let at = source.find(".setGlobalPrefix(")?; + quoted_values(&source[at..source[at..].find(')').map_or(source.len(), |end| at + end)]) + .into_iter() + .next() +} + fn append_next_app_routes( observations: &mut SourceObservationCollector<'_>, source_path: &str, @@ -470,18 +516,28 @@ fn next_route_export_method(line: &str) -> Option { } #[derive(Debug)] -struct Statement { - text: String, - lines: SourceLineRange, +pub(crate) struct Statement { + pub(crate) text: String, + pub(crate) lines: SourceLineRange, } fn statements(source: &str) -> Vec { + split_statements(source, false) +} + +/// Splits Go source into statements, also ending one at each opening or closing block brace so +/// closures and blocks are visible as separate statements. +fn go_statements(source: &str) -> Vec { + split_statements(source, true) +} + +fn split_statements(source: &str, split_blocks: bool) -> Vec { let mut output = Vec::new(); let mut text = String::new(); let mut start = 1_u32; let mut depth = 0_i32; for (index, line) in source.lines().enumerate() { - let trimmed = line.split("//").next().unwrap_or_default().trim(); + let trimmed = strip_line_comment(line).trim(); if trimmed.is_empty() { continue; } @@ -492,7 +548,8 @@ fn statements(source: &str) -> Vec { } text.push_str(trimmed); depth += delimiter_delta(trimmed); - if depth <= 0 || trimmed.ends_with(';') { + let block_boundary = split_blocks && (trimmed.ends_with('{') || trimmed.starts_with('}')); + if depth <= 0 || trimmed.ends_with(';') || block_boundary { output.push(Statement { text: std::mem::take(&mut text), lines: SourceLineRange { @@ -545,25 +602,31 @@ fn append_js_route( if method.is_none() { return; } - output.push(http_observation( + let mut observation = http_observation( language, framework, SourceRole::Provider, - method, + method.clone(), first_literal_after_call(&statement.text), - argument_identifier(&statement.text, 2), + route_symbol(&statement.text, method.as_deref()), statement.lines, - )); + ); + observation.router = METHODS + .iter() + .find_map(|(name, _)| receiver_before(&statement.text, name)) + .filter(|receiver| receivers.contains(*receiver)) + .map(|receiver| SymbolRef::Local(receiver.to_owned())); + output.push(observation); } fn append_receiver_route( output: &mut SourceObservationCollector<'_>, - language: SourceLanguage, framework: SourceFramework, statement: &Statement, uppercase: bool, + scope: &GoRouterScope, ) { - let method = METHODS.iter().find_map(|(name, method)| { + let Some((call, method)) = METHODS.iter().find_map(|(name, method)| { let name = if uppercase { name.to_ascii_uppercase() } else { @@ -575,20 +638,58 @@ fn append_receiver_route( statement .text .contains(&format!(".{name}(")) - .then_some((*method).to_owned()) - }); - if method.is_none() { + .then(|| (name, (*method).to_owned())) + }) else { return; - } - output.push(http_observation( - language, + }; + let mut observation = http_observation( + SourceLanguage::Go, framework, SourceRole::Provider, - method, + Some(method.clone()), first_literal_after_call(&statement.text), - argument_identifier(&statement.text, 2), + route_symbol(&statement.text, Some(&method)), statement.lines, - )); + ); + observation.router = + receiver_before(&statement.text, &call).map(|receiver| scope.reference(receiver)); + output.push(observation); +} + +/// Records Go 1.22 `ServeMux` patterns such as `mux.HandleFunc("GET /orders/{id}", handler)`. +fn append_go_pattern_route( + output: &mut SourceObservationCollector<'_>, + statement: &Statement, + scope: &GoRouterScope, +) { + let Some(call) = ["HandleFunc", "Handle"] + .into_iter() + .find(|call| statement.text.contains(&format!(".{call}("))) + else { + return; + }; + let Some(pattern) = first_literal_after_call(&statement.text) else { + return; + }; + let Some((method, path)) = pattern.split_once(' ') else { + return; + }; + let Some(method) = canonical_method(method) else { + return; + }; + let mut observation = http_observation( + SourceLanguage::Go, + SourceFramework::GoNetHttp, + SourceRole::Provider, + Some(method.clone()), + Some(path.trim().to_owned()), + route_symbol(&statement.text, Some(&method)), + statement.lines, + ); + observation.router = receiver_before(&statement.text, call) + .filter(|receiver| *receiver != "http") + .map(|receiver| scope.reference(receiver)); + output.push(observation); } fn http_observation( @@ -601,10 +702,9 @@ fn http_observation( lines: SourceLineRange, ) -> SourceObservation { let literal_missing = literal.is_none(); - let unmapped_authority = role == SourceRole::Consumer - && literal - .as_deref() - .is_some_and(|value| value.starts_with("http://") || value.starts_with("https://")); + let authority = (role == SourceRole::Consumer) + .then(|| literal.as_deref().and_then(crate::routes::url_authority)) + .flatten(); let path = literal.and_then(|value| literal_path(&value)); let mut warnings = Vec::new(); if method.is_none() { @@ -615,15 +715,11 @@ fn http_observation( } else if path.is_none() { warnings.push(SourceWarning::UnsupportedLiteralPath); } - if unmapped_authority { - warnings.push(SourceWarning::UnmappedAuthority); - } if role == SourceRole::Provider && symbol_name.is_none() { warnings.push(SourceWarning::MissingSymbol); } let confirmed = method.is_some() && path.is_some() - && !unmapped_authority && (role != SourceRole::Provider || symbol_name.is_some()); SourceObservation { language, @@ -634,6 +730,11 @@ fn http_observation( symbol_name, related_symbol: None, related_path: None, + authority, + router: None, + mount_parent: None, + url: None, + call: None, lines, status: if confirmed { SourceEpistemicStatus::Confirmed @@ -668,6 +769,11 @@ fn incomplete_observation( symbol_name, related_symbol: None, related_path: None, + authority: None, + router: None, + mount_parent: None, + url: None, + call: None, lines, status: SourceEpistemicStatus::Incomplete, confidence: 0.0, @@ -675,15 +781,24 @@ fn incomplete_observation( } } -fn method_call(text: &str, prefix: &str) -> Option { - METHODS.iter().find_map(|(name, method)| { - text.contains(&format!("{prefix}{name}(")) - .then_some((*method).to_owned()) - .or_else(|| { - text.contains(&format!("{prefix}{}(", name.to_ascii_uppercase())) - .then_some((*method).to_owned()) - }) - }) +/// Removes a trailing `//` comment that starts outside string literals. +fn strip_line_comment(line: &str) -> &str { + let mut quote = None; + let mut previous = '\0'; + for (index, character) in line.char_indices() { + match quote { + Some(active) if character == active && previous != '\\' => quote = None, + None if matches!(character, '\'' | '"' | '`') => quote = Some(character), + None if character == '/' && previous == '/' => return &line[..index - 1], + Some(_) | None => {} + } + previous = if previous == '\\' && character == '\\' { + '\0' + } else { + character + }; + } + line } fn first_literal_after_call(text: &str) -> Option { @@ -696,17 +811,11 @@ fn literal_after(text: &str, marker: &str) -> Option { quoted_value(&text[start..]) } -fn nth_literal(text: &str, target: usize) -> Option { - quoted_values(text) - .into_iter() - .nth(target.saturating_sub(1)) -} - fn quoted_value(text: &str) -> Option { quoted_values(text).into_iter().next() } -fn quoted_values(text: &str) -> Vec { +pub(crate) fn quoted_values(text: &str) -> Vec { let mut values = Vec::new(); let mut quote = None; let mut start = 0; @@ -724,15 +833,6 @@ fn quoted_values(text: &str) -> Vec { values } -fn literal_method_argument(text: &str) -> Option { - nth_literal(text, 1).and_then(|value| canonical_method(&value)) -} - -fn object_method(text: &str) -> Option { - let start = text.find("method")? + "method".len(); - quoted_value(&text[start..]).and_then(|method| canonical_method(&method)) -} - fn assigned_receivers(source: &str, factory: &str) -> BTreeSet { source .lines() @@ -750,17 +850,94 @@ fn assigned_receivers(source: &str, factory: &str) -> BTreeSet { .collect() } -fn argument_identifier(text: &str, target: usize) -> Option { +/// Handler symbol of a route registration: its last argument as a dotted name, unwrapped from +/// single-argument wrapper calls such as `asyncHandler(controller.list)`, or `METHOD path` for an +/// inline function, which has no name of its own. +fn route_symbol(text: &str, method: Option<&str>) -> Option { + let arguments = top_level_arguments(text)?; + if arguments.len() < 2 { + return None; + } + let mut handler = arguments.last()?.trim(); + loop { + let unprefixed = handler + .strip_prefix("async") + .map_or(handler, str::trim_start); + let inline = ["function", "(", "func(", "func "] + .iter() + .any(|prefix| unprefixed.starts_with(prefix)) + || unprefixed + .split_once("=>") + .is_some_and(|(parameter, _)| is_dotted_name(parameter.trim())); + if inline { + let path = first_literal_after_call(text)?; + return Some(match method { + Some(method) => format!("{method} {path}"), + None => path, + }); + } + let name = handler.trim_start_matches('&'); + if is_dotted_name(name) { + return Some(name.to_owned()); + } + let (callee, rest) = handler.split_once('(')?; + let inner = rest.strip_suffix(')')?.trim(); + if !is_dotted_name(callee.trim()) + || inner.is_empty() + || top_level_arguments(&format!("({inner})"))?.len() != 1 + { + return None; + } + handler = inner; + } +} + +fn is_dotted_name(text: &str) -> bool { + !text.is_empty() + && !text.starts_with('.') + && !text.ends_with('.') + && !text.starts_with(|character: char| character.is_ascii_digit()) + && text.chars().all(|character| { + character.is_ascii_alphanumeric() || matches!(character, '_' | '$' | '.') + }) +} + +/// Top-level arguments of the first call in `text`, ignoring nested groups and string contents. +fn top_level_arguments(text: &str) -> Option> { let open = text.find('(')?; - let close = text.rfind(')')?; - let argument = text[open + 1..close].split(',').nth(target - 1)?.trim(); - let identifier = argument - .trim_start_matches('&') - .split(|character: char| !character.is_ascii_alphanumeric() && character != '_') - .next() - .unwrap_or_default(); - (!identifier.is_empty() && !identifier.starts_with(['"', '\'', '`'])) - .then(|| identifier.to_owned()) + let mut arguments = Vec::new(); + let mut depth = 0_u32; + let mut quote = None; + let mut escaped = false; + let mut start = open + 1; + for (offset, character) in text[open + 1..].char_indices() { + let index = open + 1 + offset; + if let Some(delimiter) = quote { + if escaped { + escaped = false; + } else if character == '\\' { + escaped = true; + } else if character == delimiter { + quote = None; + } + continue; + } + match character { + '"' | '\'' | '`' => quote = Some(character), + '(' | '[' | '{' => depth += 1, + ')' | ']' | '}' if depth == 0 => { + arguments.push(&text[start..index]); + return Some(arguments); + } + ')' | ']' | '}' => depth -= 1, + ',' if depth == 0 => { + arguments.push(&text[start..index]); + start = index + 1; + } + _ => {} + } + } + None } fn canonical_method(value: &str) -> Option { @@ -789,6 +966,29 @@ fn literal_path(value: &str) -> Option { )) } +/// Whether the annotation on `index` precedes a class or interface declaration. +fn java_annotates_type(lines: &[&str], index: usize) -> bool { + if !lines[index].trim_start().starts_with('@') { + return false; + } + lines + .iter() + .skip(index + 1) + .map(|line| line.trim()) + .find(|line| !line.is_empty() && !line.starts_with('@') && !line.starts_with("//")) + .is_some_and(|declaration| { + declaration + .split_whitespace() + .any(|token| matches!(token, "class" | "interface" | "record")) + }) +} + +fn feign_path(line: &str) -> Option { + let at = line.find("path")?; + let rest = line[at + "path".len()..].trim_start().strip_prefix('=')?; + quoted_value(rest) +} + fn java_mapping(line: &str) -> Option<(Option, Option)> { for (annotation, method) in [ ("@DeleteMapping", "DELETE"), @@ -822,8 +1022,12 @@ fn nest_mapping(line: &str) -> Option<(String, Option)> { ("Put", "PUT"), ] { let marker = format!("@{name}("); - if line.contains(&marker) { - return Some((method.to_owned(), literal_after(line, &marker))); + if let Some(at) = line.find(&marker) { + let arguments = &line[at + marker.len()..]; + let literal = (!arguments.trim_start().starts_with(')')) + .then(|| literal_after(line, &marker)) + .flatten(); + return Some((method.to_owned(), literal)); } } None @@ -910,6 +1114,68 @@ mod tests { )); } + #[test] + fn route_handlers_should_be_named_through_members_middleware_wrappers_and_inline_functions() { + let express = r#" +import express from "express"; +const app = express(); +app.get("/orders", auth, orders.list); +app.post("/orders", asyncHandler(orders.create)); +app.delete("/orders/:id", async (req, res) => { res.sendStatus(204); }); +"#; + let gin = r#" +package main +import "github.com/gin-gonic/gin" +func main() { + r := gin.Default() + r.GET("/orders/:id", handlers.GetOrder) +} +"#; + let handlers = parse_typescript_source(express) + .into_iter() + .chain(parse_go_source(gin)) + .filter(|item| item.role == SourceRole::Provider) + .map(|item| item.symbol_name.unwrap_or_default()) + .collect::>(); + + assert_eq!( + handlers, + [ + "orders.list", + "orders.create", + "DELETE /orders/:id", + "handlers.GetOrder" + ] + ); + } + + #[test] + fn spring_handlers_should_be_the_method_declared_after_the_mapping() { + let source = r#" +import org.springframework.web.bind.annotation.*; + +@RestController +@RequestMapping("/orders") +class OrdersController { + @GetMapping("/{id}") + @Operation( + summary = "Read one order (by id)", + description = "Returns the order") + @PreAuthorize("hasRole('reader')") + public Order read(@PathVariable String id) { return null; } + + @PostMapping public Order create(@RequestBody Order order) { return order; } +} +"#; + let handlers = parse_java_source(source) + .into_iter() + .filter(|item| item.role == SourceRole::Provider) + .map(|item| item.symbol_name.unwrap_or_default()) + .collect::>(); + + assert_eq!(handlers, ["read", "create"]); + } + #[test] fn typescript_should_extract_fetch_axios_express_fastify_and_nestjs() { let source = r#" @@ -941,6 +1207,33 @@ listUsers() {} ); } + #[test] + fn absolute_urls_should_survive_comment_stripping_and_record_their_authority() { + let source = "export async function load() {\n await fetch('http://Orders-API:8080/v1/orders/42'); // primary\n await fetch(\"https://payments.example.com/v1/orders/42\");\n}\n"; + + let result = parse_typescript_source(source); + let calls = result + .iter() + .map(|item| (item.status, item.path.as_deref(), item.authority.as_deref())) + .collect::>(); + + assert_eq!( + calls, + [ + ( + SourceEpistemicStatus::Confirmed, + Some("/v1/orders/42"), + Some("orders-api:8080") + ), + ( + SourceEpistemicStatus::Confirmed, + Some("/v1/orders/42"), + Some("payments.example.com") + ), + ] + ); + } + #[test] fn multiline_fetch_should_preserve_literal_method_and_path() { let result = parse_typescript_source(include_str!( diff --git a/crates/code-system-graph-core/src/source_routers.rs b/crates/code-system-graph-core/src/source_routers.rs new file mode 100644 index 0000000..50b6809 --- /dev/null +++ b/crates/code-system-graph-core/src/source_routers.rs @@ -0,0 +1,648 @@ +//! Router declarations and mounts for the line-oriented TypeScript, JavaScript, and Go extractors. +//! +//! Routes record the router they are registered on, and mounts record which router is attached +//! under which prefix. [`crate::compose_router_mounts`] resolves both per repository. + +use std::collections::{BTreeMap, BTreeSet}; + +use crate::router_mounts::mount_observation; +use crate::source_polyglot::{Statement, quoted_values}; +use crate::{SourceFramework, SourceLanguage, SourceObservation, SymbolRef}; + +/// Module bindings of one script file: local name to `(module specifier, imported name)`. +pub(crate) type ScriptImports = BTreeMap; + +/// Returns the identifier immediately before `.{call}(` in `text`. +pub(crate) fn receiver_before<'a>(text: &'a str, call: &str) -> Option<&'a str> { + let at = text.find(&format!(".{call}("))?; + let start = text[..at] + .rfind(|character: char| !is_identifier(character)) + .map_or(0, |index| index + 1); + let receiver = &text[start..at]; + (!receiver.is_empty()).then_some(receiver) +} + +fn is_identifier(character: char) -> bool { + character.is_ascii_alphanumeric() || character == '_' || character == '$' +} + +/// Top-level comma-separated arguments of the first call opened at `open`. +pub(crate) fn call_arguments(text: &str, open: usize) -> Vec<&str> { + let mut arguments = Vec::new(); + let mut depth = 0_i32; + let mut quote = None; + let mut start = open + 1; + for (index, character) in text.char_indices().skip_while(|(index, _)| *index <= open) { + match quote { + Some(active) if character == active => quote = None, + Some(_) => {} + None => match character { + '\'' | '"' | '`' => quote = Some(character), + '(' | '[' | '{' => depth += 1, + ')' | ']' | '}' if depth == 0 => { + arguments.push(text[start..index].trim()); + return arguments + .into_iter() + .filter(|value| !value.is_empty()) + .collect(); + } + ')' | ']' | '}' => depth -= 1, + ',' if depth == 0 => { + arguments.push(text[start..index].trim()); + start = index + 1; + } + _ => {} + }, + } + } + // A statement split at a block opener, such as `r.Route("/x", func(r chi.Router) {`. + arguments.push(text[start..].trim()); + arguments + .into_iter() + .filter(|value| !value.is_empty()) + .collect() +} + +fn string_literal(argument: &str) -> Option { + let first = argument.chars().next()?; + if !matches!(first, '\'' | '"' | '`') || !argument.ends_with(first) || argument.len() < 2 { + return None; + } + let value = &argument[1..argument.len() - 1]; + (!value.contains("${")).then(|| value.to_owned()) +} + +fn simple_identifier(value: &str) -> Option<&str> { + (!value.is_empty() && value.chars().all(is_identifier)).then_some(value) +} + +/// Discovers relative ES module and `CommonJS` bindings. +pub(crate) fn script_imports(statements: &[Statement]) -> ScriptImports { + let mut imports = ScriptImports::new(); + for statement in statements { + let text = statement.text.as_str(); + if let Some(rest) = text.strip_prefix("import ") + && let Some((clause, module)) = rest.rsplit_once(" from ") + && let Some(module) = quoted_values(module).into_iter().next() + { + record_import_clause(&mut imports, clause, &module); + continue; + } + let Some(open) = text.find("require(") else { + continue; + }; + let Some(module) = quoted_values(&text[open..]).into_iter().next() else { + continue; + }; + let Some((left, right)) = text.split_once('=') else { + continue; + }; + let binding = left + .trim() + .trim_start_matches("const ") + .trim_start_matches("let ") + .trim_start_matches("var ") + .trim(); + let member = right + .split_once(").") + .and_then(|(_, member)| simple_identifier(member.trim().trim_end_matches(';'))); + if let Some(fields) = binding + .strip_prefix('{') + .and_then(|value| value.strip_suffix('}')) + { + for field in fields.split(',') { + let (imported, local) = field + .split_once(':') + .map_or((field.trim(), field.trim()), |(imported, local)| { + (imported.trim(), local.trim()) + }); + if let (Some(imported), Some(local)) = + (simple_identifier(imported), simple_identifier(local)) + { + imports.insert(local.to_owned(), (module.clone(), imported.to_owned())); + } + } + } else if let Some(local) = simple_identifier(binding) { + imports.insert( + local.to_owned(), + (module, member.unwrap_or("default").to_owned()), + ); + } + } + imports.retain(|_, (module, _)| module.starts_with('.')); + imports +} + +fn record_import_clause(imports: &mut ScriptImports, clause: &str, module: &str) { + let clause = clause.trim(); + let (default, named) = match clause.split_once('{') { + Some((default, named)) => (default.trim().trim_end_matches(','), Some(named)), + None => (clause, None), + }; + if let Some(namespace) = default.trim().strip_prefix("* as ") { + if let Some(local) = simple_identifier(namespace.trim()) { + imports.insert(local.to_owned(), (module.to_owned(), "*".to_owned())); + } + } else if let Some(local) = simple_identifier(default.trim()) { + imports.insert(local.to_owned(), (module.to_owned(), "default".to_owned())); + } + for item in named + .and_then(|named| named.split('}').next()) + .into_iter() + .flat_map(|named| named.split(',')) + { + let item = item.trim().trim_start_matches("type "); + let (imported, local) = item + .split_once(" as ") + .map_or((item, item), |(imported, local)| { + (imported.trim(), local.trim()) + }); + if let (Some(imported), Some(local)) = + (simple_identifier(imported), simple_identifier(local)) + { + imports.insert(local.to_owned(), (module.to_owned(), imported.to_owned())); + } + } +} + +/// Resolves a router expression written in a script file. +fn script_reference( + expression: &str, + receivers: &BTreeSet, + imports: &ScriptImports, +) -> Option { + let expression = expression.trim().trim_end_matches("()"); + if let Some(open) = expression.find("require(") { + let module = quoted_values(&expression[open..]).into_iter().next()?; + let member = expression + .split_once(").") + .and_then(|(_, member)| simple_identifier(member)) + .unwrap_or("default"); + return module.starts_with('.').then(|| SymbolRef::Import { + module, + name: member.to_owned(), + }); + } + let (head, member) = expression + .split_once('.') + .map_or((expression, None), |(head, member)| (head, Some(member))); + let head = simple_identifier(head)?; + if let Some((module, imported)) = imports.get(head) { + let name = match (imported.as_str(), member) { + ("*", Some(member)) => simple_identifier(member)?.to_owned(), + ("*", None) => "default".to_owned(), + (imported, None) => imported.to_owned(), + (_, Some(_)) => return None, + }; + return Some(SymbolRef::Import { + module: module.clone(), + name, + }); + } + (member.is_none() && receivers.contains(head)).then(|| SymbolRef::Local(head.to_owned())) +} + +/// Emits Express `use` mounts, default-export aliases, and router factory returns. +pub(crate) fn script_router_observations( + statements: &[Statement], + language: SourceLanguage, + receivers: &BTreeSet, + imports: &ScriptImports, +) -> Vec { + let mut output = Vec::new(); + let mut functions = Vec::<(String, i32)>::new(); + let mut depth = 0_i32; + for statement in statements { + let text = statement.text.as_str(); + if let Some(name) = declared_function(text) { + functions.push((name.to_owned(), depth)); + } + for receiver in receivers { + let marker = format!("{receiver}.use("); + let Some(at) = text.find(&marker) else { + continue; + }; + if at > 0 + && text[..at] + .ends_with(|character: char| is_identifier(character) || character == '.') + { + continue; + } + let arguments = call_arguments(text, at + marker.len() - 1); + let (prefix, routers) = match arguments.split_first() { + Some((first, rest)) => match string_literal(first) { + Some(prefix) => (Some(prefix), rest), + None => (None, arguments.as_slice()), + }, + None => continue, + }; + for router in routers { + if let Some(child) = script_reference(router, receivers, imports) { + output.push(mount_observation( + language, + SourceFramework::Express, + child, + Some(SymbolRef::Local(receiver.clone())), + prefix.as_deref(), + statement.lines, + )); + } + } + } + let exported = text + .strip_prefix("export default ") + .or_else(|| text.strip_prefix("module.exports = ")) + .or_else(|| text.strip_prefix("module.exports=")); + if let Some(exported) = + exported.and_then(|value| simple_identifier(value.trim().trim_end_matches(';'))) + && receivers.contains(exported) + { + output.push(alias(language, exported, "default", statement)); + } + if let Some(returned) = text + .strip_prefix("return ") + .and_then(|value| simple_identifier(value.trim().trim_end_matches(';'))) + && receivers.contains(returned) + && let Some((function, _)) = functions.last() + { + output.push(alias(language, returned, function, statement)); + } + depth += brace_delta(text); + while functions.last().is_some_and(|(_, opened)| depth <= *opened) { + functions.pop(); + } + } + output +} + +fn alias( + language: SourceLanguage, + router: &str, + name: &str, + statement: &Statement, +) -> SourceObservation { + mount_observation( + language, + SourceFramework::Express, + SymbolRef::Local(router.to_owned()), + Some(SymbolRef::Local(name.to_owned())), + None, + statement.lines, + ) +} + +fn declared_function(text: &str) -> Option<&str> { + let text = text + .trim_start_matches("export ") + .trim_start_matches("default ") + .trim_start_matches("async "); + if let Some(rest) = text.strip_prefix("function ") { + return simple_identifier(rest.split('(').next()?.trim()); + } + let rest = text + .strip_prefix("const ") + .or_else(|| text.strip_prefix("let "))?; + let (name, value) = rest.split_once('=')?; + (value.contains("=>") && text.trim_end().ends_with('{')) + .then(|| simple_identifier(name.trim())) + .flatten() +} + +/// Net `{` minus `}` outside string literals. +pub(crate) fn brace_delta(text: &str) -> i32 { + let mut quote = None; + let mut delta = 0; + let mut previous = '\0'; + for character in text.chars() { + match quote { + Some(active) if character == active && previous != '\\' => quote = None, + None if matches!(character, '\'' | '"' | '`') => quote = Some(character), + None if character == '{' => delta += 1, + None if character == '}' => delta -= 1, + Some(_) | None => {} + } + previous = character; + } + delta +} + +const GO_ROUTER_TYPES: [&str; 7] = [ + "*gin.Engine", + "*gin.RouterGroup", + "gin.IRouter", + "gin.IRoutes", + "chi.Router", + "*chi.Mux", + "*http.ServeMux", +]; + +const GO_ROUTER_FACTORIES: [&str; 5] = [ + "gin.Default(", + "gin.New(", + "chi.NewRouter(", + "chi.NewMux(", + "http.NewServeMux(", +]; + +/// Lexical router scope of one Go statement. +#[derive(Debug, Default)] +pub(crate) struct GoRouterScope { + function: Option, + parameters: BTreeSet, + frames: Vec<(String, String, i32)>, + package_routers: BTreeSet, + function_routers: BTreeSet, + depth: i32, +} + +impl GoRouterScope { + /// Advances the scope past `statement` and returns the mounts it declares. + pub(crate) fn observe( + &mut self, + statement: &Statement, + framework: SourceFramework, + ) -> Vec { + let text = statement.text.as_str(); + if self.depth == 0 + && let Some((name, parameters)) = go_function_signature(text) + { + self.function = Some(name); + self.parameters = parameters; + self.function_routers.clear(); + } + let mut output = Vec::new(); + let mount = |child: SymbolRef, parent: SymbolRef, prefix: Option<&str>| { + mount_observation( + SourceLanguage::Go, + framework, + child, + Some(parent), + prefix, + statement.lines, + ) + }; + if let Some((assigned, value)) = go_assignment(text) { + let is_router = GO_ROUTER_FACTORIES + .iter() + .any(|factory| value.contains(factory)); + let group = self.receiver_call(value, "Group"); + if is_router || group.is_some() { + if self.function.is_some() && self.depth > 0 { + self.function_routers.insert(assigned.to_owned()); + } else { + self.package_routers.insert(assigned.to_owned()); + } + } + if let Some((parent, arguments)) = group { + let prefix = arguments + .first() + .and_then(|argument| string_literal(argument)); + output.push(mount( + self.reference(assigned), + self.reference(parent), + prefix.as_deref(), + )); + } + } + output.extend(self.closure_frame(statement, framework)); + if let Some((parent, arguments)) = self.receiver_call(text, "Mount") + && let [prefix, child] = arguments.as_slice() + && let Some(prefix) = string_literal(prefix) + && let Some(child) = self.expression_reference(child) + { + output.push(mount(child, self.reference(parent), Some(&prefix))); + } + if let Some((parent, arguments)) = self.receiver_call(text, "Handle") + && let [_, handler] = arguments.as_slice() + && let Some(open) = handler.find("http.StripPrefix(") + && let [prefix, child] = + call_arguments(handler, open + "http.StripPrefix".len()).as_slice() + && let Some(prefix) = string_literal(prefix) + && let Some(child) = self.expression_reference(child) + { + output.push(mount(child, self.reference(parent), Some(&prefix))); + } + output.extend(self.function_call_mounts(text, framework, statement)); + if let Some(returned) = text.strip_prefix("return ").and_then(simple_identifier) + && self.is_router(returned) + && let Some(function) = &self.function + { + output.push(mount( + self.reference(returned), + SymbolRef::Function(function.clone()), + None, + )); + } + self.depth += brace_delta(text); + self.frames.retain(|(_, _, depth)| self.depth >= *depth); + if self.depth <= 0 { + self.depth = 0; + self.function = None; + self.parameters.clear(); + self.function_routers.clear(); + } + output + } + + /// Opens a Chi `Route` or `Group` closure frame and returns its mount. + fn closure_frame( + &mut self, + statement: &Statement, + framework: SourceFramework, + ) -> Option { + let text = statement.text.as_str(); + let variable = go_closure_router(text)?; + let (parent, prefix) = if let Some((parent, arguments)) = self.receiver_call(text, "Route") + { + let prefix = arguments + .first() + .and_then(|argument| string_literal(argument))?; + (parent, Some(prefix)) + } else { + (self.receiver_call(text, "Group")?.0, None) + }; + let key = self.frame_key(&variable, statement); + let observation = mount_observation( + SourceLanguage::Go, + framework, + SymbolRef::Local(key.clone()), + Some(self.reference(parent)), + prefix.as_deref(), + statement.lines, + ); + self.frames.push((variable, key, self.depth + 1)); + Some(observation) + } + + /// Router registered by `receiver` in the current scope. + pub(crate) fn reference(&self, receiver: &str) -> SymbolRef { + if let Some((_, key, _)) = self + .frames + .iter() + .rev() + .find(|(variable, _, _)| variable == receiver) + { + return SymbolRef::Local(key.clone()); + } + match &self.function { + Some(function) if self.parameters.contains(receiver) => { + SymbolRef::Function(function.clone()) + } + Some(function) if !self.package_routers.contains(receiver) => { + SymbolRef::Local(format!("{function}.{receiver}")) + } + _ => SymbolRef::Local(receiver.to_owned()), + } + } + + fn is_router(&self, name: &str) -> bool { + self.parameters.contains(name) + || self.function_routers.contains(name) + || self.package_routers.contains(name) + || self.frames.iter().any(|(variable, _, _)| variable == name) + } + + fn frame_key(&self, variable: &str, statement: &Statement) -> String { + format!( + "{}.{variable}@{}", + self.function.as_deref().unwrap_or_default(), + statement.lines.start + ) + } + + fn receiver_call<'t>(&self, text: &'t str, call: &str) -> Option<(&'t str, Vec<&'t str>)> { + let receiver = receiver_before(text, call)?; + if !self.is_router(receiver) { + return None; + } + let open = text.find(&format!("{receiver}.{call}("))? + receiver.len() + call.len() + 1; + Some((receiver, call_arguments(text, open))) + } + + fn expression_reference(&self, expression: &str) -> Option { + let expression = expression.trim(); + if let Some(name) = simple_identifier(expression) { + return self.is_router(name).then(|| self.reference(name)); + } + let callee = expression.strip_suffix("()")?; + callee + .split('.') + .all(|segment| simple_identifier(segment).is_some()) + .then(|| SymbolRef::Call(callee.to_owned())) + } + + /// Mounts a router function called with a router argument, such as `routes.Register(v1)`. + fn function_call_mounts( + &self, + text: &str, + framework: SourceFramework, + statement: &Statement, + ) -> Vec { + let Some(open) = text.find('(') else { + return Vec::new(); + }; + let callee = text[..open].trim(); + if callee.is_empty() + || callee.starts_with("func") + || !callee + .split('.') + .all(|segment| simple_identifier(segment).is_some()) + || callee + .split('.') + .next() + .is_some_and(|head| self.is_router(head)) + { + return Vec::new(); + } + call_arguments(text, open) + .into_iter() + .filter_map(|argument| { + if let Some(name) = simple_identifier(argument) { + return self.is_router(name).then(|| (self.reference(name), None)); + } + let (parent, arguments) = self.receiver_call(argument, "Group")?; + Some(( + self.reference(parent), + arguments + .first() + .and_then(|argument| string_literal(argument)), + )) + }) + .map(|(parent, prefix)| { + mount_observation( + SourceLanguage::Go, + framework, + SymbolRef::Call(callee.to_owned()), + Some(parent), + prefix.as_deref(), + statement.lines, + ) + }) + .collect() + } +} + +fn go_function_signature(text: &str) -> Option<(String, BTreeSet)> { + let rest = text.strip_prefix("func ")?; + let rest = if rest.starts_with('(') { + rest.split_once(')')?.1.trim_start() + } else { + rest + }; + let (name, after) = rest.split_once('(')?; + let name = simple_identifier(name.trim())?; + let parameters = after.split(')').next().unwrap_or_default(); + let mut routers = BTreeSet::new(); + let mut pending = Vec::new(); + for parameter in parameters.split(',') { + let mut parts = parameter.split_whitespace(); + let Some(variable) = parts.next() else { + continue; + }; + pending.push(variable.to_owned()); + if let Some(kind) = parts.next() { + if GO_ROUTER_TYPES.contains(&kind) { + routers.extend(pending.drain(..)); + } else { + pending.clear(); + } + } + } + Some((name.to_owned(), routers)) +} + +fn go_assignment(text: &str) -> Option<(&str, &str)> { + let (left, right) = text.split_once(":=").or_else(|| { + let (left, right) = text.split_once('=')?; + (!left.ends_with(['!', '<', '>', '='])).then_some((left, right)) + })?; + let left = left.trim().trim_start_matches("var ").trim(); + Some((simple_identifier(left)?, right.trim())) +} + +fn go_closure_router(text: &str) -> Option { + let rest = &text[text.find("func(")? + "func(".len()..]; + let mut parts = rest.split(')').next()?.split_whitespace(); + let variable = simple_identifier(parts.next()?)?; + parts + .next() + .is_some_and(|kind| GO_ROUTER_TYPES.contains(&kind)) + .then(|| variable.to_owned()) +} + +#[cfg(test)] +mod tests { + use super::{call_arguments, receiver_before}; + + #[test] + fn call_arguments_should_split_top_level_arguments_only() { + let text = "app.use('/api', auth({ a: 1, b: [2, 3] }), require('./x'))"; + let open = text.find('(').unwrap_or_default(); + + assert_eq!( + call_arguments(text, open), + ["'/api'", "auth({ a: 1, b: [2, 3] })", "require('./x')"] + ); + assert_eq!(receiver_before(" v1.GET(\"/x\", h)", "GET"), Some("v1")); + } +} diff --git a/crates/code-system-graph-core/src/test_links.rs b/crates/code-system-graph-core/src/test_links.rs index 6e7b224..04873a3 100644 --- a/crates/code-system-graph-core/src/test_links.rs +++ b/crates/code-system-graph-core/src/test_links.rs @@ -1,13 +1,8 @@ -use std::collections::BTreeMap; - use code_system_graph_model::{ - Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, Node, NodeId, NodeKind, Provenance, RepoId, stable_id + Evidence, EvidenceId, Node, NodeId, NodeKind, Provenance, RepoId, stable_id }; -use crate::linker::{http_link_ambiguity, sort_http_ambiguities}; -use crate::{ - BoundaryRole, ContractImplementationConfig, HttpBoundary, HttpLinkResolution, IntegrationTestConfig, LinkError, normalize_http_path -}; +use crate::{CallScope, ContractImplementationConfig, IntegrationTestConfig, normalize_http_path}; /// Declared cross-language test case and its validated HTTP target. #[derive(Debug, Clone, PartialEq)] @@ -18,6 +13,8 @@ pub struct DeclaredTestCase { pub method: String, /// Canonical target HTTP path. pub path: String, + /// Repositories allowed to provide the validated operation. + pub scope: CallScope, /// Evidence from the test source declaration. pub evidence: Evidence, } @@ -63,6 +60,7 @@ pub fn declared_test_case( }, method, path, + scope: CallScope::Workspace, evidence: Evidence { id: EvidenceId::new(stable_id("evidence", &evidence_key)), repo_id: Some(repo_id), @@ -124,200 +122,14 @@ pub fn declared_implementation( } } -/// Links declared tests to exact HTTP providers with bilateral evidence. -/// -/// Tests with no observed provider remain unlinked instead of inventing a target. -/// -/// Duplicate providers remain fail-closed for compatibility. Use -/// [`link_declared_tests_with_ambiguities`] to preserve ambiguities as data while continuing with -/// unrelated contracts. -/// -/// # Errors -/// -/// Returns [`LinkError::AmbiguousProvider`] instead of silently omitting an ambiguous relationship. -pub fn link_declared_tests( - tests: &[DeclaredTestCase], - boundaries: &[HttpBoundary], -) -> Result, LinkError> { - link_declared_tests_with_ambiguities(tests, boundaries).into_legacy_result() -} - -/// Links declared tests while preserving duplicate-provider decisions. -#[must_use] -pub fn link_declared_tests_with_ambiguities( - tests: &[DeclaredTestCase], - boundaries: &[HttpBoundary], -) -> HttpLinkResolution { - let mut providers: BTreeMap<(&str, &str), Vec<&HttpBoundary>> = BTreeMap::new(); - for provider in boundaries - .iter() - .filter(|boundary| boundary.role == BoundaryRole::Provider) - { - let candidates = providers - .entry((&provider.method, &provider.path)) - .or_default(); - if let Some(existing) = candidates - .iter() - .position(|candidate| candidate.node.id == provider.node.id) - { - if provider.evidence.confidence > candidates[existing].evidence.confidence { - candidates[existing] = provider; - } - } else { - candidates.push(provider); - } - } - let mut edges = Vec::new(); - let mut ambiguities = Vec::new(); - for test in tests { - let Some(candidates) = providers.get(&(test.method.as_str(), test.path.as_str())) else { - continue; - }; - if candidates.len() > 1 { - ambiguities.push(http_link_ambiguity( - &test.method, - &test.path, - candidates.iter().map(|candidate| &candidate.node.id), - )); - continue; - } - let provider = candidates[0]; - let edge_key = format!( - "{}:validates:{}", - test.node.id.as_str(), - provider.node.id.as_str() - ); - edges.push(Edge { - id: EdgeId::new(stable_id("edge", &edge_key)), - source: test.node.id.clone(), - target: provider.node.id.clone(), - kind: EdgeKind::Validates, - confidence: test.evidence.confidence.min(provider.evidence.confidence), - status: consensus_status(test.evidence.confidence.min(provider.evidence.confidence)), - evidence: vec![test.evidence.id.clone(), provider.evidence.id.clone()], - }); - } - edges.sort_by(|left, right| left.id.cmp(&right.id)); - sort_http_ambiguities(&mut ambiguities); - HttpLinkResolution { edges, ambiguities } -} - -/// Links HTTP provider contracts to declared source implementations. -/// -/// Duplicate providers remain fail-closed for compatibility. Use -/// [`link_declared_implementations_with_ambiguities`] to preserve ambiguities as data while -/// continuing with unrelated contracts. -/// -/// # Errors -/// -/// Returns [`LinkError::AmbiguousProvider`] instead of silently omitting an ambiguous relationship. -pub fn link_declared_implementations( - implementations: &[DeclaredImplementation], - boundaries: &[HttpBoundary], -) -> Result, LinkError> { - link_declared_implementations_with_ambiguities(implementations, boundaries).into_legacy_result() -} - -/// Links declared implementations while preserving duplicate-provider decisions. -#[must_use] -pub fn link_declared_implementations_with_ambiguities( - implementations: &[DeclaredImplementation], - boundaries: &[HttpBoundary], -) -> HttpLinkResolution { - let mut edges = BTreeMap::::new(); - let mut ambiguities = Vec::new(); - for implementation in implementations { - let candidates = boundaries - .iter() - .filter(|boundary| { - boundary.role == BoundaryRole::Provider - && boundary.node.repo_id == implementation.node.repo_id - && boundary.method == implementation.method - && boundary.path == implementation.path - }) - .fold( - BTreeMap::::new(), - |mut candidates, boundary| { - candidates - .entry(boundary.node.id.clone()) - .and_modify(|existing| { - if boundary.evidence.confidence > existing.evidence.confidence { - *existing = boundary; - } - }) - .or_insert(boundary); - candidates - }, - ) - .into_values() - .collect::>(); - if candidates.len() > 1 { - ambiguities.push(http_link_ambiguity( - &implementation.method, - &implementation.path, - candidates.iter().map(|candidate| &candidate.node.id), - )); - continue; - } - let Some(provider) = candidates.first() else { - continue; - }; - let edge_key = format!( - "{}:implemented_by:{}", - provider.node.id.as_str(), - implementation.node.id.as_str() - ); - let confidence = provider - .evidence - .confidence - .min(implementation.evidence.confidence); - let edge = Edge { - id: EdgeId::new(stable_id("edge", &edge_key)), - source: provider.node.id.clone(), - target: implementation.node.id.clone(), - kind: EdgeKind::ImplementedBy, - confidence, - status: consensus_status(confidence), - evidence: vec![ - provider.evidence.id.clone(), - implementation.evidence.id.clone(), - ], - }; - edges - .entry(edge.id.clone()) - .and_modify(|existing| merge_equivalent_implementation_edge(existing, &edge)) - .or_insert(edge); - } - sort_http_ambiguities(&mut ambiguities); - HttpLinkResolution { - edges: edges.into_values().collect(), - ambiguities, - } -} - -fn merge_equivalent_implementation_edge(existing: &mut Edge, candidate: &Edge) { - existing.confidence = existing.confidence.max(candidate.confidence); - existing.status = consensus_status(existing.confidence); - existing.evidence.extend(candidate.evidence.iter().cloned()); - existing.evidence.sort(); - existing.evidence.dedup(); -} - -fn consensus_status(confidence: f32) -> EpistemicStatus { - if confidence >= 1.0 { - EpistemicStatus::Confirmed - } else { - EpistemicStatus::Inferred - } -} - #[cfg(test)] mod tests { + use code_system_graph_model::{EdgeKind, EpistemicStatus, RepoId}; - use super::{declared_test_case, link_declared_implementations, link_declared_tests}; + use super::declared_test_case; use crate::{ - HttpContractConfig, IntegrationTestConfig, extract_openapi, parse_rust_source, source_observations_to_graph + AuthorityMap, HttpContractConfig, IntegrationTestConfig, extract_openapi, link_http_routes, parse_rust_source, source_observations_to_graph }; #[test] @@ -343,8 +155,8 @@ mod tests { ); let result = providers .map_err(|error| error.to_string()) - .and_then(|providers| { - link_declared_tests(&[test], &providers).map_err(|error| error.to_string()) + .map(|providers| { + link_http_routes(&providers, &[test], &[], &AuthorityMap::new()).edges }); assert!(matches!( @@ -370,8 +182,16 @@ async fn handler() {} ); let facts = source_observations_to_graph(&repo, "src/routes.rs", "content:routes", &observations); - let edges = link_declared_implementations(&facts.implementations, &facts.boundaries) - .expect("both exact declarations should remain linkable"); + let edges = link_http_routes( + &facts.boundaries, + &[], + &facts.implementations, + &AuthorityMap::new(), + ) + .edges + .into_iter() + .filter(|edge| edge.kind == EdgeKind::ImplementedBy) + .collect::>(); let mut consensus = edges .iter() .filter_map(|edge| { @@ -409,8 +229,16 @@ async fn handler() {} ); let facts = source_observations_to_graph(&repo, "src/routes.rs", "content:routes", &observations); - let edges = link_declared_implementations(&facts.implementations, &facts.boundaries) - .expect("matching declarations should merge into one implementation edge"); + let edges = link_http_routes( + &facts.boundaries, + &[], + &facts.implementations, + &AuthorityMap::new(), + ) + .edges + .into_iter() + .filter(|edge| edge.kind == EdgeKind::ImplementedBy) + .collect::>(); assert!(matches!( edges.as_slice(), diff --git a/crates/code-system-graph-core/src/url_template.rs b/crates/code-system-graph-core/src/url_template.rs new file mode 100644 index 0000000..3e9d634 --- /dev/null +++ b/crates/code-system-graph-core/src/url_template.rs @@ -0,0 +1,618 @@ +//! Client URL templates assembled from literals, constants, format strings, and parameters. +//! +//! Extractors evaluate URL expressions into [`UrlTemplate`]s. A template yields an exact client +//! path when its runtime values are confined to the scheme, the authority, or whole path segments; +//! values in whole segments become path parameters. Templates that depend on parameters of the +//! enclosing function are instantiated at call sites by repository-level composition. + +use std::collections::BTreeMap; + +use crate::{UrlPart, UrlTemplate}; + +/// Placeholder for one runtime value while a template is flattened to text. +const HOLE: char = '\u{E000}'; + +/// Calls whose result is their single argument rendered as text. +pub(crate) const CONVERSIONS: [&str; 9] = [ + "String", + "encodeURIComponent", + "encodeURI", + "Itoa", + "FormatInt", + "valueOf", + "toString", + "Sprint", + "PathEscape", +]; + +impl UrlTemplate { + /// Template of one literal. + #[must_use] + pub(crate) fn text(value: &str) -> Self { + let mut template = Self::default(); + template.push_text(value); + template + } + + /// Template of one runtime value. + #[must_use] + pub(crate) fn part(part: UrlPart) -> Self { + let mut template = Self::default(); + template.push(part); + template + } + + pub(crate) fn push_text(&mut self, value: &str) { + if value.is_empty() { + return; + } + if let Some(UrlPart::Text(last)) = self.parts.last_mut() { + last.push_str(value); + } else { + self.parts.push(UrlPart::Text(value.to_owned())); + } + } + + pub(crate) fn push(&mut self, part: UrlPart) { + match part { + UrlPart::Text(value) => self.push_text(&value), + part => self.parts.push(part), + } + } + + pub(crate) fn extend(&mut self, other: Self) { + for part in other.parts { + self.push(part); + } + } + + /// Literal value when the template has no runtime parts. + #[must_use] + pub(crate) fn as_literal(&self) -> Option<&str> { + match self.parts.as_slice() { + [] => Some(""), + [UrlPart::Text(value)] => Some(value), + _ => None, + } + } + + #[must_use] + pub(crate) fn has_parameters(&self) -> bool { + self.parts + .iter() + .any(|part| matches!(part, UrlPart::Parameter { .. })) + } + + /// Replaces parameters with the templates bound to their name or position. + /// + /// Unbound parameters become anonymous runtime values. + #[must_use] + pub(crate) fn bind(&self, arguments: &BoundArguments) -> Self { + let mut bound = Self::default(); + for part in &self.parts { + match part { + UrlPart::Parameter { name, index } => { + match arguments + .by_name + .get(name) + .or_else(|| arguments.by_index.get(index)) + { + Some(value) => bound.extend(value.clone()), + None => bound.push(UrlPart::Value(Some(name.clone()))), + } + } + part => bound.push(part.clone()), + } + } + bound + } + + /// Client literal with the same meaning for the client URL parser, or `None` when the path + /// depends on runtime values that are not whole path segments. + #[must_use] + pub(crate) fn client_literal(&self) -> Option { + if let Some(literal) = self.as_literal() { + return Some(literal.to_owned()); + } + let mut text = String::new(); + let mut names = Vec::new(); + for part in &self.parts { + match part { + UrlPart::Text(value) => text.push_str(value), + UrlPart::Parameter { name, .. } => { + text.push(HOLE); + names.push(Some(name.as_str())); + } + UrlPart::Value(name) => { + text.push(HOLE); + names.push(name.as_deref()); + } + } + } + let (origin, path) = split_origin(&text)?; + let path = path.split(['?', '#']).next().unwrap_or_default(); + if !path.is_empty() && !path.starts_with('/') { + return None; + } + let mut holes = names.into_iter(); + let mut origin_holes = origin + .chars() + .filter(|character| *character == HOLE) + .count(); + while origin_holes > 0 { + holes.next(); + origin_holes -= 1; + } + let mut literal = if origin.contains(HOLE) { + String::new() + } else { + origin.to_owned() + }; + if path.is_empty() { + literal.push('/'); + } + for (position, segment) in path.split('/').enumerate() { + if position > 0 { + literal.push('/'); + } + if !segment.contains(HOLE) { + literal.push_str(segment); + continue; + } + if segment.chars().count() != 1 { + return None; + } + let name = holes.next().flatten().map_or("value", parameter_name); + literal.push('{'); + literal.push_str(name); + literal.push('}'); + } + Some(literal) + } +} + +/// Arguments bound to a wrapper's parameters at one call site. +#[derive(Debug, Default)] +pub(crate) struct BoundArguments { + pub(crate) by_name: BTreeMap, + pub(crate) by_index: BTreeMap, +} + +/// Last identifier segment of a runtime value, used as a path parameter name. +fn parameter_name(name: &str) -> &str { + let name = name.rsplit(['.', ':']).next().unwrap_or(name); + if !name.is_empty() + && name + .chars() + .all(|character| character.is_ascii_alphanumeric() || character == '_') + { + name + } else { + "value" + } +} + +/// Splits flattened text into its scheme and authority and the path that follows. +/// +/// A leading runtime value directly followed by `/` is a base URL and contributes no path. A +/// runtime value that follows a literal authority starts the path unless it continues a host +/// label, a port, or user information. +fn split_origin(text: &str) -> Option<(&str, &str)> { + let after_scheme = text + .find("://") + .filter(|position| { + text[..*position] + .chars() + .all(|character| character == HOLE || character.is_ascii_alphabetic()) + }) + .map(|position| position + 3) + .or_else(|| text.starts_with("//").then_some(2)); + if let Some(start) = after_scheme { + let mut end = text[start..] + .find(['/', '?', '#']) + .map_or(text.len(), |offset| start + offset); + if let Some(hole) = text[start..end].find(HOLE).map(|offset| start + offset) + && hole > start + && !text[..hole].ends_with(['.', '-', ':', '@']) + { + end = hole; + } + if end == start { + return None; + } + return Some((&text[..end], &text[end..])); + } + if text.starts_with('/') { + return Some(("", text)); + } + let mut characters = text.chars(); + (characters.next() == Some(HOLE) && characters.next() == Some('/')) + .then(|| (&text[..HOLE.len_utf8()], &text[HOLE.len_utf8()..])) +} + +/// Expands a Python `str.format` or Rust `format!` string. +/// +/// `{}` takes the next positional argument, `{0}` a numbered one, and `{name}` a named one; +/// `resolve_inline` resolves names that are not passed explicitly, such as f-string expressions +/// and Rust inline arguments. `{{` and `}}` are literal braces. +pub(crate) fn brace_format( + format: &str, + positional: &[UrlTemplate], + named: &BTreeMap, + resolve_inline: &dyn Fn(&str) -> UrlTemplate, +) -> UrlTemplate { + let mut template = UrlTemplate::default(); + let mut next = 0_usize; + let mut characters = format.chars().peekable(); + while let Some(character) = characters.next() { + match character { + '{' if characters.peek() == Some(&'{') => { + characters.next(); + template.push_text("{"); + } + '}' if characters.peek() == Some(&'}') => { + characters.next(); + template.push_text("}"); + } + '{' => { + let mut field = String::new(); + let mut depth = 1_u32; + for inner in characters.by_ref() { + match inner { + '{' => depth += 1, + '}' => { + depth -= 1; + if depth == 0 { + break; + } + } + _ => {} + } + field.push(inner); + } + let name = field + .split(['!', ':', '=']) + .next() + .unwrap_or_default() + .trim(); + let value = if name.is_empty() { + let value = positional.get(next).cloned(); + next += 1; + value + } else if let Ok(position) = name.parse::() { + positional.get(position).cloned() + } else { + named + .get(name) + .cloned() + .or_else(|| Some(resolve_inline(name))) + }; + template.extend(value.unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(None)))); + } + character => { + let mut buffer = [0_u8; 4]; + template.push_text(character.encode_utf8(&mut buffer)); + } + } + } + template +} + +/// Expands a `printf`-style format string, as used by Go `fmt.Sprintf` and Java `String.format`. +pub(crate) fn printf_format(format: &str, arguments: &[UrlTemplate]) -> UrlTemplate { + let mut template = UrlTemplate::default(); + let mut next = 0_usize; + let mut characters = format.chars().peekable(); + while let Some(character) = characters.next() { + if character != '%' { + let mut buffer = [0_u8; 4]; + template.push_text(character.encode_utf8(&mut buffer)); + continue; + } + if characters.peek() == Some(&'%') { + characters.next(); + template.push_text("%"); + continue; + } + while characters + .peek() + .is_some_and(|flag| matches!(flag, '-' | '+' | '#' | ' ' | '0'..='9' | '.')) + { + characters.next(); + } + if characters.next().is_none() { + break; + } + let value = arguments + .get(next) + .cloned() + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(None))); + next += 1; + template.extend(value); + } + template +} + +/// Evaluates a JavaScript, TypeScript, Go, or Java string expression written as source text. +/// +/// Supported forms are string and template literals, `+` concatenation, parentheses, plain or +/// dotted identifiers, `fmt.Sprintf`, and `String.format`. Any other form yields `None`. +pub(crate) fn text_expression( + expression: &str, + resolve: &dyn Fn(&str) -> UrlTemplate, +) -> Option { + let mut template = UrlTemplate::default(); + for operand in split_top_level(expression.trim(), '+') { + template.extend(text_operand(operand.trim(), resolve)?); + } + (!template.parts.is_empty()).then_some(template) +} + +fn text_operand(operand: &str, resolve: &dyn Fn(&str) -> UrlTemplate) -> Option { + let first = operand.chars().next()?; + if matches!(first, '"' | '\'') { + let value = operand.strip_prefix(first)?.strip_suffix(first)?; + return (!value.contains(first) || value.contains('\\')).then(|| UrlTemplate::text(value)); + } + if first == '`' { + return template_literal(operand.strip_prefix('`')?.strip_suffix('`')?, resolve); + } + if first == '(' && operand.ends_with(')') { + return text_expression(&operand[1..operand.len() - 1], resolve); + } + for (call, printf) in [ + ("fmt.Sprintf(", true), + ("String.format(", true), + ("fmt.Sprint(", false), + ] { + if let Some(arguments) = operand + .strip_prefix(call) + .and_then(|rest| rest.strip_suffix(')')) + { + let arguments = split_top_level(arguments, ','); + let values = arguments + .iter() + .map(|argument| { + text_expression(argument, resolve) + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(None))) + }) + .collect::>(); + if !printf { + let mut joined = UrlTemplate::default(); + for value in values { + joined.extend(value); + } + return Some(joined); + } + let format = values.first()?.as_literal()?.to_owned(); + return Some(printf_format(&format, &values[1..])); + } + } + if let Some((callee, argument)) = operand + .strip_suffix(')') + .and_then(|call| call.split_once('(')) + && CONVERSIONS.contains(&callee.rsplit('.').next().unwrap_or(callee)) + { + return text_expression(argument, resolve); + } + let identifier = operand + .strip_suffix(".toString()") + .or_else(|| operand.strip_suffix(".String()")) + .unwrap_or(operand); + let is_identifier = identifier.split('.').all(|segment| { + !segment.is_empty() + && segment.chars().all(|character| { + character.is_ascii_alphanumeric() || matches!(character, '_' | '$') + }) + }); + is_identifier.then(|| resolve(identifier)) +} + +pub(crate) fn template_literal( + body: &str, + resolve: &dyn Fn(&str) -> UrlTemplate, +) -> Option { + let mut template = UrlTemplate::default(); + let mut rest = body; + while let Some(start) = rest.find("${") { + template.push_text(&rest[..start]); + let inner = &rest[start + 2..]; + let mut depth = 1_u32; + let end = inner.char_indices().find_map(|(offset, character)| { + match character { + '{' => depth += 1, + '}' => { + depth -= 1; + if depth == 0 { + return Some(offset); + } + } + _ => {} + } + None + })?; + template.extend( + text_expression(&inner[..end], resolve) + .unwrap_or_else(|| UrlTemplate::part(UrlPart::Value(None))), + ); + rest = &inner[end + 1..]; + } + template.push_text(rest); + Some(template) +} + +/// Splits `text` at `separator` outside quotes, template literals, and brackets. +pub(crate) fn split_top_level(text: &str, separator: char) -> Vec<&str> { + let mut parts = Vec::new(); + let mut depth = 0_i32; + let mut quote = None::; + let mut escaped = false; + let mut start = 0; + for (index, character) in text.char_indices() { + if let Some(active) = quote { + if escaped { + escaped = false; + } else if character == '\\' { + escaped = true; + } else if character == active { + quote = None; + } + continue; + } + match character { + '"' | '\'' | '`' => quote = Some(character), + '(' | '[' | '{' => depth += 1, + ')' | ']' | '}' => depth -= 1, + character if character == separator && depth == 0 => { + parts.push(&text[start..index]); + start = index + character.len_utf8(); + } + _ => {} + } + } + parts.push(&text[start..]); + parts +} + +#[cfg(test)] +mod tests { + use std::collections::BTreeMap; + + use super::{BoundArguments, brace_format, printf_format, text_expression}; + use crate::{UrlPart, UrlTemplate}; + + fn parameter(name: &str, index: usize) -> UrlPart { + UrlPart::Parameter { + name: name.to_owned(), + index, + } + } + + fn template(parts: Vec) -> UrlTemplate { + let mut template = UrlTemplate::default(); + for part in parts { + template.push(part); + } + template + } + + #[test] + fn runtime_values_in_origin_or_whole_segments_should_keep_exact_paths() { + let cases = [ + ( + vec![ + UrlPart::Value(Some("BASE_URL".to_owned())), + UrlPart::Text("/orders/".to_owned()), + UrlPart::Value(Some("order.id".to_owned())), + ], + Some("/orders/{id}"), + ), + ( + vec![ + UrlPart::Text("http://".to_owned()), + UrlPart::Value(Some("host".to_owned())), + UrlPart::Text(":8080/v1/users?page=".to_owned()), + UrlPart::Value(None), + ], + Some("/v1/users"), + ), + ( + vec![ + UrlPart::Text("https://api.test/items/".to_owned()), + parameter("item_id", 0), + ], + Some("https://api.test/items/{item_id}"), + ), + ( + vec![ + UrlPart::Text("/files/".to_owned()), + parameter("name", 0), + UrlPart::Text(".json".to_owned()), + ], + None, + ), + (vec![parameter("url", 0)], None), + ( + vec![UrlPart::Value(None), UrlPart::Text("orders".to_owned())], + None, + ), + ( + vec![ + UrlPart::Text("http://localhost:8080".to_owned()), + parameter("path", 0), + ], + None, + ), + ( + vec![ + UrlPart::Text("https://api.".to_owned()), + UrlPart::Value(Some("domain".to_owned())), + UrlPart::Text("/v1/items".to_owned()), + ], + Some("/v1/items"), + ), + ]; + + for (parts, expected) in cases { + assert_eq!(template(parts).client_literal().as_deref(), expected); + } + } + + #[test] + fn parameters_should_bind_by_name_or_position() { + let wrapper = template(vec![ + UrlPart::Value(Some("BASE".to_owned())), + parameter("path", 0), + ]); + let mut arguments = BoundArguments::default(); + arguments + .by_index + .insert(0, UrlTemplate::text("/api/orders/7")); + + assert_eq!( + wrapper.bind(&arguments).client_literal().as_deref(), + Some("/api/orders/7") + ); + let absolute = template(vec![ + UrlPart::Text("http://localhost:8080".to_owned()), + parameter("path", 0), + ]); + let mut named = BoundArguments::default(); + named + .by_name + .insert("path".to_owned(), UrlTemplate::text("/payments/1")); + assert_eq!( + absolute.bind(&named).client_literal().as_deref(), + Some("http://localhost:8080/payments/1") + ); + } + + #[test] + fn format_strings_should_expand_into_templates() { + let resolve = |name: &str| UrlTemplate::part(UrlPart::Value(Some(name.to_owned()))); + let base = UrlTemplate::text("http://orders:8080"); + let python = brace_format( + "{}/orders/{order_id}/items/{{literal}}", + std::slice::from_ref(&base), + &BTreeMap::new(), + &resolve, + ); + let go = printf_format("%s/v1/users/%d", &[base, resolve("id")]); + let script = text_expression("`${API}/users/${user.id}` + '/roles'", &resolve); + + assert_eq!( + python.client_literal().as_deref(), + Some("http://orders:8080/orders/{order_id}/items/{literal}") + ); + assert_eq!( + go.client_literal().as_deref(), + Some("http://orders:8080/v1/users/{id}") + ); + assert_eq!( + script + .and_then(|template| template.client_literal()) + .as_deref(), + Some("/users/{id}/roles") + ); + } +} diff --git a/crates/code-system-graph-core/tests/codegraph_provider_e2e.rs b/crates/code-system-graph-core/tests/codegraph_provider_e2e.rs index 082380a..a0bd369 100644 --- a/crates/code-system-graph-core/tests/codegraph_provider_e2e.rs +++ b/crates/code-system-graph-core/tests/codegraph_provider_e2e.rs @@ -77,7 +77,7 @@ async fn provider_should_probe_mcp_and_execute_all_supported_operations() { .expect("probe should succeed"); assert_eq!(capability.status, ProviderStatus::Available); - assert_eq!(capability.version.as_deref(), Some("1.5.0")); + assert_eq!(capability.version.as_deref(), Some("1.6.1")); assert!(capability.operations.iter().any(|operation| { operation.operation == ProviderOperation::LocalContext && operation.transport == ProviderTransport::Mcp @@ -139,7 +139,7 @@ async fn provider_should_probe_mcp_and_execute_all_supported_operations() { }) .await .expect("affected-tests query should succeed") - .expect("CodeGraph 1.5 supports affected tests"); + .expect("CodeGraph 1.6 supports affected tests"); assert_eq!(tests.affected_tests, vec!["tests/anchor.rs"]); } @@ -224,16 +224,16 @@ async fn cli_should_enforce_cancellation_and_output_caps() { #[tokio::test] #[ignore = "requires a user-installed and explicitly indexed CodeGraph checkout"] -async fn live_codegraph_1_5_should_match_the_declared_public_contract() { +async fn live_codegraph_1_6_should_match_the_declared_public_contract() { let project_path = PathBuf::from(env!("CARGO_MANIFEST_DIR")) .join("../..") .canonicalize() .expect("workspace root should be canonicalized"); let provider = CodeGraphProvider::new(CodeGraphConfig::default()).expect("default config should be valid"); - let request = ProviderRequest { + let request = || ProviderRequest { repo_id: RepoId::new("repo:live-smoke"), - project_path, + project_path: project_path.clone(), budget: ProviderBudget { timeout: Duration::from_secs(10), max_output_bytes: 1024 * 1024, @@ -242,11 +242,11 @@ async fn live_codegraph_1_5_should_match_the_declared_public_contract() { cancellation: CancellationToken::new(), }; let capability = provider - .probe(request) + .probe(request()) .await .expect("live CodeGraph probe should succeed"); - assert_eq!(capability.version.as_deref(), Some("1.5.0")); + assert_eq!(capability.version.as_deref(), Some("1.6.1")); assert!(matches!( capability.status, ProviderStatus::Available | ProviderStatus::Stale @@ -255,4 +255,45 @@ async fn live_codegraph_1_5_should_match_the_declared_public_contract() { operation.operation == ProviderOperation::LocalContext && operation.transport == ProviderTransport::Mcp })); + + let symbols = provider + .resolve_symbols(ResolveSymbolsRequest { + request: request(), + query: "CodeGraphProvider".to_owned(), + }) + .await + .expect("live symbol query should parse"); + assert!( + symbols + .symbols + .iter() + .any(|symbol| symbol.name == "CodeGraphProvider") + ); + provider + .get_local_neighbors(LocalNeighborsRequest { + request: request(), + symbol: "supports_cli_contract".to_owned(), + direction: LocalNeighborDirection::Callers, + }) + .await + .expect("live callers query should parse"); + provider + .get_local_impact(LocalImpactRequest { + request: request(), + symbol: "supports_cli_contract".to_owned(), + max_depth: 2, + }) + .await + .expect("live impact query should parse"); + provider + .get_affected_tests(AffectedTestsRequest { + request: request(), + changed_files: vec![ + "crates/code-system-graph-core/src/codegraph/contract.rs".to_owned(), + ], + max_depth: 3, + }) + .await + .expect("live affected-tests query should parse") + .expect("CodeGraph 1.6 supports affected tests"); } diff --git a/crates/code-system-graph-core/tests/scale_performance.rs b/crates/code-system-graph-core/tests/scale_performance.rs index 7a99eb3..8a80f81 100644 --- a/crates/code-system-graph-core/tests/scale_performance.rs +++ b/crates/code-system-graph-core/tests/scale_performance.rs @@ -1,4 +1,4 @@ -//! Release-mode scale acceptance for the `Code System Graph` 1.0.0 workstation targets. +//! Release-mode scale acceptance for the `Code System Graph` workstation targets. use std::collections::BTreeMap; use std::time::{Duration, Instant}; diff --git a/crates/code-system-graph-core/tests/source_compatibility.rs b/crates/code-system-graph-core/tests/source_compatibility.rs deleted file mode 100644 index 12283d7..0000000 --- a/crates/code-system-graph-core/tests/source_compatibility.rs +++ /dev/null @@ -1,33 +0,0 @@ -//! Compile-time guards for the intentionally breaking 1.1.0 public structs. - -use code_system_graph_core::{ExecutionPolicy, ExecutionPolicyOverrides, RepositoryConfig}; - -#[test] -fn version_1_1_0_public_structs_should_remain_constructible() { - let repository = RepositoryConfig { - path: "../service".to_owned(), - openapi: None, - http_consumers: None, - integration_tests: None, - implementations: None, - excludes: None, - include_defaults: None, - }; - let overrides = ExecutionPolicyOverrides::default(); - let policy = ExecutionPolicy { - max_scan_wall_time_ms: 1, - max_no_progress_time_ms: 1, - max_codegraph_sync_wall_time_ms_per_repo: 1, - max_worker_memory_bytes: 1, - graceful_termination_ms: 1, - watch_idle_timeout_ms: 1, - max_watch_session_wall_time_ms: 1, - min_watch_rescan_interval_ms: 1, - max_checkpoint_cache_bytes: 1, - ..ExecutionPolicy::default() - }; - - assert_eq!(repository.path, "../service"); - assert!(overrides.max_scan_wall_time_ms.is_none()); - assert_eq!(policy.max_scan_wall_time_ms, 1); -} diff --git a/crates/code-system-graph-hooks/Cargo.toml b/crates/code-system-graph-hooks/Cargo.toml index 8ff2353..3692753 100644 --- a/crates/code-system-graph-hooks/Cargo.toml +++ b/crates/code-system-graph-hooks/Cargo.toml @@ -21,14 +21,14 @@ disabled-strategies = ["quick-install", "compile"] pkg-fmt = "zip" [dependencies] -atomic-write-file = "0.3.0" -blake3 = "1.8.5" -libc = "0.2" +atomic-write-file = "0.3.1" +blake3 = "1.8.7" +libc = "0.2.190" nix = { version = "0.31.3", features = ["fs"] } schemars = "1.2.2" serde = { version = "1.0.229", features = ["derive"] } serde_json = "1.0.151" -thiserror = "2.0.19" +thiserror = "2.0.21" [lints] workspace = true diff --git a/crates/code-system-graph-model/Cargo.toml b/crates/code-system-graph-model/Cargo.toml index a3f4ad7..865e06b 100644 --- a/crates/code-system-graph-model/Cargo.toml +++ b/crates/code-system-graph-model/Cargo.toml @@ -12,12 +12,10 @@ keywords.workspace = true categories.workspace = true [dependencies] -blake3 = "1.8.5" -camino = { version = "1.2.5", features = ["serde1"] } +blake3 = "1.8.7" schemars = "1.2.2" semver = { version = "1.0.28", features = ["serde"] } serde = { version = "1.0.229", features = ["derive"] } -time = { version = "0.3.54", features = ["serde"] } [lints] workspace = true diff --git a/crates/code-system-graph-model/src/lib.rs b/crates/code-system-graph-model/src/lib.rs index 2f06fa5..86efac3 100644 --- a/crates/code-system-graph-model/src/lib.rs +++ b/crates/code-system-graph-model/src/lib.rs @@ -48,7 +48,7 @@ string_id!( /// Lossless platform encoding used for a native filesystem path. #[derive( - Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize, JsonSchema, + Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize, JsonSchema, )] #[serde(rename_all = "snake_case")] pub enum NativePathEncoding { @@ -61,7 +61,9 @@ pub enum NativePathEncoding { } /// Lossless native path plus a diagnostic-only display form. -#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize, JsonSchema)] +#[derive( + Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize, JsonSchema, +)] pub struct NativePath { /// Platform-specific lossless encoding. pub encoding: NativePathEncoding, @@ -179,8 +181,6 @@ pub enum ArtifactChangeKind { Modified, /// Artifact is absent from the current scan. Deleted, - /// Artifact fingerprint is unchanged. - Unchanged, } /// Planned incremental action for one artifact identity. @@ -622,6 +622,82 @@ pub struct RepositoryCoverageGap { pub reason: String, } +/// Why an HTTP consumer or test call has no provider edge. +#[derive( + Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize, JsonSchema, +)] +#[serde(rename_all = "snake_case")] +pub enum HttpLinkGapReason { + /// No workspace provider matches the method and path within the call's scope. + NoProvider, + /// Several equally specific providers match; the candidates are listed. + Ambiguous, + /// The call names a host outside the workspace. + External, +} + +impl HttpLinkGapReason { + /// Stable persisted name. + #[must_use] + pub const fn as_str(self) -> &'static str { + match self { + Self::NoProvider => "no_provider", + Self::Ambiguous => "ambiguous", + Self::External => "external", + } + } + + /// Parses a persisted name. + #[must_use] + pub fn parse(value: &str) -> Option { + match value { + "no_provider" => Some(Self::NoProvider), + "ambiguous" => Some(Self::Ambiguous), + "external" => Some(Self::External), + _ => None, + } + } +} + +/// An HTTP consumer or test call that produced no provider edge. +#[derive( + Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize, JsonSchema, +)] +pub struct HttpLinkGap { + /// Consumer boundary or test case node. + pub caller: NodeId, + /// HTTP method of the call. + pub method: String, + /// Concrete or templated path of the call. + pub path: String, + /// Why the call has no provider edge. + pub reason: HttpLinkGapReason, + /// Equally specific providers of an ambiguous call, in stable order. + pub candidates: Vec, +} + +/// Counts of HTTP consumer and test calls by link outcome. +#[derive(Debug, Clone, Copy, Default, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] +pub struct HttpLinkCoverage { + /// Calls linked to exactly one provider. + pub linked: u64, + /// Calls without a matching provider. + pub no_provider: u64, + /// Calls with several equally specific providers. + pub ambiguous: u64, + /// Calls to hosts outside the workspace, which are not coverage gaps. + pub external: u64, +} + +/// HTTP link outcome of one published graph. +#[derive(Debug, Clone, Default, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] +pub struct HttpLinkReport { + /// Counts of calls by outcome. + pub coverage: HttpLinkCoverage, + /// Calls without a provider edge, in stable order. + pub gaps: Vec, +} + /// Status of a public tool result. #[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, JsonSchema)] #[serde(rename_all = "snake_case")] diff --git a/crates/code-system-graph-store-sqlite/Cargo.toml b/crates/code-system-graph-store-sqlite/Cargo.toml index a736e04..6efac5a 100644 --- a/crates/code-system-graph-store-sqlite/Cargo.toml +++ b/crates/code-system-graph-store-sqlite/Cargo.toml @@ -12,20 +12,20 @@ keywords.workspace = true categories = ["database", "development-tools"] [dependencies] -blake3 = "1.8.5" -code-system-graph-model = { version = "1.1.0", path = "../code-system-graph-model" } -rusqlite = { version = "0.40.1", features = ["backup", "bundled", "hooks"] } +blake3 = "1.8.7" +code-system-graph-model = { version = "1.2.0", path = "../code-system-graph-model" } +rusqlite = { version = "0.40.2", features = ["backup", "bundled", "hooks"] } same-file = "1.0.6" serde_json = "1.0.151" sysinfo = { version = "0.39.6", default-features = false, features = ["system"] } tempfile = "3.27.0" -thiserror = "2.0.19" +thiserror = "2.0.21" [lints] workspace = true [target.'cfg(unix)'.dependencies] -libc = "0.2" +libc = "0.2.190" [target.'cfg(windows)'.dependencies] windows-sys = { version = "0.61.2", features = [ diff --git a/crates/code-system-graph-store-sqlite/migrations/0001_initial.sql b/crates/code-system-graph-store-sqlite/migrations/0001_initial.sql index 1063eb3..0350275 100644 --- a/crates/code-system-graph-store-sqlite/migrations/0001_initial.sql +++ b/crates/code-system-graph-store-sqlite/migrations/0001_initial.sql @@ -1,7 +1,7 @@ PRAGMA foreign_keys = ON; CREATE TABLE IF NOT EXISTS schema_metadata ( - version INTEGER PRIMARY KEY, + schema_id TEXT PRIMARY KEY, instance_id TEXT NOT NULL, applied_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP ); @@ -53,42 +53,57 @@ CREATE INDEX IF NOT EXISTS workspace_repositories_repo_idx ON workspace_repositories(repo_id); CREATE TABLE IF NOT EXISTS repo_snapshots ( - id TEXT PRIMARY KEY, - workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, - created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, - is_current INTEGER NOT NULL DEFAULT 0 CHECK (is_current IN (0, 1)) + workspace_name TEXT PRIMARY KEY REFERENCES workspaces(name) ON DELETE CASCADE, + id TEXT NOT NULL UNIQUE, + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP ); -CREATE UNIQUE INDEX IF NOT EXISTS one_current_snapshot_per_workspace -ON repo_snapshots(workspace_name) -WHERE is_current = 1; - CREATE UNIQUE INDEX IF NOT EXISTS repo_snapshots_id_workspace_idx ON repo_snapshots(id, workspace_name); CREATE TABLE IF NOT EXISTS nodes ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + node_rowid INTEGER PRIMARY KEY, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, id TEXT NOT NULL, kind TEXT NOT NULL, repo_id TEXT, stable_key TEXT NOT NULL, label TEXT NOT NULL, - PRIMARY KEY (snapshot_id, id) + UNIQUE (workspace_name, id) ); -CREATE INDEX IF NOT EXISTS nodes_kind_idx ON nodes(kind); +CREATE INDEX IF NOT EXISTS nodes_kind_idx ON nodes(workspace_name, kind); CREATE INDEX IF NOT EXISTS nodes_repo_idx ON nodes(repo_id); -CREATE INDEX IF NOT EXISTS nodes_stable_key_idx ON nodes(stable_key); +CREATE INDEX IF NOT EXISTS nodes_stable_key_idx ON nodes(workspace_name, stable_key); CREATE VIRTUAL TABLE IF NOT EXISTS nodes_fts USING fts5( - snapshot_id UNINDEXED, - node_id UNINDEXED, label, stable_key ); +CREATE TRIGGER IF NOT EXISTS nodes_fts_insert +AFTER INSERT ON nodes +BEGIN + INSERT INTO nodes_fts(rowid, label, stable_key) + VALUES (NEW.node_rowid, NEW.label, NEW.stable_key); +END; + +CREATE TRIGGER IF NOT EXISTS nodes_fts_update +AFTER UPDATE OF label, stable_key ON nodes +BEGIN + DELETE FROM nodes_fts WHERE rowid = OLD.node_rowid; + INSERT INTO nodes_fts(rowid, label, stable_key) + VALUES (NEW.node_rowid, NEW.label, NEW.stable_key); +END; + +CREATE TRIGGER IF NOT EXISTS nodes_fts_delete +AFTER DELETE ON nodes +BEGIN + DELETE FROM nodes_fts WHERE rowid = OLD.node_rowid; +END; + CREATE TABLE IF NOT EXISTS evidence ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, id TEXT NOT NULL, repo_id TEXT, file_path TEXT, @@ -105,53 +120,83 @@ CREATE TABLE IF NOT EXISTS evidence ( confidence REAL NOT NULL CHECK (confidence >= 0.0 AND confidence <= 1.0), content_hash TEXT, observed_at_commit TEXT, - PRIMARY KEY (snapshot_id, id) -); + PRIMARY KEY (workspace_name, id) +) WITHOUT ROWID; -CREATE INDEX IF NOT EXISTS evidence_snapshot_file_lines_idx -ON evidence(snapshot_id, repo_id, file_path, start_line, end_line); +CREATE INDEX IF NOT EXISTS evidence_file_lines_idx +ON evidence(workspace_name, repo_id, file_path, start_line, end_line); CREATE TABLE IF NOT EXISTS edges ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL, id TEXT NOT NULL, source_node_id TEXT NOT NULL, target_node_id TEXT NOT NULL, kind TEXT NOT NULL, confidence REAL NOT NULL CHECK (confidence >= 0.0 AND confidence <= 1.0), epistemic_status TEXT NOT NULL, - PRIMARY KEY (snapshot_id, id), - FOREIGN KEY (snapshot_id, source_node_id) REFERENCES nodes(snapshot_id, id), - FOREIGN KEY (snapshot_id, target_node_id) REFERENCES nodes(snapshot_id, id) -); - -CREATE INDEX IF NOT EXISTS edges_source_idx ON edges(source_node_id); -CREATE INDEX IF NOT EXISTS edges_target_idx ON edges(target_node_id); + PRIMARY KEY (workspace_name, id), + FOREIGN KEY (workspace_name, source_node_id) + REFERENCES nodes(workspace_name, id) ON DELETE CASCADE, + FOREIGN KEY (workspace_name, target_node_id) + REFERENCES nodes(workspace_name, id) ON DELETE CASCADE +) WITHOUT ROWID; + +CREATE INDEX IF NOT EXISTS edges_source_idx ON edges(workspace_name, source_node_id); +CREATE INDEX IF NOT EXISTS edges_target_idx ON edges(workspace_name, target_node_id); CREATE INDEX IF NOT EXISTS edges_kind_idx ON edges(kind); -CREATE INDEX IF NOT EXISTS edges_confidence_idx ON edges(confidence); CREATE TABLE IF NOT EXISTS edge_evidence ( - snapshot_id TEXT NOT NULL, + workspace_name TEXT NOT NULL, edge_id TEXT NOT NULL, evidence_id TEXT NOT NULL, - PRIMARY KEY (snapshot_id, edge_id, evidence_id), - FOREIGN KEY (snapshot_id, edge_id) REFERENCES edges(snapshot_id, id) ON DELETE CASCADE, - FOREIGN KEY (snapshot_id, evidence_id) REFERENCES evidence(snapshot_id, id) ON DELETE CASCADE -); + PRIMARY KEY (workspace_name, edge_id, evidence_id), + FOREIGN KEY (workspace_name, edge_id) + REFERENCES edges(workspace_name, id) ON DELETE CASCADE, + FOREIGN KEY (workspace_name, evidence_id) + REFERENCES evidence(workspace_name, id) ON DELETE CASCADE +) WITHOUT ROWID; + +CREATE INDEX IF NOT EXISTS edge_evidence_evidence_idx +ON edge_evidence(workspace_name, evidence_id); CREATE TABLE IF NOT EXISTS repository_snapshot_freshness ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, repo_id TEXT NOT NULL REFERENCES repositories(id), checkout_id TEXT NOT NULL REFERENCES repository_checkouts(id), head_commit TEXT, manifest_hash TEXT NOT NULL, state TEXT NOT NULL, reason TEXT, - PRIMARY KEY (snapshot_id, repo_id, checkout_id) + PRIMARY KEY (workspace_name, repo_id, checkout_id) ); +CREATE TABLE IF NOT EXISTS http_link_coverage ( + workspace_name TEXT PRIMARY KEY REFERENCES workspaces(name) ON DELETE CASCADE, + linked INTEGER NOT NULL CHECK (linked >= 0), + no_provider INTEGER NOT NULL CHECK (no_provider >= 0), + ambiguous INTEGER NOT NULL CHECK (ambiguous >= 0), + external INTEGER NOT NULL CHECK (external >= 0) +) STRICT; + +CREATE TABLE IF NOT EXISTS http_link_gaps ( + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, + caller_node_id TEXT NOT NULL + CHECK (length(CAST(caller_node_id AS BLOB)) BETWEEN 1 AND 2048), + method TEXT NOT NULL + CHECK (length(CAST(method AS BLOB)) BETWEEN 1 AND 32), + path TEXT NOT NULL + CHECK (length(CAST(path AS BLOB)) BETWEEN 1 AND 4096), + reason TEXT NOT NULL + CHECK (reason IN ('no_provider', 'ambiguous', 'external')), + candidates_json TEXT NOT NULL + CHECK (json_valid(candidates_json)) + CHECK (json_type(candidates_json) = 'array'), + PRIMARY KEY (workspace_name, caller_node_id, method, path, reason) +) STRICT, WITHOUT ROWID; + CREATE TABLE IF NOT EXISTS extractor_runs ( - id TEXT PRIMARY KEY, - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, + id TEXT NOT NULL, repo_id TEXT NOT NULL REFERENCES repositories(id), checkout_id TEXT REFERENCES repository_checkouts(id), extractor TEXT NOT NULL, @@ -160,12 +205,10 @@ CREATE TABLE IF NOT EXISTS extractor_runs ( discovered_files INTEGER NOT NULL DEFAULT 0, parsed_files INTEGER NOT NULL DEFAULT 0, skipped_files INTEGER NOT NULL DEFAULT 0, - elapsed_ms INTEGER NOT NULL DEFAULT 0 + elapsed_ms INTEGER NOT NULL DEFAULT 0, + PRIMARY KEY (workspace_name, id) ); -CREATE INDEX IF NOT EXISTS extractor_runs_snapshot_idx -ON extractor_runs(snapshot_id); - CREATE TABLE IF NOT EXISTS audit_events ( id TEXT PRIMARY KEY, workspace_name TEXT REFERENCES workspaces(name) ON DELETE SET NULL, @@ -176,7 +219,7 @@ CREATE TABLE IF NOT EXISTS audit_events ( ); CREATE TABLE IF NOT EXISTS artifact_fingerprints ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, repo_id TEXT NOT NULL REFERENCES repositories(id), checkout_id TEXT NOT NULL REFERENCES repository_checkouts(id), path_encoding TEXT NOT NULL, @@ -186,38 +229,17 @@ CREATE TABLE IF NOT EXISTS artifact_fingerprints ( content_hash TEXT NOT NULL, size_bytes INTEGER NOT NULL CHECK (size_bytes >= 0), PRIMARY KEY ( - snapshot_id, + workspace_name, repo_id, checkout_id, path_encoding, relative_path, extractor ) -); - -CREATE INDEX IF NOT EXISTS artifact_fingerprints_checkout_idx -ON artifact_fingerprints(checkout_id, extractor); - -CREATE TABLE IF NOT EXISTS extractor_run_inputs ( - run_id TEXT NOT NULL REFERENCES extractor_runs(id) ON DELETE CASCADE, - repo_id TEXT NOT NULL, - checkout_id TEXT NOT NULL, - path_encoding TEXT NOT NULL, - relative_path BLOB NOT NULL, - extractor TEXT NOT NULL, - content_hash TEXT NOT NULL, - PRIMARY KEY ( - run_id, - repo_id, - checkout_id, - path_encoding, - relative_path, - extractor - ) -); +) WITHOUT ROWID; CREATE TABLE IF NOT EXISTS extractor_batches ( - snapshot_id TEXT NOT NULL REFERENCES repo_snapshots(id) ON DELETE CASCADE, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, repo_id TEXT NOT NULL REFERENCES repositories(id), checkout_id TEXT NOT NULL REFERENCES repository_checkouts(id), path_encoding TEXT NOT NULL, @@ -230,9 +252,10 @@ CREATE TABLE IF NOT EXISTS extractor_batches ( budget_fingerprint TEXT NOT NULL, source_was_lossy INTEGER NOT NULL CHECK (source_was_lossy IN (0, 1)), output_count INTEGER NOT NULL CHECK (output_count >= 0), + payload_hash TEXT NOT NULL, payload BLOB NOT NULL CHECK (json_valid(CAST(payload AS TEXT))), PRIMARY KEY ( - snapshot_id, + workspace_name, repo_id, checkout_id, path_encoding, @@ -241,11 +264,9 @@ CREATE TABLE IF NOT EXISTS extractor_batches ( ) ); -CREATE INDEX IF NOT EXISTS extractor_batches_source_idx -ON extractor_batches(repo_id, checkout_id, extractor); - CREATE TABLE IF NOT EXISTS community_snapshots ( - snapshot_id TEXT PRIMARY KEY REFERENCES repo_snapshots(id) ON DELETE CASCADE, + snapshot_id TEXT PRIMARY KEY, + workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, engine_version TEXT NOT NULL, algorithm TEXT NOT NULL CHECK (json_valid(algorithm)) @@ -255,8 +276,8 @@ CREATE TABLE IF NOT EXISTS community_snapshots ( CHECK (length(CAST(config_json AS BLOB)) <= 65536) ); -CREATE INDEX IF NOT EXISTS community_snapshots_snapshot_idx -ON community_snapshots(snapshot_id); +CREATE INDEX IF NOT EXISTS community_snapshots_workspace_idx +ON community_snapshots(workspace_name); CREATE TABLE IF NOT EXISTS communities ( snapshot_id TEXT NOT NULL, @@ -272,12 +293,6 @@ CREATE TABLE IF NOT EXISTS communities ( FOREIGN KEY (snapshot_id) REFERENCES community_snapshots(snapshot_id) ON DELETE CASCADE ); -CREATE INDEX IF NOT EXISTS communities_snapshot_idx -ON communities(snapshot_id); - -CREATE INDEX IF NOT EXISTS communities_snapshot_community_idx -ON communities(snapshot_id, id); - CREATE TABLE IF NOT EXISTS community_memberships ( snapshot_id TEXT NOT NULL, community_id TEXT NOT NULL, @@ -286,22 +301,11 @@ CREATE TABLE IF NOT EXISTS community_memberships ( PRIMARY KEY (snapshot_id, community_id, node_id), UNIQUE (snapshot_id, community_id, member_order), FOREIGN KEY (snapshot_id, community_id) - REFERENCES communities(snapshot_id, id) ON DELETE CASCADE, - FOREIGN KEY (snapshot_id, node_id) - REFERENCES nodes(snapshot_id, id) ON DELETE CASCADE + REFERENCES communities(snapshot_id, id) ON DELETE CASCADE ); -CREATE INDEX IF NOT EXISTS community_memberships_snapshot_idx -ON community_memberships(snapshot_id); - -CREATE INDEX IF NOT EXISTS community_memberships_community_idx -ON community_memberships(snapshot_id, community_id); - -CREATE INDEX IF NOT EXISTS community_memberships_member_idx -ON community_memberships(snapshot_id, node_id); - CREATE TABLE IF NOT EXISTS manual_links ( - snapshot_id TEXT NOT NULL, + workspace_name TEXT NOT NULL, id TEXT NOT NULL CHECK (length(CAST(id AS BLOB)) BETWEEN 1 AND 2048) CHECK (length(trim(id)) > 0), @@ -324,24 +328,14 @@ CREATE TABLE IF NOT EXISTS manual_links ( CHECK (json_valid(CAST(decision_json AS TEXT))), config_version INTEGER NOT NULL CHECK (config_version BETWEEN 1 AND 2147483647), - PRIMARY KEY (snapshot_id, id), - UNIQUE (snapshot_id, source_node_id, target_node_id, kind), - FOREIGN KEY (snapshot_id) REFERENCES repo_snapshots(id) ON DELETE CASCADE, - FOREIGN KEY (snapshot_id, source_node_id) - REFERENCES nodes(snapshot_id, id) ON DELETE CASCADE, - FOREIGN KEY (snapshot_id, target_node_id) - REFERENCES nodes(snapshot_id, id) ON DELETE CASCADE + PRIMARY KEY (workspace_name, id), + UNIQUE (workspace_name, source_node_id, target_node_id, kind), + FOREIGN KEY (workspace_name, source_node_id) + REFERENCES nodes(workspace_name, id) ON DELETE CASCADE, + FOREIGN KEY (workspace_name, target_node_id) + REFERENCES nodes(workspace_name, id) ON DELETE CASCADE ) STRICT, WITHOUT ROWID; -CREATE INDEX IF NOT EXISTS manual_links_source_idx -ON manual_links(snapshot_id, source_node_id); - -CREATE INDEX IF NOT EXISTS manual_links_target_idx -ON manual_links(snapshot_id, target_node_id); - -CREATE INDEX IF NOT EXISTS manual_links_disposition_idx -ON manual_links(snapshot_id, disposition, kind); - CREATE TABLE IF NOT EXISTS provider_capabilities ( workspace_name TEXT NOT NULL REFERENCES workspaces(name) ON DELETE CASCADE, @@ -431,9 +425,3 @@ BEGIN LIMIT -1 OFFSET 1024 ); END; - -CREATE TRIGGER IF NOT EXISTS repo_snapshots_delete_fts -AFTER DELETE ON repo_snapshots -BEGIN - DELETE FROM nodes_fts WHERE snapshot_id = OLD.id; -END; diff --git a/crates/code-system-graph-store-sqlite/src/access_lock.rs b/crates/code-system-graph-store-sqlite/src/access_lock.rs index 311d346..64cd031 100644 --- a/crates/code-system-graph-store-sqlite/src/access_lock.rs +++ b/crates/code-system-graph-store-sqlite/src/access_lock.rs @@ -63,7 +63,10 @@ fn canonical_access_path(database_path: &Path) -> Result { "database path must include a file name", ), })?; - let parent = database_path.parent().unwrap_or_else(|| Path::new(".")); + let parent = database_path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); if fs::symlink_metadata(database_path).is_err() { fs::create_dir_all(parent).map_err(|source| StoreError::Io { path: parent.to_path_buf(), diff --git a/crates/code-system-graph-store-sqlite/src/backup_restore.rs b/crates/code-system-graph-store-sqlite/src/backup_restore.rs index 0ff7459..1625f93 100644 --- a/crates/code-system-graph-store-sqlite/src/backup_restore.rs +++ b/crates/code-system-graph-store-sqlite/src/backup_restore.rs @@ -8,7 +8,7 @@ use rusqlite::{Connection, ErrorCode, OpenFlags, OptionalExtension}; use same_file::Handle; use super::{ - LATEST_SCHEMA_VERSION, RestoreReport, StoreError, StoreLock, open_read_only_connection, schema_version, validate_exact_schema + RestoreReport, StoreError, StoreLock, open_read_only_connection, stored_schema_id, validate_exact_schema }; #[cfg(windows)] use crate::file_permissions::current_user_sid_string; @@ -61,7 +61,10 @@ fn copy_database_file_snapshot( destination: &Path, ) -> Result { let destination = canonical_destination_path(destination)?; - let parent = destination.parent().unwrap_or_else(|| Path::new(".")); + let parent = destination + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); let directory = OwnedStagingDirectory::create(parent).map_err(|source| StoreError::Io { path: parent.to_path_buf(), source, @@ -278,7 +281,10 @@ fn create_private_staging_directory(_path: &Path) -> std::io::Result<()> { impl StagedDatabase { fn create(destination: &Path) -> Result<(Self, Connection), StoreError> { - let parent = destination.parent().unwrap_or_else(|| Path::new(".")); + let parent = destination + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); let directory = OwnedStagingDirectory::create(parent).map_err(|source| StoreError::Io { path: parent.to_path_buf(), source, @@ -402,7 +408,10 @@ fn canonical_destination_path(path: &Path) -> Result { "database path must include a file name", ), })?; - let parent = path.parent().unwrap_or_else(|| Path::new(".")); + let parent = path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); fs::create_dir_all(parent).map_err(|source| StoreError::Io { path: parent.to_path_buf(), source, @@ -783,17 +792,17 @@ pub(super) fn restore_database( configure_staged_restore(&destination)?; validate_backup(&destination, staging.path())?; restrict_store_permissions(staging.path())?; - schema_version(&destination) + stored_schema_id(&destination) })(); drop(destination); let restore_result = match restore_result { - Ok(schema_version) => source.finish().map(|()| schema_version), + Ok(schema_id) => source.finish().map(|()| schema_id), Err(error) => Err(error), }; drop(source); - let schema_version = match restore_result { - Ok(schema_version) => schema_version, + let schema_id = match restore_result { + Ok(schema_id) => schema_id, Err(operation) => return Err(staging.cleanup_error(operation)), }; @@ -842,7 +851,7 @@ pub(super) fn restore_database( Ok(RestoreReport { source_path: backup_path.to_path_buf(), safety_backup_path, - schema_version, + schema_id, }) } @@ -1002,7 +1011,10 @@ fn canonical_path_for_comparison(path: &Path) -> std::io::Result { Ok(canonical) => Ok(canonical), Err(source) if source.kind() == std::io::ErrorKind::NotFound => { let file_name = path.file_name().ok_or(source)?; - let parent = path.parent().unwrap_or_else(|| Path::new(".")); + let parent = path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); Ok(fs::canonicalize(parent)?.join(file_name)) } Err(source) => Err(source), @@ -1067,10 +1079,7 @@ fn validate_backup(connection: &Connection, path: &Path) -> Result<(), StoreErro validate_exact_schema(connection).map_err(|error| match error { StoreError::InvalidSchema => StoreError::InvalidBackup { path: path.to_path_buf(), - reason: format!( - "schema objects do not match the exact supported contract version \ - {LATEST_SCHEMA_VERSION}" - ), + reason: "schema objects do not match the exact supported schema".to_owned(), }, other => other, }) diff --git a/crates/code-system-graph-store-sqlite/src/backup_restore/tests.rs b/crates/code-system-graph-store-sqlite/src/backup_restore/tests.rs index 8392ef3..8ada9ec 100644 --- a/crates/code-system-graph-store-sqlite/src/backup_restore/tests.rs +++ b/crates/code-system-graph-store-sqlite/src/backup_restore/tests.rs @@ -299,9 +299,13 @@ fn restore_should_preserve_replaced_database_as_safety_backup() ( restored.snapshot_id, replaced.snapshot_id, - report.schema_version, + report.schema_id.as_str(), ), - ("snapshot:before".to_owned(), "snapshot:after".to_owned(), 2) + ( + "snapshot:before".to_owned(), + "snapshot:after".to_owned(), + crate::schema_identity() + ) ); Ok(()) } @@ -712,10 +716,10 @@ fn restore_should_reject_backup_with_foreign_key_violations() connection.execute_batch( "PRAGMA foreign_keys = OFF; INSERT INTO edges( - snapshot_id, id, source_node_id, target_node_id, kind, confidence, + workspace_name, id, source_node_id, target_node_id, kind, confidence, epistemic_status ) VALUES ( - 'snapshot:0', 'edge:corrupt', 'missing:a', 'missing:b', + 'commerce', 'edge:corrupt', 'missing:a', 'missing:b', '\"calls_remote\"', 1.0, '\"confirmed\"' ); PRAGMA foreign_keys = ON;", @@ -755,10 +759,10 @@ fn restore_should_copy_from_pinned_source_snapshot() -> Result<(), Box PathBuf { } pub(crate) fn prepare_database_file(path: &Path) -> Result<(), StoreError> { - if let Some(parent) = path.parent() { + if let Some(parent) = path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + { fs::create_dir_all(parent).map_err(|source| StoreError::Io { path: parent.to_path_buf(), source, diff --git a/crates/code-system-graph-store-sqlite/src/lib.rs b/crates/code-system-graph-store-sqlite/src/lib.rs index 717907d..0aa4ca5 100644 --- a/crates/code-system-graph-store-sqlite/src/lib.rs +++ b/crates/code-system-graph-store-sqlite/src/lib.rs @@ -6,35 +6,42 @@ mod access_lock; mod backup_restore; mod file_permissions; +mod publication; mod schema_contract; use std::collections::{BTreeMap, BTreeSet}; use std::fs::{self, OpenOptions}; use std::io::Write; use std::path::{Path, PathBuf}; +use std::sync::LazyLock; use std::time::{Duration, SystemTime, UNIX_EPOCH}; use access_lock::StoreAccessLock; use backup_restore::{backup_connection, restore_database}; use code_system_graph_model::{ - ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityId, CommunityMetrics, CommunitySnapshot, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, LinkDecision, LinkStatus, NativePath, Node, NodeId, NodeKind, RepoFreshness, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, WorkspaceId, WorkspaceRecord, contains_unsafe_metadata_characters, stable_id + ArtifactChange, ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityId, CommunityMetrics, CommunitySnapshot, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, HttpLinkCoverage, HttpLinkGap, HttpLinkGapReason, HttpLinkReport, LinkDecision, LinkStatus, NativePath, Node, NodeId, NodeKind, RepoFreshness, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, WorkspaceId, WorkspaceRecord, contains_unsafe_metadata_characters, stable_id }; pub use file_permissions::set_owner_only_file; use file_permissions::{ SQLITE_ARTIFACT_SUFFIXES, artifact_path, prepare_database_file, restrict_store_permissions }; +use publication::{PublicationInput, publish_current_graph, replace_manual_links}; use rusqlite::{Connection, OpenFlags, OptionalExtension, params}; use schema_contract::validate_exact_schema; use sysinfo::{Pid, ProcessesToUpdate, System}; use thiserror::Error; const INITIAL_SCHEMA: &str = include_str!("../migrations/0001_initial.sql"); -const LATEST_SCHEMA_VERSION: i64 = 2; +static SCHEMA_ID: LazyLock = + LazyLock::new(|| stable_id("schema", &INITIAL_SCHEMA.replace("\r\n", "\n"))); -/// Returns the newest on-disk schema version supported by this binary. +/// Returns the identity of the exact on-disk schema supported by this binary. +/// +/// The identity is derived from the embedded schema definition, so any schema change produces a +/// new identity and databases created by other builds are rejected. #[must_use] -pub const fn latest_schema_version() -> i64 { - LATEST_SCHEMA_VERSION +pub fn schema_identity() -> &'static str { + SCHEMA_ID.as_str() } /// Source-free runtime capabilities observed from one open store connection. @@ -257,6 +264,21 @@ pub struct SnapshotBatch<'a> { pub community_snapshot: Option<&'a CommunitySnapshot>, } +/// Artifact rows of a candidate that differ from the current graph of its workspace. +/// +/// Publishing with a delta writes only these rows; every other stored fingerprint and extractor +/// batch must already equal the candidate, which holds when the delta was planned against the +/// stored fingerprints under the writer lock. +#[derive(Debug, Clone, Copy)] +pub struct ArtifactDelta<'a> { + /// Added or modified fingerprints. + pub upserted_fingerprints: &'a [&'a ArtifactFingerprint], + /// Extractor batches that are not stored in their current form. + pub upserted_batches: &'a [&'a StoredExtractorBatch], + /// Artifacts removed since the current graph; both their fingerprint and batch are deleted. + pub removed: &'a [ArtifactChange], +} + /// Result of restoring a database backup. #[derive(Debug, Clone, PartialEq, Eq)] pub struct RestoreReport { @@ -264,8 +286,8 @@ pub struct RestoreReport { pub source_path: PathBuf, /// Safety backup of the replaced database, when it existed. pub safety_backup_path: Option, - /// Exact schema version restored. - pub schema_version: i64, + /// Exact schema identity restored. + pub schema_id: String, } /// Compact persisted workspace registry entry. @@ -514,7 +536,7 @@ impl SqliteStore { Ok(()) } - /// Opens an exact 1.1.0 store or initializes a new empty database. + /// Opens an exact-schema store or initializes a new empty database. /// /// # Errors /// @@ -529,7 +551,7 @@ impl SqliteStore { }) } - /// Creates a validated online backup of the exact 1.1.0 schema. + /// Creates a validated online backup of the exact embedded schema. /// /// A source on read-only media is treated as immutable only when no `SQLite` sidecars exist. /// @@ -541,7 +563,7 @@ impl SqliteStore { backup_restore::backup_file(access_lock.database_path(), destination) } - /// Restores a validated backup with the exact 1.1.0 schema. + /// Restores a validated backup with the exact embedded schema. /// /// The existing destination is first preserved as a non-overwriting safety backup. /// Restore is refused while another [`SqliteStore`] has the destination open. @@ -581,7 +603,7 @@ impl SqliteStore { /// /// # Errors /// - /// Returns [`StoreError`] if the database is absent, corrupt, or not the exact 1.1.0 schema. + /// Returns [`StoreError`] if the database is absent, corrupt, or not the exact embedded schema. pub fn open_read_only(path: impl AsRef) -> Result { let path = path.as_ref(); let access_lock = StoreAccessLock::shared(path)?; @@ -617,13 +639,13 @@ impl SqliteStore { }) } - /// Returns the exact initial schema version recorded by this store. + /// Returns the schema identity recorded by this store. /// /// # Errors /// /// Returns [`StoreError`] when schema metadata cannot be read. - pub fn schema_version(&self) -> Result { - schema_version(&self.connection) + pub fn schema_id(&self) -> Result { + stored_schema_id(&self.connection) } /// Returns the opaque identity that binds disposable operational state to this database. @@ -784,8 +806,10 @@ impl SqliteStore { &self, workspace: &str, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.load_freshness_snapshot(&snapshot_id) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.load_freshness_snapshot(&snapshot_id) + }) } /// Loads per-repository freshness from one immutable graph snapshot. @@ -797,38 +821,172 @@ impl SqliteStore { &self, snapshot_id: &str, ) -> Result, StoreError> { - self.require_snapshot(snapshot_id)?; - let mut statement = self.connection.prepare( - "SELECT repo_id, checkout_id, head_commit, manifest_hash, state, reason - FROM repository_snapshot_freshness - WHERE snapshot_id = ?1 - ORDER BY repo_id, checkout_id", - )?; + self.consistent_read(|| { + let workspace = self.require_snapshot(snapshot_id)?; + let mut statement = self.connection.prepare( + "SELECT repo_id, checkout_id, head_commit, manifest_hash, state, reason + FROM repository_snapshot_freshness + WHERE workspace_name = ?1 + ORDER BY repo_id, checkout_id", + )?; + let rows = statement + .query_map([workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + row.get::<_, Option>(5)?, + )) + })? + .collect::, _>>()?; + rows.into_iter() + .map( + |(repo_id, checkout_id, head_commit, manifest_hash, state, reason)| { + Ok(RepoFreshness { + repo_id: RepoId::new(repo_id), + checkout_id: CheckoutId::new(checkout_id), + head_commit, + manifest_hash, + state: serde_json::from_str::(&state)?, + reason, + }) + }, + ) + .collect() + }) + } + + /// Loads the HTTP link coverage of the current graph with at most `gap_limit` gaps: calls + /// without a provider first, then ambiguous calls, then external calls. The second value is + /// the total number of stored gaps. + /// + /// # Errors + /// + /// Returns [`StoreError`] when no current snapshot exists or stored coverage is invalid. + pub fn load_http_link_report( + &self, + workspace: &str, + gap_limit: usize, + ) -> Result<(HttpLinkReport, usize), StoreError> { + self.consistent_read(|| { + self.current_snapshot_id(workspace)?; + let counts = self + .connection + .query_row( + "SELECT linked, no_provider, ambiguous, external + FROM http_link_coverage WHERE workspace_name = ?1", + [workspace], + |row| { + Ok([ + row.get::<_, i64>(0)?, + row.get::<_, i64>(1)?, + row.get::<_, i64>(2)?, + row.get::<_, i64>(3)?, + ]) + }, + ) + .optional()? + .unwrap_or_default(); + let count = |value: i64| { + u64::try_from(value).map_err(|_| StoreError::IntegerOutOfRange { + field: "http_link_coverage", + value: i128::from(value), + }) + }; + let coverage = HttpLinkCoverage { + linked: count(counts[0])?, + no_provider: count(counts[1])?, + ambiguous: count(counts[2])?, + external: count(counts[3])?, + }; + let total = self.connection.query_row( + "SELECT COUNT(*) FROM http_link_gaps WHERE workspace_name = ?1", + [workspace], + |row| row.get::<_, i64>(0), + )?; + let limit = i64::try_from(gap_limit).unwrap_or(i64::MAX); + let gaps = self.http_link_gaps( + "SELECT caller_node_id, method, path, reason, candidates_json + FROM http_link_gaps + WHERE workspace_name = ?1 + ORDER BY CASE reason + WHEN 'no_provider' THEN 0 WHEN 'ambiguous' THEN 1 ELSE 2 END, + caller_node_id, method, path + LIMIT ?2", + params![workspace, limit], + )?; + Ok(( + HttpLinkReport { coverage, gaps }, + usize::try_from(total).unwrap_or(usize::MAX), + )) + }) + } + + /// Loads the HTTP link gaps of the given callers in the current graph. + /// + /// # Errors + /// + /// Returns [`StoreError`] when no current snapshot exists or stored gaps are invalid. + pub fn load_http_link_gaps_for_callers( + &self, + workspace: &str, + callers: &[NodeId], + ) -> Result, StoreError> { + if callers.is_empty() { + return Ok(Vec::new()); + } + let callers = + serde_json::to_string(&callers.iter().map(NodeId::as_str).collect::>())?; + self.consistent_read(|| { + self.current_snapshot_id(workspace)?; + self.http_link_gaps( + "SELECT caller_node_id, method, path, reason, candidates_json + FROM http_link_gaps + WHERE workspace_name = ?1 + AND caller_node_id IN (SELECT value FROM json_each(?2)) + ORDER BY caller_node_id, method, path, reason", + params![workspace, callers], + ) + }) + } + + fn http_link_gaps( + &self, + sql: &str, + parameters: impl rusqlite::Params, + ) -> Result, StoreError> { + let mut statement = self.connection.prepare(sql)?; let rows = statement - .query_map([snapshot_id], |row| { + .query_map(parameters, |row| { Ok(( row.get::<_, String>(0)?, row.get::<_, String>(1)?, - row.get::<_, Option>(2)?, + row.get::<_, String>(2)?, row.get::<_, String>(3)?, row.get::<_, String>(4)?, - row.get::<_, Option>(5)?, )) })? .collect::, _>>()?; rows.into_iter() - .map( - |(repo_id, checkout_id, head_commit, manifest_hash, state, reason)| { - Ok(RepoFreshness { - repo_id: RepoId::new(repo_id), - checkout_id: CheckoutId::new(checkout_id), - head_commit, - manifest_hash, - state: serde_json::from_str::(&state)?, - reason, - }) - }, - ) + .map(|(caller, method, path, reason, candidates)| { + let reason = HttpLinkGapReason::parse(&reason).ok_or_else(|| { + StoreError::InvalidGraphSnapshot(format!( + "unknown HTTP link gap reason `{reason}`" + )) + })?; + Ok(HttpLinkGap { + caller: NodeId::new(caller), + method, + path, + reason, + candidates: serde_json::from_str::>(&candidates)? + .into_iter() + .map(NodeId::new) + .collect(), + }) + }) .collect() } @@ -841,61 +999,63 @@ impl SqliteStore { &self, workspace: &str, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - let mut statement = self.connection.prepare( - "SELECT - repo_id, checkout_id, path_encoding, relative_path, path_display, - extractor, content_hash, size_bytes - FROM artifact_fingerprints - WHERE snapshot_id = ?1 - ORDER BY repo_id, checkout_id, path_encoding, relative_path, extractor", - )?; - let rows = statement - .query_map([snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, Vec>(3)?, - row.get::<_, String>(4)?, - row.get::<_, String>(5)?, - row.get::<_, String>(6)?, - row.get::<_, i64>(7)?, - )) - })? - .collect::, _>>()?; - rows.into_iter() - .map( - |( - repo_id, - checkout_id, - encoding, - bytes, - display, - extractor, - content_hash, - stored_size_bytes, - )| { - Ok(ArtifactFingerprint { - repo_id: RepoId::new(repo_id), - checkout_id: CheckoutId::new(checkout_id), - path: NativePath { - encoding: serde_json::from_str(&encoding)?, - bytes, - display, - }, + self.consistent_read(|| { + self.current_snapshot_id(workspace)?; + let mut statement = self.connection.prepare( + "SELECT + repo_id, checkout_id, path_encoding, relative_path, path_display, + extractor, content_hash, size_bytes + FROM artifact_fingerprints + WHERE workspace_name = ?1 + ORDER BY repo_id, checkout_id, path_encoding, relative_path, extractor", + )?; + let rows = statement + .query_map([workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, String>(2)?, + row.get::<_, Vec>(3)?, + row.get::<_, String>(4)?, + row.get::<_, String>(5)?, + row.get::<_, String>(6)?, + row.get::<_, i64>(7)?, + )) + })? + .collect::, _>>()?; + rows.into_iter() + .map( + |( + repo_id, + checkout_id, + encoding, + bytes, + display, extractor, content_hash, - size_bytes: u64::try_from(stored_size_bytes).map_err(|_| { - StoreError::IntegerOutOfRange { - field: "artifact_fingerprints.size_bytes", - value: i128::from(stored_size_bytes), - } - })?, - }) - }, - ) - .collect() + stored_size_bytes, + )| { + Ok(ArtifactFingerprint { + repo_id: RepoId::new(repo_id), + checkout_id: CheckoutId::new(checkout_id), + path: NativePath { + encoding: serde_json::from_str(&encoding)?, + bytes, + display, + }, + extractor, + content_hash, + size_bytes: u64::try_from(stored_size_bytes).map_err(|_| { + StoreError::IntegerOutOfRange { + field: "artifact_fingerprints.size_bytes", + value: i128::from(stored_size_bytes), + } + })?, + }) + }, + ) + .collect() + }) } /// Loads reusable source-owned extractor outputs from the current snapshot. @@ -923,81 +1083,114 @@ impl SqliteStore { workspace: &str, maximum_payload_bytes: u64, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - let maximum_payload_bytes = i64::try_from(maximum_payload_bytes).unwrap_or(i64::MAX); - let mut statement = self.connection.prepare( - "SELECT - repo_id, checkout_id, path_encoding, relative_path, path_display, - extractor, content_hash, size_bytes, extractor_version, budget_fingerprint, - source_was_lossy, output_count, payload - FROM extractor_batches - WHERE snapshot_id = ?1 AND length(payload) <= ?2 - ORDER BY repo_id, checkout_id, path_encoding, relative_path, extractor", - )?; - let rows = statement - .query_map(params![snapshot_id, maximum_payload_bytes], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, Vec>(3)?, - row.get::<_, String>(4)?, - row.get::<_, String>(5)?, - row.get::<_, String>(6)?, - row.get::<_, i64>(7)?, - row.get::<_, String>(8)?, - row.get::<_, String>(9)?, - row.get::<_, bool>(10)?, - row.get::<_, i64>(11)?, - row.get::<_, Vec>(12)?, - )) - })? - .collect::, _>>()?; - rows.into_iter() - .map( - |( - repo_id, - checkout_id, - encoding, - bytes, - display, - extractor, - content_hash, - stored_size_bytes, - extractor_version, - budget_fingerprint, - source_was_lossy, - stored_output_count, - payload, - )| { - Ok(StoredExtractorBatch { - source: ArtifactFingerprint { - repo_id: RepoId::new(repo_id), - checkout_id: CheckoutId::new(checkout_id), - path: NativePath { - encoding: serde_json::from_str(&encoding)?, - bytes, - display, - }, - extractor, - content_hash, - size_bytes: stored_metric_to_u64( - "extractor_batches.size_bytes", - stored_size_bytes, - )?, - }, + self.load_current_extractor_batches_where(workspace, maximum_payload_bytes, None) + } + + /// Loads the current batches that can carry persisted extraction degradations: batches whose + /// source was decoded lossily and batches owned by `extractor`. + /// + /// Oversized rows are omitted exactly as in + /// [`Self::load_current_extractor_batches_with_limit`]. + /// + /// # Errors + /// + /// Returns [`StoreError`] when no current snapshot exists or stored values are invalid. + pub fn load_current_degradation_batches( + &self, + workspace: &str, + maximum_payload_bytes: u64, + extractor: &str, + ) -> Result, StoreError> { + self.load_current_extractor_batches_where(workspace, maximum_payload_bytes, Some(extractor)) + } + + fn load_current_extractor_batches_where( + &self, + workspace: &str, + maximum_payload_bytes: u64, + degradation_extractor: Option<&str>, + ) -> Result, StoreError> { + self.consistent_read(|| { + self.current_snapshot_id(workspace)?; + let maximum_payload_bytes = i64::try_from(maximum_payload_bytes).unwrap_or(i64::MAX); + let mut statement = self.connection.prepare( + "SELECT + repo_id, checkout_id, path_encoding, relative_path, path_display, + extractor, content_hash, size_bytes, extractor_version, budget_fingerprint, + source_was_lossy, output_count, payload + FROM extractor_batches + WHERE workspace_name = ?1 AND length(payload) <= ?2 + AND (?3 IS NULL OR source_was_lossy = 1 OR extractor = ?3) + ORDER BY repo_id, checkout_id, path_encoding, relative_path, extractor", + )?; + let rows = statement + .query_map( + params![workspace, maximum_payload_bytes, degradation_extractor], + |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, String>(2)?, + row.get::<_, Vec>(3)?, + row.get::<_, String>(4)?, + row.get::<_, String>(5)?, + row.get::<_, String>(6)?, + row.get::<_, i64>(7)?, + row.get::<_, String>(8)?, + row.get::<_, String>(9)?, + row.get::<_, bool>(10)?, + row.get::<_, i64>(11)?, + row.get::<_, Vec>(12)?, + )) + }, + )? + .collect::, _>>()?; + rows.into_iter() + .map( + |( + repo_id, + checkout_id, + encoding, + bytes, + display, + extractor, + content_hash, + stored_size_bytes, extractor_version, budget_fingerprint, source_was_lossy, - output_count: stored_metric_to_u64( - "extractor_batches.output_count", - stored_output_count, - )?, + stored_output_count, payload, - }) - }, - ) - .collect() + )| { + Ok(StoredExtractorBatch { + source: ArtifactFingerprint { + repo_id: RepoId::new(repo_id), + checkout_id: CheckoutId::new(checkout_id), + path: NativePath { + encoding: serde_json::from_str(&encoding)?, + bytes, + display, + }, + extractor, + content_hash, + size_bytes: stored_metric_to_u64( + "extractor_batches.size_bytes", + stored_size_bytes, + )?, + }, + extractor_version, + budget_fingerprint, + source_was_lossy, + output_count: stored_metric_to_u64( + "extractor_batches.output_count", + stored_output_count, + )?, + payload, + }) + }, + ) + .collect() + }) } /// Loads extractor run metrics from the current snapshot. @@ -1009,76 +1202,80 @@ impl SqliteStore { &self, workspace: &str, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - let mut statement = self.connection.prepare( - "SELECT - id, repo_id, checkout_id, extractor, extractor_version, status, - discovered_files, parsed_files, skipped_files, elapsed_ms - FROM extractor_runs - WHERE snapshot_id = ?1 - ORDER BY repo_id, checkout_id, extractor, id", - )?; - let rows = statement - .query_map([&snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, String>(3)?, - row.get::<_, String>(4)?, - row.get::<_, String>(5)?, - row.get::<_, i64>(6)?, - row.get::<_, i64>(7)?, - row.get::<_, i64>(8)?, - row.get::<_, i64>(9)?, - )) - })? - .collect::, _>>()?; - rows.into_iter() - .map( - |( - id, - repo_id, - checkout_id, - extractor, - extractor_version, - status, - discovered_files, - parsed_files, - skipped_files, - elapsed_ms, - )| { - Ok(ExtractorRun { + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + let mut statement = self.connection.prepare( + "SELECT + id, repo_id, checkout_id, extractor, extractor_version, status, + discovered_files, parsed_files, skipped_files, elapsed_ms + FROM extractor_runs + WHERE workspace_name = ?1 + ORDER BY repo_id, checkout_id, extractor, id", + )?; + let rows = statement + .query_map([workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, String>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + row.get::<_, String>(5)?, + row.get::<_, i64>(6)?, + row.get::<_, i64>(7)?, + row.get::<_, i64>(8)?, + row.get::<_, i64>(9)?, + )) + })? + .collect::, _>>()?; + rows.into_iter() + .map( + |( id, - snapshot_id: snapshot_id.clone(), - repo_id: RepoId::new(repo_id), - checkout_id: CheckoutId::new(checkout_id), + repo_id, + checkout_id, extractor, extractor_version, - status: serde_json::from_str(&status)?, - discovered_files: stored_metric_to_u64( - "extractor_runs.discovered_files", - discovered_files, - )?, - parsed_files: stored_metric_to_u64( - "extractor_runs.parsed_files", - parsed_files, - )?, - skipped_files: stored_metric_to_u64( - "extractor_runs.skipped_files", - skipped_files, - )?, - elapsed_ms: stored_metric_to_u64("extractor_runs.elapsed_ms", elapsed_ms)?, - }) - }, - ) - .collect() + status, + discovered_files, + parsed_files, + skipped_files, + elapsed_ms, + )| { + Ok(ExtractorRun { + id, + snapshot_id: snapshot_id.clone(), + repo_id: RepoId::new(repo_id), + checkout_id: CheckoutId::new(checkout_id), + extractor, + extractor_version, + status: serde_json::from_str(&status)?, + discovered_files: stored_metric_to_u64( + "extractor_runs.discovered_files", + discovered_files, + )?, + parsed_files: stored_metric_to_u64( + "extractor_runs.parsed_files", + parsed_files, + )?, + skipped_files: stored_metric_to_u64( + "extractor_runs.skipped_files", + skipped_files, + )?, + elapsed_ms: stored_metric_to_u64( + "extractor_runs.elapsed_ms", + elapsed_ms, + )?, + }) + }, + ) + .collect() + }) } - /// Replaces manual link declarations for one existing snapshot in a transaction. + /// Replaces manual link declarations of the current snapshot in a transaction. /// - /// Records for other snapshots are retained, preserving historical declarations. Passing an - /// empty slice clears only the selected snapshot's declarations. + /// Passing an empty slice clears the declarations. /// /// # Errors /// @@ -1090,7 +1287,7 @@ impl SqliteStore { records: &[ManualLinkRecord], ) -> Result<(), StoreError> { validate_manual_links(snapshot_id, records, None)?; - self.require_snapshot(snapshot_id)?; + let workspace = self.require_snapshot(snapshot_id)?; let transaction = self.connection.transaction()?; for record in records { for (field, node_id) in [ @@ -1099,9 +1296,9 @@ impl SqliteStore { ] { let exists = transaction.query_row( "SELECT EXISTS( - SELECT 1 FROM nodes WHERE snapshot_id = ?1 AND id = ?2 + SELECT 1 FROM nodes WHERE workspace_name = ?1 AND id = ?2 )", - params![snapshot_id, node_id], + params![workspace, node_id], |row| row.get::<_, bool>(0), )?; if !exists { @@ -1113,16 +1310,12 @@ impl SqliteStore { } } } - transaction.execute( - "DELETE FROM manual_links WHERE snapshot_id = ?1", - [snapshot_id], - )?; - insert_manual_links(&transaction, records)?; + replace_manual_links(&transaction, &workspace, records)?; transaction.commit()?; Ok(()) } - /// Loads manual link declarations for one immutable snapshot in stable identifier order. + /// Loads manual link declarations of the current snapshot in stable identifier order. /// /// # Errors /// @@ -1132,72 +1325,74 @@ impl SqliteStore { &self, snapshot_id: &str, ) -> Result, StoreError> { - self.require_snapshot(snapshot_id)?; - let mut statement = self.connection.prepare( - "SELECT id, source_node_id, target_node_id, kind, disposition, reason, decision_json, - config_version - FROM manual_links - WHERE snapshot_id = ?1 - ORDER BY id", - )?; - let rows = statement - .query_map([snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, String>(3)?, - row.get::<_, String>(4)?, - row.get::<_, String>(5)?, - row.get::<_, Vec>(6)?, - row.get::<_, i64>(7)?, - )) - })? - .collect::, _>>()?; - let records = rows - .into_iter() - .map( - |(id, source, target, kind, disposition, reason, decision_json, config_version)| { - let disposition = - ManualLinkDisposition::from_stored(&disposition).ok_or_else(|| { + self.consistent_read(|| { + let workspace = self.require_snapshot(snapshot_id)?; + let mut statement = self.connection.prepare( + "SELECT id, source_node_id, target_node_id, kind, disposition, reason, decision_json, + config_version + FROM manual_links + WHERE workspace_name = ?1 + ORDER BY id", + )?; + let rows = statement + .query_map([workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, String>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + row.get::<_, String>(5)?, + row.get::<_, Vec>(6)?, + row.get::<_, i64>(7)?, + )) + })? + .collect::, _>>()?; + let records = rows + .into_iter() + .map( + |(id, source, target, kind, disposition, reason, decision_json, config_version)| { + let disposition = + ManualLinkDisposition::from_stored(&disposition).ok_or_else(|| { + malformed_stored_data( + format!("manual link `{id}` in snapshot `{snapshot_id}`"), + format!("unknown disposition `{disposition}`"), + ) + })?; + let config_version = u32::try_from(config_version).map_err(|_| { + StoreError::IntegerOutOfRange { + field: "manual_links.config_version", + value: i128::from(config_version), + } + })?; + let decision = serde_json::from_slice(&decision_json).map_err(|error| { malformed_stored_data( - format!("manual link `{id}` in snapshot `{snapshot_id}`"), - format!("unknown disposition `{disposition}`"), + format!("manual link `{id}` decision in snapshot `{snapshot_id}`"), + error.to_string(), ) })?; - let config_version = u32::try_from(config_version).map_err(|_| { - StoreError::IntegerOutOfRange { - field: "manual_links.config_version", - value: i128::from(config_version), - } - })?; - let decision = serde_json::from_slice(&decision_json).map_err(|error| { - malformed_stored_data( - format!("manual link `{id}` decision in snapshot `{snapshot_id}`"), - error.to_string(), - ) - })?; - Ok(ManualLinkRecord { - id, - snapshot_id: snapshot_id.to_owned(), - source_node_id: NodeId::new(source), - target_node_id: NodeId::new(target), - kind, - disposition, - reason, - decision, - config_version, - }) - }, - ) - .collect::, StoreError>>()?; - validate_manual_links(snapshot_id, &records, None).map_err(|error| { - malformed_stored_data( - format!("manual links in snapshot `{snapshot_id}`"), - error.to_string(), - ) - })?; - Ok(records) + Ok(ManualLinkRecord { + id, + snapshot_id: snapshot_id.to_owned(), + source_node_id: NodeId::new(source), + target_node_id: NodeId::new(target), + kind, + disposition, + reason, + decision, + config_version, + }) + }, + ) + .collect::, StoreError>>()?; + validate_manual_links(snapshot_id, &records, None).map_err(|error| { + malformed_stored_data( + format!("manual links in snapshot `{snapshot_id}`"), + error.to_string(), + ) + })?; + Ok(records) + }) } /// Inserts or replaces one source-free provider capability report. @@ -1459,26 +1654,28 @@ impl SqliteStore { &self, workspace: &str, ) -> Result { - let snapshot_id = self.current_snapshot_id(workspace)?; - let (node_count, edge_count, evidence_count) = self.connection.query_row( - "SELECT - (SELECT COUNT(*) FROM nodes WHERE snapshot_id = ?1), - (SELECT COUNT(*) FROM edges WHERE snapshot_id = ?1), - (SELECT COUNT(*) FROM evidence WHERE snapshot_id = ?1)", - [&snapshot_id], - |row| { - Ok(( - row.get::<_, i64>(0)?, - row.get::<_, i64>(1)?, - row.get::<_, i64>(2)?, - )) - }, - )?; - Ok(StoredSnapshotSummary { - snapshot_id, - node_count: count_to_usize("nodes", node_count)?, - edge_count: count_to_usize("edges", edge_count)?, - evidence_count: count_to_usize("evidence", evidence_count)?, + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + let (node_count, edge_count, evidence_count) = self.connection.query_row( + "SELECT + (SELECT COUNT(*) FROM nodes WHERE workspace_name = ?1), + (SELECT COUNT(*) FROM edges WHERE workspace_name = ?1), + (SELECT COUNT(*) FROM evidence WHERE workspace_name = ?1)", + [workspace], + |row| { + Ok(( + row.get::<_, i64>(0)?, + row.get::<_, i64>(1)?, + row.get::<_, i64>(2)?, + )) + }, + )?; + Ok(StoredSnapshotSummary { + snapshot_id, + node_count: count_to_usize("nodes", node_count)?, + edge_count: count_to_usize("edges", edge_count)?, + evidence_count: count_to_usize("evidence", evidence_count)?, + }) }) } @@ -1508,10 +1705,20 @@ impl SqliteStore { where F: FnMut(u64), { - self.publish_snapshot_with_progress_and_coverage(batch, &[], progress) + self.publish_snapshot_with_progress_and_coverage( + batch, + None, + &[], + &HttpLinkReport::default(), + progress, + ) } - /// Publishes one atomic snapshot with repository-scoped coverage gaps. + /// Publishes one atomic snapshot with repository-scoped coverage gaps and the HTTP link + /// coverage of its graph. + /// + /// With `artifact_delta`, only the listed artifact rows are written; without it, every + /// stored fingerprint and extractor batch is compared with the candidate. /// /// # Errors /// @@ -1519,7 +1726,9 @@ impl SqliteStore { pub fn publish_snapshot_with_progress_and_coverage( &mut self, batch: SnapshotBatch<'_>, + artifact_delta: Option>, coverage_gaps: &[RepositoryCoverageGap], + http_links: &HttpLinkReport, mut progress: F, ) -> Result<(), StoreError> where @@ -1543,52 +1752,25 @@ impl SqliteStore { validate_community_snapshot(snapshot_id, nodes, community_snapshot)?; let transaction = self.connection.transaction()?; upsert_registry(&transaction, workspace)?; - transaction.execute( - "DELETE FROM query_cache WHERE workspace_name = ?1", - [&workspace.name], - )?; - transaction.execute( - "UPDATE repo_snapshots SET is_current = 0 WHERE workspace_name = ?1", - [&workspace.name], - )?; - transaction.execute("DELETE FROM repo_snapshots WHERE id = ?1", [snapshot_id])?; - transaction.execute( - "INSERT INTO repo_snapshots(id, workspace_name, is_current) VALUES (?1, ?2, 0)", - params![snapshot_id, workspace.name], - )?; - - insert_graph( - &transaction, - snapshot_id, - nodes, - edges, - evidence, - &mut progress, - )?; - insert_manual_links(&transaction, manual_links)?; - progress(u64::try_from(manual_links.len()).unwrap_or(u64::MAX)); - insert_incremental_state( + publish_current_graph( &transaction, - snapshot_id, - fingerprints, - extractor_runs, + &PublicationInput { + workspace, + snapshot_id, + nodes, + edges, + evidence, + fingerprints, + extractor_batches, + artifact_delta, + extractor_runs, + manual_links, + community_snapshot, + coverage_gaps, + http_links, + }, &mut progress, )?; - insert_extractor_batches(&transaction, snapshot_id, extractor_batches, &mut progress)?; - insert_freshness( - &transaction, - snapshot_id, - workspace, - extractor_batches, - coverage_gaps, - )?; - if let Some(community_snapshot) = community_snapshot { - insert_community_snapshot(&transaction, community_snapshot, &mut progress)?; - } - transaction.execute( - "UPDATE repo_snapshots SET is_current = 1 WHERE id = ?1", - [snapshot_id], - )?; transaction.commit()?; Ok(()) } @@ -1602,8 +1784,10 @@ impl SqliteStore { &self, workspace: &str, ) -> Result<(Vec, Vec), StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.load_graph_snapshot(&snapshot_id) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.load_graph_snapshot(&snapshot_id) + }) } /// Loads nodes and edges from one immutable graph snapshot. @@ -1618,79 +1802,81 @@ impl SqliteStore { &self, snapshot_id: &str, ) -> Result<(Vec, Vec), StoreError> { - self.require_snapshot(snapshot_id)?; - let mut node_statement = self.connection.prepare( - "SELECT id, kind, repo_id, stable_key, label - FROM nodes WHERE snapshot_id = ?1 ORDER BY id", - )?; - let node_rows = node_statement - .query_map([&snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, Option>(2)?, - row.get::<_, String>(3)?, - row.get::<_, String>(4)?, - )) - })? - .collect::, _>>()?; - let nodes = node_rows - .into_iter() - .map(|(id, kind, repo_id, stable_key, label)| { - Ok(Node { - id: NodeId::new(id), - kind: serde_json::from_str::(&kind)?, - repo_id: repo_id.map(RepoId::new), - stable_key, - label, + self.consistent_read(|| { + let workspace = self.require_snapshot(snapshot_id)?; + let mut node_statement = self.connection.prepare( + "SELECT id, kind, repo_id, stable_key, label + FROM nodes WHERE workspace_name = ?1 ORDER BY id", + )?; + let node_rows = node_statement + .query_map([&workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + )) + })? + .collect::, _>>()?; + let nodes = node_rows + .into_iter() + .map(|(id, kind, repo_id, stable_key, label)| { + Ok(Node { + id: NodeId::new(id), + kind: serde_json::from_str::(&kind)?, + repo_id: repo_id.map(RepoId::new), + stable_key, + label, + }) }) - }) - .collect::, StoreError>>()?; + .collect::, StoreError>>()?; - let mut edge_statement = self.connection.prepare( - "SELECT id, source_node_id, target_node_id, kind, confidence, epistemic_status - FROM edges WHERE snapshot_id = ?1 ORDER BY id", - )?; - let edge_rows = edge_statement - .query_map([&snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, String>(3)?, - row.get::<_, f32>(4)?, - row.get::<_, String>(5)?, - )) - })? - .collect::, _>>()?; - let mut evidence_statement = self.connection.prepare( - "SELECT edge_id, evidence_id FROM edge_evidence - WHERE snapshot_id = ?1 ORDER BY edge_id, evidence_id", - )?; - let mut evidence_by_edge = BTreeMap::>::new(); - for result in evidence_statement.query_map([&snapshot_id], |row| { - Ok((row.get::<_, String>(0)?, row.get::<_, String>(1)?)) - })? { - let (edge_id, evidence_id) = result?; - evidence_by_edge - .entry(edge_id) - .or_default() - .push(EvidenceId::new(evidence_id)); - } - let mut edges = Vec::with_capacity(edge_rows.len()); - for (id, source, target, kind, confidence, status) in edge_rows { - let evidence = evidence_by_edge.remove(&id).unwrap_or_default(); - edges.push(Edge { - id: EdgeId::new(id), - source: NodeId::new(source), - target: NodeId::new(target), - kind: serde_json::from_str::(&kind)?, - confidence, - status: serde_json::from_str::(&status)?, - evidence, - }); - } - Ok((nodes, edges)) + let mut edge_statement = self.connection.prepare( + "SELECT id, source_node_id, target_node_id, kind, confidence, epistemic_status + FROM edges WHERE workspace_name = ?1 ORDER BY id", + )?; + let edge_rows = edge_statement + .query_map([&workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, String>(2)?, + row.get::<_, String>(3)?, + row.get::<_, f32>(4)?, + row.get::<_, String>(5)?, + )) + })? + .collect::, _>>()?; + let mut evidence_statement = self.connection.prepare( + "SELECT edge_id, evidence_id FROM edge_evidence + WHERE workspace_name = ?1 ORDER BY edge_id, evidence_id", + )?; + let mut evidence_by_edge = BTreeMap::>::new(); + for result in evidence_statement.query_map([&workspace], |row| { + Ok((row.get::<_, String>(0)?, row.get::<_, String>(1)?)) + })? { + let (edge_id, evidence_id) = result?; + evidence_by_edge + .entry(edge_id) + .or_default() + .push(EvidenceId::new(evidence_id)); + } + let mut edges = Vec::with_capacity(edge_rows.len()); + for (id, source, target, kind, confidence, status) in edge_rows { + let evidence = evidence_by_edge.remove(&id).unwrap_or_default(); + edges.push(Edge { + id: EdgeId::new(id), + source: NodeId::new(source), + target: NodeId::new(target), + kind: serde_json::from_str::(&kind)?, + confidence, + status: serde_json::from_str::(&status)?, + evidence, + }); + } + Ok((nodes, edges)) + }) } /// Loads a bounded, ordered page of current nodes for exact kinds and returns the exact total. @@ -1705,8 +1891,10 @@ impl SqliteStore { kinds: &[NodeKind], limit: usize, ) -> Result<(usize, Vec), StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.load_nodes_by_kinds_snapshot(&snapshot_id, kinds, limit) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.load_nodes_by_kinds_snapshot(&snapshot_id, kinds, limit) + }) } /// Loads a bounded, ordered page of nodes for exact kinds from one immutable snapshot. @@ -1721,66 +1909,69 @@ impl SqliteStore { kinds: &[NodeKind], limit: usize, ) -> Result<(usize, Vec), StoreError> { - if kinds.is_empty() || limit == 0 { - return Ok((0, Vec::new())); - } - self.require_snapshot(snapshot_id)?; - let kind_values = kinds - .iter() - .map(serde_json::to_string) - .collect::, _>>()?; - let placeholders = (2..=kind_values.len() + 1) - .map(|index| format!("?{index}")) - .collect::>() - .join(", "); - let mut parameters = Vec::::with_capacity(kind_values.len() + 2); - parameters.push(snapshot_id.to_owned().into()); - parameters.extend(kind_values.into_iter().map(rusqlite::types::Value::Text)); - let count_sql = format!( - "SELECT COUNT(*) FROM nodes WHERE snapshot_id = ?1 AND kind IN ({placeholders})" - ); - let count = self.connection.query_row( - &count_sql, - rusqlite::params_from_iter(parameters.iter()), - |row| row.get::<_, i64>(0), - )?; - let total = count_to_usize("nodes", count)?; - let limit = i64::try_from(limit).map_err(|_| StoreError::IntegerOutOfRange { - field: "nodes.limit", - value: i128::try_from(limit).unwrap_or(i128::MAX), - })?; - let limit_parameter = parameters.len() + 1; - parameters.push(rusqlite::types::Value::Integer(limit)); - let select_sql = format!( - "SELECT id, kind, repo_id, stable_key, label FROM nodes - WHERE snapshot_id = ?1 AND kind IN ({placeholders}) - ORDER BY id LIMIT ?{limit_parameter}" - ); - let mut statement = self.connection.prepare(&select_sql)?; - let rows = statement - .query_map(rusqlite::params_from_iter(parameters.iter()), |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, Option>(2)?, - row.get::<_, String>(3)?, - row.get::<_, String>(4)?, - )) - })? - .collect::, _>>()?; - let nodes = rows - .into_iter() - .map(|(id, kind, repo_id, stable_key, label)| { - Ok(Node { - id: NodeId::new(id), - kind: serde_json::from_str::(&kind)?, - repo_id: repo_id.map(RepoId::new), - stable_key, - label, + self.consistent_read(|| { + if kinds.is_empty() || limit == 0 { + return Ok((0, Vec::new())); + } + let workspace = self.require_snapshot(snapshot_id)?; + let kind_values = kinds + .iter() + .map(serde_json::to_string) + .collect::, _>>()?; + let placeholders = (2..=kind_values.len() + 1) + .map(|index| format!("?{index}")) + .collect::>() + .join(", "); + let mut parameters = + Vec::::with_capacity(kind_values.len() + 2); + parameters.push(workspace.into()); + parameters.extend(kind_values.into_iter().map(rusqlite::types::Value::Text)); + let count_sql = format!( + "SELECT COUNT(*) FROM nodes WHERE workspace_name = ?1 AND kind IN ({placeholders})" + ); + let count = self.connection.query_row( + &count_sql, + rusqlite::params_from_iter(parameters.iter()), + |row| row.get::<_, i64>(0), + )?; + let total = count_to_usize("nodes", count)?; + let limit = i64::try_from(limit).map_err(|_| StoreError::IntegerOutOfRange { + field: "nodes.limit", + value: i128::try_from(limit).unwrap_or(i128::MAX), + })?; + let limit_parameter = parameters.len() + 1; + parameters.push(rusqlite::types::Value::Integer(limit)); + let select_sql = format!( + "SELECT id, kind, repo_id, stable_key, label FROM nodes + WHERE workspace_name = ?1 AND kind IN ({placeholders}) + ORDER BY id LIMIT ?{limit_parameter}" + ); + let mut statement = self.connection.prepare(&select_sql)?; + let rows = statement + .query_map(rusqlite::params_from_iter(parameters.iter()), |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + )) + })? + .collect::, _>>()?; + let nodes = rows + .into_iter() + .map(|(id, kind, repo_id, stable_key, label)| { + Ok(Node { + id: NodeId::new(id), + kind: serde_json::from_str::(&kind)?, + repo_id: repo_id.map(RepoId::new), + stable_key, + label, + }) }) - }) - .collect::, StoreError>>()?; - Ok((total, nodes)) + .collect::, StoreError>>()?; + Ok((total, nodes)) + }) } /// Loads evidence metadata for the current workspace snapshot. @@ -1789,8 +1980,10 @@ impl SqliteStore { /// /// Returns [`StoreError`] when the current snapshot is absent or stored tags are malformed. pub fn load_current_evidence(&self, workspace: &str) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.load_evidence_snapshot(&snapshot_id) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.load_evidence_snapshot(&snapshot_id) + }) } /// Loads one exact evidence record from the current workspace snapshot without scanning the @@ -1804,102 +1997,33 @@ impl SqliteStore { workspace: &str, evidence_id: &str, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - let row = self - .connection - .query_row( - "SELECT id, repo_id, file_path, start_line, end_line, extractor, extractor_version, - provenance, confidence, observed_at_commit, content_hash - FROM evidence WHERE snapshot_id = ?1 AND id = ?2", - params![snapshot_id, evidence_id], - |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, Option>(1)?, - row.get::<_, Option>(2)?, - row.get::<_, Option>(3)?, - row.get::<_, Option>(4)?, - row.get::<_, String>(5)?, - row.get::<_, String>(6)?, - row.get::<_, String>(7)?, - row.get::<_, f32>(8)?, - row.get::<_, Option>(9)?, - row.get::<_, Option>(10)?, - )) - }, - ) - .optional()?; - row.map( - |( - id, - repo_id, - file_path, - start_line, - end_line, - extractor, - extractor_version, - provenance, - confidence, - observed_at_commit, - content_hash, - )| { - let provenance = serde_json::from_str(&provenance).map_err(|error| { - malformed_stored_data(format!("evidence `{id}`"), error.to_string()) - })?; - Ok(Evidence { - id: EvidenceId::new(id), - repo_id: repo_id.map(RepoId::new), - file_path, - start_line, - end_line, - extractor, - extractor_version, - provenance, - confidence, - observed_at_commit, - content_hash, - note: None, - }) - }, - ) - .transpose() - } - - /// Loads evidence metadata for one immutable graph snapshot. - /// - /// Line ranges, extractor versions, commits, and notes predate their normalized storage - /// columns and therefore remain absent in this projection. - /// - /// # Errors - /// - /// Returns [`StoreError::SnapshotMissing`] when the snapshot is absent or - /// [`StoreError::MalformedStoredData`] when a stored provenance tag is invalid. - pub fn load_evidence_snapshot(&self, snapshot_id: &str) -> Result, StoreError> { - self.require_snapshot(snapshot_id)?; - let mut statement = self.connection.prepare( - "SELECT id, repo_id, file_path, start_line, end_line, extractor, extractor_version, - provenance, confidence, observed_at_commit, content_hash - FROM evidence WHERE snapshot_id = ?1 ORDER BY id", - )?; - let rows = statement - .query_map([snapshot_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, Option>(1)?, - row.get::<_, Option>(2)?, - row.get::<_, Option>(3)?, - row.get::<_, Option>(4)?, - row.get::<_, String>(5)?, - row.get::<_, String>(6)?, - row.get::<_, String>(7)?, - row.get::<_, f32>(8)?, - row.get::<_, Option>(9)?, - row.get::<_, Option>(10)?, - )) - })? - .collect::, _>>()?; - rows.into_iter() - .map( + self.consistent_read(|| { + self.current_snapshot_id(workspace)?; + let row = self + .connection + .query_row( + "SELECT id, repo_id, file_path, start_line, end_line, extractor, extractor_version, + provenance, confidence, observed_at_commit, content_hash + FROM evidence WHERE workspace_name = ?1 AND id = ?2", + params![workspace, evidence_id], + |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, Option>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, Option>(3)?, + row.get::<_, Option>(4)?, + row.get::<_, String>(5)?, + row.get::<_, String>(6)?, + row.get::<_, String>(7)?, + row.get::<_, f32>(8)?, + row.get::<_, Option>(9)?, + row.get::<_, Option>(10)?, + )) + }, + ) + .optional()?; + row.map( |( id, repo_id, @@ -1932,7 +2056,79 @@ impl SqliteStore { }) }, ) - .collect() + .transpose() + }) + } + + /// Loads evidence metadata for one current graph snapshot. + /// + /// Evidence notes are not persisted and remain absent in this projection. + /// + /// # Errors + /// + /// Returns [`StoreError::SnapshotMissing`] when the snapshot is absent or + /// [`StoreError::MalformedStoredData`] when a stored provenance tag is invalid. + pub fn load_evidence_snapshot(&self, snapshot_id: &str) -> Result, StoreError> { + self.consistent_read(|| { + let workspace = self.require_snapshot(snapshot_id)?; + let mut statement = self.connection.prepare( + "SELECT id, repo_id, file_path, start_line, end_line, extractor, extractor_version, + provenance, confidence, observed_at_commit, content_hash + FROM evidence WHERE workspace_name = ?1 ORDER BY id", + )?; + let rows = statement + .query_map([workspace], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, Option>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, Option>(3)?, + row.get::<_, Option>(4)?, + row.get::<_, String>(5)?, + row.get::<_, String>(6)?, + row.get::<_, String>(7)?, + row.get::<_, f32>(8)?, + row.get::<_, Option>(9)?, + row.get::<_, Option>(10)?, + )) + })? + .collect::, _>>()?; + rows.into_iter() + .map( + |( + id, + repo_id, + file_path, + start_line, + end_line, + extractor, + extractor_version, + provenance, + confidence, + observed_at_commit, + content_hash, + )| { + let provenance = serde_json::from_str(&provenance).map_err(|error| { + malformed_stored_data(format!("evidence `{id}`"), error.to_string()) + })?; + Ok(Evidence { + id: EvidenceId::new(id), + repo_id: repo_id.map(RepoId::new), + file_path, + start_line, + end_line, + extractor, + extractor_version, + provenance, + confidence, + observed_at_commit, + content_hash, + note: None, + }) + }, + ) + .collect() + }) } /// Searches current snapshot node labels and stable keys using bounded FTS5. @@ -1975,8 +2171,10 @@ impl SqliteStore { query: &str, limit: usize, ) -> Result, StoreError> { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.search_snapshot_nodes_ranked(&snapshot_id, query, limit) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.search_snapshot_nodes_ranked(&snapshot_id, query, limit) + }) } /// Searches nodes in one immutable snapshot and includes the native FTS5 relevance score. @@ -1991,73 +2189,119 @@ impl SqliteStore { query: &str, limit: usize, ) -> Result, StoreError> { - self.require_snapshot(snapshot_id)?; - let query = query.trim(); - if query.is_empty() { - return Err(StoreError::InvalidSearchQuery( - "query must not be empty".to_owned(), - )); - } - if query.len() > NODE_SEARCH_QUERY_MAX_BYTES { - return Err(StoreError::InvalidSearchQuery(format!( - "query must not exceed {NODE_SEARCH_QUERY_MAX_BYTES} UTF-8 bytes" - ))); - } - if !(1..=500).contains(&limit) { - return Err(StoreError::InvalidSearchQuery( - "limit must be between 1 and 500".to_owned(), - )); - } - let phrase = bounded_fts_disjunction(query); - let limit = i64::try_from(limit).map_err(|_| StoreError::IntegerOutOfRange { - field: "search.limit", - value: i128::try_from(limit).unwrap_or(i128::MAX), - })?; - let mut statement = self.connection.prepare( - "SELECT n.id, n.kind, n.repo_id, n.stable_key, n.label, bm25(nodes_fts) - FROM nodes_fts - JOIN nodes n - ON n.snapshot_id = nodes_fts.snapshot_id AND n.id = nodes_fts.node_id - WHERE nodes_fts.snapshot_id = ?1 - AND nodes_fts MATCH ?2 - ORDER BY bm25(nodes_fts), n.id - LIMIT ?3", - )?; - let rows = statement - .query_map(params![snapshot_id, phrase, limit], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, Option>(2)?, - row.get::<_, String>(3)?, - row.get::<_, String>(4)?, - row.get::<_, f64>(5)?, - )) - })? - .collect::, _>>()?; - rows.into_iter() - .map(|(id, kind, repo_id, stable_key, label, fts_rank)| { - if !fts_rank.is_finite() { - return Err(malformed_stored_data( - format!("FTS hit `{id}`"), - "rank is not finite", - )); - } - let kind = serde_json::from_str(&kind).map_err(|error| { - malformed_stored_data(format!("node `{id}`"), error.to_string()) - })?; - Ok(StoredNodeSearchHit { - node: Node { - id: NodeId::new(id), - kind, - repo_id: repo_id.map(RepoId::new), - stable_key, - label, - }, - fts_rank, + self.consistent_read(|| { + let workspace = self.require_snapshot(snapshot_id)?; + let query = query.trim(); + if query.is_empty() { + return Err(StoreError::InvalidSearchQuery( + "query must not be empty".to_owned(), + )); + } + if query.len() > NODE_SEARCH_QUERY_MAX_BYTES { + return Err(StoreError::InvalidSearchQuery(format!( + "query must not exceed {NODE_SEARCH_QUERY_MAX_BYTES} UTF-8 bytes" + ))); + } + if !(1..=500).contains(&limit) { + return Err(StoreError::InvalidSearchQuery( + "limit must be between 1 and 500".to_owned(), + )); + } + let phrase = bounded_fts_disjunction(query); + let limit = i64::try_from(limit).map_err(|_| StoreError::IntegerOutOfRange { + field: "search.limit", + value: i128::try_from(limit).unwrap_or(i128::MAX), + })?; + let mut statement = self.connection.prepare( + "SELECT n.id, n.kind, n.repo_id, n.stable_key, n.label, bm25(nodes_fts) + FROM nodes_fts + JOIN nodes n ON n.node_rowid = nodes_fts.rowid + WHERE n.workspace_name = ?1 + AND nodes_fts MATCH ?2 + ORDER BY bm25(nodes_fts), n.id + LIMIT ?3", + )?; + let rows = statement + .query_map(params![workspace, phrase, limit], |row| { + Ok(( + row.get::<_, String>(0)?, + row.get::<_, String>(1)?, + row.get::<_, Option>(2)?, + row.get::<_, String>(3)?, + row.get::<_, String>(4)?, + row.get::<_, f64>(5)?, + )) + })? + .collect::, _>>()?; + rows.into_iter() + .map(|(id, kind, repo_id, stable_key, label, fts_rank)| { + if !fts_rank.is_finite() { + return Err(malformed_stored_data( + format!("FTS hit `{id}`"), + "rank is not finite", + )); + } + let kind = serde_json::from_str(&kind).map_err(|error| { + malformed_stored_data(format!("node `{id}`"), error.to_string()) + })?; + Ok(StoredNodeSearchHit { + node: Node { + id: NodeId::new(id), + kind, + repo_id: repo_id.map(RepoId::new), + stable_key, + label, + }, + fts_rank, + }) + }) + .collect() + }) + } + + /// Returns the current snapshot identifier, or `None` before the first publication. + /// + /// # Errors + /// + /// Returns [`StoreError`] when the query fails. + pub fn find_current_snapshot_id(&self, workspace: &str) -> Result, StoreError> { + match self.current_snapshot_id(workspace) { + Ok(snapshot_id) => Ok(Some(snapshot_id)), + Err(StoreError::CurrentSnapshotMissing(_)) => Ok(None), + Err(error) => Err(error), + } + } + + /// Counts the communities of the current snapshot without loading their memberships. + /// + /// Returns `None` when no current snapshot or community analysis exists. + /// + /// # Errors + /// + /// Returns [`StoreError`] when the query fails or the count is out of range. + pub fn current_community_count(&self, workspace: &str) -> Result, StoreError> { + self.consistent_read(|| { + let Some(snapshot_id) = self.find_current_snapshot_id(workspace)? else { + return Ok(None); + }; + let count = self + .connection + .query_row( + "SELECT (SELECT COUNT(*) FROM communities WHERE snapshot_id = ?1) + FROM community_snapshots WHERE snapshot_id = ?1", + [&snapshot_id], + |row| row.get::<_, i64>(0), + ) + .optional()?; + count + .map(|count| { + usize::try_from(count).map_err(|_| StoreError::IntegerOutOfRange { + field: "communities.count", + value: i128::from(count), + }) }) - }) - .collect() + .transpose() + }) } /// Loads the community analysis associated with the current workspace snapshot. @@ -2070,30 +2314,46 @@ impl SqliteStore { &self, workspace: &str, ) -> Result { - let snapshot_id = self.current_snapshot_id(workspace)?; - self.load_community_snapshot(&snapshot_id) + self.consistent_read(|| { + let snapshot_id = self.current_snapshot_id(workspace)?; + self.load_community_snapshot(&snapshot_id) + }) } - /// Loads one immutable community analysis in deterministic identifier order. + /// Loads one retained community analysis in deterministic identifier order. + /// + /// The analyses of the current and the previous snapshot of each workspace are retained. /// /// # Errors /// - /// Returns [`StoreError::SnapshotMissing`] when the graph snapshot does not exist, - /// [`StoreError::CommunitySnapshotMissing`] when it has no analysis, or - /// [`StoreError::MalformedStoredData`] when persisted rows violate domain invariants. + /// Returns [`StoreError::SnapshotMissing`] when neither a graph nor an analysis exists for the + /// snapshot, [`StoreError::CommunitySnapshotMissing`] when the current graph has no analysis, + /// or [`StoreError::MalformedStoredData`] when persisted rows violate domain invariants. pub fn load_community_snapshot( &self, snapshot_id: &str, ) -> Result { - self.require_snapshot(snapshot_id)?; - load_community_snapshot(&self.connection, snapshot_id) + self.consistent_read(|| { + let retained = self + .connection + .query_row( + "SELECT 1 FROM community_snapshots WHERE snapshot_id = ?1", + [snapshot_id], + |_| Ok(()), + ) + .optional()?; + if retained.is_none() { + self.require_snapshot(snapshot_id)?; + } + load_community_snapshot(&self.connection, snapshot_id) + }) } - /// Loads an immutable community analysis only when it belongs to the requested workspace. + /// Loads a retained community analysis only when it belongs to the requested workspace. /// /// # Errors /// - /// Returns [`StoreError::SnapshotMissing`] when the snapshot is absent or belongs to another + /// Returns [`StoreError::SnapshotMissing`] when the analysis is absent or belongs to another /// workspace, preserving workspace isolation at read-only delivery boundaries. pub fn load_workspace_community_snapshot( &self, @@ -2103,7 +2363,7 @@ impl SqliteStore { let belongs = self .connection .query_row( - "SELECT 1 FROM repo_snapshots WHERE id = ?1 AND workspace_name = ?2", + "SELECT 1 FROM community_snapshots WHERE snapshot_id = ?1 AND workspace_name = ?2", params![snapshot_id, workspace], |_| Ok(()), ) @@ -2112,11 +2372,36 @@ impl SqliteStore { load_community_snapshot(&self.connection, snapshot_id) } + /// Runs related reads in one deferred transaction so they observe the same publication. + /// + /// Include snapshot selection and all dependent reads in the callback: a concurrent writer + /// can replace the current snapshot after this transaction ends. Nested calls reuse the + /// existing transaction. The callback must only perform reads on this connection. + /// + /// # Errors + /// + /// Returns errors from the callback or transaction setup/commit. A failed callback rolls back + /// a transaction started here; an existing transaction remains owned by its caller. + pub fn consistent_read(&self, read: impl FnOnce() -> Result) -> Result + where + E: From, + { + if !self.connection.is_autocommit() { + return read(); + } + let transaction = self + .connection + .unchecked_transaction() + .map_err(StoreError::from)?; + let value = read()?; + transaction.commit().map_err(StoreError::from)?; + Ok(value) + } + fn current_snapshot_id(&self, workspace: &str) -> Result { self.connection .query_row( - "SELECT id FROM repo_snapshots - WHERE workspace_name = ?1 AND is_current = 1", + "SELECT id FROM repo_snapshots WHERE workspace_name = ?1", [workspace], |row| row.get::<_, String>(0), ) @@ -2124,16 +2409,16 @@ impl SqliteStore { .ok_or_else(|| StoreError::CurrentSnapshotMissing(workspace.to_owned())) } - fn require_snapshot(&self, snapshot_id: &str) -> Result<(), StoreError> { - let exists = self - .connection + /// Resolves a current snapshot identifier to the workspace that owns it. + fn require_snapshot(&self, snapshot_id: &str) -> Result { + self.connection .query_row( - "SELECT 1 FROM repo_snapshots WHERE id = ?1", + "SELECT workspace_name FROM repo_snapshots WHERE id = ?1", [snapshot_id], - |_| Ok(()), + |row| row.get::<_, String>(0), ) - .optional()?; - exists.ok_or_else(|| StoreError::SnapshotMissing(snapshot_id.to_owned())) + .optional()? + .ok_or_else(|| StoreError::SnapshotMissing(snapshot_id.to_owned())) } /// Runs `SQLite`'s quick integrity check. @@ -2145,7 +2430,7 @@ impl SqliteStore { let result = self .connection .query_row("PRAGMA quick_check", [], |row| row.get::<_, String>(0))?; - if result != "ok" || self.schema_version()? != LATEST_SCHEMA_VERSION { + if result != "ok" || self.schema_id()? != schema_identity() { return Ok(false); } let foreign_key_violation = self @@ -2598,6 +2883,7 @@ fn validate_community_snapshot( fn insert_community_snapshot( transaction: &rusqlite::Transaction<'_>, + workspace: &str, snapshot: &CommunitySnapshot, progress: &mut F, ) -> Result<(), StoreError> @@ -2617,10 +2903,12 @@ where COMMUNITY_CONFIG_JSON_MAX_BYTES, )?; transaction.execute( - "INSERT INTO community_snapshots(snapshot_id, engine_version, algorithm, config_json) - VALUES (?1, ?2, ?3, ?4)", + "INSERT INTO community_snapshots( + snapshot_id, workspace_name, engine_version, algorithm, config_json + ) VALUES (?1, ?2, ?3, ?4, ?5)", params![ snapshot.snapshot_id, + workspace, snapshot.engine_version, algorithm, config_json @@ -2893,20 +3181,6 @@ fn load_community_members( "membership order is not contiguous", )); } - let node_exists = connection - .query_row( - "SELECT 1 FROM nodes WHERE snapshot_id = ?1 AND id = ?2", - params![snapshot_id, node_id], - |_| Ok(()), - ) - .optional()? - .is_some(); - if !node_exists { - return Err(malformed_stored_data( - entity, - format!("membership references missing node `{node_id}`"), - )); - } members.push(NodeId::new(node_id)); } Ok(members) @@ -3064,438 +3338,110 @@ fn validate_manual_link_decision(record: &ManualLinkRecord) -> Result<(), StoreE || record.decision.relation != relation || record.decision.status != expected_status || record.decision.reasons.first() != Some(&record.reason) - { - return Err(StoreError::InvalidPersistenceRecord(format!( - "manual link `{}` decision does not match its persisted declaration", - record.id - ))); - } - Ok(()) -} - -fn validate_provider_capability_record( - record: &ProviderCapabilityRecord, -) -> Result<(), StoreError> { - validate_safe_metadata( - "workspace name", - &record.workspace_name, - 1, - MANUAL_LINK_ID_MAX_BYTES, - )?; - validate_safe_metadata( - "repository identifier", - record.repo_id.as_str(), - 1, - MANUAL_LINK_ID_MAX_BYTES, - )?; - validate_safe_metadata( - "provider name", - &record.provider, - 1, - PROVIDER_COMPONENT_MAX_BYTES, - )?; - validate_safe_metadata( - "provider version", - &record.provider_version, - 1, - PROVIDER_COMPONENT_MAX_BYTES, - )?; - if record.capabilities.len() > PROVIDER_CAPABILITY_COUNT_MAX { - return Err(StoreError::InvalidPersistenceRecord(format!( - "provider capability count exceeds {PROVIDER_CAPABILITY_COUNT_MAX}" - ))); - } - let mut capability_names = BTreeSet::new(); - for capability in &record.capabilities { - validate_safe_metadata( - "provider capability name", - capability, - 1, - PROVIDER_COMPONENT_MAX_BYTES, - )?; - if !capability_names.insert(capability.as_str()) { - return Err(StoreError::InvalidPersistenceRecord(format!( - "provider capability `{capability}` is duplicated" - ))); - } - } - let capabilities_json = serde_json::to_string(&record.capabilities)?; - if capabilities_json.len() > PROVIDER_CAPABILITIES_JSON_MAX_BYTES { - return Err(StoreError::InvalidPersistenceRecord(format!( - "provider capability JSON exceeds {PROVIDER_CAPABILITIES_JSON_MAX_BYTES} bytes" - ))); - } - metric_to_i64( - "provider_capabilities.observed_at_unix_ms", - record.observed_at_unix_ms, - )?; - Ok(()) -} - -fn validate_query_cache_record(record: &QueryCacheRecord) -> Result<(), StoreError> { - validate_safe_metadata( - "workspace name", - &record.workspace_name, - 1, - MANUAL_LINK_ID_MAX_BYTES, - )?; - validate_safe_metadata( - "snapshot identifier", - &record.snapshot_id, - 1, - MANUAL_LINK_ID_MAX_BYTES, - )?; - validate_safe_metadata( - "query cache input fingerprint", - &record.input_fingerprint, - 1, - QUERY_CACHE_FINGERPRINT_MAX_BYTES, - )?; - if record.result_summary_json.is_empty() - || record.result_summary_json.len() > QUERY_CACHE_RESULT_MAX_BYTES - || serde_json::from_slice::(&record.result_summary_json).is_err() - { - return Err(StoreError::InvalidPersistenceRecord(format!( - "query cache result must be valid JSON between 1 and {QUERY_CACHE_RESULT_MAX_BYTES} bytes" - ))); - } - if record - .expires_at_unix_ms - .is_some_and(|expires| expires < record.stored_at_unix_ms) - { - return Err(StoreError::InvalidPersistenceRecord( - "query cache expiry precedes storage timestamp".to_owned(), - )); - } - Ok(()) -} - -fn insert_manual_links( - transaction: &rusqlite::Transaction<'_>, - records: &[ManualLinkRecord], -) -> Result<(), StoreError> { - for record in records { - transaction.execute( - "INSERT INTO manual_links( - snapshot_id, id, source_node_id, target_node_id, kind, disposition, reason, - decision_json, config_version - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9)", - params![ - record.snapshot_id, - record.id, - record.source_node_id.as_str(), - record.target_node_id.as_str(), - record.kind, - record.disposition.as_str(), - record.reason, - serde_json::to_vec(&record.decision)?, - i64::from(record.config_version), - ], - )?; - } - Ok(()) -} - -fn insert_graph( - transaction: &rusqlite::Transaction<'_>, - snapshot_id: &str, - nodes: &[Node], - edges: &[Edge], - evidence: &[Evidence], - progress: &mut F, -) -> Result<(), StoreError> -where - F: FnMut(u64), -{ - for node in nodes { - transaction.execute( - "INSERT INTO nodes( - snapshot_id, id, kind, repo_id, stable_key, label - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6)", - params![ - snapshot_id, - node.id.as_str(), - serde_json::to_string(&node.kind)?, - node.repo_id.as_ref().map(RepoId::as_str), - node.stable_key, - node.label, - ], - )?; - progress(1); - transaction.execute( - "INSERT INTO nodes_fts(snapshot_id, node_id, label, stable_key) - VALUES (?1, ?2, ?3, ?4)", - params![snapshot_id, node.id.as_str(), node.label, node.stable_key], - )?; - progress(1); - } - for item in evidence { - transaction.execute( - "INSERT INTO evidence( - snapshot_id, id, repo_id, file_path, start_line, end_line, extractor, - extractor_version, provenance, confidence, observed_at_commit, content_hash - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12)", - params![ - snapshot_id, - item.id.as_str(), - item.repo_id.as_ref().map(RepoId::as_str), - item.file_path, - item.start_line, - item.end_line, - item.extractor, - item.extractor_version, - serde_json::to_string(&item.provenance)?, - item.confidence, - item.observed_at_commit, - item.content_hash, - ], - )?; - } - for edge in edges { - transaction.execute( - "INSERT INTO edges( - snapshot_id, id, source_node_id, target_node_id, kind, confidence, - epistemic_status - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)", - params![ - snapshot_id, - edge.id.as_str(), - edge.source.as_str(), - edge.target.as_str(), - serde_json::to_string(&edge.kind)?, - edge.confidence, - serde_json::to_string(&edge.status)?, - ], - )?; - progress(1); - for evidence_id in &edge.evidence { - transaction.execute( - "INSERT INTO edge_evidence(snapshot_id, edge_id, evidence_id) - VALUES (?1, ?2, ?3)", - params![snapshot_id, edge.id.as_str(), evidence_id.as_str()], - )?; - progress(1); - } - } - Ok(()) -} - -fn insert_extractor_batches( - transaction: &rusqlite::Transaction<'_>, - snapshot_id: &str, - batches: &[StoredExtractorBatch], - progress: &mut F, -) -> Result<(), StoreError> -where - F: FnMut(u64), -{ - for batch in batches { - transaction.execute( - "INSERT INTO extractor_batches( - snapshot_id, repo_id, checkout_id, path_encoding, relative_path, path_display, - extractor, content_hash, size_bytes, extractor_version, budget_fingerprint, - source_was_lossy, output_count, payload - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12, ?13, ?14)", - params![ - snapshot_id, - batch.source.repo_id.as_str(), - batch.source.checkout_id.as_str(), - serde_json::to_string(&batch.source.path.encoding)?, - batch.source.path.bytes, - batch.source.path.display, - batch.source.extractor, - batch.source.content_hash, - metric_to_i64("extractor_batches.size_bytes", batch.source.size_bytes)?, - batch.extractor_version, - batch.budget_fingerprint, - batch.source_was_lossy, - metric_to_i64("extractor_batches.output_count", batch.output_count)?, - batch.payload, - ], - )?; - progress(1); - } - Ok(()) -} - -fn insert_incremental_state( - transaction: &rusqlite::Transaction<'_>, - snapshot_id: &str, - fingerprints: &[ArtifactFingerprint], - extractor_runs: &[ExtractorRun], - progress: &mut F, -) -> Result<(), StoreError> -where - F: FnMut(u64), -{ - for fingerprint in fingerprints { - let size_bytes = metric_to_i64("artifact_fingerprints.size_bytes", fingerprint.size_bytes)?; - transaction.execute( - "INSERT INTO artifact_fingerprints( - snapshot_id, repo_id, checkout_id, path_encoding, relative_path, - path_display, extractor, content_hash, size_bytes - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9)", - params![ - snapshot_id, - fingerprint.repo_id.as_str(), - fingerprint.checkout_id.as_str(), - serde_json::to_string(&fingerprint.path.encoding)?, - fingerprint.path.bytes, - fingerprint.path.display, - fingerprint.extractor, - fingerprint.content_hash, - size_bytes, - ], - )?; - progress(1); - } - for run in extractor_runs { - insert_extractor_run(transaction, snapshot_id, run, fingerprints)?; - progress(1); - } - Ok(()) -} - -fn insert_extractor_run( - transaction: &rusqlite::Transaction<'_>, - snapshot_id: &str, - run: &ExtractorRun, - fingerprints: &[ArtifactFingerprint], -) -> Result<(), StoreError> { - transaction.execute( - "INSERT INTO extractor_runs( - id, snapshot_id, repo_id, extractor, extractor_version, status, - discovered_files, parsed_files, skipped_files, elapsed_ms, checkout_id - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11)", - params![ - run.id, - snapshot_id, - run.repo_id.as_str(), - run.extractor, - run.extractor_version, - serde_json::to_string(&run.status)?, - metric_to_i64("extractor_runs.discovered_files", run.discovered_files)?, - metric_to_i64("extractor_runs.parsed_files", run.parsed_files)?, - metric_to_i64("extractor_runs.skipped_files", run.skipped_files)?, - metric_to_i64("extractor_runs.elapsed_ms", run.elapsed_ms)?, - run.checkout_id.as_str(), - ], - )?; - for fingerprint in fingerprints.iter().filter(|fingerprint| { - fingerprint.repo_id == run.repo_id - && fingerprint.checkout_id == run.checkout_id - && fingerprint.extractor == run.extractor - }) { - transaction.execute( - "INSERT INTO extractor_run_inputs( - run_id, repo_id, checkout_id, path_encoding, relative_path, - extractor, content_hash - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)", - params![ - run.id, - fingerprint.repo_id.as_str(), - fingerprint.checkout_id.as_str(), - serde_json::to_string(&fingerprint.path.encoding)?, - fingerprint.path.bytes, - fingerprint.extractor, - fingerprint.content_hash, - ], - )?; + { + return Err(StoreError::InvalidPersistenceRecord(format!( + "manual link `{}` decision does not match its persisted declaration", + record.id + ))); } Ok(()) } -fn insert_freshness( - transaction: &rusqlite::Transaction<'_>, - snapshot_id: &str, - workspace: &WorkspaceRecord, - extractor_batches: &[StoredExtractorBatch], - coverage_gaps: &[RepositoryCoverageGap], +fn validate_provider_capability_record( + record: &ProviderCapabilityRecord, ) -> Result<(), StoreError> { - let mut unmapped_authorities = BTreeMap::<&RepoId, usize>::new(); - for batch in extractor_batches.iter().filter(|batch| { - batch - .source - .extractor - .starts_with("code-system-graph.source.") - }) { - let count = source_warning_count(&batch.payload, "unmapped_authority")?; - if count > 0 { - *unmapped_authorities - .entry(&batch.source.repo_id) - .or_default() += count; - } + validate_safe_metadata( + "workspace name", + &record.workspace_name, + 1, + MANUAL_LINK_ID_MAX_BYTES, + )?; + validate_safe_metadata( + "repository identifier", + record.repo_id.as_str(), + 1, + MANUAL_LINK_ID_MAX_BYTES, + )?; + validate_safe_metadata( + "provider name", + &record.provider, + 1, + PROVIDER_COMPONENT_MAX_BYTES, + )?; + validate_safe_metadata( + "provider version", + &record.provider_version, + 1, + PROVIDER_COMPONENT_MAX_BYTES, + )?; + if record.capabilities.len() > PROVIDER_CAPABILITY_COUNT_MAX { + return Err(StoreError::InvalidPersistenceRecord(format!( + "provider capability count exceeds {PROVIDER_CAPABILITY_COUNT_MAX}" + ))); } - for repository in &workspace.repositories { - let unmapped_count = unmapped_authorities - .get(&repository.id) - .copied() - .unwrap_or_default(); - let mut coverage_reasons = coverage_gaps - .iter() - .filter(|gap| gap.repo_id == repository.id) - .map(|gap| gap.reason.clone()) - .collect::>(); - if unmapped_count > 0 { - coverage_reasons.push(format!( - "{unmapped_count} absolute HTTP consumer URL{} lacked an explicit workspace authority mapping and {} not linked.", - if unmapped_count == 1 { "" } else { "s" }, - if unmapped_count == 1 { "was" } else { "were" } - )); - } - coverage_reasons.sort(); - coverage_reasons.dedup(); - let (freshness_state, reason) = if !coverage_reasons.is_empty() { - if repository.working_tree_dirty { - coverage_reasons.push( - "The repository working tree differed from HEAD when scanned.".to_owned(), - ); - } - ( - RepoFreshnessState::Partial, - Some(coverage_reasons.join(" ")), - ) - } else if repository.working_tree_dirty { - (RepoFreshnessState::WorkingTreeChanged, None) - } else { - (RepoFreshnessState::Fresh, None) - }; - transaction.execute( - "INSERT INTO repository_snapshot_freshness( - snapshot_id, repo_id, checkout_id, head_commit, manifest_hash, state, reason - ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)", - params![ - snapshot_id, - repository.id.as_str(), - repository.checkout_id.as_str(), - repository.head_commit, - workspace.manifest_hash, - serde_json::to_string(&freshness_state)?, - reason, - ], + let mut capability_names = BTreeSet::new(); + for capability in &record.capabilities { + validate_safe_metadata( + "provider capability name", + capability, + 1, + PROVIDER_COMPONENT_MAX_BYTES, )?; + if !capability_names.insert(capability.as_str()) { + return Err(StoreError::InvalidPersistenceRecord(format!( + "provider capability `{capability}` is duplicated" + ))); + } } - transaction.execute( - "DELETE FROM provider_capabilities - WHERE workspace_name = ?1 - AND NOT EXISTS ( - SELECT 1 FROM workspace_repositories - WHERE workspace_name = ?1 - AND repo_id = provider_capabilities.repo_id - )", - [&workspace.name], + let capabilities_json = serde_json::to_string(&record.capabilities)?; + if capabilities_json.len() > PROVIDER_CAPABILITIES_JSON_MAX_BYTES { + return Err(StoreError::InvalidPersistenceRecord(format!( + "provider capability JSON exceeds {PROVIDER_CAPABILITIES_JSON_MAX_BYTES} bytes" + ))); + } + metric_to_i64( + "provider_capabilities.observed_at_unix_ms", + record.observed_at_unix_ms, )?; Ok(()) } -fn source_warning_count(payload: &[u8], warning: &str) -> Result { - let observations = serde_json::from_slice::>(payload)?; - Ok(observations - .iter() - .filter_map(|observation| observation.get("warnings")?.as_array()) - .flatten() - .filter(|value| value.as_str() == Some(warning)) - .count()) +fn validate_query_cache_record(record: &QueryCacheRecord) -> Result<(), StoreError> { + validate_safe_metadata( + "workspace name", + &record.workspace_name, + 1, + MANUAL_LINK_ID_MAX_BYTES, + )?; + validate_safe_metadata( + "snapshot identifier", + &record.snapshot_id, + 1, + MANUAL_LINK_ID_MAX_BYTES, + )?; + validate_safe_metadata( + "query cache input fingerprint", + &record.input_fingerprint, + 1, + QUERY_CACHE_FINGERPRINT_MAX_BYTES, + )?; + if record.result_summary_json.is_empty() + || record.result_summary_json.len() > QUERY_CACHE_RESULT_MAX_BYTES + || serde_json::from_slice::(&record.result_summary_json).is_err() + { + return Err(StoreError::InvalidPersistenceRecord(format!( + "query cache result must be valid JSON between 1 and {QUERY_CACHE_RESULT_MAX_BYTES} bytes" + ))); + } + if record + .expires_at_unix_ms + .is_some_and(|expires| expires < record.stored_at_unix_ms) + { + return Err(StoreError::InvalidPersistenceRecord( + "query cache expiry precedes storage timestamp".to_owned(), + )); + } + Ok(()) } fn metric_to_i64(field: &'static str, value: u64) -> Result { @@ -3523,26 +3469,26 @@ fn initialize_empty_schema(connection: &mut Connection) -> Result<(), StoreError let transaction = connection.transaction()?; transaction.execute_batch(INITIAL_SCHEMA)?; transaction.execute( - "INSERT INTO schema_metadata(version, instance_id) + "INSERT INTO schema_metadata(schema_id, instance_id) VALUES (?1, lower(hex(randomblob(32))))", - [LATEST_SCHEMA_VERSION], + [schema_identity()], )?; transaction.commit()?; Ok(()) } -fn schema_version(connection: &Connection) -> Result { - Ok(connection.query_row( - "SELECT COALESCE(MAX(version), 0) FROM schema_metadata", - [], - |row| row.get(0), - )?) +fn stored_schema_id(connection: &Connection) -> Result { + Ok( + connection.query_row("SELECT schema_id FROM schema_metadata", [], |row| { + row.get(0) + })?, + ) } fn database_instance_id(connection: &Connection) -> Result { Ok(connection.query_row( - "SELECT instance_id FROM schema_metadata WHERE version = ?1", - [LATEST_SCHEMA_VERSION], + "SELECT instance_id FROM schema_metadata WHERE schema_id = ?1", + [schema_identity()], |row| row.get(0), )?) } @@ -3752,7 +3698,7 @@ mod tests { use std::time::Duration; use code_system_graph_model::{ - ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityId, CommunityMetrics, CommunityScope, CommunitySnapshot, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, ExtractorRunStatus, NativePath, NativePathEncoding, Node, NodeId, NodeKind, Provenance, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, WorkspaceId, WorkspaceRecord + ArtifactFingerprint, CheckoutId, Community, CommunityAlgorithm, CommunityConfig, CommunityId, CommunityMetrics, CommunityScope, CommunitySnapshot, Edge, EdgeId, EdgeKind, EpistemicStatus, Evidence, EvidenceId, ExtractorRun, ExtractorRunStatus, HttpLinkCoverage, HttpLinkGap, HttpLinkGapReason, HttpLinkReport, NativePath, NativePathEncoding, Node, NodeId, NodeKind, Provenance, RepoFreshnessState, RepoId, RepositoryCoverageGap, RepositoryRecord, StoredExtractorBatch, WorkspaceId, WorkspaceRecord }; use rusqlite::params; @@ -4085,13 +4031,20 @@ mod tests { community_snapshot: None, }) .expect("initial publication"); + let relabeled = nodes + .iter() + .map(|node| Node { + label: format!("{} renamed", node.label), + ..node.clone() + }) + .collect::>(); let mut rows = 0_u64; let interrupted = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { let _ = store.publish_snapshot_with_progress( SnapshotBatch { workspace: &workspace, snapshot_id: "snapshot:interrupted", - nodes: &nodes, + nodes: &relabeled, edges: &edges, evidence: &evidence, fingerprints: &[], @@ -4115,6 +4068,57 @@ mod tests { .snapshot_id, "snapshot:previous" ); + let mut expected_nodes = nodes; + expected_nodes.sort_by(|left, right| left.id.cmp(&right.id)); + assert_eq!( + store + .load_current_graph("commerce") + .expect("previous graph") + .0, + expected_nodes + ); + } + + #[test] + fn republication_should_write_only_changed_rows_and_keep_fts_current() + -> Result<(), Box> { + let mut store = SqliteStore::in_memory()?; + let workspace = workspace(); + let (nodes, edges, evidence) = fixture(); + let batch = |snapshot_id, nodes| SnapshotBatch { + workspace: &workspace, + snapshot_id, + nodes, + edges: &edges, + evidence: &evidence, + fingerprints: &[], + extractor_batches: &[], + extractor_runs: &[], + manual_links: &[], + community_snapshot: None, + }; + store.publish_snapshot(batch("snapshot:first", &nodes))?; + let mut unchanged_rows = 0_u64; + store.publish_snapshot_with_progress(batch("snapshot:second", &nodes), |rows| { + unchanged_rows += rows; + })?; + let mut relabeled = nodes.clone(); + relabeled[0].label = "Checkout gateway".to_owned(); + let mut changed_rows = 0_u64; + store.publish_snapshot_with_progress(batch("snapshot:third", &relabeled), |rows| { + changed_rows += rows; + })?; + let hits = store.search_current_nodes("commerce", "gateway", 10)?; + let fts_rows = store + .connection + .query_row("SELECT COUNT(*) FROM nodes_fts", [], |row| { + row.get::<_, i64>(0) + })?; + + assert_eq!((unchanged_rows, changed_rows), (0, 1)); + assert_eq!(hits.len(), 1); + assert_eq!(fts_rows, 2); + Ok(()) } #[test] @@ -4314,7 +4318,7 @@ mod tests { } #[test] - fn historical_graph_and_community_snapshots_should_remain_loadable() { + fn previous_community_snapshot_should_remain_loadable_until_next_publication() { let mut store = match SqliteStore::in_memory() { Ok(store) => store, Err(error) => panic!("test store must initialize: {error}"), @@ -4322,6 +4326,7 @@ mod tests { let workspace = workspace(); let (nodes, edges, evidence) = fixture(); let historical_communities = community_fixture("snapshot:historical"); + let current_communities = community_fixture("snapshot:current"); let first = store.publish_snapshot(SnapshotBatch { workspace: &workspace, snapshot_id: "snapshot:historical", @@ -4345,19 +4350,39 @@ mod tests { extractor_batches: &[], extractor_runs: &[], manual_links: &[], - community_snapshot: None, + community_snapshot: Some(¤t_communities), }); assert!(second.is_ok(), "current fixture failed: {second:?}"); - let result = store - .load_graph_snapshot("snapshot:historical") - .and_then(|graph| Ok((graph, store.load_community_snapshot("snapshot:historical")?))); assert!(matches!( - result, - Ok(((stored_nodes, stored_edges), stored_communities)) - if stored_nodes.len() == 2 - && stored_edges.len() == 1 - && stored_communities == historical_communities + store.load_graph_snapshot("snapshot:historical"), + Err(StoreError::SnapshotMissing(_)) + )); + assert!(matches!( + store.load_workspace_community_snapshot("commerce", "snapshot:historical"), + Ok(stored) if stored == historical_communities + )); + let third_communities = community_fixture("snapshot:third"); + let third = store.publish_snapshot(SnapshotBatch { + workspace: &workspace, + snapshot_id: "snapshot:third", + nodes: &nodes, + edges: &edges, + evidence: &evidence, + fingerprints: &[], + extractor_batches: &[], + extractor_runs: &[], + manual_links: &[], + community_snapshot: Some(&third_communities), + }); + assert!(third.is_ok(), "third fixture failed: {third:?}"); + assert!(matches!( + store.load_community_snapshot("snapshot:historical"), + Err(StoreError::SnapshotMissing(_)) + )); + assert!(matches!( + store.load_community_snapshot("snapshot:current"), + Ok(stored) if stored == current_communities )); } @@ -4419,13 +4444,13 @@ mod tests { "PRAGMA foreign_keys = OFF; INSERT INTO workspaces(name, manifest_hash, id) VALUES ('corrupt', 'hash', 'workspace:corrupt'); - INSERT INTO repo_snapshots(id, workspace_name, is_current) - VALUES ('snapshot:corrupt', 'corrupt', 1); + INSERT INTO repo_snapshots(id, workspace_name) + VALUES ('snapshot:corrupt', 'corrupt'); INSERT INTO edges( - snapshot_id, id, source_node_id, target_node_id, kind, confidence, + workspace_name, id, source_node_id, target_node_id, kind, confidence, epistemic_status ) VALUES ( - 'snapshot:corrupt', 'edge:corrupt', 'missing:a', 'missing:b', + 'corrupt', 'edge:corrupt', 'missing:a', 'missing:b', '\"calls_remote\"', 1.0, '\"confirmed\"' ); PRAGMA foreign_keys = ON;", @@ -4438,9 +4463,9 @@ mod tests { #[test] fn fresh_database_should_apply_initial_schema() { - let result = SqliteStore::in_memory().and_then(|store| store.schema_version()); + let result = SqliteStore::in_memory().and_then(|store| store.schema_id()); - assert!(matches!(result, Ok(2))); + assert!(matches!(result, Ok(schema_id) if schema_id == super::schema_identity())); } #[test] @@ -4462,13 +4487,18 @@ mod tests { 'community_snapshots', 'manual_links', 'provider_capabilities', - 'query_cache' + 'query_cache', + 'http_link_coverage', + 'http_link_gaps' )", [], |row| row.get::<_, i64>(0), )?; - assert_eq!((store.schema_version()?, table_count), (2, 12)); + assert_eq!( + (store.schema_id()?.as_str(), table_count), + (super::schema_identity(), 14) + ); Ok(()) } @@ -4480,7 +4510,10 @@ mod tests { super::validate_exact_schema(&connection)?; super::validate_exact_schema(&connection)?; - assert_eq!(super::schema_version(&connection)?, 2); + assert_eq!( + super::stored_schema_id(&connection)?, + super::schema_identity() + ); Ok(()) } @@ -4500,9 +4533,6 @@ mod tests { WHERE type = 'index' AND name IN ( 'repo_snapshots_id_workspace_idx', - 'manual_links_source_idx', - 'manual_links_target_idx', - 'manual_links_disposition_idx', 'provider_capabilities_repo_provider_idx', 'query_cache_snapshot_idx', 'query_cache_workspace_expiry_idx', @@ -4517,7 +4547,10 @@ mod tests { AND name IN ( 'provider_capabilities_require_registration', 'provider_capabilities_update_require_registration', - 'query_cache_bound_workspace_entries' + 'query_cache_bound_workspace_entries', + 'nodes_fts_insert', + 'nodes_fts_update', + 'nodes_fts_delete' )", [], |row| row.get::<_, i64>(0), @@ -4537,23 +4570,23 @@ mod tests { assert_eq!( (strict_tables, indexes, triggers, foreign_keys), - ((3, 3), 8, 3, 9) + ((3, 3), 5, 6, 8) ); Ok(()) } #[test] - fn exact_schema_should_reject_unknown_version() { + fn exact_schema_should_reject_unknown_schema_identity() { let connection = match rusqlite::Connection::open_in_memory() { Ok(connection) => connection, Err(error) => panic!("test connection must initialize: {error}"), }; let setup = connection.execute_batch( "CREATE TABLE schema_metadata ( - version INTEGER PRIMARY KEY, + schema_id TEXT PRIMARY KEY, applied_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP ); - INSERT INTO schema_metadata(version) VALUES (999);", + INSERT INTO schema_metadata(schema_id) VALUES ('schema:unknown');", ); assert!(setup.is_ok(), "newer schema fixture failed: {setup:?}"); let result = SqliteStore::from_connection(connection); @@ -4562,13 +4595,13 @@ mod tests { } #[test] - fn exact_schema_should_reject_structurally_current_v1_database() + fn exact_schema_should_reject_structurally_identical_database_with_other_identity() -> Result<(), Box> { let mut connection = rusqlite::Connection::open_in_memory()?; super::initialize_empty_schema(&mut connection)?; connection.execute( - "UPDATE schema_metadata SET version = 1 WHERE version = ?1", - [super::LATEST_SCHEMA_VERSION], + "UPDATE schema_metadata SET schema_id = 'schema:other' WHERE schema_id = ?1", + [super::schema_identity()], )?; let result = SqliteStore::from_connection(connection); @@ -4631,7 +4664,7 @@ mod tests { let connection = rusqlite::Connection::open(&database)?; connection.execute_batch( "CREATE TABLE schema_metadata ( - version INTEGER PRIMARY KEY, + schema_id TEXT PRIMARY KEY, applied_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP );", )?; @@ -4751,10 +4784,74 @@ mod tests { let database = temporary.path().join("store.db"); let initial = SqliteStore::open(&database)?; - assert_eq!(initial.schema_version()?, 2); + assert_eq!(initial.schema_id()?, super::schema_identity()); drop(initial); let repeated = SqliteStore::open(&database)?; - assert_eq!(repeated.schema_version()?, 2); + assert_eq!(repeated.schema_id()?, super::schema_identity()); + Ok(()) + } + + #[test] + fn composite_reads_should_keep_graph_and_evidence_during_publication() + -> Result<(), Box> { + let temporary = tempfile::tempdir()?; + let database = temporary.path().join("composite.db"); + let workspace = workspace(); + let (mut nodes, edges, evidence) = fixture(); + nodes.sort_by(|left, right| left.id.cmp(&right.id)); + let mut writer = SqliteStore::open(&database)?; + writer.publish_snapshot(SnapshotBatch { + workspace: &workspace, + snapshot_id: "snapshot:before", + nodes: &nodes, + edges: &edges, + evidence: &evidence, + fingerprints: &[], + extractor_batches: &[], + extractor_runs: &[], + manual_links: &[], + community_snapshot: None, + })?; + let reader = SqliteStore::open_read_only(&database)?; + let mut next_nodes = nodes.clone(); + next_nodes[0].label = "POST /updated".to_owned(); + let mut next_evidence = evidence.clone(); + next_evidence[0].content_hash = Some("updated".to_owned()); + + let (read_nodes, read_edges, read_evidence) = reader.consistent_read(|| { + let snapshot = reader.current_snapshot_summary("commerce")?; + let (read_nodes, read_edges) = reader.load_graph_snapshot(&snapshot.snapshot_id)?; + // Publish through a separate connection after the graph read has finished. + writer.publish_snapshot(SnapshotBatch { + workspace: &workspace, + snapshot_id: "snapshot:after", + nodes: &next_nodes, + edges: &edges, + evidence: &next_evidence, + fingerprints: &[], + extractor_batches: &[], + extractor_runs: &[], + manual_links: &[], + community_snapshot: None, + })?; + let read_evidence = reader.load_evidence_snapshot(&snapshot.snapshot_id)?; + Ok::<_, StoreError>((read_nodes, read_edges, read_evidence)) + })?; + assert_eq!(read_nodes, nodes); + assert_eq!(read_edges, edges); + assert_eq!(read_evidence, evidence); + assert!(reader.connection.is_autocommit()); + assert_eq!(reader.load_current_graph("commerce")?.0, next_nodes); + assert_eq!(reader.load_current_evidence("commerce")?, next_evidence); + assert!(matches!( + reader.load_evidence_snapshot("snapshot:before"), + Err(StoreError::SnapshotMissing(_)) + )); + + let failed = reader.consistent_read(|| reader.load_graph_snapshot("snapshot:missing")); + assert!(matches!(failed, Err(StoreError::SnapshotMissing(_)))); + assert!(reader.connection.is_autocommit()); + assert_eq!(reader.load_current_graph("commerce")?.0, next_nodes); Ok(()) } @@ -4810,7 +4907,7 @@ mod tests { .and_then(|result| result) }); - assert!(reader_results.into_iter().all(|result| result.is_ok())); + assert_eq!(reader_results, [Ok(()), Ok(())]); Ok(()) } @@ -4942,65 +5039,6 @@ mod tests { )); } - #[test] - fn published_freshness_preserves_dirty_and_unmapped_authority_coverage() { - let mut store = SqliteStore::in_memory().expect("test store"); - let mut workspace = workspace(); - workspace.repositories[0].working_tree_dirty = true; - let (nodes, edges, evidence) = fixture(); - let batch = StoredExtractorBatch { - source: ArtifactFingerprint { - repo_id: RepoId::new("repo:web"), - checkout_id: CheckoutId::new("checkout:web"), - path: NativePath { - encoding: NativePathEncoding::Utf8, - bytes: b"src/client.rs".to_vec(), - display: "src/client.rs".to_owned(), - }, - extractor: "code-system-graph.source.rust".to_owned(), - content_hash: "content:http".to_owned(), - size_bytes: 1, - }, - extractor_version: "1.1.0".to_owned(), - budget_fingerprint: "budget".to_owned(), - source_was_lossy: false, - output_count: 1, - payload: br#"[{"warnings":["unmapped_authority"]}]"#.to_vec(), - }; - - store - .publish_snapshot(SnapshotBatch { - workspace: &workspace, - snapshot_id: "snapshot:coverage", - nodes: &nodes, - edges: &edges, - evidence: &evidence, - fingerprints: &[], - extractor_batches: &[batch], - extractor_runs: &[], - manual_links: &[], - community_snapshot: None, - }) - .expect("publish"); - let freshness = store.load_current_freshness("commerce").expect("freshness"); - - assert_eq!(freshness[0].state, RepoFreshnessState::Partial); - assert!(freshness[0].reason.as_deref().is_some_and(|reason| { - reason.contains("authority mapping") && reason.contains("working tree") - })); - } - - #[test] - fn source_warning_count_should_ignore_matching_symbol_text() { - let payload = br#"[{"symbol_name":"unmapped_authority","warnings":[]}]"#; - - assert_eq!( - super::source_warning_count(payload, "unmapped_authority") - .expect("valid source payload"), - 0 - ); - } - #[test] fn unresolved_repository_dependency_gap_should_persist_as_partial_freshness() { let mut store = SqliteStore::in_memory().expect("test store"); @@ -5010,6 +5048,30 @@ mod tests { repo_id: RepoId::new("repo:web"), reason: "An explicit repository dependency could not be matched exactly.".to_owned(), }]; + let http_links = HttpLinkReport { + coverage: HttpLinkCoverage { + linked: 3, + no_provider: 1, + ambiguous: 1, + external: 0, + }, + gaps: vec![ + HttpLinkGap { + caller: NodeId::new("test:orders"), + method: "GET".to_owned(), + path: "/missing".to_owned(), + reason: HttpLinkGapReason::NoProvider, + candidates: Vec::new(), + }, + HttpLinkGap { + caller: NodeId::new("consumer:web"), + method: "GET".to_owned(), + path: "/orders/{id}".to_owned(), + reason: HttpLinkGapReason::Ambiguous, + candidates: vec![NodeId::new("provider:a"), NodeId::new("provider:b")], + }, + ], + }; store .publish_snapshot_with_progress_and_coverage( @@ -5025,12 +5087,23 @@ mod tests { manual_links: &[], community_snapshot: None, }, + None, &gaps, + &http_links, |_| {}, ) .expect("publish"); let freshness = store.load_current_freshness("commerce").expect("freshness"); + let stored_links = store + .load_http_link_report("commerce", 10) + .expect("HTTP link report"); + + let caller_gaps = store + .load_http_link_gaps_for_callers("commerce", &[NodeId::new("consumer:web")]) + .expect("caller gaps"); + assert_eq!(caller_gaps, http_links.gaps[1..]); + assert_eq!(stored_links, (http_links, 2)); assert_eq!(freshness[0].state, RepoFreshnessState::Partial); assert!( freshness[0] @@ -5095,7 +5168,7 @@ mod tests { } #[test] - fn manual_links_should_round_trip_and_preserve_snapshot_history() + fn manual_links_should_round_trip_with_the_current_snapshot() -> Result<(), Box> { let mut store = SqliteStore::in_memory()?; let workspace = workspace(); @@ -5129,13 +5202,13 @@ mod tests { current_links.truncate(1); store.persist_manual_links("snapshot:current-links", ¤t_links)?; - let stored_historical = store.load_manual_links("snapshot:historical-links")?; let stored_current = store.load_manual_links("snapshot:current-links")?; - assert_eq!( - (stored_historical, stored_current), - (historical_links, current_links) - ); + assert!(matches!( + store.load_manual_links("snapshot:historical-links"), + Err(StoreError::SnapshotMissing(_)) + )); + assert_eq!(stored_current, current_links); Ok(()) } diff --git a/crates/code-system-graph-store-sqlite/src/publication.rs b/crates/code-system-graph-store-sqlite/src/publication.rs new file mode 100644 index 0000000..e010d0c --- /dev/null +++ b/crates/code-system-graph-store-sqlite/src/publication.rs @@ -0,0 +1,992 @@ +//! Delta publication of the current workspace graph. +//! +//! Each workspace owns exactly one current graph. Publication compares the candidate with the +//! stored rows by identifier and content digest, then writes only inserted, changed, and removed +//! rows inside the caller's transaction. Readers observe either the previous or the new graph. + +use std::collections::{BTreeSet, HashMap}; + +use code_system_graph_model::{ + ArtifactFingerprint, CommunitySnapshot, Edge, Evidence, EvidenceId, ExtractorRun, HttpLinkReport, Node, NodeId, RepoFreshnessState, RepoId, RepositoryCoverageGap, StoredExtractorBatch, WorkspaceRecord +}; +use rusqlite::types::Type; +use rusqlite::{OptionalExtension, Row, Transaction, params}; + +use super::{ + ArtifactDelta, ManualLinkRecord, StoreError, insert_community_snapshot, metric_to_i64 +}; + +type Digest = [u8; 32]; +type ArtifactRowKey = (String, String, String, Vec, String); + +/// Candidate graph rows that replace the current workspace graph. +pub(crate) struct PublicationInput<'a> { + pub(crate) workspace: &'a WorkspaceRecord, + pub(crate) snapshot_id: &'a str, + pub(crate) nodes: &'a [Node], + pub(crate) edges: &'a [Edge], + pub(crate) evidence: &'a [Evidence], + pub(crate) fingerprints: &'a [ArtifactFingerprint], + pub(crate) extractor_batches: &'a [StoredExtractorBatch], + pub(crate) artifact_delta: Option>, + pub(crate) extractor_runs: &'a [ExtractorRun], + pub(crate) manual_links: &'a [ManualLinkRecord], + pub(crate) community_snapshot: Option<&'a CommunitySnapshot>, + pub(crate) coverage_gaps: &'a [RepositoryCoverageGap], + pub(crate) http_links: &'a HttpLinkReport, +} + +/// Writes the candidate as the current graph of its workspace. +pub(crate) fn publish_current_graph( + transaction: &Transaction<'_>, + input: &PublicationInput<'_>, + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let workspace = input.workspace.name.as_str(); + let previous_snapshot_id = transaction + .query_row( + "SELECT id FROM repo_snapshots WHERE workspace_name = ?1", + [workspace], + |row| row.get::<_, String>(0), + ) + .optional()?; + transaction.execute( + "DELETE FROM query_cache WHERE workspace_name = ?1", + [workspace], + )?; + transaction.execute( + "INSERT INTO repo_snapshots(workspace_name, id) VALUES (?1, ?2) + ON CONFLICT(workspace_name) DO UPDATE SET + id = excluded.id, + created_at = CURRENT_TIMESTAMP", + params![workspace, input.snapshot_id], + )?; + + let removed_evidence = upsert_evidence(transaction, workspace, input.evidence, progress)?; + let removed_nodes = upsert_nodes(transaction, workspace, input.nodes, progress)?; + replace_changed_edges(transaction, workspace, input.edges, progress)?; + replace_manual_links(transaction, workspace, input.manual_links)?; + progress(u64::try_from(input.manual_links.len()).unwrap_or(u64::MAX)); + delete_by_id( + transaction, + "DELETE FROM nodes WHERE workspace_name = ?1 AND id = ?2", + workspace, + &removed_nodes, + progress, + )?; + delete_by_id( + transaction, + "DELETE FROM evidence WHERE workspace_name = ?1 AND id = ?2", + workspace, + &removed_evidence, + progress, + )?; + + if let Some(delta) = input.artifact_delta { + apply_artifact_delta(transaction, workspace, delta, progress)?; + } else { + upsert_fingerprints(transaction, workspace, input.fingerprints, progress)?; + upsert_extractor_batches(transaction, workspace, input.extractor_batches, progress)?; + } + replace_extractor_runs(transaction, workspace, input.extractor_runs, progress)?; + replace_freshness(transaction, input.workspace, input.coverage_gaps)?; + replace_http_links(transaction, workspace, input.http_links)?; + + if let Some(community_snapshot) = input.community_snapshot { + transaction.execute( + "DELETE FROM community_snapshots WHERE snapshot_id = ?1", + [&community_snapshot.snapshot_id], + )?; + insert_community_snapshot(transaction, workspace, community_snapshot, progress)?; + } + let retained_previous = previous_snapshot_id + .as_deref() + .filter(|previous| *previous != input.snapshot_id) + .unwrap_or(input.snapshot_id); + transaction.execute( + "DELETE FROM community_snapshots + WHERE workspace_name = ?1 AND snapshot_id NOT IN (?2, ?3)", + params![workspace, input.snapshot_id, retained_previous], + )?; + Ok(()) +} + +struct DigestHasher(blake3::Hasher); + +impl DigestHasher { + fn new() -> Self { + Self(blake3::Hasher::new()) + } + + fn text(&mut self, value: &str) -> &mut Self { + self.bytes(value.as_bytes()) + } + + fn optional_text(&mut self, value: Option<&str>) -> &mut Self { + if let Some(value) = value { + self.0.update(&[1]); + self.text(value) + } else { + self.0.update(&[0]); + self + } + } + + fn bytes(&mut self, value: &[u8]) -> &mut Self { + self.0.update(&(value.len() as u64).to_le_bytes()); + self.0.update(value); + self + } + + fn integer(&mut self, value: Option) -> &mut Self { + match value { + Some(value) => { + self.0.update(&[1]); + self.0.update(&value.to_le_bytes()); + } + None => { + self.0.update(&[0]); + } + } + self + } + + fn finish(&self) -> Digest { + *self.0.finalize().as_bytes() + } +} + +fn node_digest(kind: &str, repo_id: Option<&str>, stable_key: &str, label: &str) -> Digest { + DigestHasher::new() + .text(kind) + .optional_text(repo_id) + .text(stable_key) + .text(label) + .finish() +} + +fn upsert_nodes( + transaction: &Transaction<'_>, + workspace: &str, + nodes: &[Node], + progress: &mut F, +) -> Result, StoreError> +where + F: FnMut(u64), +{ + let mut stored = HashMap::::new(); + { + let mut statement = transaction.prepare_cached( + "SELECT id, kind, repo_id, stable_key, label FROM nodes WHERE workspace_name = ?1", + )?; + let mut rows = statement.query([workspace])?; + while let Some(row) = rows.next()? { + let digest = node_digest( + column_text(row, 1)?, + optional_column_text(row, 2)?, + column_text(row, 3)?, + column_text(row, 4)?, + ); + stored.insert(row.get(0)?, digest); + } + } + let mut insert = transaction.prepare_cached( + "INSERT INTO nodes(workspace_name, id, kind, repo_id, stable_key, label) + VALUES (?1, ?2, ?3, ?4, ?5, ?6) + ON CONFLICT(workspace_name, id) DO UPDATE SET + kind = excluded.kind, + repo_id = excluded.repo_id, + stable_key = excluded.stable_key, + label = excluded.label", + )?; + for node in nodes { + let kind = serde_json::to_string(&node.kind)?; + let repo_id = node.repo_id.as_ref().map(RepoId::as_str); + let digest = node_digest(&kind, repo_id, &node.stable_key, &node.label); + if stored.remove(node.id.as_str()) == Some(digest) { + continue; + } + insert.execute(params![ + workspace, + node.id.as_str(), + kind, + repo_id, + node.stable_key, + node.label, + ])?; + progress(1); + } + Ok(sorted_keys(stored)) +} + +fn evidence_digest(item: &Evidence, provenance: &str) -> Digest { + DigestHasher::new() + .optional_text(item.repo_id.as_ref().map(RepoId::as_str)) + .optional_text(item.file_path.as_deref()) + .integer(item.start_line.map(i64::from)) + .integer(item.end_line.map(i64::from)) + .text(&item.extractor) + .text(&item.extractor_version) + .text(provenance) + .integer(Some(i64::from(item.confidence.to_bits()))) + .optional_text(item.content_hash.as_deref()) + .optional_text(item.observed_at_commit.as_deref()) + .finish() +} + +fn upsert_evidence( + transaction: &Transaction<'_>, + workspace: &str, + evidence: &[Evidence], + progress: &mut F, +) -> Result, StoreError> +where + F: FnMut(u64), +{ + let mut stored = HashMap::::new(); + { + let mut statement = transaction.prepare_cached( + "SELECT id, repo_id, file_path, start_line, end_line, extractor, extractor_version, + provenance, confidence, content_hash, observed_at_commit + FROM evidence WHERE workspace_name = ?1", + )?; + let mut rows = statement.query([workspace])?; + while let Some(row) = rows.next()? { + let confidence = row.get::<_, f32>(8)?; + let digest = DigestHasher::new() + .optional_text(optional_column_text(row, 1)?) + .optional_text(optional_column_text(row, 2)?) + .integer(row.get::<_, Option>(3)?) + .integer(row.get::<_, Option>(4)?) + .text(column_text(row, 5)?) + .text(column_text(row, 6)?) + .text(column_text(row, 7)?) + .integer(Some(i64::from(confidence.to_bits()))) + .optional_text(optional_column_text(row, 9)?) + .optional_text(optional_column_text(row, 10)?) + .finish(); + stored.insert(row.get(0)?, digest); + } + } + let mut insert = transaction.prepare_cached( + "INSERT INTO evidence( + workspace_name, id, repo_id, file_path, start_line, end_line, extractor, + extractor_version, provenance, confidence, observed_at_commit, content_hash + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12) + ON CONFLICT(workspace_name, id) DO UPDATE SET + repo_id = excluded.repo_id, + file_path = excluded.file_path, + start_line = excluded.start_line, + end_line = excluded.end_line, + extractor = excluded.extractor, + extractor_version = excluded.extractor_version, + provenance = excluded.provenance, + confidence = excluded.confidence, + observed_at_commit = excluded.observed_at_commit, + content_hash = excluded.content_hash", + )?; + for item in evidence { + let provenance = serde_json::to_string(&item.provenance)?; + let digest = evidence_digest(item, &provenance); + if stored.remove(item.id.as_str()) == Some(digest) { + continue; + } + insert.execute(params![ + workspace, + item.id.as_str(), + item.repo_id.as_ref().map(RepoId::as_str), + item.file_path, + item.start_line, + item.end_line, + item.extractor, + item.extractor_version, + provenance, + item.confidence, + item.observed_at_commit, + item.content_hash, + ])?; + progress(1); + } + Ok(sorted_keys(stored)) +} + +fn edge_digest<'a>( + source: &str, + target: &str, + kind: &str, + confidence: f32, + status: &str, + evidence: impl Iterator, +) -> Digest { + let mut hasher = DigestHasher::new(); + hasher + .text(source) + .text(target) + .text(kind) + .integer(Some(i64::from(confidence.to_bits()))) + .text(status); + for evidence_id in evidence { + hasher.text(evidence_id); + } + hasher.finish() +} + +fn replace_changed_edges( + transaction: &Transaction<'_>, + workspace: &str, + edges: &[Edge], + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut stored_evidence = HashMap::>::new(); + { + let mut statement = transaction.prepare_cached( + "SELECT edge_id, evidence_id FROM edge_evidence + WHERE workspace_name = ?1 ORDER BY edge_id, evidence_id", + )?; + let mut rows = statement.query([workspace])?; + while let Some(row) = rows.next()? { + stored_evidence + .entry(row.get(0)?) + .or_default() + .push(row.get(1)?); + } + } + let mut stored = HashMap::::new(); + { + let mut statement = transaction.prepare_cached( + "SELECT id, source_node_id, target_node_id, kind, confidence, epistemic_status + FROM edges WHERE workspace_name = ?1", + )?; + let mut rows = statement.query([workspace])?; + while let Some(row) = rows.next()? { + let id = row.get::<_, String>(0)?; + let evidence = stored_evidence.remove(&id).unwrap_or_default(); + let digest = edge_digest( + column_text(row, 1)?, + column_text(row, 2)?, + column_text(row, 3)?, + row.get::<_, f32>(4)?, + column_text(row, 5)?, + evidence.iter().map(String::as_str), + ); + stored.insert(id, digest); + } + } + let mut changed = Vec::new(); + for edge in edges { + let kind = serde_json::to_string(&edge.kind)?; + let status = serde_json::to_string(&edge.status)?; + let mut evidence = edge + .evidence + .iter() + .map(EvidenceId::as_str) + .collect::>(); + evidence.sort_unstable(); + let digest = edge_digest( + edge.source.as_str(), + edge.target.as_str(), + &kind, + edge.confidence, + &status, + evidence.into_iter(), + ); + match stored.remove(edge.id.as_str()) { + Some(previous) if previous == digest => {} + Some(_) => changed.push((edge, kind, status, true)), + None => changed.push((edge, kind, status, false)), + } + } + let mut delete = + transaction.prepare_cached("DELETE FROM edges WHERE workspace_name = ?1 AND id = ?2")?; + for removed in sorted_keys(stored) { + delete.execute(params![workspace, removed])?; + progress(1); + } + let mut insert_edge = transaction.prepare_cached( + "INSERT INTO edges( + workspace_name, id, source_node_id, target_node_id, kind, confidence, + epistemic_status + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)", + )?; + let mut insert_evidence = transaction.prepare_cached( + "INSERT INTO edge_evidence(workspace_name, edge_id, evidence_id) VALUES (?1, ?2, ?3)", + )?; + for (edge, kind, status, existed) in changed { + if existed { + delete.execute(params![workspace, edge.id.as_str()])?; + } + insert_edge.execute(params![ + workspace, + edge.id.as_str(), + edge.source.as_str(), + edge.target.as_str(), + kind, + edge.confidence, + status, + ])?; + progress(1); + for evidence_id in &edge.evidence { + insert_evidence.execute(params![workspace, edge.id.as_str(), evidence_id.as_str()])?; + progress(1); + } + } + Ok(()) +} + +fn insert_manual_links( + transaction: &Transaction<'_>, + workspace: &str, + records: &[ManualLinkRecord], +) -> Result<(), StoreError> { + let mut statement = transaction.prepare_cached( + "INSERT INTO manual_links( + workspace_name, id, source_node_id, target_node_id, kind, disposition, reason, + decision_json, config_version + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9)", + )?; + for record in records { + statement.execute(params![ + workspace, + record.id, + record.source_node_id.as_str(), + record.target_node_id.as_str(), + record.kind, + record.disposition.as_str(), + record.reason, + serde_json::to_vec(&record.decision)?, + i64::from(record.config_version), + ])?; + } + Ok(()) +} + +/// Replaces manual link declarations of the current workspace graph. +pub(crate) fn replace_manual_links( + transaction: &Transaction<'_>, + workspace: &str, + records: &[ManualLinkRecord], +) -> Result<(), StoreError> { + transaction.execute( + "DELETE FROM manual_links WHERE workspace_name = ?1", + [workspace], + )?; + insert_manual_links(transaction, workspace, records) +} + +fn delete_by_id( + transaction: &Transaction<'_>, + sql: &str, + workspace: &str, + ids: &[String], + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut statement = transaction.prepare_cached(sql)?; + for id in ids { + statement.execute(params![workspace, id])?; + progress(1); + } + Ok(()) +} + +fn artifact_key(fingerprint: &ArtifactFingerprint) -> Result { + Ok(( + fingerprint.repo_id.as_str().to_owned(), + fingerprint.checkout_id.as_str().to_owned(), + serde_json::to_string(&fingerprint.path.encoding)?, + fingerprint.path.bytes.clone(), + fingerprint.extractor.clone(), + )) +} + +fn load_artifact_digests( + transaction: &Transaction<'_>, + sql: &str, + workspace: &str, + digest_columns: usize, +) -> Result, StoreError> { + let mut stored = HashMap::new(); + let mut statement = transaction.prepare_cached(sql)?; + let mut rows = statement.query([workspace])?; + while let Some(row) = rows.next()? { + let mut hasher = DigestHasher::new(); + for index in 5..5 + digest_columns { + match row.get_ref(index)? { + rusqlite::types::ValueRef::Integer(value) => { + hasher.integer(Some(value)); + } + _ => { + hasher.text(column_text(row, index)?); + } + } + } + stored.insert( + ( + row.get(0)?, + row.get(1)?, + row.get(2)?, + row.get(3)?, + row.get(4)?, + ), + hasher.finish(), + ); + } + Ok(stored) +} + +fn delete_artifact_rows( + transaction: &Transaction<'_>, + table: &str, + workspace: &str, + stored: HashMap, + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut removed = stored.into_keys().collect::>(); + removed.sort_unstable(); + let mut statement = transaction.prepare_cached(&format!( + "DELETE FROM {table} + WHERE workspace_name = ?1 AND repo_id = ?2 AND checkout_id = ?3 + AND path_encoding = ?4 AND relative_path = ?5 AND extractor = ?6" + ))?; + for (repo_id, checkout_id, encoding, path, extractor) in removed { + statement.execute(params![ + workspace, + repo_id, + checkout_id, + encoding, + path, + extractor + ])?; + progress(1); + } + Ok(()) +} + +const UPSERT_FINGERPRINT_SQL: &str = "INSERT INTO artifact_fingerprints( + workspace_name, repo_id, checkout_id, path_encoding, relative_path, + path_display, extractor, content_hash, size_bytes + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9) + ON CONFLICT( + workspace_name, repo_id, checkout_id, path_encoding, relative_path, extractor + ) DO UPDATE SET + path_display = excluded.path_display, + content_hash = excluded.content_hash, + size_bytes = excluded.size_bytes"; + +const UPSERT_BATCH_SQL: &str = "INSERT INTO extractor_batches( + workspace_name, repo_id, checkout_id, path_encoding, relative_path, path_display, + extractor, content_hash, size_bytes, extractor_version, budget_fingerprint, + source_was_lossy, output_count, payload_hash, payload + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12, ?13, ?14, ?15) + ON CONFLICT( + workspace_name, repo_id, checkout_id, path_encoding, relative_path, extractor + ) DO UPDATE SET + path_display = excluded.path_display, + content_hash = excluded.content_hash, + size_bytes = excluded.size_bytes, + extractor_version = excluded.extractor_version, + budget_fingerprint = excluded.budget_fingerprint, + source_was_lossy = excluded.source_was_lossy, + output_count = excluded.output_count, + payload_hash = excluded.payload_hash, + payload = excluded.payload"; + +fn write_fingerprint( + insert: &mut rusqlite::CachedStatement<'_>, + workspace: &str, + fingerprint: &ArtifactFingerprint, +) -> Result<(), StoreError> { + insert.execute(params![ + workspace, + fingerprint.repo_id.as_str(), + fingerprint.checkout_id.as_str(), + serde_json::to_string(&fingerprint.path.encoding)?, + fingerprint.path.bytes, + fingerprint.path.display, + fingerprint.extractor, + fingerprint.content_hash, + metric_to_i64("artifact_fingerprints.size_bytes", fingerprint.size_bytes)?, + ])?; + Ok(()) +} + +fn write_batch( + insert: &mut rusqlite::CachedStatement<'_>, + workspace: &str, + batch: &StoredExtractorBatch, + payload_hash: &str, +) -> Result<(), StoreError> { + insert.execute(params![ + workspace, + batch.source.repo_id.as_str(), + batch.source.checkout_id.as_str(), + serde_json::to_string(&batch.source.path.encoding)?, + batch.source.path.bytes, + batch.source.path.display, + batch.source.extractor, + batch.source.content_hash, + metric_to_i64("extractor_batches.size_bytes", batch.source.size_bytes)?, + batch.extractor_version, + batch.budget_fingerprint, + batch.source_was_lossy, + metric_to_i64("extractor_batches.output_count", batch.output_count)?, + payload_hash, + batch.payload, + ])?; + Ok(()) +} + +/// Writes only the artifact rows that changed since the current graph. +fn apply_artifact_delta( + transaction: &Transaction<'_>, + workspace: &str, + delta: ArtifactDelta<'_>, + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut insert_fingerprint = transaction.prepare_cached(UPSERT_FINGERPRINT_SQL)?; + for fingerprint in delta.upserted_fingerprints { + write_fingerprint(&mut insert_fingerprint, workspace, fingerprint)?; + progress(1); + } + let mut insert_batch = transaction.prepare_cached(UPSERT_BATCH_SQL)?; + for batch in delta.upserted_batches { + let payload_hash = blake3::hash(&batch.payload).to_hex().to_string(); + write_batch(&mut insert_batch, workspace, batch, &payload_hash)?; + progress(1); + } + for table in ["artifact_fingerprints", "extractor_batches"] { + let mut delete = transaction.prepare_cached(&format!( + "DELETE FROM {table} + WHERE workspace_name = ?1 AND repo_id = ?2 AND checkout_id = ?3 + AND path_encoding = ?4 AND relative_path = ?5 AND extractor = ?6" + ))?; + for removed in delta.removed { + delete.execute(params![ + workspace, + removed.repo_id.as_str(), + removed.checkout_id.as_str(), + serde_json::to_string(&removed.path.encoding)?, + removed.path.bytes, + removed.extractor, + ])?; + progress(1); + } + } + Ok(()) +} + +fn upsert_fingerprints( + transaction: &Transaction<'_>, + workspace: &str, + fingerprints: &[ArtifactFingerprint], + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut stored = load_artifact_digests( + transaction, + "SELECT repo_id, checkout_id, path_encoding, relative_path, extractor, + path_display, content_hash, size_bytes + FROM artifact_fingerprints WHERE workspace_name = ?1", + workspace, + 3, + )?; + let mut insert = transaction.prepare_cached(UPSERT_FINGERPRINT_SQL)?; + for fingerprint in fingerprints { + let size_bytes = metric_to_i64("artifact_fingerprints.size_bytes", fingerprint.size_bytes)?; + let digest = DigestHasher::new() + .text(&fingerprint.path.display) + .text(&fingerprint.content_hash) + .integer(Some(size_bytes)) + .finish(); + let key = artifact_key(fingerprint)?; + if stored.remove(&key) == Some(digest) { + continue; + } + write_fingerprint(&mut insert, workspace, fingerprint)?; + progress(1); + } + delete_artifact_rows( + transaction, + "artifact_fingerprints", + workspace, + stored, + progress, + ) +} + +fn upsert_extractor_batches( + transaction: &Transaction<'_>, + workspace: &str, + batches: &[StoredExtractorBatch], + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + let mut stored = load_artifact_digests( + transaction, + "SELECT repo_id, checkout_id, path_encoding, relative_path, extractor, + path_display, content_hash, size_bytes, extractor_version, budget_fingerprint, + source_was_lossy, output_count, payload_hash + FROM extractor_batches WHERE workspace_name = ?1", + workspace, + 8, + )?; + let mut insert = transaction.prepare_cached(UPSERT_BATCH_SQL)?; + for batch in batches { + let size_bytes = metric_to_i64("extractor_batches.size_bytes", batch.source.size_bytes)?; + let output_count = metric_to_i64("extractor_batches.output_count", batch.output_count)?; + let payload_hash = blake3::hash(&batch.payload).to_hex().to_string(); + let digest = DigestHasher::new() + .text(&batch.source.path.display) + .text(&batch.source.content_hash) + .integer(Some(size_bytes)) + .text(&batch.extractor_version) + .text(&batch.budget_fingerprint) + .integer(Some(i64::from(batch.source_was_lossy))) + .integer(Some(output_count)) + .text(&payload_hash) + .finish(); + let key = artifact_key(&batch.source)?; + if stored.remove(&key) == Some(digest) { + continue; + } + write_batch(&mut insert, workspace, batch, &payload_hash)?; + progress(1); + } + delete_artifact_rows( + transaction, + "extractor_batches", + workspace, + stored, + progress, + ) +} + +fn replace_extractor_runs( + transaction: &Transaction<'_>, + workspace: &str, + runs: &[ExtractorRun], + progress: &mut F, +) -> Result<(), StoreError> +where + F: FnMut(u64), +{ + transaction.execute( + "DELETE FROM extractor_runs WHERE workspace_name = ?1", + [workspace], + )?; + let mut statement = transaction.prepare_cached( + "INSERT INTO extractor_runs( + workspace_name, id, repo_id, extractor, extractor_version, status, + discovered_files, parsed_files, skipped_files, elapsed_ms, checkout_id + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11)", + )?; + for run in runs { + statement.execute(params![ + workspace, + run.id, + run.repo_id.as_str(), + run.extractor, + run.extractor_version, + serde_json::to_string(&run.status)?, + metric_to_i64("extractor_runs.discovered_files", run.discovered_files)?, + metric_to_i64("extractor_runs.parsed_files", run.parsed_files)?, + metric_to_i64("extractor_runs.skipped_files", run.skipped_files)?, + metric_to_i64("extractor_runs.elapsed_ms", run.elapsed_ms)?, + run.checkout_id.as_str(), + ])?; + progress(1); + } + Ok(()) +} + +/// Persisted key of one HTTP link gap: caller, method, path, reason, and candidates JSON. +type HttpLinkGapRow = (String, String, String, String, String); + +/// Replaces the HTTP link coverage of `workspace`, writing only rows that changed. +fn replace_http_links( + transaction: &Transaction<'_>, + workspace: &str, + report: &HttpLinkReport, +) -> Result<(), StoreError> { + let coverage = report.coverage; + let count = |value: u64, field: &'static str| { + i64::try_from(value).map_err(|_| StoreError::IntegerOutOfRange { + field, + value: i128::from(value), + }) + }; + transaction.execute( + "INSERT INTO http_link_coverage(workspace_name, linked, no_provider, ambiguous, external) + VALUES (?1, ?2, ?3, ?4, ?5) + ON CONFLICT(workspace_name) DO UPDATE SET + linked = excluded.linked, + no_provider = excluded.no_provider, + ambiguous = excluded.ambiguous, + external = excluded.external + WHERE linked != excluded.linked + OR no_provider != excluded.no_provider + OR ambiguous != excluded.ambiguous + OR external != excluded.external", + params![ + workspace, + count(coverage.linked, "http_link_coverage.linked")?, + count(coverage.no_provider, "http_link_coverage.no_provider")?, + count(coverage.ambiguous, "http_link_coverage.ambiguous")?, + count(coverage.external, "http_link_coverage.external")?, + ], + )?; + let mut current = BTreeSet::new(); + for gap in &report.gaps { + let within_bounds = (1..=2048).contains(&gap.caller.as_str().len()) + && (1..=32).contains(&gap.method.len()) + && (1..=4096).contains(&gap.path.len()); + if !within_bounds { + continue; + } + let candidates = gap + .candidates + .iter() + .map(NodeId::as_str) + .collect::>(); + current.insert(( + gap.caller.as_str().to_owned(), + gap.method.clone(), + gap.path.clone(), + gap.reason.as_str().to_owned(), + serde_json::to_string(&candidates)?, + )); + } + let previous = { + let mut statement = transaction.prepare_cached( + "SELECT caller_node_id, method, path, reason, candidates_json + FROM http_link_gaps WHERE workspace_name = ?1", + )?; + statement + .query_map([workspace], |row| { + Ok(( + row.get(0)?, + row.get(1)?, + row.get(2)?, + row.get(3)?, + row.get(4)?, + )) + })? + .collect::, _>>()? + }; + let mut delete = transaction.prepare_cached( + "DELETE FROM http_link_gaps + WHERE workspace_name = ?1 AND caller_node_id = ?2 AND method = ?3 AND path = ?4 + AND reason = ?5", + )?; + for (caller, method, path, reason, _) in previous.difference(¤t) { + delete.execute(params![workspace, caller, method, path, reason])?; + } + let mut insert = transaction.prepare_cached( + "INSERT INTO http_link_gaps( + workspace_name, caller_node_id, method, path, reason, candidates_json + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6)", + )?; + for (caller, method, path, reason, candidates) in current.difference(&previous) { + insert.execute(params![workspace, caller, method, path, reason, candidates])?; + } + Ok(()) +} + +fn replace_freshness( + transaction: &Transaction<'_>, + workspace: &WorkspaceRecord, + coverage_gaps: &[RepositoryCoverageGap], +) -> Result<(), StoreError> { + transaction.execute( + "DELETE FROM repository_snapshot_freshness WHERE workspace_name = ?1", + [&workspace.name], + )?; + let mut insert = transaction.prepare_cached( + "INSERT INTO repository_snapshot_freshness( + workspace_name, repo_id, checkout_id, head_commit, manifest_hash, state, reason + ) VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)", + )?; + for repository in &workspace.repositories { + let mut coverage_reasons = coverage_gaps + .iter() + .filter(|gap| gap.repo_id == repository.id) + .map(|gap| gap.reason.clone()) + .collect::>(); + coverage_reasons.sort(); + coverage_reasons.dedup(); + let (freshness_state, reason) = if !coverage_reasons.is_empty() { + if repository.working_tree_dirty { + coverage_reasons.push( + "The repository working tree differed from HEAD when scanned.".to_owned(), + ); + } + ( + RepoFreshnessState::Partial, + Some(coverage_reasons.join(" ")), + ) + } else if repository.working_tree_dirty { + (RepoFreshnessState::WorkingTreeChanged, None) + } else { + (RepoFreshnessState::Fresh, None) + }; + insert.execute(params![ + workspace.name, + repository.id.as_str(), + repository.checkout_id.as_str(), + repository.head_commit, + workspace.manifest_hash, + serde_json::to_string(&freshness_state)?, + reason, + ])?; + } + transaction.execute( + "DELETE FROM provider_capabilities + WHERE workspace_name = ?1 + AND NOT EXISTS ( + SELECT 1 FROM workspace_repositories + WHERE workspace_name = ?1 + AND repo_id = provider_capabilities.repo_id + )", + [&workspace.name], + )?; + Ok(()) +} + +fn column_text<'row>(row: &'row Row<'_>, index: usize) -> rusqlite::Result<&'row str> { + row.get_ref(index)? + .as_str() + .map_err(|error| rusqlite::Error::FromSqlConversionFailure(index, Type::Text, error.into())) +} + +fn optional_column_text<'row>( + row: &'row Row<'_>, + index: usize, +) -> rusqlite::Result> { + row.get_ref(index)? + .as_str_or_null() + .map_err(|error| rusqlite::Error::FromSqlConversionFailure(index, Type::Text, error.into())) +} + +fn sorted_keys(map: HashMap) -> Vec { + let mut keys = map.into_keys().collect::>(); + keys.sort_unstable(); + keys +} diff --git a/crates/code-system-graph-store-sqlite/src/schema_contract.rs b/crates/code-system-graph-store-sqlite/src/schema_contract.rs index 76b00f7..e84d522 100644 --- a/crates/code-system-graph-store-sqlite/src/schema_contract.rs +++ b/crates/code-system-graph-store-sqlite/src/schema_contract.rs @@ -1,6 +1,6 @@ use rusqlite::Connection; -use super::{INITIAL_SCHEMA, LATEST_SCHEMA_VERSION, StoreError, schema_version}; +use super::{INITIAL_SCHEMA, StoreError, schema_identity, stored_schema_id}; type SchemaContractEntry = (String, String, String, String); @@ -9,8 +9,8 @@ pub(super) fn validate_exact_schema(connection: &Connection) -> Result<(), Store if schema_contract(connection)? != expected { return Err(StoreError::InvalidSchema); } - match schema_version(connection) { - Ok(LATEST_SCHEMA_VERSION) => Ok(()), + match stored_schema_id(connection) { + Ok(schema_id) if schema_id == schema_identity() => Ok(()), Ok(_) | Err(_) => Err(StoreError::InvalidSchema), } } diff --git a/docs/AGENT_SETUP.md b/docs/AGENT_SETUP.md index af2b0b3..abdf0f1 100644 --- a/docs/AGENT_SETUP.md +++ b/docs/AGENT_SETUP.md @@ -44,7 +44,7 @@ csgraph plugin create \ Omit `--codegraph` to keep `explore` unavailable, or add `--codegraph-binary /absolute/path/to/codegraph` with the opt-in flag. Generation requires a valid manifest and an existing matching snapshot. Stale snapshots are allowed and reported. `csgraph -1.1.0` must be in the client's `PATH`, and no platform binary is bundled. Repeating the command is a +1.2.0` must be in the client's `PATH`, and no platform binary is bundled. Repeating the command is a no-op only for an exactly identical directory; conflicts, additional files, and symlinks fail without partial writes. diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 3230727..6a9bd3b 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -1,15 +1,16 @@ # Code System Graph Architecture -Code System Graph builds a federated graph of repository boundaries, contracts, deployments, ownership, -evidence, and immutable snapshots. It complements repository-local symbol and call graphs instead -of duplicating them. +Code System Graph builds a federated graph of repository boundaries, contracts, deployments, +ownership, and evidence, and keeps one current graph per workspace. It complements +repository-local symbol and call graphs instead of duplicating them. ## Design principles - Dependencies point inward toward domain types and application-owned ports. - Extraction is conservative: ambiguous or dynamic observations remain explicit and unlinked. - Durable data is source-free and secret-safe. -- Published snapshots are immutable and become visible atomically. +- Each publication replaces the current graph atomically; readers see the previous or the new + graph, never a mix. - Query results expose freshness, coverage, evidence, ambiguity, and truncation. - External providers are optional, bounded adapters rather than sources of graph truth. @@ -76,7 +77,9 @@ malformed locations, unsafe metadata, and unbounded fields before publication. Tree-sitter parsers and narrow framework recognizers extract literal or conservatively resolvable facts. Source-owned output batches are keyed by repository, checkout, lossless path, and extractor. Their payloads contain structured observations and evidence locators, never source bodies. -Unchanged compatible batches can be reused while affected link neighborhoods are recomputed. +Unchanged compatible batches are reused, and every relationship is relinked from all current +batches, so incremental and full scans publish identical graphs. A scan whose snapshot identity +matches the current snapshot keeps it without loading any batch. Extractors cover package and HTTP contracts, events, GraphQL, protobuf and gRPC, SQL and ORM data, Docker Compose, Kubernetes, Helm, Terraform and OpenTofu, Markdown and ADR references, CODEOWNERS, @@ -95,9 +98,9 @@ Snapshot-scoped records preserve the declaration, reason, resolution, and histor ### Infrastructure adapters -SQLite stores the registry, immutable snapshots, source-free extraction batches, communities, -manual links, provider capabilities, and bounded query summaries. A single blocking writer works -with WAL readers. See `DATA_MODEL.md` and ADR 0003 for storage details. +SQLite stores the registry, the current graph of each workspace, source-free extraction batches, +communities, manual links, provider capabilities, and bounded query summaries. A single blocking +writer works with WAL readers. See `DATA_MODEL.md` and ADR 0003 for storage details. The local Git adapter emits repository state, native paths, changed-file layers, and line positions without retaining diff bodies. GitHub and Bitbucket Cloud inspection is opt-in and uses @@ -113,18 +116,25 @@ returned by the provider remain ephemeral. ## Scan and publication flow 1. Validate the workspace manifest and canonicalize repositories beneath allowed roots. -2. Fingerprint relevant artifacts and plan added, modified, deleted, and reusable inputs. -3. Extract structured facts and evidence into source-owned batches. +2. Fingerprint relevant artifacts once per physical file and plan added, modified, deleted, and + reusable inputs. A file whose size, modification time, and, on Unix, device, inode, and change + time match the operational stat cache reuses its content hash without a read; files modified + within two seconds of the cached observation are always read. +3. Extract structured facts and evidence into source-owned batches. Each changed file is read once + and shared by all of its extractors; files are processed by at most `maxExtractionWorkers` + threads and merged in artifact-key order, so the result never depends on scheduling. Completed + batches are checkpointed in one sidecar transaction. 4. Normalize identities and perform deterministic bilateral linking. 5. Apply exact manual additions and suppressions. -6. Validate the complete candidate graph and persist it in one transaction. -7. Publish the snapshot only after the transaction commits. +6. Validate the complete candidate graph and write only its difference from the current graph in + one transaction. +7. Publish the new snapshot ID only after the transaction commits. -If scanning fails, the last valid snapshot remains queryable. +If scanning fails, the last published graph remains queryable. ## Query and change flow -1. Resolve query anchors against the selected immutable snapshot. +1. Resolve query anchors against the current graph of the selected workspace. 2. Validate freshness and graph consistency. 3. Traverse only within declared direction, depth, result, and time bounds. 4. Optionally request bounded local context from CodeGraph. @@ -142,7 +152,7 @@ false-safe conclusion. Tokio coordinates cancellation and bounded asynchronous work. Repository, extractor, and provider concurrency use explicit semaphores. Shutdown cancels work, terminates child processes, drains -diagnostics, and leaves only fully published snapshots visible. +diagnostics, and leaves only fully published graphs visible. Operational and release details are documented in `INTERFACES.md`, `PERFORMANCE.md`, `INSTALLATION.md`, `RELEASE.md`, and ADR 0009. diff --git a/docs/CHANGES.md b/docs/CHANGES.md index 2a7ed74..ef63114 100644 --- a/docs/CHANGES.md +++ b/docs/CHANGES.md @@ -57,7 +57,7 @@ Rate limits are returned to the caller; Code System Graph does not sleep or retr Provider responses normalize metadata, refs, changed-file counts, CI/check state, reviews, approvals, warnings, and rate-limit metadata. Patch bodies and raw responses are discarded. Structured ETag/TTL caching retains only this source-free normalized result. Bitbucket Data Center -is intentionally not implemented in 1.0.0. +is intentionally not implemented. `csgraph pr list` returns one provider page of source-free summaries for GitHub or Bitbucket Cloud. It uses the same enablement, per-call consent, credential, HTTPS, rate-limit, cancellation, diff --git a/docs/CLI.md b/docs/CLI.md index 74a3be2..3dfa5d9 100644 --- a/docs/CLI.md +++ b/docs/CLI.md @@ -40,6 +40,11 @@ csgraph plugin create --output [--mcp-server-name --routing csgraph plugin uninstall --output --mcp-server-name --routing-skill ``` +`status` reports `http_links`: the counts of consumer and test calls that were linked, that have +no provider, that are ambiguous, and that target an external host, the total number of unlinked +calls in `gap_count`, and up to 50 of them in `gaps` (calls without a provider first, then +ambiguous calls with their candidate providers, then external calls). + Normal scan reuses unchanged source-owned batches. `--repo` recomputes only the selected alias and retains the other repositories' previous batches. `--force` invalidates reuse for the selected scope. Doctor reports unavailable observations as unknown rather than healthy. @@ -52,12 +57,16 @@ alias. `scan` and plain `sync` each perform one pass and exit. Neither command installs a watcher or background service. `sync` is incremental, but first runs `codegraph sync --quiet` with direct -process arguments for every selected repository that already contains `.codegraph/`. Missing or +process arguments for every selected repository that already contains `.codegraph/`, running up +to four repositories concurrently (bounded by `executionPolicy.maxExtractionWorkers`). Missing or failed CodeGraph indexes are explicit per-repository results and do not block the native atomic snapshot. Use `--no-codegraph` to skip that layer. `sync --watch` performs an initial pass and then automatically repeats while its foreground process is running. It emits one JSON object per completed pass, -coalesces event bursts, and observes the manifest plus all declared checkout trees. It uses the +coalesces event bursts, and observes the manifest plus all declared checkout trees. Each pass after +the initial one discovers and synchronizes only the repositories touched by the coalesced events +and reuses every other repository's published batches; manifest edits, ignore-rule changes, event +overflow, and watcher errors widen the pass to the whole workspace. It uses the operating system's native backend on Linux/Unix, macOS, and Windows, automatically falls back to polling if native watcher setup fails, and selects polling for repositories on WSL Windows mounts. Set `--poll-interval-ms ` to force polling for network or virtual filesystems. The same effective @@ -103,8 +112,7 @@ contain the exact `/.local/` rule. The ignored `.local/code-system-graph/` directory contains the runtime binding and an ownership receipt with hashes for managed skill files. Both record the generating package version, exact -executable fingerprint, and build-time source commit/dirty state when available. Receipts and -bindings generated by 1.0.x are rejected. Existing-plugin mode +executable fingerprint, and build-time source commit/dirty state when available. Existing-plugin mode reports `changed: false` for an identical integration. `--replace-generated` atomically replaces only recognized local state and never adopts an unmanaged directory. Updating a changed versioned skill requires `plugin uninstall` followed by `plugin create`. Uninstall requires that receipt and @@ -126,6 +134,9 @@ csgraph contracts list|show|validate|diff|explain-link ... csgraph export --format json|graphml|markdown ... ``` +Query results include `link_gaps`: the unlinked HTTP calls among the returned entities, with the +reason and any candidate providers. + Every operation has server-side bounds. Local Git collection and hosted PR requests propagate SIGINT/SIGTERM cancellation to their child process or HTTP request. diff --git a/docs/CODEGRAPH_INTEGRATION.md b/docs/CODEGRAPH_INTEGRATION.md index 3c94772..fe27990 100644 --- a/docs/CODEGRAPH_INTEGRATION.md +++ b/docs/CODEGRAPH_INTEGRATION.md @@ -175,23 +175,30 @@ All persisted resources and other tools remain source-free. ## Compatibility matrix -The minimum and currently tested production contract is CodeGraph 1.5.0. The structured CLI -adapter accepts the 1.5.x contract family; later versions are incompatible until fixtures and -contract tests are added. MCP capability discovery remains name/schema-driven, but unknown -structured output is never parsed speculatively. +The minimum and currently tested production contract is CodeGraph 1.6.1. The structured CLI +adapter accepts 1.6.1 and later 1.6.x patch releases; earlier versions, pre-releases, and later +minor versions are reported as incompatible until fixtures and contract tests are added. MCP +capability discovery remains name/schema-driven, but unknown structured output is never parsed +speculatively. | Adapter | Version tested | Capabilities | Result | | --- | --- | --- | --- | -| Public MCP | 1.5.0 | `codegraph_explore`; `query`, `maxFiles`, `projectPath`; protocol `2024-11-05` | Passing | -| Public CLI JSON | 1.5.0 | status, query, callers, callees, impact, affected | Passing | -| Public CLI text | 1.5.0 | explore context fallback | Passing, opaque and bounded | -| Process fake | 1.5.0 contract | success, invalid MCP, timeout, cancellation, output cap | Passing | -| Live host smoke | 1.5.0 | initialize, tools/list, status, capability mapping | Passing | +| Public MCP | 1.6.1 | `codegraph_explore`; `query`, `maxFiles`, `projectPath`; protocol `2024-11-05` | Passing | +| Public CLI JSON | 1.6.1 | status, query, callers, callees, impact, affected | Passing | +| Public CLI text | 1.6.1 | explore context fallback, file context | Passing, opaque and bounded | +| Process fake | 1.6.1 contract | success, invalid MCP, timeout, cancellation, output cap | Passing | +| Live host smoke | 1.6.1 | initialize, tools/list, status, capability mapping, query, callers, impact, affected | Passing | -The versioned fixtures are under `fixtures/codegraph/1.5.0/`; degradation fixtures cover missing -indexes, stale indexes, missing optional tools, invalid MCP responses, and oversized output. +The versioned fixtures under `fixtures/codegraph/1.6.1/` are captured from CodeGraph 1.6.1 and +parsed by the contract tests; degradation fixtures cover missing indexes, stale indexes, missing +optional tools, invalid MCP responses, and oversized output. The live smoke test is ignored by +default and runs against an explicitly indexed checkout of this repository: -Operation selection for 1.5.0 is: +```bash +cargo test -p code-system-graph-core --test codegraph_provider_e2e -- --ignored +``` + +Operation selection for 1.6.1 is: - local context: MCP `codegraph_explore`, then CLI `explore` fallback; - symbol resolution: CLI `query --json`; @@ -206,7 +213,8 @@ tool are separate states. Federated boundaries remain queryable, but local detai coverage warnings prevent false-safe conclusions. `sync` does not start a CodeGraph daemon and does not reuse CodeGraph's private lock or database -format. Its own watcher observes source trees and invokes one-shot public syncs serially. The +format. Its own watcher observes source trees and invokes one-shot public syncs for up to four +repositories at a time (bounded by `executionPolicy.maxExtractionWorkers`). The short-lived MCP child still starts with `--no-watch`, avoiding duplicate hidden watchers per scan. CodeGraph database/WAL changes under `.codegraph/` are excluded from csgraph watch events. diff --git a/docs/COMMUNITIES.md b/docs/COMMUNITIES.md index 17218d4..6446899 100644 --- a/docs/COMMUNITIES.md +++ b/docs/COMMUNITIES.md @@ -63,7 +63,7 @@ Each community reports: - incomplete inputs and other limitations. Labels use only bounded structural terms from repository, service, contract, and central-node -labels. LLM-generated names are outside the factual 1.0.0 engine. +labels. LLM-generated names are outside the factual engine. ## Snapshot deltas diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index ffdb0d4..ced18d7 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -58,6 +58,9 @@ repos: includeDefaults: - vendor/internal-sdk/** useGitignore: true + authorities: + - orders-api:8080 + - orders.internal httpConsumers: - method: POST path: /orders @@ -73,10 +76,30 @@ repos: | `excludes` | Additional repository-relative paths must be omitted from automatic discovery | | `includeDefaults` | A specific path inside a default dependency or build exclusion must be discovered | | `useGitignore` | Repository-contained `.gitignore` rules should filter automatic discovery | +| `authorities` | Other repositories call this one through absolute URLs whose host is not a Compose or Kubernetes service | These fields add explicit evidence. They are not required for supported, unambiguous source patterns. +## HTTP authorities + +An absolute URL such as `http://orders-api:8080/v1/orders/42` is linked by its path, and its +`host:port` authority decides which repositories may provide it: + +- `localhost`, `127.0.0.0/8`, `::1`, `0.0.0.0`, and `host.docker.internal` resolve across the + workspace with the usual scope rules, so a test that calls a locally started service links to + the repository that serves the route. +- An authority listed in a repository's `authorities` restricts resolution to that repository. An + entry without a port matches every port of the host. Entries are `host` or `host:port` values + without scheme or path, are compared case-insensitively, and must be unique across repositories. +- Compose services, Kubernetes deployments and Services, their declared service names, and host + aliases name the repository that declares them. Cluster DNS names such as + `orders.shop.svc.cluster.local` resolve to the service name. When the declaring repository has no + matching route, resolution falls back to the workspace; a name declared by several repositories + is not used. +- Any other host is external: the call is kept as a consumer boundary, is not linked, and is + counted as external in the HTTP link report rather than as a call without a provider. + ## Discovery exclusions `excludes` and `includeDefaults` accept repository-relative globs with `*`, `?`, and `**`. The @@ -307,6 +330,7 @@ executionPolicy: maxWatchSessionWallTimeMs: 86400000 minWatchRescanIntervalMs: 10000 maxCheckpointCacheBytes: 10737418240 + maxExtractionWorkers: 8 maxExploreWallTimeMs: 8000 maxExploreCodeGraphOperations: 8 @@ -331,10 +355,14 @@ executionPolicy: ``` The defaults allow six hours and 16 GiB per worker, five minutes without verified completed work, -eight idle hours and 24 total hours per watcher session, and 10 GiB of historical checkpoint cache. -The cache is not preallocated. Completed batches and a fully validated candidate graph may be -resumed from the owner-only operational sidecar; that candidate remains invisible to every query -surface until one atomic publication transaction succeeds. +eight idle hours and 24 total hours per watcher session, and 10 GiB of checkpoint cache. The cache +is not preallocated. Completed extractor batches are checkpointed in the owner-only operational +sidecar and reused by the next pass; the graph remains invisible to every query surface until one +atomic publication transaction succeeds. + +`maxExtractionWorkers` bounds concurrent file extraction. The effective value never exceeds the +host's available parallelism, and the published graph is identical for every value: results are +merged in artifact-key order, never in completion order. It accepts `1` through `256`. All values must be positive and representable except `maxCodeGraphCorroborationAnchorsPerRepo`, which also accepts `-1` for unlimited. Its default is diff --git a/docs/DATA_MODEL.md b/docs/DATA_MODEL.md index 5b592e7..9d3f462 100644 --- a/docs/DATA_MODEL.md +++ b/docs/DATA_MODEL.md @@ -1,8 +1,7 @@ # Federated Data Model -Code System Graph stores normalized repository-boundary facts and their evidence. The initial 1.0 -database schema is the single definitive schema. Migration support will be designed only after a -released schema exists. +Code System Graph stores normalized repository-boundary facts and their evidence. Each binary +embeds one exact database schema; a database is accepted only when it matches that schema. ## Identity @@ -33,8 +32,8 @@ Repository paths are canonicalized before registration. The manifest directory i allowed root; paths outside it require an explicit `allowedRoots` entry. Canonical checkout and Git common-directory paths must remain beneath an allowed root. -Workspace removal cascades snapshots and alias mappings, then garbage-collects repositories and -checkouts no longer referenced by another workspace. +Workspace removal cascades the current graph and alias mappings, then garbage-collects repositories +and checkouts no longer referenced by another workspace. Manifest updates are constrained to the top-level `repos` mapping. Preview validates the complete resulting manifest and registry without writing. Commit verifies the original content fingerprint, @@ -43,24 +42,30 @@ comments, ordering, line endings, and unrelated bytes. ## Snapshots, graph, and evidence -`repo_snapshots` is published atomically per workspace. Nodes, edges, evidence, edge-evidence links, -community data, manual-link records, and per-checkout freshness are written in the publication -transaction. If any constraint or write fails, the previous current snapshot remains visible. +Each workspace owns exactly one current graph. `repo_snapshots` holds one row per workspace whose +snapshot ID identifies the published inputs. Publication compares the candidate with the stored +graph by identifier and content digest, then inserts, updates, and deletes only the differing nodes, +edges, evidence, edge-evidence links, fingerprints, and extractor batches in one transaction. +Manual-link records, extractor runs, and per-checkout freshness are rewritten in the same +transaction. If any constraint or write fails, the previous graph remains visible. An unchanged +republication writes no graph rows, so database size stays proportional to the current graph. -Nodes and edges are snapshot-scoped. Foreign keys reject dangling graph edges and missing evidence. -FTS5 indexes bounded labels and stable keys without source bodies. Search is scoped to the current -workspace snapshot and accepts quoted, bounded input rather than raw FTS syntax. +Graph tables are keyed by workspace. Foreign keys reject dangling graph edges and missing evidence. +FTS5 indexes bounded labels and stable keys without source bodies and is maintained by triggers on +node rows. Search is scoped to the workspace and accepts quoted, bounded input rather than raw FTS +syntax. Reads addressed by snapshot ID resolve that ID to its workspace and fail when the snapshot is +no longer current. Evidence records retain source ownership, hashes, line locators, extraction metadata, and bounded notes. Source bodies, diff bodies, credentials, configuration values, connection strings, and external-provider responses are not graph evidence. -Integrity checks combine SQLite `quick_check`, exact schema-version validation, and +Integrity checks combine SQLite `quick_check`, exact schema-identity validation, and `foreign_key_check`. ## Extraction artifacts -`artifact_fingerprints` stores bounded BLAKE3 content hashes keyed by snapshot, repository, +`artifact_fingerprints` stores bounded BLAKE3 content hashes keyed by workspace, repository, checkout, lossless relative path, and extractor. Scan planning compares current and previous keys: - current only: added; @@ -69,17 +74,18 @@ checkout, lossless relative path, and extractor. Scan planning compares current - same identity and hash: unchanged. A rename is represented as delete plus add unless an extractor can establish semantic continuity. -`extractor_runs` records versioned counts and status, while `extractor_run_inputs` binds each run to -exact content hashes. If merged configuration and all fingerprints are unchanged, scanning reuses -the published snapshot without running extractors or writing SQLite. +`extractor_runs` records versioned counts and status per repository, checkout, and extractor; the +fingerprints of the same scope are the exact run inputs. If merged configuration and all +fingerprints are unchanged, scanning reuses the published snapshot without running extractors or +writing SQLite. Repository-local `.code-system-graph.yaml` configuration participates in the workspace fingerprint, so a configuration change invalidates freshness even when contract artifacts are byte-identical. -Versioned output batches are keyed by snapshot, repository, checkout, lossless source path, and +Versioned output batches are keyed by workspace, repository, checkout, lossless source path, and extractor. Payloads contain structured observations only, and output counts are checked while -decoding. Compatible unchanged batches are reused; add, replace, and delete actions identify the -link neighborhoods to recompute. +decoding. Each batch stores a BLAKE3 payload hash, so publication never decodes unchanged payloads. Compatible unchanged batches are reused, and relationships are relinked from all current +batches on every scan. ## Federated entities @@ -111,18 +117,55 @@ test_case --validates--> http_operation --implemented_by--> symbol_ref ``` `test_case` identities include repository, language, framework, source path, and test name. -`symbol_ref` identities include repository, language, source path, and declared symbol. +`symbol_ref` identities include repository, language, source path, and declared symbol. A test +callback without a function name, such as a Jest `it` block, declares the symbol formed by its +`describe` chain and title joined with ` > `. A `validates` edge requires test evidence and provider-contract evidence. An `implemented_by` edge -requires provider-contract evidence plus implementation declaration or source evidence. Missing -providers remain unlinked, and duplicate providers remain ambiguous. Optional CodeGraph +requires provider-contract evidence plus implementation declaration or source evidence. + +HTTP operations are identified by method and canonical route shape. Parameter syntax and names are +not part of the identity: `{id}`, `:id`, ``, `{id:int}`, `{id:[0-9]+}`, and `[id]` are one +parameter segment, and `{*rest}`, `{path...}`, ``, `[...slug]`, and `*rest` are one +catch-all segment. `calls_remote`, `validates`, and `implemented_by` resolve through one per-method +route index: + +- provider routes carry their full path: extractors record the router each route is registered on + and every router mount with its prefix, and mounts are composed per repository across files + before linking; +- consumer calls are composed per repository before linking: a call whose URL depends on a + parameter of its function is instantiated at each call site that binds it, and a test owns the + exact calls of the functions and fixtures it reaches, with evidence at the call site in the test; + a call made through a parameter is kept only when the parameter resolves to an in-process test + client through a fixture or every caller; +- a concrete path such as `/orders/42` matches the `/orders/{id}` template; +- the most specific template wins, comparing segments left to right (static, then parameter, then + catch-all); +- equally specific providers are narrowed by scope: an explicit repository restriction (an + authority declared in the manifest or inferred from Compose and Kubernetes services, or an + in-process test client), then the caller's own repository, then a unique provider in the + workspace; +- loopback hosts resolve across the workspace, and calls to other unknown hosts are external and + never linked; +- remaining ties are reported as ambiguities with their candidates, and implementation anchors + resolve only inside their own repository. + +Missing providers remain unlinked. Every scan publishes an HTTP link report with the workspace +graph: the number of consumer and test calls (counted per calling node) that were linked, that have +no provider, that are ambiguous, and that target an external host, plus one gap row per unlinked +call with its method, path, reason, and, for ambiguities, the candidate providers. The report is +replaced with the graph and read by `status`, `query`, and the MCP coverage resource. + +Optional CodeGraph corroboration can strengthen an existing implementation edge after exact path, name, and line agreement; it does not create unsupported federated relationships. ## Communities and derived views -Each graph snapshot may own one immutable `CommunitySnapshot` plus normalized communities and -memberships. It records engine version, algorithm, seed, scope, resolution, confidence threshold, +Each graph snapshot may own one `CommunitySnapshot` plus normalized communities and memberships. +The analyses of the current and the previous snapshot are retained so deltas can be compared; older +analyses are deleted on publication. Communities are recomputed only when graph topology changes; +otherwise the previous analysis is carried forward under the new snapshot ID. It records engine version, algorithm, seed, scope, resolution, confidence threshold, relationship weights, and iteration bound. Community rows retain deterministic labels, central nodes, metrics, contract boundaries, label evidence, and limitations. @@ -148,8 +191,8 @@ metadata, and one deterministic fingerprint. Tokens, patches, and raw responses ## Manual links and bounded runtime data `manual_links` stores a declaration ID, exact resolved source and target node IDs, edge kind, -addition or suppression disposition, required reason, versioned link decision, manifest -configuration version, and owning immutable snapshot. +addition or suppression disposition, required reason, versioned link decision, and manifest +configuration version. Each manifest `manualLinks` endpoint must exactly match a node ID or stable key in the candidate snapshot and resolve to one node. No match is unresolved; multiple matches are ambiguous. An @@ -158,14 +201,13 @@ optional `contract` value supplies audit context and does not change endpoint ma An addition replaces an automatic edge with the same source, target, and edge kind and creates one confirmed edge backed only by manual provenance. A suppression removes automatic edges matching that exact triple and fails when none exists. Extractor observations and unrelated edges remain -intact. Updating or removing a declaration changes current state while preserving prior immutable -snapshot history. +intact. Updating or removing a declaration changes the current graph on the next publication. `provider_capabilities` stores normalized capability names for one workspace, repository, provider name, and provider version, plus the observation timestamp. It does not store provider responses. -`query_cache` stores bounded JSON result summaries keyed by workspace, immutable snapshot, and -normalized input fingerprint. Each summary is capped at 1 MiB, and each workspace retains at most +`query_cache` stores bounded JSON result summaries keyed by workspace, current snapshot, and +normalized input fingerprint. Publication clears the workspace cache. Each summary is capped at 1 MiB, and each workspace retains at most 1,024 entries. Administrative MCP mutations append a separate source-free JSONL audit record containing version, @@ -186,11 +228,14 @@ Unknown, unavailable, corrupt, or partial inputs never produce a fresh or safe c ## Schema, locking, and recovery -The definitive 1.1.0 schema version (`2`) and an opaque database instance identity are recorded in -`schema_metadata`. The complete SQLite schema must match the 1.1.0 definition exactly. Version 1 -databases from the 1.0.x line and any other development schema are rejected with an instruction to remove the disposable database -and run a full scan. Operational checkpoints are bound to the instance identity, so a sidecar from -another or rebuilt database is discarded rather than reused. +`schema_metadata` records the schema identity, a BLAKE3 digest of the embedded schema definition, +and an opaque database instance identity. The complete SQLite schema and the recorded identity must +match the binary exactly; any other database is rejected with an instruction to remove the +disposable database and run a full scan. The owner-only `.work.db` sidecar holds +operational state only: checkpointed extractor batches (metadata plus the raw payload BLOB, evicted +least-recently-used under `maxCheckpointCacheBytes`), the per-file stat cache used to skip reading +unchanged files, and watcher leases. It is bound to the database instance identity, so a sidecar +from another or rebuilt database is discarded rather than reused. Restore validates integrity and exact schema compatibility, preserves the destination as a separate safety backup, and restores through SQLite's backup API without migration. diff --git a/docs/EXTRACTOR_COVERAGE.md b/docs/EXTRACTOR_COVERAGE.md index 7858edd..52d9bc5 100644 --- a/docs/EXTRACTOR_COVERAGE.md +++ b/docs/EXTRACTOR_COVERAGE.md @@ -128,7 +128,10 @@ schemas outside the extracted file set, or infer application handlers from gener and truncates oversized column/index detail without aborting the workspace scan. - Factory Boy `Meta.model` links from a factory symbol to its statically imported Python model. - Literal SELECT/WITH/INSERT/UPDATE/DELETE observations across supported source languages with - reader/writer roles and enclosing-symbol anchors. + reader/writer roles and enclosing-symbol anchors. JavaScript under `theme`, `themes`, `vendor`, + `vendors`, `third_party`, `third-party`, or `bower_components` directories, files named + `*.min.js`, `*-min.js`, `*.bundle.js`, or `*.chunk.js`, and minified content (at least 4 KiB with + lines averaging 500 bytes or more) are not scanned for SQL. - Exact linking to one unambiguous normalized table identity shared across languages and repositories. @@ -142,6 +145,8 @@ literal values are discarded before persistence. - Docker Compose services, images, ports, dependencies, and environment key names. - Kubernetes workloads, Services, Ingresses, static resources, selectors, and Secret key names. - Helm templates and values with static names retained and templated values marked incomplete. + A template is a `.yaml`, `.yml`, or `.tpl` file under a `templates` directory whose parent holds + `Chart.yaml`; Jinja `*.j2` files and other `templates` directories are not Helm. - Terraform/OpenTofu literal resource declarations and dependency metadata. - Repository `deploys`, deployment `provides`, resource dependency, and config-key links from direct declarations. @@ -165,12 +170,15 @@ credentials, connection strings, and secret material are never stored. ## Incremental and linking behavior Extractor outputs are stored per repository, checkout, lossless source path, extractor, content -hash, and extractor version. Identical batches are reused. Adds, replacements, and deletions -recompute only affected method/path neighborhoods; a missing previous deletion batch fails closed. -`--force` never reuses a planned replacement, and internal extractor revisions invalidate cached -outputs even when source content is unchanged. -Duplicate providers remain ambiguous. Consumers, tests, provider contracts, and implementation -symbols require direct evidence before edges are created. +hash, and extractor version. Identical batches are reused. Relationships are relinked from the +complete set of current batches on every scan, so an incremental scan publishes exactly the graph a +full scan would. `--force` never reuses a planned replacement, and internal extractor revisions +invalidate cached outputs even when source content is unchanged. A source file whose event, +GraphQL, generated-protobuf, or literal-SQL scan finds no facts keeps a persisted batch without +outputs and adds no artifact node to the graph; declared AsyncAPI, GraphQL, protobuf, and data artifacts always +appear. +Equally specific providers that no scope rule separates remain ambiguous. Consumers, tests, provider +contracts, and implementation symbols require direct evidence before edges are created. Optional scan-time CodeGraph corroboration uses only public MCP/CLI operations. Exact symbol path/name/line matches add provenance to an existing source-derived implementation edge; they diff --git a/docs/INSTALLATION.md b/docs/INSTALLATION.md index 24acbea..067ad76 100644 --- a/docs/INSTALLATION.md +++ b/docs/INSTALLATION.md @@ -5,7 +5,7 @@ for maintainers are in [Release engineering](RELEASE.md). ## Current availability -Code System Graph `1.1.0` is published on [crates.io](https://crates.io/crates/code-system-graph) +Code System Graph `1.2.0` is published on [crates.io](https://crates.io/crates/code-system-graph) and [GitHub Releases](https://github.com/dertin/code-system-graph/releases). Native release CI validates Linux x86_64/ARM64, macOS x86_64/ARM64, and Windows x86_64 before their archives are published. @@ -49,7 +49,7 @@ Do not substitute an unofficial download URL or third-party binary mirror. Requirements: -- the latest stable Rust toolchain; the minimum supported Rust version (MSRV) is 1.97.1; +- the latest stable Rust toolchain; the minimum supported Rust version (MSRV) is 1.96.0; - Cargo. ```bash @@ -62,7 +62,7 @@ This builds and installs the same binaries as the binstall path. Verify with the Requirements: -- the latest stable Rust toolchain; MSRV is 1.97.1; +- the latest stable Rust toolchain; MSRV is 1.96.0; - Cargo and Git; - a trusted checkout of this repository. @@ -84,8 +84,8 @@ For a Linux x86_64 archive: ```bash sha256sum --ignore-missing --check SHA256SUMS -tar -xzf code-system-graph-x86_64-unknown-linux-gnu-v1.1.0.tgz -PREFIX="$HOME/.local" ./code-system-graph-x86_64-unknown-linux-gnu-v1.1.0/install.sh +tar -xzf code-system-graph-x86_64-unknown-linux-gnu-v1.2.0.tgz +PREFIX="$HOME/.local" ./code-system-graph-x86_64-unknown-linux-gnu-v1.2.0/install.sh ``` Replace the target in the archive name with `x86_64-apple-darwin`, @@ -93,7 +93,7 @@ Replace the target in the archive name with `x86_64-apple-darwin`, same installer. The Windows archive is a ZIP file. Verify `SHA256SUMS`, extract -`code-system-graph-x86_64-pc-windows-msvc-v1.1.0.zip`, and add its `bin` directory containing +`code-system-graph-x86_64-pc-windows-msvc-v1.2.0.zip`, and add its `bin` directory containing `csgraph.exe` and `code-system-graph-hooks.exe` to `PATH`. `PREFIX` defaults to `$HOME/.local`. The installer places binaries under `$PREFIX/bin`, installed @@ -153,13 +153,10 @@ The package installer preserves replaced binaries under: $PREFIX/share/code-system-graph/backups/ ``` -Version 1.1.0 deliberately rejects 1.0.x databases, query caches, checkpoints, generated plugin -bindings, and delivery contracts. Archive or remove the old `.code-system-graph` database and -operational sidecars, regenerate any Agent Plugin/binding, then run a complete scan. Backup and -restore accept only the exact 1.1.0 schema and never migrate older state. Direct MCP registrations -must add `--config`; MCP consumers receive one human-oriented Markdown text block plus typed -`structuredContent` as its complete machine-readable mirror, while HTTP/CLI consumers must accept -delivery schema v2. +A database is accepted only when its schema exactly matches the running binary. When `csgraph` +reports an incompatible database, remove the `.code-system-graph` database and its operational +sidecars, regenerate any Agent Plugin binding with `plugin create --replace-generated`, and run a +complete scan. ## Uninstall @@ -175,7 +172,7 @@ cargo uninstall code-system-graph-hooks Run `uninstall.sh` from the verified extracted package with the same prefix: ```bash -PREFIX="$HOME/.local" ./code-system-graph-x86_64-unknown-linux-gnu-v1.1.0/uninstall.sh +PREFIX="$HOME/.local" ./code-system-graph-x86_64-unknown-linux-gnu-v1.2.0/uninstall.sh ``` Before uninstalling either installation type, remove any optional agent hooks: diff --git a/docs/INTERFACES.md b/docs/INTERFACES.md index 06de4c9..be8dfe4 100644 --- a/docs/INTERFACES.md +++ b/docs/INTERFACES.md @@ -92,5 +92,5 @@ alias and exact staged state; advisory mode fails open. secrets, provider patches, and raw provider responses are never returned. - GraphML and Markdown exports are deterministic views, not import or round-trip formats. - Doctor reports unknown when an observation is absent; absence never becomes healthy. -- HTTP does not expose administrative, pull-request, or source-context routes in 1.1.0. +- HTTP does not expose administrative, pull-request, or source-context routes. - Hook guidance cannot guarantee that a host follows the suggested provider routing. diff --git a/docs/MCP.md b/docs/MCP.md index 3f22285..27091af 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -135,7 +135,10 @@ read only after substituting a concrete stable ID. Every resource uses `text/mar catalog embeds each generated JSON Schema in a closed fenced `json` block. Resource item and byte limits come from the immutable workspace policy. Every resource collection reports its `total`, retained count, and truncation state in Markdown; status, coverage, and freshness use -collection-specific names for the same metadata. +collection-specific names for the same metadata. The coverage resource also reports HTTP link +coverage (linked, without a provider, ambiguous, and external calls) and the bounded list of +unlinked calls, and query results list the unlinked HTTP calls among their entities under +"Unlinked HTTP calls". ## Administrative profile diff --git a/docs/PERFORMANCE.md b/docs/PERFORMANCE.md index b1e8fb2..36aff33 100644 --- a/docs/PERFORMANCE.md +++ b/docs/PERFORMANCE.md @@ -2,8 +2,9 @@ ## Scope -Code System Graph has explicit release-mode acceptance workloads for graph analytics and 100-repository -incremental scanning. The measurements below were collected locally on Linux x86_64. They are +Code System Graph has explicit release-mode acceptance workloads for graph analytics, 100-repository +incremental scanning, and 200-repository synchronization. The measurements below were collected +locally on Linux x86_64. They are release evidence, not service-level objectives or guarantees for other hardware, operating systems, repository contents, or graph topologies. @@ -27,11 +28,11 @@ The measured operations are: Measured Linux x86_64 result: -- query p95: 48 ms, below the 500 ms gate; +- query p95: 43 ms, below the 500 ms gate; - trace p95: 0 ms, below the 1 second gate; -- summary impact p95: 824 ms, below the 2 second gate; -- connected-components analysis: 337 ms; -- resident memory: 497,444 KiB. +- summary impact p95: 704 ms, below the 2 second gate; +- connected-components analysis: 313 ms; +- resident memory: 401,108 KiB. With five samples, p95 selects the slowest observed sample. The zero-millisecond trace result means the elapsed duration rounded down when represented as whole milliseconds; it does not mean the @@ -49,11 +50,11 @@ cargo test -p code-system-graph --test scale_registry_e2e --release -- --ignored After registry discovery optimization, the measured Linux x86_64 result was: - initially discovered inputs: 600; -- targeted incremental scan: 165 ms; +- targeted incremental scan: 145 ms; - changed inputs: 5; - unchanged inputs retained: 595; - unchanged repositories retained: at least 99; -- status p95 over ten samples: 11 ms, below the 200 ms gate. +- status p95 over ten samples: 33 ms, below the 200 ms gate. The changed-input count includes the source file and the source-owned extractor observations affected by that repository change; it is not a count of changed repositories. @@ -68,21 +69,97 @@ representative extractor invocation per file, then performs a complete persisted cargo test -p code-system-graph --test extraction_scale_e2e --release --locked -- --ignored --nocapture ``` -Measured Linux x86_64 result on August 1, 2026: +Measured Linux x86_64 result on October 2, 2026: - files: 5,000; - discovered artifact-extractor invocations: 9,000, because JavaScript files intentionally run through multiple focused extractors; -- complete scan: 47,029 ms; -- representative per-artifact extraction p50: 10 microseconds; -- representative per-artifact extraction p95: 33 microseconds; -- representative per-artifact extraction p99: 45 microseconds; -- test-process peak resident memory from Linux `VmHWM`: 89,060 KiB. +- complete scan: 1,190 ms; +- representative per-artifact extraction p50: 3 microseconds; +- representative per-artifact extraction p95: 10 microseconds; +- representative per-artifact extraction p99: 14 microseconds; +- test-process peak resident memory from Linux `VmHWM`: 23,820 KiB. The memory value comes from the release test process after the scan. It excludes Cargo and compiler processes. The per-artifact samples measure parser extraction only; the complete-scan figure also includes discovery, fingerprinting, linking, community analysis, and SQLite publication. +## Synchronization workload + +The synchronization acceptance test generates 200 repositories with 250 files each (50,000 files): +an Express server, a `fetch` client of the next repository, a Python `requests` test, and +TypeScript, Python, and Go modules. Every file is written with a modification time older than the +stat-cache racy window. The test runs a cold scan, an unchanged scan, a scan after adding one +route to one server file, and 20 further syncs that alternate that file between two contents: + +```text +cargo test -p code-system-graph --test sync_scale_e2e --release -- --ignored --nocapture +``` + +Set `CODE_SYSTEM_GRAPH_SCALE_WORKLOAD` to a directory to generate the workload there instead of a +temporary directory, so another build can scan the same files. The test enforces these gates: + +- the unchanged scan reads 0 content bytes and reuses the published snapshot; +- the one-file sync publishes no more rows than the changed repository segment owns: its nodes, + every edge touching them, its evidence, fingerprints, extractor batches, and extractor runs; +- the database (main file plus WAL) after the 20 syncs stays within 5% of its size after the + one-file sync. + +Measured Linux x86_64 result on October 2, 2026: + +- cold scan: 8,204 ms, 15,481,200 content bytes read, peak worker resident memory 695,578,624 + bytes; +- unchanged scan: 680 ms, 0 content bytes read, 50,000 stat-cache hits; +- one-file sync: 2,034 ms, 634 content bytes read, 2,022 published rows against a 2,532-row + segment; +- database: 299,958,272 bytes after the one-file sync and 299,970,560 bytes after 20 more syncs. + +An unchanged scan compares the snapshot identity, a hash of the manifest, the extraction contract, +the budgets, and every artifact fingerprint, with the current snapshot, so it loads neither the +previous fingerprints nor any extractor batch. An incremental scan publishes artifact rows as a +delta planned against the stored fingerprints instead of comparing every stored row. + +`ExecutionSummary` reports `contentBytesRead` and `publishedRows` for every scan, so the same +figures are available from `csgraph scan` output outside the test. + +### Comparison with 1.1.0 + +Both releases scanned subsets of the same generated workload with the release `csgraph scan` +binary under `/usr/bin/time`. The table lists wall time and the maximum resident set size of the +process tree. Each subset keeps 250 files per repository. + +| Repositories | 1.1.0 cold | 1.2.0 cold | 1.1.0 unchanged | 1.2.0 unchanged | 1.1.0 peak RSS | 1.2.0 peak RSS | +| ---: | ---: | ---: | ---: | ---: | ---: | ---: | +| 10 | 48.6 s | 0.34 s | 12.4 s | 0.18 s | 98,184 KiB | 60,560 KiB | +| 20 | 188.7 s | 0.58 s | 40.5 s | 0.18 s | 171,212 KiB | 96,440 KiB | +| 40 | 742.4 s | 1.24 s | 150.2 s | 0.29 s | 316,880 KiB | 165,992 KiB | +| 200 | not run | 8.16 s | not run | 0.74 s | not run | 709,640 KiB | + +1.1.0 cold-scan time grows quadratically with the workload, so the 200-repository run was not +attempted. At 40 repositories the 1.1.0 database occupies 275,894,272 bytes and the 1.2.0 database +58,691,584 bytes. The peak resident memory of a 1.2.0 unchanged scan is 79,696 KiB at 40 +repositories and 309,396 KiB at 200. + +### Manual validation on a product workspace + +A private eight-repository product workspace (FastAPI backend with pytest suites, TypeScript web +client, Rust services, deployment, infrastructure, and documentation) was scanned by both releases +into a temporary database: + +- cold scan: 3.33 s with 1.1.0 and 0.35 s with 1.2.0; the 1.2.0 unchanged scan takes 0.24 s, + reads 0 content bytes, and peaks at 24,068 KiB; +- graph: 1,664 nodes and 1,618 edges with 1.1.0, 851 nodes and 836 edges with 1.2.0, because + source files without facts no longer add per-file artifact nodes; +- HTTP relations: `validates` grows from 0 to 36 and `calls_remote` from 5 to 17, while + `implemented_by` stays at 88; +- link report: 88 linked calls, 10 calls without a provider, all from Rust integration tests posting + to a bare `hyper` service that declares no routes, and 1 external call to a third-party host; +- database: 13,172,736 bytes with 1.1.0 and 6,381,568 bytes with 1.2.0; +- cold-scan peak resident memory: 43,340 KiB with 1.1.0 and 53,072 KiB with 1.2.0 at the default + eight extraction workers, or 43,400 KiB with `maxExtractionWorkers: 1` at 0.59 s. On workspaces + this small the per-thread allocator arenas outweigh the per-artifact savings; capping glibc + arenas saved about 5 MB here while slowing the 40-repository workload, so the default stays. + ## Interpretation and reproducibility - Always use `--release`; scale acceptance tests reject debug builds. diff --git a/docs/RELEASE.md b/docs/RELEASE.md index a0dd3cb..e256e90 100644 --- a/docs/RELEASE.md +++ b/docs/RELEASE.md @@ -1,10 +1,11 @@ -# Code System Graph 1.1.0 Release - -Code System Graph 1.1.0 is an intentionally breaking agent-delivery release. MCP tools pair -bounded Markdown with typed agent delivery schema v5, resources are bounded Markdown, typed -HTTP/CLI envelopes use delivery schema v2, Explore performs -multi-step ephemeral source/flow enrichment under one global policy, and Query supplies grounded -next actions. The source tree and package version are `1.1.0`. Continuous integration +# Code System Graph 1.2.0 Release + +Code System Graph 1.2.0 is a synchronization and linking release. Each workspace keeps one current +graph that every scan updates by delta publication; unchanged files are recognized from a per-file +stat cache without being read, changed files are read once for all of their extractors, and +extraction runs in parallel with a deterministic merge. One route engine links consumers, tests, +and implementations across languages by canonical HTTP method and path, and every scan reports the +calls it could not link. The source tree and package version are `1.2.0`. Continuous integration runs on GitHub at `https://github.com/dertin/code-system-graph`. Platform claims below require native build, test, packaging, and archive-smoke evidence from the @@ -26,12 +27,19 @@ target. ## Included capabilities -Code System Graph 1.1.0 includes: +Code System Graph 1.2.0 includes: - multi-repository workspace registration with lossless native-path identity; - incremental, source-free extraction for package, HTTP, event, GraphQL, RPC, data, infrastructure, documentation, ownership, configuration, test, and implementation boundaries; -- deterministic linking, exact manual links and suppressions, immutable snapshots, and freshness; +- one current graph per workspace published as a delta, a per-file stat cache, one read per changed + file, and parallel extraction whose output is identical for every worker count; +- cross-language HTTP linking by canonical route shape with concrete-path matching, router + prefixes composed across files, evaluated client URLs, client wrappers, test helpers and + fixtures, in-process test clients, and per-repository `authorities`; +- an HTTP link report of linked, provider-less, ambiguous, and external calls in `status`, query + results, and MCP coverage; +- deterministic linking, exact manual links and suppressions, and freshness; - bounded search, trace, community analysis, compatibility, impact, and change analysis; - opt-in CodeGraph integration through public MCP or CLI contracts; - configurable per-repository corroboration bounds and opt-in Git-native ignore discovery; @@ -47,26 +55,10 @@ Code System Graph 1.1.0 includes: - exact-schema SQLite backup, restore, and integrity validation; - deterministic Linux package archives, CycloneDX SBOMs, and SHA-256 checksums. -## Breaking migration from 1.0.x - -- Databases created by 1.0.x are rejected. Remove or archive the old database and operational - sidecars, then perform a complete 1.1.0 scan; there is no in-place migration or legacy mode. -- Regenerate Agent Plugins and local bindings. Binding and ownership contracts are version 2 and - older generated state is not accepted as a runtime binding. -- Direct `csgraph mcp` invocations must add `--config `; `--binding` mode resolves - the manifest recorded in the binding. -- MCP consumers receive one Markdown text block plus typed `structuredContent`. They must not - expect a JSON text block, the CLI/HTTP envelope inside `content`, or result `outputSchema` - metadata. -- HTTP and CLI consumers must accept `schema_version: 2`; Explore data is now `ExploreReport`, with - source in `source_markdown` rather than `LocalContextResult.content`. -- Workspace manifests may omit all new policy fields to use the documented defaults. Inputs such - as Explore `max_files` can request less work but cannot exceed the effective global policy. - ## Local validation Run the main Rust validation from a clean checkout. This requires stable, nightly with the rustfmt -and Clippy components, and Rust 1.97.1: +and Clippy components, and Rust 1.96.0: ```text cargo +nightly fmt --all -- --check @@ -74,7 +66,7 @@ cargo +nightly clippy --workspace --exclude code-system-graph-fuzz --all-targets cargo +stable check --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked cargo +stable test --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked RUSTDOCFLAGS="-D warnings" cargo +stable doc --workspace --exclude code-system-graph-fuzz --all-features --no-deps --locked -cargo +1.97.1 check --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked +cargo +1.96.0 check --workspace --exclude code-system-graph-fuzz --all-targets --all-features --locked ``` Dependency and source-policy tools can then run against the locked workspace: @@ -105,12 +97,12 @@ See `PERFORMANCE.md` for workload definitions, measurements, and interpretation. ## Package validation Creating the Linux x86_64 artifacts requires the latest stable Rust toolchain, the target, Syft, -GNU tar, and SHA-256 tooling. The workspace MSRV remains 1.97.1 and is validated separately: +GNU tar, and SHA-256 tooling. The workspace MSRV remains 1.96.0 and is validated separately: ```text SOURCE_DATE_EPOCH=0 scripts/package-release.sh x86_64-unknown-linux-gnu -scripts/smoke-install.sh dist/code-system-graph-x86_64-unknown-linux-gnu-v1.1.0 -sha256sum --check dist/code-system-graph-x86_64-unknown-linux-gnu-v1.1.0.sha256 +scripts/smoke-install.sh dist/code-system-graph-x86_64-unknown-linux-gnu-v1.2.0 +sha256sum --check dist/code-system-graph-x86_64-unknown-linux-gnu-v1.2.0.sha256 ``` The package contains `csgraph`, `code-system-graph-hooks`, the visible Agent integration template @@ -144,7 +136,7 @@ clean `main` branch aligned with `origin/main`, Cargo credentials for crates.io, selects `prepare`: ```text -.github/workflows/release.sh 1.1.0 prepare +.github/workflows/release.sh 1.2.0 prepare ``` Preparation runs the complete publish-readiness suite and dry-runs all five packages without @@ -152,7 +144,7 @@ creating a tag, publishing a crate, or dispatching a workflow. To perform the ir pass `publish` explicitly: ```text -.github/workflows/release.sh 1.1.0 publish +.github/workflows/release.sh 1.2.0 publish ``` Publish mode verifies that the workspace repository matches `origin`, creates and pushes the diff --git a/docs/SUPPORTED_TECHNOLOGIES.md b/docs/SUPPORTED_TECHNOLOGIES.md index 20cfbd3..2fe876e 100644 --- a/docs/SUPPORTED_TECHNOLOGIES.md +++ b/docs/SUPPORTED_TECHNOLOGIES.md @@ -25,11 +25,34 @@ guessed. | Language | HTTP clients | HTTP servers | Database access | Recognized tests | | --- | --- | --- | --- | --- | -| TypeScript / JavaScript | Fetch, Axios | Express, Fastify, NestJS, Next.js App Router | Literal SQL | Not recognized | -| Python | requests, HTTPX, aiohttp, static method registries | FastAPI, Flask | psycopg/psycopg2, PyMySQL, SQLAlchemy ORM, Alembic, literal SQL | pytest, unittest, Factory Boy model links | -| Go | `net/http` | `net/http`, Gin, Chi | Literal SQL | Not recognized | -| Java | WebClient, Feign | Spring MVC | Literal SQL | Not recognized | -| Rust | Reqwest | Axum, Actix Web; advisory Utoipa/OpenAPI operations | SQLx, `mysql_async`, Diesel, literal SQL | Built-in tests, Tokio tests, rstest | +| TypeScript / JavaScript | Fetch, Axios | Express, Fastify, NestJS, Next.js App Router | Literal SQL | Jest, Vitest, Mocha, and Playwright `describe`/`it`/`test`; supertest, Playwright `request` | +| Python | requests, HTTPX, aiohttp, static method registries | FastAPI, Flask | psycopg/psycopg2, PyMySQL, SQLAlchemy ORM, Alembic, literal SQL | pytest, unittest, Factory Boy model links; `TestClient`, Flask `test_client()`, HTTPX with `app=` or an ASGI/WSGI transport | +| Go | `net/http` | `net/http`, Gin, Chi | Literal SQL | `TestX(t *testing.T)`; `httptest.NewRequest`, `httptest.NewServer` with `srv.URL` | +| Java | WebClient, Feign, RestTemplate | Spring MVC | Literal SQL | JUnit `@Test`, `@ParameterizedTest`, `@RepeatedTest`; MockMvc, RestAssured, `WebTestClient`, `TestRestTemplate` | +| Rust | Reqwest | Axum, Actix Web; advisory Utoipa/OpenAPI operations | SQLx, `mysql_async`, Diesel, literal SQL | Built-in, Tokio, Actix, and rstest tests; Axum `Request` builders sent with `oneshot`, Actix `test::TestRequest` | + +Server routes carry the full path a request reaches, including router prefixes declared in other +files of the same repository: + +| Framework | Composed prefixes | +| --- | --- | +| FastAPI | `APIRouter(prefix=)` and `include_router(router, prefix=)`, including imported routers | +| Flask | `Blueprint(url_prefix=)` and `register_blueprint(bp, url_prefix=)` | +| Express | `router.use(prefix, router)`, imported, default-exported, `module.exports`, and factory routers | +| NestJS | `@Controller(path)` and `app.setGlobalPrefix(prefix)` | +| Spring MVC and Feign | Class-level `@RequestMapping` and `@FeignClient(path=)` | +| Gin and Chi | `Group`, `Route`, `Mount`, `http.StripPrefix`, and router functions called with a group | +| Axum | `nest` and `merge` of local, returned, and module-path routers | +| Actix Web | `web::scope`, `service`, and `configure` of local and module-path services | + +Route handlers are the function a route registration names, whatever it is called: the last +argument of `app.get("/x", auth, orders.list)` or `r.GET("/x", handlers.Get)`, unwrapped from a +single-argument wrapper such as `asyncHandler(orders.create)`, FastAPI `add_api_route` and Flask +`add_url_rule` endpoints, and the Spring method declared after a mapping annotation. An inline +handler function is identified by its method and path. + +Prefix chains deeper than eight mounts, mount cycles, and routers that do not resolve to exactly +one declaration in the repository keep the route path declared next to the handler. JavaScript and TypeScript are separate parser inputs but share the same focused framework recognition. Each supported language is parsed structurally before framework-specific patterns are @@ -73,8 +96,47 @@ TypeScript Fetch call -> implemented by a Rust Axum handler ``` -Tests with a supported, exact HTTP target can also be connected to the contract they validate. -Dynamic URL construction, wrapper functions, and multiple providers are not guessed. +Tests with a supported HTTP target can also be connected to the contract they validate. A concrete +request path such as `/orders/42` reaches the `/orders/{id}` template whatever parameter syntax the +provider framework uses; the most specific template wins, and equally specific providers are +narrowed to the caller's repository before being reported as ambiguous. + +Parameters embedded in a segment retain their literal constraints: `/files/{name}.json` matches +`/files/readme.json` and rejects `/files/readme.xml`. Static segments rank ahead of mixed segments, +then whole-segment parameters and catch-alls. Overlapping mixed templates remain ambiguous. + +Client URLs are evaluated from the expression that builds them: + +- constants and variables bound earlier in the same function or file, including `this.` and `self.` + fields with one value; +- concatenation, Python f-strings, `%` and `.format`, Rust `format!` and `concat!`, JavaScript + template literals, Go `fmt.Sprintf`, and Java `String.format`; +- conversions such as `str()`, `String()`, `encodeURIComponent`, `strconv.Itoa`, and + `String.valueOf`. + +A runtime value in the scheme or authority leaves the path exact and the authority unknown. A +runtime value that fills a whole path segment becomes a path parameter. Any other runtime value +makes the path dynamic, and the call is reported without a path. + +A client call whose URL depends on a parameter of its function is a wrapper. Each call site that +binds the parameter, in the same repository and through relative imports, package imports, or +module paths, gains the instantiated call. Wrappers that forward the parameter to another wrapper +are followed for up to three hops. A test is linked to every exact client call made by the +functions it calls and the pytest fixtures it requests, including fixtures in an enclosing +`conftest.py`, for up to three calls deep. Calls that do not resolve to exactly one function in the +repository are not followed. + +A test without a named function, such as a Jest or Playwright `it`/`test` callback, is identified by +its file, its `describe` chain, and its title, and it also reaches the `beforeEach`, `beforeAll`, and +Mocha `before` blocks of its enclosing suites. A request made through a parameter, such as +`client.get("/orders/42")` in a helper or a test that receives `client`, is attributed to an +in-process test client only when that parameter is a pytest fixture returning one, or receives one +from every caller that resolves; otherwise it is not reported. + +In-process test clients (`TestClient`, Flask `test_client()`, supertest, `httptest`, MockMvc, +RestAssured, `WebTestClient`, `TestRestTemplate`, Axum `oneshot`, and Actix `test::TestRequest`) +exercise the application of their own repository, so their requests resolve only against providers +in that repository, even when another repository serves the same route. Rust executable-route evidence and Utoipa/OpenAPI annotations are additive, not mutually exclusive. An exact Actix or Axum declaration is confirmed executable evidence. A Utoipa operation @@ -242,7 +304,7 @@ exact member manifest is independently discovered. | --- | --- | | Docker Compose | Services, images, ports, dependencies, environment key names | | Kubernetes | Workloads, Services, Ingresses, static resources, selectors, Secret key names | -| Helm | Templates and values; templated identities remain incomplete | +| Helm | `Chart.yaml`, values, and YAML or `.tpl` templates of a chart; templated identities remain incomplete | | Terraform / OpenTofu | Literal resources and dependency metadata | | Documentation | README, runbook, RFC, ADR, headings, and explicit links | | Ownership | Ordered CODEOWNERS rules and explicit owner identities | @@ -278,7 +340,7 @@ See [Configuration](CONFIGURATION.md) for precedence, examples, and safety rules Code System Graph does not: - infer dynamic routes, topics, SQL, package coordinates, or generated-code roles; -- choose between duplicate providers; +- choose between equally specific providers that no scope rule separates; - execute application code, GraphQL schemas, Terraform, Helm, or `protoc`; - persist source bodies, SQL literal values, credentials, or configuration values; - replace a repository-local symbol and call graph. diff --git a/docs/adr/0003-sqlite-storage.md b/docs/adr/0003-sqlite-storage.md index e0e78f7..49ed5df 100644 --- a/docs/adr/0003-sqlite-storage.md +++ b/docs/adr/0003-sqlite-storage.md @@ -11,26 +11,27 @@ installation, concurrent readers, and deterministic recovery. ## Decision Code System Graph uses bundled SQLite through `rusqlite`. Every connection enables foreign keys, WAL, -and a busy timeout. The unpublished 1.0.0 binary embeds one definitive initial schema and records its -exact version in `schema_metadata`. +and a busy timeout. Each binary embeds one exact schema and records its identity, a BLAKE3 digest of +the schema definition, in `schema_metadata`. -Extractor output is immutable data until an application transaction atomically replaces the -batch keyed by repository snapshot, extractor, and source artifact. A snapshot becomes current -only after all selected extractors and linkers succeed. The previous valid snapshot therefore -survives interruption. +Each workspace owns one current graph. Extractor output is immutable data until an application +transaction publishes the candidate graph as a delta against the stored rows: only inserted, +changed, and removed rows are written. A snapshot becomes current only after all selected +extractors and linkers succeed. The previous valid graph therefore survives interruption, and the +database does not grow with the number of scans. SQLite calls execute through a dedicated blocking boundary. The schema uses relational columns for stable and queried attributes; JSON is reserved for genuinely variant attributes. FTS5 indexes user-facing labels and descriptions. Repository and checkout paths use a platform tag plus lossless native bytes; lossy display paths -are diagnostic-only. Before the first public release, only the definitive 1.0.0 schema is accepted; -incompatible local databases are rebuilt. Online backups use SQLite's backup API and do not +are diagnostic-only. Only the exact embedded schema is accepted; other local databases are removed +and rebuilt by a full scan. Online backups use SQLite's backup API and do not overwrite an existing destination. A restrictive sidecar lock provides single-writer ownership and bounded stale-lock recovery. Restore validates the selected source's exact schema and preserves the replaced database as a -separately named safety backup. Migration policy will be designed after a schema has shipped. +separately named safety backup. Writer-lock metadata includes PID and process start time. Recovery requires both expiry and proof that the original process instance is no longer alive, preventing long-running writers or reused diff --git a/docs/adr/0004-codegraph-provider-boundary.md b/docs/adr/0004-codegraph-provider-boundary.md index 7445f86..c68fba1 100644 --- a/docs/adr/0004-codegraph-provider-boundary.md +++ b/docs/adr/0004-codegraph-provider-boundary.md @@ -35,8 +35,8 @@ upgrade commands. ## Implementation evidence -The provider contract is validated against CodeGraph 1.5.0. The official Rust MCP client -negotiates initialization and discovers tools dynamically; 1.5.0 exposes `codegraph_explore` +The provider contract is validated against CodeGraph 1.6.1. The official Rust MCP client +negotiates initialization and discovers tools dynamically; 1.6.1 exposes `codegraph_explore` with explicit `projectPath`. Structured symbol, neighbor, impact, and affected-test operations use versioned CLI JSON contracts. MCP starts with `--no-watch`, timeout opens a per-repository MCP circuit, and CLI diff --git a/docs/adr/0009-release-engineering.md b/docs/adr/0009-release-engineering.md index b52b9e8..7ff7536 100644 --- a/docs/adr/0009-release-engineering.md +++ b/docs/adr/0009-release-engineering.md @@ -19,7 +19,7 @@ continuous validation on `main`, while release artifacts remain gated behind exp ## Decision The workspace and release artifacts use version `1.0.0`. Release builds use the latest stable Rust -toolchain and the committed lockfile. Every crate declares an MSRV of 1.97.1, which is checked +toolchain and the committed lockfile. Every crate declares an MSRV of 1.96.0, which is checked separately. Formatting and strict Clippy checks use the latest nightly toolchain, while compilation, tests, documentation, and release gates use stable. diff --git a/fixtures/codegraph/1.5.0/affected.json b/fixtures/codegraph/1.5.0/affected.json deleted file mode 100644 index 7963d89..0000000 --- a/fixtures/codegraph/1.5.0/affected.json +++ /dev/null @@ -1,9 +0,0 @@ -{ - "changedFiles": [ - "crates/code-system-graph-core/src/provider.rs" - ], - "affectedTests": [ - "crates/code-system-graph-core/tests/provider_e2e.rs" - ], - "totalDependentsTraversed": 6 -} diff --git a/fixtures/codegraph/1.5.0/impact.json b/fixtures/codegraph/1.5.0/impact.json deleted file mode 100644 index e9d407a..0000000 --- a/fixtures/codegraph/1.5.0/impact.json +++ /dev/null @@ -1,20 +0,0 @@ -{ - "symbol": "scan_workspace", - "depth": 2, - "nodeCount": 2, - "edgeCount": 1, - "affected": [ - { - "name": "scan_workspace", - "kind": "function", - "filePath": "crates/code-system-graph-cli/src/lib.rs", - "startLine": 270 - }, - { - "name": "scan_should_publish_snapshot", - "kind": "function", - "filePath": "crates/code-system-graph-cli/tests/http_trace_e2e.rs", - "startLine": 63 - } - ] -} diff --git a/fixtures/codegraph/1.5.0/neighbors.json b/fixtures/codegraph/1.5.0/neighbors.json deleted file mode 100644 index 76ebc3a..0000000 --- a/fixtures/codegraph/1.5.0/neighbors.json +++ /dev/null @@ -1,11 +0,0 @@ -{ - "symbol": "scan_workspace", - "callers": [ - { - "name": "scan_should_publish_snapshot", - "kind": "function", - "filePath": "crates/code-system-graph-cli/tests/http_trace_e2e.rs", - "startLine": 63 - } - ] -} diff --git a/fixtures/codegraph/1.5.0/query.json b/fixtures/codegraph/1.5.0/query.json deleted file mode 100644 index a9f1dc6..0000000 --- a/fixtures/codegraph/1.5.0/query.json +++ /dev/null @@ -1,15 +0,0 @@ -[ - { - "node": { - "id": "function:fixture", - "kind": "function", - "name": "scan_workspace", - "qualifiedName": "scan_workspace", - "filePath": "crates/code-system-graph-cli/src/lib.rs", - "language": "rust", - "startLine": 270, - "endLine": 278 - }, - "score": 19.5 - } -] diff --git a/fixtures/codegraph/1.5.0/status.json b/fixtures/codegraph/1.5.0/status.json deleted file mode 100644 index 06c8a2b..0000000 --- a/fixtures/codegraph/1.5.0/status.json +++ /dev/null @@ -1,24 +0,0 @@ -{ - "initialized": true, - "version": "1.5.0", - "projectPath": "/fixture/repository", - "indexPath": "/fixture/repository/.codegraph", - "lastIndexed": "2026-07-29T00:00:00.000Z", - "fileCount": 24, - "nodeCount": 649, - "edgeCount": 1827, - "pendingChanges": { - "added": 0, - "modified": 0, - "removed": 0 - }, - "worktreeMismatch": null, - "index": { - "builtWithVersion": "1.5.0", - "builtWithExtractionVersion": 24, - "currentExtractionVersion": 24, - "reindexRecommended": false, - "state": "complete", - "pendingRefs": 0 - } -} diff --git a/fixtures/codegraph/1.5.0/tools-list.json b/fixtures/codegraph/1.5.0/tools-list.json deleted file mode 100644 index e1991f1..0000000 --- a/fixtures/codegraph/1.5.0/tools-list.json +++ /dev/null @@ -1,24 +0,0 @@ -[ - { - "name": "codegraph_explore", - "description": "Explore relevant symbols, source, and call paths.", - "inputSchema": { - "type": "object", - "properties": { - "query": { - "type": "string" - }, - "maxFiles": { - "type": "number", - "default": 12 - }, - "projectPath": { - "type": "string" - } - }, - "required": [ - "query" - ] - } - } -] diff --git a/fixtures/codegraph/1.6.1/affected.json b/fixtures/codegraph/1.6.1/affected.json new file mode 100644 index 0000000..c521c9a --- /dev/null +++ b/fixtures/codegraph/1.6.1/affected.json @@ -0,0 +1,9 @@ +{ + "changedFiles": [ + "src/provider.ts" + ], + "affectedTests": [ + "tests/provider.test.ts" + ], + "totalDependentsTraversed": 1 +} diff --git a/fixtures/codegraph/1.6.1/impact.json b/fixtures/codegraph/1.6.1/impact.json new file mode 100644 index 0000000..668c9a1 --- /dev/null +++ b/fixtures/codegraph/1.6.1/impact.json @@ -0,0 +1,83 @@ +{ + "symbol": "discover", + "depth": 2, + "targets": [ + { + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1, + "id": "function:fixture-1", + "qualifiedName": "discover", + "language": "rust" + } + ], + "ambiguous": false, + "aggregation": "definition", + "filteredOut": false, + "definitions": [ + { + "definition": { + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1, + "id": "function:fixture-1", + "qualifiedName": "discover", + "language": "rust" + }, + "roots": [ + "function:fixture-1" + ], + "nodeCount": 2, + "edgeCount": 1, + "affected": [ + { + "id": "function:fixture-1", + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1 + }, + { + "id": "function:fixture-2", + "name": "scan_workspace", + "kind": "function", + "filePath": "src/lib.rs", + "startLine": 3 + } + ], + "edges": [ + { + "source": "function:fixture-2", + "target": "function:fixture-1", + "kind": "calls", + "metadata": { + "confidence": 0.9, + "resolvedBy": "import", + "refName": "provider::discover" + }, + "line": 4, + "column": 4, + "provenance": null + } + ] + } + ], + "nodeCount": 2, + "edgeCount": 1, + "affected": [ + { + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1 + }, + { + "name": "scan_workspace", + "kind": "function", + "filePath": "src/lib.rs", + "startLine": 3 + } + ] +} diff --git a/fixtures/codegraph/1.5.0/initialize.json b/fixtures/codegraph/1.6.1/initialize.json similarity index 51% rename from fixtures/codegraph/1.5.0/initialize.json rename to fixtures/codegraph/1.6.1/initialize.json index 018a69a..c6f07c1 100644 --- a/fixtures/codegraph/1.5.0/initialize.json +++ b/fixtures/codegraph/1.6.1/initialize.json @@ -5,6 +5,7 @@ }, "serverInfo": { "name": "codegraph", - "version": "1.5.0" - } + "version": "1.6.1" + }, + "instructions": "# Codegraph — code intelligence over an indexed knowledge graph" } diff --git a/fixtures/codegraph/1.6.1/neighbors.json b/fixtures/codegraph/1.6.1/neighbors.json new file mode 100644 index 0000000..99f9d1e --- /dev/null +++ b/fixtures/codegraph/1.6.1/neighbors.json @@ -0,0 +1,77 @@ +{ + "symbol": "discover", + "targets": [ + { + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1, + "id": "function:fixture-1", + "qualifiedName": "discover", + "language": "rust" + } + ], + "ambiguous": false, + "aggregation": "definition", + "filteredOut": false, + "definitions": [ + { + "definition": { + "name": "discover", + "kind": "function", + "filePath": "src/provider.rs", + "startLine": 1, + "id": "function:fixture-1", + "qualifiedName": "discover", + "language": "rust" + }, + "roots": [ + "function:fixture-1" + ], + "callers": [ + { + "id": "function:fixture-2", + "name": "scan_workspace", + "kind": "function", + "filePath": "src/lib.rs", + "startLine": 3, + "relationships": [ + "calls" + ] + } + ], + "edges": [ + { + "source": "function:fixture-2", + "target": "function:fixture-1", + "kind": "calls", + "metadata": { + "confidence": 0.9, + "resolvedBy": "import", + "refName": "provider::discover" + }, + "line": 4, + "column": 4, + "provenance": null + } + ], + "total": 1, + "limit": 10, + "truncated": false + } + ], + "callers": [ + { + "name": "scan_workspace", + "kind": "function", + "filePath": "src/lib.rs", + "startLine": 3, + "relationships": [ + "calls" + ] + } + ], + "total": 1, + "limit": 10, + "truncated": false +} diff --git a/fixtures/codegraph/1.6.1/query.json b/fixtures/codegraph/1.6.1/query.json new file mode 100644 index 0000000..f852af0 --- /dev/null +++ b/fixtures/codegraph/1.6.1/query.json @@ -0,0 +1,24 @@ +[ + { + "node": { + "id": "function:fixture", + "kind": "function", + "name": "scan_workspace", + "qualifiedName": "scan_workspace", + "filePath": "src/lib.rs", + "language": "rust", + "startLine": 3, + "endLine": 5, + "startColumn": 0, + "endColumn": 1, + "signature": "(root: &str) -> usize", + "visibility": "public", + "isExported": false, + "isAsync": false, + "isStatic": false, + "isAbstract": false, + "updatedAt": 1790899200000 + }, + "score": 91.66 + } +] diff --git a/fixtures/codegraph/1.6.1/status.json b/fixtures/codegraph/1.6.1/status.json new file mode 100644 index 0000000..16ae907 --- /dev/null +++ b/fixtures/codegraph/1.6.1/status.json @@ -0,0 +1,36 @@ +{ + "initialized": true, + "version": "1.6.1", + "projectPath": "/fixture/repository", + "indexPath": "/fixture/repository/.codegraph", + "lastIndexed": "2026-10-02T00:00:00.000Z", + "fileCount": 3, + "nodeCount": 7, + "edgeCount": 5, + "dbSizeBytes": 176128, + "walSizeBytes": 8272, + "backend": "node-sqlite", + "journalMode": "wal", + "nodesByKind": { + "file": 3, + "function": 3, + "import": 1 + }, + "languages": [ + "rust" + ], + "pendingChanges": { + "added": 0, + "modified": 0, + "removed": 0 + }, + "worktreeMismatch": null, + "index": { + "builtWithVersion": "1.6.1", + "builtWithExtractionVersion": 27, + "currentExtractionVersion": 27, + "reindexRecommended": false, + "state": "complete", + "pendingRefs": 0 + } +} diff --git a/fixtures/codegraph/1.6.1/tools-list.json b/fixtures/codegraph/1.6.1/tools-list.json new file mode 100644 index 0000000..9b8f04e --- /dev/null +++ b/fixtures/codegraph/1.6.1/tools-list.json @@ -0,0 +1,33 @@ +[ + { + "name": "codegraph_explore", + "description": "PRIMARY TOOL — call FIRST for almost any question OR before an edit: how does X work, architecture, a bug, where/what is X, surveying an area, or the symbols you are about to change.", + "inputSchema": { + "type": "object", + "properties": { + "query": { + "type": "string" + }, + "maxFiles": { + "type": "number", + "default": 12 + }, + "projectPath": { + "type": "string" + } + }, + "required": [ + "query" + ] + }, + "annotations": { + "readOnlyHint": true, + "destructiveHint": false, + "idempotentHint": true, + "openWorldHint": false + }, + "_meta": { + "anthropic/alwaysLoad": true + } + } +] diff --git a/fixtures/codegraph/degraded/missing-index-status.json b/fixtures/codegraph/degraded/missing-index-status.json index fecf3f4..26d33c7 100644 --- a/fixtures/codegraph/degraded/missing-index-status.json +++ b/fixtures/codegraph/degraded/missing-index-status.json @@ -1,6 +1,6 @@ { "initialized": false, - "version": "1.5.0", + "version": "1.6.1", "projectPath": "/fixture/missing", "indexPath": "/fixture/missing/.codegraph", "lastIndexed": null diff --git a/fixtures/codegraph/degraded/stale-index-status.json b/fixtures/codegraph/degraded/stale-index-status.json index 2066584..0fea136 100644 --- a/fixtures/codegraph/degraded/stale-index-status.json +++ b/fixtures/codegraph/degraded/stale-index-status.json @@ -1,6 +1,6 @@ { "initialized": true, - "version": "1.5.0", + "version": "1.6.1", "projectPath": "/fixture/stale", "indexPath": "/fixture/stale/.codegraph", "lastIndexed": "2026-07-28T00:00:00.000Z", @@ -11,7 +11,11 @@ }, "worktreeMismatch": null, "index": { + "builtWithVersion": "1.6.1", + "builtWithExtractionVersion": 27, + "currentExtractionVersion": 27, "reindexRecommended": false, - "state": "complete" + "state": "complete", + "pendingRefs": 0 } } diff --git a/fixtures/codegraph/fake/codegraph.py b/fixtures/codegraph/fake/codegraph.py index 5d24139..914a003 100644 --- a/fixtures/codegraph/fake/codegraph.py +++ b/fixtures/codegraph/fake/codegraph.py @@ -6,6 +6,22 @@ import sys import time +ANCHOR = { + "name": "anchor", + "kind": "function", + "filePath": "src/lib.rs", + "startLine": 7, + "id": "function:fixture", + "qualifiedName": "fixture::anchor", + "language": "rust", +} +NEIGHBOR = { + "name": "neighbor", + "kind": "function", + "filePath": "src/neighbor.rs", + "startLine": 9, +} + def emit(value): print(json.dumps(value, separators=(",", ":")), flush=True) @@ -28,7 +44,7 @@ def mcp_server(mode): "result": { "protocolVersion": "2024-11-05", "capabilities": {"tools": {}}, - "serverInfo": {"name": "codegraph", "version": "1.5.0"}, + "serverInfo": {"name": "codegraph", "version": "1.6.1"}, }, } ) @@ -75,14 +91,14 @@ def main(): mode = pathlib.Path(sys.argv[0]).name.removeprefix("codegraph-") command = sys.argv[1] if len(sys.argv) > 1 else "" if command == "--version": - print("1.5.0") + print("1.6.1") elif command == "serve": mcp_server(mode) elif command == "status": emit( { "initialized": True, - "version": "1.5.0", + "version": "1.6.1", "pendingChanges": {"added": 0, "modified": 0, "removed": 0}, "worktreeMismatch": None, "index": {"reindexRecommended": False, "state": "complete"}, @@ -113,14 +129,14 @@ def main(): emit( { "symbol": "anchor", - command: [ - { - "name": "neighbor", - "kind": "function", - "filePath": "src/neighbor.rs", - "startLine": 9, - } - ], + "targets": [ANCHOR], + "ambiguous": False, + "aggregation": "definition", + "filteredOut": False, + command: [NEIGHBOR], + "total": 1, + "limit": 20, + "truncated": False, } ) elif command == "impact": @@ -128,15 +144,13 @@ def main(): { "symbol": "anchor", "depth": 2, + "targets": [ANCHOR], + "ambiguous": False, + "aggregation": "definition", + "filteredOut": False, "nodeCount": 1, - "affected": [ - { - "name": "neighbor", - "kind": "function", - "filePath": "src/neighbor.rs", - "startLine": 9, - } - ], + "edgeCount": 0, + "affected": [NEIGHBOR], } ) elif command == "affected": diff --git a/fixtures/cross-language-matrix/actix/src/items.rs b/fixtures/cross-language-matrix/actix/src/items.rs new file mode 100644 index 0000000..142a0ab --- /dev/null +++ b/fixtures/cross-language-matrix/actix/src/items.rs @@ -0,0 +1,9 @@ +use actix_web::{web, HttpResponse}; + +pub fn configure(cfg: &mut web::ServiceConfig) { + cfg.route("/items/{id}", web::get().to(read_item)); +} + +async fn read_item(id: web::Path) -> HttpResponse { + HttpResponse::Ok().body(id.into_inner()) +} diff --git a/fixtures/cross-language-matrix/actix/src/main.rs b/fixtures/cross-language-matrix/actix/src/main.rs new file mode 100644 index 0000000..ffaec11 --- /dev/null +++ b/fixtures/cross-language-matrix/actix/src/main.rs @@ -0,0 +1,11 @@ +mod items; + +use actix_web::{web, App, HttpServer}; + +#[actix_web::main] +async fn main() -> std::io::Result<()> { + HttpServer::new(|| App::new().service(web::scope("/actix").configure(items::configure))) + .bind(("0.0.0.0", 8080))? + .run() + .await +} diff --git a/fixtures/cross-language-matrix/axum/src/items.rs b/fixtures/cross-language-matrix/axum/src/items.rs new file mode 100644 index 0000000..6afaae8 --- /dev/null +++ b/fixtures/cross-language-matrix/axum/src/items.rs @@ -0,0 +1,11 @@ +use axum::extract::Path; +use axum::routing::get; +use axum::Router; + +pub fn router() -> Router { + Router::new().route("/items/:id", get(read_item)) +} + +async fn read_item(Path(id): Path) -> String { + id +} diff --git a/fixtures/cross-language-matrix/axum/src/main.rs b/fixtures/cross-language-matrix/axum/src/main.rs new file mode 100644 index 0000000..45f2f59 --- /dev/null +++ b/fixtures/cross-language-matrix/axum/src/main.rs @@ -0,0 +1,10 @@ +mod items; + +use axum::Router; + +#[tokio::main] +async fn main() { + let app = Router::new().nest("/axum", items::router()); + let listener = tokio::net::TcpListener::bind("0.0.0.0:3000").await.unwrap(); + axum::serve(listener, app).await.unwrap(); +} diff --git a/fixtures/cross-language-matrix/chi/main.go b/fixtures/cross-language-matrix/chi/main.go new file mode 100644 index 0000000..6c95a0b --- /dev/null +++ b/fixtures/cross-language-matrix/chi/main.go @@ -0,0 +1,20 @@ +package main + +import ( + "net/http" + + "github.com/go-chi/chi/v5" +) + +func getItem(w http.ResponseWriter, r *http.Request) {} + +func health(w http.ResponseWriter, r *http.Request) {} + +func main() { + r := chi.NewRouter() + r.Get("/health", health) + r.Route("/chi/items", func(r chi.Router) { + r.Get("/{id}", getItem) + }) + http.ListenAndServe(":8080", r) +} diff --git a/fixtures/cross-language-matrix/code-system-graph.yaml b/fixtures/cross-language-matrix/code-system-graph.yaml new file mode 100644 index 0000000..086b3c7 --- /dev/null +++ b/fixtures/cross-language-matrix/code-system-graph.yaml @@ -0,0 +1,46 @@ +version: 1 +name: cross-language-matrix +repos: + fastapi: + path: fastapi + authorities: [fastapi-svc] + flask: + path: flask + authorities: [flask-svc] + express: + path: express + authorities: [express-svc] + nest: + path: nest + authorities: [nest-svc] + next: + path: next + authorities: [next-svc] + spring: + path: spring + authorities: [spring-svc] + gin: + path: gin + authorities: [gin-svc] + chi: + path: chi + authorities: [chi-svc] + nethttp: + path: nethttp + authorities: [nethttp-svc] + axum: + path: axum + authorities: [axum-svc] + actix: + path: actix + authorities: [actix-svc] + tests-python: + path: tests-python + tests-ts: + path: tests-ts + tests-go: + path: tests-go + tests-java: + path: tests-java + tests-rust: + path: tests-rust diff --git a/fixtures/cross-language-matrix/express/src/app.js b/fixtures/cross-language-matrix/express/src/app.js new file mode 100644 index 0000000..aa60e0d --- /dev/null +++ b/fixtures/cross-language-matrix/express/src/app.js @@ -0,0 +1,7 @@ +const express = require("express"); +const items = require("./routes/items"); + +const app = express(); +app.use("/express/items", items); + +module.exports = app; diff --git a/fixtures/cross-language-matrix/express/src/controllers/items.js b/fixtures/cross-language-matrix/express/src/controllers/items.js new file mode 100644 index 0000000..1227c77 --- /dev/null +++ b/fixtures/cross-language-matrix/express/src/controllers/items.js @@ -0,0 +1,3 @@ +exports.read = (req, res) => { + res.json({ id: req.params.id }); +}; diff --git a/fixtures/cross-language-matrix/express/src/routes/items.js b/fixtures/cross-language-matrix/express/src/routes/items.js new file mode 100644 index 0000000..f673e0b --- /dev/null +++ b/fixtures/cross-language-matrix/express/src/routes/items.js @@ -0,0 +1,7 @@ +const express = require("express"); +const items = require("../controllers/items"); + +const router = express.Router(); +router.get("/:id", items.read); + +module.exports = router; diff --git a/fixtures/cross-language-matrix/fastapi/app/main.py b/fixtures/cross-language-matrix/fastapi/app/main.py new file mode 100644 index 0000000..e782970 --- /dev/null +++ b/fixtures/cross-language-matrix/fastapi/app/main.py @@ -0,0 +1,6 @@ +from fastapi import FastAPI + +from app.routers import items + +app = FastAPI() +app.include_router(items.router, prefix="/fastapi") diff --git a/fixtures/cross-language-matrix/fastapi/app/routers/items.py b/fixtures/cross-language-matrix/fastapi/app/routers/items.py new file mode 100644 index 0000000..1dfeaf7 --- /dev/null +++ b/fixtures/cross-language-matrix/fastapi/app/routers/items.py @@ -0,0 +1,8 @@ +from fastapi import APIRouter + +router = APIRouter(prefix="/items") + + +@router.get("/{item_id}") +def read_item(item_id: int): + return {"id": item_id} diff --git a/fixtures/cross-language-matrix/flask/app/__init__.py b/fixtures/cross-language-matrix/flask/app/__init__.py new file mode 100644 index 0000000..e38c78a --- /dev/null +++ b/fixtures/cross-language-matrix/flask/app/__init__.py @@ -0,0 +1,9 @@ +from flask import Flask + +from app import items + + +def create_app(): + app = Flask(__name__) + app.register_blueprint(items.bp, url_prefix="/flask/items") + return app diff --git a/fixtures/cross-language-matrix/flask/app/items.py b/fixtures/cross-language-matrix/flask/app/items.py new file mode 100644 index 0000000..0378faf --- /dev/null +++ b/fixtures/cross-language-matrix/flask/app/items.py @@ -0,0 +1,8 @@ +from flask import Blueprint + +bp = Blueprint("items", __name__) + + +@bp.route("/", methods=["GET"]) +def read_item(item_id): + return {"id": item_id} diff --git a/fixtures/cross-language-matrix/gin/cmd/server/main.go b/fixtures/cross-language-matrix/gin/cmd/server/main.go new file mode 100644 index 0000000..1654864 --- /dev/null +++ b/fixtures/cross-language-matrix/gin/cmd/server/main.go @@ -0,0 +1,16 @@ +package main + +import ( + "github.com/gin-gonic/gin" + "example.com/gin/internal/items" +) + +func health(c *gin.Context) {} + +func main() { + r := gin.Default() + r.GET("/health", health) + api := r.Group("/gin") + items.Register(api.Group("/items")) + r.Run() +} diff --git a/fixtures/cross-language-matrix/gin/internal/items/items.go b/fixtures/cross-language-matrix/gin/internal/items/items.go new file mode 100644 index 0000000..decbc07 --- /dev/null +++ b/fixtures/cross-language-matrix/gin/internal/items/items.go @@ -0,0 +1,11 @@ +package items + +import "github.com/gin-gonic/gin" + +func Register(rg *gin.RouterGroup) { + rg.GET("/:id", getItem) +} + +func getItem(c *gin.Context) { + c.JSON(200, gin.H{"id": c.Param("id")}) +} diff --git a/fixtures/cross-language-matrix/nest/src/items.controller.ts b/fixtures/cross-language-matrix/nest/src/items.controller.ts new file mode 100644 index 0000000..e33fda3 --- /dev/null +++ b/fixtures/cross-language-matrix/nest/src/items.controller.ts @@ -0,0 +1,9 @@ +import { Controller, Get, Param } from "@nestjs/common"; + +@Controller("items") +export class ItemsController { + @Get(":id") + findOne(@Param("id") id: string) { + return { id }; + } +} diff --git a/fixtures/cross-language-matrix/nest/src/main.ts b/fixtures/cross-language-matrix/nest/src/main.ts new file mode 100644 index 0000000..12a7ced --- /dev/null +++ b/fixtures/cross-language-matrix/nest/src/main.ts @@ -0,0 +1,9 @@ +import { NestFactory } from "@nestjs/core"; +import { AppModule } from "./app.module"; + +async function bootstrap() { + const app = await NestFactory.create(AppModule); + app.setGlobalPrefix("nest"); + await app.listen(3000); +} +bootstrap(); diff --git a/fixtures/cross-language-matrix/nethttp/main.go b/fixtures/cross-language-matrix/nethttp/main.go new file mode 100644 index 0000000..718617d --- /dev/null +++ b/fixtures/cross-language-matrix/nethttp/main.go @@ -0,0 +1,11 @@ +package main + +import "net/http" + +func getItem(w http.ResponseWriter, r *http.Request) {} + +func main() { + mux := http.NewServeMux() + mux.HandleFunc("GET /nethttp/items/{id}", getItem) + http.ListenAndServe(":8080", mux) +} diff --git a/fixtures/cross-language-matrix/next/app/api/next/items/[id]/route.ts b/fixtures/cross-language-matrix/next/app/api/next/items/[id]/route.ts new file mode 100644 index 0000000..e98013f --- /dev/null +++ b/fixtures/cross-language-matrix/next/app/api/next/items/[id]/route.ts @@ -0,0 +1,3 @@ +export async function GET(request: Request, { params }: { params: { id: string } }) { + return Response.json({ id: params.id }); +} diff --git a/fixtures/cross-language-matrix/spring/src/main/java/com/matrix/ItemsController.java b/fixtures/cross-language-matrix/spring/src/main/java/com/matrix/ItemsController.java new file mode 100644 index 0000000..722e931 --- /dev/null +++ b/fixtures/cross-language-matrix/spring/src/main/java/com/matrix/ItemsController.java @@ -0,0 +1,15 @@ +package com.matrix; + +import org.springframework.web.bind.annotation.GetMapping; +import org.springframework.web.bind.annotation.PathVariable; +import org.springframework.web.bind.annotation.RequestMapping; +import org.springframework.web.bind.annotation.RestController; + +@RestController +@RequestMapping("/spring/items") +public class ItemsController { + @GetMapping("/{id}") + public Item read(@PathVariable String id) { + return new Item(id); + } +} diff --git a/fixtures/cross-language-matrix/tests-go/matrix_test.go b/fixtures/cross-language-matrix/tests-go/matrix_test.go new file mode 100644 index 0000000..23b621f --- /dev/null +++ b/fixtures/cross-language-matrix/tests-go/matrix_test.go @@ -0,0 +1,64 @@ +package matrix + +import ( + "net/http" + "testing" +) + +const ( + fastapiURL = "http://fastapi-svc:8000" + flaskURL = "http://flask-svc:5000" + expressURL = "http://express-svc:3000" + nestURL = "http://nest-svc:3000" + nextURL = "http://next-svc:3000" + springURL = "http://spring-svc:8080" + ginURL = "http://gin-svc:8080" + chiURL = "http://chi-svc:8080" + nethttpURL = "http://nethttp-svc:8080" + axumURL = "http://axum-svc:3000" + actixURL = "http://actix-svc:8080" +) + +func TestFastapiItem(t *testing.T) { + http.Get(fastapiURL + "/fastapi/items/42") +} + +func TestFlaskItem(t *testing.T) { + http.Get(flaskURL + "/flask/items/42") +} + +func TestExpressItem(t *testing.T) { + http.Get(expressURL + "/express/items/42") +} + +func TestNestItem(t *testing.T) { + http.Get(nestURL + "/nest/items/42") +} + +func TestNextItem(t *testing.T) { + http.Get(nextURL + "/api/next/items/42") +} + +func TestSpringItem(t *testing.T) { + http.Get(springURL + "/spring/items/42") +} + +func TestGinItem(t *testing.T) { + http.Get(ginURL + "/gin/items/42") +} + +func TestChiItem(t *testing.T) { + http.Get(chiURL + "/chi/items/42") +} + +func TestNethttpItem(t *testing.T) { + http.Get(nethttpURL + "/nethttp/items/42") +} + +func TestAxumItem(t *testing.T) { + http.Get(axumURL + "/axum/items/42") +} + +func TestActixItem(t *testing.T) { + http.Get(actixURL + "/actix/items/42") +} diff --git a/fixtures/cross-language-matrix/tests-java/src/test/java/com/matrix/MatrixTest.java b/fixtures/cross-language-matrix/tests-java/src/test/java/com/matrix/MatrixTest.java new file mode 100644 index 0000000..8b4bc66 --- /dev/null +++ b/fixtures/cross-language-matrix/tests-java/src/test/java/com/matrix/MatrixTest.java @@ -0,0 +1,75 @@ +package com.matrix; + +import org.junit.jupiter.api.Test; +import org.springframework.web.client.RestTemplate; + +class MatrixTest { + private static final String FASTAPI_URL = "http://fastapi-svc:8000"; + private static final String FLASK_URL = "http://flask-svc:5000"; + private static final String EXPRESS_URL = "http://express-svc:3000"; + private static final String NEST_URL = "http://nest-svc:3000"; + private static final String NEXT_URL = "http://next-svc:3000"; + private static final String SPRING_URL = "http://spring-svc:8080"; + private static final String GIN_URL = "http://gin-svc:8080"; + private static final String CHI_URL = "http://chi-svc:8080"; + private static final String NETHTTP_URL = "http://nethttp-svc:8080"; + private static final String AXUM_URL = "http://axum-svc:3000"; + private static final String ACTIX_URL = "http://actix-svc:8080"; + + private final RestTemplate restTemplate = new RestTemplate(); + + @Test + void readsFastapiItem() { + restTemplate.getForObject(FASTAPI_URL + "/fastapi/items/42", String.class); + } + + @Test + void readsFlaskItem() { + restTemplate.getForObject(FLASK_URL + "/flask/items/42", String.class); + } + + @Test + void readsExpressItem() { + restTemplate.getForObject(EXPRESS_URL + "/express/items/42", String.class); + } + + @Test + void readsNestItem() { + restTemplate.getForObject(NEST_URL + "/nest/items/42", String.class); + } + + @Test + void readsNextItem() { + restTemplate.getForObject(NEXT_URL + "/api/next/items/42", String.class); + } + + @Test + void readsSpringItem() { + restTemplate.getForObject(SPRING_URL + "/spring/items/42", String.class); + } + + @Test + void readsGinItem() { + restTemplate.getForObject(GIN_URL + "/gin/items/42", String.class); + } + + @Test + void readsChiItem() { + restTemplate.getForObject(CHI_URL + "/chi/items/42", String.class); + } + + @Test + void readsNethttpItem() { + restTemplate.getForObject(NETHTTP_URL + "/nethttp/items/42", String.class); + } + + @Test + void readsAxumItem() { + restTemplate.getForObject(AXUM_URL + "/axum/items/42", String.class); + } + + @Test + void readsActixItem() { + restTemplate.getForObject(ACTIX_URL + "/actix/items/42", String.class); + } +} diff --git a/fixtures/cross-language-matrix/tests-python/tests/test_matrix.py b/fixtures/cross-language-matrix/tests-python/tests/test_matrix.py new file mode 100644 index 0000000..54c44da --- /dev/null +++ b/fixtures/cross-language-matrix/tests-python/tests/test_matrix.py @@ -0,0 +1,70 @@ +import httpx +import requests + +FASTAPI_URL = "http://fastapi-svc:8000" +FLASK_URL = "http://flask-svc:5000" +EXPRESS_URL = "http://express-svc:3000" +NEST_URL = "http://nest-svc:3000" +NEXT_URL = "http://next-svc:3000" +SPRING_URL = "http://spring-svc:8080" +GIN_URL = "http://gin-svc:8080" +CHI_URL = "http://chi-svc:8080" +NETHTTP_URL = "http://nethttp-svc:8080" +AXUM_URL = "http://axum-svc:3000" +ACTIX_URL = "http://actix-svc:8080" + + +def test_fastapi_item(): + requests.get(f"{FASTAPI_URL}/fastapi/items/42") + + +def test_flask_item(): + httpx.get(f"{FLASK_URL}/flask/items/42") + + +def test_express_item(): + requests.get(f"{EXPRESS_URL}/express/items/42") + + +def test_nest_item(): + httpx.get(f"{NEST_URL}/nest/items/42") + + +def test_next_item(): + requests.get(f"{NEXT_URL}/api/next/items/42") + + +def test_spring_item(): + httpx.get(f"{SPRING_URL}/spring/items/42") + + +def test_gin_item(): + requests.get(f"{GIN_URL}/gin/items/42") + + +def test_chi_item(): + httpx.get(f"{CHI_URL}/chi/items/42") + + +def test_nethttp_item(): + requests.get(f"{NETHTTP_URL}/nethttp/items/42") + + +def test_axum_item(): + httpx.get(f"{AXUM_URL}/axum/items/42") + + +def test_actix_item(): + requests.get(f"{ACTIX_URL}/actix/items/42") + + +def test_missing_item(): + requests.get("http://localhost:9000/missing/items/42") + + +def test_health(): + requests.get("http://localhost:8080/health") + + +def test_public_status(): + requests.get("https://status.example.com/api/v2/status") diff --git a/fixtures/cross-language-matrix/tests-rust/tests/matrix.rs b/fixtures/cross-language-matrix/tests-rust/tests/matrix.rs new file mode 100644 index 0000000..362e24f --- /dev/null +++ b/fixtures/cross-language-matrix/tests-rust/tests/matrix.rs @@ -0,0 +1,66 @@ +const FASTAPI_URL: &str = "http://fastapi-svc:8000"; +const FLASK_URL: &str = "http://flask-svc:5000"; +const EXPRESS_URL: &str = "http://express-svc:3000"; +const NEST_URL: &str = "http://nest-svc:3000"; +const NEXT_URL: &str = "http://next-svc:3000"; +const SPRING_URL: &str = "http://spring-svc:8080"; +const GIN_URL: &str = "http://gin-svc:8080"; +const CHI_URL: &str = "http://chi-svc:8080"; +const NETHTTP_URL: &str = "http://nethttp-svc:8080"; +const AXUM_URL: &str = "http://axum-svc:3000"; +const ACTIX_URL: &str = "http://actix-svc:8080"; + +#[tokio::test] +async fn reads_fastapi_item() { + reqwest::get(format!("{FASTAPI_URL}/fastapi/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_flask_item() { + reqwest::get(format!("{FLASK_URL}/flask/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_express_item() { + reqwest::get(format!("{EXPRESS_URL}/express/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_nest_item() { + reqwest::get(format!("{NEST_URL}/nest/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_next_item() { + reqwest::get(format!("{NEXT_URL}/api/next/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_spring_item() { + reqwest::get(format!("{SPRING_URL}/spring/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_gin_item() { + reqwest::get(format!("{GIN_URL}/gin/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_chi_item() { + reqwest::get(format!("{CHI_URL}/chi/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_nethttp_item() { + reqwest::get(format!("{NETHTTP_URL}/nethttp/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_axum_item() { + reqwest::get(format!("{AXUM_URL}/axum/items/42")).await.unwrap(); +} + +#[tokio::test] +async fn reads_actix_item() { + reqwest::get(format!("{ACTIX_URL}/actix/items/42")).await.unwrap(); +} diff --git a/fixtures/cross-language-matrix/tests-ts/test/matrix.test.ts b/fixtures/cross-language-matrix/tests-ts/test/matrix.test.ts new file mode 100644 index 0000000..5d1fcda --- /dev/null +++ b/fixtures/cross-language-matrix/tests-ts/test/matrix.test.ts @@ -0,0 +1,59 @@ +import axios from "axios"; + +const FASTAPI_URL = "http://fastapi-svc:8000"; +const FLASK_URL = "http://flask-svc:5000"; +const EXPRESS_URL = "http://express-svc:3000"; +const NEST_URL = "http://nest-svc:3000"; +const NEXT_URL = "http://next-svc:3000"; +const SPRING_URL = "http://spring-svc:8080"; +const GIN_URL = "http://gin-svc:8080"; +const CHI_URL = "http://chi-svc:8080"; +const NETHTTP_URL = "http://nethttp-svc:8080"; +const AXUM_URL = "http://axum-svc:3000"; +const ACTIX_URL = "http://actix-svc:8080"; + +describe("matrix", () => { + it("reads fastapi", async () => { + await fetch(`${FASTAPI_URL}/fastapi/items/42`); + }); + + it("reads flask", async () => { + await axios.get(`${FLASK_URL}/flask/items/42`); + }); + + it("reads express", async () => { + await fetch(`${EXPRESS_URL}/express/items/42`); + }); + + it("reads nest", async () => { + await axios.get(`${NEST_URL}/nest/items/42`); + }); + + it("reads next", async () => { + await fetch(`${NEXT_URL}/api/next/items/42`); + }); + + it("reads spring", async () => { + await axios.get(`${SPRING_URL}/spring/items/42`); + }); + + it("reads gin", async () => { + await fetch(`${GIN_URL}/gin/items/42`); + }); + + it("reads chi", async () => { + await axios.get(`${CHI_URL}/chi/items/42`); + }); + + it("reads nethttp", async () => { + await fetch(`${NETHTTP_URL}/nethttp/items/42`); + }); + + it("reads axum", async () => { + await axios.get(`${AXUM_URL}/axum/items/42`); + }); + + it("reads actix", async () => { + await fetch(`${ACTIX_URL}/actix/items/42`); + }); +}); diff --git a/fuzz/Cargo.toml b/fuzz/Cargo.toml index 69dee15..3af1ed6 100644 --- a/fuzz/Cargo.toml +++ b/fuzz/Cargo.toml @@ -12,8 +12,8 @@ cargo-fuzz = true [dependencies] libfuzzer-sys = "0.4.13" -code-system-graph-core = { version = "1.1.0", path = "../crates/code-system-graph-core" } -code-system-graph-model = { version = "1.1.0", path = "../crates/code-system-graph-model" } +code-system-graph-core = { version = "1.2.0", path = "../crates/code-system-graph-core" } +code-system-graph-model = { version = "1.2.0", path = "../crates/code-system-graph-model" } serde_json = "1.0.151" [[bin]]