diff --git a/.streamlit/config.toml b/.streamlit/config.toml index 77c2fe1c..db8655b1 100644 --- a/.streamlit/config.toml +++ b/.streamlit/config.toml @@ -4,3 +4,7 @@ primaryColor = "#d8ff63" backgroundColor = "#0d0f12" secondaryBackgroundColor = "#171a1f" textColor = "#f7f7f2" + +[server] +# Multipart framing needs transport headroom; the JUnit widget/parser stay at 1 MiB. +maxUploadSize = 2 diff --git a/CHANGELOG.md b/CHANGELOG.md index 573ca678..9c821a27 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -21,6 +21,16 @@ Development version: `0.2.4.dev0`. Public install remains the immutable v0.2.3 r ### Post-release engineering +- Added a bounded, local JUnit XML adapter that accepts bytes only, rejects unsafe or over-limit + structures, persists only sanitized exact-head/provenance-bound results, and requires explicit + human criterion mapping. Imported results remain external non-gating context: they never become + E1–E4, observed CI, runtime verification, a reviewer decision, final acceptance, or correctness. + CLI inspection/import, Streamlit save/reopen, validated JSON/Markdown/HTML/CSV exports, + deterministic comparison, and an installed-wheel loopback-only Chromium round trip share the + same core contracts. Raw XML, output bodies, failure bodies, paths, URLs, and attachments are + neither persisted nor exported. +- Capped retained JUnit imports at 20 per review and each warning or limitation at 1,000 + characters each. A new analysis cannot inherit imports from an earlier criteria revision. - Consolidated the owner decision handoff around blocker-first unresolved decisions and a direct pre-matrix criterion handoff. This is a bounded Stage 2 workflow clarification, not a claim of acceptance-criteria correctness or runtime verification. diff --git a/README.md b/README.md index 30ba8ddc..cdf942b9 100644 --- a/README.md +++ b/README.md @@ -262,6 +262,49 @@ scopeproof compare PREVIOUS_REVIEW_ID CURRENT_REVIEW_ID \ --storage-dir .scopeproof/reviews ``` +Current `0.2.4.dev0` source can also inspect and append one bounded local JUnit XML artifact +without executing target-repository code. First inspect the sanitized scope IDs: + +```bash +scopeproof inspect-junit results.xml +``` + +Create a strict mapping document that records an explicit human relationship rather than +inferring one from test names: + +```json +{ + "schema_version": "junit-mapping-v1", + "artifact_sha256": "COPY_THE_64_CHARACTER_DIGEST_FROM_INSPECT_JUNIT", + "selections": [ + {"scope_id": "suite-0001", "criterion_id": "AC-01"} + ] +} +``` + +Then append the validated import to the exact-head saved review: + +```bash +scopeproof import-junit REVIEW_ID results.xml \ + --mapping junit-mapping.json \ + --importer "Asserted reviewer name or role" \ + --storage-dir .scopeproof/reviews +``` + +The adapter accepts at most 1 MiB, 100 suites, 5,000 cases, and 20,000 XML elements. Each review +can retain at most 20 imports, and each retained warning or asserted limitation is capped at 1,000 +characters. It accepts UTF-8 only and rejects DTDs, entities, non-declaration processing +instructions, XInclude, remote references, unsupported nesting, and ambiguous result markers. It +stores computed statuses, +stable local scope IDs, the artifact SHA-256, exact review and criteria provenance, explicit +mappings, an asserted importer, warnings, and limitations. Raw XML, stdout, stderr, properties, +failure bodies, commands, paths, URLs, and attachments are discarded; path- or URL-like +suite/class/test names are replaced with deterministic redacted labels. The CLI mapping digest +must match the selected artifact before any saved-review mutation. An import is external, +non-gating context—not E1, E2, E3, E4, observed CI, runtime verification, human acceptance, final +acceptance, or proof that a criterion passed. Failed inspection, mapping, validation, or storage +does not mutate the saved review. + `resolve` records one human criterion decision and never executes PR code. Static candidates never become runtime evidence through `resolve`; accepting below a criterion's required evidence level requires a non-empty reviewer note. `verify-runtime` is the only CLI @@ -270,12 +313,15 @@ atomically links a human-supplied E3/E4 runtime record to its manual-verificatio does not run or independently verify the cited artifact. Final acceptance remains fail-closed until the deterministic prerequisites are satisfied; use `--revoke` to append a revocation. `compare` validates both saved reviews, reports candidate changes without carrying decisions -forward, and refuses to overwrite an existing output file. +forward, reports imported-artifact and mapping changes separately, and refuses to overwrite an +existing output file. Changed imported context can require a previous human decision to be +reviewed again, but it never carries, creates, or changes that decision or either gate. CSV exports neutralize leading spreadsheet-formula characters in scalar text cells. Fields that can contain multiple values (`ingestion_warnings`, `skipped_files`, `evidence_links`, -`missing_evidence`, `runtime_artifacts`, and `runtime_result`) are JSON arrays inside their CSV -cells so delimiters in repository or reviewer text do not destroy provenance. +`missing_evidence`, `runtime_artifacts`, `runtime_result`, and imported-JUnit fields) are JSON +arrays inside their CSV cells so delimiters in repository or reviewer text do not destroy +provenance. Anonymous public-repository access is the default. `--token` is optional and can increase GitHub's free rate limit, but it is not required or persisted. The CLI never comments on the pull request, @@ -312,8 +358,9 @@ The six review sections are: 5. Evidence Matrix. 6. Summary & Export. -Criterion Review contains the selected criterion evidence, external verification, and human -resolution controls. Summary & Export provides the Markdown, JSON, and CSV review records. +Criterion Review contains the selected criterion evidence, separate external JUnit context, +external verification, and human resolution controls. Summary & Export provides the Markdown, +JSON, and CSV review records. ### Durable local review workflow @@ -329,7 +376,8 @@ safe local record IDs in deterministic order, while an empty store retains manua The app validates the selected record when it is opened and refuses a configured review path that is a symbolic link or another existing non-directory. This app-owned local directory prevents a browser input from selecting arbitrary file paths. Records preserve the review SHAs, criteria -revisions, evidence, findings, resolution history, and gate decision. They never contain the +revisions, evidence, bounded imported-JUnit envelopes, findings, resolution history, and gate +decision. They never contain raw JUnit XML or the optional GitHub token. A reopened review prepares its public PR URL and bounded unchanged-candidate paths for a one-click current-head check rather than silently reusing old evidence. Records also preserve whether public repository visibility was verified; legacy records without that fact @@ -342,6 +390,9 @@ reviewer can inspect what moved or changed before recording a new decision. Exac candidates remain inspectable in a collapsed section, and the validated comparison can be downloaded as Markdown or JSON. This comparison does not prove criterion satisfaction or carry a prior human decision forward. +Imported JUnit artifacts are compared separately by digest and explicit mapping signature as +Unchanged, Added, Removed, or Mapping modified. This projection does not reinterpret test names, +copy decisions, or make imported results a gate input. From the CLI, run `scopeproof list` to return the safe local review IDs in the default `.scopeproof/reviews` directory; add `--storage-dir PATH` only when earlier CLI commands used that diff --git a/ROADMAP.md b/ROADMAP.md index 28d490a9..dfd66b41 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -59,6 +59,12 @@ places the selected criterion evidence and controls before the secondary evidenc bounded Stage 2 follow-up. This remains Stage 2 engineering work and does not claim customer validation. +The current bounded JUnit adapter slice adds the first separately typed, non-executing evidence +adapter across the core, CLI, local workbench, saved records, exports, comparison, and installed +browser proof. It accepts only bounded local bytes, stores sanitized exact-head and criteria-bound +context, requires explicit human mapping, and never changes gate inputs or creates E1–E4, CI, +runtime, decision, acceptance, correctness, or customer-validation claims. + ### Verification and evidence boundaries - Engineering checks do not prove acceptance-criteria correctness. @@ -298,6 +304,11 @@ the selected criterion evidence and controls before the secondary evidence matri Stage 2 follow-up preserves every evidence boundary and keeps optional external research secondary. +The bounded JUnit adapter is the first implemented non-executing adapter. Its locally supplied +results remain exact-head/provenance-bound, separately rendered, explicitly mapped, and +non-gating. Coverage summaries, contract reports, deployment records, and other adapter families +remain separately scoped future decisions. + The [Stage 2 productization packet](docs/commercialization/stage2-readiness-packet.md) is the operating boundary. External commercial discovery is optional and separate from owner-led productization. It is not required to continue Stage 2 and needs separate owner authorization diff --git a/apps/web/app.py b/apps/web/app.py index 2809ec3e..8a614725 100644 --- a/apps/web/app.py +++ b/apps/web/app.py @@ -51,6 +51,13 @@ InvalidPullRequestUrl, parse_pr_url, ) +from scopeproof_core.importers.junit import ( + MAX_JUNIT_BYTES, + JUnitImportError, + JUnitMappingSelection, + build_junit_evidence_import, + parse_junit_artifact, +) from scopeproof_core.presentation import ( EvidenceStatus, criterion_coverage_rows, @@ -75,6 +82,7 @@ ResolutionEventStatus, acceptance_requires_comment, append_external_verification, + append_junit_evidence_import, append_resolution, attach_analysis, can_record_final_acceptance, @@ -85,6 +93,7 @@ ) from scopeproof_core.schemas.models import ( CONSTRUCTED_DEMO_CRITERIA_SOURCE_URI, + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, RULESET_VERSION, CheckState, Criterion, @@ -229,6 +238,8 @@ "criteria_source_draft": None, "criteria_source_widget_sync_pending": None, "source_widget_sync_pending": None, + "junit_artifact_upload_version": 0, + "junit_mapping_artifact_sha256": None, } for state_key, default in _STATE_DEFAULTS.items(): if state_key not in st.session_state: @@ -829,6 +840,15 @@ def _clear_requirements_draft() -> None: st.session_state["requirements_input"] = st.session_state["source_text"] +def _junit_artifact_upload_key() -> str: + version = int(st.session_state["junit_artifact_upload_version"]) + return ( + "junit_artifact_upload" + if version == 0 + else f"junit_artifact_upload_{version}" + ) + + def _criterion_detail_draft_pending() -> bool: runtime_text_keys = ( "runtime_artifact_reference", @@ -838,7 +858,13 @@ def _criterion_detail_draft_pending() -> bool: "runtime_reviewer", "runtime_limitations", ) - return any( + junit_pending = bool( + st.session_state.get(_junit_artifact_upload_key()) + or str(st.session_state.get("junit_importer", "")).strip() + or str(st.session_state.get("junit_limitations", "")).strip() + or st.session_state.get("junit_mapping_scopes", []) + ) + return junit_pending or any( bool(str(st.session_state.get(key, ""))) for key in runtime_text_keys ) or ( st.session_state.get("runtime_evidence_level", EvidenceLevel.E3) @@ -870,11 +896,21 @@ def _clear_resolution_draft() -> None: st.session_state.pop("manual_evidence_level", None) +def _clear_junit_import_draft() -> None: + st.session_state.pop(_junit_artifact_upload_key(), None) + st.session_state["junit_artifact_upload_version"] += 1 + st.session_state["junit_importer"] = "" + st.session_state["junit_limitations"] = "" + st.session_state["junit_mapping_scopes"] = [] + st.session_state["junit_mapping_artifact_sha256"] = None + + def _clear_criterion_detail_drafts() -> bool: """Clear unsaved target-specific inputs and report whether any draft existed.""" had_pending_input = _criterion_detail_draft_pending() _clear_runtime_evidence_draft() _clear_resolution_draft() + _clear_junit_import_draft() return had_pending_input @@ -988,6 +1024,8 @@ def _render_ingestion_limitations(source: PullRequestSnapshot | Review | None) - st.session_state["delete_saved_review_confirmed"] = False if st.session_state.pop("runtime_evidence_form_reset_pending", False): _clear_runtime_evidence_draft() +if st.session_state.pop("junit_import_form_reset_pending", False): + _clear_junit_import_draft() if st.session_state.pop("resolution_form_reset_pending", False): _clear_resolution_draft() if st.session_state.pop("criteria_draft_reset_pending", False): @@ -2052,6 +2090,32 @@ def _render_ingestion_limitations(source: PullRequestSnapshot | Review | None) - _render_comparison_reference( "Current candidate", evidence_change.current ) + if comparison.junit_import_changes: + st.markdown("**Imported external test result changes**") + st.caption(JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION) + st.caption("No prior decision was carried forward.") + for junit_change in comparison.junit_import_changes: + with st.container(border=True): + st.text( + "Imported test result: " + f"{_status_label(junit_change.kind.value)}" + ) + st.code(junit_change.artifact_sha256) + for label, reference in ( + ("Previous import", junit_change.previous), + ("Current import", junit_change.current), + ): + if reference is None: + continue + st.text( + f"{label}: {reference.import_id} · asserted importer: " + f"{reference.asserted_importer}" + ) + for mapping in reference.mappings: + st.text( + f"{mapping.criterion_id}: " + + ", ".join(mapping.test_case_ids) + ) if comparison.changed_finding_statuses: st.markdown("**Changed criterion findings**") for change in comparison.changed_finding_statuses: @@ -2084,7 +2148,7 @@ def _render_ingestion_limitations(source: PullRequestSnapshot | Review | None) - st.warning( "Prior decisions must be revisited for: " + ", ".join(comparison.criteria_requiring_decision_review) - + ". ScopeProof never carries acceptance to a changed head." + + ". ScopeProof does not carry a prior decision forward automatically." ) st.caption( "Ruleset changed between reviews." @@ -2415,6 +2479,219 @@ def _render_ingestion_limitations(source: PullRequestSnapshot | Review | None) - ) st.rerun() + junit_import_save_notice = st.session_state.pop( + "junit_import_save_notice", None + ) + if junit_import_save_notice is not None: + st.success(junit_import_save_notice) + + with st.expander("Import external JUnit results", expanded=False): + st.caption(JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION) + st.caption( + "ScopeProof reads bounded local XML bytes only. It does not run tests, " + "execute target-repository code, follow artifact references, or treat the " + "import as E1, E2, E3, E4, correctness, or acceptance." + ) + exact_head_ready = bool( + len(bundle.review.head_sha) == 40 + and all(character in "0123456789abcdef" for character in bundle.review.head_sha) + ) + if not exact_head_ready: + st.caption( + "An exact 40-character reviewed head is required before import." + ) + uploaded_junit = st.file_uploader( + "Local JUnit XML artifact", + type=["xml"], + accept_multiple_files=False, + key=_junit_artifact_upload_key(), + max_upload_size=1, + ) + junit_importer = st.text_input( + "Asserted JUnit importer (required)", key="junit_importer" + ) + junit_limitations = st.text_area( + "Additional JUnit limitations (optional; one per line)", + key="junit_limitations", + ) + parsed_junit = None + available_scope_ids: list[str] = [] + scope_labels: dict[str, str] = {} + if uploaded_junit is not None: + if uploaded_junit.size > MAX_JUNIT_BYTES: + st.error( + "JUnit artifact could not be inspected. It exceeds the local " + "import byte limit and remains unsaved." + ) + else: + try: + parsed_junit = parse_junit_artifact(uploaded_junit.getvalue()) + except (JUnitImportError, TypeError, ValueError): + st.error( + "JUnit artifact could not be inspected. It is malformed, " + "unsafe, unsupported, or over a bounded parser limit; the " + "review remains unchanged." + ) + else: + previous_mapping_digest = st.session_state.get( + "junit_mapping_artifact_sha256" + ) + if ( + previous_mapping_digest is not None + and previous_mapping_digest + != parsed_junit.artifact_sha256 + ): + st.session_state["junit_mapping_scopes"] = [] + st.session_state["junit_mapping_artifact_sha256"] = ( + parsed_junit.artifact_sha256 + ) + st.caption("Sanitized JUnit preview") + st.text( + "Computed results: " + f"{parsed_junit.totals.total} total · " + f"{parsed_junit.totals.passed} passed · " + f"{parsed_junit.totals.failures} failed · " + f"{parsed_junit.totals.errors} errors · " + f"{parsed_junit.totals.skipped} skipped" + ) + for suite in parsed_junit.suites: + available_scope_ids.append(suite.suite_id) + scope_labels[suite.suite_id] = ( + f"{suite.suite_id} · suite · {suite.suite_name}" + ) + st.text(scope_labels[suite.suite_id]) + for case in suite.test_cases: + available_scope_ids.append(case.test_case_id) + scope_labels[case.test_case_id] = ( + f"{case.test_case_id} · {case.status.value} · " + f"{case.test_name}" + ) + st.text(scope_labels[case.test_case_id]) + if parsed_junit.parser_warnings: + st.caption("Parser warnings") + for warning in parsed_junit.parser_warnings: + st.text(warning) + junit_mapping_scopes = st.multiselect( + "Map JUnit scopes to the selected criterion", + options=available_scope_ids, + format_func=lambda scope_id: scope_labels.get(scope_id, scope_id), + key="junit_mapping_scopes", + disabled=parsed_junit is None, + ) + st.caption( + f"Selected mapping target: {selected_id}. ScopeProof never infers this " + "relationship from test names." + ) + junit_import_candidate = None + junit_import_draft_present = bool( + exact_head_ready + and review_state is not None + and parsed_junit is not None + and junit_importer.strip() + and junit_mapping_scopes + ) + if junit_import_draft_present: + assert review_state is not None + assert uploaded_junit is not None + try: + junit_import_candidate = build_junit_evidence_import( + review_state, + uploaded_junit.getvalue(), + [ + JUnitMappingSelection( + scope_id=scope_id, + criterion_id=selected_id, + ) + for scope_id in junit_mapping_scopes + ], + importer=junit_importer, + limitations=[ + line.strip() + for line in junit_limitations.splitlines() + if line.strip() + ], + ) + append_junit_evidence_import( + review_state, junit_import_candidate + ) + except (JUnitImportError, TypeError, ValueError): + junit_import_candidate = None + if st.button( + "Save imported JUnit results", + key="save_junit_import", + disabled=junit_import_candidate is None, + ): + assert review_state is not None + assert junit_import_candidate is not None + try: + review_state = append_junit_evidence_import( + review_state, junit_import_candidate + ) + except (JUnitImportError, TypeError, ValueError): + st.error( + "JUnit results could not be saved. The artifact, mapping, or " + "active review identity is invalid; the review remains unchanged." + ) + else: + st.session_state["review_state"] = review_state + st.session_state["bundle"] = review_state.bundle + bundle = review_state.bundle + st.session_state["junit_import_form_reset_pending"] = True + st.session_state["junit_import_save_notice"] = ( + "Imported JUnit results appended as external non-gating context." + ) + st.rerun() + + selected_junit_imports = [ + (evidence_import, mapping) + for evidence_import in bundle.junit_evidence_imports + for mapping in evidence_import.criterion_mappings + if mapping.criterion_id == selected_id + ] + with st.expander( + f"Recorded imported JUnit results ({len(selected_junit_imports)})", + expanded=False, + ): + if not selected_junit_imports: + st.caption( + "No external JUnit results are mapped to this criterion." + ) + for evidence_import, mapping in selected_junit_imports: + cases_by_id = { + item.test_case_id: item for item in evidence_import.test_cases + } + with st.container(border=True): + st.caption("External non-gating import ID") + st.code(evidence_import.import_id, language=None) + st.caption("Artifact SHA-256") + st.code(evidence_import.artifact_sha256, language=None) + st.caption("Bound repository and pull request") + st.text( + f"{evidence_import.repository} · PR #{evidence_import.pr_number}" + ) + st.caption("Bound exact head") + st.code(evidence_import.head_sha, language=None) + st.caption("Asserted importer") + st.text(evidence_import.imported_by) + st.caption("Imported at (UTC)") + st.text( + evidence_import.model_dump(mode="json")["imported_at"] + ) + st.caption("Explicitly mapped sanitized test cases") + for case_id in mapping.test_case_ids: + case = cases_by_id[case_id] + st.text( + f"{case.test_case_id} · {case.status.value} · " + f"{case.test_name}" + ) + if evidence_import.parser_warnings: + st.caption("Parser warnings") + for warning in evidence_import.parser_warnings: + st.text(warning) + st.caption("Limitations") + for limitation in evidence_import.limitations: + st.text(limitation) + runtime_evidence_save_notice = st.session_state.pop("runtime_evidence_save_notice", None) if runtime_evidence_save_notice is not None: st.success(runtime_evidence_save_notice) diff --git a/apps/web/launcher.py b/apps/web/launcher.py index 6dcbee3b..c566b075 100644 --- a/apps/web/launcher.py +++ b/apps/web/launcher.py @@ -52,6 +52,7 @@ def main(argv: list[str] | None = None) -> int: f"--server.address={args.host}", f"--server.port={args.port}", f"--server.headless={str(args.headless).lower()}", + "--server.maxUploadSize=2", "--theme.base=dark", "--theme.primaryColor=#d8ff63", "--theme.backgroundColor=#0d0f12", diff --git a/docs/commercialization/stage2-readiness-packet.md b/docs/commercialization/stage2-readiness-packet.md index 78d74c7f..1d689328 100644 --- a/docs/commercialization/stage2-readiness-packet.md +++ b/docs/commercialization/stage2-readiness-packet.md @@ -32,6 +32,17 @@ Stage 2 may improve: Every result remains bounded to its actual evidence. Tests and CI are engineering evidence, not target-repository runtime proof, accessibility conformance, customer use, demand, or adoption. +### Current bounded adapter slice + +The first owner-led adapter accepts one bounded local JUnit XML artifact without executing target +code or following references. It persists only sanitized case names and statuses, a byte digest, +exact review and confirmed-criteria identity, explicit human mappings, an asserted importer, +warnings, and limitations. It is external non-gating context and cannot become E1–E4, observed CI, +runtime verification, a reviewer decision, final acceptance, correctness, or customer validation. +One review retains at most 20 imports, each retained warning or asserted limitation is capped at +1,000 characters, and a new analysis never inherits an earlier criteria revision's imports. +Other adapter families remain unimplemented and require their own bounded design and owner scope. + This stage does not authorize outreach, participant contact, a merge, release, tag, or package publication, R-002 retuning, R-003 generation, billing, accounts, private-repository support, hosted source processing, generic code review, security scanning, automatic fixes, or paid APIs. diff --git a/docs/development-environment.md b/docs/development-environment.md index 227ad2c5..fc7ffde8 100644 --- a/docs/development-environment.md +++ b/docs/development-environment.md @@ -99,4 +99,4 @@ rename; that hostile local-account scenario remains unsupported and must not be ## Known-good UI baseline -The checked-in lock currently resolves Streamlit 1.59.1, which passes ScopeProof's complete AppTest suite. ScopeProof requires Streamlit 1.52 or newer because the workbench relies on click-time deferred download generation to revalidate saved review truth immediately before export. During this work, Streamlit 1.57.0 exposed a testing-interface regression; that observation is why the lock is the reproducible baseline rather than a claim that every version in the supported range behaves identically. CI still installs the newest versions allowed by `pyproject.toml` in the compatibility and verification lanes so future incompatibilities remain visible without a scheduled monitor or notification workflow. +The checked-in lock currently resolves Streamlit 1.59.1, which passes ScopeProof's complete AppTest suite. ScopeProof requires Streamlit 1.53 or newer because 1.53 is the first supported release that combines click-time deferred download generation with the per-widget upload limit used to keep JUnit artifacts at exactly 1 MiB while allowing multipart transport overhead. During this work, Streamlit 1.57.0 exposed a testing-interface regression; that observation is why the lock is the reproducible baseline rather than a claim that every version in the supported range behaves identically. CI still installs the newest versions allowed by `pyproject.toml` in the compatibility and verification lanes so future incompatibilities remain visible without a scheduled monitor or notification workflow. diff --git a/docs/releases/v0.2.3-status-and-next-stages.md b/docs/releases/v0.2.3-status-and-next-stages.md index 553a229b..ee520aba 100644 --- a/docs/releases/v0.2.3-status-and-next-stages.md +++ b/docs/releases/v0.2.3-status-and-next-stages.md @@ -197,6 +197,7 @@ code-review comments, scan security, or automatically fix a PR. | Fail-closed gate | Complete ingestion, current criteria, observed CI, evidence findings, runtime requirements, current decisions, exact runtime links, and a deterministic malformed-input preflight | Duplicate IDs, coverage mismatches, foreign decisions, provenance contradictions, and legacy-unlinked decisions cannot support Ready | | Persistence and privacy | Pydantic-validated version 4 local JSON; deterministic version 1–3 migration without invented runtime links or criteria-source provenance; safe paths, reopen, deletion, no persisted token | Legacy records remain fail-closed until source reconfirmation; local storage is not secure erasure or hosted collaboration | | Re-review comparison | Unchanged, Relocated, Modified, Added, and Removed evidence; changed-first UI, collapsed unchanged references, Markdown/JSON comparison exports, and affected decisions requiring review | No acceptance is silently carried to a changed head or inferred from Unchanged candidates | +| Imported external test results | Bounded local JUnit XML inspection/import in the shared core, CLI, Streamlit, versioned saved records, JSON/Markdown/HTML/CSV exports, mapping-aware comparison, and installed-wheel Chromium round trip | Sanitized, exact-head and criteria-bound external context only; never E1–E4, observed CI, runtime verification, a human decision, final acceptance, correctness, or customer validation | | Exports | Pydantic-backed JSON, Markdown, CSV, and HTML with runtime identity, linked/unlinked state, and exact criteria-source provenance | Exported content is a review record, not certification | | Alpha evidence | One-time outcome capture only from a fully revalidated saved review with matching public-GitHub origin, PR, exact head, criteria, and provenance | Demo, fixture, research, legacy-unknown, mutated, or non-public origins contribute zero qualifying alpha evidence | | Engineering benchmarks | 12-case constructed acceptance benchmark, two-case comparison benchmark, R-001 research case, and frozen 20-case R-002 research baseline | All contribute zero Stage 1 credit | @@ -331,6 +332,14 @@ the selected criterion evidence and controls before the secondary evidence matri Stage 2 follow-up preserves every evidence boundary and keeps optional external research secondary. +The bounded JUnit adapter is the first implemented non-executing evidence adapter. It accepts +bounded local bytes only, requires exact review and confirmed-criteria binding plus explicit human +mapping, persists no raw XML or output/failure bodies, and remains separate from every gate, +runtime, decision, acceptance, and correctness claim. A review retains at most 20 imports, each +retained warning or asserted limitation is capped at 1,000 characters, and a new analysis never +inherits an earlier criteria revision's imports. Other adapter families remain future owner +decisions. + External commercial discovery is optional and separate from owner-led productization. It is not required for Stage 2 and requires separate owner authorization before outreach or participant contact. Stage 2 does not authorize a merge, release, tag, package publication, R-002 retuning, @@ -354,6 +363,15 @@ requirements import, or a commercial license. None is currently authorized. Missing external evidence remains missing; engineering work does not turn into customer evidence. +## Delivered Stage 2 foundations + +- The exact-SHA informational GitHub Check lifecycle is implemented as an opt-in, neutral-only, + trusted-base surface. It is not a required branch-protection check and creates no runtime, + correctness, accessibility, demand, adoption, customer-validation, or Stage 1 evidence. +- The bounded JUnit XML adapter is implemented as local, sanitized, exact-head/provenance-bound, + explicitly mapped, non-gating context. It does not execute tests or target-repository code and + does not generalize to other report types. + ## Prioritized post-release decision candidates These are proposals, not authorized implementation. Each must preserve @@ -363,15 +381,11 @@ target repository code. ### Pilot-critical candidates -1. **Exact-SHA GitHub Check lifecycle:** design one low-noise check bound to the - current PR head, invalidate stale prior conclusions on `synchronize`, and - keep the current Action informational until independent use justifies - promotion. -2. **Non-executing evidence adapters:** design validated import records for - JUnit-style results, coverage summaries, contract reports, build/deployment - records, and externally supplied runtime attestations. Imported data remains - evidence with provenance, not proof of correctness. -3. **Authenticated reviewer identity:** current runtime records now retain +1. **Additional non-executing evidence adapters:** separately design validated import records for + coverage summaries, contract reports, build/deployment records, or externally supplied runtime + attestations. The implemented JUnit adapter does not authorize or generalize these formats; + imported data must remain provenance-bound context, not proof of correctness. +2. **Authenticated reviewer identity:** current runtime records now retain review-scoped identity and attribution, but authentication, optional expiry, and externally signed attestations remain future decisions without accounts, silent overrides, or automatic approval. diff --git a/docs/superpowers/plans/2026-08-20-junit-evidence-adapter.md b/docs/superpowers/plans/2026-08-20-junit-evidence-adapter.md new file mode 100644 index 00000000..ab1d107d --- /dev/null +++ b/docs/superpowers/plans/2026-08-20-junit-evidence-adapter.md @@ -0,0 +1,649 @@ +# Bounded JUnit Evidence Adapter Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Import bounded local JUnit XML into a validated, exact-head, +non-gating evidence record that is visible in CLI, Streamlit, exports, and +comparison without executing target-repository code. + +**Architecture:** A new core importer accepts bytes and produces sanitized +Pydantic data. A lifecycle transition appends a versioned `JUnitEvidenceImport` +to the active `ReviewBundle` while proving all gate and human-decision inputs +are unchanged. CLI and Streamlit call the same importer and lifecycle functions; +storage and exports continue to revalidate complete review state. + +**Tech Stack:** Python 3.11+, Pydantic 2, standard-library +`xml.etree.ElementTree`, argparse, Streamlit, pytest, Playwright, Ruff, uv. + +**Spec:** `docs/superpowers/specs/2026-08-20-junit-evidence-adapter-design.md` + +## Global Constraints + +- Accept at most 1,048,576 artifact bytes, 100 suites, 5,000 cases, and 20,000 XML elements. +- Accept UTF-8 only and reject DTD, entity, XInclude, remote, path, or executable processing. +- Persist no raw XML, failure bodies, stdout, stderr, properties, commands, paths, URLs, or attachments from the artifact. +- Require repository, PR, exact 40-character lowercase hexadecimal head, criteria revision, normalized criteria digest, and exact criteria-source provenance binding. +- Imported results are never E1, E2, E3, E4, CI, a human resolution, final acceptance, or a deterministic gate input. +- Every saved or exported object is Pydantic-revalidated; failed imports do not mutate state. +- Full and comparison JSON preserve a fixed typed boundary for external source, + non-gating effect, non-execution, digest scope, asserted identity, and mapping + limitations. +- Browser persistence checks reopen the exact review ID reported by the save notice. +- Preserve `.coverage 2` exactly and never stage, package, modify, rename, or delete it. +- Keep Stage 1 closed at 0/5, 0/3, 0/3, 0/3, and 0/2; keep Stage 2 active; do not begin Stage 3. + +--- + +### Task 1: Persisted import contracts + +**Files:** +- Modify: `scopeproof_core/schemas/models.py` +- Modify: `scopeproof_core/schemas/__init__.py` +- Create: `tests/schemas/test_junit_evidence_import.py` +- Modify: `tests/schemas/test_review_bundle_integrity.py` + +**Interfaces:** +- Consumes: `CriteriaSourceProvenance`, `ReviewBundle`, `ReviewState`, `normalized_criteria_sha256`. +- Produces: `JUnitCaseStatus`, `JUnitCaseResult`, `JUnitResultTotals`, + `JUnitCriterionMapping`, `JUnitEvidenceImport`, and + `JUnitImportMutationMetadata`. + +- [ ] **Step 1: Write the failing contract tests** + +Create literal fixtures with exact head `"a" * 40` and assert: + +```python +def test_junit_import_requires_exact_review_and_criteria_identity() -> None: + payload = valid_junit_import_payload() + record = JUnitEvidenceImport.model_validate(payload) + assert record.schema_version == "junit-import-v1" + assert record.totals.total == 2 + assert record.criterion_mappings[0].test_case_ids == [ + "suite-0001-case-0001" + ] + + +def test_review_bundle_rejects_junit_import_from_another_head() -> None: + bundle = exact_head_bundle() + bundle.junit_evidence_imports = [junit_import(head_sha="b" * 40)] + with pytest.raises(ValidationError, match="JUnit import identity"): + ReviewBundle.model_validate(bundle.model_dump(mode="python")) +``` + +Add parametrized failures for partial identity, non-exact SHA, unknown criterion, +unknown case ID, duplicate case IDs, inconsistent totals, duplicate import IDs, +duplicate artifact digests, mismatched revision/digest/provenance, blank asserted +importer, naive timestamp, blank warning/limitation, and extra fields. + +- [ ] **Step 2: Run tests and confirm the missing-contract failure** + +Run: + +```bash +uv run pytest tests/schemas/test_junit_evidence_import.py tests/schemas/test_review_bundle_integrity.py -q +``` + +Expected: collection fails because the new types and bundle field do not exist. + +- [ ] **Step 3: Add minimal strict Pydantic models and bundle validation** + +Use `ConfigDict(extra="forbid", frozen=True)` on nested import records, exact +digest/SHA patterns, timezone normalization, sorted-unique validators, totals +consistency, and cross-reference checks in `ReviewBundle.validate_cross_references`. +Add: + +```python +junit_evidence_imports: list[JUnitEvidenceImport] = Field(default_factory=list) +``` + +Validate all import identity/provenance fields against the owning bundle and +require unique import IDs and artifact digests. + +- [ ] **Step 4: Run the schema tests to green** + +Run the Step 2 command. Expected: all selected tests pass. + +- [ ] **Step 5: Commit the contracts** + +```bash +git add scopeproof_core/schemas/models.py scopeproof_core/schemas/__init__.py tests/schemas/test_junit_evidence_import.py tests/schemas/test_review_bundle_integrity.py +git commit -m "feat: define imported JUnit evidence contracts" +``` + +### Task 2: Bounded bytes-only parser and explicit mapping builder + +**Files:** +- Create: `scopeproof_core/importers/__init__.py` +- Create: `scopeproof_core/importers/junit.py` +- Create: `tests/importers/test_junit.py` + +**Interfaces:** +- Consumes: `ReviewState`, `JUnitEvidenceImport`, confirmed criteria and provenance. +- Produces: + +```python +def parse_junit_artifact(artifact_bytes: bytes) -> ParsedJUnitArtifact: ... +def build_junit_evidence_import( + state: ReviewState, + artifact_bytes: bytes, + selections: list[JUnitMappingSelection], + *, + importer: str, + limitations: list[str] | None = None, + imported_at: datetime | None = None, + import_id: str | None = None, +) -> JUnitEvidenceImport: ... +``` + +- [ ] **Step 1: Write parser success and safety failures first** + +Use literal byte fixtures and assert sanitized output: + +```python +def test_parser_returns_sanitized_cases_and_ignores_output_bodies() -> None: + parsed = parse_junit_artifact( + b'' + b'secret output' + ) + assert parsed.totals.model_dump() == { + "total": 1, "passed": 1, "failures": 0, "errors": 0, "skipped": 0 + } + assert parsed.suites[0].cases[0].test_case_id == "suite-0001-case-0001" + assert "secret output" not in parsed.model_dump_json() +``` + +Parametrize exact-boundary acceptance and one-over rejection for bytes, suites, +cases, and elements. Add failures for non-bytes, non-UTF-8, non-UTF-8 XML +declarations, DTD, internal/external entities, non-declaration processing +instructions, XInclude namespace elements, unsupported roots, nested suites, +cases outside suites, missing test names, multiple result markers, and malformed +XML. Assert error messages contain no supplied secret strings. + +- [ ] **Step 2: Verify the parser tests fail because the module is absent** + +```bash +uv run pytest tests/importers/test_junit.py -q +``` + +Expected: import/collection failure for `scopeproof_core.importers.junit`. + +- [ ] **Step 3: Implement the minimal bounded parser** + +Check byte and encoding boundaries before `ElementTree.fromstring`. Reject +forbidden constructs before parsing, then enforce local tag names, direct-child +structure, element/suite/case limits, bounded names, and at most one direct +result marker. Compute totals from cases. Persist only sanitized names/statuses +and deterministic warnings for declared-count mismatches or discarded output. + +- [ ] **Step 4: Add failing builder tests** + +Assert suite and case selectors expand deterministically and require explicit +valid mappings: + +```python +def test_builder_expands_explicit_suite_mapping_and_binds_review() -> None: + state = exact_head_state() + record = build_junit_evidence_import( + state, + TWO_CASE_XML, + [JUnitMappingSelection(scope_id="suite-0001", criterion_id="AC-01")], + importer="QA owner", + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + import_id="import-001", + ) + assert record.artifact_sha256 == sha256(TWO_CASE_XML).hexdigest() + assert record.criterion_mappings[0].test_case_ids == [ + "suite-0001-case-0001", "suite-0001-case-0002" + ] + assert record.repository == state.review.repository + assert record.head_sha == state.review.head_sha +``` + +Add failures for no active analysis, unconfirmed/missing provenance, non-exact +head, blank importer, no selections, unknown scope, unknown criterion, empty +suite mapping, and blank limitations. + +- [ ] **Step 5: Implement the minimal builder and run tests to green** + +Build canonical mappings from parsed IDs, copy exact review/criteria identity, +add the three fixed evidence-boundary limitations, append normalized user +limitations, and return a fully revalidated frozen envelope. + +Run: + +```bash +uv run pytest tests/importers/test_junit.py tests/schemas/test_junit_evidence_import.py -q +uv run ruff check scopeproof_core/importers tests/importers +``` + +- [ ] **Step 6: Commit the importer** + +```bash +git add scopeproof_core/importers scopeproof_core/schemas tests/importers tests/schemas/test_junit_evidence_import.py +git commit -m "feat: parse bounded JUnit evidence bytes" +``` + +### Task 3: Atomic lifecycle and saved-review behavior + +**Files:** +- Modify: `scopeproof_core/reviews/lifecycle.py` +- Modify: `scopeproof_core/reviews/__init__.py` +- Modify: `tests/reviews/test_lifecycle.py` +- Modify: `tests/storage/test_json_store.py` + +**Interfaces:** +- Consumes: `build_junit_evidence_import`, `validated_review_state`. +- Produces: + +```python +def append_junit_evidence_import( + state: ReviewState, evidence_import: JUnitEvidenceImport +) -> ReviewState: ... +``` + +- [ ] **Step 1: Write lifecycle red tests** + +Test successful deep-copy append and literal equality of the pre/post gate, +findings, resolutions, runtime evidence, final acceptance, and resolution +history. Add atomic failures for changed repository/PR/head, changed criteria +revision/digest/provenance, unknown criterion/case, duplicate ID/digest, mutated +input models, and missing active bundle. + +```python +def test_junit_import_append_is_non_gating_and_does_not_alias_input() -> None: + state = exact_head_state() + record = junit_import_for(state) + updated = append_junit_evidence_import(state, record) + assert updated.bundle.junit_evidence_imports == [record] + assert updated.bundle.gate == state.bundle.gate + assert updated.bundle.resolutions == state.bundle.resolutions + assert state.bundle.junit_evidence_imports == [] +``` + +- [ ] **Step 2: Run the lifecycle slice and observe the missing transition** + +```bash +uv run pytest tests/reviews/test_lifecycle.py -q -k junit +``` + +Expected: failure because `append_junit_evidence_import` is absent. + +- [ ] **Step 3: Implement the transition and run to green** + +Revalidate state and record, verify all active relationships, append a deep +copy, revalidate the resulting state, and explicitly compare unchanged gate and +human/runtime fields before returning. + +- [ ] **Step 4: Add storage red-green coverage** + +Persist and reopen a state containing an import; assert record version 4, +nested `junit-import-v1`, exact equality, no raw XML, and successful exports. +Downgrade a fixture by deleting `junit_evidence_imports`; assert it reopens as +an empty list. Use `JsonReviewStore.mutate` with a failing transition and assert +the file bytes and fingerprint are unchanged. + +Run: + +```bash +uv run pytest tests/reviews/test_lifecycle.py tests/storage/test_json_store.py -q -k 'junit or imported_test' +``` + +- [ ] **Step 5: Commit lifecycle and storage behavior** + +```bash +git add scopeproof_core/reviews tests/reviews/test_lifecycle.py tests/storage/test_json_store.py +git commit -m "feat: append imported test evidence atomically" +``` + +### Task 4: Shared CLI inspection and import + +**Files:** +- Modify: `scopeproof_core/cli.py` +- Modify: `tests/cli/test_cli.py` + +**Interfaces:** +- Consumes: `parse_junit_artifact`, `build_junit_evidence_import`, + `append_junit_evidence_import`, `JsonReviewStore.mutate`. +- Produces: `inspect-junit` and `import-junit` CLI commands. + +- [ ] **Step 1: Write CLI failing tests** + +Create a strict `junit-mapping-v1` JSON fixture and assert `inspect-junit` +outputs sanitized JSON with no raw output. Assert `import-junit` appends one +record and emits validated metadata. Bind the mapping to the exact inspected +artifact SHA-256. Test malformed mapping, digest mismatch, wrong schema, +extra fields, unsafe XML, stale review identity, duplicate artifact, missing +file, and store failure; each failed command must leave record bytes unchanged. + +```python +def test_import_junit_persists_one_non_gating_record(tmp_path: Path, capsys) -> None: + review_id = save_exact_head_review(tmp_path) + result = main([ + "import-junit", review_id, str(write_xml(tmp_path)), + "--mapping", str(write_mapping(tmp_path)), + "--importer", "QA owner", "--storage-dir", str(tmp_path / "reviews"), + ]) + assert result == 0 + loaded = JsonReviewStore(tmp_path / "reviews").load(review_id) + assert len(loaded.bundle.junit_evidence_imports) == 1 + assert loaded.bundle.gate.verdict is GateVerdict.NEEDS_REVIEW +``` + +- [ ] **Step 2: Run CLI tests and confirm parser rejects the new commands** + +```bash +uv run pytest tests/cli/test_cli.py -q -k junit +``` + +Expected: argparse rejects `inspect-junit` and `import-junit`. + +- [ ] **Step 3: Implement strict mapping parsing and command handlers** + +Define a strict Pydantic `JUnitMappingDocument` with required literal +`junit-mapping-v1` and required artifact SHA-256. Read only bounded regular +local artifact files in binary mode in the CLI adapter and reject a digest mismatch. Use +`JsonReviewStore.mutate` for the persisted transition. Print only sanitized +Pydantic JSON metadata with an explicit external non-gating boundary; never +print raw XML, supplied invalid values, or local paths from XML. + +- [ ] **Step 4: Run CLI and adjacent storage tests to green** + +```bash +uv run pytest tests/cli/test_cli.py tests/storage/test_json_store.py -q -k 'junit or import_junit or inspect_junit' +uv run ruff check scopeproof_core/cli.py tests/cli/test_cli.py +``` + +- [ ] **Step 5: Commit CLI parity** + +```bash +git add scopeproof_core/cli.py tests/cli/test_cli.py +git commit -m "feat: expose bounded JUnit imports in CLI" +``` + +### Task 5: Streamlit import, preview, and reopen flow + +**Files:** +- Modify: `apps/web/app.py` +- Modify: `apps/web/view_models.py` only if presentation projection is reusable. +- Modify: `tests/apps/test_streamlit_app.py` +- Modify: `tests/apps/test_view_models.py` only if `view_models.py` changes. +- Modify: `tests/browser/test_packaged_workbench.py` + +**Interfaces:** +- Consumes: shared core parser, builder, lifecycle transition. +- Produces: one separate selected-criterion JUnit import expander and recorded + import display. + +- [ ] **Step 1: Write AppTest failures for the new workflow** + +Start from `analyzed_exact_head_standard_demo`. Upload a one-suite fixture, +enter importer, preview sanitized scope IDs, choose `suite-0001`, save to the +selected criterion, and assert the import is separate from E3/E4 controls. Test +disabled save, parser error, invalid scope, oversized upload, failed transition, +form reset, save/reopen display, and unchanged gate/human/runtime state. + +Assert hostile suite/test/importer text appears only through inert Streamlit +text/code elements, never Markdown. Assert ignored XML output is absent from all +visible elements and session-state models. + +Also add the installed-wheel exact-head import/save/reopen/download browser case +described in Task 7 before changing the UI. + +- [ ] **Step 2: Run AppTest and confirm the controls are absent** + +```bash +uv run pytest tests/apps/test_streamlit_app.py -q -k junit +uv run pytest -m browser tests/browser/test_packaged_workbench.py::test_installed_wheel_junit_import_round_trip -q +``` + +Expected: both commands fail because the new uploader/controls are absent. + +- [ ] **Step 3: Implement the minimal shared-core UI** + +Place `Import external JUnit results` before the E3/E4 expander. Check uploaded +size before reading bytes. Preview with `parse_junit_artifact`; use the stable +scope IDs as multiselect options and the currently selected criterion as the +explicit mapping target. Apply `append_junit_evidence_import` only on Save. +Render saved imports with `st.text`, `st.code`, and existing inert reference +helpers, never raw Markdown from artifact values. + +- [ ] **Step 4: Run AppTest and nearby browser-independent UI tests** + +```bash +uv run pytest tests/apps/test_streamlit_app.py tests/apps/test_web_app.py tests/apps/test_view_models.py -q +uv run pytest -m browser tests/browser/test_packaged_workbench.py::test_installed_wheel_junit_import_round_trip -q +uv run ruff check apps/web tests/apps +``` + +- [ ] **Step 5: Commit the workbench flow** + +```bash +git add apps/web/app.py apps/web/view_models.py tests/apps/test_streamlit_app.py tests/apps/test_view_models.py tests/browser/test_packaged_workbench.py +git commit -m "feat: add JUnit import review workflow" +``` + +Stage only files that actually changed. + +### Task 6: Safe exports and comparison projection + +**Files:** +- Modify: `scopeproof_core/reporting/exporters.py` +- Modify: `scopeproof_core/reviews/comparison.py` +- Modify: `tests/reporting/test_exporters.py` +- Modify: `tests/reporting/test_html_export.py` +- Modify: `tests/reporting/test_comparison_exports.py` +- Modify: `tests/reviews/test_comparison.py` +- Modify: `scopeproof_core/evals/comparison_runner.py` only if the typed output changes require fixture assertions. + +**Interfaces:** +- Consumes: validated `JUnitEvidenceImport` lists. +- Produces: safe Markdown/HTML/CSV sections and typed comparison changes keyed + by artifact digest and mapping signature. + +- [ ] **Step 1: Write export red tests** + +Assert JSON, Markdown, HTML, and CSV include the artifact digest, exact head, +asserted importer, mapped criterion/case/status, warnings, and limitations. +Assert they exclude a sentinel raw XML string, ignored stdout, failure body, and +local artifact path. Use formula/HTML/Markdown payloads in all imported names +and assert CSV neutralization plus inert Markdown/HTML escaping. + +- [ ] **Step 2: Run export tests and observe missing import sections** + +```bash +uv run pytest tests/reporting/test_exporters.py tests/reporting/test_html_export.py -q -k junit +``` + +- [ ] **Step 3: Implement safe export projections and run to green** + +Reuse `_render_markdown_code`, `html.escape`, and `_csv_text`. Do not serialize +the source XML or any discarded parser field. + +- [ ] **Step 4: Write comparison red tests** + +Assert unchanged, added, removed, and same-digest mapping-modified projections. +Tamper imported repository/head/provenance/mapping data and assert comparison +fails before output. Assert changed imported context adds only previously +resolved affected criteria to `criteria_requiring_decision_review`, copies no +resolution, and changes neither gate. + +- [ ] **Step 5: Implement typed import comparison and run the full slice** + +```bash +uv run pytest tests/reviews/test_comparison.py tests/reporting/test_comparison_exports.py tests/evals/test_comparison_runner.py -q +uv run ruff check scopeproof_core/reporting scopeproof_core/reviews tests/reporting tests/reviews/test_comparison.py +``` + +- [ ] **Step 6: Commit exports and comparison** + +```bash +git add scopeproof_core/reporting/exporters.py scopeproof_core/reviews/comparison.py scopeproof_core/evals/comparison_runner.py tests/reporting tests/reviews/test_comparison.py tests/evals/test_comparison_runner.py +git commit -m "feat: export and compare imported test evidence" +``` + +Stage only files that actually changed. + +### Task 7: Installed-wheel browser completion and product documentation + +**Files:** +- Modify: `tests/browser/test_packaged_workbench.py` +- Modify: `README.md` +- Modify: `ROADMAP.md` +- Modify: `CHANGELOG.md` +- Modify: `docs/releases/v0.2.3-status-and-next-stages.md` +- Modify: `docs/commercialization/stage2-readiness-packet.md` +- Create or modify focused repository contracts in `tests/test_repository_contracts.py` only for authoritative machine-checkable status relationships. + +**Interfaces:** +- Consumes: installed wheel, local review store, browser workbench. +- Produces: exact-head import/save/reopen/download browser proof and truthful + product status. + +- [ ] **Step 1: Complete the installed-wheel browser regression started in Task 5** + +Confirm the failing test written in Task 5 builds an exact-head synthetic saved +review with installed ScopeProof code in the temporary HOME. It must launch the +installed wheel with loopback-only networking, reopen the review, upload a +one-suite XML file, map `suite-0001`, save, verify the non-gating label, +save/reopen, and download JSON and Markdown. Both downloads must contain the +SHA-256 digest and exclude raw XML/output sentinels. + +- [ ] **Step 2: Run the browser regression and observe the absent UI** + +```bash +uv run pytest -m browser tests/browser/test_packaged_workbench.py -q +``` + +Expected after Tasks 1–6: all browser cases pass with zero external requests, +console errors, or page errors. + +- [ ] **Step 3: Update product truth documents** + +Document the adapter as externally supplied, non-gating engineering context. +Keep version `0.2.4.dev0`, latest release v0.2.3, Stage 1 counts zero, Stage 2 +active, Stage 3 gated, reviewer identity asserted, and platform/accessibility +limits unchanged. Move the exact-SHA informational Check lifecycle from the +future-candidate list to delivered Stage 2 work because PR #196 already merged. + +- [ ] **Step 4: Add behavior-level documentation contracts only where needed** + +If an authoritative status relationship needs a contract, parse the exact +section and assert the state relationship rather than merely searching the +whole document for words. Do not test ordinary explanatory prose. + +- [ ] **Step 5: Run documentation, browser, and repository-contract checks** + +```bash +uv run pytest tests/test_repository_contracts.py -q +uv run pytest -m browser tests/browser/test_packaged_workbench.py -q +uv run ruff check . +git diff --check +``` + +- [ ] **Step 6: Commit browser and documentation evidence** + +```bash +git add tests/browser/test_packaged_workbench.py README.md ROADMAP.md CHANGELOG.md docs/releases/v0.2.3-status-and-next-stages.md docs/commercialization/stage2-readiness-packet.md tests/test_repository_contracts.py +git commit -m "docs: record bounded external test imports" +``` + +Stage only files that actually changed. + +### Task 8: Complete verification and exact-head review + +**Files:** +- Modify only confirmed defects found by verification, each with a failing regression first. + +**Interfaces:** +- Consumes: final feature branch. +- Produces: reproducible engineering evidence and a reviewed exact head. + +- [ ] **Step 1: Run formatting, complete suite, and coverage** + +```bash +uv run ruff check . +uv run python -m pytest --cov=scopeproof_core --cov=apps --cov-report=term-missing:skip-covered --cov-fail-under=95 -q +uv run pytest tests/test_repository_contracts.py -q +git diff --check +``` + +- [ ] **Step 2: Run deterministic benchmarks** + +```bash +uv run scopeproof-eval +uv run scopeproof-compare-eval +``` + +Require every mismatch and must-have False Ready count to equal zero. + +- [ ] **Step 3: Build and compare two wheels** + +Build in two clean temporary output directories with the repository's standard +build command, compare SHA-256 values byte-for-byte, inspect archive entries, +and assert `.coverage`, `.scopeproof`, `.superpowers`, worktrees, raw test +artifacts, and local state are absent. + +- [ ] **Step 4: Verify clean installation and runtime lanes** + +Install the wheel into a clean environment, run dependency validation, compare +source/installed versions, run both CLIs and installed benchmarks, start exact +loopback health and confirm listener shutdown, and rerun installed-wheel +Chromium. Run the supported local Python lanes available on the host; hosted +Python and Windows conclusions come from the PR checks. + +- [ ] **Step 5: Audit the exact diff and commits** + +```bash +git status --short +git diff --check origin/main...HEAD +git log --oneline --decorate origin/main..HEAD +git diff --stat origin/main...HEAD +``` + +Confirm only named in-scope files changed and the original checkout still has +the exact `.coverage 2` SHA-256, size, mtime, and inode. + +- [ ] **Step 6: Obtain an independent read-only review** + +Review exact range `f586d90b72a14fd19d5a0add01f3d05532a88955..HEAD` +against the approved spec. Require explicit Critical/Important/Minor findings, +read-only diff inspection, and a merge-readiness verdict. For every actionable +Critical or Important issue, write a failing regression, implement the minimal +fix, rerun affected and full checks, commit intentionally, and repeat exact-head +review until none remain. + +### Task 9: Publish the ready PR and monitor hosted checks + +**Files:** +- No source files unless a hosted check exposes a confirmed in-scope defect. + +**Interfaces:** +- Consumes: independently reviewed exact head. +- Produces: ready PR `feat: import bounded external test evidence`. + +- [ ] **Step 1: Push the exact branch** + +```bash +git push -u origin codex/junit-evidence-adapter +``` + +- [ ] **Step 2: Open a ready-for-review PR against `main`** + +The description must summarize the evidence taxonomy, parser bounds, explicit +mapping, atomic lifecycle, CLI/UI parity, safe exports/comparison, browser proof, +verification, independent review, unsupported environments, preserved +`.coverage 2`, Stage 1 zero counts, and the no-release/no-Stage-3 boundary. + +- [ ] **Step 3: Monitor every available check to a terminal conclusion** + +Diagnose failures systematically. Fix only confirmed in-scope defects with a +failing regression first, commit and push the repair, repeat independent review +for the changed exact head, and wait for replacement checks. + +- [ ] **Step 4: Final handoff without merging** + +Report PR URL, exact base/head/tree, commits, complete verification results, +hosted conclusions, review findings, unsupported environments, preserved local +artifact proof, Stage 1 counts, and the owner decision: merge or hold/request +changes. Do not merge, release, tag, publish, conduct outreach, generate R-003, +retune R-002, or start Stage 3/4. diff --git a/docs/superpowers/specs/2026-08-20-junit-evidence-adapter-design.md b/docs/superpowers/specs/2026-08-20-junit-evidence-adapter-design.md new file mode 100644 index 00000000..76c5aaef --- /dev/null +++ b/docs/superpowers/specs/2026-08-20-junit-evidence-adapter-design.md @@ -0,0 +1,263 @@ +# Bounded JUnit Evidence Adapter Design + +**Date:** 2026-08-20 +**Stage:** Owner-led Stage 2 productization +**Status:** Approved for implementation by the owner +**Target branch:** `codex/junit-evidence-adapter` + +## Objective + +Add one local, non-executing adapter that imports bounded JUnit-style XML as +provenance-bound external test-result context. The adapter helps a reviewer +inspect supplied test results without converting them into correctness, +runtime-verification, acceptance, CI, or customer-validation claims. + +The adapter never runs target-repository code, follows references, fetches a +URL, opens an artifact path found inside XML, or persists raw XML. Failed +parsing, mapping, validation, or storage leaves the saved review unchanged. + +## Evidence taxonomy + +Imported JUnit results are a new orthogonal record type named +`JUnitEvidenceImport`. They are not `EvidenceItem` candidates and therefore do +not become E1 or E2. They are not `RuntimeEvidence` and therefore do not become +E3 or E4. They are not `CIObservation`, `HumanResolution`, or final acceptance. + +The deterministic gate continues to consume only the existing review, +criteria, findings, and current human resolutions. Adding an import cannot +create, rescue, or justify a Ready verdict. A previously valid verdict may be +displayed beside an import only because the pre-existing gate inputs still +justify it; the import remains explicitly non-gating. + +Every product surface labels the record as externally supplied and states: + +- ScopeProof did not execute the tests or target-repository code. +- The artifact digest proves only which bytes were imported. +- The importer identity is asserted, not authenticated. +- Explicit human mapping is organizational context, not proof that a criterion + passed. + +## Architecture + +### Persisted schemas + +`scopeproof_core/schemas/models.py` owns the persisted Pydantic types: + +- `JUnitCaseStatus`: `passed`, `failure`, `error`, or `skipped`. +- `JUnitCaseResult`: stable document-order IDs, bounded suite/class/test names, + and one status. Failure bodies, stdout, stderr, properties, commands, paths, + URLs, and attachments are never persisted. Path- or URL-like names are + replaced with deterministic redacted labels. +- `JUnitResultTotals`: total, passed, failure, error, and skipped counts whose + sum must equal total. +- `JUnitCriterionMapping`: one confirmed criterion ID and a sorted unique list + of stable test-case IDs. +- `JUnitEvidenceImport`: a frozen `junit-import-v1` envelope containing import + ID, review identity, exact head, criteria revision and source provenance, + artifact digest, sanitized case results, totals, explicit mappings, asserted + importer metadata, warnings, limitations, and a required fixed typed trust + boundary. The boundary is preserved in full JSON and comparison JSON so + machine consumers receive the same non-gating semantics as human-readable + surfaces. +- `JUnitImportMutationMetadata`: validated CLI result metadata with an explicit + `externally_supplied_non_gating` boundary beside any unchanged review verdict. + +`ReviewBundle.junit_evidence_imports` is an append-only list with an empty +default so historical records remain readable without inventing imports. +Bundle validation enforces: + +- repository, PR, and exact head match the owning review; +- criteria revision and normalized digest match the active bundle; +- copied source provenance exactly matches the review; +- every mapped criterion and case exists; +- import IDs and artifact digests are unique; +- imports never appear in static evidence, runtime evidence, resolutions, or + gate cross-references. + +The outer local-review record stays at version 4 because an absent +`junit_evidence_imports` field has one unambiguous meaning: no imported JUnit +record. The nested import and CLI mapping document each require their explicit +version discriminator; neither silently defaults an unversioned payload to v1. + +### Parser and import service + +Create `scopeproof_core/importers/junit.py`. Its public interface is: + +```python +MAX_JUNIT_BYTES = 1_048_576 +MAX_JUNIT_SUITES = 100 +MAX_JUNIT_CASES = 5_000 +MAX_JUNIT_ELEMENTS = 20_000 +MAX_JUNIT_IMPORTS_PER_REVIEW = 20 +MAX_JUNIT_NOTE_LENGTH = 1_000 + +def parse_junit_artifact(artifact_bytes: bytes) -> ParsedJUnitArtifact: ... + +def build_junit_evidence_import( + state: ReviewState, + artifact_bytes: bytes, + selections: list[JUnitMappingSelection], + *, + importer: str, + limitations: list[str] | None = None, + imported_at: datetime | None = None, + import_id: str | None = None, +) -> JUnitEvidenceImport: ... +``` + +`ParsedJUnitArtifact`, `ParsedJUnitSuite`, and `JUnitMappingSelection` are +strict Pydantic boundary types. Scope IDs are deterministic document-order +identifiers: `suite-0001` and `suite-0001-case-0001`. A suite selection expands +to all cases in that suite; a case selection expands to that one case. The +persisted record contains only resolved case-to-criterion mappings, so no +selector needs to be reinterpreted after import. + +At least one mapping selection and at least one mapped case are required. +Mappings are never inferred from suite, class, or test names. Duplicate pairs +are canonicalized; unknown, empty, or conflicting selectors fail closed. One +resolved case cannot belong to more than one criterion in the same import. + +### XML safety and boundedness + +The parser accepts bytes, checks the byte limit before decoding, and accepts +UTF-8 only. UTF-8 BOM is allowed; any other declared encoding is rejected. +Before tree construction it rejects XML containing a document type, +entities, processing instructions other than the XML declaration, or obvious +XInclude markup. After parsing it rejects any XInclude namespace element and +enforces element, suite, and test-case limits. + +Accepted roots are one `testsuite` or one `testsuites` with direct +`testsuite` children. Nested suites, test cases outside suites, multiple result +markers, unknown result-marker structures, and missing test names fail closed. +Observed statuses come only from zero or one direct `failure`, `error`, or +`skipped` marker. Zero markers means passed. + +Declared JUnit counts are not trusted. ScopeProof computes totals from parsed +cases. A non-negative declared count that differs from the observed count adds +a deterministic warning. Invalid numeric declarations fail closed. Presence of +ignored `system-out`, `system-err`, or `properties` adds one bounded warning; +their contents are never retained. + +### Lifecycle and atomic persistence + +Add `append_junit_evidence_import(state, record)` to the core lifecycle. It +revalidates both objects, verifies the active bundle and all identity fields, +rejects duplicate artifact digests and IDs, appends a deep copy, and confirms +that findings, resolutions, runtime evidence, final acceptance, and the +deterministic gate are byte-for-byte unchanged. + +The CLI uses `JsonReviewStore.mutate`, so parsing and lifecycle validation run +inside the serialized read-transition-write boundary. Any exception prevents +the replacement write. Streamlit applies the same lifecycle function to its +validated session state; the existing local-save path performs persistence. + +Criteria revision moves the prior active bundle, including its import records, +to analysis history. A later active analysis begins with no imports. There is +no automatic carry-forward or remapping. + +### CLI + +Add two commands: + +```text +scopeproof inspect-junit ARTIFACT +scopeproof import-junit REVIEW_ID ARTIFACT --mapping MAPPING.json \ + --importer "Asserted name" [--limitation TEXT ...] [--storage-dir PATH] +``` + +`inspect-junit` prints the validated sanitized parser result, including scope +IDs and computed totals. It never persists anything. + +The strict mapping document is: + +```json +{ + "schema_version": "junit-mapping-v1", + "artifact_sha256": "COPY_THE_64_CHARACTER_DIGEST_FROM_INSPECT_JUNIT", + "selections": [ + {"scope_id": "suite-0001", "criterion_id": "AC-01"} + ] +} +``` + +`import-junit` reads the explicitly named local artifact and mapping files in +binary mode on every supported platform, +requires the mapping digest to match the bounded artifact bytes, builds the +record through the shared core service, applies the atomic lifecycle transition, +and prints `JUnitImportMutationMetadata` with an explicit external non-gating +boundary. It never persists file paths or raw XML. + +### Streamlit + +Add a collapsed section named `Import external JUnit results` beside the +selected criterion controls, visually separate from candidate evidence and +`Record optional external verification (E3/E4)`. + +The user uploads one XML file, enters an asserted importer, previews sanitized +suites/cases and warnings, explicitly selects one or more scope IDs, and maps +them to the already selected criterion. Save is disabled until the artifact, +importer, and mapping are valid. Successful save appends the import in session, +clears upload/mapping form state, and shows the non-gating boundary. Failure +shows a bounded error and leaves the state unchanged. + +Recorded imports are shown by selected criterion and display digest, exact +review identity, mapped sanitized cases/statuses, asserted importer, timestamp, +warnings, and limitations. Raw XML and ignored output never render. + +### Exports and comparison + +JSON naturally includes validated import envelopes. Markdown and HTML add an +`Imported external test results` section using inert escaped text. CSV adds +criterion-level imported artifact digests, mapped case IDs/statuses, asserted +importers, warnings, and limitations with spreadsheet-formula neutralization. + +Comparison revalidates both bundles and projects imported records by artifact +digest plus explicit mapping signature. It reports added, removed, unchanged, +or mapping-modified imports. Changed imported context adds affected criteria to +`criteria_requiring_decision_review` only when the previous bundle has a human +resolution for that criterion. It does not copy a resolution or alter either +gate. + +### Browser proof + +The installed-wheel Chromium regression creates a synthetic exact-head saved +review using installed ScopeProof code, opens it in the packaged workbench, +uploads a small local JUnit fixture, explicitly maps its suite to a criterion, +saves the import, verifies the displayed boundary, saves/reopens the review, +and downloads JSON and Markdown containing the artifact digest but not raw XML +or ignored output. Browser networking remains loopback-only. + +## Error behavior + +Public errors are deterministic categories rather than parser internals: + +- artifact too large; +- unsupported encoding; +- forbidden XML construct; +- malformed or unsupported JUnit structure; +- suite/case/element limit exceeded; +- invalid or unknown mapping scope; +- review identity, criteria, or provenance mismatch; +- duplicate or conflicting import; +- stale saved-state mutation. + +No error includes raw XML, failure text, stdout/stderr, local file contents, or +credentials. + +## Verification and non-goals + +The implementation requires focused red-green tests plus the complete suite, +95% coverage gate, repository contracts, deterministic and comparison +benchmarks, reproducible wheels, clean installation, installed CLIs and +benchmarks, workbench health, installed-wheel Chromium, supported Python lanes, +hosted Windows, final diff audit, and independent review. + +This slice does not authenticate reviewers, add accounts, host source or +artifacts, support private repositories, make the GitHub Action required, +release or publish `0.2.4`, retune R-002, generate R-003, begin Stage 3, reopen +Stage 1, or claim accessibility, customer, demand, adoption, or correctness +evidence. + +Stage 1 remains closed as not pursued at 0/5 qualifying reviews, 0/3 +independent practitioners, 0/3 public repositories, 0/3 independently observed +under-ten-minute completions, and 0/2 reuse-intent signals. diff --git a/pyproject.toml b/pyproject.toml index ce22a836..9ad990f3 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -11,7 +11,7 @@ requires-python = ">=3.11" dependencies = [ "httpx>=0.27,<1", "pydantic>=2.8,<3", - "streamlit>=1.52,<2", + "streamlit>=1.53,<2", ] [project.urls] diff --git a/scopeproof_core/cli.py b/scopeproof_core/cli.py index fb220320..0cf795ad 100644 --- a/scopeproof_core/cli.py +++ b/scopeproof_core/cli.py @@ -4,10 +4,15 @@ import argparse import json +import os +import stat from datetime import UTC, datetime +from hashlib import sha256 from pathlib import Path from uuid import uuid4 +from pydantic import ValidationError + from scopeproof_core.alpha.models import AlphaFrictionStage, AlphaOutcome, ParticipantRole from scopeproof_core.alpha.rehearsal import initialize_alpha_rehearsal from scopeproof_core.alpha.rehearsal_storage import JsonAlphaRehearsalStore @@ -29,6 +34,12 @@ from scopeproof_core.evals.runner import run_bundled_benchmark from scopeproof_core.gates.evaluator import evaluate_gate from scopeproof_core.github.client import GitHubClient, GitHubIngestionError +from scopeproof_core.importers.junit import ( + MAX_JUNIT_BYTES, + JUnitMappingDocument, + build_junit_evidence_import, + parse_junit_artifact, +) from scopeproof_core.reporting.exporters import ( export_comparison_json, export_comparison_markdown, @@ -42,6 +53,7 @@ from scopeproof_core.reviews.lifecycle import ( acceptance_requires_comment, append_external_verification, + append_junit_evidence_import, append_resolution, new_review_state, ) @@ -51,6 +63,7 @@ Criterion, EvidenceLevel, HumanDecision, + JUnitImportMutationMetadata, LifecycleMutationMetadata, PullRequestSnapshot, ResearchContext, @@ -72,6 +85,8 @@ from scopeproof_core.verification.service import build_findings from scopeproof_core.version import __version__ +_BINARY = getattr(os, "O_BINARY", 0) + EXPORT_RENDERERS = { "json": export_json, "markdown": export_markdown, @@ -445,6 +460,103 @@ def _compare(args: argparse.Namespace) -> int: return 0 +def _read_bounded_regular_file( + path: Path, *, max_bytes: int, label: str +) -> bytes: + """Read one regular local file without buffering beyond its explicit budget.""" + + flags = os.O_RDONLY | _BINARY + flags |= getattr(os, "O_CLOEXEC", 0) + flags |= getattr(os, "O_NONBLOCK", 0) + flags |= getattr(os, "O_NOFOLLOW", 0) + descriptor = os.open(path, flags) + try: + metadata = os.fstat(descriptor) + if not stat.S_ISREG(metadata.st_mode): + raise ValueError(f"{label} must be a regular file") + if metadata.st_size > max_bytes: + raise ValueError(f"{label} exceeds the byte limit") + chunks: list[bytes] = [] + remaining = max_bytes + 1 + while remaining: + chunk = os.read(descriptor, min(65_536, remaining)) + if not chunk: + break + chunks.append(chunk) + remaining -= len(chunk) + artifact_bytes = b"".join(chunks) + if len(artifact_bytes) > max_bytes: + raise ValueError(f"{label} exceeds the byte limit") + return artifact_bytes + finally: + os.close(descriptor) + + +def _read_bounded_junit_artifact(path: Path) -> bytes: + return _read_bounded_regular_file( + path, + max_bytes=MAX_JUNIT_BYTES, + label="JUnit artifact", + ) + + +def _inspect_junit(args: argparse.Namespace) -> int: + """Print a sanitized bytes-only JUnit projection without persistence.""" + + parsed = parse_junit_artifact(_read_bounded_junit_artifact(Path(args.artifact))) + print(parsed.model_dump_json(indent=2)) + return 0 + + +def _import_junit(args: argparse.Namespace) -> int: + """Atomically append one validated external test-result import.""" + + mapping_bytes = _read_bounded_regular_file( + Path(args.mapping), + max_bytes=MAX_JUNIT_BYTES, + label="JUnit mapping", + ) + try: + mapping = JUnitMappingDocument.model_validate_json(mapping_bytes) + except ValidationError: + raise ValueError("JUnit mapping document is invalid") from None + artifact_bytes = _read_bounded_junit_artifact(Path(args.artifact)) + if sha256(artifact_bytes).hexdigest() != mapping.artifact_sha256: + raise ValueError("JUnit mapping digest does not match the selected artifact") + store = JsonReviewStore(Path(args.storage_dir)) + imported = None + + def transition(state: ReviewState) -> ReviewState: + nonlocal imported + imported = build_junit_evidence_import( + state, + artifact_bytes, + mapping.selections, + importer=args.importer, + limitations=args.limitation, + ) + return append_junit_evidence_import(state, imported) + + updated, path = store.mutate(args.review_id, transition) + if imported is None or updated.bundle is None: + raise ValueError("JUnit import transition did not produce an active record") + metadata = JUnitImportMutationMetadata( + review_id=updated.review.review_id, + record=str(path), + head_sha=updated.review.head_sha, + import_id=imported.import_id, + artifact_sha256=imported.artifact_sha256, + mapped_criterion_ids=sorted( + mapping.criterion_id for mapping in imported.criterion_mappings + ), + totals=imported.totals, + evidence_boundary="externally_supplied_non_gating", + verdict=updated.bundle.gate.verdict, + ) + print(metadata.model_dump_json()) + return 0 + + def _validate_action_evidence(args: argparse.Namespace) -> int: """Validate owner-supplied external Action evidence without contacting GitHub.""" @@ -680,6 +792,23 @@ def _parser() -> argparse.ArgumentParser: compare.add_argument("--output") compare.add_argument("--storage-dir", default=".scopeproof/reviews") compare.set_defaults(handler=_compare) + inspect_junit = commands.add_parser( + "inspect-junit", + help="Inspect bounded local JUnit XML without saving or executing it", + ) + inspect_junit.add_argument("artifact", help="Explicit local JUnit XML file") + inspect_junit.set_defaults(handler=_inspect_junit) + import_junit = commands.add_parser( + "import-junit", + help="Append mapped external JUnit results to one exact-head saved review", + ) + import_junit.add_argument("review_id") + import_junit.add_argument("artifact", help="Explicit local JUnit XML file") + import_junit.add_argument("--mapping", required=True, help="Strict mapping JSON file") + import_junit.add_argument("--importer", required=True, help="Asserted importer identity") + import_junit.add_argument("--limitation", action="append", default=[]) + import_junit.add_argument("--storage-dir", default=".scopeproof/reviews") + import_junit.set_defaults(handler=_import_junit) benchmark = commands.add_parser("benchmark", help="Run every labelled local benchmark case") benchmark.set_defaults(handler=lambda _: _benchmark()) comparison_benchmark = commands.add_parser( diff --git a/scopeproof_core/importers/__init__.py b/scopeproof_core/importers/__init__.py new file mode 100644 index 00000000..addece8a --- /dev/null +++ b/scopeproof_core/importers/__init__.py @@ -0,0 +1,29 @@ +"""Bounded, non-executing import adapters for externally supplied evidence.""" + +from scopeproof_core.importers.junit import ( + MAX_JUNIT_BYTES, + MAX_JUNIT_CASES, + MAX_JUNIT_ELEMENTS, + MAX_JUNIT_SUITES, + JUnitImportError, + JUnitMappingDocument, + JUnitMappingSelection, + ParsedJUnitArtifact, + ParsedJUnitSuite, + build_junit_evidence_import, + parse_junit_artifact, +) + +__all__ = [ + "MAX_JUNIT_BYTES", + "MAX_JUNIT_CASES", + "MAX_JUNIT_ELEMENTS", + "MAX_JUNIT_SUITES", + "JUnitImportError", + "JUnitMappingDocument", + "JUnitMappingSelection", + "ParsedJUnitArtifact", + "ParsedJUnitSuite", + "build_junit_evidence_import", + "parse_junit_artifact", +] diff --git a/scopeproof_core/importers/junit.py b/scopeproof_core/importers/junit.py new file mode 100644 index 00000000..ab0cb77a --- /dev/null +++ b/scopeproof_core/importers/junit.py @@ -0,0 +1,534 @@ +"""Parse bounded JUnit XML bytes without executing or dereferencing artifact content.""" + +from __future__ import annotations + +import re +from collections import defaultdict +from datetime import UTC, datetime +from hashlib import sha256 +from typing import Literal +from uuid import uuid4 +from xml.etree import ElementTree + +from pydantic import BaseModel, ConfigDict, Field, ValidationError, field_validator + +from scopeproof_core.criteria.confirmation import normalized_criteria_sha256 +from scopeproof_core.gates.validation import validated_review_state +from scopeproof_core.schemas.models import ( + MAX_JUNIT_IMPORTS_PER_REVIEW, + JUnitCaseResult, + JUnitCaseStatus, + JUnitCriterionMapping, + JUnitEvidenceBoundary, + JUnitEvidenceImport, + JUnitResultTotals, + ReviewState, + junit_name_is_path_or_url_like, +) + +MAX_JUNIT_BYTES = 1_048_576 +MAX_JUNIT_SUITES = 100 +MAX_JUNIT_CASES = 5_000 +MAX_JUNIT_ELEMENTS = 20_000 +MAX_JUNIT_NAME_LENGTH = 512 + +_EXACT_HEAD = re.compile(r"^[a-f0-9]{40}$") +_XML_ENCODING = re.compile( + r"<\?xml\b[^>]*\bencoding\s*=\s*(['\"])([^'\"]+)\1", + re.IGNORECASE, +) +_OTHER_PROCESSING_INSTRUCTION = re.compile( + r"<\?(?!xml(?:\s|\?>))", + re.IGNORECASE, +) +_STABLE_SCOPE = r"^suite-\d{4}(?:-case-\d{4})?$" +_DISCARDED_WARNING = ( + "JUnit properties and output content were discarded during import." +) +_COUNT_WARNING = ( + "Declared JUnit counts differed from the sanitized observed results." +) +_REDACTED_NAME_WARNING = ( + "Path- or URL-like JUnit names were redacted during import." +) +_FIXED_LIMITATIONS = ( + "ScopeProof did not execute the imported tests or target-repository code.", + "The artifact digest does not prove criterion correctness or runtime behavior.", + "The asserted importer identity is not authenticated.", + "Explicit human mapping is organizational context, not proof that a criterion passed.", +) + + +class JUnitImportError(ValueError): + """A bounded public error that never exposes artifact contents.""" + + +class ParsedJUnitSuite(BaseModel): + """One sanitized suite projection with deterministic document-order IDs.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + suite_id: str = Field(pattern=r"^suite-\d{4}$") + suite_name: str = Field(min_length=1, max_length=MAX_JUNIT_NAME_LENGTH) + test_cases: list[JUnitCaseResult] = Field(default_factory=list) + + +class ParsedJUnitArtifact(BaseModel): + """Validated parser output containing no raw XML or ignored bodies.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + schema_version: Literal["junit-parsed-v1"] = "junit-parsed-v1" + artifact_sha256: str = Field(pattern=r"^[a-f0-9]{64}$") + artifact_format: Literal["junit_xml"] = "junit_xml" + suites: list[ParsedJUnitSuite] = Field(max_length=MAX_JUNIT_SUITES) + totals: JUnitResultTotals + parser_warnings: list[str] = Field(default_factory=list, max_length=100) + + +class JUnitMappingSelection(BaseModel): + """One explicit human choice before selectors are resolved to case IDs.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + scope_id: str = Field(pattern=_STABLE_SCOPE) + criterion_id: str = Field(min_length=1) + + @field_validator("criterion_id") + @classmethod + def normalize_criterion_id(cls, value: str) -> str: + normalized = value.strip() + if not normalized: + raise ValueError("criterion ID must contain non-whitespace text") + return normalized + + +class JUnitMappingDocument(BaseModel): + """Strict local mapping-file contract for CLI imports.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + schema_version: Literal["junit-mapping-v1"] + artifact_sha256: str = Field(pattern=r"^[a-f0-9]{64}$") + selections: list[JUnitMappingSelection] = Field(min_length=1, max_length=5_000) + + +def _local_name(tag: object) -> str: + if not isinstance(tag, str): + raise JUnitImportError("JUnit XML contains an unsupported element tag.") + if tag.startswith("{http://www.w3.org/2001/XInclude}"): + raise JUnitImportError("JUnit XML must not contain XInclude elements.") + if tag.startswith("{"): + raise JUnitImportError("JUnit XML namespaces are unsupported.") + return tag + + +def _bounded_name(value: str | None, *, fallback: str | None = None) -> str: + normalized = (value or "").strip() + if not normalized and fallback is not None: + normalized = fallback + if not normalized: + raise JUnitImportError("JUnit test cases require a non-blank test name.") + if len(normalized) > MAX_JUNIT_NAME_LENGTH: + raise JUnitImportError("JUnit names exceed the supported length.") + return normalized + + +def _optional_bounded_name(value: str | None) -> str | None: + if value is None: + return None + normalized = value.strip() + if not normalized: + return None + if len(normalized) > MAX_JUNIT_NAME_LENGTH: + raise JUnitImportError("JUnit names exceed the supported length.") + return normalized + + +def _sanitized_required_name( + value: str | None, + *, + fallback: str | None = None, + redacted_fallback: str, +) -> tuple[str, bool]: + normalized = _bounded_name(value, fallback=fallback) + if junit_name_is_path_or_url_like(normalized): + return redacted_fallback, True + return normalized, False + + +def _sanitized_optional_name( + value: str | None, *, redacted_fallback: str +) -> tuple[str | None, bool]: + normalized = _optional_bounded_name(value) + if normalized is not None and junit_name_is_path_or_url_like(normalized): + return redacted_fallback, True + return normalized, False + + +def _declared_count(element: ElementTree.Element, name: str) -> int | None: + raw = element.attrib.get(name) + if raw is None: + return None + try: + value = int(raw) + except (TypeError, ValueError): + raise JUnitImportError("JUnit declared counts must be non-negative integers.") from None + if value < 0: + raise JUnitImportError("JUnit declared counts must be non-negative integers.") + return value + + +def _totals_for_cases(cases: list[JUnitCaseResult]) -> JUnitResultTotals: + counts = { + JUnitCaseStatus.PASSED: 0, + JUnitCaseStatus.FAILURE: 0, + JUnitCaseStatus.ERROR: 0, + JUnitCaseStatus.SKIPPED: 0, + } + for item in cases: + counts[item.status] += 1 + return JUnitResultTotals( + total=len(cases), + passed=counts[JUnitCaseStatus.PASSED], + failures=counts[JUnitCaseStatus.FAILURE], + errors=counts[JUnitCaseStatus.ERROR], + skipped=counts[JUnitCaseStatus.SKIPPED], + ) + + +def _declared_counts_differ( + element: ElementTree.Element, totals: JUnitResultTotals +) -> bool: + declared = { + "tests": _declared_count(element, "tests"), + "failures": _declared_count(element, "failures"), + "errors": _declared_count(element, "errors"), + "skipped": _declared_count(element, "skipped"), + } + observed = { + "tests": totals.total, + "failures": totals.failures, + "errors": totals.errors, + "skipped": totals.skipped, + } + return any(value is not None and value != observed[name] for name, value in declared.items()) + + +def _validate_discarded_wrapper( + element: ElementTree.Element, wrapper_name: str +) -> None: + """Reject element structure that would otherwise disappear inside ignored content.""" + + if wrapper_name == "properties": + for child in element: + if _local_name(child.tag) != "property" or len(child): + raise JUnitImportError( + "JUnit discarded content contains an unsupported structure." + ) + return + if len(element): + raise JUnitImportError( + "JUnit discarded content contains an unsupported structure." + ) + + +def _parse_case( + element: ElementTree.Element, + *, + suite_id: str, + suite_name: str, + case_number: int, +) -> tuple[JUnitCaseResult, bool, bool]: + result_markers: list[str] = [] + discarded = False + for child in element: + name = _local_name(child.tag) + if name in {"failure", "error", "skipped"}: + if len(child): + raise JUnitImportError( + "JUnit test result contains an unsupported structure." + ) + result_markers.append(name) + elif name in {"properties", "system-out", "system-err"}: + _validate_discarded_wrapper(child, name) + discarded = True + else: + raise JUnitImportError("JUnit test case contains an unsupported result structure.") + if len(result_markers) > 1: + raise JUnitImportError("JUnit test case contains multiple result markers.") + marker = result_markers[0] if result_markers else "passed" + class_name, class_name_redacted = _sanitized_optional_name( + element.attrib.get("classname"), redacted_fallback="Redacted class name" + ) + test_name, test_name_redacted = _sanitized_required_name( + element.attrib.get("name"), + redacted_fallback=f"Redacted test name {case_number:04d}", + ) + try: + result = JUnitCaseResult( + test_case_id=f"{suite_id}-case-{case_number:04d}", + suite_id=suite_id, + suite_name=suite_name, + class_name=class_name, + test_name=test_name, + status=JUnitCaseStatus(marker), + ) + except ValidationError: + raise JUnitImportError("JUnit names exceed the supported length or shape.") from None + return result, discarded, class_name_redacted or test_name_redacted + + +def _parse_suite( + element: ElementTree.Element, suite_number: int +) -> tuple[ParsedJUnitSuite, bool, bool, bool]: + suite_id = f"suite-{suite_number:04d}" + suite_name, names_redacted = _sanitized_required_name( + element.attrib.get("name"), + fallback=f"Unnamed suite {suite_number:04d}", + redacted_fallback=f"Redacted suite name {suite_number:04d}", + ) + test_cases: list[JUnitCaseResult] = [] + discarded = False + for child in element: + name = _local_name(child.tag) + if name == "testcase": + if len(test_cases) >= MAX_JUNIT_CASES: + raise JUnitImportError("JUnit artifact exceeds the test-case limit.") + case, case_discarded, case_name_redacted = _parse_case( + child, + suite_id=suite_id, + suite_name=suite_name, + case_number=len(test_cases) + 1, + ) + test_cases.append(case) + discarded = discarded or case_discarded + names_redacted = names_redacted or case_name_redacted + elif name == "testsuite": + raise JUnitImportError("Nested JUnit test suites are unsupported.") + elif name in {"properties", "system-out", "system-err"}: + _validate_discarded_wrapper(child, name) + discarded = True + else: + raise JUnitImportError("JUnit test suite contains an unsupported structure.") + totals = _totals_for_cases(test_cases) + differs = _declared_counts_differ(element, totals) + return ( + ParsedJUnitSuite( + suite_id=suite_id, + suite_name=suite_name, + test_cases=test_cases, + ), + discarded, + differs, + names_redacted, + ) + + +def parse_junit_artifact(artifact_bytes: bytes) -> ParsedJUnitArtifact: + """Return a bounded sanitized projection without interpreting external references.""" + + if not isinstance(artifact_bytes, bytes): + raise TypeError("JUnit artifact input must be bytes") + if not artifact_bytes: + raise JUnitImportError("JUnit artifact is empty.") + if len(artifact_bytes) > MAX_JUNIT_BYTES: + raise JUnitImportError("JUnit artifact exceeds the byte limit.") + try: + text = artifact_bytes.decode("utf-8-sig") + except UnicodeDecodeError: + raise JUnitImportError("JUnit artifact must use UTF-8 encoding.") from None + encoding_match = _XML_ENCODING.search(text) + if encoding_match is not None and encoding_match.group(2).lower().replace("_", "-") not in { + "utf-8", + "utf8", + }: + raise JUnitImportError("JUnit artifact must use UTF-8 encoding.") + upper_text = text.upper() + if " MAX_JUNIT_ELEMENTS: + raise JUnitImportError("JUnit artifact exceeds the element limit.") + for element in elements: + _local_name(element.tag) + + root_name = _local_name(root.tag) + if root_name == "testsuite": + suite_elements = [root] + root_discarded = False + elif root_name == "testsuites": + suite_elements = [] + root_discarded = False + for child in root: + child_name = _local_name(child.tag) + if child_name == "testsuite": + suite_elements.append(child) + elif child_name in {"properties", "system-out", "system-err"}: + _validate_discarded_wrapper(child, child_name) + root_discarded = True + else: + raise JUnitImportError( + "JUnit testsuites root requires direct testsuite children." + ) + else: + raise JUnitImportError("JUnit artifact root must be testsuite or testsuites.") + if len(suite_elements) > MAX_JUNIT_SUITES: + raise JUnitImportError("JUnit artifact exceeds the suite limit.") + + suites: list[ParsedJUnitSuite] = [] + discarded = root_discarded + declared_mismatch = False + names_redacted = False + total_cases = 0 + for suite_number, element in enumerate(suite_elements, start=1): + suite, suite_discarded, suite_mismatch, suite_names_redacted = _parse_suite( + element, suite_number + ) + total_cases += len(suite.test_cases) + if total_cases > MAX_JUNIT_CASES: + raise JUnitImportError("JUnit artifact exceeds the test-case limit.") + suites.append(suite) + discarded = discarded or suite_discarded + declared_mismatch = declared_mismatch or suite_mismatch + names_redacted = names_redacted or suite_names_redacted + + all_cases = [item for suite in suites for item in suite.test_cases] + totals = _totals_for_cases(all_cases) + if root_name == "testsuites": + declared_mismatch = declared_mismatch or _declared_counts_differ(root, totals) + warnings: list[str] = [] + if discarded: + warnings.append(_DISCARDED_WARNING) + if declared_mismatch: + warnings.append(_COUNT_WARNING) + if names_redacted: + warnings.append(_REDACTED_NAME_WARNING) + return ParsedJUnitArtifact( + artifact_sha256=sha256(artifact_bytes).hexdigest(), + suites=suites, + totals=totals, + parser_warnings=warnings, + ) + + +def build_junit_evidence_import( + state: ReviewState, + artifact_bytes: bytes, + selections: list[JUnitMappingSelection], + *, + importer: str, + limitations: list[str] | None = None, + imported_at: datetime | None = None, + import_id: str | None = None, +) -> JUnitEvidenceImport: + """Bind sanitized external test results to one exact active review snapshot.""" + + state = validated_review_state(state) + if state.bundle is None: + raise ValueError("JUnit import requires an active analysis") + bundle = state.bundle + if len(bundle.junit_evidence_imports) >= MAX_JUNIT_IMPORTS_PER_REVIEW: + raise JUnitImportError( + "Review already contains the maximum number of JUnit imports." + ) + provenance = bundle.review.criteria_source_provenance + if provenance is None or not bundle.review.criteria_confirmed: + raise ValueError("JUnit import requires confirmed criteria provenance") + if _EXACT_HEAD.fullmatch(bundle.review.head_sha) is None: + raise ValueError("JUnit import requires an exact 40-character head SHA") + normalized_importer = importer.strip() + if not normalized_importer: + raise ValueError("JUnit importer must contain non-whitespace text") + supplied_limitations = [] if limitations is None else [item.strip() for item in limitations] + if any(not item for item in supplied_limitations): + raise ValueError("JUnit limitations must contain non-whitespace text") + if not selections: + raise JUnitImportError("JUnit import requires at least one explicit mapping.") + validated_selections = [ + JUnitMappingSelection.model_validate(item.model_dump(mode="python")) + for item in selections + ] + parsed = parse_junit_artifact(artifact_bytes) + if any( + existing.artifact_sha256 == parsed.artifact_sha256 + for existing in bundle.junit_evidence_imports + ): + raise JUnitImportError("Duplicate JUnit artifact digest is already imported.") + resolved_import_id = import_id or str(uuid4()) + if any(existing.import_id == resolved_import_id for existing in bundle.junit_evidence_imports): + raise JUnitImportError("Duplicate JUnit import ID is already recorded.") + + known_criteria = {criterion.criterion_id for criterion in bundle.criteria} + scopes: dict[str, list[str]] = {} + test_cases: list[JUnitCaseResult] = [] + for suite in parsed.suites: + suite_case_ids = [item.test_case_id for item in suite.test_cases] + scopes[suite.suite_id] = suite_case_ids + for item in suite.test_cases: + scopes[item.test_case_id] = [item.test_case_id] + test_cases.append(item) + mapped: dict[str, set[str]] = defaultdict(set) + case_owners: dict[str, str] = {} + for selection in validated_selections: + if selection.criterion_id not in known_criteria: + raise JUnitImportError("JUnit mapping references an unknown criterion.") + case_ids = scopes.get(selection.scope_id) + if case_ids is None: + raise JUnitImportError("JUnit import references an unknown mapping scope.") + if not case_ids: + raise JUnitImportError("JUnit mapping scope contains no test cases.") + for case_id in case_ids: + existing_owner = case_owners.get(case_id) + if existing_owner is not None and existing_owner != selection.criterion_id: + raise JUnitImportError( + "JUnit mapping must not assign one test case to multiple criteria." + ) + case_owners[case_id] = selection.criterion_id + mapped[selection.criterion_id].update(case_ids) + if not mapped or not any(mapped.values()): + raise JUnitImportError("JUnit import requires at least one explicit mapping.") + criterion_mappings = [ + JUnitCriterionMapping( + criterion_id=criterion_id, + test_case_ids=sorted(case_ids), + ) + for criterion_id, case_ids in sorted(mapped.items()) + ] + try: + return JUnitEvidenceImport( + schema_version="junit-import-v1", + evidence_boundary=JUnitEvidenceBoundary( + source="externally_supplied", + gate_effect="non_gating", + execution="not_executed_by_scopeproof", + artifact_digest_scope="imported_bytes_only", + importer_identity="asserted_not_authenticated", + criterion_mapping="organizational_context_not_proof", + ), + import_id=resolved_import_id, + repository=bundle.review.repository, + pr_number=bundle.review.pr_number, + head_sha=bundle.review.head_sha, + criteria_revision_number=state.criteria_revision.number, + confirmed_criteria_sha256=normalized_criteria_sha256(bundle.criteria), + criteria_source_provenance=provenance.model_copy(deep=True), + artifact_sha256=parsed.artifact_sha256, + imported_by=normalized_importer, + imported_at=imported_at or datetime.now(UTC), + totals=parsed.totals, + test_cases=sorted(test_cases, key=lambda item: item.test_case_id), + criterion_mappings=criterion_mappings, + parser_warnings=parsed.parser_warnings, + limitations=list(dict.fromkeys((*_FIXED_LIMITATIONS, *supplied_limitations))), + ) + except ValidationError: + raise JUnitImportError("JUnit import metadata is invalid.") from None diff --git a/scopeproof_core/reporting/exporters.py b/scopeproof_core/reporting/exporters.py index 08e5e6b9..47375d52 100644 --- a/scopeproof_core/reporting/exporters.py +++ b/scopeproof_core/reporting/exporters.py @@ -28,10 +28,13 @@ ReviewComparison, ) from scopeproof_core.schemas.models import ( + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, CriteriaSourceProvenance, CriterionRetrievalDiagnostic, EvidenceItem, HumanDecision, + JUnitCaseResult, + JUnitEvidenceImport, ReviewBundle, ReviewState, ) @@ -169,12 +172,14 @@ def _retrieval_diagnostic_html( ) exact_identifiers = ( ", ".join( - f"{html.escape(identifier)}" for identifier in diagnostic.exact_identifiers + f"{html.escape(identifier)}" + for identifier in diagnostic.exact_identifiers ) or "None" ) evidence_types = ( - ", ".join(html.escape(item.value) for item in diagnostic.searched_evidence_types) or "None" + ", ".join(html.escape(item.value) for item in diagnostic.searched_evidence_types) + or "None" ) return "".join( [ @@ -197,6 +202,137 @@ def _retrieval_diagnostic_html( ) +def _junit_mapping_cases( + evidence_import: JUnitEvidenceImport, criterion_id: str +) -> list[JUnitCaseResult]: + mapped_ids = { + case_id + for mapping in evidence_import.criterion_mappings + if mapping.criterion_id == criterion_id + for case_id in mapping.test_case_ids + } + return [ + item for item in evidence_import.test_cases if item.test_case_id in mapped_ids + ] + + +def _junit_import_markdown(bundle: ReviewBundle) -> list[str]: + lines = [ + "## Imported External Test Results", + "", + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, + "", + ] + if not bundle.junit_evidence_imports: + return [*lines, "No external JUnit results were imported.", ""] + for evidence_import in bundle.junit_evidence_imports: + lines.extend( + [ + f"### Import {_render_markdown_code(evidence_import.import_id)}", + "", + f"- Artifact SHA-256: {_render_markdown_code(evidence_import.artifact_sha256)}", + f"- Bound head: {_render_markdown_code(evidence_import.head_sha)}", + f"- Asserted importer: {_render_markdown_code(evidence_import.imported_by)}", + "- Computed totals: " + f"{evidence_import.totals.total} total; " + f"{evidence_import.totals.passed} passed; " + f"{evidence_import.totals.failures} failed; " + f"{evidence_import.totals.errors} errors; " + f"{evidence_import.totals.skipped} skipped.", + "- Explicit mappings:", + ] + ) + cases_by_id = { + item.test_case_id: item for item in evidence_import.test_cases + } + for mapping in evidence_import.criterion_mappings: + lines.append(f" - Criterion {_render_markdown_code(mapping.criterion_id)}") + for case_id in mapping.test_case_ids: + case = cases_by_id[case_id] + lines.append( + " - " + f"{_render_markdown_code(case.test_case_id)} · " + f"{_render_markdown_code(case.status.value)} · " + f"{_render_markdown_code(case.suite_name)} · " + f"{_render_markdown_code(case.test_name)}" + ) + if evidence_import.parser_warnings: + lines.append("- Parser warnings:") + lines.extend( + f" - {_render_markdown_code(item)}" + for item in evidence_import.parser_warnings + ) + lines.append("- Limitations:") + lines.extend( + f" - {_render_markdown_code(item)}" for item in evidence_import.limitations + ) + lines.append("") + return lines + + +def _junit_import_html(bundle: ReviewBundle) -> list[str]: + lines = [ + "

Imported external test results

", + f'

{html.escape(JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION)}

', + ] + if not bundle.junit_evidence_imports: + return [*lines, "

No external JUnit results were imported.

"] + for evidence_import in bundle.junit_evidence_imports: + cases_by_id = { + item.test_case_id: item for item in evidence_import.test_cases + } + lines.extend( + [ + f"

Import {html.escape(evidence_import.import_id)}

", + "", + ] + ) + return lines + + def export_json(bundle: ExportableReview) -> str: """Return canonical, diff-friendly JSON without adapter state or credentials.""" payload = _validated_exportable(bundle).model_dump(mode="json") @@ -279,6 +415,50 @@ def export_comparison_markdown(comparison: ReviewComparison) -> str: lines.append("- Review the current evidence before recording a new decision.") lines.append("") + lines.extend( + [ + "## Imported External Test Result Changes", + "", + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, + "", + ] + ) + if not comparison.junit_import_changes: + lines.extend(["No imported JUnit context was present in either review.", ""]) + for change in comparison.junit_import_changes: + lines.extend( + [ + "### " + f"{_render_markdown_code(change.artifact_sha256)} — " + f"{_escape_markdown_text(change.kind.value.replace('_', ' ').title())}", + "", + ] + ) + for label, reference in ( + ("Previous import", change.previous), + ("Current import", change.current), + ): + if reference is None: + continue + lines.extend( + [ + f"- **{label}:** {_render_markdown_code(reference.import_id)}", + f" - Bound head: {_render_markdown_code(reference.head_sha)}", + " - Asserted importer: " + f"{_render_markdown_code(reference.asserted_importer)}", + " - Explicit mappings:", + ] + ) + for mapping in reference.mappings: + lines.append( + f" - {_render_markdown_code(mapping.criterion_id)}: " + + ", ".join( + _render_markdown_code(case_id) + for case_id in mapping.test_case_ids + ) + ) + lines.append("") + if comparison.changed_finding_statuses: lines.extend(["## Changed Criterion Findings", ""]) for change in comparison.changed_finding_statuses: @@ -310,8 +490,8 @@ def export_comparison_markdown(comparison: ReviewComparison) -> str: ], "", ( - "ScopeProof never carries acceptance to a changed head. Review current " - "evidence and record a new decision." + "ScopeProof does not carry a prior decision forward automatically. Review " + "the current evidence and imported context, then record a new decision." ), "", ] @@ -558,6 +738,7 @@ def export_markdown(bundle: ExportableReview) -> str: ] ) + lines.extend(_junit_import_markdown(bundle)) lines.extend( [ "## Runtime Verification Boundary", @@ -750,6 +931,12 @@ def export_csv(bundle: ExportableReview) -> str: "runtime_pr_numbers", "runtime_head_shas", "manual_runtime_evidence_id", + "junit_artifact_digests", + "junit_mapped_cases", + "junit_evidence_boundary", + "junit_importers", + "junit_parser_warnings", + "junit_limitations", ] output = io.StringIO(newline="") writer = csv.DictWriter(output, fieldnames=fieldnames, lineterminator="\r\n") @@ -762,6 +949,14 @@ def export_csv(bundle: ExportableReview) -> str: runtime_items = [ item for item in bundle.runtime_evidence if item.criterion_id == criterion.criterion_id ] + junit_imports = [ + item + for item in bundle.junit_evidence_imports + if any( + mapping.criterion_id == criterion.criterion_id + for mapping in item.criterion_mappings + ) + ] writer.writerow( { "review_id": _csv_text(bundle.review.review_id), @@ -895,6 +1090,64 @@ def export_csv(bundle: ExportableReview) -> str: ) if resolution is not None and resolution.decision is HumanDecision.MANUALLY_VERIFIED else "", + "junit_artifact_digests": json.dumps( + [item.artifact_sha256 for item in junit_imports], + ensure_ascii=False, + ), + "junit_mapped_cases": json.dumps( + [ + { + "import_id": evidence_import.import_id, + "artifact_sha256": evidence_import.artifact_sha256, + "imported_by": evidence_import.imported_by, + "test_case_id": case.test_case_id, + "status": case.status.value, + "suite_name": case.suite_name, + "test_name": case.test_name, + } + for evidence_import in junit_imports + for case in _junit_mapping_cases( + evidence_import, criterion.criterion_id + ) + ], + ensure_ascii=False, + sort_keys=True, + ), + "junit_evidence_boundary": _csv_text( + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION + ) + if junit_imports + else "", + "junit_importers": json.dumps( + [item.imported_by for item in junit_imports], + ensure_ascii=False, + ), + "junit_parser_warnings": json.dumps( + [ + { + "import_id": item.import_id, + "artifact_sha256": item.artifact_sha256, + "warning": warning, + } + for item in junit_imports + for warning in item.parser_warnings + ], + ensure_ascii=False, + sort_keys=True, + ), + "junit_limitations": json.dumps( + [ + { + "import_id": item.import_id, + "artifact_sha256": item.artifact_sha256, + "limitation": limitation, + } + for item in junit_imports + for limitation in item.limitations + ], + ensure_ascii=False, + sort_keys=True, + ), } ) return output.getvalue() @@ -1153,6 +1406,7 @@ def export_html(value: ExportableReview) -> str: if bundle.runtime_evidence else [] ), + *_junit_import_html(bundle), "

Runtime Verification Boundary

", "

" + ( diff --git a/scopeproof_core/reviews/__init__.py b/scopeproof_core/reviews/__init__.py index ff7ec0c3..2f4e5600 100644 --- a/scopeproof_core/reviews/__init__.py +++ b/scopeproof_core/reviews/__init__.py @@ -8,6 +8,7 @@ from scopeproof_core.reviews.lifecycle import ( ResolutionEventStatus, append_external_verification, + append_junit_evidence_import, append_resolution, append_runtime_evidence, attach_analysis, @@ -23,6 +24,7 @@ "ResolutionEventStatus", "ReviewComparison", "append_external_verification", + "append_junit_evidence_import", "append_resolution", "append_runtime_evidence", "attach_analysis", diff --git a/scopeproof_core/reviews/comparison.py b/scopeproof_core/reviews/comparison.py index 9f751b08..2b3b9f2a 100644 --- a/scopeproof_core/reviews/comparison.py +++ b/scopeproof_core/reviews/comparison.py @@ -16,6 +16,8 @@ FindingStatus, GateVerdict, HumanDecision, + JUnitEvidenceBoundary, + JUnitEvidenceImport, ReviewBundle, ReviewInputOrigin, ) @@ -143,12 +145,111 @@ class ResolutionChange(BaseModel): current_decision: HumanDecision | None +class JUnitImportChangeKind(StrEnum): + """Relationship between one bounded imported artifact across reviews.""" + + UNCHANGED = "unchanged" + MAPPING_MODIFIED = "mapping_modified" + ADDED = "added" + REMOVED = "removed" + + +class JUnitMappingReference(BaseModel): + """Validated comparison projection of an explicit human mapping.""" + + criterion_id: str = Field(min_length=1) + test_case_ids: list[str] = Field(min_length=1) + + @field_validator("test_case_ids") + @classmethod + def _case_ids_are_canonical(cls, value: list[str]) -> list[str]: + if value != sorted(set(value)): + raise ValueError("JUnit comparison case IDs must be sorted and unique") + return value + + +class JUnitImportReference(BaseModel): + """Sanitized immutable reference used only for comparison reporting.""" + + import_id: str = Field(min_length=1) + artifact_sha256: str = Field(pattern=r"^[0-9a-f]{64}$") + head_sha: str = Field(pattern=r"^[0-9a-f]{40}$") + asserted_importer: str = Field(min_length=1) + evidence_boundary: JUnitEvidenceBoundary = Field( + default_factory=JUnitEvidenceBoundary + ) + mappings: list[JUnitMappingReference] = Field(min_length=1) + + @classmethod + def from_import(cls, evidence_import: JUnitEvidenceImport) -> JUnitImportReference: + """Project only sanitized identity and explicit mappings.""" + + return cls( + import_id=evidence_import.import_id, + artifact_sha256=evidence_import.artifact_sha256, + head_sha=evidence_import.head_sha, + asserted_importer=evidence_import.imported_by, + evidence_boundary=evidence_import.evidence_boundary.model_copy(deep=True), + mappings=[ + JUnitMappingReference( + criterion_id=mapping.criterion_id, + test_case_ids=mapping.test_case_ids, + ) + for mapping in evidence_import.criterion_mappings + ], + ) + + +class JUnitImportChange(BaseModel): + """One artifact-digest relationship with explicit mapping projections.""" + + artifact_sha256: str = Field(pattern=r"^[0-9a-f]{64}$") + kind: JUnitImportChangeKind + previous: JUnitImportReference | None = None + current: JUnitImportReference | None = None + + @model_validator(mode="after") + def _references_match_kind_and_digest(self) -> JUnitImportChange: + if self.kind in { + JUnitImportChangeKind.UNCHANGED, + JUnitImportChangeKind.MAPPING_MODIFIED, + } and (self.previous is None or self.current is None): + raise ValueError("paired JUnit import changes require both references") + if self.kind is JUnitImportChangeKind.ADDED and ( + self.previous is not None or self.current is None + ): + raise ValueError("added JUnit import changes require only a current reference") + if self.kind is JUnitImportChangeKind.REMOVED and ( + self.previous is None or self.current is not None + ): + raise ValueError("removed JUnit import changes require only a previous reference") + for reference in (self.previous, self.current): + if reference is not None and reference.artifact_sha256 != self.artifact_sha256: + raise ValueError("JUnit import change digest must match its references") + if ( + self.kind is JUnitImportChangeKind.UNCHANGED + and self.previous is not None + and self.current is not None + and self.previous.mappings != self.current.mappings + ): + raise ValueError("unchanged JUnit imports require identical mappings") + if ( + self.kind is JUnitImportChangeKind.MAPPING_MODIFIED + and self.previous is not None + and self.current is not None + and self.previous.mappings == self.current.mappings + ): + raise ValueError("mapping-modified JUnit imports require changed mappings") + return self + + class ReviewComparison(BaseModel): previous_head_sha: str current_head_sha: str evidence_changes: list[EvidenceChange] changed_finding_statuses: list[FindingStatusChange] changed_human_resolutions: list[ResolutionChange] + junit_import_changes: list[JUnitImportChange] = Field(default_factory=list) previous_gate: GateVerdict current_gate: GateVerdict ruleset_version_changed: bool @@ -422,6 +523,56 @@ def _resolution_decisions(bundle: ReviewBundle) -> dict[str, HumanDecision]: return {resolution.criterion_id: resolution.decision for resolution in bundle.resolutions} +def _compare_junit_imports( + previous_items: list[JUnitEvidenceImport], + current_items: list[JUnitEvidenceImport], +) -> list[JUnitImportChange]: + previous_by_digest = {item.artifact_sha256: item for item in previous_items} + current_by_digest = {item.artifact_sha256: item for item in current_items} + changes: list[JUnitImportChange] = [] + for digest in sorted(set(previous_by_digest) | set(current_by_digest)): + previous_item = previous_by_digest.get(digest) + current_item = current_by_digest.get(digest) + previous_reference = ( + JUnitImportReference.from_import(previous_item) + if previous_item is not None + else None + ) + current_reference = ( + JUnitImportReference.from_import(current_item) + if current_item is not None + else None + ) + if previous_reference is None: + kind = JUnitImportChangeKind.ADDED + elif current_reference is None: + kind = JUnitImportChangeKind.REMOVED + elif previous_reference.mappings == current_reference.mappings: + kind = JUnitImportChangeKind.UNCHANGED + else: + kind = JUnitImportChangeKind.MAPPING_MODIFIED + changes.append( + JUnitImportChange( + artifact_sha256=digest, + kind=kind, + previous=previous_reference, + current=current_reference, + ) + ) + return changes + + +def _changed_junit_criteria(changes: list[JUnitImportChange]) -> set[str]: + return { + mapping.criterion_id + for change in changes + if change.kind is not JUnitImportChangeKind.UNCHANGED + for reference in (change.previous, change.current) + if reference is not None + for mapping in reference.mappings + } + + def _criteria_source_identity(bundle: ReviewBundle) -> tuple[str, str | None, str, str]: provenance = bundle.review.criteria_source_provenance if provenance is None: @@ -498,6 +649,9 @@ def compare_reviews(previous: ReviewBundle, current: ReviewBundle) -> ReviewComp if previous_resolutions.get(criterion_id) != current_resolutions.get(criterion_id) ] evidence_changes = _compare_evidence(previous.evidence, current.evidence) + junit_import_changes = _compare_junit_imports( + previous.junit_evidence_imports, current.junit_evidence_imports + ) changed_criterion_ids = { change.criterion_id for change in evidence_changes @@ -506,6 +660,7 @@ def compare_reviews(previous: ReviewBundle, current: ReviewBundle) -> ReviewComp changed_criterion_ids.update( change.criterion_id for change in changed_findings ) + changed_criterion_ids.update(_changed_junit_criteria(junit_import_changes)) if ( previous.review.head_sha != current.review.head_sha or previous.review.ruleset_version != current.review.ruleset_version @@ -517,6 +672,7 @@ def compare_reviews(previous: ReviewBundle, current: ReviewBundle) -> ReviewComp evidence_changes=evidence_changes, changed_finding_statuses=changed_findings, changed_human_resolutions=changed_resolutions, + junit_import_changes=junit_import_changes, previous_gate=previous.gate.verdict, current_gate=current.gate.verdict, ruleset_version_changed=( diff --git a/scopeproof_core/reviews/lifecycle.py b/scopeproof_core/reviews/lifecycle.py index 4b4dd790..5ddcba37 100644 --- a/scopeproof_core/reviews/lifecycle.py +++ b/scopeproof_core/reviews/lifecycle.py @@ -12,6 +12,7 @@ from scopeproof_core.resolution_events import current_resolutions, final_acceptance from scopeproof_core.review_policy import acceptance_requires_comment from scopeproof_core.schemas.models import ( + MAX_JUNIT_IMPORTS_PER_REVIEW, CheckState, CriteriaRevision, CriteriaSourceProvenance, @@ -19,10 +20,12 @@ EvidenceLevel, HumanDecision, IngestionState, + JUnitEvidenceImport, ResolutionEvent, ReviewBundle, ReviewState, RuntimeEvidence, + normalized_criteria_sha256, ) @@ -48,6 +51,8 @@ def new_review_state(bundle: ReviewBundle) -> ReviewState: raise ValueError("initial analysis bundle must not contain human resolutions") if bundle.review.final_acceptance: raise ValueError("initial analysis bundle must not contain final acceptance") + if bundle.junit_evidence_imports: + raise ValueError("initial analysis bundle must not contain JUnit imports") bundle = validated_review_bundle(bundle) if bundle.review.criteria_source_provenance is None: raise ValueError( @@ -162,6 +167,8 @@ def attach_analysis(state: ReviewState, bundle: ReviewBundle) -> ReviewState: raise ValueError("attached analysis must not contain human resolutions") if bundle.review.final_acceptance: raise ValueError("attached analysis must not contain final acceptance") + if bundle.junit_evidence_imports: + raise ValueError("attached analysis must not contain JUnit imports") bundle = validated_review_bundle(bundle) if ( state.criteria_revision.source_provenance is None @@ -334,6 +341,86 @@ def append_runtime_evidence(state: ReviewState, evidence: RuntimeEvidence) -> Re return validated_review_state(state.model_copy(update={"bundle": bundle})) +def append_junit_evidence_import( + state: ReviewState, + evidence_import: JUnitEvidenceImport, +) -> ReviewState: + """Append non-gating external test context to one exact active review.""" + + state = _validated_state(state) + if state.bundle is None: + raise ValueError("JUnit import requires an active analysis") + bundle = state.bundle + if len(bundle.junit_evidence_imports) >= MAX_JUNIT_IMPORTS_PER_REVIEW: + raise ValueError( + f"review may contain at most {MAX_JUNIT_IMPORTS_PER_REVIEW} JUnit imports" + ) + evidence_import = JUnitEvidenceImport.model_validate( + evidence_import.model_dump(mode="python") + ) + if ( + evidence_import.repository, + evidence_import.pr_number, + evidence_import.head_sha, + ) != ( + state.review.repository, + state.review.pr_number, + state.review.head_sha, + ): + raise ValueError("JUnit import must match the active review identity") + if evidence_import.criteria_revision_number != state.criteria_revision.number: + raise ValueError("JUnit import criteria revision must match the active revision") + if evidence_import.confirmed_criteria_sha256 != normalized_criteria_sha256( + state.criteria_revision.criteria + ): + raise ValueError("JUnit import criteria digest must match the active revision") + if ( + state.criteria_revision.source_provenance is None + or evidence_import.criteria_source_provenance + != state.criteria_revision.source_provenance + ): + raise ValueError("JUnit import criteria provenance must match the active revision") + known_criteria = { + criterion.criterion_id for criterion in state.criteria_revision.criteria + } + if any( + mapping.criterion_id not in known_criteria + for mapping in evidence_import.criterion_mappings + ): + raise ValueError("JUnit import mappings must reference active criteria") + if any( + existing.artifact_sha256 == evidence_import.artifact_sha256 + for existing in bundle.junit_evidence_imports + ): + raise ValueError("JUnit artifact is already imported") + if any( + existing.import_id == evidence_import.import_id + for existing in bundle.junit_evidence_imports + ): + raise ValueError("JUnit import ID is already recorded") + + unchanged_gate = bundle.gate.model_copy(deep=True) + unchanged_findings = [item.model_copy(deep=True) for item in bundle.findings] + unchanged_resolutions = [item.model_copy(deep=True) for item in bundle.resolutions] + unchanged_runtime = [item.model_copy(deep=True) for item in bundle.runtime_evidence] + unchanged_events = [item.model_copy(deep=True) for item in state.resolution_events] + unchanged_final_acceptance = state.review.final_acceptance + updated_bundle = bundle.model_copy(deep=True) + updated_bundle.junit_evidence_imports.append(evidence_import.model_copy(deep=True)) + updated = validated_review_state(state.model_copy(update={"bundle": updated_bundle})) + assert updated.bundle is not None + if ( + updated.bundle.gate != unchanged_gate + or updated.bundle.findings != unchanged_findings + or updated.bundle.resolutions != unchanged_resolutions + or updated.bundle.runtime_evidence != unchanged_runtime + or updated.resolution_events != unchanged_events + or updated.review.final_acceptance is not unchanged_final_acceptance + ): + raise ValueError("JUnit import must not alter deterministic or human review truth") + return updated + + def append_external_verification( state: ReviewState, evidence: RuntimeEvidence, diff --git a/scopeproof_core/schemas/models.py b/scopeproof_core/schemas/models.py index 91ea37e0..74210f4c 100644 --- a/scopeproof_core/schemas/models.py +++ b/scopeproof_core/schemas/models.py @@ -106,9 +106,21 @@ def require_verified_public_origin( _SHA256_PATTERN = re.compile(r"^[a-f0-9]{64}$") +_EXACT_HEAD_PATTERN = r"^[a-f0-9]{40}$" +MAX_JUNIT_IMPORTS_PER_REVIEW = 20 +MAX_JUNIT_NOTE_LENGTH = 1_000 +_PATH_OR_URI_LIKE = re.compile( + r"(?:[/\\]|(? bool: + """Return whether a JUnit display name could disclose a path or URI.""" + + return _PATH_OR_URI_LIKE.search(value) is not None _CRITERIA_SOURCE_URI_ERROR = ( "source URI must be an HTTPS URL or " "scopeproof://constructed-demo/acceptance-criteria" @@ -1090,6 +1102,263 @@ def validate_manual_level(self) -> RuntimeEvidence: return self +class JUnitCaseStatus(StringEnum): + """Sanitized externally supplied JUnit result state.""" + + PASSED = "passed" + FAILURE = "failure" + ERROR = "error" + SKIPPED = "skipped" + + +JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION = ( + "Imported test results are externally supplied, non-gating context. " + "ScopeProof did not execute the tests or target-repository code; the artifact " + "digest covers imported bytes only; importer identity is asserted, not " + "authenticated; and criterion mapping is organizational context, not proof." +) + + +class JUnitCaseResult(BaseModel): + """One bounded result projection without raw XML or output bodies.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + test_case_id: str = Field(pattern=r"^suite-\d{4}-case-\d{4}$") + suite_id: str = Field(pattern=r"^suite-\d{4}$") + suite_name: str = Field(max_length=512) + class_name: str | None = Field(default=None, max_length=512) + test_name: str = Field(max_length=512) + status: JUnitCaseStatus + + @field_validator("suite_name", "test_name") + @classmethod + def require_non_blank_names(cls, value: str) -> str: + normalized = value.strip() + if not normalized: + raise ValueError("must contain non-whitespace text") + if junit_name_is_path_or_url_like(normalized): + raise ValueError("JUnit names must not contain path- or URL-like text") + return normalized + + @field_validator("class_name") + @classmethod + def normalize_optional_class_name(cls, value: str | None) -> str | None: + if value is None: + return None + normalized = value.strip() + if not normalized: + raise ValueError("class name must contain non-whitespace text") + if junit_name_is_path_or_url_like(normalized): + raise ValueError("JUnit names must not contain path- or URL-like text") + return normalized + + @model_validator(mode="after") + def require_case_to_belong_to_suite(self) -> JUnitCaseResult: + if not self.test_case_id.startswith(f"{self.suite_id}-case-"): + raise ValueError("JUnit test case ID must belong to its suite ID") + return self + + +class JUnitResultTotals(BaseModel): + """Computed JUnit result totals; artifact-declared totals are not trusted.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + total: int = Field(ge=0) + passed: int = Field(ge=0) + failures: int = Field(ge=0) + errors: int = Field(ge=0) + skipped: int = Field(ge=0) + + @model_validator(mode="after") + def require_categories_to_sum_to_total(self) -> JUnitResultTotals: + categorized = self.passed + self.failures + self.errors + self.skipped + if categorized != self.total: + raise ValueError("JUnit result categories must sum to total") + return self + + +class JUnitCriterionMapping(BaseModel): + """Explicit human mapping from sanitized test cases to one criterion.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + criterion_id: str = Field(min_length=1) + test_case_ids: list[str] = Field(min_length=1) + + @field_validator("criterion_id") + @classmethod + def require_non_blank_criterion_id(cls, value: str) -> str: + normalized = value.strip() + if not normalized: + raise ValueError("criterion ID must contain non-whitespace text") + return normalized + + @field_validator("test_case_ids") + @classmethod + def require_canonical_case_ids(cls, value: list[str]) -> list[str]: + if value != sorted(set(value)): + raise ValueError("mapped test case IDs must be sorted and unique") + if any(re.fullmatch(r"suite-\d{4}-case-\d{4}", item) is None for item in value): + raise ValueError("mapped test case IDs must use stable JUnit case IDs") + return value + + +class JUnitEvidenceBoundary(BaseModel): + """Fixed machine-readable trust semantics for every external JUnit import.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + source: Literal["externally_supplied"] = "externally_supplied" + gate_effect: Literal["non_gating"] = "non_gating" + execution: Literal["not_executed_by_scopeproof"] = "not_executed_by_scopeproof" + artifact_digest_scope: Literal["imported_bytes_only"] = "imported_bytes_only" + importer_identity: Literal["asserted_not_authenticated"] = ( + "asserted_not_authenticated" + ) + criterion_mapping: Literal["organizational_context_not_proof"] = ( + "organizational_context_not_proof" + ) + + +class JUnitEvidenceImport(BaseModel): + """Versioned external test-result context that never enters gate truth.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + schema_version: Literal["junit-import-v1"] + evidence_boundary: JUnitEvidenceBoundary + import_id: LocalReviewId + repository: str = Field(pattern=GITHUB_REPOSITORY_PATTERN) + pr_number: int = Field(gt=0) + head_sha: str = Field(pattern=_EXACT_HEAD_PATTERN) + criteria_revision_number: Annotated[StrictInt, Field(gt=0)] + confirmed_criteria_sha256: str + criteria_source_provenance: CriteriaSourceProvenance + artifact_sha256: str + artifact_format: Literal["junit_xml"] = "junit_xml" + imported_by: str = Field(max_length=256) + imported_at: datetime + totals: JUnitResultTotals + test_cases: list[JUnitCaseResult] = Field(min_length=1, max_length=5_000) + criterion_mappings: list[JUnitCriterionMapping] = Field(min_length=1) + parser_warnings: list[str] = Field(default_factory=list, max_length=100) + limitations: list[str] = Field(min_length=1, max_length=100) + + @field_validator("confirmed_criteria_sha256", "artifact_sha256") + @classmethod + def validate_sha256_digest(cls, value: str) -> str: + if not _SHA256_PATTERN.fullmatch(value): + raise ValueError("must be a lowercase SHA-256 digest") + return value + + @field_validator("imported_by") + @classmethod + def normalize_importer(cls, value: str) -> str: + normalized = value.strip() + if not normalized: + raise ValueError("imported_by must contain non-whitespace text") + return normalized + + @field_validator("imported_at") + @classmethod + def normalize_import_timestamp(cls, value: datetime) -> datetime: + if value.tzinfo is None or value.utcoffset() is None: + raise ValueError("imported_at must be timezone-aware") + return value.astimezone(UTC) + + @field_validator("parser_warnings", "limitations") + @classmethod + def require_unique_non_blank_notes(cls, value: list[str]) -> list[str]: + normalized = [item.strip() for item in value] + if any(not item for item in normalized): + raise ValueError("notes must contain non-whitespace text") + if any(len(item) > MAX_JUNIT_NOTE_LENGTH for item in normalized): + raise ValueError( + f"notes must be at most {MAX_JUNIT_NOTE_LENGTH} characters each" + ) + if len(normalized) != len(set(normalized)): + raise ValueError("notes must be unique") + return normalized + + @model_validator(mode="after") + def validate_result_and_mapping_cross_references(self) -> JUnitEvidenceImport: + case_ids = [item.test_case_id for item in self.test_cases] + if case_ids != sorted(set(case_ids)): + raise ValueError("JUnit test case IDs must be sorted and unique") + observed = { + JUnitCaseStatus.PASSED: 0, + JUnitCaseStatus.FAILURE: 0, + JUnitCaseStatus.ERROR: 0, + JUnitCaseStatus.SKIPPED: 0, + } + for item in self.test_cases: + observed[item.status] += 1 + if ( + self.totals.total, + self.totals.passed, + self.totals.failures, + self.totals.errors, + self.totals.skipped, + ) != ( + len(self.test_cases), + observed[JUnitCaseStatus.PASSED], + observed[JUnitCaseStatus.FAILURE], + observed[JUnitCaseStatus.ERROR], + observed[JUnitCaseStatus.SKIPPED], + ): + raise ValueError("JUnit totals must match sanitized test case results") + mapping_criteria = [item.criterion_id for item in self.criterion_mappings] + if mapping_criteria != sorted(set(mapping_criteria)): + raise ValueError("JUnit criterion mappings must be sorted and unique") + known_case_ids = set(case_ids) + if any( + case_id not in known_case_ids + for mapping in self.criterion_mappings + for case_id in mapping.test_case_ids + ): + raise ValueError("mapped test case IDs must resolve") + mapped_case_ids = [ + case_id + for mapping in self.criterion_mappings + for case_id in mapping.test_case_ids + ] + if len(mapped_case_ids) != len(set(mapped_case_ids)): + raise ValueError("one JUnit test case must not map to multiple criteria") + return self + + +class JUnitImportMutationMetadata(BaseModel): + """Validated CLI output for one persisted JUnit import mutation.""" + + model_config = ConfigDict(extra="forbid", frozen=True) + + review_id: LocalReviewId + record: str = Field(min_length=1) + head_sha: str = Field(pattern=_EXACT_HEAD_PATTERN) + import_id: LocalReviewId + artifact_sha256: str + mapped_criterion_ids: list[str] = Field(min_length=1) + totals: JUnitResultTotals + evidence_boundary: Literal["externally_supplied_non_gating"] + verdict: GateVerdict + + @field_validator("artifact_sha256") + @classmethod + def validate_artifact_digest(cls, value: str) -> str: + if not _SHA256_PATTERN.fullmatch(value): + raise ValueError("must be a lowercase SHA-256 digest") + return value + + @field_validator("mapped_criterion_ids") + @classmethod + def require_canonical_criteria(cls, value: list[str]) -> list[str]: + if value != sorted(set(value)) or any(not item.strip() for item in value): + raise ValueError("mapped criterion IDs must be sorted unique IDs") + return value + + class Finding(BaseModel): criterion_id: str status: FindingStatus @@ -1295,6 +1564,9 @@ class ReviewBundle(BaseModel): default_factory=list ) runtime_evidence: list[RuntimeEvidence] = Field(default_factory=list) + junit_evidence_imports: list[JUnitEvidenceImport] = Field( + default_factory=list, max_length=MAX_JUNIT_IMPORTS_PER_REVIEW + ) findings: list[Finding] resolutions: list[HumanResolution] = Field(default_factory=list) gate: GateDecision @@ -1437,6 +1709,38 @@ def validate_cross_references(self) -> ReviewBundle: if item.runtime_evidence_id is not None } + junit_import_ids = [item.import_id for item in self.junit_evidence_imports] + if len(junit_import_ids) != len(set(junit_import_ids)): + raise ValueError("JUnit import IDs must be unique") + junit_artifact_digests = [ + item.artifact_sha256 for item in self.junit_evidence_imports + ] + if len(junit_artifact_digests) != len(set(junit_artifact_digests)): + raise ValueError("JUnit artifact digests must be unique") + criteria_digest = normalized_criteria_sha256(self.criteria) + for item in self.junit_evidence_imports: + if ( + item.repository, + item.pr_number, + item.head_sha, + ) != ( + self.review.repository, + self.review.pr_number, + self.review.head_sha, + ): + raise ValueError("JUnit import identity must match the owning review") + if item.criteria_revision_number != self.criteria_revision_number: + raise ValueError("JUnit import criteria revision must match the bundle") + if item.confirmed_criteria_sha256 != criteria_digest: + raise ValueError("JUnit import criteria digest must match the bundle") + if item.criteria_source_provenance != self.review.criteria_source_provenance: + raise ValueError("JUnit import criteria provenance must match the review") + if any( + mapping.criterion_id not in known_criteria + for mapping in item.criterion_mappings + ): + raise ValueError("JUnit import mappings must reference known criteria") + resolution_ids = [resolution.criterion_id for resolution in self.resolutions] if len(resolution_ids) != len(set(resolution_ids)): raise ValueError("resolution criterion IDs must be unique") diff --git a/tests/apps/test_streamlit_app.py b/tests/apps/test_streamlit_app.py index 191b330b..06d5de7b 100644 --- a/tests/apps/test_streamlit_app.py +++ b/tests/apps/test_streamlit_app.py @@ -21,12 +21,18 @@ from scopeproof_core.demo import load_demo_snapshot from scopeproof_core.gates.evaluator import evaluate_gate from scopeproof_core.github.client import GitHubNetworkError, GitHubPaginationError +from scopeproof_core.importers.junit import ( + JUnitMappingSelection, + build_junit_evidence_import, +) from scopeproof_core.reviews.lifecycle import ( append_external_verification, + append_junit_evidence_import, append_resolution, ) from scopeproof_core.schemas.models import ( CONSTRUCTED_DEMO_CRITERIA_SOURCE_URI, + JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, RULESET_VERSION, CheckState, CIObservation, @@ -3545,6 +3551,55 @@ def test_same_head_reanalysis_exposes_unchanged_candidates_and_comparison_export assert all(not button.disabled for button in comparison_downloads.values()) +def test_comparison_view_shows_removed_external_junit_import_as_non_gating() -> None: + app = analyzed_exact_head_standard_demo(new_app()) + current_state = app.session_state["review_state"] + criterion_id = current_state.criteria_revision.criteria[0].criterion_id + current_state = append_resolution( + current_state, + ResolutionEvent( + criterion_id=criterion_id, + decision=HumanDecision.ACCEPTED, + comment="Reviewed before the imported context changed.", + ), + ) + imported = build_junit_evidence_import( + current_state, + b'', + [ + JUnitMappingSelection( + scope_id="suite-0001", + criterion_id=criterion_id, + ) + ], + importer="QA owner", + import_id="junit-comparison-import", + ) + previous_state = append_junit_evidence_import(current_state, imported) + assert previous_state.bundle is not None + app.session_state["comparison_base_bundle"] = previous_state.bundle + app.session_state["review_state"] = current_state + app.session_state["bundle"] = current_state.bundle + + app = app.run() + + assert app.exception == [] + rendered = "\n".join( + item.value for item in [*app.markdown, *app.caption, *app.text, *app.code] + ) + assert "Imported external test result changes" in rendered + assert "Removed" in rendered + assert imported.artifact_sha256 in rendered + assert "externally supplied, non-gating context" in rendered.lower() + assert "did not execute" in rendered.lower() + assert "imported bytes only" in rendered.lower() + assert "asserted, not authenticated" in rendered.lower() + assert "organizational context, not proof" in rendered.lower() + warnings = "\n".join(item.value for item in app.warning) + assert "does not carry a prior decision forward automatically" in warnings + assert "changed head" not in warnings + + def test_ineligible_comparison_base_is_cleared_without_hiding_current_analysis() -> None: app = analyzed_demo(new_app()) current_bundle = app.session_state["bundle"].model_copy(deep=True) @@ -5123,3 +5178,175 @@ def test_successful_runtime_evidence_save_clears_form_and_prevents_accidental_re assert "External verification and reviewer decision recorded together." in [ item.value for item in app.success ] + + +def test_junit_import_maps_uploaded_suite_without_changing_gate_or_decisions() -> None: + app = analyzed_exact_head_standard_demo(new_app()) + before = app.session_state["review_state"].model_copy(deep=True) + assert before.bundle is not None + assert ( + app.file_uploader(key="junit_artifact_upload").proto.max_upload_size_mb + == 1 + ) + + app = app.file_uploader(key="junit_artifact_upload").upload( + "results.xml", + b'', + "application/xml", + ).run() + app = app.text_input(key="junit_importer").set_value("QA owner").run() + app = app.multiselect(key="junit_mapping_scopes").set_value( + ["suite-0001"] + ).run() + app = app.button(key="save_junit_import").click().run() + + updated = app.session_state["review_state"] + assert updated.bundle is not None + assert len(updated.bundle.junit_evidence_imports) == 1 + imported = updated.bundle.junit_evidence_imports[0] + assert imported.criterion_mappings[0].criterion_id == app.session_state[ + "selected_criterion" + ] + assert imported.test_cases[0].test_case_id == "suite-0001-case-0001" + assert updated.bundle.gate == before.bundle.gate + assert updated.bundle.resolutions == before.bundle.resolutions + assert updated.bundle.runtime_evidence == before.bundle.runtime_evidence + assert updated.review.final_acceptance is before.review.final_acceptance + assert JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION in [ + item.value for item in app.caption + ] + assert app.file_uploader(key="junit_artifact_upload_1").value is None + assert app.text_input(key="junit_importer").value == "" + assert app.multiselect(key="junit_mapping_scopes").value == [] + assert app.button(key="save_junit_import").disabled is True + + +def test_replacing_junit_upload_clears_mapping_even_when_scope_ids_match() -> None: + app = analyzed_exact_head_standard_demo(new_app()) + app = app.file_uploader(key="junit_artifact_upload").upload( + "first.xml", + b'', + "application/xml", + ).run() + app = app.multiselect(key="junit_mapping_scopes").set_value( + ["suite-0001"] + ).run() + + app = app.file_uploader(key="junit_artifact_upload").upload( + "second.xml", + b'', + "application/xml", + ).run() + + assert app.multiselect(key="junit_mapping_scopes").value == [] + + +def test_junit_preview_and_saved_values_render_inertly_without_raw_output() -> None: + app = analyzed_exact_head_standard_demo(new_app()) + hostile_name = "" + raw_output = "RAW-JUNIT-OUTPUT-SENTINEL" + xml = ( + f'", ">")}">' + f'{raw_output}' + ).encode() + + app = app.file_uploader(key="junit_artifact_upload").upload( + "results.xml", xml, "application/xml" + ).run() + + assert app.exception == [] + assert hostile_name not in [item.value for item in app.markdown] + assert raw_output not in "\n".join( + [ + *(item.value for item in app.text), + *(item.value for item in app.code), + *(item.value for item in app.caption), + *(item.value for item in app.markdown), + ] + ) + assert "JUnit properties and output content were discarded during import." in [ + item.value for item in app.text + ] + + +def test_junit_parser_failure_leaves_review_unchanged_and_hides_artifact_text() -> None: + app = analyzed_exact_head_standard_demo(new_app()) + before = app.session_state["review_state"].model_copy(deep=True) + unsafe = ( + b']>' + b'' + ) + + app = app.file_uploader(key="junit_artifact_upload").upload( + "unsafe.xml", unsafe, "application/xml" + ).run() + + assert app.session_state["review_state"] == before + rendered_errors = "\n".join(item.value for item in app.error) + assert "could not be inspected" in rendered_errors + assert "PRIVATE-SENTINEL" not in rendered_errors + assert app.button(key="save_junit_import").disabled is True + + +def test_junit_import_requires_exact_head_and_explicit_mapping() -> None: + app = analyzed_demo(new_app()) + app = app.file_uploader(key="junit_artifact_upload").upload( + "results.xml", + b'', + "application/xml", + ).run() + app = app.text_input(key="junit_importer").set_value("QA owner").run() + + assert app.button(key="save_junit_import").disabled is True + assert "An exact 40-character reviewed head is required before import." in [ + item.value for item in app.caption + ] + + +@pytest.mark.parametrize( + ("xml", "importer", "mapping"), + [ + (b'', "QA owner", ["suite-0001"]), + ( + b'', + "x" * 257, + ["suite-0001"], + ), + ], +) +def test_junit_import_save_stays_disabled_until_full_draft_is_valid( + xml: bytes, + importer: str, + mapping: list[str], +) -> None: + app = analyzed_exact_head_standard_demo(new_app()) + app = app.file_uploader(key="junit_artifact_upload").upload( + "results.xml", xml, "application/xml" + ).run() + app = app.text_input(key="junit_importer").set_value(importer).run() + app = app.multiselect(key="junit_mapping_scopes").set_value(mapping).run() + + assert app.button(key="save_junit_import").disabled is True + + +def test_junit_import_save_stays_disabled_for_already_imported_artifact() -> None: + xml = b'' + app = analyzed_exact_head_standard_demo(new_app()) + app = app.file_uploader(key="junit_artifact_upload").upload( + "results.xml", xml, "application/xml" + ).run() + app = app.text_input(key="junit_importer").set_value("QA owner").run() + app = app.multiselect(key="junit_mapping_scopes").set_value( + ["suite-0001"] + ).run() + app = app.button(key="save_junit_import").click().run() + + app = app.file_uploader(key="junit_artifact_upload_1").upload( + "results.xml", xml, "application/xml" + ).run() + app = app.text_input(key="junit_importer").set_value("QA owner").run() + app = app.multiselect(key="junit_mapping_scopes").set_value( + ["suite-0001"] + ).run() + + assert app.button(key="save_junit_import").disabled is True diff --git a/tests/apps/test_web_launcher.py b/tests/apps/test_web_launcher.py index f653317a..320ec058 100644 --- a/tests/apps/test_web_launcher.py +++ b/tests/apps/test_web_launcher.py @@ -56,7 +56,7 @@ def fake_run(command: list[str], *, check: bool) -> subprocess.CompletedProcess[ ) assert len(calls) == 1 command, check = calls[0] - assert command[:12] == [ + assert command[:13] == [ sys.executable, "-m", "streamlit", @@ -64,13 +64,14 @@ def fake_run(command: list[str], *, check: bool) -> subprocess.CompletedProcess[ "--server.address=127.0.0.2", "--server.port=8765", "--server.headless=false", + "--server.maxUploadSize=2", "--theme.base=dark", "--theme.primaryColor=#d8ff63", "--theme.backgroundColor=#0d0f12", "--theme.secondaryBackgroundColor=#171a1f", "--theme.textColor=#f7f7f2", ] - assert Path(command[12]).resolve() == Path("apps/web/app.py").resolve() + assert Path(command[13]).resolve() == Path("apps/web/app.py").resolve() assert check is False diff --git a/tests/browser/test_packaged_workbench.py b/tests/browser/test_packaged_workbench.py index 2f386d1f..660c769e 100644 --- a/tests/browser/test_packaged_workbench.py +++ b/tests/browser/test_packaged_workbench.py @@ -1,11 +1,14 @@ from __future__ import annotations import os +import re import signal import socket import subprocess import sys import time +from collections.abc import Callable +from hashlib import sha256 from importlib.metadata import version from importlib.util import find_spec from pathlib import Path @@ -13,7 +16,19 @@ from urllib.request import urlopen import pytest -from playwright.sync_api import Locator, Page, Route, expect, sync_playwright +from playwright.sync_api import ( + Download, + Locator, + Page, + Route, + expect, + sync_playwright, +) +from playwright.sync_api import ( + TimeoutError as PlaywrightTimeoutError, +) + +from scopeproof_core.schemas.models import JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION pytestmark = pytest.mark.browser @@ -59,6 +74,40 @@ outlineWidth: parseFloat(style.outlineWidth), }; }""" +ARM_STREAMLIT_RERUN_OBSERVER_SCRIPT = """() => { + if (window.__scopeproofRerunObserver) { + window.__scopeproofRerunObserver.disconnect(); + } + window.__scopeproofRerunObserved = false; + const observer = new MutationObserver(mutations => { + for (const mutation of mutations) { + if ( + mutation.attributeName === "data-test-script-state" && + ( + mutation.oldValue === "notRunning" || + mutation.target.getAttribute("data-test-script-state") !== + "notRunning" + ) + ) { + window.__scopeproofRerunObserved = true; + } + } + }); + observer.observe(document.documentElement, { + attributes: true, + attributeFilter: ["data-test-script-state"], + attributeOldValue: true, + subtree: true, + }); + window.__scopeproofRerunObserver = observer; +}""" +STREAMLIT_RERUN_FINISHED_SCRIPT = """() => { + const app = document.querySelector('[data-testid="stApp"]'); + return ( + window.__scopeproofRerunObserved === true && + app?.getAttribute("data-test-script-state") === "notRunning" + ); +}""" def _run(*command: str, cwd: Path | None = None) -> None: @@ -223,6 +272,237 @@ def _activate_with_keyboard( page.keyboard.press(key) +def _run_and_wait_for_streamlit_rerun( + page: Page, action: Callable[[], None] +) -> None: + page.evaluate(ARM_STREAMLIT_RERUN_OBSERVER_SCRIPT) + action() + page.wait_for_function(STREAMLIT_RERUN_FINISHED_SCRIPT, timeout=10_000) + + +def _choose_combobox_option( + page: Page, + combobox: Locator, + *, + option_name: str, + settled_control: Locator | None = None, +) -> None: + for attempt in range(3): + value_is_selected = combobox.input_value() == option_name + control_is_settled = ( + settled_control is None or settled_control.is_enabled() + ) + if value_is_selected and control_is_settled: + return + try: + if value_is_selected: + combobox.fill("") + combobox.click() + combobox.fill(option_name) + option = page.get_by_role("option", name=option_name, exact=True) + expect(option).to_be_visible(timeout=5_000) + _run_and_wait_for_streamlit_rerun( + page, + lambda selected_option=option: selected_option.click(timeout=5_000), + ) + expect(combobox).to_have_value(option_name, timeout=5_000) + if settled_control is not None: + expect(settled_control).to_be_enabled(timeout=5_000) + return + except PlaywrightTimeoutError: + if attempt == 2: + raise + + +def _download_with_retry(page: Page, *, button_name: str) -> Download: + for attempt in range(3): + download_button = page.get_by_role( + "button", name=button_name, exact=True + ) + expect(download_button).to_be_visible(timeout=5_000) + expect(download_button).to_be_enabled(timeout=5_000) + try: + with page.expect_download(timeout=5_000) as download_info: + download_button.click() + return download_info.value + except PlaywrightTimeoutError: + if attempt == 2: + raise + raise AssertionError("download retry loop exhausted without a terminal result") + + +def _click_until_text_visible( + page: Page, *, button_name: str, outcome_text: str +) -> None: + outcome = page.get_by_text(outcome_text, exact=True) + for attempt in range(3): + if outcome.is_visible(): + return + button = page.get_by_role("button", name=button_name, exact=True) + try: + expect(button).to_be_visible(timeout=5_000) + expect(button).to_be_enabled(timeout=5_000) + button.click(timeout=5_000) + expect(outcome).to_be_visible(timeout=5_000) + return + except (AssertionError, PlaywrightTimeoutError): + if outcome.is_visible(): + return + if attempt == 2: + raise + + +def test_choose_combobox_option_retries_after_detached_option( + monkeypatch: pytest.MonkeyPatch, +) -> None: + class FakeCombobox: + def __init__(self) -> None: + self.value = "" + self.clicks = 0 + self.fills: list[str] = [] + + def input_value(self) -> str: + return self.value + + def click(self) -> None: + self.clicks += 1 + + def fill(self, value: str) -> None: + self.fills.append(value) + + class FakeControl: + def __init__(self) -> None: + self.enabled = False + + def is_enabled(self) -> bool: + return self.enabled + + class FakeOption: + def __init__(self, combobox: FakeCombobox, control: FakeControl) -> None: + self.combobox = combobox + self.control = control + self.clicks = 0 + + def click(self, *, timeout: int | None = None) -> None: + self.clicks += 1 + self.combobox.value = "junit-browser-review" + if self.clicks == 1: + raise PlaywrightTimeoutError("element was detached from the DOM") + self.control.enabled = True + + class FakePage: + def __init__(self, option: FakeOption) -> None: + self.option = option + + def evaluate(self, script: str) -> None: + return None + + def get_by_role(self, *args: object, **kwargs: object) -> FakeOption: + return self.option + + def wait_for_function(self, script: str, *, timeout: int) -> None: + return None + + class FakeExpectation: + def __init__(self, target: object) -> None: + self.target = target + + def to_be_visible(self, *, timeout: int | None = None) -> None: + return None + + def to_have_value(self, expected: str, *, timeout: int | None = None) -> None: + assert isinstance(self.target, FakeCombobox) + assert self.target.value == expected + + def to_be_enabled(self, *, timeout: int | None = None) -> None: + assert isinstance(self.target, FakeControl) + if not self.target.enabled: + raise PlaywrightTimeoutError("dependent control did not settle") + + combobox = FakeCombobox() + control = FakeControl() + option = FakeOption(combobox, control) + monkeypatch.setattr( + sys.modules[__name__], "expect", lambda target: FakeExpectation(target) + ) + + _choose_combobox_option( + FakePage(option), # type: ignore[arg-type] + combobox, # type: ignore[arg-type] + option_name="junit-browser-review", + settled_control=control, # type: ignore[arg-type] + ) + + assert option.clicks == 2 + assert combobox.clicks == 2 + assert combobox.fills == [ + "junit-browser-review", + "", + "junit-browser-review", + ] + assert control.enabled is True + + +def test_download_with_retry_reissues_click_after_missed_event( + monkeypatch: pytest.MonkeyPatch, +) -> None: + class FakeButton: + def __init__(self) -> None: + self.clicks = 0 + + def click(self) -> None: + self.clicks += 1 + + class FakeDownload: + suggested_filename = "review.md" + + class FakeDownloadEvent: + def __init__(self, attempt: int) -> None: + self.attempt = attempt + self.value = FakeDownload() + + def __enter__(self) -> FakeDownloadEvent: + return self + + def __exit__(self, *args: object) -> None: + if self.attempt == 1: + raise PlaywrightTimeoutError("download event was not observed") + + class FakePage: + def __init__(self, button: FakeButton) -> None: + self.button = button + self.download_attempts = 0 + + def get_by_role(self, *args: object, **kwargs: object) -> FakeButton: + return self.button + + def expect_download(self, *, timeout: int) -> FakeDownloadEvent: + self.download_attempts += 1 + return FakeDownloadEvent(self.download_attempts) + + class FakeExpectation: + def to_be_visible(self, *, timeout: int | None = None) -> None: + return None + + def to_be_enabled(self, *, timeout: int | None = None) -> None: + return None + + button = FakeButton() + page = FakePage(button) + monkeypatch.setattr( + sys.modules[__name__], "expect", lambda target: FakeExpectation() + ) + + download = _download_with_retry( + page, # type: ignore[arg-type] + button_name="Download Markdown", + ) + + assert download.suggested_filename == "review.md" + assert page.download_attempts == 2 + assert button.clicks == 2 + + def _exercise_primary_path( page: Page, base_url: str, *, verify_persistence_and_downloads: bool ) -> None: @@ -267,6 +547,9 @@ def _exercise_primary_path( expect(page.get_by_text("Missing evidence", exact=True).first).to_be_visible() expect(page.get_by_text("Review status: Action required", exact=True)).to_be_visible() expect(page.get_by_text("Evidence status:", exact=False).first).to_be_visible() + expect( + page.get_by_text("Import external JUnit results", exact=True) + ).to_be_visible() export_controls = ( ("Download Markdown", ".md"), ("Download JSON", ".json"), @@ -280,22 +563,27 @@ def _exercise_primary_path( if verify_persistence_and_downloads: save_notice = page.get_by_text("Review saved automatically. ID:", exact=False) expect(save_notice).to_be_visible() + saved_id_match = re.search( + r"Review saved automatically\. ID: ([A-Za-z0-9_-]+)\.", + save_notice.inner_text(), + ) + assert saved_id_match is not None + saved_review_id = saved_id_match.group(1) - markdown_export = page.get_by_role("button", name="Download Markdown", exact=True) - with page.expect_download() as download_info: - markdown_export.click() - download = download_info.value + download = _download_with_retry(page, button_name="Download Markdown") assert download.suggested_filename.endswith(".md") assert b"head-demo-002" in download.path().read_bytes() page.get_by_text("Resume a saved review", exact=True).click() - expect(page.get_by_text("saved local review found", exact=False)).to_be_visible() + expect(page.get_by_text(re.compile(r"saved local reviews? found"))).to_be_visible() saved_review = page.get_by_role("combobox", name="Saved review ID", exact=True) - saved_review.locator("..").get_by_role("button", name="Open", exact=True).click() - page.keyboard.press("ArrowDown") - page.keyboard.press("Enter") reopen = page.get_by_role("button", name="Reopen local review", exact=True) - expect(reopen).to_be_enabled() + _choose_combobox_option( + page, + saved_review, + option_name=saved_review_id, + settled_control=reopen, + ) reopen.click() expect( page.get_by_text( @@ -332,6 +620,116 @@ def _exercise_primary_path( assert page.evaluate("document.body.scrollWidth <= window.innerWidth") +def _exercise_junit_import_round_trip( + page: Page, + base_url: str, + *, + artifact_path: Path, + artifact_digest: str, +) -> None: + page.goto(base_url, wait_until="domcontentloaded") + page.get_by_text("Resume a saved review", exact=True).click() + saved_review = page.get_by_role("combobox", name="Saved review ID", exact=True) + reopen = page.get_by_role("button", name="Reopen local review", exact=True) + _choose_combobox_option( + page, + saved_review, + option_name="junit-browser-review", + settled_control=reopen, + ) + reopen.click() + expect( + page.get_by_text( + "Review reopened from local storage after validation.", exact=True + ) + ).to_be_visible() + + junit_expander = page.get_by_text("Import external JUnit results", exact=True) + junit_expander.click() + artifact_input = page.get_by_label( + "Local JUnit XML artifact", exact=True + ).locator("input[type=file]") + _run_and_wait_for_streamlit_rerun( + page, lambda: artifact_input.set_input_files(artifact_path) + ) + preview = page.get_by_text( + "Computed results: 1 total · 1 passed · 0 failed · 0 errors · 0 skipped", + exact=True, + ) + expect(preview).to_have_count(1) + if not preview.is_visible(): + junit_expander.click() + expect(preview).to_be_visible() + importer = page.get_by_label("Asserted JUnit importer (required)", exact=True) + importer.fill("Packaged browser reviewer") + _run_and_wait_for_streamlit_rerun(page, lambda: importer.press("Tab")) + if not preview.is_visible(): + junit_expander.click() + expect(preview).to_be_visible() + mapping = page.get_by_role( + "combobox", name="Map JUnit scopes to the selected criterion", exact=True + ) + expect(mapping).to_be_enabled() + mapping.click() + mapping.fill("suite-0001") + suite_option = page.get_by_text( + "suite-0001 · suite · unit", exact=True + ).last + expect(suite_option).to_be_visible() + _run_and_wait_for_streamlit_rerun(page, suite_option.click) + page.keyboard.press("Escape") + expect(preview).to_have_count(1) + if not preview.is_visible(): + junit_expander.click() + expect( + page.get_by_label("Asserted JUnit importer (required)", exact=True) + ).to_have_value("Packaged browser reviewer") + save = page.get_by_role( + "button", name="Save imported JUnit results", exact=True + ) + expect(save).to_be_enabled() + expect(save).to_be_visible() + _click_until_text_visible( + page, + button_name="Save imported JUnit results", + outcome_text=( + "Imported JUnit results appended as external non-gating context." + ), + ) + boundary = page.get_by_text(JUNIT_EVIDENCE_BOUNDARY_DESCRIPTION, exact=True) + if not boundary.is_visible(): + junit_expander.click() + expect(boundary).to_be_visible() + expect(page.get_by_text("Review saved automatically. ID:", exact=False)).to_be_visible() + + page.get_by_text("Resume a saved review", exact=True).click() + saved_review = page.get_by_role("combobox", name="Saved review ID", exact=True) + reopen = page.get_by_role("button", name="Reopen local review", exact=True) + _choose_combobox_option( + page, + saved_review, + option_name="junit-browser-review", + settled_control=reopen, + ) + reopen.click() + expect( + page.get_by_text( + "Review reopened from local storage after validation.", exact=True + ) + ).to_be_visible() + expect( + page.get_by_text("Recorded imported JUnit results (1)", exact=True) + ).to_be_visible() + + for label, suffix in (("Download Markdown", ".md"), ("Download JSON", ".json")): + download = _download_with_retry(page, button_name=label) + assert download.suggested_filename.endswith(suffix) + downloaded_bytes = download.path().read_bytes() + assert artifact_digest.encode() in downloaded_bytes + assert b"RAW-JUNIT-OUTPUT-SENTINEL" not in downloaded_bytes + assert b"FAILURE-BODY-SENTINEL" not in downloaded_bytes + + def test_installed_wheel_primary_path_in_chromium( tmp_path: Path, request: pytest.FixtureRequest ) -> None: @@ -381,6 +779,36 @@ def test_installed_wheel_primary_path_in_chromium( ) _run(str(environment_python), "-c", "import playwright, scopeproof_core, streamlit") + review_store_dir = home_dir / ".scopeproof" / "reviews" + _run( + str(environment_python), + "-c", + ( + "from pathlib import Path; import sys; " + "from scopeproof_core.demo import build_demo_review; " + "from scopeproof_core.reviews.lifecycle import new_review_state; " + "from scopeproof_core.schemas.models import ReviewBundle; " + "from scopeproof_core.storage.json_store import JsonReviewStore; " + "bundle=build_demo_review().model_copy(deep=True); " + "bundle.review.review_id='junit-browser-review'; " + "bundle.review.head_sha='a'*40; " + "bundle.criteria_revision_number=1; " + "[(setattr(item, 'commit_sha', bundle.review.head_sha), " + "setattr(item, 'permalink', item.permalink.replace('head-demo-002', " + "bundle.review.head_sha))) for item in bundle.evidence]; " + "bundle=ReviewBundle.model_validate(bundle.model_dump(mode='python')); " + "JsonReviewStore(Path(sys.argv[1])).save(new_review_state(bundle))" + ), + str(review_store_dir), + ) + junit_artifact = runtime_dir / "junit-results.xml" + junit_artifact_bytes = ( + b'' + b"RAW-JUNIT-OUTPUT-SENTINEL" + ) + junit_artifact.write_bytes(junit_artifact_bytes) + junit_artifact_digest = sha256(junit_artifact_bytes).hexdigest() + port = _available_port() base_url = f"http://127.0.0.1:{port}" log_path = runtime_dir / "scopeproof-web.log" @@ -438,8 +866,15 @@ def test_installed_wheel_primary_path_in_chromium( _exercise_primary_path( page, base_url, - verify_persistence_and_downloads=viewport_index == 0, + verify_persistence_and_downloads=viewport_index == 1, ) + if viewport_index == 0: + _exercise_junit_import_round_trip( + page, + base_url, + artifact_path=junit_artifact, + artifact_digest=junit_artifact_digest, + ) context.close() finally: browser.close() diff --git a/tests/cli/test_cli.py b/tests/cli/test_cli.py index 08526106..5fd95bac 100644 --- a/tests/cli/test_cli.py +++ b/tests/cli/test_cli.py @@ -35,11 +35,13 @@ Criterion, EvidenceLevel, HumanDecision, + JUnitImportMutationMetadata, LifecycleMutationMetadata, PullRequestSnapshot, RepositoryVisibility, ResolutionEvent, ReviewInputOrigin, + ReviewState, RuntimeEvidence, ) from scopeproof_core.storage.json_store import JsonReviewStore @@ -2840,3 +2842,364 @@ def test_alpha_friction_requires_stage(tmp_path: Path, capsys) -> None: assert error.value.code == 2 assert "friction stage" in capsys.readouterr().err + + +def save_exact_head_cli_review(storage: Path) -> ReviewState: + state = new_review_state(build_demo_review()).model_copy(deep=True) + state.review.head_sha = "a" * 40 + assert state.bundle is not None + state.bundle.review.head_sha = "a" * 40 + state = ReviewState.model_validate(state.model_dump(mode="python")) + JsonReviewStore(storage).save(state) + return state + + +def write_junit_cli_files(tmp_path: Path, criterion_id: str) -> tuple[Path, Path]: + artifact = tmp_path / "results.xml" + artifact.write_bytes( + b'' + ) + mapping = tmp_path / "mapping.json" + mapping.write_text( + json.dumps( + { + "schema_version": "junit-mapping-v1", + "artifact_sha256": sha256(artifact.read_bytes()).hexdigest(), + "selections": [ + {"scope_id": "suite-0001", "criterion_id": criterion_id} + ], + } + ), + encoding="utf-8", + ) + return artifact, mapping + + +def test_inspect_junit_prints_only_sanitized_bounded_json( + tmp_path: Path, capsys +) -> None: + artifact = tmp_path / "results.xml" + artifact.write_bytes( + b'' + b'RAW-FAILURE-SENTINEL' + b'RAW-OUTPUT-SENTINEL' + ) + + assert main(["inspect-junit", str(artifact)]) == 0 + + output = capsys.readouterr().out + payload = json.loads(output) + assert payload["suites"][0]["suite_id"] == "suite-0001" + assert payload["suites"][0]["test_cases"][0]["status"] == "failure" + assert payload["totals"]["total"] == 1 + assert "RAW-FAILURE-SENTINEL" not in output + assert "RAW-OUTPUT-SENTINEL" not in output + + +def test_import_junit_persists_one_non_gating_record(tmp_path: Path, capsys) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + assert state.bundle is not None + original_gate = state.bundle.gate + + assert main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA owner", + "--limitation", + "Synthetic local artifact", + "--storage-dir", + str(storage), + ] + ) == 0 + + metadata = JUnitImportMutationMetadata.model_validate_json( + capsys.readouterr().out + ) + loaded = JsonReviewStore(storage).load(state.review.review_id) + assert loaded.bundle is not None + assert len(loaded.bundle.junit_evidence_imports) == 1 + imported = loaded.bundle.junit_evidence_imports[0] + assert metadata.import_id == imported.import_id + assert metadata.artifact_sha256 == sha256(artifact.read_bytes()).hexdigest() + assert metadata.head_sha == "a" * 40 + assert metadata.evidence_boundary == "externally_supplied_non_gating" + assert loaded.bundle.gate == original_gate + assert loaded.bundle.runtime_evidence == [] + assert loaded.bundle.resolutions == [] + + +def test_import_junit_failure_preserves_saved_record_bytes( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + command = [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA", + "--storage-dir", + str(storage), + ] + assert main(command) == 0 + capsys.readouterr() + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + + with pytest.raises(SystemExit) as duplicate: + main(command) + + assert duplicate.value.code == 2 + assert "already imported" in capsys.readouterr().err.lower() + assert path.read_bytes() == before + + +def test_import_junit_rejects_extra_mapping_fields_without_mutation( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + payload = json.loads(mapping.read_text(encoding="utf-8")) + payload["raw_xml"] = "TOP-SECRET-MAPPING-CONTENT" + mapping.write_text(json.dumps(payload), encoding="utf-8") + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + + with pytest.raises(SystemExit) as error: + main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA", + "--storage-dir", + str(storage), + ] + ) + + assert error.value.code == 2 + stderr = capsys.readouterr().err + assert "mapping document is invalid" in stderr.lower() + assert "TOP-SECRET-MAPPING-CONTENT" not in stderr + assert path.read_bytes() == before + + +def test_import_junit_rejects_unversioned_mapping_without_mutation( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + payload = json.loads(mapping.read_text(encoding="utf-8")) + payload.pop("schema_version") + mapping.write_text(json.dumps(payload), encoding="utf-8") + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + + with pytest.raises(SystemExit) as error: + main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA", + "--storage-dir", + str(storage), + ] + ) + + assert error.value.code == 2 + assert "mapping document is invalid" in capsys.readouterr().err.lower() + assert path.read_bytes() == before + + +def test_import_junit_bounds_invalid_importer_error_without_mutation( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + secret = "SECRET-CREDENTIAL-" + "x" * 280 + + with pytest.raises(SystemExit) as error: + main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + secret, + "--storage-dir", + str(storage), + ] + ) + + assert error.value.code == 2 + stderr = capsys.readouterr().err + assert "metadata is invalid" in stderr.lower() + assert "SECRET-CREDENTIAL" not in stderr + assert secret not in stderr + assert path.read_bytes() == before + + +def test_import_junit_rejects_oversized_mapping_without_mutation( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + mapping.write_bytes(b"x" * 1_048_577) + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + + with pytest.raises(SystemExit) as error: + main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA", + "--storage-dir", + str(storage), + ] + ) + + assert error.value.code == 2 + assert "mapping exceeds the byte limit" in capsys.readouterr().err.lower() + assert path.read_bytes() == before + + +def test_import_junit_rejects_mapping_for_changed_artifact_without_mutation( + tmp_path: Path, capsys +) -> None: + storage = tmp_path / "reviews" + state = save_exact_head_cli_review(storage) + artifact, mapping = write_junit_cli_files( + tmp_path, state.criteria_revision.criteria[0].criterion_id + ) + artifact.write_bytes( + b'' + ) + path = storage / f"{state.review.review_id}.json" + before = path.read_bytes() + + with pytest.raises(SystemExit) as error: + main( + [ + "import-junit", + state.review.review_id, + str(artifact), + "--mapping", + str(mapping), + "--importer", + "QA", + "--storage-dir", + str(storage), + ] + ) + + assert error.value.code == 2 + assert "digest" in capsys.readouterr().err.lower() + assert path.read_bytes() == before + + +def test_inspect_junit_rejects_oversized_regular_file_before_parsing( + tmp_path: Path, capsys +) -> None: + artifact = tmp_path / "oversized.xml" + artifact.write_bytes(b"x" * 1_048_577) + + with pytest.raises(SystemExit) as error: + main(["inspect-junit", str(artifact)]) + + assert error.value.code == 2 + assert "byte limit" in capsys.readouterr().err.lower() + + +def test_junit_file_reader_requests_binary_mode( + tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + artifact = tmp_path / "results.xml" + artifact.write_bytes( + b'' + ) + binary_flag = 1 << 29 + observed_flags: list[int] = [] + original_open = cli_module.os.open + + def recording_open(path, flags, *args, **kwargs): + observed_flags.append(flags) + return original_open(path, flags & ~binary_flag, *args, **kwargs) + + monkeypatch.setattr(cli_module, "_BINARY", binary_flag, raising=False) + monkeypatch.setattr(cli_module.os, "open", recording_open) + + assert cli_module._read_bounded_junit_artifact(artifact) == artifact.read_bytes() + assert observed_flags + assert all(flags & binary_flag for flags in observed_flags) + + +def test_inspect_junit_rejects_non_regular_artifact(tmp_path: Path, capsys) -> None: + artifact = tmp_path / "artifact-directory" + artifact.mkdir() + + with pytest.raises(SystemExit) as error: + main(["inspect-junit", str(artifact)]) + + assert error.value.code == 2 + assert "regular file" in capsys.readouterr().err.lower() + + +def test_inspect_junit_rejects_unsafe_xml_without_echoing_artifact( + tmp_path: Path, capsys +) -> None: + artifact = tmp_path / "unsafe.xml" + artifact.write_bytes( + b']>' + b'' + ) + + with pytest.raises(SystemExit) as error: + main(["inspect-junit", str(artifact)]) + + assert error.value.code == 2 + stderr = capsys.readouterr().err + assert "forbidden XML construct" in stderr + assert "PRIVATE-SENTINEL" not in stderr diff --git a/tests/importers/test_junit.py b/tests/importers/test_junit.py new file mode 100644 index 00000000..88c75331 --- /dev/null +++ b/tests/importers/test_junit.py @@ -0,0 +1,514 @@ +from datetime import UTC, datetime +from hashlib import sha256 + +import pytest +from pydantic import ValidationError + +import scopeproof_core.importers.junit as junit_module +from scopeproof_core.demo import build_demo_review +from scopeproof_core.importers.junit import ( + JUnitImportError, + JUnitMappingDocument, + JUnitMappingSelection, + build_junit_evidence_import, + parse_junit_artifact, +) +from scopeproof_core.reviews.lifecycle import new_review_state +from scopeproof_core.schemas.models import JUnitCaseStatus, ReviewBundle, ReviewState + +HEAD_SHA = "a" * 40 +SIMPLE_XML = b'' +TWO_CASE_XML = ( + b'' + b'' + b'' + b'secret body' + b'' +) + + +def exact_head_state() -> ReviewState: + bundle = build_demo_review().model_copy(deep=True) + bundle.review.head_sha = HEAD_SHA + bundle = ReviewBundle.model_validate(bundle.model_dump(mode="python")) + return new_review_state(bundle) + + +def first_criterion_id(state: ReviewState) -> str: + assert state.bundle is not None + return state.bundle.criteria[0].criterion_id + + +def test_parser_returns_sanitized_cases_and_computed_totals() -> None: + parsed = parse_junit_artifact(TWO_CASE_XML) + + assert parsed.artifact_sha256 == sha256(TWO_CASE_XML).hexdigest() + assert parsed.totals.model_dump() == { + "total": 2, + "passed": 1, + "failures": 1, + "errors": 0, + "skipped": 0, + } + assert [suite.suite_id for suite in parsed.suites] == ["suite-0001"] + assert [item.test_case_id for item in parsed.suites[0].test_cases] == [ + "suite-0001-case-0001", + "suite-0001-case-0002", + ] + assert parsed.suites[0].test_cases[1].status is JUnitCaseStatus.FAILURE + serialized = parsed.model_dump_json() + assert "secret body" not in serialized + assert "boom" not in serialized + + +def test_parser_discards_output_and_properties_with_one_bounded_warning() -> None: + secret = "SENTINEL-OUTPUT-DO-NOT-PERSIST" + xml = ( + '' + f'{secret}' + 'also hidden' + ).encode() + + parsed = parse_junit_artifact(xml) + + assert parsed.parser_warnings == [ + "JUnit properties and output content were discarded during import." + ] + assert secret not in parsed.model_dump_json() + assert "hidden" not in parsed.model_dump_json() + + +@pytest.mark.parametrize( + "xml", + [ + ( + b'' + b'' + b'' + ), + ( + b'' + b'' + b'' + ), + ( + b'' + b'' + b'' + ), + ], +) +def test_parser_rejects_structural_results_hidden_inside_discarded_wrappers( + xml: bytes, +) -> None: + with pytest.raises(JUnitImportError, match="unsupported structure"): + parse_junit_artifact(xml) + + +def test_parser_redacts_path_and_url_like_names_before_persistence() -> None: + xml = ( + b'' + b'' + ) + + parsed = parse_junit_artifact(xml) + + suite = parsed.suites[0] + case = suite.test_cases[0] + assert suite.suite_name == "Redacted suite name 0001" + assert case.suite_name == "Redacted suite name 0001" + assert case.class_name == "Redacted class name" + assert case.test_name == "Redacted test name 0001" + assert parsed.parser_warnings == [ + "Path- or URL-like JUnit names were redacted during import." + ] + serialized = parsed.model_dump_json() + assert "ci.example.test" not in serialized + assert "workspace" not in serialized + assert "agent" not in serialized + + +@pytest.mark.parametrize( + "unsafe_name", + [ + "mailto:secret@example.test", + "data:,TOP-SECRET", + "urn:example:private", + "C:relative-secret.xml", + "case for mailto:secret@example.test", + "artifact C:relative-secret.xml", + ], +) +def test_parser_redacts_scheme_like_names(unsafe_name: str) -> None: + parsed = parse_junit_artifact( + f''.encode() + ) + + assert parsed.suites[0].test_cases[0].test_name == "Redacted test name 0001" + assert unsafe_name not in parsed.model_dump_json() + + +def test_parser_reports_declared_count_mismatches_without_trusting_them() -> None: + parsed = parse_junit_artifact( + b'' + b'' + ) + + assert parsed.totals.total == 1 + assert parsed.totals.passed == 1 + assert parsed.parser_warnings == [ + "Declared JUnit counts differed from the sanitized observed results." + ] + + +@pytest.mark.parametrize( + ("xml", "message"), + [ + (b'', "UTF-8"), + (b'', "forbidden XML construct"), + ( + b']>' + b'', + "forbidden XML construct", + ), + (b'', "processing"), + ( + b'' + b'', + "XInclude", + ), + (b'', "root"), + (b'', "direct testsuite"), + ( + b'' + b'', + "(?i)nested", + ), + (b'', "test name"), + ( + b'' + b'', + "multiple result", + ), + (b'', "malformed"), + ], +) +def test_parser_rejects_unsafe_or_ambiguous_xml_without_leaking_input( + xml: bytes, message: str +) -> None: + with pytest.raises(JUnitImportError, match=message) as exc_info: + parse_junit_artifact(xml) + + assert "file:///etc/passwd" not in str(exc_info.value) + assert "example.test" not in str(exc_info.value) + + +def test_parser_rejects_non_bytes_invalid_utf8_and_blank_input() -> None: + with pytest.raises(TypeError, match="bytes"): + parse_junit_artifact("") # type: ignore[arg-type] + with pytest.raises(JUnitImportError, match="UTF-8"): + parse_junit_artifact(b"\xff\xfe") + with pytest.raises(JUnitImportError, match="empty"): + parse_junit_artifact(b"") + + +def test_parser_enforces_byte_suite_case_and_element_budgets( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr(junit_module, "MAX_JUNIT_BYTES", len(SIMPLE_XML) - 1) + with pytest.raises(JUnitImportError, match="byte limit"): + parse_junit_artifact(SIMPLE_XML) + + monkeypatch.setattr(junit_module, "MAX_JUNIT_BYTES", 1_048_576) + monkeypatch.setattr(junit_module, "MAX_JUNIT_SUITES", 1) + with pytest.raises(JUnitImportError, match="suite limit"): + parse_junit_artifact( + b'' + ) + + monkeypatch.setattr(junit_module, "MAX_JUNIT_CASES", 1) + with pytest.raises(JUnitImportError, match="test-case limit"): + parse_junit_artifact( + b'' + b'' + ) + + monkeypatch.setattr(junit_module, "MAX_JUNIT_CASES", 5_000) + monkeypatch.setattr(junit_module, "MAX_JUNIT_ELEMENTS", 2) + with pytest.raises(JUnitImportError, match="element limit"): + parse_junit_artifact( + b'' + b'' + ) + + +def test_mapping_selection_is_strict_and_non_blank() -> None: + selection = JUnitMappingSelection( + scope_id="suite-0001", criterion_id="AC-01" + ) + assert selection.scope_id == "suite-0001" + with pytest.raises(ValidationError): + JUnitMappingSelection.model_validate( + {"scope_id": "suite-0001", "criterion_id": "AC-01", "extra": True} + ) + with pytest.raises(ValidationError, match="non-whitespace"): + JUnitMappingSelection(scope_id="suite-0001", criterion_id=" ") + + +def test_mapping_document_requires_exact_artifact_digest() -> None: + payload = { + "schema_version": "junit-mapping-v1", + "selections": [{"scope_id": "suite-0001", "criterion_id": "AC-01"}], + } + + with pytest.raises(ValidationError, match="artifact_sha256"): + JUnitMappingDocument.model_validate(payload) + + document = JUnitMappingDocument.model_validate( + {**payload, "artifact_sha256": "a" * 64} + ) + assert document.artifact_sha256 == "a" * 64 + + +def test_mapping_document_requires_explicit_schema_version() -> None: + with pytest.raises(ValidationError, match="schema_version"): + JUnitMappingDocument.model_validate( + { + "artifact_sha256": "a" * 64, + "selections": [ + {"scope_id": "suite-0001", "criterion_id": "AC-01"} + ], + } + ) + + +def test_builder_expands_explicit_suite_mapping_and_binds_review() -> None: + state = exact_head_state() + criterion_id = first_criterion_id(state) + + record = build_junit_evidence_import( + state, + TWO_CASE_XML, + [JUnitMappingSelection(scope_id="suite-0001", criterion_id=criterion_id)], + importer=" QA owner ", + limitations=[" Browser lane not supplied. "], + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + import_id="import-001", + ) + + assert record.artifact_sha256 == sha256(TWO_CASE_XML).hexdigest() + assert record.criterion_mappings[0].test_case_ids == [ + "suite-0001-case-0001", + "suite-0001-case-0002", + ] + assert record.repository == state.review.repository + assert record.pr_number == state.review.pr_number + assert record.head_sha == state.review.head_sha + assert record.criteria_revision_number == state.criteria_revision.number + assert record.imported_by == "QA owner" + assert record.limitations[-1] == "Browser lane not supplied." + assert all( + "did not execute" in item or "not" in item.lower() + for item in record.limitations[:3] + ) + + +def test_builder_expands_case_mapping_and_canonicalizes_duplicate_pairs() -> None: + state = exact_head_state() + criterion_id = first_criterion_id(state) + selection = JUnitMappingSelection( + scope_id="suite-0001-case-0002", criterion_id=criterion_id + ) + + record = build_junit_evidence_import( + state, + TWO_CASE_XML, + [selection, selection], + importer="QA", + ) + + assert record.criterion_mappings[0].test_case_ids == [ + "suite-0001-case-0002" + ] + + +def test_builder_rejects_one_case_mapped_to_multiple_criteria() -> None: + state = exact_head_state() + assert state.bundle is not None + first, second = state.bundle.criteria[:2] + + with pytest.raises(JUnitImportError, match="multiple criteria"): + build_junit_evidence_import( + state, + SIMPLE_XML, + [ + JUnitMappingSelection( + scope_id="suite-0001", criterion_id=first.criterion_id + ), + JUnitMappingSelection( + scope_id="suite-0001-case-0001", + criterion_id=second.criterion_id, + ), + ], + importer="QA", + ) + + +def test_builder_bounds_envelope_validation_errors_without_echoing_input() -> None: + state = exact_head_state() + secret = "SECRET-CREDENTIAL-" + "x" * 280 + + with pytest.raises(JUnitImportError, match="metadata is invalid") as exc_info: + build_junit_evidence_import( + state, + SIMPLE_XML, + [ + JUnitMappingSelection( + scope_id="suite-0001", + criterion_id=first_criterion_id(state), + ) + ], + importer=secret, + ) + + assert secret not in str(exc_info.value) + assert "SECRET-CREDENTIAL" not in str(exc_info.value) + + +@pytest.mark.parametrize( + ("selections", "importer", "message"), + [ + ([], "QA", "explicit mapping"), + ( + [JUnitMappingSelection(scope_id="suite-9999", criterion_id="AC-01")], + "QA", + "unknown mapping scope", + ), + ( + [ + JUnitMappingSelection( + scope_id="suite-0001", criterion_id="AC-UNKNOWN" + ) + ], + "QA", + "unknown criterion", + ), + ([JUnitMappingSelection(scope_id="suite-0001", criterion_id="AC-01")], " ", "importer"), + ], +) +def test_builder_rejects_missing_or_invalid_human_mapping( + selections: list[JUnitMappingSelection], importer: str, message: str +) -> None: + state = exact_head_state() + if selections and selections[0].criterion_id == "AC-01": + selections = [ + selection.model_copy( + update={"criterion_id": first_criterion_id(state)} + ) + for selection in selections + ] + + with pytest.raises((JUnitImportError, ValueError), match=message): + build_junit_evidence_import( + state, + SIMPLE_XML, + selections, + importer=importer, + ) + + +def test_builder_requires_active_confirmed_exact_head_review() -> None: + state = exact_head_state() + criterion_id = first_criterion_id(state) + mapping = [ + JUnitMappingSelection(scope_id="suite-0001", criterion_id=criterion_id) + ] + + no_bundle = state.model_copy(update={"bundle": None}) + with pytest.raises(ValueError, match="active analysis"): + build_junit_evidence_import(no_bundle, SIMPLE_XML, mapping, importer="QA") + + non_exact = state.model_copy(deep=True) + non_exact.review.head_sha = "constructed-head" + assert non_exact.bundle is not None + non_exact.bundle.review.head_sha = "constructed-head" + with pytest.raises(ValueError, match="exact 40-character"): + build_junit_evidence_import(non_exact, SIMPLE_XML, mapping, importer="QA") + + +def test_builder_rejects_blank_limitations_without_exposing_artifact() -> None: + state = exact_head_state() + mapping = [ + JUnitMappingSelection( + scope_id="suite-0001", criterion_id=first_criterion_id(state) + ) + ] + with pytest.raises(ValueError, match="limitations"): + build_junit_evidence_import( + state, + SIMPLE_XML, + mapping, + importer="QA", + limitations=[""], + ) + + +def test_builder_rejects_overlong_limitations_without_exposing_artifact() -> None: + state = exact_head_state() + mapping = [ + JUnitMappingSelection( + scope_id="suite-0001", criterion_id=first_criterion_id(state) + ) + ] + + with pytest.raises(JUnitImportError, match="metadata is invalid"): + build_junit_evidence_import( + state, + SIMPLE_XML, + mapping, + importer="QA", + limitations=["x" * 1_001], + ) + + assert state.bundle is not None + assert state.bundle.junit_evidence_imports == [] + + +def test_builder_rejects_review_at_aggregate_import_cap() -> None: + state = exact_head_state() + mapping = [ + JUnitMappingSelection( + scope_id="suite-0001", criterion_id=first_criterion_id(state) + ) + ] + record = build_junit_evidence_import( + state, + SIMPLE_XML, + mapping, + importer="QA", + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + import_id="import-template", + ) + assert state.bundle is not None + state.bundle.junit_evidence_imports = [ + record.model_copy( + update={ + "import_id": f"import-{index + 1:03d}", + "artifact_sha256": f"{index + 1:064x}", + } + ) + for index in range(20) + ] + capped = ReviewState.model_validate(state.model_dump(mode="python")) + + with pytest.raises(JUnitImportError, match="maximum number"): + build_junit_evidence_import( + capped, + SIMPLE_XML, + mapping, + importer="QA", + ) + + assert capped.bundle is not None + assert len(capped.bundle.junit_evidence_imports) == 20 diff --git a/tests/reporting/test_comparison_exports.py b/tests/reporting/test_comparison_exports.py index 0f71201b..e55d0c87 100644 --- a/tests/reporting/test_comparison_exports.py +++ b/tests/reporting/test_comparison_exports.py @@ -11,6 +11,10 @@ EvidenceChange, EvidenceChangeKind, EvidenceReference, + JUnitImportChange, + JUnitImportChangeKind, + JUnitImportReference, + JUnitMappingReference, ResolutionChange, ReviewComparison, ) @@ -117,7 +121,8 @@ def test_comparison_markdown_shows_two_sides_and_evidence_boundary() -> None: assert "review the current evidence" in report.lower() assert "Prior Decisions Requiring Review" in report assert "AC\\-01" in report - assert "never carries acceptance to a changed head" in report + assert "does not carry a prior decision forward automatically" in report + assert "changed head" not in report def test_comparison_exports_do_not_carry_a_previous_decision_into_current() -> None: @@ -144,6 +149,67 @@ def test_comparison_exports_do_not_carry_a_previous_decision_into_current() -> N assert "does not carry forward a prior human decision" in report +def test_comparison_exports_show_inert_non_gating_junit_mapping_changes() -> None: + comparison = example_comparison() + artifact_digest = "a" * 64 + previous = JUnitImportReference( + import_id="import-old", + artifact_sha256=artifact_digest, + head_sha="b" * 40, + asserted_importer="", + mappings=[ + JUnitMappingReference( + criterion_id="AC-01", + test_case_ids=["suite-0001-case-0001"], + ) + ], + ) + current = JUnitImportReference( + import_id="import-new", + artifact_sha256=artifact_digest, + head_sha="c" * 40, + asserted_importer="=owner-new ", + mappings=[ + JUnitMappingReference( + criterion_id="AC-01", + test_case_ids=[ + "suite-0001-case-0001", + "suite-0001-case-0002", + ], + ) + ], + ) + comparison.junit_import_changes = [ + JUnitImportChange( + artifact_sha256=artifact_digest, + kind=JUnitImportChangeKind.MAPPING_MODIFIED, + previous=previous, + current=current, + ) + ] + + payload = json.loads(export_comparison_json(comparison)) + report = export_comparison_markdown(comparison) + + assert payload["junit_import_changes"][0]["kind"] == "mapping_modified" + assert payload["junit_import_changes"][0]["previous"]["evidence_boundary"][ + "criterion_mapping" + ] == "organizational_context_not_proof" + assert payload["junit_import_changes"][0]["current"]["evidence_boundary"][ + "gate_effect" + ] == "non_gating" + assert artifact_digest in report + assert "Imported External Test Result Changes" in report + assert "externally supplied, non-gating context" in report + assert "artifact digest covers imported bytes only" in report + assert "importer identity is asserted, not authenticated" in report + assert "criterion mapping is organizational context, not proof" in report + assert "" not in report + assert "" not in report + assert "<owner-old>" in report + assert "<unsafe>" in report + + def test_comparison_markdown_escapes_repository_controlled_text() -> None: comparison = example_comparison() unsafe = reference( diff --git a/tests/reporting/test_exporters.py b/tests/reporting/test_exporters.py index f94e99ae..d6c9d79d 100644 --- a/tests/reporting/test_exporters.py +++ b/tests/reporting/test_exporters.py @@ -5,7 +5,10 @@ import pytest -from scopeproof_core.criteria.confirmation import build_criteria_source_provenance +from scopeproof_core.criteria.confirmation import ( + build_criteria_source_provenance, + normalized_criteria_sha256, +) from scopeproof_core.gates import validation as gate_validation from scopeproof_core.gates.evaluator import evaluate_gate from scopeproof_core.reporting.exporters import ( @@ -31,6 +34,8 @@ HumanDecision, HumanResolution, IngestionState, + JUnitEvidenceBoundary, + JUnitEvidenceImport, RepositoryVisibility, ResearchContext, ResolutionEvent, @@ -42,6 +47,72 @@ ) +def add_junit_import(bundle: ReviewBundle) -> ReviewBundle: + bundle = bundle.model_copy(deep=True) + bundle.review.head_sha = "a" * 40 + bundle.evidence[0].commit_sha = bundle.review.head_sha + bundle.evidence[0].permalink = ( + "https://github.com/acme/widget/blob/" + f"{bundle.review.head_sha}/src/export.py#L42-L42" + ) + bundle.criteria_revision_number = 1 + provenance = bundle.review.criteria_source_provenance + assert provenance is not None + bundle.junit_evidence_imports = [ + JUnitEvidenceImport( + schema_version="junit-import-v1", + evidence_boundary=JUnitEvidenceBoundary(), + import_id="junit-import-001", + repository=bundle.review.repository, + pr_number=bundle.review.pr_number, + head_sha=bundle.review.head_sha, + criteria_revision_number=1, + confirmed_criteria_sha256=normalized_criteria_sha256(bundle.criteria), + criteria_source_provenance=provenance, + artifact_sha256="b" * 64, + imported_by="=ASSERTED ", + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + totals={ + "total": 2, + "passed": 1, + "failures": 1, + "errors": 0, + "skipped": 0, + }, + test_cases=[ + { + "test_case_id": "suite-0001-case-0001", + "suite_id": "suite-0001", + "suite_name": "", + "class_name": None, + "test_name": "=test_pass", + "status": "passed", + }, + { + "test_case_id": "suite-0001-case-0002", + "suite_id": "suite-0001", + "suite_name": "", + "class_name": "tests.", + "test_name": "test_fail **claim**", + "status": "failure", + }, + ], + criterion_mappings=[ + { + "criterion_id": "AC-01", + "test_case_ids": [ + "suite-0001-case-0001", + "suite-0001-case-0002", + ], + } + ], + parser_warnings=["@warning "], + limitations=["+external result only "], + ) + ] + return ReviewBundle.model_validate(bundle.model_dump(mode="python")) + + def example_bundle() -> ReviewBundle: review = Review( review_id="review-1", @@ -154,6 +225,94 @@ def rebind_criteria_source_provenance(bundle: ReviewBundle) -> None: ) +def test_junit_import_exports_are_complete_inert_and_non_gating() -> None: + bundle = add_junit_import(example_bundle()) + second_import = bundle.junit_evidence_imports[0].model_copy( + update={ + "import_id": "junit-import-002", + "artifact_sha256": "c" * 64, + "imported_by": "second owner", + "parser_warnings": ["second warning"], + "limitations": ["second limitation"], + } + ) + bundle.junit_evidence_imports.append(second_import) + bundle = ReviewBundle.model_validate(bundle.model_dump(mode="python")) + + json_report = export_json(bundle) + markdown = export_markdown(bundle) + csv_row = next(csv.DictReader(io.StringIO(export_csv(bundle)))) + html_report = export_html(bundle) + rendered = "\n".join((json_report, markdown, str(csv_row), html_report)) + + assert "b" * 64 in rendered + assert bundle.review.head_sha in rendered + assert "suite-0001-case-0001" in rendered + assert "passed" in rendered + assert "=ASSERTED " in json_report + json_payload = json.loads(json_report) + assert json_payload["junit_evidence_imports"][0]["evidence_boundary"] == { + "source": "externally_supplied", + "gate_effect": "non_gating", + "execution": "not_executed_by_scopeproof", + "artifact_digest_scope": "imported_bytes_only", + "importer_identity": "asserted_not_authenticated", + "criterion_mapping": "organizational_context_not_proof", + } + assert "## Imported External Test Results" in markdown + assert "Imported external test results" in html_report + assert "RAW-JUNIT-OUTPUT-SENTINEL" not in rendered + assert "FAILURE-BODY-SENTINEL" not in rendered + assert "/private/local/results.xml" not in rendered + assert "" not in markdown + assert "" not in html_report + assert "<suite>" in markdown + assert "<suite>" in html_report + assert csv_row["junit_artifact_digests"] == json.dumps(["b" * 64, "c" * 64]) + csv_cases = json.loads(csv_row["junit_mapped_cases"]) + assert { + (item["import_id"], item["artifact_sha256"], item["imported_by"]) + for item in csv_cases + } == { + ("junit-import-001", "b" * 64, "=ASSERTED "), + ("junit-import-002", "c" * 64, "second owner"), + } + assert "externally supplied, non-gating context" in csv_row[ + "junit_evidence_boundary" + ].lower() + assert "did not execute" in csv_row["junit_evidence_boundary"].lower() + assert "imported bytes only" in csv_row["junit_evidence_boundary"].lower() + assert "asserted, not authenticated" in csv_row["junit_evidence_boundary"].lower() + assert "organizational context, not proof" in csv_row["junit_evidence_boundary"].lower() + assert csv_row["junit_importers"].startswith("[") + assert json.loads(csv_row["junit_importers"])[0] == "=ASSERTED " + assert json.loads(csv_row["junit_parser_warnings"]) == [ + { + "artifact_sha256": "b" * 64, + "import_id": "junit-import-001", + "warning": "@warning ", + }, + { + "artifact_sha256": "c" * 64, + "import_id": "junit-import-002", + "warning": "second warning", + }, + ] + assert json.loads(csv_row["junit_limitations"]) == [ + { + "artifact_sha256": "b" * 64, + "import_id": "junit-import-001", + "limitation": "+external result only ", + }, + { + "artifact_sha256": "c" * 64, + "import_id": "junit-import-002", + "limitation": "second limitation", + }, + ] + assert bundle.gate.verdict.value in rendered + + def example_state(): bundle = example_bundle() resolution = bundle.resolutions[0] diff --git a/tests/reviews/test_comparison.py b/tests/reviews/test_comparison.py index 8e75ca52..8cb2fcbd 100644 --- a/tests/reviews/test_comparison.py +++ b/tests/reviews/test_comparison.py @@ -3,7 +3,10 @@ import pytest from pydantic import ValidationError -from scopeproof_core.criteria.confirmation import build_criteria_source_provenance +from scopeproof_core.criteria.confirmation import ( + build_criteria_source_provenance, + normalized_criteria_sha256, +) from scopeproof_core.gates.evaluator import evaluate_gate from scopeproof_core.reviews.comparison import ( EvidenceChange, @@ -25,6 +28,8 @@ HumanDecision, HumanResolution, IngestionState, + JUnitEvidenceBoundary, + JUnitEvidenceImport, RepositoryVisibility, Review, ReviewBundle, @@ -111,6 +116,67 @@ def bundle_with(*items: EvidenceItem, head_sha: str) -> ReviewBundle: ) +def with_junit_import( + bundle: ReviewBundle, + *, + artifact_digest: str, + mapped_case_ids: list[str], +) -> ReviewBundle: + bundle = bundle.model_copy(deep=True) + bundle.criteria_revision_number = 1 + provenance = bundle.review.criteria_source_provenance + assert provenance is not None + bundle.junit_evidence_imports = [ + JUnitEvidenceImport( + schema_version="junit-import-v1", + evidence_boundary=JUnitEvidenceBoundary(), + import_id=f"import-{artifact_digest[0]}", + repository=bundle.review.repository, + pr_number=bundle.review.pr_number, + head_sha=bundle.review.head_sha, + criteria_revision_number=1, + confirmed_criteria_sha256=normalized_criteria_sha256(bundle.criteria), + criteria_source_provenance=provenance, + artifact_sha256=artifact_digest, + imported_by="Fixture owner", + imported_at=bundle.review.created_at, + totals={ + "total": 2, + "passed": 1, + "failures": 1, + "errors": 0, + "skipped": 0, + }, + test_cases=[ + { + "test_case_id": "suite-0001-case-0001", + "suite_id": "suite-0001", + "suite_name": "unit", + "class_name": None, + "test_name": "test_one", + "status": "passed", + }, + { + "test_case_id": "suite-0001-case-0002", + "suite_id": "suite-0001", + "suite_name": "unit", + "class_name": None, + "test_name": "test_two", + "status": "failure", + }, + ], + criterion_mappings=[ + { + "criterion_id": "AC-01", + "test_case_ids": mapped_case_ids, + } + ], + limitations=["External non-gating context."], + ) + ] + return ReviewBundle.model_validate(bundle.model_dump(mode="python")) + + def test_comparison_preserves_legacy_unlinked_manual_verification_as_needs_review() -> None: previous = bundle_with(head_sha="old") current = bundle_with(head_sha="new") @@ -613,3 +679,90 @@ def test_comparison_relationship_rejects_non_exact_legacy_unknown_head() -> None with pytest.raises(ValueError, match="exact head SHAs"): compare_reviews(previous, current) + + +def test_comparison_projects_unchanged_added_removed_and_mapping_modified_junit_imports() -> None: + head = "a" * 40 + base = bundle_with(head_sha=head) + unchanged_previous = with_junit_import( + base, artifact_digest="b" * 64, mapped_case_ids=["suite-0001-case-0001"] + ) + unchanged_current = with_junit_import( + base, artifact_digest="b" * 64, mapped_case_ids=["suite-0001-case-0001"] + ) + + unchanged = compare_reviews(unchanged_previous, unchanged_current) + assert [item.kind.value for item in unchanged.junit_import_changes] == [ + "unchanged" + ] + assert unchanged.junit_import_changes[0].artifact_sha256 == "b" * 64 + + modified_current = with_junit_import( + base, + artifact_digest="b" * 64, + mapped_case_ids=["suite-0001-case-0001", "suite-0001-case-0002"], + ) + modified = compare_reviews(unchanged_previous, modified_current) + assert [item.kind.value for item in modified.junit_import_changes] == [ + "mapping_modified" + ] + assert modified.junit_import_changes[0].previous is not None + assert modified.junit_import_changes[0].current is not None + + different_current = with_junit_import( + base, artifact_digest="c" * 64, mapped_case_ids=["suite-0001-case-0001"] + ) + different = compare_reviews(unchanged_previous, different_current) + assert [item.kind.value for item in different.junit_import_changes] == [ + "removed", + "added", + ] + + +def test_changed_junit_mapping_requires_review_only_for_previously_resolved_criteria() -> None: + head = "a" * 40 + base = bundle_with( + evidence("EV-AC-01", sha=head, criterion_id="AC-01"), + evidence("EV-AC-02", sha=head, criterion_id="AC-02"), + head_sha=head, + ) + base.resolutions = [ + HumanResolution( + criterion_id=criterion_id, + decision=HumanDecision.ACCEPTED, + comment="Owner reviewed existing evidence.", + ) + for criterion_id in ("AC-01", "AC-02") + ] + base.gate = evaluate_gate( + base.review, base.criteria, base.findings, base.resolutions + ) + previous = with_junit_import( + base, artifact_digest="d" * 64, mapped_case_ids=["suite-0001-case-0001"] + ) + current = with_junit_import( + base, + artifact_digest="d" * 64, + mapped_case_ids=["suite-0001-case-0001", "suite-0001-case-0002"], + ) + + comparison = compare_reviews(previous, current) + + assert comparison.criteria_requiring_decision_review == ["AC-01"] + assert previous.resolutions == current.resolutions + assert previous.gate == current.gate + + +def test_comparison_revalidates_junit_import_relationships_before_projection() -> None: + head = "a" * 40 + valid = with_junit_import( + bundle_with(head_sha=head), + artifact_digest="e" * 64, + mapped_case_ids=["suite-0001-case-0001"], + ) + tampered = valid.model_copy(deep=True) + imported = tampered.junit_evidence_imports[0] + object.__setattr__(imported, "head_sha", "f" * 40) + + with pytest.raises(ValidationError, match="JUnit import identity"): + compare_reviews(valid, tampered) diff --git a/tests/reviews/test_lifecycle.py b/tests/reviews/test_lifecycle.py index 7537ea79..e3a7425e 100644 --- a/tests/reviews/test_lifecycle.py +++ b/tests/reviews/test_lifecycle.py @@ -5,11 +5,16 @@ from scopeproof_core.criteria.confirmation import build_criteria_source_provenance from scopeproof_core.gates.evaluator import evaluate_gate +from scopeproof_core.importers.junit import ( + JUnitMappingSelection, + build_junit_evidence_import, +) from scopeproof_core.reviews import attach_analysis from scopeproof_core.reviews.lifecycle import ( ResolutionEventStatus, acceptance_requires_comment, append_external_verification, + append_junit_evidence_import, append_resolution, append_runtime_evidence, can_record_final_acceptance, @@ -33,6 +38,7 @@ ResolutionEvent, Review, ReviewBundle, + ReviewState, RuntimeEvidence, normalized_criteria_sha256, source_text_sha256, @@ -349,6 +355,28 @@ def test_attach_analysis_preserves_reanalysis_lineage() -> None: assert attached.bundle.criteria_revision_number == 2 +def test_attach_analysis_rejects_preloaded_junit_imports() -> None: + state = exact_head_state() + revised = revise_criteria( + state, + [Criterion(criterion_id="AC-01", text="Export filtered CSV")], + "Export filtered CSV", + ) + confirmed = confirm_pending_revision(revised) + incoming = analysis_bundle_for(confirmed) + analyzed = attach_analysis(confirmed, incoming) + record = junit_import_for(analyzed) + incoming.criteria_revision_number = confirmed.criteria_revision.number + incoming.junit_evidence_imports = [record] + + with pytest.raises( + ValueError, match="attached analysis must not contain JUnit imports" + ): + attach_analysis(confirmed, incoming) + + assert confirmed.bundle is None + + def test_skipped_analysis_history_records_exact_criteria_revisions() -> None: revision_one = initial_state() revision_two = confirm_pending_revision( @@ -1541,3 +1569,145 @@ def test_appended_runtime_evidence_does_not_alias_the_supplied_object() -> None: assert updated.bundle is not None assert updated.bundle.runtime_evidence[0].result == "passed" assert updated.bundle.runtime_evidence[0].limitations == ["Browser only"] + + +def exact_head_state() -> ReviewState: + state = initial_state().model_copy(deep=True) + state.review.head_sha = "a" * 40 + assert state.bundle is not None + state.bundle.review.head_sha = "a" * 40 + return ReviewState.model_validate(state.model_dump(mode="python")) + + +def junit_import_for(state: ReviewState, *, import_id: str = "import-001"): + return build_junit_evidence_import( + state, + b'', + [ + JUnitMappingSelection( + scope_id="suite-0001", + criterion_id="AC-01", + ) + ], + importer="QA owner", + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + import_id=import_id, + ) + + +def test_junit_import_append_is_non_gating_and_does_not_alias_input() -> None: + state = exact_head_state() + record = junit_import_for(state) + assert state.bundle is not None + original_gate = state.bundle.gate.model_copy(deep=True) + original_findings = [item.model_copy(deep=True) for item in state.bundle.findings] + original_resolutions = list(state.bundle.resolutions) + original_runtime = list(state.bundle.runtime_evidence) + original_events = list(state.resolution_events) + original_final_acceptance = state.review.final_acceptance + + updated = append_junit_evidence_import(state, record) + record.limitations.append("Caller mutation") + + assert updated.bundle is not None + assert len(updated.bundle.junit_evidence_imports) == 1 + assert "Caller mutation" not in updated.bundle.junit_evidence_imports[0].limitations + assert state.bundle.junit_evidence_imports == [] + assert updated.bundle.gate == original_gate + assert updated.bundle.findings == original_findings + assert updated.bundle.resolutions == original_resolutions + assert updated.bundle.runtime_evidence == original_runtime + assert updated.resolution_events == original_events + assert updated.review.final_acceptance is original_final_acceptance + + +def test_new_review_state_rejects_preexisting_junit_imports() -> None: + state = exact_head_state() + imported = append_junit_evidence_import(state, junit_import_for(state)) + assert imported.bundle is not None + + with pytest.raises( + ValueError, match="initial analysis bundle must not contain JUnit imports" + ): + new_review_state(imported.bundle) + + +def test_junit_import_append_rejects_aggregate_cap_atomically() -> None: + state = exact_head_state() + record = junit_import_for(state) + assert state.bundle is not None + state.bundle.junit_evidence_imports = [ + record.model_copy( + update={ + "import_id": f"import-{index + 1:03d}", + "artifact_sha256": f"{index + 1:064x}", + } + ) + for index in range(20) + ] + capped = ReviewState.model_validate(state.model_dump(mode="python")) + overflow = record.model_copy( + update={"import_id": "import-overflow", "artifact_sha256": "f" * 64} + ) + + with pytest.raises(ValueError, match="at most 20"): + append_junit_evidence_import(capped, overflow) + + assert capped.bundle is not None + assert len(capped.bundle.junit_evidence_imports) == 20 + + +@pytest.mark.parametrize( + ("field", "value", "message"), + [ + ("repository", "other/repository", "active review identity"), + ("pr_number", 2, "active review identity"), + ("head_sha", "b" * 40, "active review identity"), + ("criteria_revision_number", 2, "criteria revision"), + ("confirmed_criteria_sha256", "f" * 64, "criteria digest"), + ], +) +def test_junit_import_append_rejects_stale_or_foreign_relationship_atomically( + field: str, value: object, message: str +) -> None: + state = exact_head_state() + record = junit_import_for(state).model_copy(update={field: value}) + + with pytest.raises(ValueError, match=message): + append_junit_evidence_import(state, record) + + assert state.bundle is not None + assert state.bundle.junit_evidence_imports == [] + + +def test_junit_import_append_rejects_duplicate_id_or_artifact_atomically() -> None: + state = exact_head_state() + record = junit_import_for(state) + imported = append_junit_evidence_import(state, record) + assert imported.bundle is not None + + with pytest.raises(ValueError, match="already imported"): + append_junit_evidence_import(imported, record) + + conflicting_id = record.model_copy( + update={"artifact_sha256": "e" * 64} + ) + with pytest.raises(ValueError, match="import ID"): + append_junit_evidence_import(imported, conflicting_id) + + assert len(imported.bundle.junit_evidence_imports) == 1 + + +def test_junit_import_append_requires_active_analysis() -> None: + state = exact_head_state() + record = junit_import_for(state) + pending = revise_criteria( + state, + [Criterion(criterion_id="AC-01", text="Export filtered CSV")], + "Export filtered CSV", + ) + + with pytest.raises(ValueError, match="active analysis"): + append_junit_evidence_import(pending, record) + + assert pending.bundle is None diff --git a/tests/schemas/test_junit_evidence_import.py b/tests/schemas/test_junit_evidence_import.py new file mode 100644 index 00000000..2c7f23e8 --- /dev/null +++ b/tests/schemas/test_junit_evidence_import.py @@ -0,0 +1,315 @@ +from copy import deepcopy +from datetime import UTC, datetime + +import pytest +from pydantic import ValidationError + +from scopeproof_core.criteria.confirmation import normalized_criteria_sha256 +from scopeproof_core.demo import build_demo_review +from scopeproof_core.schemas.models import ( + JUnitEvidenceImport, + ReviewBundle, +) + +HEAD_SHA = "a" * 40 +OTHER_HEAD_SHA = "b" * 40 +ARTIFACT_SHA256 = "c" * 64 + + +def exact_head_bundle() -> ReviewBundle: + bundle = build_demo_review().model_copy(deep=True) + bundle.review.head_sha = HEAD_SHA + bundle.criteria_revision_number = 1 + return ReviewBundle.model_validate(bundle.model_dump(mode="python")) + + +def valid_import_payload(bundle: ReviewBundle | None = None) -> dict[str, object]: + bundle = bundle or exact_head_bundle() + provenance = bundle.review.criteria_source_provenance + assert provenance is not None + criterion_id = bundle.criteria[0].criterion_id + return { + "schema_version": "junit-import-v1", + "import_id": "import-001", + "repository": bundle.review.repository, + "pr_number": bundle.review.pr_number, + "head_sha": bundle.review.head_sha, + "criteria_revision_number": bundle.criteria_revision_number, + "confirmed_criteria_sha256": normalized_criteria_sha256(bundle.criteria), + "criteria_source_provenance": provenance.model_dump(mode="python"), + "artifact_sha256": ARTIFACT_SHA256, + "artifact_format": "junit_xml", + "evidence_boundary": { + "source": "externally_supplied", + "gate_effect": "non_gating", + "execution": "not_executed_by_scopeproof", + "artifact_digest_scope": "imported_bytes_only", + "importer_identity": "asserted_not_authenticated", + "criterion_mapping": "organizational_context_not_proof", + }, + "imported_by": "QA owner", + "imported_at": datetime(2026, 8, 20, tzinfo=UTC), + "totals": { + "total": 2, + "passed": 1, + "failures": 1, + "errors": 0, + "skipped": 0, + }, + "test_cases": [ + { + "test_case_id": "suite-0001-case-0001", + "suite_id": "suite-0001", + "suite_name": "unit", + "class_name": "tests.WidgetTests", + "test_name": "test_export", + "status": "passed", + }, + { + "test_case_id": "suite-0001-case-0002", + "suite_id": "suite-0001", + "suite_name": "unit", + "class_name": None, + "test_name": "test_error", + "status": "failure", + }, + ], + "criterion_mappings": [ + { + "criterion_id": criterion_id, + "test_case_ids": ["suite-0001-case-0001"], + } + ], + "parser_warnings": ["Declared failure count differed from observed results."], + "limitations": [ + "ScopeProof imported externally supplied results and did not execute tests." + ], + } + + +def test_junit_import_accepts_strict_exact_identity_and_sanitized_results() -> None: + record = JUnitEvidenceImport.model_validate(valid_import_payload()) + + assert record.schema_version == "junit-import-v1" + assert record.head_sha == HEAD_SHA + assert record.totals.total == 2 + assert record.criterion_mappings[0].test_case_ids == [ + "suite-0001-case-0001" + ] + assert record.model_dump_json().count("test_error") == 1 + + +@pytest.mark.parametrize("required_field", ["schema_version", "evidence_boundary"]) +def test_junit_import_requires_explicit_envelope_fields(required_field: str) -> None: + payload = valid_import_payload() + payload.pop(required_field) + + with pytest.raises(ValidationError, match=required_field): + JUnitEvidenceImport.model_validate(payload) + + +@pytest.mark.parametrize( + ("path", "value", "message"), + [ + (("schema_version",), "junit-import-v2", "junit-import-v1"), + (("evidence_boundary", "gate_effect"), "gating", "non_gating"), + (("head_sha",), "short", "40"), + (("artifact_sha256",), "A" * 64, "SHA-256"), + (("imported_by",), " ", "non-whitespace"), + (("parser_warnings",), [""], "non-whitespace"), + (("limitations",), [""], "non-whitespace"), + (("test_cases", 0, "test_name"), "", "non-whitespace"), + ], +) +def test_junit_import_rejects_malformed_boundary_fields( + path: tuple[str | int, ...], value: object, message: str +) -> None: + payload = deepcopy(valid_import_payload()) + target: object = payload + for key in path[:-1]: + target = target[key] # type: ignore[index] + target[path[-1]] = value # type: ignore[index] + + with pytest.raises(ValidationError, match=message): + JUnitEvidenceImport.model_validate(payload) + + +def test_junit_import_rejects_naive_timestamp_and_extra_fields() -> None: + payload = valid_import_payload() + payload["imported_at"] = datetime(2026, 8, 20) + payload["raw_xml"] = "" + + with pytest.raises(ValidationError) as exc_info: + JUnitEvidenceImport.model_validate(payload) + + rendered = str(exc_info.value) + assert "timezone-aware" in rendered + assert "raw_xml" in rendered + + +def test_junit_import_rejects_overlong_note_entries() -> None: + payload = valid_import_payload() + payload["limitations"] = ["x" * 1_001] + + with pytest.raises(ValidationError, match="at most 1000"): + JUnitEvidenceImport.model_validate(payload) + + +def test_review_bundle_caps_accumulated_junit_imports() -> None: + bundle = exact_head_bundle() + payload = bundle.model_dump(mode="python") + imports: list[dict[str, object]] = [] + for index in range(21): + imported = deepcopy(valid_import_payload(bundle)) + imported["import_id"] = f"import-{index + 1:03d}" + imported["artifact_sha256"] = f"{index + 1:064x}" + imports.append(imported) + payload["junit_evidence_imports"] = imports + + with pytest.raises(ValidationError, match="at most 20"): + ReviewBundle.model_validate(payload) + + +def test_junit_import_rejects_inconsistent_totals() -> None: + payload = valid_import_payload() + payload["totals"] = { + "total": 2, + "passed": 2, + "failures": 1, + "errors": 0, + "skipped": 0, + } + + with pytest.raises(ValidationError, match="sum to total"): + JUnitEvidenceImport.model_validate(payload) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("suite_name", "/private/workspace/unit"), + ("class_name", "C:\\agent\\tests.Widget"), + ("test_name", "https://ci.example.test/jobs/42"), + ("test_name", "mailto:secret@example.test"), + ("test_name", "case for mailto:secret@example.test"), + ("test_name", "artifact C:relative-secret.xml"), + ], +) +def test_persisted_junit_case_rejects_path_or_url_like_names( + field: str, value: str +) -> None: + payload = valid_import_payload() + payload["test_cases"][0][field] = value # type: ignore[index] + + with pytest.raises(ValidationError, match="path- or URL-like"): + JUnitEvidenceImport.model_validate(payload) + + +def test_junit_import_rejects_unknown_or_duplicate_case_mapping() -> None: + unknown = valid_import_payload() + unknown["criterion_mappings"] = [ + { + "criterion_id": exact_head_bundle().criteria[0].criterion_id, + "test_case_ids": ["suite-9999-case-9999"], + } + ] + with pytest.raises(ValidationError, match="mapped test case IDs must resolve"): + JUnitEvidenceImport.model_validate(unknown) + + duplicate = valid_import_payload() + duplicate["criterion_mappings"] = [ + { + "criterion_id": exact_head_bundle().criteria[0].criterion_id, + "test_case_ids": [ + "suite-0001-case-0001", + "suite-0001-case-0001", + ], + } + ] + with pytest.raises(ValidationError, match="sorted and unique"): + JUnitEvidenceImport.model_validate(duplicate) + + +def test_junit_import_rejects_one_case_mapped_to_multiple_criteria() -> None: + bundle = exact_head_bundle() + first, second = bundle.criteria[:2] + payload = valid_import_payload(bundle) + payload["criterion_mappings"] = [ + { + "criterion_id": first.criterion_id, + "test_case_ids": ["suite-0001-case-0001"], + }, + { + "criterion_id": second.criterion_id, + "test_case_ids": ["suite-0001-case-0001"], + }, + ] + + with pytest.raises(ValidationError, match="multiple criteria"): + JUnitEvidenceImport.model_validate(payload) + + +def test_review_bundle_accepts_matching_import_and_preserves_legacy_absence() -> None: + bundle = exact_head_bundle() + payload = bundle.model_dump(mode="python") + payload["junit_evidence_imports"] = [valid_import_payload(bundle)] + + reopened = ReviewBundle.model_validate(payload) + + assert reopened.junit_evidence_imports[0].artifact_sha256 == ARTIFACT_SHA256 + legacy_payload = bundle.model_dump(mode="python") + legacy_payload.pop("junit_evidence_imports", None) + assert ReviewBundle.model_validate(legacy_payload).junit_evidence_imports == [] + + +@pytest.mark.parametrize( + ("field", "value", "message"), + [ + ("repository", "other/repository", "JUnit import identity"), + ("pr_number", 999, "JUnit import identity"), + ("head_sha", OTHER_HEAD_SHA, "JUnit import identity"), + ("criteria_revision_number", 2, "criteria revision"), + ("confirmed_criteria_sha256", "d" * 64, "criteria digest"), + ], +) +def test_review_bundle_rejects_import_from_another_review_or_criteria_snapshot( + field: str, value: object, message: str +) -> None: + bundle = exact_head_bundle() + imported = valid_import_payload(bundle) + imported[field] = value + payload = bundle.model_dump(mode="python") + payload["junit_evidence_imports"] = [imported] + + with pytest.raises(ValidationError, match=message): + ReviewBundle.model_validate(payload) + + +def test_review_bundle_rejects_unknown_criterion_and_duplicate_import_identity() -> None: + bundle = exact_head_bundle() + imported = valid_import_payload(bundle) + imported["criterion_mappings"] = [ + { + "criterion_id": "AC-UNKNOWN", + "test_case_ids": ["suite-0001-case-0001"], + } + ] + payload = bundle.model_dump(mode="python") + payload["junit_evidence_imports"] = [imported] + with pytest.raises(ValidationError, match="known criteria"): + ReviewBundle.model_validate(payload) + + first = valid_import_payload(bundle) + second = deepcopy(first) + second["import_id"] = "import-002" + duplicate_digest = bundle.model_dump(mode="python") + duplicate_digest["junit_evidence_imports"] = [first, second] + with pytest.raises(ValidationError, match="artifact digests must be unique"): + ReviewBundle.model_validate(duplicate_digest) + + second["artifact_sha256"] = "e" * 64 + second["import_id"] = first["import_id"] + duplicate_id = bundle.model_dump(mode="python") + duplicate_id["junit_evidence_imports"] = [first, second] + with pytest.raises(ValidationError, match="import IDs must be unique"): + ReviewBundle.model_validate(duplicate_id) diff --git a/tests/storage/test_json_store.py b/tests/storage/test_json_store.py index fb128cd8..d8e7a680 100644 --- a/tests/storage/test_json_store.py +++ b/tests/storage/test_json_store.py @@ -4,7 +4,7 @@ import os from concurrent.futures import ThreadPoolExecutor from copy import deepcopy -from datetime import timedelta +from datetime import UTC, datetime, timedelta from hashlib import sha256 from pathlib import Path from threading import Event, Lock @@ -17,10 +17,15 @@ from scopeproof_core.criteria.confirmation import build_criteria_source_provenance from scopeproof_core.demo import build_demo_review from scopeproof_core.gates.evaluator import evaluate_gate +from scopeproof_core.importers.junit import ( + JUnitMappingSelection, + build_junit_evidence_import, +) from scopeproof_core.reporting.exporters import export_html, export_markdown from scopeproof_core.reviews.comparison import compare_reviews from scopeproof_core.reviews.lifecycle import ( append_external_verification, + append_junit_evidence_import, append_resolution, append_runtime_evidence, attach_analysis, @@ -1842,3 +1847,94 @@ def test_load_rejects_a_review_record_symlink_that_escapes_the_store(tmp_path: P with pytest.raises(FileNotFoundError): store.load("review-1") + + +def exact_head_review_state(review_id: str = "junit-review"): + state = review_state(review_id).model_copy(deep=True) + state.review.head_sha = "a" * 40 + assert state.bundle is not None + state.bundle.review.head_sha = "a" * 40 + return type(state).model_validate(state.model_dump(mode="python")) + + +def imported_junit_state(review_id: str = "junit-review"): + state = exact_head_review_state(review_id) + record = build_junit_evidence_import( + state, + ( + b'' + b'RAW-FAILURE-SENTINEL' + ), + [ + JUnitMappingSelection( + scope_id="suite-0001", + criterion_id=state.criteria_revision.criteria[0].criterion_id, + ) + ], + importer="QA owner", + imported_at=datetime(2026, 8, 20, tzinfo=UTC), + import_id="import-001", + ) + return append_junit_evidence_import(state, record) + + +def test_junit_import_round_trips_in_record_version_four_without_raw_xml( + tmp_path: Path, +) -> None: + store = JsonReviewStore(tmp_path) + state = imported_junit_state() + + path = store.save(state) + payload = json.loads(path.read_text(encoding="utf-8")) + reopened = store.load(state.review.review_id) + + assert payload["record_version"] == 4 + assert payload["state"]["bundle"]["junit_evidence_imports"][0][ + "schema_version" + ] == "junit-import-v1" + assert b"RAW-FAILURE-SENTINEL" not in path.read_bytes() + assert reopened == state + + +def test_legacy_record_without_junit_import_field_reopens_as_empty( + tmp_path: Path, +) -> None: + store = JsonReviewStore(tmp_path) + state = exact_head_review_state() + path = store.save(state) + payload = json.loads(path.read_text(encoding="utf-8")) + payload["state"]["bundle"].pop("junit_evidence_imports", None) + path.write_text(json.dumps(payload), encoding="utf-8") + + reopened = store.load(state.review.review_id) + + assert reopened.bundle is not None + assert reopened.bundle.junit_evidence_imports == [] + + +def test_failed_junit_mutation_preserves_saved_record_bytes(tmp_path: Path) -> None: + store = JsonReviewStore(tmp_path) + state = exact_head_review_state() + path = store.save(state) + before = path.read_bytes() + + def stale_transition(current): + record = build_junit_evidence_import( + current, + b'', + [ + JUnitMappingSelection( + scope_id="suite-0001", + criterion_id=current.criteria_revision.criteria[0].criterion_id, + ) + ], + importer="QA", + import_id="import-stale", + ).model_copy(update={"head_sha": "b" * 40}) + return append_junit_evidence_import(current, record) + + with pytest.raises(ValueError, match="active review identity"): + store.mutate(state.review.review_id, stale_transition) + + assert path.read_bytes() == before + assert store.load(state.review.review_id) == state diff --git a/tests/test_repository_contracts.py b/tests/test_repository_contracts.py index 37ef06e2..72eddb34 100644 --- a/tests/test_repository_contracts.py +++ b/tests/test_repository_contracts.py @@ -428,10 +428,10 @@ def test_ci_runs_lint_tests_and_benchmark() -> None: assert "scopeproof_core.evals.runner" in workflow -def test_streamlit_floor_supports_click_time_deferred_exports() -> None: +def test_streamlit_floor_supports_per_widget_junit_upload_limits() -> None: project = tomllib.loads(Path("pyproject.toml").read_text(encoding="utf-8")) - assert "streamlit>=1.52,<2" in project["project"]["dependencies"] + assert "streamlit>=1.53,<2" in project["project"]["dependencies"] def test_locked_development_environment_is_documented_and_verified() -> None: @@ -447,6 +447,7 @@ def test_locked_development_environment_is_documented_and_verified() -> None: assert "uv run pytest" in guide assert "uv run scopeproof benchmark" in guide assert "Streamlit 1.59.1" in guide + assert "Streamlit 1.53 or newer" in guide assert "Streamlit 1.57.0" in guide assert "testing-interface regression" in guide assert "locked-environment:" in workflow @@ -634,6 +635,7 @@ def test_product_surfaces_share_the_supported_theme_and_alpha_action_hierarchy() "secondaryBackgroundColor": "#171a1f", "textColor": "#f7f7f2", } + assert config["server"]["maxUploadSize"] == 2 assert "):focus-visible" in app assert "[data-testid=\"stAppViewContainer\"]" in app assert "@media (prefers-reduced-motion: reduce)" in app @@ -3779,3 +3781,21 @@ def test_informational_check_status_preserves_product_evidence_boundaries() -> N "0/2 reuse-intent signals", ): assert count in normalized + + +def test_stage2_status_separates_delivered_junit_and_check_foundations_from_candidates() -> None: + status = Path("docs/releases/v0.2.3-status-and-next-stages.md").read_text( + encoding="utf-8" + ) + delivered = status.split("## Delivered Stage 2 foundations", maxsplit=1)[1].split( + "## Prioritized post-release decision candidates", maxsplit=1 + )[0] + candidates = status.split( + "## Prioritized post-release decision candidates", maxsplit=1 + )[1].split("## Next executable queue", maxsplit=1)[0] + + assert "exact-SHA informational GitHub Check lifecycle is implemented" in delivered + assert "bounded JUnit XML adapter is implemented" in delivered + assert "Exact-SHA GitHub Check lifecycle:" not in candidates + assert "JUnit-style results" not in candidates + assert "Additional non-executing evidence adapters" in candidates diff --git a/uv.lock b/uv.lock index 84523652..c0f2ad7f 100644 --- a/uv.lock +++ b/uv.lock @@ -1394,7 +1394,7 @@ requires-dist = [ { name = "pytest", marker = "extra == 'dev'", specifier = ">=9.0.3,<10" }, { name = "pytest-cov", marker = "extra == 'dev'", specifier = ">=6,<7" }, { name = "ruff", marker = "extra == 'dev'", specifier = ">=0.5,<1" }, - { name = "streamlit", specifier = ">=1.52,<2" }, + { name = "streamlit", specifier = ">=1.53,<2" }, ] provides-extras = ["dev", "research"]