Skip to content

Use CommonMark parsing for Markdown content parity - #155

Merged
dacharyc merged 2 commits into
mainfrom
fix/152-commonmark-content-parity
Sep 28, 2026
Merged

dacharyc merged 2 commits into
mainfrom
fix/152-commonmark-content-parity

Conversation

@dacharyc

Copy link
Copy Markdown
Member

Summary

Closes #152.

  • Replace the Markdown-side text-stripping passes and line-based list/table parser with the shared CommonMark/GFM AST, parsed once per compared page.
  • Preserve literal fenced, indented, and inline code; extract rendered prose and image alt text; exclude comments, reference definitions, and link metadata. Retain historical bare placeholders such as <REGION>.
  • Count direct items in the largest list and data rows in the largest GFM table. Keep multiline items intact and deduplicate their complete normalized rendered text, including nested content. Code examples do not contribute list/table counts.
  • Add 20 check-level regressions and document extraction semantics and intentional counting corrections.

Compatibility

  • HTML-side extraction, normalization, comparison thresholds, public result shapes, audience exclusions, and request/sampling behavior remain unchanged.
  • Parity disables the shared helper's vendor-table fence normalization so valid fence metadata containing | cannot turn literal code into prose. Discovery and pagination retain their existing default behavior.
  • Deduplication now uses complete rendered item text rather than the first source line; wrapping, formatting, and link destinations alone do not make an item distinct. Item-count diagnostics remain informational.
  • Heading-quality extraction, link discovery, and fence-validity diagnostics are not migrated by this PR.

Verification

  • Full test suite: 1,926 tests passing on Node 22.22.0 and Node 24.21.0, including all 84 parity tests.
  • Version references, lint/typecheck, and build pass on both runtimes; repository format check passes on Node 22.
  • New regressions were exercised against the previous implementation before replacement.
  • Request assertions verify one HTML fetch per uncached page and cache reuse.
  • Confirmed the HTML extractor, normalization function, and item-count comparison logic are unchanged.

Replace Markdown text stripping and repeated-item parsing with the shared CommonMark/GFM AST. Preserve literal code, historical placeholders, and existing HTML extraction and comparison thresholds.

Add 20 regressions covering rendered text, code containers, metadata exclusions, multiline catalogs, nested lists, and GFM tables. Document extraction semantics and keep vendor-fence normalization unchanged for discovery consumers.

Refs #152.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Embedded SVG markup can cause false parity failures, and the score-changing update lacks the required scoring-version bump.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)
What changed in this PR

Migrates Markdown parity analysis to the shared CommonMark/GFM AST for consistent text and structure extraction.

Changes:

  • Replaces regex-based Markdown extraction and counting with AST traversal.
  • Adds regressions for code, prose, lists, tables, and caching.
  • Documents the revised extraction and deduplication semantics.
File Description
src/​checks/​observability/​markdown-content-parity.ts Implements AST-based parity extraction and counting.
src/​helpers/​parse-markdown.ts Adds optional vendor-fence normalization.
test/​unit/​checks/​markdown-content-parity.test.ts Adds CommonMark parity regressions.
docs/​checks/​observability.md Documents extraction and counting behavior.
working-notes/​parity-check-notes.md Records implementation decisions.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/checks/observability/markdown-content-parity.ts Outdated
Comment thread src/checks/observability/markdown-content-parity.ts
Extract Markdown text and embedded markup through a shared HTML fragment so paired SVG, MathML, and custom elements are not mistaken for literal placeholders. Escape Markdown text to preserve code and retain standalone unknown placeholders using parsed source ranges.

Add 10 regressions covering inline markup, visible text, standalone placeholders, and non-rendered SVG titles. Document the extraction semantics.

Addresses review feedback on #155.
@dacharyc
dacharyc merged commit 4baba53 into main Sep 28, 2026
3 checks passed
@dacharyc
dacharyc deleted the fix/152-commonmark-content-parity branch September 28, 2026 14:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Use CommonMark parsing for Markdown content-parity text and structure

2 participants