You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Use CommonMark parsing for Markdown content-parity text and structure #152
Follow-up to merged PR #150, which introduced CommonMark parsing with GFM table support. Replace the hand-written Markdown text and repeated-item parsers in markdown-content-parity with AST-based extraction while preserving the existing comparison policy.
Recommended sequence: implement shared parsing in #151 first, then reuse that infrastructure here. Do not use a link-discovery projection as parity input: discovery excludes code, but parity must retain literal code text to match HTML pre and code content.
extractMarkdownText(): multiple ordered passes protecting fenced blocks, inline code, escaped punctuation, and numbered headings with placeholders, then stripping formatting and restoring content.
protectCodeSpansAndEscapes(): hand-written CommonMark inline handling.
countMarkdownItems(): separate line-based fence, list, indentation, and pipe-table parsing.
The repeated parser logic is a maintenance risk. This issue proposes a migration, not a claim that every existing path is broken. Historical regressions in #7, #89, #90, #91, #106, and #110 are compatibility evidence and must remain covered.
Scope
Replace only the Markdown-side syntax interpretation needed for text extraction and list/table counting. Reuse existing tests in test/unit/checks/markdown-content-parity.test.ts and existing fixtures/helpers. Inspect working-notes/parity-check-notes.md for prior decisions before changing behavior.
Acceptance Criteria
Establish focused characterization tests before replacing extraction. Keep the historical field regressions passing.
Preserve literal fenced/indented/inline code text, including Markdown-looking syntax, backslashes, and angle-bracket placeholders. Handle backtick and tilde fences, nested containers, and inline delimiter lengths through the parser.
Extract rendered prose consistently for ATX/setext headings, nested lists, blockquotes, emphasis, inline/reference links, escaped punctuation, and character references. Define image-alt and embedded-HTML handling from existing parity semantics; exclude non-rendered comments/definitions.
Verify equivalent HTML/Markdown fixtures under wrapping, LF/CRLF, and supported formatting variations without losing substantive text or inventing missing-content reports.
Preserve intended repeated-item accounting: the largest list with direct items rather than nested bullets counted again, item deduplication, and the largest GFM table with header/delimiter rows excluded. Add cases for multiline items, nested lists, escaped pipes, and code containing list/table examples.
Verify check-level outcomes for complete content, actual missing content, duplicates, and paginated catalogs. Do not make missing content pass by discarding meaningful text.
Preserve public result shapes, HTML-side extraction, comparison thresholds, normalization policy, noise/audience exclusions, and sampling/request budgets. Any necessary behavior correction must be explicit and regression-tested.
Remove obsolete Markdown parser passes once covered by the replacement, rather than leaving parallel implementations.
Run affected tests and all repository CI gates.
Boundaries
This is not the noise-filter redesign in #87. Do not change parity scoring, HTML rendering heuristics, link discovery, heading-quality extraction, or fence-validity diagnostics. CommonMark accepts an unclosed fence at EOF; the separate fence-validity check intentionally flags it and must retain source-level diagnostics.
Share parsing infrastructure with #151 and #150, but keep the parity-specific text/count interpretation local. No implementation is included in this issue.
Purpose
Follow-up to merged PR #150, which introduced CommonMark parsing with GFM table support. Replace the hand-written Markdown text and repeated-item parsers in
markdown-content-paritywith AST-based extraction while preserving the existing comparison policy.Recommended sequence: implement shared parsing in #151 first, then reuse that infrastructure here. Do not use a link-discovery projection as parity input: discovery excludes code, but parity must retain literal code text to match HTML
preandcodecontent.Current Implementation
src/checks/observability/markdown-content-parity.tscontains:extractMarkdownText(): multiple ordered passes protecting fenced blocks, inline code, escaped punctuation, and numbered headings with placeholders, then stripping formatting and restoring content.protectCodeSpansAndEscapes(): hand-written CommonMark inline handling.countMarkdownItems(): separate line-based fence, list, indentation, and pipe-table parsing.The repeated parser logic is a maintenance risk. This issue proposes a migration, not a claim that every existing path is broken. Historical regressions in #7, #89, #90, #91, #106, and #110 are compatibility evidence and must remain covered.
Scope
Replace only the Markdown-side syntax interpretation needed for text extraction and list/table counting. Reuse existing tests in
test/unit/checks/markdown-content-parity.test.tsand existing fixtures/helpers. Inspectworking-notes/parity-check-notes.mdfor prior decisions before changing behavior.Acceptance Criteria
Boundaries
This is not the noise-filter redesign in #87. Do not change parity scoring, HTML rendering heuristics, link discovery, heading-quality extraction, or fence-validity diagnostics. CommonMark accepts an unclosed fence at EOF; the separate fence-validity check intentionally flags it and must retain source-level diagnostics.
Share parsing infrastructure with #151 and #150, but keep the parity-specific text/count interpretation local. No implementation is included in this issue.