Skip to content

Use CommonMark parsing for Markdown content-parity text and structure #152

Description

@dacharyc

Purpose

Follow-up to merged PR #150, which introduced CommonMark parsing with GFM table support. Replace the hand-written Markdown text and repeated-item parsers in markdown-content-parity with AST-based extraction while preserving the existing comparison policy.

Recommended sequence: implement shared parsing in #151 first, then reuse that infrastructure here. Do not use a link-discovery projection as parity input: discovery excludes code, but parity must retain literal code text to match HTML pre and code content.

Current Implementation

src/checks/observability/markdown-content-parity.ts contains:

  • extractMarkdownText(): multiple ordered passes protecting fenced blocks, inline code, escaped punctuation, and numbered headings with placeholders, then stripping formatting and restoring content.
  • protectCodeSpansAndEscapes(): hand-written CommonMark inline handling.
  • countMarkdownItems(): separate line-based fence, list, indentation, and pipe-table parsing.

The repeated parser logic is a maintenance risk. This issue proposes a migration, not a claim that every existing path is broken. Historical regressions in #7, #89, #90, #91, #106, and #110 are compatibility evidence and must remain covered.

Scope

Replace only the Markdown-side syntax interpretation needed for text extraction and list/table counting. Reuse existing tests in test/unit/checks/markdown-content-parity.test.ts and existing fixtures/helpers. Inspect working-notes/parity-check-notes.md for prior decisions before changing behavior.

Acceptance Criteria

  • Establish focused characterization tests before replacing extraction. Keep the historical field regressions passing.
  • Preserve literal fenced/indented/inline code text, including Markdown-looking syntax, backslashes, and angle-bracket placeholders. Handle backtick and tilde fences, nested containers, and inline delimiter lengths through the parser.
  • Extract rendered prose consistently for ATX/setext headings, nested lists, blockquotes, emphasis, inline/reference links, escaped punctuation, and character references. Define image-alt and embedded-HTML handling from existing parity semantics; exclude non-rendered comments/definitions.
  • Verify equivalent HTML/Markdown fixtures under wrapping, LF/CRLF, and supported formatting variations without losing substantive text or inventing missing-content reports.
  • Preserve intended repeated-item accounting: the largest list with direct items rather than nested bullets counted again, item deduplication, and the largest GFM table with header/delimiter rows excluded. Add cases for multiline items, nested lists, escaped pipes, and code containing list/table examples.
  • Verify check-level outcomes for complete content, actual missing content, duplicates, and paginated catalogs. Do not make missing content pass by discarding meaningful text.
  • Preserve public result shapes, HTML-side extraction, comparison thresholds, normalization policy, noise/audience exclusions, and sampling/request budgets. Any necessary behavior correction must be explicit and regression-tested.
  • Remove obsolete Markdown parser passes once covered by the replacement, rather than leaving parallel implementations.
  • Run affected tests and all repository CI gates.

Boundaries

This is not the noise-filter redesign in #87. Do not change parity scoring, HTML rendering heuristics, link discovery, heading-quality extraction, or fence-validity diagnostics. CommonMark accepts an unclosed fence at EOF; the separate fence-validity check intentionally flags it and must retain source-level diagnostics.

Share parsing infrastructure with #151 and #150, but keep the parity-specific text/count interpretation local. No implementation is included in this issue.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions