Skip to content

Use CommonMark parsing for section-header-quality Markdown headings #153

Description

@dacharyc

Purpose

Follow-up to merged PR #150, which introduced CommonMark parsing with GFM table support. Replace the Markdown heading regex in section-header-quality with parser-backed heading extraction while preserving existing HTML/MDX and callout behavior.

Recommended sequence: reuse any shared parser helper established by #151. This is a separate bounded migration and does not depend on completing the content-parity migration.

Current Implementation and Risks

src/checks/content-structure/section-header-quality.ts uses extractHeaders() to combine HTML headings from node-html-parser with Markdown matches from MD_HEADING_RE.

The Markdown pass scans raw source for ATX headings. It does not recognize setext headings or distinguish a heading from matching text inside a code example. It also leaves Markdown inline formatting in the extracted heading text. Add reproductions before changing behavior; these risks were identified by inspection, unlike the independently reproduced link-extraction bugs tracked in #151.

The HTML path intentionally excludes headings in callout/admonition containers using ancestor tags, ARIA roles, classes, and data attributes. That behavior fixed #51 and must be preserved.

Implementation Surface

  • src/checks/content-structure/section-header-quality.ts: extractHeaders(), MD_HEADING_RE, and isCalloutHeading().
  • test/unit/checks/section-header-quality.test.ts: extend existing check-level fixtures.
  • Inspect the tabbed-content-serialization panel contract and HTML/MDX examples before deciding how parsed Markdown and HTML headings are combined.

Acceptance Criteria

  • Add focused regressions for missed setext headings, ATX-looking code/comment examples, closing ATX hashes, and formatted heading labels.
  • Extract heading text from actual Markdown heading nodes. Handle emphasis, links, inline code, escaped punctuation, and entities according to rendered-text semantics.
  • Exclude heading-like text in fenced/indented code and HTML comments, including nested containers and indented comments. Verify LF/CRLF behavior.
  • Preserve HTML heading extraction and the callout/admonition exclusions from Callout headings trigger "section-header-quality: tab headers don't distinguish between variants" #51. Respect ancestor context in mixed HTML/Markdown or MDX panels; a second extraction path must not reintroduce excluded headings or count a heading twice.
  • Test equivalent HTML and Markdown panels and the intended outcomes for contextual versus repeated generic headers, including cross-group comparisons.
  • Preserve check dependencies, public details/results, classification thresholds, and request behavior. Keep the work limited to accurately extracting header text and structure.
  • Remove the replaced Markdown heading regex, reuse shared parser infrastructure where appropriate, and run affected tests plus repository CI gates.

Boundaries

Do not replace HTML parsing with CommonMark, introduce a wholesale MDX migration, or redefine which supplementary callout headings count as structural headers. Do not migrate content-start detection, soft-404 heuristics, llms.txt structural validation, link extraction, content parity, or fence validity in this issue.

The parser determines Markdown structure; the existing check still decides whether a section label distinguishes tab variants. No implementation is included in this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions