Skip to content

Use CommonMark parsing for shared Markdown link extraction #151

Description

@dacharyc

Purpose

Follow-up to #149 and merged PR #150. Replace overlapping hand-written Markdown link extraction with parser-backed extraction, without refactoring unrelated checks. PR #150 introduced mdast-util-from-markdown, mdast-util-gfm-table, and micromark-extension-gfm-table for pagination detection.

This is the recommended first parser migration. Share parsing infrastructure, not consumer-specific interpretation: discovery excludes code links, while a later content-parity migration must retain code text.

Verified Gaps

Probed during the #150 investigation:

Input extractMarkdownLinks() scanRawLinks()
[guide](/guide_(intro)) Truncates destination to /guide_(intro Correct destination
[guide][ref] with [ref]: /guide No link Correct reference link
Four-space-indented [example](/example) Includes example Includes example
<!-- [hidden](/hidden) --> Includes hidden link Includes hidden link

The pagination fix masks code/comments before calling the shared scanner, but other consumers do not get that protection.

Implementation Surface

  • src/checks/content-discoverability/llms-txt-valid.ts: extractMarkdownLinks().
  • src/helpers/classify-markdown-links.ts: scanRawLinks(), reference resolution, code blanking, and downstream classification.
  • src/helpers/detect-pagination.ts: reuse parser infrastructure where appropriate and retain its distinct bare-URL/path and declaration policies.
  • Inspect consumers in discovery, Markdown fetching, llms.txt link checks, coverage, and markdown-link-portability; avoid maintaining a second extractor for them.
  • Reuse existing tests in test/unit/helpers/classify-markdown-links.test.ts, test/unit/checks/llms-txt-valid.test.ts, and test/unit/helpers/get-page-urls.test.ts, plus affected consumer tests.

Acceptance Criteria

  • Add regressions for the verified gaps before changing extraction.
  • Correctly handle inline and reference links, nested destinations, titles, escaped syntax, images/badges, and entity decoding. Audit raw versus parser-decoded destinations to prevent double decoding.
  • Exclude links in inline/fenced/indented code and HTML comments, including nested containers, one-to-three-space comment indentation, unclosed comments, and LF/CRLF.
  • Preserve exported signatures and result contracts, document order, source offsets/end positions, and the same-length blanked contract used by pagination. Test offsets against original source, including CRLF and non-ASCII text.
  • Preserve intentional classification, deduplication, malformed-destination diagnostics, image separation, and autolink/bare-URL policies. Document any intentional behavior correction rather than silently changing consumer policy.
  • Verify discovery and check-level outcomes, not only extraction. Code/comment links must not generate requests or false broken-link diagnostics; newly recognized real links must participate normally.
  • Preserve scope filters, sampling/verification limits, and curated-input semantics. Add no content-sniffing requests or deeper traversal. Corrected candidates may change total requests; pin intended requests in regression tests rather than promising identical runtime.
  • Keep the Fix false pagination failures in hard-wrapped Markdown #150 pagination regressions passing and run repository CI gates.

Boundaries

Do not migrate content-parity text/count extraction, heading extraction, or code-fence validity here. Changing llms.txt structural validation rules is also separate from replacing its link extractor. Preserve existing MDX/vendor defenses unless a tested replacement is established. Consult existing fixtures and comments before removing behavior that deliberately differs from plain CommonMark.

No implementation is included in this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions