You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #149 and merged PR #150. Replace overlapping hand-written Markdown link extraction with parser-backed extraction, without refactoring unrelated checks. PR #150 introduced mdast-util-from-markdown, mdast-util-gfm-table, and micromark-extension-gfm-table for pagination detection.
This is the recommended first parser migration. Share parsing infrastructure, not consumer-specific interpretation: discovery excludes code links, while a later content-parity migration must retain code text.
src/helpers/classify-markdown-links.ts: scanRawLinks(), reference resolution, code blanking, and downstream classification.
src/helpers/detect-pagination.ts: reuse parser infrastructure where appropriate and retain its distinct bare-URL/path and declaration policies.
Inspect consumers in discovery, Markdown fetching, llms.txt link checks, coverage, and markdown-link-portability; avoid maintaining a second extractor for them.
Reuse existing tests in test/unit/helpers/classify-markdown-links.test.ts, test/unit/checks/llms-txt-valid.test.ts, and test/unit/helpers/get-page-urls.test.ts, plus affected consumer tests.
Acceptance Criteria
Add regressions for the verified gaps before changing extraction.
Correctly handle inline and reference links, nested destinations, titles, escaped syntax, images/badges, and entity decoding. Audit raw versus parser-decoded destinations to prevent double decoding.
Exclude links in inline/fenced/indented code and HTML comments, including nested containers, one-to-three-space comment indentation, unclosed comments, and LF/CRLF.
Preserve exported signatures and result contracts, document order, source offsets/end positions, and the same-length blanked contract used by pagination. Test offsets against original source, including CRLF and non-ASCII text.
Preserve intentional classification, deduplication, malformed-destination diagnostics, image separation, and autolink/bare-URL policies. Document any intentional behavior correction rather than silently changing consumer policy.
Verify discovery and check-level outcomes, not only extraction. Code/comment links must not generate requests or false broken-link diagnostics; newly recognized real links must participate normally.
Preserve scope filters, sampling/verification limits, and curated-input semantics. Add no content-sniffing requests or deeper traversal. Corrected candidates may change total requests; pin intended requests in regression tests rather than promising identical runtime.
Do not migrate content-parity text/count extraction, heading extraction, or code-fence validity here. Changing llms.txt structural validation rules is also separate from replacing its link extractor. Preserve existing MDX/vendor defenses unless a tested replacement is established. Consult existing fixtures and comments before removing behavior that deliberately differs from plain CommonMark.
Purpose
Follow-up to #149 and merged PR #150. Replace overlapping hand-written Markdown link extraction with parser-backed extraction, without refactoring unrelated checks. PR #150 introduced
mdast-util-from-markdown,mdast-util-gfm-table, andmicromark-extension-gfm-tablefor pagination detection.This is the recommended first parser migration. Share parsing infrastructure, not consumer-specific interpretation: discovery excludes code links, while a later content-parity migration must retain code text.
Verified Gaps
Probed during the #150 investigation:
extractMarkdownLinks()scanRawLinks()[guide](/guide_(intro))/guide_(intro[guide][ref]with[ref]: /guide[example](/example)<!-- [hidden](/hidden) -->The pagination fix masks code/comments before calling the shared scanner, but other consumers do not get that protection.
Implementation Surface
src/checks/content-discoverability/llms-txt-valid.ts:extractMarkdownLinks().src/helpers/classify-markdown-links.ts:scanRawLinks(), reference resolution, code blanking, and downstream classification.src/helpers/detect-pagination.ts: reuse parser infrastructure where appropriate and retain its distinct bare-URL/path and declaration policies.markdown-link-portability; avoid maintaining a second extractor for them.test/unit/helpers/classify-markdown-links.test.ts,test/unit/checks/llms-txt-valid.test.ts, andtest/unit/helpers/get-page-urls.test.ts, plus affected consumer tests.Acceptance Criteria
blankedcontract used by pagination. Test offsets against original source, including CRLF and non-ASCII text.Boundaries
Do not migrate content-parity text/count extraction, heading extraction, or code-fence validity here. Changing llms.txt structural validation rules is also separate from replacing its link extractor. Preserve existing MDX/vendor defenses unless a tested replacement is established. Consult existing fixtures and comments before removing behavior that deliberately differs from plain CommonMark.
No implementation is included in this issue.