Skip to content

Use CommonMark for shared Markdown link extraction - #154

Merged
dacharyc merged 2 commits into
mainfrom
fix/commonmark-link-extraction-151
Sep 28, 2026
Merged

dacharyc merged 2 commits into
mainfrom
fix/commonmark-link-extraction-151

Conversation

@dacharyc

Copy link
Copy Markdown
Member

Summary

Closes #151.

  • Replace the llms.txt link regex and handwritten portability scanner with shared CommonMark parsing using the existing GFM table extension.
  • Recognize inline and reference links, balanced destinations, escaped syntax, and entities decoded exactly once. Exclude code, HTML comments, and standalone images from navigation while keeping portability's image reporting separate.
  • Preserve original-source UTF-16 offsets, exclusive end positions, same-length blanking, malformed-URL diagnostics, vendor table-fence defenses, and consumer-specific autolink/bare-URL policies.
  • Keep scope filters, URL normalization, traversal depth, sampling strategies, curated-list semantics, and verification limits unchanged. No per-discovered-URL probes are added; corrected candidates can change normal request totals.
  • Document intentional corrections and the extraction contract in the check documentation and page-discovery notes.

Regression Coverage

  • Reproduced the four reported gaps before replacing extraction.
  • Added LF/CRLF, non-ASCII offsets, nested code/comments, reference precedence, linked badges, entity decoding, and pagination continuation tests.
  • Added check-level outcomes and exact request assertions for discovery, coverage, Markdown fetching, link resolution, portability, and random/deterministic/curated sampling.
  • Kept the Fix false pagination failures in hard-wrapped Markdown #150 pagination regressions passing. Heading extraction, content parity, code-fence validity, and llms.txt structural rules remain outside this migration.

Validation

  • npm test: 1,906 tests pass across 68 files.
  • npm run lint: passed.
  • npm run build: passed.
  • npm run version:check: passed.
  • Prettier checks for all changed files: passed.
  • The repository-wide format check flagged only an unrelated, untracked local .vscode/settings.json; that file is not included in this PR.

Validation was run locally on Node 25.2.1; CI covers Node 22 and 24.

Replace handwritten link extraction with the shared CommonMark/GFM parser. Preserve source offsets, vendor table defenses, classification policies, and request budgets.

Add extraction and consumer regressions covering references, nested destinations, code and comments, entity decoding, images, and sampling. Document intentional behavior corrections.

Fixes #151

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Score-affecting behavior changed without the repository-required scoring-version and synchronized documentation updates.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Replaces handwritten Markdown link extraction with shared CommonMark/GFM parsing across discovery, validation, portability, and pagination.

Changes:

  • Adds parser-backed extraction with correct references, decoding, offsets, and exclusions.
  • Preserves consumer-specific URL and pagination behavior.
  • Adds broad regression coverage and documentation.
File Description
src/​helpers/​parse-markdown.ts Adds shared CommonMark/GFM parser.
src/​helpers/​classify-markdown-links.ts Replaces manual link scanning with AST extraction.
src/​helpers/​detect-pagination.ts Reuses shared parser and scanner.
src/​checks/​content-discoverability/​llms-txt-valid.ts Uses shared extraction for llms.txt links.
docs/​checks/​content-structure.md Documents portability extraction rules.
docs/​checks/​content-discoverability.md Documents llms.txt link handling.
working-notes/​page-discovery-notes.md Records migration decisions and behavior.
test/​unit/​helpers/​classify-markdown-links.test.ts Covers extraction contracts and edge cases.
test/​unit/​helpers/​detect-pagination.test.ts Covers reference decoding and offsets.
test/​unit/​helpers/​get-markdown-content.test.ts Covers fetching and limits.
test/​unit/​helpers/​get-page-urls.test.ts Covers discovery and sampling behavior.
test/​unit/​checks/​markdown-link-portability.test.ts Covers navigation-only portability checks.
test/​unit/​checks/​llms-txt-valid.test.ts Covers validation with parsed links.
test/​unit/​checks/​llms-txt-links-resolve.test.ts Covers reference-link resolution.
test/​unit/​checks/​llms-txt-links-markdown.test.ts Covers Markdown checks for references.
test/​unit/​checks/​llms-txt-coverage.test.ts Covers parsed-link coverage behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread working-notes/page-discovery-notes.md
@dacharyc
dacharyc merged commit cf3742c into main Sep 28, 2026
3 checks passed
@dacharyc
dacharyc deleted the fix/commonmark-link-extraction-151 branch September 28, 2026 13:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Use CommonMark parsing for shared Markdown link extraction

2 participants