Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ docs/.vitepress/cache/
docs/deploy
.DS_Store
.claude/
.vscode/

# Raw output from working-notes/run-parity-baseline.sh (snapshots go stale; keep the tables in the notes instead)
parity-results/
Expand Down
6 changes: 6 additions & 0 deletions docs/checks/content-discoverability.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,12 @@ Whether your `llms.txt` follows the [llmstxt.org](https://llmstxt.org/) structur

A well-structured `llms.txt` gives agents a reliable map of the documentation. Inconsistent implementations reduce its value, though even a non-standard file with useful links is better than nothing.

### What is measured

AFDocs recognizes both inline links (`[Guide](/guide)`) and reference links (`[Guide][intro]` with an `[intro]: /guide` definition). It uses CommonMark parsing, including links in Markdown tables. Links inside code examples or HTML comments do not count. Standalone images do not count as navigation links, but a linked image or badge contributes its enclosing link.

Discovery, coverage, and the llms.txt link checks use this same extraction. Destinations with balanced parentheses, such as `/guide_(intro)`, are preserved. Markdown escapes and HTML entities are decoded once, so `/search?a=1&amp;b=2` is checked as `/search?a=1&b=2`. Plain URL text and angle-bracket autolinks such as `<https://example.com/guide>` are not included; use inline or reference links for index entries.

### Results

| Result | Condition |
Expand Down
4 changes: 3 additions & 1 deletion docs/checks/content-structure.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,7 +121,9 @@ The check reads the same markdown responses the other markdown checks already fe

A link whose address is not a valid URL at all fails the page rather than being counted as a well-formed absolute link.

Same-document fragment links (`#anchor` with no path) are exempt: they resolve within the content the agent already holds, and rewriting them to absolute URLs adds nothing. Links with a non-HTTP scheme (`mailto:`, `tel:`) are exempt for the same reason. Links inside fenced code blocks and inline code are ignored, so documentation that shows example markdown is not graded on its examples. Image references are classified and reported but never affect the result: they point at assets rather than at documentation an agent navigates to.
Same-document fragment links (`#anchor` with no path) are exempt: they resolve within the content the agent already holds, and rewriting them to absolute URLs adds nothing. Links with a non-HTTP scheme (`mailto:`, `tel:`) are exempt for the same reason. Links inside fenced or indented code blocks, inline code, and HTML comments are ignored, so documentation that shows example markdown is not graded on its examples. Image references are classified and reported but never affect the result: they point at assets rather than at documentation an agent navigates to.

Link destinations follow CommonMark rules: escaped punctuation and HTML entities are decoded once before resolving the URL. Plain URL text and angle-bracket autolinks are not included in the portability tally.

Relative links are resolved against the URL that served the markdown, not the page URL. For a site serving `/docs/api` as `/md/docs/api.md`, `guide.md` means `/md/docs/guide.md`.

Expand Down
13 changes: 5 additions & 8 deletions src/checks/content-discoverability/llms-txt-valid.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import { registerCheck } from '../registry.js';
import { getLlmsTxtFilesForAnalysis } from '../../helpers/llms-txt.js';
import { scanRawLinks } from '../../helpers/classify-markdown-links.js';
import type { CheckContext, CheckResult } from '../../types.js';

interface ValidationResult {
Expand All @@ -11,15 +12,11 @@ interface ValidationResult {
issues: string[];
}

/** Extract markdown links from text: [name](url) */
/** Extract inline and reference links, excluding images, code and comments. */
export function extractMarkdownLinks(content: string): Array<{ name: string; url: string }> {
const linkRegex = /\[([^\]]+)\]\(([^\s)]+)(?:\s+["'][^"']*["'])?\)/g;
const links: Array<{ name: string; url: string }> = [];
let match;
while ((match = linkRegex.exec(content)) !== null) {
links.push({ name: match[1], url: match[2] });
}
return links;
return scanRawLinks(content)
.links.filter((link) => !link.isImage && link.destination !== '')
.map((link) => ({ name: link.text, url: link.destination }));
}

function validateLlmsTxt(content: string, url: string): ValidationResult {
Expand Down
Loading
Loading