defuddle extracts clean, readable content from web pages — stripping navigation, ads, and clutter. It reads a URL, a local HTML file, or HTML piped on stdin, and emits HTML, Markdown, or JSON.
Usage: defuddle <command> [flags]
Extract and structure content from web pages.
Commands:
parse (p) Parse and extract content from a URL, HTML file, or stdin.
[<source>] URL or HTML file; reads stdin when omitted.
extractors List registered site-specific extractors.
batch Parse multiple URLs, output JSONL.
Flags:
-h, --help Show context-sensitive help.
--version Print version information and quit.
Run "defuddle <command> --help" for more information on a command.
There are no persistent/global flags beyond --version and --help; each command owns its own flags.
go install github.com/dotcommander/defuddle/cmd/defuddle@latestExtract content from a single source.
parse selects its input in this order:
- A positional argument —
defuddle parse <source>. If it starts withhttp://orhttps://it is fetched as a URL; otherwise it is read as a local file path. -as the argument — read HTML from stdin explicitly.- No argument, but HTML is piped in — stdin is used automatically.
defuddle parse https://example.com/article # URL
defuddle parse ./page.html # local file
curl -s https://example.com | defuddle parse # piped stdin (auto)
curl -s https://example.com | defuddle parse - # explicit stdinLocal input (files and stdin) is capped at 5 MiB; larger input returns an error wrapping defuddle.ErrTooLarge with the source path or stdin in the message. URL fetches cap downloaded bytes before charset decoding. Rendered snapshots over 5 MiB fail with defuddle.ErrTooLarge (exit 2 for explicit --render); no partial HTML is parsed. With --render-auto, a render-cap error falls back to the capped static fetch, just like other render-stage failures.
For a local file source, relative URLs in the extracted content resolve against a file:// URL derived from the file's absolute path (matching the TypeScript CLI's JSDOM.fromFile behavior); parsing from stdin leaves relative URLs untouched.
By default parse prints the extracted HTML content to stdout. Change the format with:
defuddle parse URL --markdown # Markdown (-m; --md is an alias)
defuddle parse URL --json # full JSON: content + all metadata
defuddle parse URL --property title # a single field, raw
defuddle parse URL --output out.html # write to a file instead of stdout (-o)When more than one is set, precedence is --property > --tables-json > --json > --markdown > default HTML. --markdown falls back to HTML if no markdown was produced.
--property accepts: content, language, title, description, domain, favicon, image, author, site, published, wordCount, parseTime, metaTags, schemaOrgData, extractorType, contentMarkdown.
See JSON output for the full object shape.
By default parse fetches the raw server HTML and does not execute JavaScript — client-rendered (SPA) pages may come back nearly empty. Pass --render (alias --js) to render the page in a headless browser first, then extract:
defuddle parse --render https://example.com/spa-articleThis requires an existing Chrome or Chromium install — defuddle drives it over CDP and bundles no browser. If Chrome is not found, point at one with --chrome-path, or install Chrome/Chromium.
# Wait for the network to settle (good for lazy-loaded content)
defuddle parse --render --render-wait networkidle https://example.com
# Point at a specific browser, cap render time, and set a user agent
defuddle parse --render \
--chrome-path "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" \
--render-timeout 45s \
--render-user-agent "MyBot/1.0" \
https://example.com--render-auto first fetches the static HTML and escalates to Chrome only when
the page looks like a JavaScript shell. If Chrome is unavailable, it falls back
to parsing the static HTML. --render-wait accepts load (default; snapshot
after the load event) or networkidle (snapshot after a brief network-idle
settle). --render-auto escalation defaults to networkidle when
--render-wait is unset, because the JS shells it escalates on hydrate after
the load event; an explicit --render-wait value overrides this on every path.
Use --render-wait-for for a stable CSS selector, or --render-settle
for a fixed post-load delay. Render tuning flags only affect a render path.
Rendering requires an http(s) URL source. With a file or stdin source the render flags cannot apply (there is no URL for Chrome to navigate), so the CLI prints a one-line warning to stderr and parses the input statically.
Charset note: pages that declare no charset anywhere (neither an HTTP header
nor a <meta> tag) are decoded by Chrome's legacy-encoding fallback on the
render path, which can garble non-ASCII text; the static fetch path assumes
UTF-8. Pages that declare any charset are unaffected.
# Custom timeout
defuddle parse https://example.com --timeout 60s
# Custom user agent
defuddle parse https://example.com --user-agent "MyBot/1.0"
# Custom headers (repeatable)
defuddle parse https://example.com -H "Authorization: Bearer token123"
defuddle parse https://example.com -H "Cookie: session=abc" -H "Accept-Language: en"
# Route through a proxy
defuddle parse https://example.com --proxy http://localhost:8080
defuddle parse https://example.com --proxy socks5://localhost:1080Headers must use the Key: Value form. Invalid headers are rejected before any HTTP request is issued. Proxy URLs accept http://, https://, and socks5:// schemes; any other scheme is rejected as invalid input (exit 2) before a request is made.
Fetch --timeout and --render-timeout are independent stage budgets. Parsing after a successful stage uses the caller context. Caller cancellation stops rendering and automatic fallback; a render-stage timeout alone still permits --render-auto static fallback. Static URL parsing resolves relative URLs against the final response URL after redirects. Rendered snapshots retain the configured source URL as their base.
# Remove all images from output
defuddle parse https://example.com --remove-images
# Force a specific content root (bypass auto-detection)
defuddle parse https://example.com --content-selector "article.post-body"
# Bypass extraction (process the body with formatting and safety)
defuddle parse https://example.com --no-clutter-removal
# Debug mode (shows pipeline and recovery steps, timings, statistics)
defuddle parse https://example.com --debug| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--json |
-j |
bool | false | Output as JSON with metadata and content |
--tables-json |
bool | false | Output detected tables as structured JSON | |
--markdown |
-m |
bool | false | Convert content to markdown format |
--md |
bool | false | Alias for --markdown |
|
--property |
-p |
string | Extract a single property (e.g. title, author) |
|
--output |
-o |
string | stdout | Write output to a file |
--user-agent |
string | Custom user agent string (default: built-in) | ||
--header |
-H |
string[] | Custom HTTP headers, Key: Value (repeatable) |
|
--timeout |
duration | 30s | HTTP request timeout | |
--proxy |
string | Proxy URL (http://, https://, socks5://) |
||
--debug |
bool | false | Enable debug output | |
--remove-images |
bool | false | Strip images from content | |
--content-selector |
string | CSS selector for content root (bypasses auto-detection) | ||
--no-clutter-removal |
bool | false | Bypass extraction and process the body | |
--render |
bool | false | Render JavaScript via headless Chrome before extracting | |
--render-auto |
bool | false | Render only pages detected as JavaScript-heavy | |
--js |
bool | false | Alias for --render |
|
--render-wait |
string | load | Render wait strategy: load or networkidle (default load; networkidle under --render-auto when unset) |
|
--render-wait-for |
string | Wait for a visible CSS selector before snapshot | ||
--render-settle |
duration | 0 | Extra post-load delay before snapshot | |
--render-user-agent |
string | User agent for the render stage (default: Chrome default) | ||
--chrome-path |
string | Path to a Chrome/Chromium executable (default: auto-detect) | ||
--render-timeout |
duration | 30s | Maximum time to spend rendering the page |
Parse multiple URLs concurrently. Reads one URL per line from stdin (default) or a file, and outputs JSONL — one JSON object per line.
# From stdin
echo -e "https://example.com/a\nhttps://example.com/b" | defuddle batch
# From file
defuddle batch --input urls.txt
# Control concurrency
defuddle batch --input urls.txt --concurrency 10
# Include markdown in output
defuddle batch --input urls.txt --markdown
# Skip failures instead of stopping
defuddle batch --input urls.txt --continue-on-error
# Bound total batch duration (0 = no overall deadline)
defuddle batch --input urls.txt --timeout 2m
# Save results
defuddle batch --input urls.txt > results.jsonlBlank lines and lines beginning with # are skipped. Each input line is bounded at 64 KiB; longer lines surface as an error rather than being silently truncated. Results are written in input order. --concurrency values below 1 are rejected (exit 2); note that the short form -c -3 is parsed as a flag by the CLI parser — use the long equals form (--concurrency=-3) to probe negative values.
batch writes one JSON object per line (JSONL) to stdout. Successful results emit the full defuddle.Result JSON on their own line. With --continue-on-error, a failed URL emits a per-line error object instead of aborting:
{"url":"https://example.com/broken","error":"<message>"}Without --continue-on-error, the first failure terminates the batch with a non-zero exit.
| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--input |
-i |
string | stdin | Read URLs from a file instead of stdin |
--concurrency |
-c |
int | 5 | Maximum concurrent requests |
--markdown |
-m |
bool | false | Include markdown in output |
--continue-on-error |
bool | false | Continue processing on individual URL errors | |
--timeout |
duration | 0 | Overall batch deadline (e.g. 30s, 2m); 0 disables |
List all registered site-specific extractors.
defuddle extractorsCheck which extractor matches a URL:
defuddle extractors --match https://github.com/dotcommander/defuddle/issues/1| Flag | Type | Description |
|---|---|---|
--match |
string | Show which extractor matches the given URL |
--json (parse) and every successful line of batch emit a defuddle.Result object.
Always present:
content, title, description, domain, favicon, image, language, parseTime, published, author, site, schemaOrgData, wordCount
Present when applicable (omitted when empty):
contentMarkdown, extractorType, variables, metaTags, debugInfo
defuddle parse https://example.com --json | jq '{title, author, wordCount}'- Extracted content is written to stdout; status notices (e.g.
Output written to <file>) and error messages go to stderr, so a piped stdout stays clean. - The exit code classifies the failure so a calling script or agent can branch without parsing stderr.
0is success; any error the CLI does not recognize stays1(unchanged from prior behavior). Recognized failure categories use codes2–6:
| Code | Category | Meaning | Example |
|---|---|---|---|
0 |
success | Command completed | — |
1 |
error | Unclassified error (fallback) | JSON marshal failure, HTTP client build failure |
2 |
validation | Bad flags, arguments, or input | Malformed -H header, unknown --property, non-HTML input, input over the 5 MiB cap, unwritable --output path |
3 |
not_found | Source file / input file missing | defuddle parse ./missing.html, batch --input missing.txt |
4 |
upstream | Fetch / HTTP / network failure | Non-2xx HTTP status, 304 Not Modified |
5 |
precondition | Operator action needed | --render with no Chrome/Chromium installed |
6 |
cancelled | Context cancelled or timeout | --timeout / --render-timeout exceeded, interrupted |
- The classification is applied to the single top-level error via
errors.Is; the error message is printed to stderr regardless of code. Example:--renderon a machine without Chrome exits5withchrome/chromium not found: install Google Chrome or Chromium, or set --chrome-path to an existing executable.
defuddle --version
# v0.16.0 (commit: 4f1c2ab, built: 2026-10-04)The output format is <version> (commit: <hash>, built: <date>). The commit hash and build date are injected at build time; when they are not injected, both print as unknown. On Go 1.24+ a plain go build inside a VCS checkout also embeds the module's VCS-stamped version, so a tagged-but-dirty tree prints something like v0.16.0+dirty (commit: unknown, built: unknown); outside a checkout (or before Go 1.24) it stays dev (commit: unknown, built: unknown).
defuddle parse https://blog.example.com/post --markdowndefuddle parse --render --render-wait networkidle https://example.com/spa-article --markdowndefuddle parse https://example.com --property titlecat urls.txt | defuddle batch --markdown --continue-on-error > results.jsonldefuddle parse https://example.com --debug --json 2>/dev/null | jq '.debugInfo'defuddle parse https://example.com/private \
-H "Authorization: Bearer mytoken" \
-H "Cookie: session=abc123"
## Generic extraction engine
Defuddle uses go-trafilatura v2.2.6 for generic article extraction, with its
native fallback enabled. Site-specific extractors, Defuddle metadata, CJK-aware
word counts, HTML safety processing, and Markdown conversion remain available.
A matching `ContentSelector` takes the first subtree before site dispatch; a
selector miss continues normal extraction.
The five removal controls (`RemoveExactSelectors`, `RemovePartialSelectors`,
`RemoveHiddenElements`, `RemoveLowScoring`, `RemoveContentPatterns`) are
deprecated compatibility fields. Individual combinations have no effect on
extraction. Setting all five to false bypasses extraction and processes the
selected subtree or body; the CLI's `--no-clutter-removal` retains this behavior.
`RemoveImages` remains effective on every path. The six processor gates retain
their defaults and control normalization independently of basic preservation.
The adapter preserves selected block and inline code, including whitespace and
language attributes, supported MathML/KaTeX/MWE math, and supported local
footnote relationships. Referenced definitions appear once in reference order.
Disabled normalization still preserves already-supported safe markup. Arbitrary
widgets, canvas equations, remote footnotes, and ambiguous IDs are outside this
contract.
Upstream failures, panics, unusable results, or malformed preservation markers
recover using a marker-free, processed body snapshot. Recovery favors retaining
content and may include page clutter. Debug processing steps report the reason.
Cancellation and acquisition, parsing, or serialization failures remain errors.
The adapter resolves links using the page/base URL and sanitizes restored
fragments and complete output.
The library remains Chrome-free. Consumers keep their existing Defuddle calls
and need dependency bumps after release. Release the library before updating
released CLI or consumer pins; workspace builds alone do not prove standalone
installation.