๐ Self-reported limits: what dsh-webfetch will quietly not do
Every plugin has edges. This is ours, written down before you hit them โ all four tools here are deliberately read-only and dependency-free, and that constraint is exactly where each limit comes from.
1. It does not run JavaScript
web_fetch / web_text / web_links / web_table read the HTML the server sends. Single-page apps that render their content client-side arrive as an empty shell โ you get the layout, not the data. There is no browser engine in the package and there never will be: if you need rendering, drive a real browser.
2. It does not authenticate for you
No cookie jar, no login flow, no OAuth. Pages behind a session will return the login page, and a 200 on the login page is still a 200 โ check the body, not just the status. (This is why web_headers exists: get the status, redirect chain and content type before spending a fetch on the body.)
3. Truncation is a feature, not a bug
Output is capped so a huge page cannot flood the context window. When a page is cut, the tool says so โ but "the text ends here" and "the page ends here" are different claims, and only the first one is guaranteed. For long documents, fetch the table of contents first and read sections.
4. web_table is a heuristic, not a parser spec
Header detection, colspan/rowspan grid expansion and row caps handle ordinary tables. Tables whose design is decorative (layout tables), whose headers live in the first row group rather than <thead>, or which nest tables inside cells can come back flattened or re-ordered. It reports what it did with the grid โ read that before trusting a number column.
5. No anti-bot work, no proxy rotation
One request per call, a normal user agent (configurable), an honest http/https proxy option. If a site returns 403, 429 or a challenge page, that is the answer โ there is no retry ladder, no stealth mode and no captcha solving. Respect the site and slow down.
6. Robots/ToS are your call
The tools do not consult robots.txt and will not stop you. That policy is the caller's, deliberately: an agent that fetches a page you asked for should not silently refuse, and an agent that scrapes a site it was asked to respect should not be able to point at the plugin.
If something above blocks a real workflow, say so in this issue and describe the page shape โ a reproducible gap is the only thing that moves a limit here.
๐ Self-reported limits: what
dsh-webfetchwill quietly not doEvery plugin has edges. This is ours, written down before you hit them โ all four tools here are deliberately read-only and dependency-free, and that constraint is exactly where each limit comes from.
1. It does not run JavaScript
web_fetch/web_text/web_links/web_tableread the HTML the server sends. Single-page apps that render their content client-side arrive as an empty shell โ you get the layout, not the data. There is no browser engine in the package and there never will be: if you need rendering, drive a real browser.2. It does not authenticate for you
No cookie jar, no login flow, no OAuth. Pages behind a session will return the login page, and a
200on the login page is still a200โ check the body, not just the status. (This is whyweb_headersexists: get the status, redirect chain and content type before spending a fetch on the body.)3. Truncation is a feature, not a bug
Output is capped so a huge page cannot flood the context window. When a page is cut, the tool says so โ but "the text ends here" and "the page ends here" are different claims, and only the first one is guaranteed. For long documents, fetch the table of contents first and read sections.
4.
web_tableis a heuristic, not a parser specHeader detection,
colspan/rowspangrid expansion and row caps handle ordinary tables. Tables whose design is decorative (layout tables), whose headers live in the first row group rather than<thead>, or which nest tables inside cells can come back flattened or re-ordered. It reports what it did with the grid โ read that before trusting a number column.5. No anti-bot work, no proxy rotation
One request per call, a normal user agent (configurable), an honest
http/httpsproxy option. If a site returns403,429or a challenge page, that is the answer โ there is no retry ladder, no stealth mode and no captcha solving. Respect the site and slow down.6. Robots/ToS are your call
The tools do not consult
robots.txtand will not stop you. That policy is the caller's, deliberately: an agent that fetches a page you asked for should not silently refuse, and an agent that scrapes a site it was asked to respect should not be able to point at the plugin.If something above blocks a real workflow, say so in this issue and describe the page shape โ a reproducible gap is the only thing that moves a limit here.