Skip to content

๐Ÿ“‰ Self-reported limits: what dsh-webfetch will quietly not do | ่‡ชๆŠฅๅฑ€้™๏ผšdsh-webfetch ไธไผšๅš็š„ไบ‹ย #10

Description

@TYEclipse

๐Ÿ“‰ Self-reported limits: what dsh-webfetch will quietly not do

Every plugin has edges. This is ours, written down before you hit them โ€” all four tools here are deliberately read-only and dependency-free, and that constraint is exactly where each limit comes from.

1. It does not run JavaScript

web_fetch / web_text / web_links / web_table read the HTML the server sends. Single-page apps that render their content client-side arrive as an empty shell โ€” you get the layout, not the data. There is no browser engine in the package and there never will be: if you need rendering, drive a real browser.

2. It does not authenticate for you

No cookie jar, no login flow, no OAuth. Pages behind a session will return the login page, and a 200 on the login page is still a 200 โ€” check the body, not just the status. (This is why web_headers exists: get the status, redirect chain and content type before spending a fetch on the body.)

3. Truncation is a feature, not a bug

Output is capped so a huge page cannot flood the context window. When a page is cut, the tool says so โ€” but "the text ends here" and "the page ends here" are different claims, and only the first one is guaranteed. For long documents, fetch the table of contents first and read sections.

4. web_table is a heuristic, not a parser spec

Header detection, colspan/rowspan grid expansion and row caps handle ordinary tables. Tables whose design is decorative (layout tables), whose headers live in the first row group rather than <thead>, or which nest tables inside cells can come back flattened or re-ordered. It reports what it did with the grid โ€” read that before trusting a number column.

5. No anti-bot work, no proxy rotation

One request per call, a normal user agent (configurable), an honest http/https proxy option. If a site returns 403, 429 or a challenge page, that is the answer โ€” there is no retry ladder, no stealth mode and no captcha solving. Respect the site and slow down.

6. Robots/ToS are your call

The tools do not consult robots.txt and will not stop you. That policy is the caller's, deliberately: an agent that fetches a page you asked for should not silently refuse, and an agent that scrapes a site it was asked to respect should not be able to point at the plugin.

If something above blocks a real workflow, say so in this issue and describe the page shape โ€” a reproducible gap is the only thing that moves a limit here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions