A Python 3.10+ library that discovers website page URLs by following HTML links. Uses only the standard library at runtime. It does not read sitemaps and excludes external sites, non-HTTP links, URL fragments, and common downloadable assets. Only successfully fetched HTML pages appear in the output.
Install directly from GitHub:
python -m pip install git+https://github.com/PersonaliAI/website-pages.gitOr clone and install locally:
git clone https://github.com/PersonaliAI/website-pages.git
cd website-pages
python -m pip install .This initial release is a synchronous, single-process crawler. It is useful for bounded website discovery; distributed crawling and JavaScript rendering are not implemented. The project is not yet published to PyPI.
from website_pages import extract_page_urls, WebsiteCrawler
urls = extract_page_urls("https://example.com/", max_pages=5000)
print("\n".join(urls))
# Inspect failures and whether the page-request limit was reached:
result = WebsiteCrawler(max_pages=5000).crawl("https://example.com/")
print(result.urls)
print(result.errors)
print(result.skipped)
print(result.truncated)website-pages https://example.com/ --max-pages 5000 > urls.txt
website-pages https://example.com/ --json > urls.json
python -m website_pages https://example.com/ --max-depth 3The starting hostname is the boundary: www.example.com, example.com, and
other subdomains are separate sites. Default-port HTTP/HTTPS transitions are
allowed; redirects to external hosts are blocked before fetching them. Start
with the site's canonical hostname when it redirects between www and non-www.
The result also contains subdomains: sorted, unique related hostnames found
in HTML links and redirects, even when those hosts are not crawled or fail to
load. This is link discovery, not DNS enumeration; unlinked subdomains may be
missing. The main domain itself is excluded from this list.
result = WebsiteCrawler(include_subdomains=True).crawl("https://postiz.com/")
print(result.urls) # HTML pages on the main domain and its subdomains
print(result.subdomains) # Hostnames, e.g. docs.postiz.comwebsite-pages https://postiz.com/ --include-subdomains --json
website-pages https://postiz.com/ --subdomains-only --jsonBy default, the main domain is the starting hostname with a leading www.
removed. When starting from another subdomain, specify the main domain:
website-pages https://docs.example.co.uk/ --root-domain example.co.uk --include-subdomainsThe package does not guess registrable domains or consult a public suffix
list. Supply the actual website domain, not a shared suffix such as co.uk.
Matching requires a dot boundary: evilpostiz.com and postiz.com.evil.com
are excluded. The original port restriction and global page limit still apply.
Each subdomain's robots.txt is checked independently.
Query strings are preserved because they may identify separate pages. Use
--drop-query or include_query=False to remove them before requesting pages.
The default limit is 1,000 attempted page requests, with 0.2 seconds between
requests, a 15-second request timeout, and a 5 MB response limit. max_depth=0
fetches only the starting page. Results are sorted and deduplicated; redirected
pages use their final URL. truncated reports pending work at the request limit,
not pages excluded by the depth setting.
The crawler checks robots.txt by default but never uses its sitemap entries.
Missing robots.txt permits crawling; authorization errors, server errors, and
unreachable robots.txt conservatively block crawling. To disable checking, use
--ignore-robots or respect_robots=False.
No link crawler can guarantee all pages: unlinked pages, login-only pages, JavaScript-generated links, and pages blocked by robots.txt are not discoverable here. This package parses server-returned HTML and does not execute JavaScript. Links in anchors, image-map areas, and frame elements are followed. Forms are not submitted. Common asset extensions are skipped even if a server serves HTML at those addresses.
python -m unittest discover -s tests -vSee CONTRIBUTING.md for development instructions. Licensed under the MIT License.