Skip to content

About

Python library and CLI for discovering website page URLs and linked subdomains without sitemaps.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

website-pages

Tests

A Python 3.10+ library that discovers website page URLs by following HTML links. Uses only the standard library at runtime. It does not read sitemaps and excludes external sites, non-HTTP links, URL fragments, and common downloadable assets. Only successfully fetched HTML pages appear in the output.

Install

Install directly from GitHub:

python -m pip install git+https://github.com/PersonaliAI/website-pages.git

Or clone and install locally:

git clone https://github.com/PersonaliAI/website-pages.git
cd website-pages
python -m pip install .

This initial release is a synchronous, single-process crawler. It is useful for bounded website discovery; distributed crawling and JavaScript rendering are not implemented. The project is not yet published to PyPI.

Python

from website_pages import extract_page_urls, WebsiteCrawler

urls = extract_page_urls("https://example.com/", max_pages=5000)
print("\n".join(urls))

# Inspect failures and whether the page-request limit was reached:
result = WebsiteCrawler(max_pages=5000).crawl("https://example.com/")
print(result.urls)
print(result.errors)
print(result.skipped)
print(result.truncated)

Command line

website-pages https://example.com/ --max-pages 5000 > urls.txt
website-pages https://example.com/ --json > urls.json
python -m website_pages https://example.com/ --max-depth 3

The starting hostname is the boundary: www.example.com, example.com, and other subdomains are separate sites. Default-port HTTP/HTTPS transitions are allowed; redirects to external hosts are blocked before fetching them. Start with the site's canonical hostname when it redirects between www and non-www.

Subdomains

The result also contains subdomains: sorted, unique related hostnames found in HTML links and redirects, even when those hosts are not crawled or fail to load. This is link discovery, not DNS enumeration; unlinked subdomains may be missing. The main domain itself is excluded from this list.

result = WebsiteCrawler(include_subdomains=True).crawl("https://postiz.com/")
print(result.urls)        # HTML pages on the main domain and its subdomains
print(result.subdomains) # Hostnames, e.g. docs.postiz.com
website-pages https://postiz.com/ --include-subdomains --json
website-pages https://postiz.com/ --subdomains-only --json

By default, the main domain is the starting hostname with a leading www. removed. When starting from another subdomain, specify the main domain:

website-pages https://docs.example.co.uk/ --root-domain example.co.uk --include-subdomains

The package does not guess registrable domains or consult a public suffix list. Supply the actual website domain, not a shared suffix such as co.uk. Matching requires a dot boundary: evilpostiz.com and postiz.com.evil.com are excluded. The original port restriction and global page limit still apply. Each subdomain's robots.txt is checked independently.

Query strings are preserved because they may identify separate pages. Use --drop-query or include_query=False to remove them before requesting pages. The default limit is 1,000 attempted page requests, with 0.2 seconds between requests, a 15-second request timeout, and a 5 MB response limit. max_depth=0 fetches only the starting page. Results are sorted and deduplicated; redirected pages use their final URL. truncated reports pending work at the request limit, not pages excluded by the depth setting.

The crawler checks robots.txt by default but never uses its sitemap entries. Missing robots.txt permits crawling; authorization errors, server errors, and unreachable robots.txt conservatively block crawling. To disable checking, use --ignore-robots or respect_robots=False.

No link crawler can guarantee all pages: unlinked pages, login-only pages, JavaScript-generated links, and pages blocked by robots.txt are not discoverable here. This package parses server-returned HTML and does not execute JavaScript. Links in anchors, image-map areas, and frame elements are followed. Forms are not submitted. Common asset extensions are skipped even if a server serves HTML at those addresses.

Run tests

python -m unittest discover -s tests -v

See CONTRIBUTING.md for development instructions. Licensed under the MIT License.

About

Python library and CLI for discovering website page URLs and linked subdomains without sitemaps.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages