Skip to content

fix(download-ref): drop .pdf suffix from arXiv PDF URLs - #67

Merged
GiggleLiu merged 1 commit into
mainfrom
fix/arxiv-pdf-url
Sep 21, 2026
Merged

GiggleLiu merged 1 commit into
mainfrom
fix/arxiv-pdf-url

Conversation

@GiggleLiu

Copy link
Copy Markdown
Member

Summary

  • arXiv returns HTTP 406 Not Acceptable for https://arxiv.org/pdf/<id>.pdf regardless of User-Agent; the suffix-free https://arxiv.org/pdf/<id> still serves the PDF.
  • fetch_metadata.py used the suffixed form at both fetch sites (arXiv manifest entries, and DOI entries with an arXiv preprint), so every arXiv PDF download failed.
  • One-line change per site; no behaviour change otherwise.

Test plan

  • python3 -c probe: suffixed URL → 406, suffix-free URL → 200 application/pdf
  • Re-ran fetch_metadata.py --download-arxiv-pdfs on a 79-paper manifest: 66 PDFs fetched (the rest blocked by a VPN-level 406 unrelated to the URL form)

🤖 Generated with Claude Code

arXiv now answers HTTP 406 Not Acceptable for https://arxiv.org/pdf/<id>.pdf
regardless of User-Agent, while the suffix-free https://arxiv.org/pdf/<id>
returns the PDF. Both fetch sites in fetch_metadata.py (arXiv manifest
entries and DOI entries with an arXiv preprint) used the suffixed form, so
every PDF download failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@GiggleLiu
GiggleLiu merged commit 6bb20ba into main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant