Skip to content

fix(download): parse quoted and extended response filenames - #65

Open
KennyMcSimpson wants to merge 1 commit into
ApodexAI:mainfrom
KennyMcSimpson:fix/download-response-filenames
Open

KennyMcSimpson wants to merge 1 commit into
ApodexAI:mainfrom
KennyMcSimpson:fix/download-response-filenames

Conversation

@KennyMcSimpson

Copy link
Copy Markdown

The download runner treats a response filename as a regular-expression match, which truncates quoted names such as filename="paper;v2.pdf". It also does not handle filename* correctly, so a server-provided Unicode filename can be lost or replaced by the fallback URL name.

Use the standard email header parser for quoted parameters and RFC 2231/5987 extended filenames. Prefer the extended filename when both forms are present, and leave literal percent sequences in ordinary filename values intact. The result still goes through the existing basename sanitization.

Validation:

  • pytest tests/test_download_filenames.py -q: 20 passed, covering quoted punctuation, extended filenames, parameter order, URL fallback, explicit-name precedence, and encoded path separators.
  • The CI Ruff command and the new test file pass lint.
  • Earlier full-suite run on this exact source tree: 4,798 passed, 5 skipped; the redirect-validation test fails because the container cannot resolve www.iana.org. That test and its fetch implementation are unchanged.
  • Pyright and import smoke passed. The external API preflight could not complete without a valid API key.

@KennyMcSimpson
KennyMcSimpson marked this pull request as ready for review October 8, 2026 11:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant