Skip to content

The PMC HTML fallback cannot succeed from a Python HTTP client: keep it, gate it, or drop it? #95

Description

@cmungall

The pmc full-text provider's HTML fallback cannot succeed from a Python HTTP client, and I do not think that is fixable within the provider. This is a question about whether to keep advertising it, not a bug report.

PMC answers requests with a bot-check interstitial carried on an HTTP 200, so no status check sees it. Measured against PMC6451728:

client / headers bytes result
requests, default 21,214 interstitial
requests, User-Agent: curl/8.7.1 21,210 interstitial
requests, honest tool UA + contact address 21,316 interstitial
requests, Accept-Encoding: identity 21,214 interstitial
curl, default 236,615 article

Giving requests curl's own User-Agent still returns the interstitial, so the discrimination is below the header layer — TLS or HTTP fingerprint. I did not pin down which, because the only way past it is impersonating a browser's fingerprint, and that is not something this library should do.

So _fetch_pmc_html returns None for every article, every time. #94 fixed its container selector, which was independently wrong — the container moved from div.article-body to <section class="body main-article-body"> — but that only means the extraction would work if the fetch ever returned an article page. It does not.

Why it matters beyond a dead code path

Two things downstream depend on the route existing:

  1. pmc sits in the default full_text_providers chain. Every reference that reaches it spends a rate_limit_delay plus a request, and the only possible outcomes are None or a TransientFullTextError on a 429. In a batch validation that is real time spent on a call that cannot succeed.
  2. The stale-HTML warning used to tell users to retry. Let a reviewed cache serve stale HTML, and find PMC's current article container #94 changed that text, because "retry when the source serves full text again" is advice a curator cannot act on for a PMC-only article. The underlying situation is still that a cached PMC HTML entry can never be repaired by this library.

Options

  • Drop the HTML fallback and let pmc be XML-only. Honest, removes a guaranteed-failing request from the default chain. Costs nothing that currently works.
  • Keep it but mark it non-functional — leave the code for a future where PMC serves plain clients again, and take it out of the default full_text_providers so it is opt-in.
  • Keep it as is and accept that one provider in the default chain is a no-op, now that its selector is at least correct.

I lean toward the second: the code is right, the access is not, and nothing distinguishes those two states in a log line today.

Worth noting the one legitimate route still works — Europe PMC's fullTextXML for OA-subset articles. This is only about the HTML page fallback for articles outside that subset, which is also the set where the licence would not permit redistribution anyway.

Found while investigating monarch-initiative/dismech#12672.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions