You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The pmc full-text provider's HTML fallback cannot succeed from a Python HTTP client, and I do not think that is fixable within the provider. This is a question about whether to keep advertising it, not a bug report.
PMC answers requests with a bot-check interstitial carried on an HTTP 200, so no status check sees it. Measured against PMC6451728:
client / headers
bytes
result
requests, default
21,214
interstitial
requests, User-Agent: curl/8.7.1
21,210
interstitial
requests, honest tool UA + contact address
21,316
interstitial
requests, Accept-Encoding: identity
21,214
interstitial
curl, default
236,615
article
Giving requests curl's own User-Agent still returns the interstitial, so the discrimination is below the header layer — TLS or HTTP fingerprint. I did not pin down which, because the only way past it is impersonating a browser's fingerprint, and that is not something this library should do.
So _fetch_pmc_html returns None for every article, every time. #94 fixed its container selector, which was independently wrong — the container moved from div.article-body to <section class="body main-article-body"> — but that only means the extraction would work if the fetch ever returned an article page. It does not.
Why it matters beyond a dead code path
Two things downstream depend on the route existing:
pmc sits in the default full_text_providers chain. Every reference that reaches it spends a rate_limit_delay plus a request, and the only possible outcomes are None or a TransientFullTextError on a 429. In a batch validation that is real time spent on a call that cannot succeed.
The stale-HTML warning used to tell users to retry.Let a reviewed cache serve stale HTML, and find PMC's current article container #94 changed that text, because "retry when the source serves full text again" is advice a curator cannot act on for a PMC-only article. The underlying situation is still that a cached PMC HTML entry can never be repaired by this library.
Options
Drop the HTML fallback and let pmc be XML-only. Honest, removes a guaranteed-failing request from the default chain. Costs nothing that currently works.
Keep it but mark it non-functional — leave the code for a future where PMC serves plain clients again, and take it out of the default full_text_providers so it is opt-in.
Keep it as is and accept that one provider in the default chain is a no-op, now that its selector is at least correct.
I lean toward the second: the code is right, the access is not, and nothing distinguishes those two states in a log line today.
Worth noting the one legitimate route still works — Europe PMC's fullTextXML for OA-subset articles. This is only about the HTML page fallback for articles outside that subset, which is also the set where the licence would not permit redistribution anyway.
The
pmcfull-text provider's HTML fallback cannot succeed from a Python HTTP client, and I do not think that is fixable within the provider. This is a question about whether to keep advertising it, not a bug report.PMC answers
requestswith a bot-check interstitial carried on an HTTP 200, so no status check sees it. Measured against PMC6451728:requests, defaultrequests,User-Agent: curl/8.7.1requests, honest tool UA + contact addressrequests,Accept-Encoding: identitycurl, defaultGiving
requestscurl's own User-Agent still returns the interstitial, so the discrimination is below the header layer — TLS or HTTP fingerprint. I did not pin down which, because the only way past it is impersonating a browser's fingerprint, and that is not something this library should do.So
_fetch_pmc_htmlreturnsNonefor every article, every time. #94 fixed its container selector, which was independently wrong — the container moved fromdiv.article-bodyto<section class="body main-article-body">— but that only means the extraction would work if the fetch ever returned an article page. It does not.Why it matters beyond a dead code path
Two things downstream depend on the route existing:
pmcsits in the defaultfull_text_providerschain. Every reference that reaches it spends arate_limit_delayplus a request, and the only possible outcomes areNoneor aTransientFullTextErroron a 429. In a batch validation that is real time spent on a call that cannot succeed.Options
pmcbe XML-only. Honest, removes a guaranteed-failing request from the default chain. Costs nothing that currently works.full_text_providersso it is opt-in.I lean toward the second: the code is right, the access is not, and nothing distinguishes those two states in a log line today.
Worth noting the one legitimate route still works — Europe PMC's
fullTextXMLfor OA-subset articles. This is only about the HTML page fallback for articles outside that subset, which is also the set where the licence would not permit redistribution anyway.Found while investigating monarch-initiative/dismech#12672.