Let a consumer trust its cache instead of re-fetching it - #97
Conversation
Staleness re-fetching suits a cache the tool owns and can rebuild. For a cache
that is committed to a repository and reviewed, it inverts: the re-fetch is the
risk, not the repair.
Three failures downstream are the same re-fetch seen from different angles:
dismech#12672 an unstamped full_text_html entry is withheld, and the repair
route it points at cannot be reached
dismech#12867 a fetch returns PMC's bot-check page on an HTTP 200 and its
<title> is compared against the curated title
dismech#12879 a re-fetch shortens 64 bodies and breaks snippets that verified,
in a pull request about something else
Each of those can be attacked on its own -- serve the stale entry anyway, detect
the interstitial, refuse a refresh that shrinks a body. I started doing exactly
that, and the third of those needs a size heuristic that this project already
tried and removed for good reasons.
`trust_cached_entries` removes the cause instead. When set, a cache entry is
served as it is on disk: no staleness re-fetch and no full-text retry, so there
is no fetch during which any of the above can happen. Off by default, so caches
still migrate for consumers that want that.
An imperfect cached entry stays imperfect, which is the trade. That is a content
problem to fix deliberately -- `cache reference --force` still rebuilds one on
purpose -- rather than something to re-litigate on every validation run.
This supersedes the CAPTCHA-detection guard proposed in #96, which is closed: it
recognised the interstitial instead of not asking for it.
just test: 1173 passed, 1 skipped; mypy clean; ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Claude finished @cmungall's task in 57s —— View job PR review:
|
Staleness re-fetching suits a cache the tool owns and can rebuild. For a cache that is committed to a repository and reviewed, it inverts: the re-fetch is the risk, not the repair.
Three failures reported downstream are the same re-fetch seen from different angles:
full_text_htmlentry is withheld, and the repair route it points at cannot be reached<title>is compared against the curated titleEach can be attacked on its own — serve the stale entry anyway, detect the interstitial, refuse a refresh that shrinks a body. I started doing exactly that, and it was the wrong shape: the third needs a size heuristic this project already tried and removed after three review rounds each found a defect in it.
The flag
trust_cached_entries(defaultFalse). When set, a cache entry is served as it is on disk — no staleness re-fetch, no full-text retry. There is no fetch during which any of the above can happen.Two lines at the call site. The default is unchanged, so caches still migrate for consumers that want that.
An imperfect cached entry stays imperfect, and that is the trade. It is a content problem to fix deliberately —
cache reference --forcestill rebuilds one on purpose — rather than something re-litigated on every validation run.Tests
5 tests: served with no network call at all for
full_text_html,full_text_xmlandabstract_only; the file byte-identical afterwards; and the default still re-fetching, so the flag is genuinely opt-in.Also here: the #94 review follow-ups
test_serve_stale_html_optin.pyonly covered_stale_fallback, which needs the source to yield nothing; a PMC-only article still returns its abstract, so the refresh reaches_preserve_cached_full_textinstead. Three tests now pin that.find_pmc_article_bodyloses its underscore and is imported at module level, since two modules share it.PMC_ARTICLE_BODY_CLASSESrecords that its order is a priority order.Supersedes #96
That PR added bot-check phrase matching and a body-size threshold to
URLSource. Closed: it recognised the interstitial instead of not asking for it.just test: 1173 passed, 1 skipped; mypy clean; ruff clean.🤖 Generated with Claude Code