When URLSource caches a PDF URL, the resulting entry's title is the URL itself. There is no title to extract from a PDF byte stream, so the reference ends up identified by its own address.
This matters for any consumer that treats title as bibliographic metadata. dismech compares reference_title in its KB against the cached title, so a URL-as-title either blocks that check or gets copied into the knowledge base as though it were the paper's name.
J-STAGE is the case that surfaced it — 53 such references in the dismech cache. A J-STAGE PDF URL has a sibling _article page carrying a citation_title meta tag:
https://www.jstage.jst.go.jp/article/<journal>/<vol>/<issue>/<article>/_pdf
https://www.jstage.jst.go.jp/article/<journal>/<vol>/<issue>/<article>/_article <- citation_title here
The general shape is not J-STAGE-specific. Publishers routinely serve a PDF at one URL and its metadata at a predictable sibling, and citation_title is a near-universal convention (Google Scholar requires it).
Suggested fix
When a url fetch yields full_text_pdf and no title distinct from the identifier, try to recover one — either from the PDF's own metadata where present, or from a landing page discovered by a documented rule. A conservative version that only fires when title == identifier would fix the whole class without changing any entry that already has a real title.
If you would rather not chase landing pages, extracting the embedded PDF /Title where it exists would still cover part of it, and a consumer could then tell "no title available" from "title is the URL".
Context
dismech currently does this as a runtime patch over URLSource.fetch. That makes two patches it still carries — this and #92 — after retiring ten whose fixes landed here. Filing rather than keeping the workaround quiet, per the rule that a patch needs an upstream issue and an exit.
When
URLSourcecaches a PDF URL, the resulting entry's title is the URL itself. There is no title to extract from a PDF byte stream, so the reference ends up identified by its own address.This matters for any consumer that treats
titleas bibliographic metadata. dismech comparesreference_titlein its KB against the cached title, so a URL-as-title either blocks that check or gets copied into the knowledge base as though it were the paper's name.J-STAGE is the case that surfaced it — 53 such references in the dismech cache. A J-STAGE PDF URL has a sibling
_articlepage carrying acitation_titlemeta tag:The general shape is not J-STAGE-specific. Publishers routinely serve a PDF at one URL and its metadata at a predictable sibling, and
citation_titleis a near-universal convention (Google Scholar requires it).Suggested fix
When a
urlfetch yieldsfull_text_pdfand no title distinct from the identifier, try to recover one — either from the PDF's own metadata where present, or from a landing page discovered by a documented rule. A conservative version that only fires whentitle == identifierwould fix the whole class without changing any entry that already has a real title.If you would rather not chase landing pages, extracting the embedded PDF
/Titlewhere it exists would still cover part of it, and a consumer could then tell "no title available" from "title is the URL".Context
dismech currently does this as a runtime patch over
URLSource.fetch. That makes two patches it still carries — this and #92 — after retiring ten whose fixes landed here. Filing rather than keeping the workaround quiet, per the rule that a patch needs an upstream issue and an exit.