Where citation data actually lives: 18 scholarly platforms measured

A reference is only as good as the data behind it. Eighteen platforms a researcher would plausibly use were asked what they declare about their own articles. Eleven answered; the most complete answer did not come from an article page at all.

3 August 2026 · 18 platforms, one pass · raw data

11/18
returned a record
10
with authors and year
8
titles confirmed by Crossref

What each platform declares

PlatformAuthorsYearDOIPagesISSNs
PubMed220170.52
PMC420200.46
Europe PMCRemote end closed connection w55.49
Springer520131.01
PLOS220231.45
BMC620230.91
Nature1320230.98
FrontiersThe server answered 404 Not Found.0.68
arXiv820170.21
bioRxivThe only title the page declares i0.71
Zenodo120230.51
SSOARThe page looks like an error messa0.33
DOAJThe server answered 403 Forbidden.0.14
Wiley020220.41
MDPIThe server answered 403 Forbidden.0.19
ScienceDirectThe server answered 403 Forbidden.0.24
DOI-Resolver220160.41
Wikipedia120060.18

The finding worth acting on

Resolving the DOI beat visiting the article page. The same work that Wiley's own page serves without page numbers came back complete through doi.org — authors, year, journal, volume, pages and ISSN — in 0.4 seconds, from a publisher whose article pages refuse server-side readers outright.

So the practical rule for anyone assembling a bibliography: if you have the DOI, use https://doi.org/…, not the link your search engine gave you. It is faster, more complete, and it works where the publisher's own page does not.

What this is good for

Keeping a source that will not survive the term

Web pages cited in student work vanish — we have measured how often. A capture keeps the page as it looked, with the retrieval time down to the second and its time zone, a checksum of the image data, and the citation record beside it. When the marker asks what the page said in August, the answer is a file rather than a memory.

Turning a reading list into records

Given a list of addresses, the endpoint returns RIS entries that import into Citavi, Zotero or EndNote without retyping — and names the ones it cannot reach, so those can be opened in a browser instead of being silently dropped.

Sources behind a login

A university licence, a library proxy, a paywalled journal: no server-side reader can follow you there, and three of the publishers measured here refuse them outright. A capture extension runs in your own session with your own access, which is why the two approaches belong together rather than competing.

Feeding a long page to a language model

One continuous sheet with real, selectable text — taken from the document rather than recognised from pixels — and page breaks that fall between lines instead of through them. What the model reads is what the page said.

What it is not good for

Limits of this measurement