A reading list of twenty sources through a citation endpoint. Ten came back as complete records with RIS and BibTeX, in eight seconds. The interesting half is the other ten — because only one of them was stopped by a bot defence. Five answered every request in full and simply had no citation data to declare.
Eine Leseliste aus zwanzig Quellen durch einen Zitations-Endpunkt. Zehn kamen als vollständige Nachweise mit RIS und BibTeX zurück, in acht Sekunden. Die interessante Hälfte sind die anderen zehn — denn nur eine davon wurde von einer Bot-Abwehr gestoppt. Fünf beantworteten jede Anfrage vollständig und hatten schlicht keine Zitationsdaten auszuweisen.
A bibliography is the part of a piece of work where a machine looks most useful and is hardest to check. Hand an assistant twenty addresses, ask for a reference list, and something plausible comes back for all twenty. The question worth measuring is not how many entries appear. It is how many are records — read from what the page declares about itself — and whether the rest are named as gaps or quietly filled in.
Each source was sent once to extract_citation on this site's
MCP endpoint, which reads the
citation metadata a page publishes about itself and returns a structured
record. Nothing was retried, nothing was chosen for whether it works. Every
address was checked in a browser before the run, so that a typo of ours could
not be counted as a failure of theirs.
Not between disciplines, and not between paid and free. It runs between pages built to be cited and pages built to be read.
| Kind of source | Complete records |
|---|---|
| reference | 2 of 2 |
| bare doi | 1 of 1 |
| publisher | 4 of 5 |
| open access | 2 of 3 |
| repository | 1 of 3 |
| preprint | 0 of 1 |
| official | 0 of 3 |
| grey lit | 0 of 1 |
| news | 0 of 1 |
Journal publishers are the easiest case, whether the article is paywalled or
open: 4/5 and 2/3. A journal page
carries citation_author, citation_title and a DOI in
its head, because it wants to be indexed. An encyclopaedia entry and a bare
DOI resolve just as cleanly.
Official statistics, chambers of commerce and newspapers are the hard case: none of the four produced a record. Not because they defend themselves — they answered every request in full — but because a statistics portal page is a topic overview, not a work, and declares no author, no date and no title of the kind a reference list needs.
The distinction matters because each cause needs a different response from whoever is writing. Lumping them together as “blocked” is what makes a citation tool feel unreliable when it is being accurate.
| Host | Kind | What happened |
|---|---|---|
www.mdpi.com | open access | answers a browser, refuses the reader |
This is the only case in twenty where a browser sees something a server-side reader is not allowed to see. It is also the only case where opening the page yourself changes the outcome — see what to do with the ten.
| Host | Kind | What happened |
|---|---|---|
www.sciencedirect.com | publisher | refuses both — HTTP 403 from a data centre address |
papers.ssrn.com | preprint | refuses both — HTTP 403 from a data centre address |
eur-lex.europa.eu | official | refuses both — HTTP 202 from a data centre address |
www.oecd.org | official | refuses both — HTTP 403 from a data centre address |
These answered 403 to a browser user agent as readily as to a reader. The common factor is not the client but the network: requests from a data centre are refused whatever they claim to be. From a home connection the same pages open normally. That is worth stating plainly, because it is the one result here that would look different measured from a different place.
| Host | Kind | What happened |
|---|---|---|
www.ssoar.info | repository | answers in full, declares no citation data |
zenodo.org | repository | answers in full, declares no citation data |
www.statistik.at | official | answers in full, declares no citation data |
www.wko.at | grey lit | answers in full, declares no citation data |
www.derstandard.at | news | answers in full, declares no citation data |
Fifty to ninety kilobytes of perfectly readable HTML, no defence of any kind,
and no citation_* metadata, no author, no publication date. A
reference for these has to be written by a person who decides what the work
is — a page of a statistics portal, an article in a newspaper, a
software release on a repository. No amount of retrying changes that, and any
tool that returns a tidy entry here has invented the missing half.
Two of the five silent cases are the trap this measurement was worth doing
for. The Zenodo record returns
"kjswedberg/kjswedberg.github.io: First Release" with an author
attached; the statistics portal returns
"Forschung, Innovation, Digitalisierung" with
STATISTIK AUSTRIA as author. Both look like results. Both come
back with complete: false.
Anything reading the title field and skipping the flag will file both as sources. The lesson is not about this endpoint — it applies to every citation service: read the completeness flag, not the title. On this endpoint that check is one field:
if not record["complete"]:
hand_back(url, record.get("warning") or "no citation data on the page")
A gap we should close on our side: in these five cases the
warning field is empty. complete: false is correct
and sufficient to act on, but a reason would be more useful than a silence,
and it is noted
as such.
The section below said a home connection should produce a higher completion rate. That claim has now been tested once — from a commercial VPN exit rather than a home line — and it mostly did not hold.
| Data centre | VPN exit | |
|---|---|---|
| Complete records | 10 | 11 |
| Stopped by a bot defence | 1 | 1 |
| Refusing every client from this address | 4 | 4 |
| Answered in full, declared nothing | 5 | 4 |
| Seconds per source | 0.4 | 0.7 |
The four network-level refusals did not move. ScienceDirect, SSRN, the OECD and EUR-Lex answered the second address exactly as they answered the first. The likely reason is that a commercial VPN exit is itself a data-centre range — so this run swapped one data centre for another rather than testing the claim. Whether a residential connection changes the outcome is still open, and this measurement does not close it.
The one source that changed is Zenodo, and not because of the network: it
returned authors, doi, title on the first run and
authors, doi, publisher, title, year on the second. The record
gained a year, which is the field that decides completeness. Either the
deposit was edited between the runs or Zenodo serves its metadata unevenly —
from outside, both look the same.
Raw data: second run. Anyone with a residential line is invited to settle the open half — issue 3.
Both runs above measured the same endpoint from two networks. A third run on 4 August 2026 measured the same twenty addresses against an improved endpoint, and the split moved again:
| 3 Aug, data centre | 3 Aug, VPN exit | 4 Aug, after identifier derivation | |
|---|---|---|---|
| Complete records | 10 | 11 | 14 |
| Seconds per source | 0.4 | 0.7 | 0.58 |
The gain came from the endpoint, not from the pages. An SSRN address carries its abstract ID and an OECD address its publication slug — both resolve to a DOI, and the registration agency (Crossref) returns the authoritative fields. A EUR-Lex address carries its CELEX number, which the Publications Office of the EU resolves in Cellar. And the endpoint now reads Zenodo's deposit metadata in full. Newly complete: Zenodo, SSRN, EUR-Lex and the OECD.
The limit is equally clear. ScienceDirect and MDPI remain walls without a derivable identifier; SSOAR's bot defence covers its API as well; and the three remaining pages — statistik.at, wko.at, derstandard.at — still declare nothing. There the extension's browser path remains the honest way: open the page yourself and cite what you saw.
Raw data: third run, retrieved 4 August 2026.
One URL per line in, one importable .ris out, with the refusals
named on stderr instead of half-imported. The
recipes page has the same thing for Claude Code,
Claude Desktop, Python and the browser.
while read -r u; do
curl -sX POST https://provinglab.dev/mcp \
-H 'content-type: application/json' \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",
\"params\":{\"name\":\"extract_citation\",\"arguments\":{\"url\":\"$u\"}}}" \
| python3 -c 'import json,sys
d = json.loads(json.load(sys.stdin)["result"]["content"][0]["text"])
sys.stdout.write(d["ris"]) if d.get("complete") else \
sys.stderr.write("hand back: " + d["url"] + "\n")'
done < reading-list.txt > literature.ris
Then Zotero → File → Import, or Citavi → Import → RIS. No
key, no account. One caveat worth knowing: this site sits behind a filter that
refuses the user agent Python's urllib sends by default. Set any
user agent of your own and it answers normally.
Ein Literaturverzeichnis ist die Stelle einer Arbeit, an der eine Maschine am nützlichsten wirkt und am schwersten zu kontrollieren ist. Gibt man einem Assistenten zwanzig Adressen und bittet um ein Literaturverzeichnis, kommt zu allen zwanzig etwas Plausibles zurück. Die Messfrage ist nicht, wie viele Einträge erscheinen. Sie lautet: Wie viele davon sind Nachweise — gelesen aus dem, was die Seite über sich selbst ausweist —, und wird der Rest als Lücke benannt oder still aufgefüllt?
Jede Quelle ging genau einmal an extract_citation am
MCP-Endpunkt dieser Seite,
der die Zitationsmetadaten liest, die eine Seite über sich selbst
veröffentlicht, und einen strukturierten Nachweis zurückgibt. Nichts wurde
wiederholt, nichts danach ausgewählt, ob es funktioniert. Jede Adresse wurde
vor dem Lauf in einem Browser geprüft, damit ein eigener Tippfehler nicht
als Versagen der Gegenseite gezählt wird.
Nicht zwischen Fachgebieten, und nicht zwischen kostenpflichtig und frei. Sie verläuft zwischen Seiten, die gebaut sind, um zitiert zu werden, und Seiten, die gebaut sind, um gelesen zu werden.
| Art der Quelle | Vollständige Nachweise |
|---|---|
| Nachschlagewerk | 2 von 2 |
| nackter DOI | 1 von 1 |
| Verlag | 4 von 5 |
| Open Access | 2 von 3 |
| Repositorium | 1 von 3 |
| Preprint | 0 von 1 |
| amtliche Stelle | 0 von 3 |
| graue Literatur | 0 von 1 |
| Presse | 0 von 1 |
Verlage von Fachzeitschriften sind der einfachste Fall, gleich ob der
Artikel hinter einer Bezahlschranke liegt oder offen steht:
4/5 und 2/3. Eine
Zeitschriftenseite trägt citation_author,
citation_title und einen DOI im Kopf, weil sie indexiert werden
will. Ein Enzyklopädie-Eintrag und ein nackter DOI lösen genauso sauber auf.
Amtliche Statistik, Wirtschaftskammern und Zeitungen sind der harte Fall: Keine der vier brachte einen Nachweis hervor. Nicht weil sie sich wehren — sie beantworteten jede Anfrage vollständig —, sondern weil eine Statistikportal-Seite ein Themenüberblick ist und kein Werk: Sie weist weder Autor noch Datum noch einen Titel aus, wie ihn ein Literaturverzeichnis braucht.
Die Unterscheidung zählt, weil jede Ursache eine andere Reaktion der schreibenden Person verlangt. Alles unter „blockiert" zusammenzufassen ist genau das, was ein Zitationswerkzeug unzuverlässig wirken lässt, während es in Wahrheit genau ist.
| Host | Art | Was geschah |
|---|---|---|
www.mdpi.com | open access | beantwortet einen Browser, verweigert dem Leser |
Das ist der einzige Fall unter zwanzig, in dem ein Browser etwas sieht, das ein serverseitiger Leser nicht sehen darf. Es ist auch der einzige Fall, in dem das eigene Öffnen der Seite das Ergebnis ändert — siehe was mit den zehn zu tun ist.
| Host | Art | Was geschah |
|---|---|---|
www.sciencedirect.com | publisher | verweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse |
papers.ssrn.com | preprint | verweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse |
eur-lex.europa.eu | official | verweigert beiden — HTTP 202 von einer Rechenzentrums-Adresse |
www.oecd.org | official | verweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse |
Diese antworteten einem Browser-User-Agent genauso bereitwillig mit 403 wie einem Leser. Der gemeinsame Faktor ist nicht der Client, sondern das Netz: Anfragen aus einem Rechenzentrum werden abgewiesen, was immer sie vorgeben zu sein. Von einem Hausanschluss aus öffnen dieselben Seiten normal. Das ist es wert, klar benannt zu werden, denn es ist das eine Ergebnis hier, das von einem anderen Ort aus gemessen anders aussehen würde.
| Host | Art | Was geschah |
|---|---|---|
www.ssoar.info | repository | antwortet vollständig, weist keine Zitationsdaten aus |
zenodo.org | repository | antwortet vollständig, weist keine Zitationsdaten aus |
www.statistik.at | official | antwortet vollständig, weist keine Zitationsdaten aus |
www.wko.at | grey lit | antwortet vollständig, weist keine Zitationsdaten aus |
www.derstandard.at | news | antwortet vollständig, weist keine Zitationsdaten aus |
Fünfzig bis neunzig Kilobyte einwandfrei lesbares HTML, keinerlei Abwehr,
und keine citation_*-Metadaten, kein Autor, kein
Veröffentlichungsdatum. Einen Nachweis dafür muss ein Mensch schreiben, der
entscheidet, was das Werk ist — eine Seite eines Statistikportals,
ein Zeitungsartikel, ein Software-Release in einem Repositorium. Kein noch
so häufiges Wiederholen ändert das, und jedes Werkzeug, das hier einen
sauberen Eintrag zurückgibt, hat die fehlende Hälfte erfunden.
Zwei der fünf stillen Fälle sind die Falle, für die sich diese Messung
gelohnt hat. Der Zenodo-Nachweis meldet
"kjswedberg/kjswedberg.github.io: First Release" mit einem
Autor; das Statistikportal meldet
"Forschung, Innovation, Digitalisierung" mit
STATISTIK AUSTRIA als Autor. Beide sehen wie Ergebnisse aus.
Beide kommen mit complete: false zurück.
Wer das Titelfeld liest und das Flag überspringt, legt beide als Quellen ab. Die Lehre betrifft nicht nur diesen Endpunkt — sie gilt für jeden Zitationsdienst: Lesen Sie das Vollständigkeits-Flag, nicht den Titel. An diesem Endpunkt ist diese Prüfung ein einziges Feld:
if not record["complete"]:
hand_back(url, record.get("warning") or "no citation data on the page")
Eine Lücke, die wir auf eigener Seite schließen sollten: In diesen fünf
Fällen ist das warning-Feld leer. complete: false
ist korrekt und reicht zum Handeln, aber ein Grund wäre nützlicher als
Schweigen — und das ist
als solches
vermerkt.
Der Abschnitt weiter unten sagte voraus, ein Hausanschluss sollte eine höhere Abschlussquote liefern. Diese Behauptung wurde inzwischen einmal geprüft — von einem kommerziellen VPN-Ausgang statt von einem Hausanschluss —, und sie hielt größtenteils nicht stand.
| Rechenzentrum | VPN-Ausgang | |
|---|---|---|
| Vollständige Nachweise | 10 | 11 |
| Von einer Bot-Abwehr gestoppt | 1 | 1 |
| Verweigern jedem Client von dieser Adresse | 4 | 4 |
| Antworteten vollständig, wiesen nichts aus | 5 | 4 |
| Sekunden pro Quelle | 0.4 | 0.7 |
Die vier Verweigerungen auf Netzebene bewegten sich nicht. ScienceDirect, SSRN, die OECD und EUR-Lex beantworteten die zweite Adresse exakt wie die erste. Der wahrscheinliche Grund: Ein kommerzieller VPN-Ausgang ist selbst ein Rechenzentrums-Bereich — dieser Lauf tauschte also ein Rechenzentrum gegen ein anderes, statt die Behauptung zu prüfen. Ob ein privater Anschluss das Ergebnis ändert, bleibt offen; diese Messung schließt die Frage nicht.
Die eine Quelle, die sich änderte, ist Zenodo, und nicht wegen des Netzes:
Sie gab im ersten Lauf authors, doi, title zurück, im zweiten
authors, doi, publisher, title, year. Der Nachweis gewann ein
Jahr — und das ist das Feld, das über Vollständigkeit entscheidet. Entweder
wurde der Datensatz zwischen den Läufen bearbeitet, oder Zenodo liefert
seine Metadaten uneinheitlich aus; von außen sehen beide gleich aus.
Rohdaten: zweiter Lauf. Wer über einen privaten Anschluss verfügt, ist eingeladen, die offene Hälfte zu klären — Issue 3.
Beide Läufe oben maßen denselben Endpunkt aus zwei Netzen. Ein dritter Lauf am 4. August 2026 maß dieselben zwanzig Adressen gegen einen verbesserten Endpunkt — und die Aufteilung wanderte erneut:
| 3.8., Rechenzentrum | 3.8., VPN-Ausgang | 4.8., nach Kennungs-Ableitung | |
|---|---|---|---|
| Vollständige Nachweise | 10 | 11 | 14 |
| Sekunden pro Quelle | 0.4 | 0.7 | 0.58 |
Der Zugewinn kam vom Endpunkt, nicht von den Seiten. Eine SSRN-Adresse trägt ihre Abstract-ID, eine OECD-Adresse ihren Publikations-Slug — beide lösen zu einem DOI auf, und die Registrierungsstelle (Crossref) liefert die autoritativen Angaben. Eine EUR-Lex-Adresse trägt ihre CELEX-Nummer, die das Amt für Veröffentlichungen der EU in Cellar auflöst. Und der Endpunkt liest die Deposit-Metadaten von Zenodo inzwischen vollständig. Neu vollständig: Zenodo, SSRN, EUR-Lex und die OECD.
Die Grenze ist ebenso klar. ScienceDirect und MDPI bleiben Wände ohne ableitbare Kennung; SSOARs Bot-Abwehr deckt auch die API ab; und die drei übrigen Seiten — statistik.at, wko.at, derstandard.at — deklarieren weiterhin nichts. Dort bleibt der Browser-Weg der Erweiterung der ehrliche Weg: die Seite selbst öffnen und das zitieren, was man gesehen hat.
Rohdaten: dritter Lauf, abgerufen am 4. August 2026.
Eine URL pro Zeile hinein, eine importierbare .ris heraus — mit
den Verweigerungen auf stderr benannt statt halb importiert. Die
Rezepte-Seite bietet dasselbe für Claude Code,
Claude Desktop, Python und den Browser.
while read -r u; do
curl -sX POST https://provinglab.dev/mcp \
-H 'content-type: application/json' \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",
\"params\":{\"name\":\"extract_citation\",\"arguments\":{\"url\":\"$u\"}}}" \
| python3 -c 'import json,sys
d = json.loads(json.load(sys.stdin)["result"]["content"][0]["text"])
sys.stdout.write(d["ris"]) if d.get("complete") else \
sys.stderr.write("hand back: " + d["url"] + "\n")'
done < reading-list.txt > literature.ris
Dann Zotero → Datei → Importieren oder Citavi → Import →
RIS. Kein Schlüssel, kein Konto. Ein Hinweis lohnt sich noch: Diese
Seite sitzt hinter einem Filter, der den User-Agent ablehnt, den Pythons
urllib standardmäßig sendet. Setzen Sie einen beliebigen
eigenen User-Agent, und sie antwortet normal.