Twenty links, ten citations: what a machine finishes and what it hands back

Zwanzig Links, zehn Nachweise: was eine Maschine erledigt und was sie zurückgibt

A reading list of twenty sources through a citation endpoint. Ten came back as complete records with RIS and BibTeX, in eight seconds. The interesting half is the other ten — because only one of them was stopped by a bot defence. Five answered every request in full and simply had no citation data to declare.

Eine Leseliste aus zwanzig Quellen durch einen Zitations-Endpunkt. Zehn kamen als vollständige Nachweise mit RIS und BibTeX zurück, in acht Sekunden. Die interessante Hälfte sind die anderen zehn — denn nur eine davon wurde von einer Bot-Abwehr gestoppt. Fünf beantworteten jede Anfrage vollständig und hatten schlicht keine Zitationsdaten auszuweisen.

3 August 2026 · 20 sources, one pass · raw data3. August 2026 · 20 Quellen, ein Durchlauf · Rohdaten

The question

A bibliography is the part of a piece of work where a machine looks most useful and is hardest to check. Hand an assistant twenty addresses, ask for a reference list, and something plausible comes back for all twenty. The question worth measuring is not how many entries appear. It is how many are records — read from what the page declares about itself — and whether the rest are named as gaps or quietly filled in.

Each source was sent once to extract_citation on this site's MCP endpoint, which reads the citation metadata a page publishes about itself and returns a structured record. Nothing was retried, nothing was chosen for whether it works. Every address was checked in a browser before the run, so that a typo of ours could not be counted as a failure of theirs.

10/20
complete records
1
stopped by a bot defence
5
no citation data on the page
0.4 s
per source

Where the split runs

Not between disciplines, and not between paid and free. It runs between pages built to be cited and pages built to be read.

Kind of sourceComplete records
reference2 of 2
bare doi1 of 1
publisher4 of 5
open access2 of 3
repository1 of 3
preprint0 of 1
official0 of 3
grey lit0 of 1
news0 of 1

Journal publishers are the easiest case, whether the article is paywalled or open: 4/5 and 2/3. A journal page carries citation_author, citation_title and a DOI in its head, because it wants to be indexed. An encyclopaedia entry and a bare DOI resolve just as cleanly.

Official statistics, chambers of commerce and newspapers are the hard case: none of the four produced a record. Not because they defend themselves — they answered every request in full — but because a statistics portal page is a topic overview, not a work, and declares no author, no date and no title of the kind a reference list needs.

The ten that came back, sorted by cause

The distinction matters because each cause needs a different response from whoever is writing. Lumping them together as “blocked” is what makes a citation tool feel unreliable when it is being accurate.

One was stopped by a bot defence

HostKindWhat happened
www.mdpi.comopen accessanswers a browser, refuses the reader

This is the only case in twenty where a browser sees something a server-side reader is not allowed to see. It is also the only case where opening the page yourself changes the outcome — see what to do with the ten.

Four refuse everyone from this address

HostKindWhat happened
www.sciencedirect.compublisherrefuses both — HTTP 403 from a data centre address
papers.ssrn.compreprintrefuses both — HTTP 403 from a data centre address
eur-lex.europa.euofficialrefuses both — HTTP 202 from a data centre address
www.oecd.orgofficialrefuses both — HTTP 403 from a data centre address

These answered 403 to a browser user agent as readily as to a reader. The common factor is not the client but the network: requests from a data centre are refused whatever they claim to be. From a home connection the same pages open normally. That is worth stating plainly, because it is the one result here that would look different measured from a different place.

Five answered in full and had nothing to declare

HostKindWhat happened
www.ssoar.inforepositoryanswers in full, declares no citation data
zenodo.orgrepositoryanswers in full, declares no citation data
www.statistik.atofficialanswers in full, declares no citation data
www.wko.atgrey litanswers in full, declares no citation data
www.derstandard.atnewsanswers in full, declares no citation data

Fifty to ninety kilobytes of perfectly readable HTML, no defence of any kind, and no citation_* metadata, no author, no publication date. A reference for these has to be written by a person who decides what the work is — a page of a statistics portal, an article in a newspaper, a software release on a repository. No amount of retrying changes that, and any tool that returns a tidy entry here has invented the missing half.

A record can carry a title and still not be one

Two of the five silent cases are the trap this measurement was worth doing for. The Zenodo record returns "kjswedberg/kjswedberg.github.io: First Release" with an author attached; the statistics portal returns "Forschung, Innovation, Digitalisierung" with STATISTIK AUSTRIA as author. Both look like results. Both come back with complete: false.

Anything reading the title field and skipping the flag will file both as sources. The lesson is not about this endpoint — it applies to every citation service: read the completeness flag, not the title. On this endpoint that check is one field:

if not record["complete"]:
    hand_back(url, record.get("warning") or "no citation data on the page")

A gap we should close on our side: in these five cases the warning field is empty. complete: false is correct and sufficient to act on, but a reason would be more useful than a silence, and it is noted as such.

Counter-measurement: a second network, and what it did not change

The section below said a home connection should produce a higher completion rate. That claim has now been tested once — from a commercial VPN exit rather than a home line — and it mostly did not hold.

Data centreVPN exit
Complete records1011
Stopped by a bot defence11
Refusing every client from this address44
Answered in full, declared nothing54
Seconds per source0.40.7

The four network-level refusals did not move. ScienceDirect, SSRN, the OECD and EUR-Lex answered the second address exactly as they answered the first. The likely reason is that a commercial VPN exit is itself a data-centre range — so this run swapped one data centre for another rather than testing the claim. Whether a residential connection changes the outcome is still open, and this measurement does not close it.

The one source that changed is Zenodo, and not because of the network: it returned authors, doi, title on the first run and authors, doi, publisher, title, year on the second. The record gained a year, which is the field that decides completeness. Either the deposit was edited between the runs or Zenodo serves its metadata unevenly — from outside, both look the same.

Raw data: second run. Anyone with a residential line is invited to settle the open half — issue 3.

Follow-up measurement: the endpoint changed, the pages did not

Both runs above measured the same endpoint from two networks. A third run on 4 August 2026 measured the same twenty addresses against an improved endpoint, and the split moved again:

3 Aug, data centre3 Aug, VPN exit4 Aug, after identifier derivation
Complete records101114
Seconds per source0.40.70.58

The gain came from the endpoint, not from the pages. An SSRN address carries its abstract ID and an OECD address its publication slug — both resolve to a DOI, and the registration agency (Crossref) returns the authoritative fields. A EUR-Lex address carries its CELEX number, which the Publications Office of the EU resolves in Cellar. And the endpoint now reads Zenodo's deposit metadata in full. Newly complete: Zenodo, SSRN, EUR-Lex and the OECD.

The limit is equally clear. ScienceDirect and MDPI remain walls without a derivable identifier; SSOAR's bot defence covers its API as well; and the three remaining pages — statistik.at, wko.at, derstandard.at — still declare nothing. There the extension's browser path remains the honest way: open the page yourself and cite what you saw.

Raw data: third run, retrieved 4 August 2026.

What this does not settle

Running it yourself

One URL per line in, one importable .ris out, with the refusals named on stderr instead of half-imported. The recipes page has the same thing for Claude Code, Claude Desktop, Python and the browser.

while read -r u; do
  curl -sX POST https://provinglab.dev/mcp \
    -H 'content-type: application/json' \
    -d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",
         \"params\":{\"name\":\"extract_citation\",\"arguments\":{\"url\":\"$u\"}}}" \
  | python3 -c 'import json,sys
d = json.loads(json.load(sys.stdin)["result"]["content"][0]["text"])
sys.stdout.write(d["ris"]) if d.get("complete") else \
  sys.stderr.write("hand back: " + d["url"] + "\n")'
done < reading-list.txt > literature.ris

Then Zotero → File → Import, or Citavi → Import → RIS. No key, no account. One caveat worth knowing: this site sits behind a filter that refuses the user agent Python's urllib sends by default. Set any user agent of your own and it answers normally.

Die Frage

Ein Literaturverzeichnis ist die Stelle einer Arbeit, an der eine Maschine am nützlichsten wirkt und am schwersten zu kontrollieren ist. Gibt man einem Assistenten zwanzig Adressen und bittet um ein Literaturverzeichnis, kommt zu allen zwanzig etwas Plausibles zurück. Die Messfrage ist nicht, wie viele Einträge erscheinen. Sie lautet: Wie viele davon sind Nachweise — gelesen aus dem, was die Seite über sich selbst ausweist —, und wird der Rest als Lücke benannt oder still aufgefüllt?

Jede Quelle ging genau einmal an extract_citation am MCP-Endpunkt dieser Seite, der die Zitationsmetadaten liest, die eine Seite über sich selbst veröffentlicht, und einen strukturierten Nachweis zurückgibt. Nichts wurde wiederholt, nichts danach ausgewählt, ob es funktioniert. Jede Adresse wurde vor dem Lauf in einem Browser geprüft, damit ein eigener Tippfehler nicht als Versagen der Gegenseite gezählt wird.

10/20
vollständige Nachweise
1
von einer Bot-Abwehr gestoppt
5
keine Zitationsdaten auf der Seite
0.4 s
pro Quelle

Wo die Trennlinie verläuft

Nicht zwischen Fachgebieten, und nicht zwischen kostenpflichtig und frei. Sie verläuft zwischen Seiten, die gebaut sind, um zitiert zu werden, und Seiten, die gebaut sind, um gelesen zu werden.

Art der QuelleVollständige Nachweise
Nachschlagewerk2 von 2
nackter DOI1 von 1
Verlag4 von 5
Open Access2 von 3
Repositorium1 von 3
Preprint0 von 1
amtliche Stelle0 von 3
graue Literatur0 von 1
Presse0 von 1

Verlage von Fachzeitschriften sind der einfachste Fall, gleich ob der Artikel hinter einer Bezahlschranke liegt oder offen steht: 4/5 und 2/3. Eine Zeitschriftenseite trägt citation_author, citation_title und einen DOI im Kopf, weil sie indexiert werden will. Ein Enzyklopädie-Eintrag und ein nackter DOI lösen genauso sauber auf.

Amtliche Statistik, Wirtschaftskammern und Zeitungen sind der harte Fall: Keine der vier brachte einen Nachweis hervor. Nicht weil sie sich wehren — sie beantworteten jede Anfrage vollständig —, sondern weil eine Statistikportal-Seite ein Themenüberblick ist und kein Werk: Sie weist weder Autor noch Datum noch einen Titel aus, wie ihn ein Literaturverzeichnis braucht.

Die zehn, die zurückkamen, nach Ursache sortiert

Die Unterscheidung zählt, weil jede Ursache eine andere Reaktion der schreibenden Person verlangt. Alles unter „blockiert" zusammenzufassen ist genau das, was ein Zitationswerkzeug unzuverlässig wirken lässt, während es in Wahrheit genau ist.

Eine wurde von einer Bot-Abwehr gestoppt

HostArtWas geschah
www.mdpi.comopen accessbeantwortet einen Browser, verweigert dem Leser

Das ist der einzige Fall unter zwanzig, in dem ein Browser etwas sieht, das ein serverseitiger Leser nicht sehen darf. Es ist auch der einzige Fall, in dem das eigene Öffnen der Seite das Ergebnis ändert — siehe was mit den zehn zu tun ist.

Vier verweigern jeder Anfrage von dieser Adresse

HostArtWas geschah
www.sciencedirect.compublisherverweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse
papers.ssrn.compreprintverweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse
eur-lex.europa.euofficialverweigert beiden — HTTP 202 von einer Rechenzentrums-Adresse
www.oecd.orgofficialverweigert beiden — HTTP 403 von einer Rechenzentrums-Adresse

Diese antworteten einem Browser-User-Agent genauso bereitwillig mit 403 wie einem Leser. Der gemeinsame Faktor ist nicht der Client, sondern das Netz: Anfragen aus einem Rechenzentrum werden abgewiesen, was immer sie vorgeben zu sein. Von einem Hausanschluss aus öffnen dieselben Seiten normal. Das ist es wert, klar benannt zu werden, denn es ist das eine Ergebnis hier, das von einem anderen Ort aus gemessen anders aussehen würde.

Fünf antworteten vollständig und hatten nichts auszuweisen

HostArtWas geschah
www.ssoar.inforepositoryantwortet vollständig, weist keine Zitationsdaten aus
zenodo.orgrepositoryantwortet vollständig, weist keine Zitationsdaten aus
www.statistik.atofficialantwortet vollständig, weist keine Zitationsdaten aus
www.wko.atgrey litantwortet vollständig, weist keine Zitationsdaten aus
www.derstandard.atnewsantwortet vollständig, weist keine Zitationsdaten aus

Fünfzig bis neunzig Kilobyte einwandfrei lesbares HTML, keinerlei Abwehr, und keine citation_*-Metadaten, kein Autor, kein Veröffentlichungsdatum. Einen Nachweis dafür muss ein Mensch schreiben, der entscheidet, was das Werk ist — eine Seite eines Statistikportals, ein Zeitungsartikel, ein Software-Release in einem Repositorium. Kein noch so häufiges Wiederholen ändert das, und jedes Werkzeug, das hier einen sauberen Eintrag zurückgibt, hat die fehlende Hälfte erfunden.

Ein Nachweis kann einen Titel tragen und trotzdem keiner sein

Zwei der fünf stillen Fälle sind die Falle, für die sich diese Messung gelohnt hat. Der Zenodo-Nachweis meldet "kjswedberg/kjswedberg.github.io: First Release" mit einem Autor; das Statistikportal meldet "Forschung, Innovation, Digitalisierung" mit STATISTIK AUSTRIA als Autor. Beide sehen wie Ergebnisse aus. Beide kommen mit complete: false zurück.

Wer das Titelfeld liest und das Flag überspringt, legt beide als Quellen ab. Die Lehre betrifft nicht nur diesen Endpunkt — sie gilt für jeden Zitationsdienst: Lesen Sie das Vollständigkeits-Flag, nicht den Titel. An diesem Endpunkt ist diese Prüfung ein einziges Feld:

if not record["complete"]:
    hand_back(url, record.get("warning") or "no citation data on the page")

Eine Lücke, die wir auf eigener Seite schließen sollten: In diesen fünf Fällen ist das warning-Feld leer. complete: false ist korrekt und reicht zum Handeln, aber ein Grund wäre nützlicher als Schweigen — und das ist als solches vermerkt.

Gegenmessung: ein zweites Netz, und was es nicht änderte

Der Abschnitt weiter unten sagte voraus, ein Hausanschluss sollte eine höhere Abschlussquote liefern. Diese Behauptung wurde inzwischen einmal geprüft — von einem kommerziellen VPN-Ausgang statt von einem Hausanschluss —, und sie hielt größtenteils nicht stand.

RechenzentrumVPN-Ausgang
Vollständige Nachweise1011
Von einer Bot-Abwehr gestoppt11
Verweigern jedem Client von dieser Adresse44
Antworteten vollständig, wiesen nichts aus54
Sekunden pro Quelle0.40.7

Die vier Verweigerungen auf Netzebene bewegten sich nicht. ScienceDirect, SSRN, die OECD und EUR-Lex beantworteten die zweite Adresse exakt wie die erste. Der wahrscheinliche Grund: Ein kommerzieller VPN-Ausgang ist selbst ein Rechenzentrums-Bereich — dieser Lauf tauschte also ein Rechenzentrum gegen ein anderes, statt die Behauptung zu prüfen. Ob ein privater Anschluss das Ergebnis ändert, bleibt offen; diese Messung schließt die Frage nicht.

Die eine Quelle, die sich änderte, ist Zenodo, und nicht wegen des Netzes: Sie gab im ersten Lauf authors, doi, title zurück, im zweiten authors, doi, publisher, title, year. Der Nachweis gewann ein Jahr — und das ist das Feld, das über Vollständigkeit entscheidet. Entweder wurde der Datensatz zwischen den Läufen bearbeitet, oder Zenodo liefert seine Metadaten uneinheitlich aus; von außen sehen beide gleich aus.

Rohdaten: zweiter Lauf. Wer über einen privaten Anschluss verfügt, ist eingeladen, die offene Hälfte zu klären — Issue 3.

Folgemessung: Der Endpunkt änderte sich, die Seiten nicht

Beide Läufe oben maßen denselben Endpunkt aus zwei Netzen. Ein dritter Lauf am 4. August 2026 maß dieselben zwanzig Adressen gegen einen verbesserten Endpunkt — und die Aufteilung wanderte erneut:

3.8., Rechenzentrum3.8., VPN-Ausgang4.8., nach Kennungs-Ableitung
Vollständige Nachweise101114
Sekunden pro Quelle0.40.70.58

Der Zugewinn kam vom Endpunkt, nicht von den Seiten. Eine SSRN-Adresse trägt ihre Abstract-ID, eine OECD-Adresse ihren Publikations-Slug — beide lösen zu einem DOI auf, und die Registrierungsstelle (Crossref) liefert die autoritativen Angaben. Eine EUR-Lex-Adresse trägt ihre CELEX-Nummer, die das Amt für Veröffentlichungen der EU in Cellar auflöst. Und der Endpunkt liest die Deposit-Metadaten von Zenodo inzwischen vollständig. Neu vollständig: Zenodo, SSRN, EUR-Lex und die OECD.

Die Grenze ist ebenso klar. ScienceDirect und MDPI bleiben Wände ohne ableitbare Kennung; SSOARs Bot-Abwehr deckt auch die API ab; und die drei übrigen Seiten — statistik.at, wko.at, derstandard.at — deklarieren weiterhin nichts. Dort bleibt der Browser-Weg der Erweiterung der ehrliche Weg: die Seite selbst öffnen und das zitieren, was man gesehen hat.

Rohdaten: dritter Lauf, abgerufen am 4. August 2026.

Was hier nicht geklärt ist

Selbst ausführen

Eine URL pro Zeile hinein, eine importierbare .ris heraus — mit den Verweigerungen auf stderr benannt statt halb importiert. Die Rezepte-Seite bietet dasselbe für Claude Code, Claude Desktop, Python und den Browser.

while read -r u; do
  curl -sX POST https://provinglab.dev/mcp \
    -H 'content-type: application/json' \
    -d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",
         \"params\":{\"name\":\"extract_citation\",\"arguments\":{\"url\":\"$u\"}}}" \
  | python3 -c 'import json,sys
d = json.loads(json.load(sys.stdin)["result"]["content"][0]["text"])
sys.stdout.write(d["ris"]) if d.get("complete") else \
  sys.stderr.write("hand back: " + d["url"] + "\n")'
done < reading-list.txt > literature.ris

Dann Zotero → Datei → Importieren oder Citavi → Import → RIS. Kein Schlüssel, kein Konto. Ein Hinweis lohnt sich noch: Diese Seite sitzt hinter einem Filter, der den User-Agent ablehnt, den Pythons urllib standardmäßig sendet. Setzen Sie einen beliebigen eigenen User-Agent, und sie antwortet normal.