A reading list goes in, citable records come out — for two thirds of it. The remaining third is refused by the publishers to any server that asks. What makes the endpoint useful is not the two thirds; it is that it names the third precisely enough to act on.
Eine Leseliste geht hinein, zitierfähige Nachweise kommen heraus — für zwei Drittel davon. Das restliche Drittel verweigern die Verlage jedem Server, der nachfragt. Was den Endpunkt nützlich macht, sind nicht die zwei Drittel, sondern dass er das letzte Drittel so genau benennt, dass man damit weiterarbeiten kann.
An agent is handed a list of addresses and asked to prepare them for a
bibliography. For each one it calls extract_citation on this
site's MCP endpoint, which
reads whatever citation data the page declares about itself and returns a
structured record with RIS and BibTeX.
Complete records — authors, year, container, identifier — each with an RIS entry that imports into Citavi, Zotero or EndNote without retyping. A journal article from Springer with five authors, a PLOS article, a PubMed Central paper, an arXiv preprint, a Wikipedia entry, a Zenodo record, and a bare DOI that resolved through the registration agency because the publisher's page refused.
Two of the eight carry no author, because the page declares none. That is the page's gap, not the reader's — and it is reported as an empty field rather than filled in with a guess. A bibliography entry that looks complete and is wrong costs more than one with a visible hole.
| Source | Reason |
|---|---|
www.mdpi.com | http-403 |
www.sciencedirect.com | http-403 |
www.ssoar.info | bot wall served to server-side readers |
eur-lex.europa.eu | near-empty response to server-side readers |
Every one of these was re-fetched from an unrelated network to check the block is real rather than a fault at this end. All four are genuine.
A citation tool that returns something for every URL sounds better and is worse. Fed a bot wall, it produces a reference whose title reads Making sure you're not a bot! — formatted, complete-looking, and worthless. We have measured that happening to an established service on two of eighteen random sources.
A refusal, by contrast, is actionable. The agent can tell its user exactly four addresses to open in a browser, where a capture extension reaches what no server can: the page is loaded in the reader's own session, with the reader's own access. The division of labour is not a workaround — it follows from who is allowed to see what.
curl -X POST https://provinglab.dev/mcp \
-H 'content-type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{
"name":"extract_citation",
"arguments":{"url":"https://arxiv.org/abs/1706.03762"}}}'
No key and no account. Please use it in proportion — it is one small endpoint, and a reading list is a handful of calls, not a crawl. A record comes back with a
source field saying where the details were read, and a
warning field that is empty when there is nothing to warn about.
Ein Agent bekommt eine Liste von Adressen und soll sie für ein
Literaturverzeichnis vorbereiten. Für jede ruft er
extract_citation am
MCP-Endpunkt dieser Seite auf.
Der Endpunkt liest die Zitationsdaten, die eine Seite über sich selbst
ausweist, und gibt einen strukturierten Nachweis mit RIS und BibTeX zurück.
Vollständige Nachweise — Autoren, Jahr, Publikationsorgan, Identifikator —, jeder mit einem RIS-Eintrag, der sich ohne Abtippen in Citavi, Zotero oder EndNote importieren lässt. Ein Zeitschriftenartikel von Springer mit fünf Autoren, ein PLOS-Artikel, ein Beitrag aus PubMed Central, ein arXiv-Preprint, ein Wikipedia-Eintrag, ein Zenodo-Datensatz und ein nackter DOI, der über die Registrierungsagentur aufgelöst wurde, weil die Seite des Verlags die Antwort verweigerte.
Zwei der acht tragen keinen Autor, weil die Seite keinen ausweist. Das ist die Lücke der Seite, nicht die des Werkzeugs — und sie wird als leeres Feld gemeldet statt mit einer Vermutung gefüllt. Ein Eintrag im Literaturverzeichnis, der vollständig aussieht und falsch ist, kostet mehr als einer mit sichtbarer Lücke.
| Quelle | Grund |
|---|---|
www.mdpi.com | http-403 |
www.sciencedirect.com | http-403 |
www.ssoar.info | Bot-Abwehr-Seite an serverseitige Leser ausgeliefert |
eur-lex.europa.eu | nahezu leere Antwort an serverseitige Leser |
Jede dieser Adressen wurde zur Kontrolle von einem unabhängigen Netz aus erneut abgerufen, damit eine Störung auf eigener Seite nicht für eine Sperre gehalten wird. Alle vier Sperren sind echt.
Ein Zitationswerkzeug, das zu jeder Adresse etwas zurückgibt, klingt besser und ist schlechter. Mit einer Bot-Abwehr gefüttert, erzeugt es einen Beleg, dessen Titel Making sure you're not a bot! lautet — formatiert, vollständig wirkend und wertlos. Gemessen haben wir das an einem etablierten Dienst, bei zwei von achtzehn zufällig gezogenen Quellen.
Eine Verweigerung dagegen lässt sich umsetzen. Der Agent kann genau die vier Adressen nennen, die im Browser zu öffnen sind — dort erreicht eine Aufnahme-Erweiterung, was kein Server erreicht: Die Seite lädt in der eigenen Sitzung, mit dem eigenen Zugang. Diese Arbeitsteilung ist kein Notbehelf, sie folgt daraus, wer was sehen darf.
curl -X POST https://provinglab.dev/mcp \
-H 'content-type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{
"name":"extract_citation",
"arguments":{"url":"https://arxiv.org/abs/1706.03762"}}}'
Kein Schlüssel, kein Konto. Bitte im Verhältnis nutzen — es ist ein kleiner
Endpunkt, und eine Leseliste ist eine Handvoll Aufrufe, kein Crawl. Der
Nachweis kommt mit einem source-Feld zurück, das angibt, wo die
Einzelheiten gelesen wurden, und einem warning-Feld, das leer
ist, wenn es nichts zu warnen gibt.