>

Who actually reads this site: 24 hours of crawler logs, published

Plenty is claimed about AI crawlers and very little is shown. Anyone running a site can see in their own logs which systems actually arrive, what they take and how often — and almost nobody publishes it. Here is one day of it, with the query that produced it.

Collected 2026-08-03 · 23.9 h window · raw data

The headline

AI systems made more requests than search engines did — 299 against 282. The search engines pulled more data (4.0 MB against 2.1 MB), because Googlebot fetches images and stylesheets that a language model has no use for.

This domain was registered on 1 August 2026. It is two days old, ranks for nothing, and has no inbound links worth counting — and seven distinct AI systems have already read it.

AI systems

Requests by AI user agent, 23.9 hours to 2026-08-03 18:04 UTC
SystemRequests DataShare
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, on behalf of a human)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Search engines, for comparison

Same window, conventional crawlers
CrawlerRequests DataShare
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot at 130 requests is within reach of Googlebot, which is not the ratio most sites see. Bingbot at 10 is the surprise in the other direction.

Twenty requests were not reading

Separated out rather than counted: requests carrying an AI user agent that went straight for files nobody links to.

Credential scans arriving under AI or crawler user agents
User agent claimedRequests Paths sought
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, on behalf of a human)5/keys.json

A user agent is not an identity. It is a string the client chooses. Whether these came from the named operators or from someone borrowing their name cannot be determined from this side, and this page does not claim to know. What can be said: they were not reading, all of them got 404, and nothing they sought exists here.

What the AI systems actually fetched

The most-requested paths across all of them are the machine-readable ones — /sitemap.xml, /llms.txt, /mcp — before the articles. That is the whole argument for maintaining those files: they are not decoration, they are what gets read first.

What this is not

These numbers cannot be checked from outside. They come from our own provider's analytics, retrieved with a token only the operator holds. Nobody can recompute them. What is published instead is the method: the query runs in tools/crawler-bericht.py, in plain text, and anyone with a Cloudflare zone can run the same one against their own.

So this is a self-report with a disclosed method, not a measurement someone could repeat. Everywhere else on this site that distinction is the point, and it would be dishonest to blur it here because the numbers happen to be flattering.

Two further limits. One day is one day — a crawler that visits weekly is invisible in it. And requests are counted at the edge, so a system reading a cached copy through an intermediary never appears at all.

Method: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, grouped by user agent over a 24-hour window, collected 2026-08-03. Requests for paths that only a scanner seeks are separated out and not counted as reading. The query is published; the underlying log is not accessible to anyone but the operator, so these figures are a self-report and are labelled as one. A user agent can be set freely and is not proof of origin. Nothing here is legal advice.

Corrections are welcome and are made in public: open an issue.

← Proving Lab · Disclaimer

Wer diese Seite wirklich liest: 24 Stunden Crawler-Protokoll, veröffentlicht

Über KI-Crawler wird viel behauptet und wenig gezeigt. Wer eine Seite betreibt, kann im eigenen Protokoll sehen, welche Systeme tatsächlich kommen, was sie holen und wie oft — und fast niemand veröffentlicht es. Hier ist ein Tag davon, mitsamt der Abfrage, die ihn erzeugt hat.

Erhoben am 03.08.2026 · Zeitfenster 23,9 h · Rohdaten

Der Befund

Die KI-Systeme haben mehr Anfragen gestellt als die Suchmaschinen — 299 gegen 282. Die Suchmaschinen haben mehr Daten gezogen (4,0 MB gegen 2,1 MB), weil Googlebot Bilder und Stilvorlagen holt, mit denen ein Sprachmodell nichts anfangen kann.

Diese Domain wurde am 1. August 2026 registriert. Sie ist zwei Tage alt, rankt für nichts und hat keine nennenswerten eingehenden Links — und sieben verschiedene KI-Systeme haben sie bereits gelesen.

KI-Systeme

Anfragen nach KI-User-Agent, 23,9 Stunden bis 2026-08-03 18:04 UTC
SystemAnfragen DatenAnteil
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, im Auftrag eines Menschen)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Suchmaschinen zum Vergleich

Dasselbe Zeitfenster, herkömmliche Crawler
CrawlerAnfragen DatenAnteil
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot liegt mit 130 Anfragen in Reichweite von Googlebot, und das ist nicht das Verhältnis, das die meisten Seiten sehen. Bingbot mit 10 ist die Überraschung in die andere Richtung.

Zwanzig Anfragen haben nicht gelesen

Herausgerechnet statt mitgezählt: Anfragen mit einem KI-User-Agent, die direkt auf Dateien zielten, die niemand verlinkt.

Zugangsdaten-Scans unter KI- oder Crawler-User-Agents
Behaupteter User-AgentAnfragen Gesuchte Pfade
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, im Auftrag eines Menschen)5/keys.json

Ein User-Agent ist keine Identität. Er ist eine Zeichenkette, die der Client wählt. Ob diese Anfragen von den genannten Betreibern kamen oder von jemandem, der sich ihren Namen leiht, lässt sich von dieser Seite aus nicht feststellen, und diese Seite behauptet nicht, es zu wissen. Sagen lässt sich: Sie haben nicht gelesen, alle bekamen 404, und nichts von dem, was sie suchten, gibt es hier.

Was die KI-Systeme tatsächlich geholt haben

Die meistgefragten Pfade über alle hinweg sind die maschinenlesbaren — /sitemap.xml, /llms.txt, /mcp — vor den Artikeln. Das ist das ganze Argument dafür, diese Dateien zu pflegen: Sie sind keine Zierde, sie sind das, was zuerst gelesen wird.

Was das nicht ist

Diese Zahlen sind von außen nicht überprüfbar. Sie stammen aus der Auswertung unseres eigenen Anbieters, abgerufen mit einem Token, das nur der Betreiber hat. Niemand kann sie nachrechnen. Veröffentlicht wird stattdessen die Methode: Die Abfrage steht in tools/crawler-bericht.py, im Klartext, und wer eine Cloudflare-Zone hat, kann dieselbe gegen die eigene laufen lassen.

Das hier ist also eine Selbstauskunft mit offengelegter Methode, keine Messung, die jemand wiederholen könnte. Überall sonst auf dieser Seite ist genau dieser Unterschied der Punkt, und es wäre unredlich, ihn hier zu verwischen, weil die Zahlen zufällig schmeichelhaft sind.

Zwei weitere Grenzen. Ein Tag ist ein Tag — ein Crawler, der wöchentlich vorbeikommt, ist darin unsichtbar. Und gezählt wird am Rand des Netzes, ein System also, das über einen Zwischenspeicher eine Kopie liest, taucht überhaupt nicht auf.

Methode: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, gruppiert nach User-Agent über ein 24-Stunden-Fenster, erhoben am 2026-08-03. Anfragen auf Pfade, die nur ein Scanner sucht, sind herausgerechnet und nicht als Lesen gezählt. Die Abfrage ist veröffentlicht; das zugrunde liegende Protokoll ist für niemanden außer dem Betreiber zugänglich, diese Zahlen sind also eine Selbstauskunft und als solche gekennzeichnet. Ein User-Agent lässt sich frei setzen und ist kein Herkunftsnachweis. Nichts davon ist eine Rechtsberatung.

Korrekturen sind willkommen und werden öffentlich gemacht: ein Issue eröffnen.

← Proving Lab · Haftungsausschluss

Quién lee realmente este sitio: 24 horas de registros de rastreadores, publicadas

Sobre los rastreadores de IA se afirma mucho y se muestra muy poco. Quien gestiona un sitio puede ver en sus propios registros qué sistemas llegan de verdad, qué se llevan y con qué frecuencia — y casi nadie lo publica. Aquí hay un día de ello, junto con la consulta que lo produjo.

Recogido el 2026-08-03 · ventana de 23,9 h · datos en bruto

El titular

Los sistemas de IA hicieron más peticiones que los motores de búsqueda — 299 frente a 282. Los motores de búsqueda se llevaron más datos (4,0 MB frente a 2,1 MB), porque Googlebot recupera imágenes y hojas de estilo que a un modelo de lenguaje no le sirven de nada.

Este dominio se registró el 1 de agosto de 2026. Tiene dos días, no posiciona para nada y no tiene enlaces entrantes dignos de contarse — y siete sistemas de IA distintos ya lo han leído.

Sistemas de IA

Peticiones por agente de usuario de IA, 23,9 horas hasta 2026-08-03 18:04 UTC
SistemaPeticiones DatosCuota
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, por encargo de una persona)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Motores de búsqueda, para comparar

La misma ventana, rastreadores convencionales
RastreadorPeticiones DatosCuota
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot, con 130 peticiones, está al alcance de Googlebot, y esa no es la proporción que ve la mayoría de los sitios. Bingbot, con 10, es la sorpresa en la dirección contraria.

Veinte peticiones no estaban leyendo

Separadas en lugar de contadas: peticiones que llevaban un agente de usuario de IA y fueron directas a archivos que nadie enlaza.

Escaneos de credenciales llegados bajo agentes de usuario de IA o de rastreadores
Agente de usuario declaradoPeticiones Rutas buscadas
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, por encargo de una persona)5/keys.json

Un agente de usuario no es una identidad. Es una cadena que elige el cliente. Si estas vinieron de los operadores nombrados o de alguien que toma prestado su nombre no puede determinarse desde este lado, y esta página no pretende saberlo. Lo que sí puede decirse: no estaban leyendo, todas recibieron 404 y nada de lo que buscaban existe aquí.

Qué recuperaron realmente los sistemas de IA

Las rutas más solicitadas por todos ellos son las legibles por máquina — /sitemap.xml, /llms.txt, /mcp — antes que los artículos. Ese es todo el argumento para mantener esos archivos: no son adorno, son lo que se lee primero.

Qué no es esto

Estas cifras no pueden comprobarse desde fuera. Proceden de la analítica de nuestro propio proveedor, obtenidas con un token que solo tiene el operador. Nadie puede recalcularlas. Lo que se publica en su lugar es el método: la consulta está en tools/crawler-bericht.py, en texto plano, y cualquiera que tenga una zona de Cloudflare puede ejecutar la misma contra la suya.

Así que esto es un autoinforme con el método revelado, no una medición que alguien pudiera repetir. En todo el resto de este sitio esa distinción es justamente el asunto, y sería deshonesto difuminarla aquí porque las cifras resulten halagadoras.

Dos límites más. Un día es un día — un rastreador que pasa semanalmente es invisible en él. Y las peticiones se cuentan en el borde, así que un sistema que lee una copia en caché a través de un intermediario no aparece en absoluto.

Método: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, agrupado por agente de usuario sobre una ventana de 24 horas, recogido el 2026-08-03. Las peticiones a rutas que solo busca un escáner se separan y no se cuentan como lectura. La consulta está publicada; el registro subyacente no es accesible para nadie salvo el operador, de modo que estas cifras son un autoinforme y se etiquetan como tal. Un agente de usuario puede fijarse libremente y no prueba el origen. Nada de esto es asesoramiento jurídico.

Las correcciones son bienvenidas y se hacen en público: abrir un issue.

← Proving Lab · Aviso legal

Qui lit vraiment ce site : 24 heures de journaux de robots, publiées

On affirme beaucoup de choses sur les robots d’IA et on en montre très peu. Qui exploite un site peut voir dans ses propres journaux quels systèmes arrivent réellement, ce qu’ils prennent et à quelle fréquence — et presque personne ne le publie. En voici une journée, avec la requête qui l’a produite.

Relevé le 2026-08-03 · fenêtre de 23,9 h · données brutes

Le constat

Les systèmes d’IA ont fait plus de requêtes que les moteurs de recherche — 299 contre 282. Les moteurs de recherche ont tiré plus de données (4,0 MB contre 2,1 MB), parce que Googlebot récupère des images et des feuilles de style dont un modèle de langue n’a aucun usage.

Ce domaine a été enregistré le 1er août 2026. Il a deux jours, ne se classe sur rien et n’a aucun lien entrant digne d’être compté — et sept systèmes d’IA distincts l’ont déjà lu.

Systèmes d’IA

Requêtes par agent utilisateur d’IA, 23,9 heures jusqu’au 2026-08-03 18:04 UTC
SystèmeRequêtes DonnéesPart
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, pour le compte d’un humain)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Moteurs de recherche, pour comparaison

Même fenêtre, robots classiques
RobotRequêtes DonnéesPart
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot, avec 130 requêtes, est à portée de Googlebot, et ce n’est pas le rapport que voient la plupart des sites. Bingbot, avec 10, est la surprise dans l’autre sens.

Vingt requêtes ne lisaient pas

Mises à part plutôt que comptées : des requêtes portant un agent utilisateur d’IA et allant droit vers des fichiers que personne ne lie.

Scans d’identifiants arrivés sous des agents utilisateurs d’IA ou de robots
Agent utilisateur déclaréRequêtes Chemins recherchés
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, pour le compte d’un humain)5/keys.json

Un agent utilisateur n’est pas une identité. C’est une chaîne que le client choisit. Savoir si celles-ci venaient des opérateurs nommés ou de quelqu’un qui emprunte leur nom ne peut pas être déterminé depuis ce côté-ci, et cette page ne prétend pas le savoir. Ce qu’on peut dire : elles ne lisaient pas, toutes ont reçu 404, et rien de ce qu’elles cherchaient n’existe ici.

Ce que les systèmes d’IA ont réellement récupéré

Les chemins les plus demandés, tous confondus, sont ceux lisibles par machine — /sitemap.xml, /llms.txt, /mcp — avant les articles. C’est tout l’argument pour entretenir ces fichiers : ce ne sont pas des ornements, ce sont eux qu’on lit en premier.

Ce que ceci n’est pas

Ces chiffres ne sont pas vérifiables de l’extérieur. Ils viennent de l’analytique de notre propre fournisseur, récupérée avec un jeton que seul l’exploitant détient. Personne ne peut les recalculer. Ce qui est publié à la place, c’est la méthode : la requête se trouve dans tools/crawler-bericht.py, en clair, et quiconque dispose d’une zone Cloudflare peut lancer la même contre la sienne.

Ceci est donc une auto-déclaration à méthode divulguée, pas une mesure que quelqu’un pourrait répéter. Partout ailleurs sur ce site cette distinction est le sujet même, et il serait malhonnête de la brouiller ici parce que les chiffres se trouvent être flatteurs.

Deux limites de plus. Un jour est un jour — un robot qui passe chaque semaine y est invisible. Et les requêtes sont comptées en bordure de réseau : un système qui lit une copie en cache via un intermédiaire n’apparaît pas du tout.

Méthode : Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, groupé par agent utilisateur sur une fenêtre de 24 heures, relevé le 2026-08-03. Les requêtes vers des chemins que seul un scanner cherche sont mises à part et ne comptent pas comme lecture. La requête est publiée ; le journal sous-jacent n’est accessible à personne d’autre qu’à l’exploitant, ces chiffres sont donc une auto-déclaration et sont signalés comme tels. Un agent utilisateur peut être fixé librement et ne prouve pas l’origine. Rien de ceci n’est un conseil juridique.

Les corrections sont bienvenues et sont faites en public : ouvrir un issue.

← Proving Lab · Avertissement

Chi legge davvero questo sito: 24 ore di registri dei crawler, pubblicate

Sui crawler di IA si afferma molto e si mostra pochissimo. Chi gestisce un sito può vedere nei propri registri quali sistemi arrivano davvero, che cosa prendono e con quale frequenza — e quasi nessuno lo pubblica. Eccone una giornata, con la query che l’ha prodotta.

Rilevato il 2026-08-03 · finestra di 23,9 h · dati grezzi

Il dato principale

I sistemi di IA hanno fatto più richieste dei motori di ricerca — 299 contro 282. I motori di ricerca hanno tirato più dati (4,0 MB contro 2,1 MB), perché Googlebot preleva immagini e fogli di stile che a un modello linguistico non servono.

Questo dominio è stato registrato il 1° agosto 2026. Ha due giorni, non si posiziona per nulla e non ha link in entrata degni di essere contati — e sette sistemi di IA distinti lo hanno già letto.

Sistemi di IA

Richieste per user agent di IA, 23,9 ore fino al 2026-08-03 18:04 UTC
SistemaRichieste DatiQuota
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, per conto di una persona)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Motori di ricerca, per confronto

Stessa finestra, crawler convenzionali
CrawlerRichieste DatiQuota
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot, con 130 richieste, è a portata di Googlebot, e non è il rapporto che vede la maggior parte dei siti. Bingbot, con 10, è la sorpresa nella direzione opposta.

Venti richieste non stavano leggendo

Messe da parte anziché contate: richieste che portavano uno user agent di IA e puntavano dritto a file che nessuno collega.

Scansioni di credenziali giunte sotto user agent di IA o di crawler
User agent dichiaratoRichieste Percorsi cercati
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, per conto di una persona)5/keys.json

Uno user agent non è un’identità. È una stringa che sceglie il client. Se queste siano venute dagli operatori nominati o da qualcuno che ne prende in prestito il nome non è determinabile da questo lato, e questa pagina non pretende di saperlo. Quello che si può dire: non stavano leggendo, tutte hanno ricevuto 404 e nulla di ciò che cercavano esiste qui.

Che cosa hanno davvero prelevato i sistemi di IA

I percorsi più richiesti nel complesso sono quelli leggibili dalla macchina — /sitemap.xml, /llms.txt, /mcp — prima degli articoli. È tutto qui l’argomento per mantenere quei file: non sono decorazione, sono ciò che viene letto per primo.

Che cosa questo non è

Questi numeri non sono verificabili dall’esterno. Vengono dall’analitica del nostro stesso fornitore, recuperata con un token che ha solo il gestore. Nessuno può ricalcolarli. Ciò che si pubblica al loro posto è il metodo: la query sta in tools/crawler-bericht.py, in chiaro, e chiunque abbia una zona Cloudflare può eseguire la stessa sulla propria.

Questa è quindi un’autodichiarazione con metodo dichiarato, non una misura che qualcuno potrebbe ripetere. In ogni altro punto di questo sito quella distinzione è il punto, e sarebbe disonesto sfumarla qui perché i numeri risultano lusinghieri.

Altri due limiti. Un giorno è un giorno — un crawler che passa ogni settimana lì è invisibile. E le richieste si contano al bordo della rete, quindi un sistema che legge una copia in cache tramite un intermediario non compare affatto.

Metodo: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, raggruppato per user agent su una finestra di 24 ore, rilevato il 2026-08-03. Le richieste verso percorsi che solo uno scanner cerca sono messe da parte e non contate come lettura. La query è pubblicata; il registro sottostante non è accessibile a nessuno tranne al gestore, quindi queste cifre sono un’autodichiarazione e come tali sono etichettate. Uno user agent può essere impostato liberamente e non è prova di origine. Nulla di qui è consulenza legale.

Le correzioni sono benvenute e vengono fatte in pubblico: aprire un issue.

← Proving Lab · Avvertenza

このサイトを実際に読んでいるのは誰か——クローラー記録 24 時間分の公開

AI クローラーについては多くが語られ、ほとんど示されていない。サイトを運用していれば、どのシステムが実際に来て、何を持ち去り、どれくらいの頻度で来るのかを、自分の記録で見ることができる。それを公開する者はほとんどいない。ここにその一日分を、生成に使ったクエリとともに置く。

取得日 2026-08-03 · 対象期間 23.9 時間 · 生データ

要点

AI システムのほうが検索エンジンより多くリクエストした。299 対 282 である。データ量では検索エンジンのほうが多く(4.0 MB 対 2.1 MB)、Googlebot が画像やスタイルシートまで取得するからで、それらは言語モデルには用がない。

このドメインは 2026 年 8 月 1 日に登録された。二日目であり、何の検索順位もなく、数えるに値する被リンクもない。それでも七つの異なる AI システムがすでに読んでいる。

AI システム

AI ユーザーエージェント別のリクエスト数、2026-08-03 18:04 UTC までの 23.9 時間
システムリクエスト データ量割合
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI、人間の依頼による)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

比較としての検索エンジン

同じ期間、従来型のクローラー
クローラーリクエスト データ量割合
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot は 130 リクエストで Googlebot に手が届く位置にあり、これはたいていのサイトが見る比率ではない。Bingbot の 10 は、逆方向の驚きである。

20 件のリクエストは読んでいなかった

数えずに切り分けたもの——AI のユーザーエージェントを名乗りながら、誰もリンクしていないファイルへ一直線に向かったリクエストである。

AI またはクローラーのユーザーエージェントで届いた認証情報スキャン
名乗ったユーザーエージェントリクエスト 求めたパス
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI、人間の依頼による)5/keys.json

ユーザーエージェントは身元ではない。クライアントが選ぶ文字列にすぎない。これらが名指しされた運営者から来たのか、その名を借りた誰かから来たのかは、こちら側からは判定できず、本ページはそれを知っているとは主張しない。言えるのは、それらは読んでいなかった、すべて 404 を受け取った、求めていたものはここに何一つ存在しない、ということである。

AI システムが実際に取得したもの

すべてを通じて最も多く求められたパスは、機械可読なもの——/sitemap.xml/llms.txt/mcp——であり、記事より先である。これらのファイルを整備する理由はそれに尽きる。飾りではなく、最初に読まれるものだ。

これが何ではないか

これらの数字は外部から検証できない。自分たちの事業者の解析から得たもので、運営者しか持たないトークンで取得した。誰も再計算できない。代わりに公開するのは方法である。クエリは tools/crawler-bericht.pyに平文で置いてあり、Cloudflare のゾーンを持つ者は誰でも、同じものを自分のゾーンに対して実行できる。

したがってこれは方法を開示した自己申告であって、誰かが再現できる測定ではない。本サイトの他のどこでもこの区別こそが要点であり、数字がたまたま好都合だからといってここでぼかすのは不誠実だろう。

さらに二つの限界がある。一日は一日でしかない——週に一度来るクローラーはそこには映らない。そしてリクエストは配信の縁で数えるので、仲介を経てキャッシュされた複製を読む系統はまったく現れない。

方法:Cloudflare GraphQL Analytics の httpRequestsAdaptiveGroups を、24 時間の窓でユーザーエージェント別に集計。取得日は 2026-08-03。スキャナーしか探さないパスへのリクエストは切り分け、読んだものとしては数えていない。クエリは公開しているが、元の記録は運営者以外に閲覧できないため、これらの数値は自己申告であり、そう明示している。ユーザーエージェントは自由に設定でき、出所の証明にはならない。ここに書かれたことは法的助言ではない。

訂正は歓迎し、公開の場で行う: issue を立てる.

← Proving Lab · 免責事項

Quem realmente lê este site: 24 horas de registros de rastreadores, publicadas

Sobre rastreadores de IA afirma-se muito e mostra-se muito pouco. Quem mantém um site pode ver nos próprios registros quais sistemas de fato chegam, o que levam e com que frequência — e quase ninguém publica isso. Aqui está um dia disso, com a consulta que o produziu.

Coletado em 2026-08-03 · janela de 23,9 h · dados brutos

O achado principal

Os sistemas de IA fizeram mais requisições do que os buscadores — 299 contra 282. Os buscadores puxaram mais dados (4,0 MB contra 2,1 MB), porque o Googlebot busca imagens e folhas de estilo que não servem para um modelo de linguagem.

Este domínio foi registrado em 1º de agosto de 2026. Tem dois dias, não posiciona para nada e não tem links de entrada dignos de contagem — e sete sistemas de IA distintos já o leram.

Sistemas de IA

Requisições por agente de usuário de IA, 23,9 horas até 2026-08-03 18:04 UTC
SistemaRequisições DadosParticipação
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, a pedido de uma pessoa)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Buscadores, para comparação

Mesma janela, rastreadores convencionais
RastreadorRequisições DadosParticipação
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

O YandexBot, com 130 requisições, está ao alcance do Googlebot, e essa não é a proporção que a maioria dos sites vê. O Bingbot, com 10, é a surpresa na direção oposta.

Vinte requisições não estavam lendo

Separadas em vez de contadas: requisições que traziam um agente de usuário de IA e foram direto a arquivos que ninguém liga.

Varreduras de credenciais chegadas sob agentes de usuário de IA ou de rastreadores
Agente de usuário declaradoRequisições Caminhos procurados
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, a pedido de uma pessoa)5/keys.json

Um agente de usuário não é uma identidade. É uma cadeia de caracteres que o cliente escolhe. Se estas vieram dos operadores nomeados ou de alguém que toma emprestado o nome deles não se pode determinar deste lado, e esta página não afirma saber. O que se pode dizer: não estavam lendo, todas receberam 404 e nada do que procuravam existe aqui.

O que os sistemas de IA de fato buscaram

Os caminhos mais requisitados no conjunto são os legíveis por máquina — /sitemap.xml, /llms.txt, /mcp — antes dos artigos. É esse todo o argumento para manter esses arquivos: não são enfeite, são o que se lê primeiro.

O que isto não é

Estes números não podem ser verificados de fora. Vêm da analítica do nosso próprio provedor, obtida com um token que só o operador tem. Ninguém pode recalculá-los. O que se publica em vez disso é o método: a consulta está em tools/crawler-bericht.py, em texto claro, e quem tiver uma zona Cloudflare pode rodar a mesma contra a sua.

Portanto isto é um autorrelato com método divulgado, não uma medição que alguém pudesse repetir. Em todo o resto deste site essa distinção é justamente a questão, e seria desonesto borrá-la aqui porque os números por acaso são lisonjeiros.

Mais dois limites. Um dia é um dia — um rastreador que passa semanalmente é invisível nele. E as requisições são contadas na borda, de modo que um sistema que lê uma cópia em cache por meio de um intermediário não aparece de jeito nenhum.

Método: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, agrupado por agente de usuário sobre uma janela de 24 horas, coletado em 2026-08-03. Requisições a caminhos que só um scanner procura são separadas e não contadas como leitura. A consulta está publicada; o registro subjacente não é acessível a ninguém além do operador, portanto estes números são um autorrelato e estão rotulados como tal. Um agente de usuário pode ser definido livremente e não é prova de origem. Nada aqui é aconselhamento jurídico.

Correções são bem-vindas e são feitas em público: abrir um issue.

← Proving Lab · Aviso legal

Кто на самом деле читает этот сайт: 24 часа журналов краулеров, опубликованные

Об ИИ-краулерах утверждают много, а показывают крайне мало. Тот, кто держит сайт, может увидеть в собственных журналах, какие системы приходят на самом деле, что забирают и как часто, — и почти никто этого не публикует. Вот один день, вместе с запросом, который его получил.

Собрано 2026-08-03 · окно 23,9 ч · исходные данные

Главное

ИИ-системы сделали больше запросов, чем поисковые машины — 299 против 282. Поисковые машины вытянули больше данных (4,0 MB против 2,1 MB), потому что Googlebot забирает изображения и таблицы стилей, которые языковой модели ни к чему.

Этот домен зарегистрирован 1 августа 2026 года. Ему два дня, он ни по чему не ранжируется и не имеет входящих ссылок, достойных подсчёта, — и семь разных ИИ-систем уже его прочитали.

ИИ-системы

Запросы по ИИ-агентам пользователя, 23,9 часа до 2026-08-03 18:04 UTC
СистемаЗапросы ДанныеДоля
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI, по поручению человека)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

Поисковые машины для сравнения

То же окно, обычные краулеры
КраулерЗапросы ДанныеДоля
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot со 130 запросами находится в пределах досягаемости Googlebot, и это не то соотношение, которое видит большинство сайтов. Bingbot с 10 — неожиданность в другую сторону.

Двадцать запросов ничего не читали

Выделено отдельно, а не засчитано: запросы с ИИ-агентом пользователя, шедшие прямо к файлам, на которые никто не ссылается.

Сканирование учётных данных под ИИ- или краулерными агентами пользователя
Заявленный агент пользователяЗапросы Искомые пути
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI, по поручению человека)5/keys.json

Агент пользователя — не удостоверение личности. Это строка, которую выбирает клиент. Пришли ли эти запросы от названных операторов или от того, кто позаимствовал их имя, с этой стороны определить нельзя, и эта страница не утверждает, что знает. Сказать можно вот что: они не читали, все получили 404, и ничего из того, что они искали, здесь нет.

Что ИИ-системы на самом деле забрали

Самые запрашиваемые пути по всем вместе — машиночитаемые: /sitemap.xml, /llms.txt, /mcp — раньше статей. В этом и весь довод за то, чтобы эти файлы вести: они не украшение, они то, что читают первым.

Чем это не является

Эти числа нельзя проверить снаружи. Они получены из аналитики нашего собственного провайдера, извлечены токеном, который есть только у оператора. Пересчитать их никто не может. Вместо этого публикуется метод: запрос лежит в tools/crawler-bericht.py, открытым текстом, и всякий, у кого есть зона Cloudflare, может выполнить такой же для своей.

Итак, это самоотчёт с раскрытым методом, а не измерение, которое кто-то мог бы повторить. Везде на этом сайте именно это различие и составляет суть, и было бы нечестно размывать его здесь только потому, что числа случайно лестны.

Ещё два ограничения. День есть день — краулер, приходящий раз в неделю, в нём невидим. И запросы считаются на границе сети, поэтому система, читающая кэшированную копию через посредника, не появляется вовсе.

Метод: Cloudflare GraphQL Analytics, httpRequestsAdaptiveGroups, сгруппировано по агенту пользователя за 24-часовое окно, собрано 2026-08-03. Запросы к путям, которые ищет только сканер, выделены отдельно и не засчитаны как чтение. Запрос опубликован; лежащий в основе журнал недоступен никому, кроме оператора, поэтому эти цифры — самоотчёт, и они так и обозначены. Агент пользователя можно задать свободно, и он не доказывает происхождение. Ничто здесь не является юридической консультацией.

Исправления приветствуются и вносятся публично: открыть issue.

← Proving Lab · Отказ от ответственности

究竟谁在读这个站点:24 小时爬虫日志,公开

关于 AI 爬虫,说的多,拿出来看的极少。任何运营站点的人都能在自己的日志里看到哪些系统真的来过、取走了什么、来得多频繁——却几乎没有人把它公开。这里是其中一天,连同产生它的那条查询。

采集于 2026-08-03 · 时间窗 23.9 h · 原始数据

要点

AI 系统发出的请求比搜索引擎更多——299 对 282。搜索引擎拉走的数据更多(4.0 MB 对 2.1 MB),因为 Googlebot 会取图片和样式表,而语言模型用不上这些。

这个域名注册于 2026 年 8 月 1 日。它只有两天大,没有任何排名,也没有值得计入的外部链接——而已经有七个不同的 AI 系统读过它。

AI 系统

按 AI 用户代理统计的请求数,截至 2026-08-03 18:04 UTC 的 23.9 小时
系统请求数 数据量占比
ClaudeBot (Anthropic)135795 kB45 %
GPTBot (OpenAI)96798 kB32 %
PerplexityBot22213 kB7 %
CCBot (Common Crawl)1669 kB5 %
ChatGPT-User (OpenAI,代表人类发起)15152 kB5 %
OAI-SearchBot (OpenAI)943 kB3 %
Amazonbot645 kB2 %

搜索引擎,用作对照

同一时间窗,传统爬虫
爬虫请求数 数据量占比
Googlebot1342658 kB48 %
YandexBot1301306 kB46 %
Bingbot1078 kB4 %
Applebot853 kB3 %

YandexBot 以 130 次请求已在 Googlebot 的射程之内,这不是大多数站点看到的比例。Bingbot 的 10 次,则是反方向上的意外。

有二十次请求不是在读

单列出来而不计入:带着 AI 用户代理、却径直奔向无人链接的文件的请求。

以 AI 或爬虫用户代理到达的凭据扫描
声称的用户代理请求数 所求路径
Bingbot5/.env.prod
PerplexityBot5/terraform.tfvars
OAI-SearchBot (OpenAI)5/.git/HEAD
ChatGPT-User (OpenAI,代表人类发起)5/keys.json

用户代理不是身份。它只是客户端自己选定的一个字符串。这些请求究竟来自被点名的运营方,还是来自借用其名号的人,从这一侧无法判定,本页也不声称知道。可以说的是:它们不是在读,全部收到 404,它们所寻找的东西这里一样都没有。

AI 系统实际取走了什么

在所有请求中被要得最多的路径是机器可读的那几个——/sitemap.xml/llms.txt/mcp——排在文章之前。维护这些文件的全部理由就在这里:它们不是装饰,它们是最先被读到的东西。

这不是什么

这些数字无法从外部核验。它们出自我们自己服务商的分析数据,用只有运营者持有的令牌取得。没有人能重算。作为替代,公开的是方法:这条查询就在 tools/crawler-bericht.py里,以明文写出;任何拥有 Cloudflare 区域的人,都可以对自己的区域运行同一条。

所以这是一份公开了方法的自述,而不是一次别人能够重复的测量。在本站其他任何地方,这个区分正是要点所在;仅仅因为数字碰巧好看就在这里把它抹糊,那是不诚实的。

还有两条界限。一天就只是一天——每周才来一次的爬虫在其中是看不见的。而且请求是在网络边缘计数的,因此经由中间层读取缓存副本的系统根本不会出现。