Don't trust your own logs
Most SEOs that have been around for even a few years know that logs are a total mess, and if you want the truth you have to dig. Every AI visibility tool on the market will tell you how often GPTBot, ClaudeBot and PerplexityBot crawled your site. Almost every one of them gets that number by reading the user agent string.
A user agent is a text field. Anybody can type anything into it. Equip yourself with a standard linux command line and you could go to the CIA website and fill their logs with fake user agents.
I have known that forever and I still built a whole case study on top of it last month. I just wanted to hard-test this in 2026.
How to actually "count" real AI bot traffic
There are three levels of rigor here and I went through all of them in one afternoon, which is how I know the middle one is a trap.
Count the user agent. This is what the tools do. It is wrong and everybody who thinks about it for ten seconds knows it is wrong.
Count Cloudflare's verified bot flag. This is what I did next and I was pleased with myself. Cloudflare checks the source network for you and stamps requests as verified. Then I ran the numbers on the car dealer (it sits behind Cloudflare) and PerplexityBot came back 100% unverified. Over two thousand requests and none verified.
Cloudflare just was not verifying those crawlers for me, so every request from the real ones got stamped unverified too. When I looked at where that traffic actually came from, most of it was Amazon Web Services, which is exactly where Perplexity genuinely lives, and it matched Perplexity's own published list. The flag was not telling me the traffic was fake. It was telling me Cloudflare had no opinion.
Check the source against the published IP Ranges. This is the only actual way to do this. OpenAI (gptbot.json), Perplexity (perplexitybot.json), Apple (applebot.json), Anthropic and the rest publish the exact IP ranges their crawlers use, as JSON files, specifically so you can do this. For Google and Bing you use forward confirmed reverse DNS (Apple supports that too, so you get two ways to check it), look up the IP, confirm the hostname belongs to them, then resolve that hostname back and check it matches. A spoofer fails that every time because they do not control the DNS.
No tool does this for you. It took a couple hundred lines of Python.
The actual numbers
I tested 3 websites over the same two weeks, September 16 to 30, 2026. A logistics company, a busy car dealer, and this site. Clients anonymized.
| Site | Claimed AI bot hits | Fake | % fake |
|---|---|---|---|
| Car dealer | 71,032 | 6,118 | 8.6% |
| Logistics company | 8,925 | 2,489 | 27.9% |
| This site | 1,865 | 1,362 | 73.0% |
Our own site at 73% is almost certainly an edge case so maybe don't read into it. We are an AI company, or at least close enough to that vertical, so we get probed by competitors and by people poking at whatever we just published. It's an outlier. Don't use it as a benchmark for your own site.
Per bot, across all three sites, the honest ones look like this: meta-externalagent 2.8% fake, Applebot 5.7%, Amazonbot 10.9%, ClaudeBot 13.8%, GPTBot 18.7%, OAI-SearchBot 20.0%. If you are a decent sized site and only tracking the big names, your numbers are in the right neighborhood, but already off by more than a schmidge.
Then it falls apart. ChatGPT-User 42% fake. PerplexityBot 50.5%. DuckAssistBot 56%. Half of what your tool calls Perplexity is not Perplexity. And three came back at 100%.
Google-Extended is 100% fake and always will be. It is a robots.txt token, not a crawler. Google never sends it as a user agent. Every single request carrying it is a lie, on every site, forever. If your tool reports Google-Extended visits, throw the tool away.
CCBot and MistralAI-User came back 100% fake on these sites. Most of it came from Google Cloud, and neither of them crawls from Google Cloud. The rest of the CCBot hits came from Amazon, which is where Common Crawl actually lives, so I checked those by hand against Common Crawl's own published list. Not one of them was on it, and Cloudflare did not verify any of them either. Those were not Common Crawl.
Bytespider I am not going to give you a number for, and I pulled it out of the table entirely. Its traffic came from Amazon, China Telecom, China Unicom and Zenlayer all at once, and ByteDance is a Chinese company, so some of that is plausibly the real crawler and my network check is too narrow to tell them apart. It scored 100% fake, which I do not believe, and Cloudflare backs me up on that one, it verified 602 Bytespider requests on the car dealer that my check called fake. Leaving it in would have inflated every number on this page. I would rather give you no number than a wrong one.
The bots don't care how big you are
I really thought the bigger, busier site would attract more spoofing. More pages, more traffic, more scrapers but that is not true.
Look at the raw counts instead of the percentages. The car dealer had 64,914 verified AI crawler hits in two weeks. The logistics company had 6,436. This site had 503. That is a 130x spread in real crawling.
The fakes: 6,118, 2,489 and 1,362. About a 4.5x spread.
Real crawler traffic scales with your site. Bigger, older, more authoritative, more pages, more real crawling. Fake traffic barely scales with any of that, so on a small site it eats a much bigger share of the pie. The car dealer is 8.6% fake. The logistics company is 27.9%. Same farm, same two weeks.
I am not going to tell you it is a fixed number of hits per site, because it plainly is not - the car dealer still took more than twice as many fake hits as anyone else. What I will say is that the smaller site pays a much higher rate for the same issue.
If you launched a site this year and your analytics is showing healthy AI crawler interest, check it before you celebrate. A quarter of it may be a botnet that found your DNS record.
What the real AI bots actually look like
Everything above is about the liars. Here is the other half, the traffic that passed verification, because I think most people have never seen this list.
| Bot | Car dealer | Logistics | This site |
|---|---|---|---|
| PetalBot | 15,500 | 286 | 63 |
| meta-externalagent | 15,494 | 0 | 22 |
| Applebot | 11,326 | 388 | 2 |
| ClaudeBot | 7,800 | 0 | 199 |
| Amazonbot | 3,740 | 2 | 37 |
| GoogleOther | 3,595 | 104 | 4 |
| GPTBot | 1,960 | 0 | 19 |
| ShapBot (Parallel) | 1,474 | 720 | 8 |
| PerplexityBot | 1,408 | 36 | 30 |
| ChatGPT-User | 1,020 | 44 | 31 |
| OAI-SearchBot | 911 | 4,825 | 85 |
| DuckAssistBot | 563 | 2 | 1 |
| Claude-User | 123 | 29 | 2 |
The thing that jumps out at me: the single biggest verified AI crawler on the car dealer is PetalBot, at 15,500 hits. That is Huawei. Nobody is writing GEO guides about ranking in Huawei. Meta is right behind it at 15,494, and Apple at 11,326, and none of those three are the ones anybody optimizes for either.
GPTBot, the one everybody actually worries about, came in seventh. And on the logistics company it was zero. OAI-SearchBot on the other hand was 4,825 there, more than everything else combined. That's ChatGPT's search index, not the training crawler.
One of these deserves a callout. ShapBot turned out to be Parallel, and they publish their IP list at docs.parallel.ai/resources/shapbot.json - ten addresses, every one a single host. Every ShapBot request on every site I checked verified clean. Not one fake. It is also the crawler that showed up in our markdown experiment as the only thing on the internet actually consuming the markdown we were serving. Small fleet, publishes its ranges, never impersonated. That is what a well behaved AI crawler looks like, and almost nobody has heard of it.
The third bucket: bots you cannot check from a log file
While I was doing this I kept hitting AI crawlers I could not place. No published IP list, no useful reverse DNS. I was ready to write them all off until I went and read their docs, and it turns out there are two very different reasons a bot lands in this bucket.
Some of them are signed, and I just could not see it. Exa, You.com and Kimi do not publish IP ranges at all. They sign every request cryptographically instead, using HTTP Message Signatures - the Web Bot Auth scheme. You fetch their public key and verify the signature, and it holds no matter what address the request came from. Exa's keys are sitting at crawler.exa.ai/.well-known/http-message-signatures-directory right now.
That is a better method than an IP list. It is also completely invisible in an access log, because nginx does not record the signature headers. So when I audit last month's traffic, a properly signed Exa request and an outright forgery look identical. You can only check it live, at the edge, as the request arrives. Cloudflare does that automatically for the bots that support it, which is the one case where its verified flag genuinely earns its keep.
And here is the thing. I have Cloudflare's own logs for two of these sites, so I could check. On the car dealer Cloudflare verified 269 of 1,162 ExaSearchBot requests. YouBot, 0 of 443. KimiBot, 0 of 342. So most of this bucket is not signed at all, and when I lined it up against the Google Cloud farm further down, the answer was obvious: 319 of those YouBot hits and 323 of the KimiBot hits came from farm addresses that were also claiming to be GPTBot, ClaudeBot and ten other crawlers the same week. The signing is real. Most of the traffic using these names is not.
| Bot | Car dealer | Logistics | This site |
|---|---|---|---|
| ExaSearchBot | 1,162 | 18 | 3 |
| YouBot | 443 | 320 | 86 |
| KimiBot | 342 | 311 | 118 |
The rest genuinely have nothing. No list, no signatures, no reverse DNS worth the name.
| Bot | Car dealer | Logistics | This site |
|---|---|---|---|
| LinkupBot | 3,528 | 21 | 0 |
| cohere-ai | 350 | 315 | 98 |
| Claude desktop app | 85 | 0 | 0 |
| MathPicDatasetCrawler | 55 | 0 | 0 |
| img2dataset | 0 | 15 | 0 |
That last group is 4,467 hits out of 89,092, about 5% of all AI bot traffic, where the honest answer is "I do not know." Even there the farm shows up: 725 of the 763 cohere-ai hits came from the same Google Cloud addresses.
img2dataset is the clearest case of why this bucket exists. It is an open source tool for building image training sets. Anyone can install it and run it. Asking whether a request is "really" img2dataset is not even a sensible question.
So there are three ways a bot can prove itself, and you need all three. A published IP list - OpenAI, Anthropic, Perplexity, Apple, Parallel. Reverse DNS - Google, Bing, Apple, Amazon, Common Crawl, Huawei. A cryptographic signature you can only check live - Exa, You.com, Kimi. Then there is the leftover, about one hit in twenty, that has no way to prove anything at all.
Counting user agents is wrong. Trusting Cloudflare's verified flag is wrong. And even doing it properly, a chunk of your AI traffic is simply unknowable after the fact.
One server / IP range running 13 different tasks (all of them lying)
The fake traffic is not scattered, it is part of one operation.
104 Google Cloud IP addresses are each pretending to be different AI crawlers. 43 of them wore sixteen different identities, and I only know it is sixteen because I went back and looked for the ones my own bot list had missed - KimiBot, cohere-ai and YouBot were all in there too. The thirteen I can name are: Amazonbot, Applebot, Bytespider, CCBot, ChatGPT-User, ClaudeBot, DuckAssistBot, GPTBot, Google-Extended, MistralAI-User, OAI-SearchBot, PerplexityBot and meta-externalagent.
It is the same kit on every site, the exact same list of names in the exact same combination, but not the same addresses. Only one IP out of 104 hit more than one of my sites. They rotate. On our own site, on completely different hosting, 6 Google Cloud addresses pretending to be Applebot accounted for 95 requests while genuine Apple accounted for 2.
Very similar footprint: renting cloud servers, cycling identities, not caring if they are id'd, and it is indiscriminate about targets. Attackers have been doing this for decades, it is just a bit different now. When I stopped reading the user agents and started reading what they actually asked for, the answer was not crawling at all. On the logistics company alone, those addresses made about 9,300 requests and 80% of them were 404s. They went after /.env, /.env.local, /.env.production, /admin/.env, /@fs/.env, /api/proc/self/environ and every GraphQL endpoint they could guess. That is not content. That is someone hunting for leaked API keys.
Which puts it in one of two buckets:
- bug bounty recon, sprayed across a huge list of domains from rented cloud
- credential harvesting, which uses the exact same toolkit with worse intent
What it is definitely not is an AI company. Nobody training a model wants your .env file.
Don't Beleive This BAD GEO Advice
Right now eveeryone is basically saying "allow gptbot/claudebot etc so you show up in AI answers." Which is kind of true, but scanners are taking advantage of this and ripping through small business websites accounting for a large portion of their traffic AND tricking you into thinking you are popular with the AI crawlers.
In short: you see Claudebot, but its actually a credential harvester or most likely some other black hat tool.
How to spot it in your own logs
You do not need a script to sanity check this. There are two ways to reveal the "truth" in a few minutes.
Requests per IP. Real crawlers come from large fleets and touch you lightly from each machine. On the logistics company, real Applebot averaged 3.2 requests per address and the fake one 7.2. With that said, this one is a hint, not proof. On the car dealer the fakes averaged 4.4 and the real Applebot 4.6, basically identical. If one address is responsible for hundreds of hits from a major crawler, sure, that is not how the major crawler works. But low numbers do not clear anybody.
Identity per IP. This is the one that actually works. Real Googlebot never sends a request as ClaudeBot. If a single address in your logs claims to be more than one company's crawler, every one of those claims is false. 101 of the 104 farm addresses claimed at least two.
The cloud provier is the real tell. None of these operators crawl from Google Cloud. If the user agent says GPTBot and the address belongs to a cloud VM provider, it is not GPTBot.
The rule I am (probably) deploying
Since none of these operators crawl from Google Cloud, the fix is narrow enough to be safe. Challenge anything from Google Cloud's network that claims to be one of them:
(ip.src.asnum in {396982 15169}) and not cf.client.bot and (
lower(http.user_agent) contains "gptbot" or
lower(http.user_agent) contains "claudebot" or
lower(http.user_agent) contains "chatgpt-user" or
lower(http.user_agent) contains "oai-searchbot" or
lower(http.user_agent) contains "perplexitybot" or
lower(http.user_agent) contains "applebot" or
lower(http.user_agent) contains "amazonbot" or
lower(http.user_agent) contains "ccbot" or
lower(http.user_agent) contains "meta-externalagent" or
lower(http.user_agent) contains "mistralai" or
lower(http.user_agent) contains "duckassistbot" or
lower(http.user_agent) contains "bytespider" or
lower(http.user_agent) contains "kimibot" or
lower(http.user_agent) contains "youbot" or
lower(http.user_agent) contains "cohere-ai"
)
A few notes on that. 396982 is Google Cloud, but some Google Cloud addresses still get announced under Google's main network (15169), so I check both. None of the operators in that list publish a single Google Cloud address. The lower() is there because contains is case sensitive and I don't want one weird capitalization to slip through. The last four are identities the same farm was wearing, so they go in too.
The not cf.client.bot part is the safety net. Kimi and You.com don't publish IP lists, so I can't promise they never crawl from Google Cloud. If Cloudflare has already verified the request (the signed ones get checked live, like I talked about above) it skips the rule.
Do NOT add ShapBot to this. Parallel's whole published list is Google Cloud addresses, so the real one would get caught.
Google-Extended gets its own rule with no network condition, because it is fake from everywhere.
Start on managed challenge rather than block and watch it for a week.