Is noindex an AI shield, or just a search instruction?
Everybody assumes that if a page is set to noindex, it stays out of the AI engines too. I was not sure that was true. noindex tells search engines one thing: do not list this page in results. But AI answer engines are not search engines, and they do not all play by the same rules.
Will an AI crawler fetch and ingest a noindex page at all? And if the page is in no index anywhere, can a link people pass around by hand - a post, a forwarded email - still drag its content into an AI answer? I asked this question espeically because of Grok, which relies so much on X/Twitter for its answers.
If the answer to both is yes, then noindex is not the shield people think it is, and sharing is its own way into the engines that has nothing to do with search.
How we built this
It is on page set to noindex with a meta-robots tag. We left it out of the sitemap, out of llms.txt, and we did not link to it from anywhere on the site. As far as the normal discovery machinery is concerned, the page does not exist.
On it sits a coined term that returns zero results anywhere before launch and lives only on this one page. We kept it out of the messages we use to share the link - those carry the bare URL and a vague teaser, nothing else. That gap is the whole point. The term shows up nowhere in the share text, so any engine that later repeats it had to fetch the page itself, not just read our posts about it.
Then we hand it out on purpose - a few channels, staggered over several days, so each share can be lined up against the crawl log. We stayed off search-owned surfaces too, so the only way search itself reaches the content is by crawling the noindex page directly.
How we are measuring it
Same two tracks as the rest of this series. Crawl confirmation comes from our edge middleware, which logs every known AI crawler, the path it asked for, and the time. Because the channels go out on different days, the log tells us more than whether the page got fetched - it tells us which share most likely set it off.
Then the regurgitation test: ask the major engines about the coined term after a crawl window. A hit means the noindex page was read and ingested anyway, despite carrying every do-not-index signal short of flat out blocking the crawler.
(note: I withheld naming the term itself until publishing this case study)
What we expected (my prediction I wrote before knowing anything)
"My bet is that at least one engine fetches the page off a shared link despite
noindex, because fetching and indexing are different pipes and the crawlers do not all treat the directive the same way. The negative case matters just as much. If nothing comes back, the bot log still tells us which of two stories it is: a fetch with no later citation means the page was read and then respected, and no fetch at all means the shares never reached a crawler. Those are very different outcomes, and we can tell them apart.
The URL and Keyword Revealed for this Case Study
The page is llmcartel.com/unlisted. The coined term on it is Klovect.
Klovect is not a word, I made it up so it could not be guessed, confabulated, or confused with something real.
The page carries a meta-robots noindex. It is not in sitemap.xml, not in llms.txt, and nothing on this site links to it. Worth being precise about one thing: robots.txt allows the crawl. The variable under test is the noindex directive, not a block. I wanted the crawlers to be able to walk right in, and then see whether the directive still kept the content out of the answers.
Bare URL / vague teaser / no mention of Klovect in any of the share text, posted to X, Facebook, and LinkedIn on June 25. That gap matters - the term appears nowhere in the shares, so any engine that repeated it had to have read the page itself rather than my post about it.
Who actually showed up
This was never a test of whether a hidden page stays hidden.
| Crawler | Verified as | Fetches | When |
|---|---|---|---|
| Applebot | Search crawler | 5 | from June 25, hours after the shares |
| Meta-ExternalAgent | AI crawler | 4 | late June |
| facebookexternalhit | Page preview | 6 | Aug 31 to Sep 15 |
| PetalBot | AI crawler | 5 | Sep 1 to Sep 26 |
| Twitterbot | Page preview | 4 | Sep 2 to Sep 26 |
| Bingbot | Search crawler | 2 | Aug 31 and Sep 2 |
Applebot got there first, within hours of the posts going up, from two different datacenters ten seconds apart.
Bingbot fetched a page carrying a noindex tag. It means "noindex" never stopped anything from reading our content. And PetalBot is categorized by Cloudflare as an AI crawler ( not a search crawler) and it was still pulling this page in late September, three months after the only links to it were a handful of social posts nobody engaged with.
Twitter and Facebook were still re-fetching this URL in late September off posts from June. A link you share does not get fetched once. It stays on a list somewhere and gets pulled again, and again, long after you have forgotten you posted it.
What the LLMs / AIs said about "klovect"
Literally zero, they did not find the term/page. I will say that when signed in, Grok "knew" about the Klovect reference but signed out, private or on a fresh account it had total amnesia.
- Copilot, July 10. "Klovect does not appear... non-existent, internal, or emerging but not indexed." It named the reason itself.
- Kimi, July 20. "I don't recognize 'klovect'... a quick search doesn't turn up any obvious results." Same session where it cited our other test page by name, which I will come back to.
- Grok, July 31. The best null in the set. I asked about Klovect. Grok shelled out to
curlto go read the CVI page, and for Klovect it ran a dedicated web search, got klove.com and an Instagram account, and gave up. It had a raw fetching tool in its hand and it still could not find the page, because nothing in any index pointed at it. - Plain search, September 28. "Klovect" returns a 1990s arcade game database, a neighbourhood in Kyiv, and a couple of Czech villages. Not us.
- Gemini did something weird. Asked about the tracking token from the page, it did not say it had never heard of it - it invented an entire backstory about our Nginx configuration, a "late June 2026 server hardening push," and bot mitigation rules, and offered to pull up the server block directives. None of that exists. It had nothing, and rather than say so it wrote fiction and offered to go deeper.
That is the same failure we caught in Vol. 03, where Copilot invented a definition for our fake metric and credited three of our competitors. When these things have no grounding, LLMs will just make shit up.
Where this falls short
I would totally redo this (and I am going to) I didnt notice these issues until after the fact.
- I have a pathetic social media following so nobody engaged with the shares. The posts got crawled, not clicked. A URL that actually goes around - real reshares, a forum thread, somebody quoting it - is a different pressure test, and it is the one I would want to run next.
What's next
If you have been treating noindex as a way to keep something away from the AI engines, it held here, for three months, against four verified crawlers and a URL I personally handed to three social networks. That is better than I expected going in. I still would not trust it.
Every one of those crawlers fetched the page and read every word on it. What the tag bought was exclusion from the index, and the index is what the engines reach for when they answer. If your content genuinely cannot be seen, it needs to be behind auth, not behind a polite request.
Read this next to our last blog post and the pair says one thing. Anything in your markup gets swallowed whole, including the parts nobody can see. And the only thing that reliably kept a page out of the answers was telling the index no.
Klovect is burned now, same as Lattice Saturation. It is in visible text on a page the crawlers read constantly, so any engine that repeats it from here proves nothing. Raw logs and query transcripts are available if you want them - ask.