Do LLM Bots Actually Use Cloudflare's Markdown for Agents?

Experiment closed - Day 25

TL;DR - No bots that are important for SEO/GEO (as of now) request Cldouflare's Markdown for Agents. I ran this for 25 days on a high-traffic car dealer client website. The feature itself works fine - send Accept: text/markdown and you get a clean markdown page, roughly 95% smaller than the HTML. Across 2.78 million requests, the named AI crawlers (GPTBot, ClaudeBot, ChatGPT-User, etc) made 121,824 requests and took markdown zero times. Real Googlebot: 87k requests, zero. The markdown that did get served went to one crawler I had never heard of (but looks like it is scraping the entire web), and to a scraper wearing a fake Googlebot costume.

The clock ran July 8 to August 1, 2026 - 25 days using Cloudflare Log Explorer, raw edge logs I query directly with SQL, which is a lot cleaner. Every number below is from that, unsampled, with my own test traffic stripped out by IP/ASN and network.

Final tally: 25 days, 2.78 million requests

Here is the whole experiment in one table. Real production traffic on the car dealer, straight off Log Explorer. "Markdown pulls" means the edge actually returned text/markdown with a 200, not just that the bot showed up.

Crawler activity on the car dealer, Log Explorer, July 8 to August 1, 2026
CrawlerRequestsMarkdown pulls
ShapBot / Shap-User4,0171,274
"Googlebot" (spoofed, 33 networks)4,1051,493
Googlebot (real, AS15169)87,7890
Bingbot54,7890
Amazonbot35,9650
PetalBot25,4960
Meta-ExternalAgent19,8130
Bytespider10,2990
ChatGPT-User8,7350
ClaudeBot7,2530
GPTBot5,5510
Applebot3,7650
OAI-SearchBot1,9980
CCBot9680
PerplexityBot9150
GoogleOther7470
YouBot6590
DuckDuckBot2680
Claude-User966
AgentRadar-Research562
AgentReadyBot491

2,779 markdown responses went out the door in 25 days. Here is where they went. ShapBot / Shap-User took 1,274 of them. It runs out of Google Cloud, Cloudflare verifies it as an AI Search crawler, and I had never heard of it before this experiment. It uses the same split OpenAI does, a -Bot that crawls and a -User that fetches on demand. It is the only genuine, sustained markdown consumer I found, and it pulled markdown on all 25 days of the run without missing one.

The other 1,493 went to something claiming to be Googlebot that is definitely not Googlebot. Real Googlebot lives on AS15169 and nowhere else. These came from 33 different hosting and proxy networks, across 1,493 distinct IP addresses - one request per IP, never repeating. That is a rotating proxy scraper wearing a costume. So the single biggest consumer of Cloudflare's Markdown for Agents on this site, by volume, was a scraper impersonating Google. If you count adoption by user-agent string you would have written that up as "Googlebot loves markdown" and been completely wrong. Count by network.

That leaves 12 markdown pulls that were neither. Six from Claude-User, which is a real person asking Claude to go look at a page - the closest thing to a live human-in-the-loop win in the whole dataset. Two from AgentRadar-Research out of the BCG Henderson Institute. One from AgentReadyBot. Three from assorted browser user-agents. Twelve. Out of 2.78 million requests.

And the number that actually settles it: the named AI crawlers made 121,824 requests and asked for markdown zero times. Not "rarely." Zero. GPTBot, ClaudeBot, ChatGPT-User, OAI-SearchBot, PerplexityBot, Bytespider, Amazonbot, Meta-ExternalAgent, Applebot, CCBot, Google-Extended, Mistral, Cohere, Diffbot. Every one of them crawled this site, some of them tens of thousands of times, and every one of them took the HTML.

What I set out to do

I enabled this a few weeks ago and started to log. I checked in on it this week (glad I did) and saw almost zero markdown hits except malware. I wanted to know whether any AI agents were actually requesting the markdown representation of the page, who they were, and how often - basically, is this worth keeping and optimizing, or is it a feature sitting there that nobody uses? I gave it 25 days to answer that.

How I checked it even works

Instead of testing an existing source, I triangulated:

  • Edge tests (curl): confirmed Accept: text/markdown returns markdown.
  • Origin logs (SSH, nginx): measured AI-bot crawl volume by user-agent.
  • Origin direct test: proved the origin only ever serves HTML, so the markdown conversion is 100% Cloudflare's edge, not the server.
  • Cloudflare GraphQL Analytics: pulled the real edge content-type split and per-bot behavior.
  • AI Crawl Control dashboard: cross-checked the feature status and fulfillment.
  • I did not touch the server config - my network admin would kill me.

The feature works - and it is tiny

A GET with Accept: text/markdown returns content-type: text/markdown with a clean markdown body (YAML front-matter plus content), roughly 95% smaller than the HTML. AI Crawl Control reported "96% of markdown requests fulfilled."

What a markdown grab actually looks like

I wanted to see if what Cloudflare was outputting actually looked worthy of the page. I looked at a VDP (a vehicle page) because that is arguably the most important type of page on a car dealer. The highlights of what it grabbed:

  • Title and description with the VIN baked into the description
  • Basic info tables (Body: SUV, Mileage: 12,234, 21 hwy / 15 city MPG, Black/Black, 3.0L V6, ZF 8-speed)
  • Full DESCRIPTION, DETAILS (the entire feature list - nav, harman/kardon, etc.), TERMS & CONDITIONS
  • Price: $24,888.00, financing CTAs, VIN XYZ / STOCK MXYZ
  • "Other vehicles you may like," full dealer NAP and hours
  • It grabbed all the JSON-LD, which contains all the images of the car
  • All of Yoast's output

Size: 23 KB markdown vs 369 KB HTML - about 16x smaller.

Does this actually matter for SEO?

When I started this I said: maybe, and probably. If LLMs can spend 96% less to get the exact same data, why wouldn't they? Twenty-five days later I have my answer, and it is no. Not today. The economics argument is sound and it did not matter even a little, because the crawlers are not making that decision on a per-request basis. They have a fetch pipeline that was written before any of this existed and it asks for HTML. Nobody at OpenAI or Anthropic or Perplexity is going to rewrite that because your edge would like to hand them something smaller.

I was wrong about the timeline. I figured a few of the big ones would be experimenting with it by now. Not one is.

Things I fixed and learned the hard way

  • Vary: Accept was missing on HTML responses. Without it, caches can serve cached HTML to a markdown-negotiating client. Fixed with a Cloudflare Response Header Transform Rule (set static Vary = Accept-Encoding, Accept). This is what broke basically everything.
  • Edge caching and markdown coexistence. Added a Cache Rule (HTML GET pages, 4h Edge TTL) that excludes Accept: text/markdown and bots, so browsers get cached HTML (HIT) while agents still get freshly-converted markdown on the same URLs. I tested this several different ways.
  • The testing gotcha (this cost me a full hour): curl -I (HEAD) returns text/html even when the feature works, because HEAD has no body to convert. Always test with GET. A false "it broke" panic came entirely from using -I.

How I am logging it (and why I reset)

Version one was a homegrown Cloudflare Worker (a transparent passthrough) that logged every Accept: text/markdown GET to Workers Analytics Engine and fired a 30-minute Discord digest when there were hits. Here is what that looked like:

Discord digest reading: Markdown for Agents, last 30 min, 9 hits - other times 7, gptbot times 1, claudebot times 1, paths slash, slash all-vehicles, slash contact
The old approach: a 30-minute Discord digest from a homegrown Worker. It worked, but it was noisy and easy to muddle.

It did the job, but it was a moving part I had to babysit, and I let the logging get tangled enough that I stopped trusting the numbers. So I tore it out and switched to Cloudflare Log Explorer instead: raw, unsampled edge requests I query directly with SQL, including edgeresponsecontenttype, which flags a genuine markdown response so I can separate content-negotiated markdown from direct .md fetches in a single query. No Worker to maintain, no digest to babysit, and the data is straight off the edge rather than a sampled dashboard. It is a lot cleaner, which is why I reset the experiment to Day 1 on July 8 and ran the whole 25 days on Log Explorer. Every number in this post came out of one SQL query against that.

Why I am calling it a failure

I am closing this one out as a negative result, and I would rather publish that than let it sit as a "live experiment" quietly going nowhere for another two months.

To be clear about what failed. The feature is not broken. Cloudflare built it correctly, it converts cleanly, it fulfilled 96% of the requests it got, and the markdown it produces off a vehicle page is genuinely good - better structured than the HTML it came from. Every technical part of this worked. What failed is the premise: that if you serve markdown, agents will come take it. They did not. Twenty-five days, 2.78 million requests, and the entire named AI crawler population took it zero times.

The honest version of the finding is that markdown negotiation is not a distribution channel right now. It is a format that is ready before the readers are. If a vendor tells you turning this on will get you into AI answers, the log file says otherwise, and I have the log file.

Would I leave it on? Yeah. It costs nothing on Pro, it did not break caching once I got Vary right, and the day GPTBot flips it will already be there. But "leave it on and forget it" is a very different recommendation than what I was hoping to write, and I am not going to dress up 0 out of 121,824 as an early signal.

Two things I would actually spend time on instead, both of which came out of this run. First, if you are measuring bot behavior at all, verify by network, not user-agent - 54% of the markdown on this site went to a fake Googlebot and I would have booked it as adoption otherwise. Second, the six Claude-User pulls are the only part of this dataset pointing anywhere. Those are real people asking an assistant to go read a page in real time, and that path is small but it is not zero, which is more than the crawlers managed.

The post is closed but the daily Log Explorer pull keeps running, because the one thing actually worth catching is the day a big crawler flips from HTML to markdown. That is a five minute update to this page, not a new experiment. If ShapBot turns out to be somebody important in six months I will look stupid, and I will write that one up too.