Do LLM Bots Actually Read JSON-LD, and Will They Cite It?

Experiment closed - Day 95

TL;DR - We built a secret page with a coined/made up phrase that only lives in the JSON-LD, invisible to anyone reading the page. Five search engines and about a month days later: they did consume it, but only the ones that actually go fetch the page.

Scope: this tests whether structured data that sits in a page's source but not its visible render gets ingested and cited by AI answer engines. Vol. 02 covered whether those engines execute JavaScript at all; this is the next layer down.

Does JSON-LD Actually Feed the Answer?

Search engines have read structured data for years. The pitch every schema vendor makes is that this same invisible layer now feeds the AI engines too, my goal was to test that.

A ton of what a page tells a machine never shows up on screen for a human (forget JS for a second)... JSON-LD, meta tags, HTML comments, stuff that sits in the source and never renders. So my question is: if a "fact" lives only in the JSON-LD, with nothing visible backing it up anywhere on the page, will an AI engine swallow/consume/remember that fact and repeat it back when you ask?

What we built

One HTML page with one made up keyword/term. The whole thing hinges on where that "framework" lives: its name and details sit only in the structured data and a few other non-visible layers. Not in the body copy, not in the headline, not in the meta description.

To keep it honest, every layer carries its own marker, so a hit points back to exactly one source:

  • A visible control - a separate term written into the body copy that you can actually see on the page. If the engines repeat this one, I know the visible text got read, so silence on the structured data means something.
  • The JSON-LD layer - the framework and one coined sub-term live only here, and that sub-term exists nowhere else on the page or anywhere on the web.
  • A few other hidden layers - a custom meta tag, an HTML comment, a display:none block - each with its own marker, so I can see whether the engines treat them differently from JSON-LD.

Every layer got a unique token. A coined sub-term cannot show up by accident - it does not exist anywhere I did not put it. So if an engine spits it back, it read the structured data. That is the whole trick.

How we are measuring success

First, did they even fetch the page? Our logs show every known AI crawler that hits the site - which bot, which path, what time.

Second, the regurgitation test. Once the crawlers have had a few days, we ask major AI's about the terms we planted. Ask about the visible control to confirm the visible text was ingested at all, then ask about the "structured-data-only term" to see if it comes back. That term exists nowhere else, so any AI that returns it read the JSON-LD. It does not need me to take any model's word for anything.

We also did this using our phones + real accounts, not LLM APIs (we did test them as well.) We also did these expirments signed in + signed out and private/not private / incognito etc. We didn't share the URL or secret phrase until just now when we published this case study.

What we expected

Leaving my original prediction up, unedited, because I want credit for the half I got right and I am not going to quietly delete the half I got wrong.

"I have no idea. AI said it should happen fast, I think it will take weeks or months. My bet is some AIs read the JSON-LD and some do not."

The "some do and some do not" part was correct. What I had completely wrong was the reason. I assumed it was a model capability thing, that the smarter engines parse structured data and the dumber ones skip it. That is not what splits them at all, and the actual answer is a lot more useful.

Here is how we structured the test

I declared this case study "complete" as of 9/28/26 so here is everything. The test page is llmcartel.com/cvi-probe. Go look at it. Then view source, because those are two different documents and that is the entire point.

I invented a fake metrics framework called the Cartel Visibility Index and gave it five made-up pillars: Citation Frequency, Entity Saturation, Lattice Saturation, Answer-Surface Share, and Prompt Coverage. That whole list lives in one place - a description string inside the page's JSON-LD, on line 64 of the source. It is not in the headline, not in the body copy, not in the meta description.

Lattice Saturation is the one that matters. I made it up specifically because you cannot guess it. "Citation Frequency" is the kind of thing any model would cough up if you asked it to invent a GEO metric - that word pairing is floating around the industry already. Nobody stumbles onto "Lattice Saturation" by accident. Before launch it returned zero results on Google.

4 other layers / things we made up for the test

  • Visible control - the body copy talks about "Mention Cadence," a term you can read on the page with your eyes
  • Custom meta tag - a claim that we have applied the framework across 47 client engagements (this was made up for this test)
  • HTML comment - an internal build number, v3.2, codename "Ridgeline."
  • A display:none div - "outputs a single 0-100 CVI Saturation Score per entity."

Quick disclaimer on that last one, because somebody is going to email me about it. A hidden div full of keywords is cloaking and you should not do it on a real page. I did this for science not to score my clients and it is just for the case study.

What the engines actually did

ClaudeBot fetched the page about 90 minutes after it went live, which looking back was probably a leak from me using Claude Code on the same server where I was running the experiment. Then I spent five weeks asking engines about it at random dates/times.

Cartel Visibility Index canary results by engine, July 10 to September 28, 2026
Date / engineFetched the page?What came back
07-10 CopilotNoInvented a definition. Cited geneo.app, upfront-ai.com, humanswith.ai. Zero llmcartel.
07-20 KimiYesCited llmcartel.com and linked the probe page directly.
07-24 GrokYesReproduced the display:none line near-verbatim.
07-26 Gemini 3.6 FlashNoPlausible definition, no source chips, cited nobody.
07-31 GrokYesAll five pillars, verbatim, in source order.
09-02 CopilotNo - served from Bing's indexReversed its July answer. Cited llmcartel repeatedly, quoted the visible control and the hidden div.
09-28 Bing / DuckDuckGon/a - search, not chatBoth rank the page #1 for the coined phrase and print the hidden div as the snippet.

The fetch column lines up perfectly and it has nothing to do with which model is smarter.

Every engine that fetched the page came back with something real from our markup.

Every engine that answered out of its own memory made up a confident, plausible, completely fabricated definition - and Copilot went further and attributed the concept to three other companies in our space. It did not say "I don't know." It handed our made-up metric to our competitors LOL.

The run that settled it

July 31, around 7pm. I ran the query on a friend's Google Pixel, logged out, in a temporary chat - no account, no history, no chance the model was echoing something from an earlier session of mine. He asked what the Cartel Visibility Index was.

Grok thought for 31 seconds, then listed all five pillars in the order they appear in our JSON-LD. Including Lattice Saturation.

And it showed its work. Grok's tool trace has this in it:

curl -s "https://llmcartel.com/cvi-probe" | head -500

That is not a headless browser rendering a page, it is a shell command piping raw HTML into a code interpreter.

Our edge logs show one IP, 34.11.74.3 on Google Cloud, user-agent curl/8.5.0, seven requests to the probe page inside that two-minute window.

That address appears exactly seven times in our entire log, all of them during this session.

Grok got the five pillar names from our markup, correctly. Then it wrote a one-line definition next to each one.

When LLM's don't know the answer to something, they make it up!

Then waited a month, and on September 2 I ran the Copilot query again.

In July it had invented a definition and handed credit to geneo.app, upfront-ai.com, and humanswith.ai. This time it cited us, six times, and quoted two of the planted layers back at me - the visible control and the hidden div.

Here is what makes that different from the Grok run. I also looked at every log on Sept 2 and no trace of Copilot.

So the hidden text was not read live off our server. It was sitting in Bing's index, put there by a single crawl forty days earlier, and it came back out when somebody asked a question.

It is not just that a live-retrieval engine can read your invisible markup. It is that the invisible markup gets stored in a mainstream search index and resurfaces weeks after the crawl.

Also I started using search engines just for fun to see what came up in AI / answers / vanilla SERPs.

Bing results for the exact phrase Cartel Visibility Index, showing llmcartel.com/cvi-probe as the top result with a snippet reading: The Cartel Visibility Index outputs a single 0-100 CVI Saturation Score per entity. Ref CARTEL-CVI-20260625-L4-HID-H8
Bing, exact phrase. That snippet is not our meta description. It is the contents of a display:none div, tracking token included.

Read the description under that result...we never wrote that! There is no meta description on the page that says any of that, and the phrase does not appear in the visible body copy. Bing pulled it out of the hidden div - the one that renders to nothing and used it as the page summary. It even kept Ref CARTEL-CVI-20260625-L4-HID-H8, which is the marker I stuck in that layer so I could tell it apart from the other four.

DuckDuckGo does the same thing and even notes "No more results found." Worth knowing that DuckDuckGo is served by Bing's index, so that second screenshot is the same index behind a different front end, not a second company confirming it.

DuckDuckGo results for the exact phrase cartel visibility index, showing llmcartel.com/cvi-probe as the only organic result with the same hidden-div snippet, followed by the message: No more results found
Same index, different front end. Note the line at the bottom.

We are the only page on the internet for that phrase, which is exactly what I built it to be, and the one result it does have is describing itself with text no visitor has ever seen.

The string "cvi" is also in the URL slug, /cvi-probe, so the acronym on its own proves nothing. The full phrase "Cartel Visibility Index" is the load-bearing part - it appears five times in the source and every one of them is a layer you cannot see, and it appears zero times in the rendered text.

Also this domain is still relatively new and has very low domain authority. It is a small company and not really on the radar of many search engines / LLM crawlers (yet) so next I am going to try this same experiment (maybe) on a client website that gets a lot of traffic.

What this means for your schema

Structured data is a real ingestion channel for AI answers. A term that no human eye can see on the page came back word-for-word in a generated answer, on a device that had never talked to us, and we logged it.

Your schema only matters when the engine goes and gets your page. If it answers from training data, your markup didn't do its job. Nothing you encode on your site influences that answer.

Bing crawled the page once, kept the invisible text, and served it months later to an engine that never came back for a second look.

So retrieval is not only a live fetch at the moment somebody asks - it can be a crawl that happened in July answering a question in September. Getting fetched once still matters. It just does not expire the way I assumed.

This kind of reframes what the purpose of schema/structured data is for. It is not a ranking input you set and forget. It is what the engine finds when it arrives, so the job is making sure it arrives - being fresh, being crawlable, being the obvious thing to fetch when someone asks about your category. Get retrieved and your markup does real work.

What's next

I want to re-run this case study on a high traffic website with lots of authority with several canaries to really see how this works.

Our next case study is already in the works and is the mirror image of this one - a page we told the search engines to ignore, testing whether AI answer engines respect that. That one is still live and I am about to publish.

If you want the raw logs or the query transcripts, ask and I will send them over.