Thirteen domains, only one cited by all three

"Is my site cited by ChatGPT?" sounds like a question with one answer. It has at least three, and they depend on how you asked.

What we compared

On August 10 we asked the same question — miglior vino bonarda oltrepò pavese, Italian for "best Bonarda wine from Oltrepò Pavese" — with the same prompt, on the same day, from the same country, through three different ways of collecting the ChatGPT answer:

  • A browser driven on the chat, which is the same surface a person sees when they open ChatGPT and type.
  • The model's official web search, called through an API by a data provider.
  • Web context injected into the prompt: an external search engine fetches the pages and hands them to the model inside the question.

All three are described as "ChatGPT with web search". All three gave a sensible answer of similar length: 5,666, 5,897 and 7,134 characters.

The sources, however, do not match

Together the three answers cite thirteen distinct domains. Only one appears in all three lists. A second appears in two. The other eleven are cited by a single method.

Cited domain browser on the chatmodel's official searchinjected context
gamberorosso.it cited cited cited
classeseoltrepopavese.it — cited cited
consorziovinioltrepo.it cited — —
callmewine.com cited — —
gamberorossointernational.com cited — —
monsupello.it cited — —
laprovinciapavese.gelocal.it — cited —
buonalombardia.regione.lombardia.it — cited —
valoritalia.it — cited —
slowfood.it — — cited
aislombardia.it — — cited
quaquarinifrancesco.com — — cited
valdamonte.it — — cited
Total 556
One row per cited domain, one column per way of collecting the answer. The only row cited in all three columns is the first.

A wine producer from Oltrepò who appeared only in the first column would come out as "visible on ChatGPT" or "invisible on ChatGPT" depending on which tool sends the report. Neither tool would be wrong.

Why they diverge

Because they are not looking at the same thing. The browser reads the surface people see. The model's official search goes through the engine the model provider really uses, but behind an API with settings of its own. The injected context does not use ChatGPT's search at all: another engine picks the pages and puts them in the model's mouth, so it is measuring ranking on that engine, not visibility on ChatGPT.

One detail makes this visible to the naked eye. The cited URLs carry a different marker in each case: in the first, the same parameter ChatGPT appends when a real user clicks a source; in the second, an API marker; in the third, no marker at all. The first is the only one that will also show up in the cited site's traffic statistics.

The cost changes three hundredfold

On the same single question, the most expensive of the three cost about three hundred times the cheapest. That is not a pricing nuance. It is the difference between checking a domain every day and checking it once a month.

And the cheapest of the three is the one that reads the surface people see. That is less reassuring than it sounds: the method depends on an automated browser that keeps working, not on a contract.

The two practical consequences

First: changing the way you measure breaks the history exactly as changing the model does. If a client's AI visibility chart goes up the month the provider changed method, visibility did not go up. The yardstick changed.

Second: a single check is not a measurement. With thirteen distinct domains and an overlap of one, the source list of a single answer is a sample, not a verdict. What counts is repetition over time and how often a domain comes back.

The part that complicates the picture

Two things we measured that make the matter less clear-cut than we have told it so far.

Web search is a permission, not a guarantee. You can ask the model to search. It decides whether to do it. On questions it answers from memory it does not search at all and cites no one — and a tool that forces search on every run overestimates visibility, because it counts citations that would never have appeared in real life.

On some systems visibility is sometimes not measurable. On the same question, one of the systems we compared answered without citing any source. It was not a parsing error: the same code, on another question, extracted seven. The honest number there is not zero, it is "not available", and those are two different things that too many dashboards draw the same way.

What to ask whoever sells you a measurement

Three questions, and you know right away what you are dealing with:

  • Which surface do you read the answer from: the chat, the model's API, or an engine of your own?
  • Do you record which model answered each single check?
  • What do you write when the system cited no one: zero, or "not available"?

If the answer to the first is just "ChatGPT", it is the same answer all three tools in this comparison would give — and on the same question, on the same day, they produced thirteen sources with only one in common.