How do you measure AI search visibility without fooling yourself?

Ask an AI engine the same question twice and you can get two different answers citing two different sources. That single fact breaks most of what is currently being sold as AI visibility measurement, and it is where any honest version has to start.
The instability is measured, not anecdotal. Researchers who sampled the same queries repeatedly across Perplexity, SearchGPT, and Gemini, daily for nine days and then at ten-minute intervals, found that citation results vary substantially between identical runs. That work is a preprint, so hold its specifics loosely, but its conclusion is the methodological heart of this whole topic: a citation share is not a fixed value you look up. It is a sample from a distribution, and many of the differences a dashboard will happily display between you and a competitor fall inside the noise floor of the measurement itself. Separate measurements point the same direction. Reword a question slightly and the cited sources change. Ask different engines and they agree with each other far less than you would expect. On one engine, the cited set turns over on a cadence measured in days.
Now hold that against the product being sold. A single number, measured once, presented as your score. If the underlying thing is a distribution that moves between two identical runs ten minutes apart, then a one-shot score is not a simplification of the truth. It is noise wearing a suit. I have looked for a published, defensible methodology behind the visibility scores currently on the market and have not found one, and the term the industry likes for the model-side version of this was coined by a marketing agency, and I have not found a published measurement method behind it either. That is not a technicality. It is the difference between an instrument and a prop.
So what does honest measurement look like? Four properties, none of them exotic.
Measure repeatedly, not once. One query run tells you what one roll of the dice showed. The shape only appears across repeated samples, which means days and weeks, not a screenshot.
Measure variations, not a single phrasing. Real customers ask the same question a dozen ways, and the cited set changes with the wording. A measurement built on one canonical phrasing is measuring the phrasing, not your visibility.
Measure across engines, and keep the results separate. The engines disagree, they run different machinery underneath, and an average across them describes nothing that exists. What you want is a per-engine picture over time.
Distinguish what happened. Being retrieved, being cited, having your claim carried accurately into the answer, and getting an action out of it are four different events, and they do not move together. A page can be cited and never visited. It can be retrieved and never cited. Deciding which of those you care about, before you measure, is most of the work.
Then report it the way you would want any number reported to you: as a range that moves, with the honest admission of how much of the movement is noise. If a report cannot tell you that, the precision on the page is decoration.
Here is the test I would put to any vendor, and it takes one email. Ask them: if you ran your measurement twice, an hour apart, how different would my score be, and how do you know? A real methodology has an answer, because measuring its own noise is the first thing a real methodology does. A prop changes the subject. You will learn everything you need from which response you get.
And notice what this means about the cheap version of this work. You do not need a platform subscription to start. Pick the ten questions that actually bring you customers. Ask them, in a few phrasings, on the two or three engines your customers use, once a week, and write down who gets named and cited. That is unglamorous, and it is already more methodologically sound than a one-shot score, because it samples over time, over phrasings, and per engine, and after a month you can see which questions you hold, which you never appear on, and which just flicker. The spreadsheet is not impressive. It is merely honest, which turns out to be the rarer property.
The pattern underneath is one I have watched for twenty years: every time a new visibility surface appears, the first products to market sell certainty about it, because certainty is what buyers want and noise is what the surface actually contains. The sellers of certainty do fine until the buyers compare notes. Measurement that admits what it does not know is slower to sell and much harder to embarrass, and in a field this young, not being embarrassed in two years is the whole game.
This essay is part of The Discovery Layer: How Businesses Get Found in the Age of AI, a book I'm writing in public. Each essay becomes a chapter, and the book expands them with the full evidence. Get on the list and I'll tell you when it ships.