New here? Start here. Read the blog →

What actually gets a page cited by AI? Here's what the published evidence shows

written by Ryan Krass

filed under AI-Mediated Discovery | Being the Answer

A small stack of printed research papers beside a tall stack of glossy brochures

There is more real research on this question than the industry lets on, and less certainty in it than the industry sells. Both halves matter. Strip away the vendor decks and what remains is a small body of published work, some of it peer reviewed, some of it preprints, that has actually tested what moves a page into an AI answer. Here is what it found, with the grades attached, because a finding without its grade is just a claim.

The single most solid finding is that citation is not ranking. The one adjudicated, peer-reviewed study in the GEO literature ran thousands of queries and measured which sources generative engines actually used. On the same query, a page ranked fifth organically gained citation share while the page ranked first lost it. That is not a fluke of one dataset. A separate large measurement of Google's AI answers found nearly a third of cited domains absent from the first page of results displayed beside them. Whatever selects sources for the answer, it is not reading the rankings off the old scoreboard. This is the bedrock finding, and it is the reason the rest of this research area exists at all.

The second finding is about what the selection favors, and it is the practical one. The same peer-reviewed work tested nine different content strategies against each other. What moved citation share: adding real statistics, citing checkable sources, including attributed quotations. Substance the machine can extract and verify. What did nothing on its own: writing in a more authoritative tone. And what actively hurt: keyword stuffing, the oldest trick in the SEO book, which scored below doing nothing at all. The exact percentages vary by study and setup, so I will not hand you a number to tape to your monitor, but the direction is consistent across the published work: checkable substance wins, styling loses, and gaming gets punished.

Third, the failure side has been studied too, and it points the same way. Work that categorized why pages fail to get cited found the dominant causes were about meaning, not authority: the content did not actually match the intent of the question, or it lacked the information the answer needed. Domain prestige was not the binding constraint. The page that loses is usually the page that never actually answered the question, which is a strangely hopeful finding if you are small, and a warning if you have been coasting on reputation.

Fourth, structure matters in one specific, tested way. Passage-level retrieval rewards self-contained, coherent paragraphs. One study found that shattering an answer into disconnected bullet fragments measurably degraded retrieval. The unit that wins is a paragraph that could be lifted out alone and still make sense. Not because machines like paragraphs aesthetically, but because a fragment stripped of context is unusable as evidence for an answer.

Now the honest ledger of what is not established, because this is where the published record gets abused. Review volume as a direct citation lever: untested for commercial queries, and I have written about that gap before. Schema markup as a citation lever: tested three separate times, nothing found, and Google says no special markup is required. Any precise universal number, a citation rate, a visibility score, a fixed reshuffle interval: the repeated-measurement research says these systems are too unstable for single numbers to mean what the dashboards imply. And the biggest open question of all, which direction the whole system is drifting, toward concentration or toward the long tail, currently has serious published evidence pointing both ways. Anyone who tells you that question is settled has not read both papers.

There is one more thing you can do with this research that nobody selling you a service will suggest: check it yourself. Most of these findings are reproducible by any reader with ten minutes. Ask an engine a question your customers actually ask. Look at what it cites. Ask again tomorrow, worded slightly differently. You will see the instability, the passage-level selection, the indifference to brand size on specific questions, all of it, live. Your results will not match mine exactly, and that variance is not a flaw in the exercise. It is the finding. A system this dynamic rewards the business that keeps showing up with checkable substance, and it forgives nobody's screenshot.

So the published evidence, graded and stacked, says something almost old-fashioned: the machine cites the page that genuinely answers the question, carries verifiable substance, and holds together as prose. It punishes the tricks. Everything else, from scores to schemas to review counts, is either refuted, unnecessary, or unproven, and you should treat anyone blurring those three words as part of the problem. The research is young, parts of it will be overturned, and I will keep reading it in public. What it supports today is on this page, with the grades showing.


This essay is part of The Discovery Layer: How Businesses Get Found in the Age of AI, a book I'm writing in public. Each essay becomes a chapter, and the book expands them with the full evidence. Get on the list and I'll tell you when it ships.

Get new posts first

Occasional posts on how businesses get found, believed, and chosen now that AI sits between them and their customers. No drip sequence, no fake urgency, unsubscribe anytime.

Ryan Krass emails you when there's something new. That's the whole system.