New here? Start here. Read the blog →

What Is Answer Engine Optimization, Really?

written by Ryan Krass

filed under AI-Mediated Discovery | Being the Answer

Rows of library card-catalogue drawers with one pulled slightly open

Answer engine optimization is the practice of influencing what an AI answer engine says about you, and whether it names you at all when someone asks a question in your category. Some consultants call it AEO, some call it GEO. That's the whole definition, stripped of the deck it usually arrives in. I've spent enough time looking at what's actually been tested in this category, versus what's being sold under its name, that I'd rather give you a skeptical read than a promotional one.

Start with the mechanism, because you can't judge the category's claims without it. A question you ask an AI system rarely stays as one question. Google has said outright that its AI Mode takes a single prompt and issues, in its own words, "a multitude of queries simultaneously," with deeper research modes running hundreds behind one request. So the first layer is fan-out: one question becomes many narrower ones, and content wins by answering a specific sub-question well, not by ranking for the broad topic around it.

The second layer is retrieval, and it doesn't happen at the level of your homepage. Engineers building these systems describe retrieving and reranking content at what they call passage or sub-document granularity. The system pulls candidate paragraphs from across the web that seem to match each sub-question, not whole pages and not whole domains.

The third layer is attribution, decided at generation time. Google's own patent on generating these summaries, US11769017B1, describes a mechanism that checks each claim the system is about to state against candidate passages by measuring how close their meaning sits to the claim, under a set threshold, and links whichever passage clears it. The system writes the sentence, then finds a passage close enough in meaning to back it up, and cites that one. That's the actual pipeline: fan-out, passage-level retrieval, embedding-based attribution. Nothing mystical, and nothing that requires buying a certified methodology to understand.

Now, the part the category gets wrong, and gets wrong often enough that I think it's worth saying plainly. A lot of what's sold as AEO or GEO is schema markup and formatting work repackaged with a new label. It's an easy thing to sell: a defined technical task, a start and an end, an invoice you can hand over. It has also been tested, more than once, specifically for whether it moves AI citation, and it hasn't. Formatting-only changes, holding the actual content constant, show no measured effect on whether a page gets cited.

Worse than that is the rise of the numeric "GEO score," a single number some tool hands you claiming to represent your AI visibility, as though it were a credit score for the machine age. I haven't found a published, replicated formula behind any of these products, and the specific structural claims I've seen tested against real data haven't held up when someone tried to reproduce them. A black-box score invites you to chase a number that doesn't trace back to anything measured. It just repackages uncertainty as a metric, which makes the uncertainty feel solved when it isn't.

The third pattern worth naming is precise lift percentages, quoted with total confidence, that don't survive anyone else trying to reproduce them. You'll see a case study claiming a specific site saw citations jump by some exact number after some specific change. Ask what else changed on that site in the same window, and usually the answer is "several things," which means the number is attributing an outcome to a cause nobody actually isolated. This is the same move the schema pitch makes: certainty a genuinely uncertain field hasn't earned yet, dressed up as a result.

So what actually seems to hold up, when you look at real measurement instead of a sales deck? Semantic alignment to the specific sub-question is the biggest one. When researchers looked at why a page that should have been a strong candidate got skipped, mismatch to the exact question being asked was the single largest documented cause, well ahead of anything structural. Page quality and clarity correlate too, in the plain sense a person would recognize reading the page. So does evidence density: real numbers and real quotes a machine can lift into a generated answer, rather than a page that describes everything in general terms. I want to grade that last one honestly: the direction, that specific evidence beats vague description, is well supported. The exact size of the effect is contested, and I won't hand you a number I can't stand behind, even though a competitor's sales page probably will.

Here's what that looks like on an actual page, not in the abstract. Say a customer's real question is "how much does it cost to install a tankless water heater." A page built for the old game buries that number three sentences into a paragraph about the company's commitment to customer service. A page built for the sub-question states the price range, the timeframe, and the zip codes it covers in one self-contained passage. It would still make sense quoted alone, with nothing above or below it for context. Same fact, same business, wildly different odds of getting cited, and the difference has nothing to do with schema, a score, or a lift percentage. It's whether the actual sentence a person needs exists somewhere plainly stated.

One more honest thing worth saying about this whole category: it doesn't even agree on its own name yet. AEO, GEO, AI SEO, LLM optimization, all competing terms for the same undefined territory, coined by different vendors trying to own the label before anyone's proven the underlying mechanics. That immaturity isn't a reason to dismiss the shift itself, the shift is real and I've written about the mechanism behind it elsewhere. It is a reason to be skeptical of anyone who shows up already selling you a certified methodology and a branded dashboard for a field that hasn't finished arguing about what it is.

If I were spending money in this category today, it would go three places. Write clearer, more specific answers to the real sub-questions your customers are asking. Keep the facts about your business consistent and verifiable everywhere a machine might look. Pass on anything sold to you as a formula, a score, or a guaranteed lift percentage. That's the boring, unglamorous version, and it's also the only version I can actually defend.

The category doesn't have its own name settled yet. Don't let anyone sell you a settled science.

Get new posts first

Occasional posts on how businesses get found, believed, and chosen now that AI sits between them and their customers. No drip sequence, no fake urgency, unsubscribe anytime.

Ryan Krass emails you when there's something new. That's the whole system.