What GEO Can and Can't Prove

What GEO Can and Can't Prove
Being known and being found are different things

A new academic survey of 45 studies checked the industry's most-cited claim against the evidence. It didn't hold up.

The number shows up everywhere in GEO (Generative Engine Optimizaton) sales material: content optimized for AI search sees up to a 40% lift in visibility. It traces back to one foundational paper, and it is real, as far as it goes. It just doesn't go nearly as far as the industry built on top of it assumes.

A new critical literature survey, reviewing 45 studies published between November 2023 and July 2026, went back to the source. The 40% figure comes from a single experiment: five documents are handed to a model already in a fixed context, one of them is rewritten, and the rewritten version claims a larger share of the words attributed to it in the answer. That's a real, measurable effect. It says nothing about whether the page gets found in the first place, whether it gets chosen over competitors outside that five-document set, or whether any of it moves a click, a lead, or a sale. The survey's own conclusion is blunt: treating this figure as a general promise about ranking in ChatGPT is not supported by the evidence, and the paper grades it explicitly as a claim to reject.

The industry treats GEO as one thing. The evidence says it's at least seven.

The survey's central argument is structural. Getting an AI system to recommend something involves a sequence of separate, mostly invisible stages: does the system decide to search at all, does it retrieve the page, does the page survive reranking into the model's working context, does the model cite it, how prominently, does the citation actually reflect what the page says, and does any of that change what a person does next. Most GEO research, and nearly all GEO marketing, measures one or two of those stages and reports the result as if it applied to the whole chain.

That matters because the stages don't move together, and can move in opposite directions. One of the more striking findings in the survey comes from a benchmark that, unusually, tested the full pipeline rather than a fixed context. Optimizing a page's body content improved how it performed once it was already in front of the model, the same effect the original 40% figure captured, but reduced how often the page got retrieved and reranked into that position at all. Averaged across the whole pipeline, the optimization made the page less likely to be cited, not more. A technique that looks like a clear win in a fixed-context test can be a net loss once the earlier stages are allowed to move.

What's actually well supported, and what isn't

The survey grades the field's claims by strength of evidence, and the pattern is worth sitting with plainly.

Strong: a document already placed in front of a model can have its citation share shifted by how it's written. Query relevance and where a source sits in the model's working context are the two most reproducible levers in the entire literature.

Moderate: specific, verifiable, dated evidence, statistics, prices, direct quotations, tends to help, though the effect depends heavily on what's actually being asked and disappears or reverses if the content isn't accurate.

Low: that any white-hat optimization technique reliably improves a brand's organic discoverability across multiple AI platforms over time. Very few studies have even tested this end to end.

Very low, close to absent: that citation, however measured, predicts clicks, conversions, or revenue. The survey found one industry study reporting a traffic lift and one academic quasi-experiment, and the latter's own placebo test came back statistically inconclusive. The causal chain most GEO pitches assume, more citation leads to more business, is, in the peer-reviewed literature, essentially unestablished.

Being known and being found are different things, at real scale

A separate finding worth its own mention: one study tested 112 real startups on whether AI systems recognized them by name versus whether those same systems surfaced them in an open, category-level discovery query, "what's a good tool for X," rather than "tell me about Company Y." ChatGPT recognized 99.4% of the startups when asked directly. It surfaced them in organic discovery queries only 3.32% of the time. Perplexity dropped from 94.3% recognition to 8.29% discovery. A model can have a brand fully encoded in its training and still almost never bring that brand up unprompted.

The survey also found that different AI platforms barely share the same sources at all. Comparing citations across ordinary Google search, Google's AI Overviews, and Gemini, overlap measured between 11% and 18%. A brand doing well on one surface has very little basis to expect the same result on another, and the paper is explicit that there's no such thing as a single, transferable GEO ranking.

Where this lands

None of this means GEO is worthless, and the survey doesn't argue that. Query relevance and context position are real, replicated, useful levers, and getting them right is worth doing. What the evidence doesn't support is treating citation as a proxy for the outcomes that actually matter commercially, being found rather than merely known, and being chosen rather than merely cited.

That's not a new argument from us. It's an independent academic literature arriving, through a completely different method, at the same structural conclusion: measuring one stage of a multi-stage process and reporting it as the whole story is where most of the industry's confidence currently outruns its evidence.


Source: Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization. arXiv preprint, version dated July 15, 2026.

AIVO Meridian