Good Statistics, Wrong Turn

Good Statistics, Wrong Turn
The paper is, by its own careful design, a study of the first prompt and nothing after it.

A rigorous citation study just raised the bar for the whole field. It still never asks whether a citation becomes a recommendation.

Petra Labs published two research papers this year on how AI assistants cite sources. Read them and it is hard not to respect the work. They resample whole model runs rather than individual citations, so they do not mistake twenty correlated citations inside one answer for twenty independent data points. They use bias-corrected bootstrap intervals rather than simple percentile ones, because the distances they measure are skewed near their ceiling and a simpler interval would misstate the uncertainty. They correct for multiple comparisons using false discovery rate control. They publish a detection threshold for how many times you need to run the same prompt before you can trust that a change is real rather than noise, and they show that at the run counts most tools use today, a lot of what looks like a shift in AI visibility is chance.

This is more careful applied statistics than most of what circulates as AI-visibility research, ours included in places. Where our own published methodology has not stated a formal noise floor for a metric, this work is a fair prompt to add one. Rigor is not a competitive move. It is simply owed to anyone reading a number and deciding what to do about it.

None of that changes what the two papers actually measured.

Every seed prompt in Petra's main study carries no brand name and no conversation history. Each one is submitted once, logged out, and answered on its own. The object of study, stated plainly in their own introduction, is how far a single question's wording moves the set of sources an assistant cites in that one exchange. There is no comparison turn. No challenge. No moment where the assistant is asked to actually choose. The paper is, by its own careful design, a study of the first prompt and nothing after it.

That is not a flaw in the work. It is a boundary the authors draw honestly, the same way we try to draw ours. But it means the paper cannot answer the question that matters commercially, which is whether being cited at that first prompt has anything to do with what the assistant recommends once a buyer has actually worked through a decision. Their own data offers a hint that it might not. The largest citation shifts they measured were not wording changes at all, but changes to who is asking and what they are asking for, a shift from a commercial question to an instructional one moved citations further than anything else in the study, with a regulated buyer persona and an added requirement close behind it. Three different ways of changing what the question is actually for, all clustered at the top.

If citations move that much just from what a buyer brings to the conversation, a first-prompt citation count was never going to be a stable proxy for a final recommendation. It is one frame from a conversation that has not finished.

We have made a related point before, using a different lens. Our own studies show that mention and recommendation are frequently decoupled. Brands named in nearly every first response are recommended a fraction of the time. Some brands never named at all are recommended more often than the household name sitting next to them. What Petra's research adds is a rigorous account of how unstable the first answer already is on its own terms, before a buyer has said a single word back. Two teams working from opposite ends arrive at a similar place. The first turn is a noisy, moving thing. Building a strategy on it, however carefully you measure it, is still building on the wrong turn.

The honest version of our own position is this. Petra Labs has shown, more carefully than we or anyone else has, that citation counting is harder to do well than the industry assumed. We think their next paper worth writing is the one that runs the same rigor forward, past the first answer, to the moment a buyer is actually choosing. That is where we have been working. It is where the revenue lives.

Full findings on how discovery and recommendation diverge, across payments, skincare, and car rental, are in our ranking studies and our open corpus.

AIVO Meridian ยท aivomeridian.com/ai-rankings