Four Questions Worth Asking Before You Trust an AI Visibility Score
A public fight between two well-funded measurement companies revealed the real problem. It isn't the number either of them is defending.
The CEO of a $19 million company told a few thousand marketers that a $1 billion competitor had a data problem. Brian Stempeck, who runs Evertune, wrote that Profound's sampling method left error bars wide enough to make the underlying scores unreliable for either measurement or optimization. Ask an AI assistant the same question once a day for a month, he argued, and a brand's visibility score can carry a margin of error of 12 points. Spend six figures lifting that score by 10, and you'd never know if it worked. The post pulled several hundred reactions in two days.
His math checked out. Sampling one prompt a small number of times does leave real uncertainty, that's basic statistics, not a competitive attack. But Profound had already run the experiment he was implicitly describing, and published the result eight days earlier. Tested against a portfolio of hundreds of different prompts, run once a day versus ten times a day, the difference in the pooled reading came out to a quarter of a percentage point. The two companies, it turned out, were both right about a narrower disagreement than the headlines suggested: one measures a wide net of different questions once each, the other measures a smaller set of questions many times. Different tradeoffs, not a data scandal.
The interesting responses didn't take either side
The public replies worth reading came from people who ignored the fight entirely and pointed at something underneath it.
One measurement builder noted that with hundreds of prompts, the average was always going to stabilize, that was never really in question. What moved the result more than any amount of repetition was which prompts were included in the first place. A brand sitting at a stable 10% citation share could be at 10% everywhere, or at 100% on some questions and zero on the rest, two completely different situations for anyone deciding where to spend, and neither number tells you which one you're looking at, or why.
A second reply widened the problem past sampling entirely: if AI systems personalize answers even lightly off user history, it's not clear any panel of anonymized test runs reflects what a real person actually sees, no matter how many times you repeat the prompt.
A third named something repetition can't touch at all. A prompt run daily for a month isn't a set of independent draws, it's a time series, and the model underneath it keeps changing. That's drift, not sampling noise, and no amount of repetition removes it, it just tightens a smaller source of error while a larger one sits untouched.
And one comment landed the whole structural problem in a single line: the company measuring your visibility is usually the same company selling you the fix for it. The personal trainer who sells you the supplements is also the one grading your progress.
The four questions that actually matter
Out of that thread came a sharper diagnostic than either company's dashboard, four questions worth putting to any vendor selling an AI visibility number, before spending against it:
What's the confidence interval around that score, not the score itself. Which prompts define your market, and who chose them. How is your content's actual effect separated from the model simply drifting on its own. And, plainly, can the vendor show the work behind the number, or only the number.
A vendor with real answers to all four is measuring something. One with a single confident figure and no error bar is selling precision it hasn't earned.
Where we've already had to answer these, and where we haven't
We publish our own measurement standard openly, so it's worth holding it against this exact bar rather than asking others to clear one we wouldn't.
On confidence intervals: no status gets assigned from a single run in our methodology. Each measurement point is a replicate set, several independent runs of the same multi-turn conversation, not just one pass, reported as a rate across the set rather than a binary yes or no.
On who picks the prompts, and whether that's auditable: the journey templates, the structured, repeatable conversation scripts used to test a brand, along with the decision-turn classification rules and any weighting, are published or available for audit, not held as a black box. Template changes are versioned, and a time series can't quietly splice across an unversioned change without disclosure.
On separating content effect from drift: measurement is longitudinal by design, not a snapshot, and every observation records the model or platform version at the time, specifically because a platform update can shift results independent of anything a brand does. Observed changes get checked against known platform-version changes before anyone treats them as a real effect of something a brand did.
On showing the work: the specification behind all of this is published for public comment, deposited with a dated DOI, openly licensed, and free to adopt or audit, not proprietary.
That's three of the four questions with a real, structural answer, not a promise. The fourth deserves the same honesty this space is currently missing. Personalization, whether two different people asking the same question ever see the same answer once location, prior conversation history, and session-level context start shaping what a model retrieves, isn't solved here either. It isn't solved anywhere in this category yet. Any vendor claiming otherwise is further ahead of the evidence than the evidence currently allows.
What this means for anyone about to spend against one of these numbers
The category will get more rigorous eventually, there's too much money moving into it for that not to happen. When it does, the advantage won't belong to whoever runs the most prompts or repeats them the most times. It'll belong to whoever is willing to tell a buyer plainly what their number can and can't do, before being asked.
Until then, the four questions are worth asking every time, regardless of who's answering them, us included.
Sources: Reporting and commentary on the Profound/Evertune sampling debate, State of Brand, July 2026.