Appeared in 96% of Journeys, Chosen in 13%

Appeared in 96% of Journeys, Chosen in 13%
What AI visibility measures, and what happens to a brand after it appears

AIVO's journey corpus now covers more than 25,000+ multi-turn probes across 250+ brands. Each probe follows a simulated buyer from an opening category question, through a series of accumulating requirements, to a final request for a recommendation. Within that corpus, 1,427 journeys isolate a single question: when a brand appears early in the conversation, where is it by the end? The brand appeared in the model's early answers in 95.7% of those journeys. It remained the final recommendation in 12.7%.

The gap between those two figures is the subject of this article. Appearance and recommendation are separate states in an AI-mediated decision, and most of what the AI search category currently measures sits on the appearance side.

A category built around the first answer

The industry organized quickly around one question for brands: does AI see me? Platforms track how often a brand appears in ChatGPT, Gemini and Perplexity, where it is positioned, which sources are cited, and which competitors appear alongside it.

The scale of investment shows the market believes the problem is real. On September 15, Profound announced a $180 million Series D at a $1.8 billion valuation, co-led by Sequoia Capital and Kleiner Perkins, less than seven months after its $96 million Series C. The company serves more than 1,000 enterprise brands and over one third of the Fortune 100. It says revenue tripled over the past six months, a figure it reported itself that has not been independently verified. Its platform draws on more than 2 billion real user prompts, and it is expanding from AI search analytics into an agentic platform that performs work across marketing organizations.

Profound is building something sophisticated, on the premise that understanding and shaping what AI systems see and say about a brand will translate into what those systems recommend. The corpus raises a more specific question: whether visibility is a sufficient proxy for what happens when the AI system has to choose.

How a decision forms across a conversation

Consider a buyer who asks an assistant for the best luxury skincare brands for sensitive mature skin. Five brands appear. A visibility platform can accurately record which appeared, which was cited and which ranked first.

The buyer then adds that the product must be fragrance-free, and the set changes. She rules out retinoids, asks which option also treats pigmentation, and specifies something she can wear under makeup each morning. Each requirement reshapes the candidates. When she finally asks which one the assistant would actually buy, the model has to choose. The brand that ranked first in the opening answer may no longer be in contention.

The first answer and the final recommendation are different events, and they call for different units of measurement. A prompt-level measure records presence, position, citation and sentiment in a single response. A journey-level measure records which brands persist as requirements accumulate, alternatives are compared and the model is finally asked to commit.

A visibility score reports visibility accurately. The difficulty arises when it is read as a forecast of choice. A brand appearing in 90% of first answers may persist through only 20% of the conversations in which buyers introduce the requirements that decide the purchase. Both figures are correct, and they describe different things.

A worked example

Take two brands in the same category. Brand A appears in 80% of relevant first answers and Brand B in 35%. On a visibility dashboard, A looks more than twice as strong.

Now follow the journeys. As buyers introduce price, compatibility, availability and quality requirements, Brand A persists to the final recommendation in 30% of the journeys it enters, while Brand B persists in 85%. Multiplied through, A's final recommendation rate across relevant journeys is roughly 24% and B's is roughly 30%. The brand with less than half the visibility is recommended more often at the end. This is an illustrative calculation, not an AI Win-Rate figure under the published methodology, but it shows how the two measures can rank the same brands in opposite order.

Why the journey is harder to measure

A single prompt is static and can be replayed thousands of times. A decision journey carries memory. Each requirement changes the context for the next answer, and a constraint introduced at turn four can eliminate an option that looked entirely viable at turn two. That makes journeys more expensive and more complex to measure. It also places them much closer to the commercial event, particularly as AI systems move from answering questions to helping people decide what to buy, which provider to use and where to go.

Search marketing trained the industry to treat ranking as a reasonable proxy for the click, because the ranking sat close to the moment of choice. In AI-mediated discovery, the answer is one step in a reasoning process. A brand can appear and then fail a new requirement, satisfy the requirement and lose on price, or survive every filter and still be passed over when the buyer asks for a single recommendation.

What the category needs to measure next

AIVO's AI Win-Rate Methodology v1.0 uses the defined decision conversation as its measurement unit. It measures the proportion of those conversations in which a brand receives the assistant's decisive recommendation. The journey research adds a diagnostic layer, recording where brands are eliminated and which requirements precede displacement. Visibility remains part of the picture as the entry condition to a decision. Win-rate measures the outcome.

The first phase of AI search measurement established whether brands appear. The next phase has to establish what happens to them afterwards, across the turns in which a buyer's requirements take shape and the model makes its choice.

AIVO Meridian