The Decision Now Forms Inside the Search

The Decision Now Forms Inside the Search
Measurement has to follow the journey

A brand can drop out of a multi-turn AI conversation for two different reasons. It can fail a requirement the user has just added, or the model can lose track of a requirement the user set earlier. The first is about the brand; the second is about the model. Measurement that samples only the first answer cannot see either failure, and a single run cannot reliably distinguish between them.

A decision built turn by turn

Generative search can operate as a multi-turn interaction. A user opens with a broad need, receives a set of candidates, then adds requirements, compares options and asks for a pick, all within one thread. The model answers each turn against everything said before. Traditional keyword search generally treats each query as a separate retrieval event, and refining it usually means writing a new query.

That changes what a brand has to achieve. Consider a moisturiser named in answer to a question about dry skin. It survives the user's next message, asking for something fragrance-free. It drops out when the user adds that nothing should feel heavy. By the time the user asks which product the model would actually choose, it is no longer in contention. The brand was recognised at the start and absent from the final recommendation.

AIVO calls this pattern the Linkage Gap: the distance between being known and being recommended. Across more than 20,000 multi-turn probes covering over 200 brands, 87.3% of brands identified earlier in the journey did not survive to the final recommendation. Earlier AIVO work treated that gap primarily as a continuity problem across turns. This article refines that account. Some displacement reflects the model losing context, and some reflects brands genuinely failing criteria the user introduced along the way.

What the final prompt leaves out

Independent research demonstrates why the distinction matters. Benjamin Tannenbaum's analysis of 8,133 real conversations found that the final prompt carried a median of about 36% of the user's own wording, and that in nearly half of conversations a stated requirement appeared only in earlier turns. A follow-up experiment held the final message constant and removed the turns before it. Across 180 sampled conversations answered by one model, the answer changed materially in 44.7% of cases, rising to 68.5% in the commercial conversations.

Those papers establish that context changes answers. The question for a brand is more specific and more commercial: at which turn does it drop out, and what preceded the exit? That is the question AIVO's probe methodology is built to answer, recording the Displacement Initiation Turn for every brand in every journey.

Two ways to lose

Models also handle accumulating context imperfectly. Laban and colleagues at Microsoft Research and Salesforce Research split fully specified tasks into pieces and revealed one piece per turn. Across 15 models, performance fell by 39% on average compared with the same task given in a single prompt. The authors found that much of the degradation reflected increased variability between runs, while peak performance fell substantially less.

That is why the two causes of displacement need to be separated. The first is criterion failure, where the brand genuinely fails a requirement the user has just introduced. A common form is what AIVO calls the Attribute Collision Gap, in which a brand's own claim, such as being highly pigmented, conflicts with what the user asked for. The second is context failure, where the model loses or misapplies a requirement the user set earlier in the conversation. The two are analytically different, and a single run cannot reliably tell them apart.

Why repeat runs matter

AIVO's replicate study, What a Citation Is Worth, ran identical prompts three times. The citations and evidence a model drew on changed in 95.8% of cases. The final brand recommendation changed in 14.6% of cases, and across 1,800 brand-level groups, 85.4% produced the same outcome in all three runs. Citation and evidence selection were substantially less stable than the final recommendation.

That result shapes the method. An exit that recurs consistently across runs is stronger evidence of criterion-driven displacement. An exit that varies across otherwise identical runs signals greater model-state instability. Tracking citations alone would miss both signals.

What to measure

Measuring a brand in a multi-turn AI journey therefore means following it to the end. Three outputs matter. Survival asks whether the brand remains in contention at each turn. Cause asks which criterion preceded the exit. Controlled probes that change one criterion at a time can then test whether that criterion caused the displacement. Stability asks whether the same journey produces the same outcome when repeated.

AIVO's next study applies this to journeys in which the user's criteria become progressively more specific. The question is whether the gap between being known and being recommended persists, widens or closes as the recommendation takes shape.

When decisions can form across turns, measurement has to follow the journey instead of sampling its opening.

References

AIVO Standard, WP-2026-14, Zenodo. DOI to be added.

Sheals, P., What a Citation Is Worth, v1.1, AIVO Standard, forthcoming on Zenodo.

Tannenbaum, B., The Prompt Is Not the Query: How Request State Evolves Across Multi-Turn AI Conversations, arXiv 2607.22392, 2026.

Tannenbaum, B., Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers, arXiv 2608.02556, 2026.

Laban, P., Hayashi, H., Zhou, Y. and Neville, J., LLMs Get Lost in Multi-Turn Conversation, ICLR 2026.

AIVO Standard working papers are DOI-anchored on Zenodo and are not peer-reviewed.

AIVO Meridian