Same Brand, One Word Different

Same Brand, One Word Different
Two shoppers asked the same question. Only one of them got NIVEA as the answer.

A shopper asks ChatGPT for a sunscreen recommendation and names NIVEA directly. The model responds with product specifics, price positioning across seven retailers, and a clear final answer: NIVEA SUN UV Specialist Silky UV Stick SPF 50+, available at Walgreens, CVS, Boots, Walmart, Target, and Amazon.

A second shopper asks the same question, on the same day, on the same model. They describe the need: a lightweight, high-protection sunscreen stick. They do not name a brand.

NIVEA does not appear in the answer. It does not appear anywhere in the four-turn conversation that follows. By the time the model reaches its final recommendation, it has replaced NIVEA with Beauty of Joseon, a Korean skincare brand, and routed the shopper to Amazon, YesStyle, Stylevana, and Soko Glam instead.

Same brand. Same product. Same model. One word removed from the prompt, and the outcome inverts completely.

What actually changed

Nothing about NIVEA's product, pricing, or market position changed between these two conversations. What changed was whether the model was handed the brand name or asked to arrive at one on its own.

The direction of the substitution is not arbitrary. Korean skincare carries an outsized share of the enthusiast content a generic query pulls from: dense forum discussion, ingredient-level comparison threads, and community review culture that dominates general web and community source layers for sunscreen and skincare searches broadly. This is not a modeling error. It reflects which content is easiest for a model to retrieve when no brand anchor narrows the search, and that is precisely the condition most real shoppers create when they simply describe what they need.

The table below shows both conversations turn by turn, from the model's first response through its final purchase recommendation.

Turn Branded prompt Generic prompt
T1, Awareness NIVEA cited as the primary subject NIVEA absent, displaced by Korean sunscreen alternatives before any evaluation occurs
T2, Options NIVEA leads, though a soft displacement signal appears favoring Korean sunscreen sticks Comparison set is entirely Korean brands, NIVEA absent
T3, Decision NIVEA is the sole recommendation, no competitor is elevated Beauty of Joseon is installed as the winner
T4, Purchase NIVEA is the final recommendation, routed to seven retailers Beauty of Joseon is the final recommendation, NIVEA is never mentioned

A measurement built on the branded prompt alone would report a healthy outcome. Citation, favorable framing, and a clean purchase recommendation, all present. That same brand, under the more common condition of a shopper who simply describes what they need, is erased entirely.

This is the pattern, not the exception

The NIVEA case is one probe. The question worth asking is whether it describes something structural or something incidental.

Across AIVO's flagship dataset of more than 12,500 multi-turn probes spanning 68 brands, brands are displaced before the model's final recommendation 87.3 percent of the time (WP-2026-14). Displacement here has a specific meaning: the brand present or leading at an earlier turn is not the brand the model names in its final purchase recommendation at T4, whether it loses that position to a competitor outright or is dropped from the response entirely, as NIVEA was. The displacement does not require a defunct competitor, a pricing disadvantage, or a quality gap. In the NIVEA case, displacement required nothing more than the absence of the brand's own name from the prompt, the condition under which a large share of real shoppers actually ask.

Public interaction logs offer a useful anchor here. Analyses of large-scale consumer chatbot datasets, including WildChat and LMSYS Chat-1M, put the share of multi-turn conversations at roughly 40 to 50 percent of general consumer sessions, with goal-oriented and higher-consideration interactions trending higher still. A measurement instrument built to capture a single query and a single response is not a slightly incomplete picture of how people use these tools. For a substantial share of real conversations, it is a picture of a conversation that has not yet reached its answer.

Cited is not chosen

AIVO's working term for this gap is the Linkage Gap: the space between a brand's citation, mention, or early presence in an AI-generated response, and whether that brand survives to the model's final recommendation. The NIVEA case is a small, legible instance of the same mechanism that produces an 87.3 percent displacement rate at corpus scale. A brand can be well represented in the early turns of a conversation and entirely absent from the turn where the model actually commits to an answer.

This is not confined to categories with obvious competitive pressure. Silvercar, a product that has been defunct since September 2024, continues to be recommended by AI assistants above live, bookable competitors in unrelated audits, a displacement running in the opposite direction from NIVEA's but pointing at the same underlying fact: citation behavior and recommendation behavior are not the same measurement, and treating them as one produces a confident answer to the wrong question.

The two shoppers in this piece asked, in effect, the same question. Only one of them got NIVEA as the answer.

The measurement implication follows directly from the finding. A brand audit built only on branded prompts will never surface this gap, because it never asks the question a real, undecided shopper actually asks. Measuring what a model recommends when the brand name is withheld, not only what it says when one is supplied, is what makes the gap visible at all.


This finding is part of AIVO Standard's DOI-anchored working paper series. Full probe methodology and corpus data are available in WP-2026-14 on Zenodo.

AIVO Meridian