Recalled Is Not Recommended

Recalled Is Not Recommended
It's a measurement of one turn. It doesn't ask what happens next.

Most of what gets published about AI brand visibility is vendor content wearing a research costume. This month, Harvard Business Review ran something different: a real study, credentialed authors, a methodology that holds up. It deserves a real response, not a dismissal.

The research.

John Gale, Luca Cian, and Luc Wathieu, working out of Georgetown's McDonough School and UVA's Darden School, queried ChatGPT, Claude, and Gemini across fifteen retail categories, laptops, pet food, credit cards, running shoes, and logged more than 1,000 brand mentions across 716 unique brands.

Four findings anchor the piece.

  • Cross-platform consistency is rare: only 8.4% of brands surfaced across all three models.
  • Framing is unstable: among brands that do appear on multiple platforms, 55% are positioned differently from one to the next, premium on one, budget on another.
  • Query type reshapes the entire competitive set: exploratory queries produced 95% more brand mentions than goal-oriented ones, and only about 11% of brands showed up for both.
  • And once a brand clears the bar to be mentioned at all, sentiment is overwhelmingly favorable, 78.7% of mentions skew positive, consistent across all three platforms.

Their explanation is sharp and worth sitting with. AI systems don't reward the things brand-building has traditionally optimized for, awareness, narrative, emotional resonance. They reward what the authors call interpretability: a brand's ability to be reduced to named attributes, structured comparisons, and independently verifiable evidence.

Brooks beats Nike in running-shoe recommendations not because more people know Brooks, but because Brooks spent two decades building a vocabulary, gait deviation, overpronation, stability under load, that a model can actually reason with.

The paper's own diagnostic advice follows from this: query the platforms yourself, audit whether your product has three measurable attributes a model could name, map your third-party evidence, and invest in shaping the vocabulary customers use to describe their problem in the first place.

None of that is wrong. It's a genuine advance on the "just be visible" thinking that still dominates most AI-marketing content, and it deserves to be taken seriously rather than folded into the pile of vendor whitepapers dressed up as research.

Where the map stops.

Read the methodology closely and the boundary is explicit, not hidden. Every probe in the study is a single prompt, a single response. What gets measured is whether a brand is mentioned, and how it's framed, in that one shot. The paper's own closing line states the finding directly: once a brand is recalled as a candidate, tone is almost always favorable, and competition is decided upstream, in whether the brand gets included at all.

That's a real and useful measurement. It's also a measurement of one turn. It doesn't ask what happens next.

We've published on this before.

In The Zero-Click Blind Spot, we wrote about the measurement layer that goes missing when a recommendation drives revenue no analytics platform was built to see.

In The Committee You Can't See, we wrote about how a buying decision that looks resolved at one point in a conversation gets quietly re-litigated by the next person, or the next turn.

The pattern in both pieces is the same one this new research doesn't test: a brand's status at the first mention and its status at the moment a decision actually gets made are two different measurements, and the gap between them is where deals are won or lost.

Across the audits we run, brands routinely open strong, present at the first turn, favorably framed, exactly the pattern this study documents, and then lose primary status by the second or third turn of an actual multi-turn conversation, once a model starts applying comparison criteria a single-prompt study never reaches.

Averaged across the brand probes in our own dataset, roughly 87% of brands present at the opening turn of a buying conversation are displaced before the model reaches a final recommendation. Inclusion and recommendation are correlated. They are not the same event, and the size of the gap between them is the whole story a single-prompt methodology can't tell.

This is also, worth noting, not an isolated study. The same author cluster behind the "share of model" metric this piece builds on, D. Dubois specifically, published research on luxury brand misreading in AI systems just weeks earlier. A real research program appears to be forming around AI-mediated brand measurement this summer, from more than one direction.

That's a good thing. It also means the field needs the layers named precisely, or the same confusion that already exists around "visibility" is going to reappear one level deeper, this time under the label "interpretability."

The second gap: diagnosis without remediation.

The paper's prescribed fix is a manual audit: query the platforms, check your attribute structure, map your third-party validation, invest in problem literacy. That's sound advice as far as it goes, and it's not nothing, most brands haven't done even this much.

The paper stops at diagnosis

It doesn't measure whether a given fix actually changes a model's behavior once applied, doesn't identify which specific missing piece of evidence is causing a specific displacement at a specific turn, and doesn't re-test after the fix to confirm the gap actually closed rather than just plausibly should have.

That last part matters more than it sounds. Publishing a stat, a certification, a named authority source is not the same as confirming that a model has actually incorporated it into how it reasons about your brand. The only way to know is to probe again after the fact, the same way you probed before, and check whether the specific turn where displacement used to happen still displaces.

Interpretability is a real property to build toward. Whether a specific brand has actually built enough of it, on a specific platform, at a specific stage of a specific buying journey, is an empirical question, and right now, for almost every brand, it's an untested one.

What this is and isn't.

This piece is a real, well-designed study of one important layer, and its finding, that structured, evidenced brands get included more consistently than brands leaning on symbolic positioning, is very likely correct and worth acting on. It answers whether a brand shows up. It doesn't answer whether a brand survives. Both questions matter. Only one of them has been measured here.


Sources: Gale, J., Cian, L., & Wathieu, L. (2026). "How to Get AI to Surface Your Brand." Harvard Business Review, June 29, 2026.

AIVO Meridian