New Product Development Was Built for a Judge That No Longer Decides
AIVO Journal — Weekend Essay
Every new product development process in consumer goods rests on an assumption so old it stopped being examined decades ago: that the entity evaluating a new product is a person who can be moved by narrative.
Concept tests measure whether a description resonates. Focus groups measure whether a story lands in a room. Brand tracking measures whether values — sustainability, heritage, craft, purpose — shift sentiment over a quarter. All of it optimizes for the same target: a human being who weighs a product partly on functional merit and partly on how the brand makes them feel about themselves for choosing it.
That target still exists. But it is no longer the only judge, and for a fast-growing share of category-defining purchase decisions, it isn't even the first one.
The New Judge Doesn't Read Values Statements
When someone asks an AI system what's actually good in a category right now — the best K-beauty sunscreen, the right sensitive-skin moisturizer, the antifungal that actually works — the model isn't retrieving a brand's positioning. It's reasoning across ingredient claims, comparative reviews, third-party testing data, and the structured evidence it can find and weigh against competitors, turn by turn, until it converges on a recommendation.
Nowhere in that process is there a slot for "we've been trusted since 1957" or "we believe in inclusive beauty for everyone." Those are values statements, built for a human who wants to feel good about a choice. A reasoning model doesn't feel good about anything. It doesn't have a self-concept a brand can flatter. It has evidence, structure, and a comparison to make — and it makes it whether or not the brand has prepared for it.
This is the finding underneath AIVO's core research: across more than 12,500 multi-turn brand probes, the large majority of brands present when a comparison opens are displaced by the time the model reaches its actual recommendation — regardless of how strong their underlying brand equity is. Presence and equity get a brand cited. They do not get it chosen. Cited is not chosen.
It's worth being precise about what kind of problem this is, because it is easy to mistake for a more familiar one. A great deal of current industry attention goes toward getting a brand mentioned by an AI system at all — appearing in the citation, showing up in the first answer, being part of the set a model considers. That is a real and worthwhile problem, and it is not the one this essay is about. Getting cited is necessary and, on its own, insufficient. The harder, less-examined problem sits one step further down the reasoning chain: once a brand is already in the conversation, does it survive the model's actual comparison to the point of being the thing recommended? Those are two different failure points, measured differently, solved differently, and a brand can be excellent at the first and still lose the second every time.
Why Traditional NPD Testing Can't See This Coming
Here is the uncomfortable part for anyone running an NPD pipeline today: a new product can pass every stage of traditional testing and still be structurally unable to survive an AI-mediated comparison, because the two processes are evaluating different things.
A concept test asks: does this resonate with a target consumer reading a description? An AI reasoning chain asks: given everything available about this product and its competitors, which one best satisfies the stated need? The first is about emotional and narrative fit. The second is about evidentiary sufficiency — does the product have specific, comparable, checkable claims that hold up against alternatives the model already knows about?
A product can win the first test decisively and fail the second one completely. We saw this recently with a well-known personal care brand's new Korean-formulation sunscreen stick — a smart, on-trend move by any traditional read of the market. Ask an AI system generically about what's good in Korean-style sunscreen, and the brand doesn't surface at all; the field belongs to newer, AI-legible challengers with sharper, more specific evidence trails. Ask about the brand by name, and it opens the conversation — then gets quietly reframed as the budget fallback within a turn or two, once the model weighs it against alternatives with more specific, comparable claims. The formulation decision was right. The traditional read of the opportunity was right. And the product still doesn't survive the conversation that, increasingly, decides the sale.
No focus group would have caught this, because a focus group was never testing for it.
What Testing Actually Needs to Measure Now
If the model reasoning through the comparison doesn't care about a positioning narrative, then testing a new product's readiness has to move upstream of narrative entirely — to whether the product has evidence a reasoning chain will actually weigh, and whether that evidence survives contact with named competitors.
That reframes new product development around two questions traditional research doesn't ask:
Does the product have a defensible position inside the reasoning chain, not just the market? This means probing how AI systems actually discuss the category — generically and by name — before launch, not after. It means finding out whether a proposed formulation, claim, or positioning line gives a model something specific enough to hold onto when a stronger competitor claim appears in the same conversation, rather than discovering that gap post-launch through declining conversion.
Where does the category actually have unclaimed ground? This is the more interesting question, because it turns the same probe data that reveals a brand's weakness into a genuine opportunity-finding tool. Running the same reasoning-chain analysis across a category doesn't just show where an existing brand is losing — it shows where an entire sub-niche has no brand winning convincingly at all. Not contested share. Unclaimed demand nobody has an established, AI-legible answer for yet. That is a fundamentally different kind of white space than anything a traditional gap analysis produces, because it's derived from where the actual decision-making system currently has no confident answer — not from where consumers say they wish something existed.
The Values Question, Reframed
None of this means brand values stop mattering. They still shape culture, hiring, loyalty, and plenty of purchase decisions that never touch an AI system at all. But treating values as the primary lever for winning a new product's way into the market assumes the decisive judge is a person who can be persuaded by identity and story. Framed that way, a growing share of category comparisons are being conducted by a judge for whom that lever doesn't exist, and the market has not yet built a corresponding practice for it.
New product development that only tests against the old judge will keep producing products that are well-loved and quietly unrecommended — decisions right on paper, invisible in the room where a growing number of purchases now actually happen. The organizations that adapt first won't be the ones with the best story. They'll be the ones who found out, before launch, whether the story was ever going to be heard at all.
AIVO Journal publishes ongoing research and commentary from the AIVO Meridian team on AI-mediated brand and purchase decisions. This piece is a working essay, not a peer-reviewed paper.