The AI Reckoning Is Not Sector Specific
AI is not creating disloyal customers. It is removing the cost of reconsideration.
AIVO Meridian has now run The AI Reckoning across insurance, banking, telecoms, wealth management, travel, SaaS and payments. Seven sectors, seven regulators, seven different economic models. Underneath all of them sits the same mechanism.
For two decades these industries built enormous value on a simple behavioral fact: reconsidering a decision is expensive, so most people don't. The insurer's renewal book, the bank's primary relationship, the telecom's out-of-contract base, the SaaS vendor's installed seat count, the payment processor's embedded integration. None of it was pure loyalty. Much of it was friction wearing loyalty's clothes, quietly priced into retention forecasts without ever being called by name.
AI is now the cheapest tool ever built for removing that friction. It does not need to persuade anyone to leave. It only needs to make it easy enough to ask, in passing, whether staying still makes sense, a smaller ask than switching, and a far more common one.
Friction disappears
In UK motor insurance, attrition runs 49 to 54 percent for policies without auto-renewal against 33 to 35 percent with it, a gap explained almost entirely by how much active effort it takes to reopen the decision, not by any difference in the underlying product. The same shape recurs everywhere this research has looked. US banking customers are moving balances away from their primary institution more often year over year, a pattern JD Power now calls soft switching precisely because the account itself rarely closes. UK telecom customers openly admit the available savings aren't worth the time spent searching for them, even where regulators have made switching itself almost frictionless. B2B software buyers report switching vendors mid-evaluation because a chatbot's guidance sent them somewhere they hadn't originally planned to look.
None of this proves AI causes churn, and none of the underlying briefs claims it does. What it shows is where the exposure actually sits: not in dissatisfaction, but in the gap between genuine preference and friction that has simply never been tested. Every board in every one of these briefs is being asked some version of the same question. How much of what we call retention is actually just a decision nobody has reopened yet, and how would we know the difference before a competitor's AI system finds out first?
Recommendation matters more than visibility
A second, harder finding emerges from controlled testing across those same sectors. Appearing in an AI answer and winning the AI's recommendation are different events, and the gap between them is large. In one cross-category study of over 1,400 probes, 87 percent of brands recognized at the first turn were displaced before the model reached its final answer.
A brand can be named, known and factually correct in a model's memory and still lose the recommendation. First-prompt visibility, the metric most of the GEO and AEO industry has organized itself around, explains only about a third of the variance in who actually gets chosen. The rest happens downstream, as the model weighs price against coverage, reputation against convenience, one plausible option against another, in exactly the stage where almost nobody in the visibility-tracking industry is currently measuring anything at all. A dashboard that reports strong first-answer presence while the customer is quietly being talked into a competitor is not an early-warning system. It is a false sense of security with a chart attached.
Measurement is unstable
A third finding should worry any board relying on a single AI-visibility snapshot: the systems being measured are not stable. In replication testing, identical prompts run under identical conditions produced meaningfully different outcomes roughly one time in five, and even the opening shortlist a model offers can shift from one otherwise identical run to the next. A single prompt shows what happened once. It cannot say whether the market moved, whether a competitor made a genuine gain, or whether the model simply landed differently that afternoon. Yet a single unreplicated prompt, run once and reported as fact, is still the basis on which most companies currently decide whether they have an AI problem at all.
Markets can tolerate imperfect measurement. They cannot tolerate incompatible measurement, and once automated systems begin executing commercial decisions at scale rather than merely informing them, measurement stops being an analytical question and becomes market infrastructure.
Governance therefore becomes necessary
Put those three findings together and the shape of the real risk becomes clear. Decisions that used to stay closed are being reopened at scale, by systems whose behavior is stochastic, using a shortlist that only weakly predicts the final answer, while the industry's dominant response has been to measure the one signal least connected to commercial outcome.
That gap is closing faster on the infrastructure side than the measurement side. Stripe, Mastercard, Visa, Adyen and Worldpay have each shipped agentic-commerce protocols in the past year, and Google now routes hotel bookings directly inside AI Mode. Each is serious infrastructure, built by capable engineering organizations moving quickly toward a real opportunity. None of it is a shared standard for how a brand's performance inside these systems should be measured, how a poor outcome should be remediated, or how the resulting transactions should be attributed back to the systems and evidence that produced them. The rails are being built. The instrumentation for what runs on them is not.
This is the pattern the open web already went through twice, first with search and again with programmatic advertising. Infrastructure arrives before governance, incompatible measurement conventions multiply because each vendor's tooling was built to answer its own product question rather than the market's, and the industry spends years afterward reconciling numbers that were never designed to agree. GEO and AEO are already showing the early symptoms. Dozens of vendors report visibility scores built on incompatible, unreplicated methodologies, most proprietary, few connected to any verified commercial outcome, and a brand can score well on three dashboards and poorly on a fourth without anyone able to say which one is right.
If agentic commerce scales the way its current infrastructure suggests, that disorder will not stay confined to marketing dashboards. It moves into real transactions, executed by agents, at machine speed, often with no human in the loop at the moment of decision. A pricing error, a stale product fact, or a misapplied fee that today might cost one lost customer could, inside an agentic channel, be replicated across thousands of transactions before anyone notices the pattern. Governance is not a compliance nicety layered on top of that shift. It is the difference between an error that costs one sale and an error that costs a quarter.
This is not a gap only one organization has noticed. Several long-established brand and research consultancies are now moving to extend their existing annual-ranking businesses into AI, packaging a flagship AI ranking alongside paid diagnostics and sponsorship tiers.
The instinct is understandable, and in most cases well intentioned. It is also, on its own, a category error. A ranking methodology built for annual brand-value surveys was designed to answer a different question than the one AI-mediated commerce is now asking. Valuing a brand once a year from a defensible but static formula is one discipline. Measuring whether a model recommends that brand under real constraints, replicated enough times to separate genuine movement from stochastic noise, with a closed loop back to remediation and verification, is another discipline entirely. Adding "AI" to an existing ranking product does not close that gap, any more than adding a mobile app closed the gap between print circulation and digital publishing two decades ago. The organizations most eager to claim authority here are, in several cases, the least equipped to build the underlying measurement themselves, which is exactly why the standard the market needs cannot simply default to whichever incumbent monetizes fastest.
Why AIVO is convening rather than competing on this point
This research gives AIVO a strong commercial position inside this shift. It also makes clear that no single vendor's proprietary scorecard can become the standard the market needs, AIVO's own included. A standard that only one company controls is not a standard. It is a walled garden with better marketing, and the market has enough experience with walled gardens by now to recognize one quickly.
That reasoning is behind current outreach to leaders across holding companies, research and measurement bodies, regulators and standards organizations in the UK and US. The conversations are deliberately broader than any single commercial relationship. The goal is not to sell a dashboard, and it is not to recruit distribution partners for one company's own rankings. It is to build consensus among organizations who will otherwise each build an incompatible private answer to the same three questions, around requirements every one of the sector briefs above independently arrives at without being asked to.
Measurement that goes beyond first-prompt mentions to the full decision journey, replicated enough to separate genuine movement from noise, and reported against a defined noise floor rather than a single run.
Remediation that closes the loop rather than stopping at diagnosis: identify what needs to change, make the change, rerun the same controlled journey, and verify the outcome improved beyond measurement noise.
Attribution that can credibly connect an AI-mediated decision back to the system and evidence that produced it, so traffic and transactions arising from agentic and conversational AI can be counted on terms the market actually agrees on. The AI Traffic Attribution Convention, currently under review with the Media Rating Council and IAB Tech Lab, is offered into this effort as a starting specification, not a finished answer.
The convening has to happen now, before agentic commerce protocols finish hardening around whatever conventions get adopted first by default. Standards set early are cheap to align around, because nobody has yet built a business case around the alternative. Standards set retroactively, after billions of dollars of transaction volume are already flowing through incompatible systems, are not; by then every participant has a commercial reason to defend the convention they already built, and consensus becomes a negotiation between entrenched interests rather than a shared design exercise. That is the difference between setting a standard and litigating one, and it is largely a matter of timing rather than persuasion.
The question underneath all seven briefs
Every AI Reckoning brief published this year ends with a version of the same question, asked of a different industry: how much of what looks like loyalty is actually friction that AI is about to remove, and could the board prove the difference before it shows up as lost revenue rather than a hypothesis? Underneath all of those sits one question for the market as a whole, not just for individual brands inside it. When AI agents are making and executing these decisions at scale, on whose terms will that be measured, corrected and counted? Nobody should assume the market answers that question well by default. The early evidence from GEO and AEO, fragmented methodologies, unreplicated scores, no shared attribution convention, suggests it is currently answering it badly, and the cost of that disorder rises every quarter more commerce moves through agentic channels.
The internet eventually converged on shared standards for search, advertising and web analytics because markets cannot function indefinitely without agreed ways of measuring performance. That convergence took years of fragmentation and disputed numbers, and arrived later than the industries involved would have preferred. AI-mediated commerce is approaching the same moment, on a faster timeline, with more money moving through fewer clicks and far less human review at the point of decision. The organizations that help establish those conventions now will shape not just how brands are evaluated, but how billions of future purchasing decisions get understood, governed and trusted. AIVO intends to be part of that room, and is inviting others, competitors included, to be there too.
AIVO Journal is the editorial arm of AIVO Standard, AIVO's DOI-anchored research programme. This article draws on AIVO Meridian's 2026 "AI Reckoning" board briefing series covering Insurance, Banking, Telecoms, Wealth & Asset Management, Travel & Hospitality, SaaS & Enterprise Technology, and Payments, together with AIVO Standard Research findings including "When AI Becomes the Decision-Maker" and "Can Trust Be Optimised?" (September 2026). See the individual briefings for full source notes. These findings are dataset-specific and DOI-anchored; they should not be read as universal constants across every brand, category or model.
Comments ()