Reputation Isn't a Snapshot. It's a Conversation.
A growing number of reputation measurement tools now claim to tell you what AI thinks of your company. That is progress. Until recently, most communications teams had no visibility at all into how generative AI models describe their institution to the people asking about it.
But the way this new category is being built matters as much as the fact that it exists, and right now, most of it is being built on a method borrowed from a world that no longer describes how people actually use AI.
RepTrak's AI as a Stakeholder, launched as part of its Compass platform, applies the same reputation, driver, and factor questions used in human survey panels directly to AI platforms, then sits the results alongside existing customer, employee, and investor data. It is a real advance over ignoring AI entirely, and the findings are worth taking seriously.
In one study across four major banks, RepTrak found that AI-generated reputation scores ran nearly 23 points lower on average than scores from the informed general public, with the gap concentrated in Conduct and Products & Services, while Innovation actually scored higher through AI than through human respondents. That is a valuable finding.
It is also, by construction, a single frame. One question, asked once, answered once, recorded as the AI's position on the company. It treats a generative model's response the way a survey treats a respondent's answer: fixed, final, and representative.
The distinction that matters is this one. RepTrak asks an AI model what it thinks and records the answer. AI RQ asks, then asks again, because a model's second answer is where its actual position shows up.
Anyone who has used ChatGPT, Gemini, Perplexity, or Grok to research a company knows that is not how these systems behave. A model's first answer is rarely its last word. Ask a follow-up. Push back. Ask what the drawbacks are, or how it compares to a competitor, and the answer moves. Sometimes it strengthens. Sometimes it reverses entirely. A single-turn measurement cannot see any of that, because it never asks the second question.
What a Four-Turn Conversation Reveals That One Question Cannot
AIVO's AI RQ methodology tests each institution across a four-turn conversation, not a single prompt, because reputation inside an AI system is not a static value. It is a trajectory.
We ran this against Bank of America and JPMorgan Chase, and the shape of the results makes the case better than any argument for the method could.
Bank of America was nearly invisible at the opening turn, absent from the AI's unprompted shortlist on three of four platforms tested. A single-question audit would likely have recorded that low visibility and stopped there. But across the conversation, Bank of America's Ethics score rose from 68 to 83, the only dimension that strengthened rather than faded as the conversation continued. Ask a model directly about Bank of America's conduct and standards, and it becomes more favorable, not less.
Its Products & Services score moved in the opposite direction. It opened at 67, held up while the model was directly defending the institution under questioning, and then dropped to 50 by the final turn, the point at which the model actually recommends a bank, and defaulted instead to Chase or Capital One. Present in the conversation. Positively framed for most of it. Still displaced at the moment that matters.
JPMorgan told a different story entirely, leading Bank of America on Trust, Vision, and Products across every platform tested, with a final AI RQ of 75.
None of that is visible in a single-question snapshot. A one-shot audit would have captured a moment, not a pattern. The pattern is the finding.
Reputation and Recommendation Are Not the Same Variable
There is a second distinction that matters here, and it goes beyond turn count.
Diagnosing how AI perceives an institution is necessary, but it is not the same as knowing what that perception does at the point a model actually makes a recommendation. A company can score well on every reputation driver an AI model is asked about and still lose the final recommendation to a competitor.
Bank of America's own results make the point directly. Its Products & Services language was favorable for most of the conversation, defended by the model under direct questioning, present, specific, and positively framed. None of that prevented the model from naming Chase and Capital One at the turn that actually counted. A reputation audit measuring sentiment alone would have scored that exchange as a success. It was a loss.
This is why AI RQ was built to sit alongside, not apart from, recommendation-stage measurement. A reputation score that cannot be connected to whether it actually shaped the outcome tells you what the model said. It does not tell you whether what the model said mattered.
Measurement Without a Response Is Half a System
The other structural gap in single-question audits is what happens after the score is delivered. A diagnostic tool that identifies a 22-point reputation gap and stops there has told a communications team something is wrong without giving them a way to know whether anything they do about it works.
AIVO's approach pairs AI RQ measurement with a remediation process built on the same Meridian infrastructure used to track recommendation-stage displacement: when a gap or a late-stage substitution is identified, the underlying cause is addressed, and the institution is reprobed on a regular cadence to confirm the correction holds. Reputation inside an AI system is not a fixed asset that gets audited once a year. It shifts as models update, as sources change, and as competitors act. A measurement approach that does not reprobe cannot tell the difference between a problem that was fixed and a problem that was simply not asked about again.
What This Means for Communications Teams
The emergence of AI as a stakeholder is not in question. Every serious reputation measurement approach now agrees on that much. The open question is what measuring it well actually requires.
A single question, asked once, can tell you what an AI model says about your institution today. It cannot tell you whether that answer holds up under scrutiny, whether it survives to the point of an actual recommendation, or whether a correction made last quarter is still working this quarter. Those are the questions comms and reputation leaders are increasingly going to be asked to answer, and they require a method built for a conversation, not a poll.
The cost of the gap is not abstract. A team working from a single-question audit can fix what the score told them to fix and have no way of knowing, until the next annual review, whether the model's position moved, held, or quietly reversed. A team working from a conversational, closed-loop measure knows within one reprobing cycle. In a stakeholder relationship that updates as often as a model's training and retrieval do, that difference compounds every quarter it goes unaddressed.
AIVO's AI RQ Banking Index, expanding on the Q1 2026 Global Banking AI Decision Index, publishes at the end of September.