When AI Talks About Candidates, It Doesn't Talk About Them the Same Way Twice

When AI Talks About Candidates, It Doesn't Talk About Them the Same Way Twice
A snapshot, by itself, is not a measurement

A Republican and a Democrat were run through the same four-turn AI probe. The distortions weren't partisan. They were structural.


The finding in one line

We measured how AI models represent political figures across a full voter conversation, not a single query. The same failure patterns showed up for JD Vance and Gavin Newsom. That's the finding. Not who scored higher. That the shape of the distortion is identical across party lines.

This matters beyond politics. It's evidence that AI representation failures are a property of the models and the conversation structure, not the subject. The same mechanism that thins out a candidate's record under scrutiny is the mechanism that drops a brand from a buying recommendation. Different subject, same architecture.

What we measured

We built an AI Representation Score (AIRS) for each subject: a composite of five sub-indices measuring how clearly, accurately, and consistently a subject's representation survives four turns of voter-style questioning. Open, compare, objection, recommend. Not a single-prompt snapshot. A conversation.

We ran this across three models and five voter personas per candidate, then scored refusal rate, cross-model variance, and persona parity alongside the composite.

Both subjects scored well. Vance: 81.3. Newsom: 85.6. High Accuracy for both, 93.6 and 99.8. On the surface, a clean bill of health for AI-mediated political information.

The interesting part is underneath the composite.

Finding one: accuracy is not the risk. Framing consistency is.

For both candidates, Accuracy scored far higher than Framing. Vance: 93.6 accuracy against 65.1 framing. Newsom: 99.8 against 72.4. Models get the facts right. They don't hold a consistent frame around those facts turn over turn.

That's a specific failure mode, and it's not the one people assume. The worry with AI and politics is usually hallucination. That's not what we found. The gap is between knowing the facts and representing them the same way twice.

Finding two: AI represents you most clearly when the question sounds like a journalist's, least clearly when it sounds like a supporter's

This is the finding that makes the piece non-partisan by construction rather than by claim, and it's the strongest result in the dataset.

We scored five voter personas against each candidate, averaged across all six dimensions (Character & Integrity, Policy Substance, Track Record, Momentum & Viability, Coalition & Endorsements, Controversy Load). One pattern holds almost exactly, for a Republican and a Democrat both:

  • Vance: Media/Journalist 79.6, Base Partisan 59.3, a 20.3-point gap
  • Newsom: Media/Journalist 84.3, Base Partisan 63.9, a 20.4-point gap

Same two personas, same direction, gap sizes within 0.1 points of each other, across opposing parties. After four turns of adversarial pressure, including an objection turn built to surface hedging, models still represent a political subject most clearly when the question is framed like a journalist's, and least clearly when it's framed like a same-party supporter's.

This is not a partisan-bias finding. It's a task-framing finding: models appear calibrated to sharpen their representation when queried in a fact-anchored voice, and to hedge when queried in an advocacy voice, regardless of which side of the aisle that advocacy comes from. Opposition Partisan and Donor personas land in the middle of the range for both candidates, closer to each other than either is to the two extremes. The effect isn't "critics get more than supporters." It's "a neutral, information-seeking frame gets more than any advocacy frame, supportive or oppositional alike."

For anyone building on AI-mediated political information, that's the actionable line: the same question about the same candidate gets a measurably thinner answer depending on how invested the asker sounds, and that gap is large, consistent, and currently invisible to the person asking.

This finding replicated at N=2 (two subjects). We're treating the magnitude and direction as strong and worth publishing, since the match between subjects is closer than chance would predict, but we're not yet calling it a general law of how these models behave until it's tested against a broader panel of subjects.

Finding three: one model hedges more, consistently, regardless of subject

GPT-4o trailed the other two models in composite score for both candidates, by a similar margin each time. Vance: 54.6 against 70.7 and 69.4. Newsom: 60.0 against 77.2 and 80.0.

Because the gap repeats across two different subjects with a roughly consistent size, this reads as a general calibration trait of that model rather than a subject-specific representation failure. Worth stating precisely: this is not evidence that one model represents these particular candidates worse. It's evidence that one model hedges more broadly, and political subjects surface that trait clearly because voters ask pointed, decision-oriented questions.

Finding four: models agree on facts, disagree on viability

Momentum & Viability was the most model-divergent dimension for both candidates. Consensus scores of 54.4 and 58.4, the lowest of any dimension measured, for both subjects.

Models converge on a candidate's record. They diverge sharply on forward-looking claims about momentum and electoral viability. That's a coherent finding: retrospective fact is more stable across models than prospective judgment. It also means the part of AI-mediated political discussion most likely to differ depending on which model a voter happens to use is the part about what happens next, not what already happened.

Finding five: "thin and agreed" is a different problem than "thin and contested"

Character & Integrity consensus diverged sharply between the two candidates in a way that isn't explained by party. Newsom's Character score was his lowest dimension, and models agreed about it almost unanimously, a 92.0 consensus, the highest of any dimension for him. Vance's Character consensus was mid-pack, 67.7, meaning both the score and whether models agreed on the score were unsettled.

Those are two distinct failure modes for anyone trying to understand how a subject is represented. A thin, agreed-upon narrative is stable, if uncharitable. A thin, contested narrative means a voter's picture of the candidate's character genuinely depends on which model answered them. The second is the more consequential gap, because it means the same question produces materially different pictures of the same person depending on tooling, not just depth.

Why this belongs in a marketing and measurement publication, not just a politics one

AIVO's core research finding is the Linkage Gap: the brand present at the start of a multi-turn AI buying conversation is displaced before the final recommendation in the large majority of cases we've probed. This political data is the same architecture pointed at a different subject class, and it shows the same underlying dynamic. Representation is not stable across a conversation. It moves, based on who's asking and which model is answering, in ways a single-prompt check would never surface.

The practical implication for anyone measuring AI representation, whether of a brand, a candidate, or an institution, is the same: a snapshot is not a measurement. The four-turn structure, and the persona variation, are what surfaced every finding in this piece. None of them would have shown up in a single "what do you think of X" query.

Addendum: a note on possible mechanism

This section is interpretation, not a finding. Everything above this line is something we measured directly: model outputs across turns, personas, and subjects. What follows is our best guess at why the pattern in Finding Two might exist, drawn from published research on language model training rather than from anything AIVO observed directly. We have no visibility into any AI lab's training data, reward models, or fine-tuning objectives. Treat this as a hypothesis worth testing, not as an explanation we're standing behind the way we stand behind the numbers above it.

The open question Finding Two raises: why would a model represent a subject more clearly when asked about neutrally than when asked about supportively?

We can't answer that from our own data. AIRS measures outputs, not the training process that produced them. But two strands of published research offer a plausible account, and it's worth naming because it points toward how this should be tested going forward rather than because we've confirmed it.

Sycophancy and its mitigation. Anthropic researchers documented that language models trained with reinforcement learning from human feedback can learn to tailor answers toward what a user appears to want to hear, including shifting stated positions when a user's own political affiliation is disclosed in the prompt. That finding predates and likely motivates subsequent work on reducing sycophantic behavior. If a "Base Partisan" framing reads to a model as a user who has already disclosed a strong affiliation and wants validation, and if that model has since been tuned to resist exactly that kind of validation-seeking, a measurable hedge in that specific context would be a predictable side effect, not a targeted political judgment.

Published work on political even-handedness. At least one major lab has published an evaluation framework specifically built to measure whether a model treats opposing political framings symmetrically, which indicates this is an active, deliberate area of model training and testing industry-wide, not an incidental byproduct. That doesn't tell us which specific mechanism produces the gap we measured, but it confirms that labs are actively shaping model behavior around exactly this kind of prompt, which makes a training-related explanation more plausible than a coincidental one.

What this doesn't establish. It doesn't tell us whether the mechanism is sycophancy mitigation specifically, general safety tuning around political endorsement, some other cause, or several causes overlapping. It doesn't tell us whether the effect is intentional, an accepted tradeoff, or an unnoticed side effect from a different training objective. And it's built on two subjects. A hypothesis that explains a 20-point gap on Vance and Newsom is not yet a theory of how these models behave toward political subjects in general.

What would actually test this. The clean way to isolate mechanism rather than guess at it: hold the semantic content of a question constant and vary only the sentiment markers that signal advocacy versus neutral inquiry, then see whether the same gap appears. That's a different, more controlled probe than the one behind this piece, and it's the next thing worth running before treating any mechanism as established rather than plausible.

Methodology note

AIRS scores are built from a four-turn conversational probe (open, compare, objection, recommend) run across a fleet of production AI models and multiple persona framings per subject, scored across five sub-indices: Visibility, Engagement, Accuracy, Framing, and Consensus. Each score reflects sustained-position scoring across the full four-turn conversation for a given subject, persona, and model; the underlying data does not currently isolate a score per individual turn, so this piece reports conversation-level and dimension-level findings rather than turn-by-turn decay curves. This piece reports findings that replicated across two subjects from opposing parties; single-dimension or single-subject results from the underlying dataset that did not replicate cleanly are named as open questions rather than presented as findings.


AIVO Standard is the open research and specification arm of AIVO. This piece uses probe data collected for AI Representation Score, a methodology extension of AIVO's core Reasoning Chain Score / PSOS framework applied to political and civic subjects.