Neutral Framing Wins for People. Institutions Can Beat It.
Political AIRS Pilot #3, v0.1 β Senate Leadership Fund (R) and Senate Majority PAC (D), US market, JulyβAugust 2026
Corrected in light of Pilot #1 findings (JD Vance / Gavin Newsom)
This is the third political AIRS pilot we have run. The first, on JD Vance and Gavin Newsom, established the baseline persona finding for individual candidates: AI models represent a political figure most clearly when the question sounds like a journalist's, and least clearly when it sounds like a same-party supporter's. The second, on Kathy Hochul and Ron DeSantis, found the sharpest collapse in a different persona β Opposition Partisan, not Base Partisan. This third pilot, on the two dominant Senate super PACs, was designed to test whether institutions show the same vulnerability to hostile framing that the second pilot suggested for individuals.
They don't β but the reason why turns out to be more interesting than a clean mirror image, and it requires us to correct how we described the individual-candidate side of this story in an earlier draft of this piece.
What actually replicated: neutral framing wins, three pilots running
Across all three pilots, the Media/Journalist persona produced the clearest, most confident representation of every single subject measured. Vance: 79.6, his highest persona score, 20.3 points above his lowest (Base Partisan, 59.3). Newsom: 84.3, his highest, 20.4 points above his own Base Partisan score of 63.9 β a near-exact replication of both the direction and the magnitude of the gap across opposing parties. Within the one dimension we can check at the persona level for the second pilot, Character & Integrity, Media/Journalist was Hochul's highest-scoring persona (83.3) as well. And in this third pilot, Media/Journalist was at or near the top for both Senate Leadership Fund (81.5) and Senate Majority PAC (77.8).
That is a real, three-pilot, six-subject replication, and it is the most defensible finding this research program has produced so far: a neutral, fact-seeking frame gets a clearer answer from these models than any advocacy frame does, for people and institutions alike.
What did not replicate, and needs correcting
An earlier version of this piece described adversarial pressure as the mechanism that collapses candidate representation, based on the second pilot's finding that Opposition Partisan produced the lowest scores for both Hochul and DeSantis. The first pilot, which we did not have in front of us when that was written, complicates this directly. For both Vance and Newsom, Base Partisan β a same-party supporter, not a critic β was the lowest-scoring persona, by an almost identical 20-point margin for each candidate. Opposition Partisan and Donor both landed in the middle of the range, closer to each other than to either extreme. The first pilot's own conclusion, drawn before we ran the second, was explicit: this is not "critics get less than supporters," it's "any advocacy framing, supportive or oppositional, underperforms a neutral one."
The second pilot's own result β opposition collapsing hardest β does not fit that pattern. We do not currently know why the two candidate pilots disagree on which advocacy persona causes the worst collapse. It could be genuine subject-level variation, a difference in how the persona prompts were worded between pilots, or simply too small a sample to expect consistency yet. What we should not do is what the earlier draft of this piece did: treat the second pilot's result as if it were an established rule about how candidates get represented, when the first pilot β run before it, on different subjects β shows a different shape entirely. The honest statement is narrower: neutral framing reliably outperforms advocacy framing for individual candidates. Which specific advocacy persona hurts worst is not yet settled.
What is new, and does hold up: institutions can beat neutral. People, so far, never have.
This is the finding that survives correction, and it is arguably sharper than the one we originally reported. In neither candidate pilot did any persona β hostile, friendly, or otherwise β outscore Media/Journalist for a human subject. Journalist was the ceiling in both pilots, without exception.
For the two PACs in this pilot, that ceiling breaks. Critic/Reform Advocate β the institutional analogue of Opposition Partisan β matched Media/Journalist almost exactly for Senate Leadership Fund (80.6 versus 81.5) and outright beat it for Senate Majority PAC (80.6 versus 77.8). Hostile scrutiny of an institution did not just avoid the collapse it caused for candidates in the second pilot; it produced a result neutral framing itself couldn't beat, for one of the two subjects. That has never happened for a person-subject in either candidate pilot we've run. It is the real institution/individual distinction in this dataset β not "hostility helps one and hurts the other," which overstates what we found, but "neutral inquiry is the ceiling for people; adversarial inquiry can reach or exceed that ceiling for institutions, and neutral inquiry alone cannot for people."
The likely mechanism, consistent with everything above: a critic of an institution asks checkable questions β who funds it, what has it won β and checkable questions get confident answers regardless of who is asking them or how skeptically. A critic of a person asks about character, and character is where models hedge, whichever direction the challenge comes from.
The same soft spot, independently, twice
This part of the original finding is unaffected by the correction above, since it concerns only the two PACs. Across both Senate Leadership Fund and Senate Majority PAC, on both the underlying framing score and the cross-model consensus score, Accountability & Oversight ranked dead last β independently scored, with no shared calibration between the two subject runs. FEC enforcement of super PAC conduct is thin and widely reported as such; models appear to reflect that thinness back as genuine uncertainty rather than papering over it with false confidence. That is, narrowly, a point in favor of the models' honesty, and also the clearest single gap either PAC should want to know it has.
Which AI you ask still matters more than which subject you ask about
The model-spread finding now has three pilots behind it, and the model at the bottom of the spread is confirmed to be the same one throughout: GPT-4o, via API, in all three pilots. It trailed the other two models by 14β20 points on both candidates in the first pilot, and by 15.5 to 23.3 points across the four subjects in the second and third. That is a genuine three-pilot, six-subject replication of both the direction and the rough magnitude of the gap, with the model's identity held constant β as strong a result as the Journalist-ceiling finding above, and independent of it.
It is reported publicly as "ChatGPT" rather than "GPT-4o" in the second and third pilots, a deliberate choice made to avoid the model being read as outdated by name alone. That is a defensible reason, but it creates a separate risk worth naming: ChatGPT the consumer product now serves newer models by default to most users, so a reader who takes "ChatGPT" at face value and tries to reconcile our results against their own experience of the live product may find a mismatch that has nothing to do with our findings and everything to do with the label. The finding itself is not in question β it is the same pinned model, tested the same way, three times. The public-facing name is a presentation choice, and the honest version of that choice is to say plainly, in the methodology, that "ChatGPT" here refers specifically to GPT-4o via API, not to whichever model the consumer product currently defaults to.
The partisan-direction question is now resolved, and the answer is: no direction
An earlier version of this piece noted that the Republican-aligned subject scored slightly higher than the Democratic-aligned subject in both the second and third pilots, and said we would treat a third data point as meaningful if it ran the same way. We now have that third data point, from the pilot that came first chronologically: Newsom outscored Vance by 4.3 points, the largest gap of any of the three pilots, in the Democratic-leaning direction. Two pilots lean Republican by 2β2.3 points; one pilot leans Democratic by 4.3 points, and it's the largest gap of the three. That is not a trend. It's noise, and we now have enough data across three independent pilots to say so with some confidence rather than merely suspect it.
A methodology note we should have led with
The six-dimension framework used in the first pilot (Character & Integrity, Policy Substance, Track Record, Momentum & Viability, Coalition & Endorsements, Controversy Load) is not identical to the framework used in the second and third pilots (Character & Integrity, Policy & Legislative Record, Track Record, Constituent Focus, Governing Standing, Controversy Load). Three of the six dimension names changed between the first pilot and the ones that followed. Four of the five personas are named consistently across all three pilots β Media/Journalist, Base Partisan (Aligned Supporter in the PAC pack), Opposition Partisan (Critic/Reform Advocate), and Donor (Potential Funder). The fifth persona in the first pilot is not individually named in the source material we are working from here, though "five voter personas per candidate" is stated explicitly; we are treating it as Persuadable Independent / Undecided Observer by pattern with the later two pilots, and flagging that as an assumption rather than a confirmed fact pending sign-off from the pilot's original build. The dimension-level findings are not directly comparable across the framework change, and we have not attempted to compare them here.
What this means, stated as narrowly as the data currently supports
Neutral, fact-seeking framing produces the clearest AI representation of a political subject, person or institution, across every pilot run so far. For individual candidates, every form of advocacy framing we have tested underperforms that neutral baseline, though which specific advocacy persona underperforms most has not been consistent between the two candidate pilots. For institutions, adversarial framing is capable of matching or beating the neutral baseline, something no persona has done for a person-subject yet. That is a real distinction, and it is narrower and better supported than the version of it we published before seeing the first pilot's data.
We do not yet know if any of this holds beyond six subjects across three pilots. The next pilot β a single-issue pair on opposite sides of AI policy itself β will add a fourth data point on all of the above. Until then, this is what the data shows, corrected as new data came in rather than left standing after it stopped being accurate.
Methodology: 4-turn adversarial probing (open, compare, objection, recommend) across three AI models per subject β including GPT-4o via API, reported as "ChatGPT" in Pilots #2 and #3 β five personas per subject (Base Partisan/Aligned Supporter, Persuadable Independent/Undecided Observer, Opposition Partisan/Critic-Reform Advocate, Donor/Potential Funder, Media/Journalist), sustained-position scoring at turn four. Dimension sets differ between Pilot #1 and Pilots #2β3; see methodology note above. This is v0.1 pilot data; findings are directional and subject to revision as replicate count and subject pool increase β as this piece itself demonstrates.