The Real Question Isn't Whether AI Wrote It

The Real Question Isn't Whether AI Wrote It
If stylistic fingerprinting is the wrong test, what's the right one?

A Weekend Essay by the Editor, Tim de Rosen

Stanley Druckenmiller's recent op-ed criticizing Treasury Secretary Scott Bessent set off a familiar ritual. An AI-detection tool flagged the piece as fully machine-generated. Commentators piled on. The Journal's editorial page editor, Paul Gigot, shrugged it off: "AI is a fact of modern life." Mr. Druckenmiller was blunter still, comparing AI to a calculator he uses without embarrassment.

Both sides are arguing past the actual problem.

No large language model is submitting op-eds under its own byline. Every AI-assisted piece of writing published today has a human name attached, a human argument behind it, and β€” in theory β€” a human editor who decided it was worth running. The question was never whether AI touched the text. By any realistic accounting, most professional writing now involves some AI assistance, from research to editing to drafting. The question is whether that assistance produced something with genuine intellectual content, or whether it produced filler dressed up as argument. Slop, or synthesis.

Detection tools like Pangram cannot answer that question, because they were never built to. They measure style β€” sentence rhythm, word choice, the telltale cadences that give away a model's hand. But style is not substance. A human could write vacuous, recycled talking points in flawless prose and pass every detector; an AI could help sharpen a genuinely novel argument and get flagged as fraudulent. Treating "100% AI-detected" as a verdict on quality is a category error, and both the Journal's critics and its defenders have been making it.

If stylistic fingerprinting is the wrong test, what's the right one?

The meaningful test comes after publication, not before it. Weak arguments disappear almost immediately. Strong ones become inputs into other people's thinking. They are quoted, challenged, extended, incorporated into reports β€” and increasingly retrieved, weighed, and synthesized by AI systems answering later questions. The relevant metric isn't whether a model once helped write the piece. It's whether anyone, or anything, continues to rely on the reasoning after it's published.

That's a higher bar than citation alone, and it's worth being precise about the difference. Citation counts are a blunt instrument β€” plenty of bad ideas get cited, some of them because they're wrong in interesting ways.

What matters is whether an argument gets incorporated into downstream reasoning: whether its claims survive being checked, whether its logic gets reused rather than merely referenced. And AI systems don't cite the way academics do. They retrieve, weight, and select among competing accounts before synthesizing an answer. That process is a more demanding test than a footnote, not a laxer one β€” an argument has to actually hold up to be drawn on, not just be findable.

Does the reasoning repeatedly survive that kind of scrutiny, or does it evaporate the moment anyone β€” or anything β€” checks it? That's the test that stylometry was never built to run.

This reframing matters because the current fight is a dead end. Newsrooms can't out-detect the models producing the text; every generation of detector gets outpaced within months, and false positives already threaten to become their own form of reputational harm β€” accusing a slow, careful human writer of being a machine is not a small mistake. Meanwhile, banning AI assistance outright is both unenforceable and increasingly beside the point, given how deeply embedded these tools already are in ordinary professional writing.

A better standard asks editors, and readers, to judge the thing that has always mattered: does this argument hold up, and does it earn a place in the conversation that follows it? That test was always available for human-written pieces too β€” good arguments got reused, bad ones got ignored β€” but it's about to become far more visible and far more measurable, because a growing share of "the conversation that follows" now happens inside AI systems synthesizing answers from what's already been written and argued.

Today the question editors ask is "who wrote this?" Before long, the more useful question will be "which of this survived?" Authorship is becoming metadata. Influence is becoming observable.

That's the deeper shift underneath this week's controversy. Before large language models, fluent writing was the scarce resource, and authorship was a reasonable proxy for effort and therefore for quality. AI has made fluent writing abundant. What's scarce now is an argument strong enough that other people β€” and increasingly other machines β€” keep relying on it after publication.

The op-ed page has always been a filter for exactly that. It just needs to stop pretending the filter is about who typed the words, and start measuring the thing that was always the actual point: not production, but contribution.


Tim de Rosen is CEO and co-founder of AIVO, Inc.