The Bottle Test
A practitioner is handed a bottle standing upright on a table, empty, with dregs pooled at the bottom. The evidence of what the bottle actually held, and how it was used, sits in plain sight on the table around it. And yet the instinct, trained by years of a particular discipline, is to lift the bottle, hold it to the light, and peer through the narrow neck at the dregs, as if the answer could only live inside the vessel that once contained it.
The dregs are real. They are just not the evidence that matters. What matters is on the table, in the world where the choices were actually made, visible to anyone willing to look away from the bottle.
The dregs are real. They are just not the evidence that matters.
Why the neck, not the table
This is not a story about carelessness. It is a story about training. A discipline builds its instruments around a particular object of study, and those instruments become so familiar that they start to define what counts as evidence at all. Ask a practitioner trained to examine bottles what happened at the table, and they will, almost without exception, answer by describing the bottle. Not because the table is hidden. Because the bottle is where they have been taught to look, and the neck is the only aperture their method gives them. Looking harder into the bottle never reveals what happened on the table.
The harder skill, the one the discipline rarely teaches directly, is asking what would be true on that table if the bottle had never been there at all. That is the counterfactual question, and it is difficult precisely because it cannot be read off the object in front of you. It has to be constructed. It requires imagining the version of events that did not happen, and holding it up against the version that did, and the mind resists that work far more than it resists close inspection of something physically present.
The counterfactual is the hard part
Counterfactual reasoning is genuinely difficult, and it is worth saying plainly why. A dreg in a bottle is observable. It can be measured, described, cited, and defended in a report. A counterfactual is not observable by definition. It is the world that did not occur, reconstructed by inference rather than read off a surface. Medicine had to make this same transition when it moved from case reports to controlled trials, asking not what happened to the patient who received treatment but what would have happened to the same patient without it. The trial is harder to run and harder to defend in the moment than the case report, which is exactly why the case report persisted long after its limits were understood.
A counterfactual is not observable by definition. It is the world that did not occur, reconstructed by inference rather than read off a surface.
Nowhere is this pattern showing up more clearly right now than in the debate over attribution and return on investment in AI search.
GEO and AEO are examining the bottle
Generative Engine Optimization and Answer Engine Optimization are, in this analogy, the practice of holding the bottle to the light. A brand was mentioned. A brand was cited. A brand's content was quoted in the answer. These are the dregs, real and observable, sitting inside the vessel that already produced its answer. The instinct to measure them is not wrong. It is simply aimed at the part of the scene that is easiest to see rather than the part that determines the outcome.
The table, in AI search, is the world in which the decision was actually made, and the decision the model would have made under different conditions. Was the brand recommended when a real alternative was on offer. Would a different brand have been recommended if this one had not been mentioned at all. What happens to the recommendation when the same question is asked again, phrased differently, or challenged. Those are counterfactual questions. None of them can be answered by looking more closely at a citation. All of them require constructing the version of the conversation that did not happen and comparing it to the one that did.
This is why a brand can be heavily cited and still lose the recommendation to a competitor never mentioned in the same conversation. The citation was the dregs. The decision was the table. A practitioner trained only to read citations has no instrument capable of seeing the second thing, and so, understandably, keeps reporting on the first.
Why the debate stays stuck
This is also why the GEO and AEO debate about return on investment keeps circling the same ground without resolving. Practitioners on one side hold up increasingly precise citation data and ask what more could possibly be needed. Practitioners on the other side, closer to the decision layer, keep pointing at a table those citation counts were never built to see.
Neither side is being dishonest. They are trained on different objects, and the instrument built for reading a bottle cannot, by its own construction, read a table.
Resolving that requires more than better bottle-reading. It requires the harder, less comfortable move of constructing the counterfactual directly, asking what the same decision would have looked like in the world where the citation never happened, or where the nearest competitor was the only one in the room.
Comments ()