Most confident claims about AI search are inherited priors nobody has tested. Ask what test produced the claim, and the field answers more openly than you would expect.
TL;DR
The tested-or-prior question is a four-word filter: was this tested, or is it a prior? It separates claims derived from controlled measurement from claims inherited from convention or analogy. The GEO Lab’s noise-floor experiment (E016) established a 22.0% combined interpretability threshold. Any citation-rate movement below that figure is indistinguishable from platform variance. Most claims circulating in generative engine optimisation do not clear that bar, because nobody measured them against one. Asked of three practitioners about three AI search claims, the question surfaced a confounded variable and an untested proxy in the first two, then a case study read backwards in the third. The failure mode was consistent: a prior repeated until repetition did the work that evidence should have done.
Failure mode one: the confounded variable
The tested-or-prior question surfaces confounds that confident advice conceals. The first claim was that content should be published on a dedicated subdomain to improve retrieval and recommendation by AI systems. It was made by a company that builds tooling in this space, so they had seen more data on it than most.
The question was whether the recommendation had been A/B tested against the same content hosted on the main domain, or whether it was a best-practice prior.
The answer came back from the co-founder: the lift was not necessarily due to the subdomain, but to information density. Then the company’s head of go-to-market added that either a subdomain or the main website works equally well, and that the subdomain is how they host and analyse content rather than a mechanism that earns retrieval.
That is a confound admitted in public, and it was admitted gracefully. The observed lift was real. The variable credited with it was not the variable doing the work. The recommendation that survives is not “move your content to a subdomain,” it is “raise the information density of your canonical pages,” which is a completely different instruction that costs nothing in site architecture.
The design that would have caught this before it became advice is a density-matched holdout. Same host, same density, subfolder against subdomain, citation rate over a fixed window. It sidesteps the caching and penalty objections that usually stall this argument, and it isolates the one variable in dispute.
Going deeper? The GEO Experiments ebook covers how to design controlled tests that isolate a single variable, including the density-matched holdout pattern described above.
Failure mode two: the untested proxy
The claim was that inspecting the internals of an open-weights model tells you something actionable about how a closed commercial model will behave, because you can observe where the open model’s representation of a brand is thin.
This is an interesting research direction, and the person making the case had been careful. The probe was narrower than the idea: a gap found in the open model’s activations is a fact about that model. It is not yet a fact about the closed system that actually serves the answer, and cross-model transfer is the assumption doing the load-bearing work.
The author conceded the point directly, and said so in his own thread: a feature in the open model is a fact about that model, and cross-model transfer is a hypothesis rather than a measurement.
That concession is what makes the research credible rather than what undermines it. A proxy is legitimate the moment somebody demonstrates that a diagnosed gap in the proxy predicts behaviour in the production system. Until then the layer generates hypotheses, which is a real contribution, and not verdicts, which is what it gets sold as downstream by people who did not read the caveats.
Notice that the failure here is not in the work. It is in the distance between what the researcher claims and what the audience hears.
Failure mode three: the case study read backwards
The claim was that a small site had been cited in AI Overviews without doing SEO, offered as a counterexample to the argument that AI visibility rests on a search foundation.
Read on its own terms, the case study says the page ranked in position three organically before it appeared in the AI Overview. AI Overviews draw heavily on pages that already rank. So the sequence is: rank first, get cited second. That is search doing the work and AI visibility riding on top of it, which is evidence for the claim it was offered to refute.
The test of “cited without SEO” is a page cited in AI answers while ranking nowhere in organic search. This was not that. The topics were also low-competition informational queries, so beating the large brands mostly meant the large brands had not entered.
Nobody was being dishonest. The case study was read in the direction the reader expected it to point, which is the most common way evidence gets misused and the hardest to catch in your own work.
Three failures, one question
A confounded variable, an untested proxy, and a case study read backwards do not look alike. They come from different disciplines and they fail for different reasons.
The same question surfaces all three, because all three share a structure. In each case a confident recommendation had been derived from an observation that could not license it, and in each case the person making the claim knew this when asked plainly. Nobody had asked.
That is the finding, and it is a finding about the field rather than about three individuals. The claims circulating in generative engine optimisation are not mostly lies. They are mostly priors that have been repeated until repetition did the work that evidence should have done, and they persist because asking feels adversarial and agreeing feels collegial.
It is the other way round. Asking someone what test produced their claim treats them as a person who might have run one. Nodding along treats them as someone whose claims do not matter enough to check.
How to ask it
Ask about the claim, never the person. “Has this been tested against X, or is it a best-practice prior?” contains no accusation and offers a dignified exit.
Accept “it is a prior” as a complete and respectable answer. If you punish honesty you will stop receiving it, and the person will simply be more confident next time.
Name the design that would settle it. This converts a challenge into a contribution, and it is the difference between scoring a point and moving the field. Every one of the three exchanges above ended with a testable design on the table.
Then apply it to your own work first, because the failure modes above are not other people’s failure modes. The most dangerous prior in your head is the one you have never been asked about, and there is nobody in your own thread to ask.
“Was this tested, or is it a prior?” has two acceptable answers and neither is an accusation. A prior stated with the confidence of a result is how the field manufactures consensus out of nothing.
Three failure modes surfaced by the same question: a confounded variable (density, not subdomain), an untested proxy (open model does not establish closed model behaviour), and a case study read backwards (ranking preceded citation).
The most dangerous prior is the one in your own work that nobody has asked about. Name the test that would settle it, then run it.
Ready to test your own claims? The controlled testing framework shows how to isolate variables and measure effects above the noise floor. See the probe query taxonomy for building a test set.
Questions? Contact The GEO Lab.
Frequently Asked Questions
How do you evaluate a claim about AI search visibility?
Ask what test produced it. A claim derived from a controlled comparison can be checked against its design. A claim derived from observation or best practice is a prior, which may still be correct but cannot support a confident recommendation until it is tested.
What is the difference between a tested claim and a prior?
A tested claim comes from a design in which the variable of interest was isolated and the outcome measured. A prior is an expectation inherited from experience, convention, or analogy. Both are useful, but only one licenses a causal recommendation.
Is asking someone to justify a claim hostile?
Not when the question is about the claim rather than the person, and when a prior is accepted as a legitimate answer. In practice, practitioners concede confounds readily when asked plainly, because most have never been asked.
Does publishing content on a subdomain improve AI retrieval?
There is no controlled evidence that it does. When asked directly, the practitioners recommending it attributed the observed lift to information density rather than to the subdomain, and confirmed that content on a main domain performs comparably.
Can you study a closed AI model by inspecting an open one?
You can generate hypotheses that way. Establishing that findings transfer requires demonstrating that a diagnosed feature in the open model predicts the behaviour of the closed system, which has not yet been shown.

