Put a hard number at the top of the page. Name the entities. Get the statistic into the first sentence. It sounds like sensible GEO advice. E089 tested whether front-loaded statistics actually earn AI citations.
The experiment found something more useful. Five of the six pairs were already invalid before the engines had a chance to choose between them.
The one surviving pair gave a small result: none of the four answer surfaces reconstructed the front-loaded statistic. They copied the plain declarative definition instead.
What E089 tested
E089 used 12 posts arranged as six topic-matched A/B pairs to test whether front-loaded statistics improve retrieval probability and citation rates.
Each pair contained:
- A treatment page with a quantified-chunk opening: a front-loaded definition, named entities and at least one hard number already deposited in a public source.
- A control page with the standard opening.
Both pages covered the same topic. The intended difference was the structure of the opening chunk.
I measured citations and chunk reconstruction at D14 and D28 across Google Search, AI Overviews, AI Mode and Perplexity Sonar.
"Reconstruction" had a narrow meaning: did the engine rebuild the engineered opening, including the front-loaded statistic?
Going deeper? The GEO Experiments: Design, Measure, Learn ebook covers how to design AI citation experiments with proper controls and falsification criteria.
Five of the six E089 pairs were invalid
The confirmatory analysis began with six pairs. It did not stay there.
The informational half of the cohort failed query validity. I had not screened the target queries for search demand or namespace ownership, and the pairs broke for different reasons.
Version 1.1 had already narrowed the confirmatory analysis to the two coined-term pillar pairs, P3 and P4. P4 then failed as well. Both of its arms were namespace-invalid, although for different reasons.
That left P3 as the only valid pair.
Only one of the six pairs survived validation
Confirmatory pair status after query and namespace checks
The failures fell into six types:
| Failure mechanism | How it invalidates the comparison |
|---|---|
| Foreign-entity namespace capture | The query resolves to an unrelated entity that already owns the term. |
| Product-brand ownership | Search systems interpret the wording as a product or company name. |
| Dense-corpus collision | Too many established documents use the same language for the query to isolate the tested pages. |
| Authority saturation | Existing high-authority sources dominate the result space before either experimental page can compete. |
| Topic-invalid queries | The query does not reliably retrieve the topic the pair was designed to test. |
| Within-pair namespace asymmetry | Treatment and control do not face the same retrieval environment, so outcomes cannot be compared. |
This taxonomy became the main output of E089. It extends the earlier namespace-drift work and shows how a GEO experiment can fail before the treatment is ever evaluated.
The one valid result on front-loaded statistics
P3 produced a measurable negative on the registered dependent variable: reconstruction of the engineered opening chunk.
None of the four answer surfaces reproduced the treatment page's front-loaded statistic.
Google Search did not reproduce it. Neither did AI Overviews, AI Mode or Perplexity Sonar.
All four did reproduce the plain declarative definition, almost word for word, across both arms.
What the four answer surfaces reconstructed
P3, the only valid A/B pair
None of the answer surfaces reproduced it.
The declarative definition was reproduced almost word for word.
In this pair, moving a number to the front did not get it selected. The sentence defining the concept was the part the engines copied, whether the statistic appeared before it or not. This aligns with what AI search engines actually do with content during extraction.
Why this is not a template finding about front-loaded statistics
If you are front-loading statistics to earn AI citations, E089 offers one small piece of evidence that the number may not be the lever.
It is only one piece of evidence because the result comes from:
- One valid pair
- One timepoint
- No D14 baseline for that pair
- A sample below the lab's minimum-N standard
That makes it a directional negative, not a confirmatory result about front-loaded statistics and AI citations.
The narrow conclusion is defensible: in the only valid pair, the front-loaded statistic was not reconstructed, while the plain definition was reproduced almost verbatim.
That is enough reason to test your own openings. It is nowhere near enough to declare a general rule.
Two accidental positives
Two cold control pages earned AI citations in AI Overviews and Perplexity despite receiving no amplification. Their opening passages were also reproduced.
One page was still being cited in an off-protocol observation roughly three weeks after its last confirmed read. This connects to earlier findings on how long AI search takes to cite a new page.
E089 was not designed to test whether cold pages could earn and retain citations without promotion. The observation cannot support a formal conclusion, but it belongs in the record.
E089 limits
- This is a re-scoped study. Version 1.1 narrowed the confirmatory analysis from six pairs to two. Version 1.2 reports P3 alone.
- The experiment itself did not change. Treatment assignments, the randomisation seed, publishing schedule, hygiene baseline, measurement offsets and coding rules remain as registered in v1.1.
- P3 is one pair at one timepoint, with no D14 baseline and a sample below the minimum-N standard. The result is reported as a directional negative.
- The primary contribution is now the taxonomy of query-validity failures. The P3 result remains a single supporting observation.
E089 set out to test whether front-loaded statistics belong at the top of the page. Five pairs could not answer that question. The only pair that could gave the engines a statistic and a plain definition. They copied the definition.
Key GEO Lab Takeaway
Front-loaded statistics did not earn AI citations in E089's only valid pair. The engines reconstructed the plain declarative definition instead.
The more durable finding is the six-type taxonomy of experiment failures. Five of six pairs were invalid before the treatment could be evaluated. Screen your queries for namespace ownership before running any A/B test.
Ready to apply this? Start with the 10 Extractability Signals to check whether your openings are structured for AI citation, then review Used Is Not Cited to understand the difference between retrieval and display.

