Front-Loaded Statistics Did Not Earn AI Citations

E089 front-loaded statistics did not earn AI citations: 0 of 4 surfaces reconstructed the statistic
Front-Loaded Statistics Did Not Earn AI Citations

Put a hard number at the top of the page. Name the entities. Get the statistic into the first sentence. It sounds like sensible GEO advice. E089 tested whether front-loaded statistics actually earn AI citations.

The experiment found something more useful. Five of the six pairs were already invalid before the engines had a chance to choose between them.

The one surviving pair gave a small result: none of the four answer surfaces reconstructed the front-loaded statistic. They copied the plain declarative definition instead.

TL;DR: E089 found no citation advantage from front-loaded statistics in its only valid A/B pair. Four answer surfaces copied the plain definition instead. Five of six pairs were invalid.

What E089 tested

E089 used 12 posts arranged as six topic-matched A/B pairs to test whether front-loaded statistics improve retrieval probability and citation rates.

Each pair contained:

  • A treatment page with a quantified-chunk opening: a front-loaded definition, named entities and at least one hard number already deposited in a public source.
  • A control page with the standard opening.

Both pages covered the same topic. The intended difference was the structure of the opening chunk.

I measured citations and chunk reconstruction at D14 and D28 across Google Search, AI Overviews, AI Mode and Perplexity Sonar.

"Reconstruction" had a narrow meaning: did the engine rebuild the engineered opening, including the front-loaded statistic?

Going deeper? The GEO Experiments: Design, Measure, Learn ebook covers how to design AI citation experiments with proper controls and falsification criteria.

Five of the six E089 pairs were invalid

The confirmatory analysis began with six pairs. It did not stay there.

The informational half of the cohort failed query validity. I had not screened the target queries for search demand or namespace ownership, and the pairs broke for different reasons.

Version 1.1 had already narrowed the confirmatory analysis to the two coined-term pillar pairs, P3 and P4. P4 then failed as well. Both of its arms were namespace-invalid, although for different reasons.

That left P3 as the only valid pair.

Only one of the six pairs survived validation

Confirmatory pair status after query and namespace checks

P1Invalid
P2Invalid
P3Valid
P4Invalid
P5Invalid
P6Invalid
Five pairs could not support the intended comparison. P3 was the only pair retained for the directional result.

The failures fell into six types:

Failure mechanismHow it invalidates the comparison
Foreign-entity namespace captureThe query resolves to an unrelated entity that already owns the term.
Product-brand ownershipSearch systems interpret the wording as a product or company name.
Dense-corpus collisionToo many established documents use the same language for the query to isolate the tested pages.
Authority saturationExisting high-authority sources dominate the result space before either experimental page can compete.
Topic-invalid queriesThe query does not reliably retrieve the topic the pair was designed to test.
Within-pair namespace asymmetryTreatment and control do not face the same retrieval environment, so outcomes cannot be compared.

This taxonomy became the main output of E089. It extends the earlier namespace-drift work and shows how a GEO experiment can fail before the treatment is ever evaluated.

The one valid result on front-loaded statistics

P3 produced a measurable negative on the registered dependent variable: reconstruction of the engineered opening chunk.

None of the four answer surfaces reproduced the treatment page's front-loaded statistic.

Google Search did not reproduce it. Neither did AI Overviews, AI Mode or Perplexity Sonar.

All four did reproduce the plain declarative definition, almost word for word, across both arms.

What the four answer surfaces reconstructed

P3, the only valid A/B pair

Front-loaded statistic 0 of 4

None of the answer surfaces reproduced it.

Plain definition Both arms

The declarative definition was reproduced almost word for word.

In P3, moving the number to the front did not get it selected. The engines copied the sentence defining the concept.

In this pair, moving a number to the front did not get it selected. The sentence defining the concept was the part the engines copied, whether the statistic appeared before it or not. This aligns with what AI search engines actually do with content during extraction.

Why this is not a template finding about front-loaded statistics

If you are front-loading statistics to earn AI citations, E089 offers one small piece of evidence that the number may not be the lever.

It is only one piece of evidence because the result comes from:

  • One valid pair
  • One timepoint
  • No D14 baseline for that pair
  • A sample below the lab's minimum-N standard

That makes it a directional negative, not a confirmatory result about front-loaded statistics and AI citations.

The narrow conclusion is defensible: in the only valid pair, the front-loaded statistic was not reconstructed, while the plain definition was reproduced almost verbatim.

That is enough reason to test your own openings. It is nowhere near enough to declare a general rule.

Two accidental positives

Two cold control pages earned AI citations in AI Overviews and Perplexity despite receiving no amplification. Their opening passages were also reproduced.

One page was still being cited in an off-protocol observation roughly three weeks after its last confirmed read. This connects to earlier findings on how long AI search takes to cite a new page.

E089 was not designed to test whether cold pages could earn and retain citations without promotion. The observation cannot support a formal conclusion, but it belongs in the record.

E089 limits

  • This is a re-scoped study. Version 1.1 narrowed the confirmatory analysis from six pairs to two. Version 1.2 reports P3 alone.
  • The experiment itself did not change. Treatment assignments, the randomisation seed, publishing schedule, hygiene baseline, measurement offsets and coding rules remain as registered in v1.1.
  • P3 is one pair at one timepoint, with no D14 baseline and a sample below the minimum-N standard. The result is reported as a directional negative.
  • The primary contribution is now the taxonomy of query-validity failures. The P3 result remains a single supporting observation.

E089 set out to test whether front-loaded statistics belong at the top of the page. Five pairs could not answer that question. The only pair that could gave the engines a statistic and a plain definition. They copied the definition.

Key GEO Lab Takeaway

Front-loaded statistics did not earn AI citations in E089's only valid pair. The engines reconstructed the plain declarative definition instead.

The more durable finding is the six-type taxonomy of experiment failures. Five of six pairs were invalid before the treatment could be evaluated. Screen your queries for namespace ownership before running any A/B test.

Ready to apply this? Start with the 10 Extractability Signals to check whether your openings are structured for AI citation, then review Used Is Not Cited to understand the difference between retrieval and display.

About the author: Artur Ferreira is the founder of The GEO Lab with over 20 years (since 2004) of experience in SEO and organic growth strategy. He developed the GEO Stack framework and leads research into Generative Engine Optimisation methodologies. Connect on X/Twitter or LinkedIn.

Have questions? Contact The GEO Lab