TL;DR: The GEO Lab’s E014 study tracked whether Perplexity, ChatGPT and Google AI Overviews cited thegeolab.net for ten fixed queries between February and September 2026. Citation of the site’s own concepts rose from 12% in June to between 92% and 96%, then held there for three consecutive months. Category queries stayed at 0%. The jump happened at the same time as an instrument change, so accumulation is still the leading explanation, but E014 cannot prove it.
For the first four months, the citation rate moved between 6% and 13%.
In August, it jumped to 48%. September produced almost the same result twice.
E014 was built to test System Memory, the fifth layer of the GEO Stack. After seven months, I can show exactly what changed and what held. I still cannot tell you with confidence what caused the change.
This is the final report. It covers the numbers, the confound that stops me making the claim I wanted to make, and four mistakes that belong in the record.
What E014 measured
E014 tracked whether thegeolab.net was cited by Perplexity, ChatGPT and Google AI Overviews across ten frozen queries. Measurements ran once a month from 28 February to 28 September 2026.
The hypothesis was simple: if System Memory builds through consistent publishing, the site’s citation rate should rise over time.
The ten queries were split into two tiers.
Tier 1 contained five questions about concepts defined by The GEO Lab: Retrieval Probability, LLM readability, the GEO Stack, Extractability and System Memory.
Tier 2 contained five broader category questions that many sites could answer, including “What is GEO (Generative Engine Optimisation)?” and “Does GEO actually work?”
Each month normally produced 50 reads. Perplexity ran three times per query, while ChatGPT and Google AI Overviews ran once each.
A read counted as cited only when a thegeolab.net URL appeared in the engine’s sources block. A link mentioned inside the answer without appearing in the sources did not count.
The full dataset, codebook and deviation log are on Zenodo. It includes the coded reads for M2 to M7, the raw M4 API responses, the raw M7 capture and the M6 control probe.
The seven-month series
| Month | Date | Combined | Tier 1 | Tier 2 |
|---|---|---|---|---|
| Baseline | 28 Feb | 3% | Not split | 0% |
| M1 | 28 Mar | 6.7% | Not split | 0% |
| M2 | 24 Apr | 12.9% (9/70) | Not split | 0% |
| M3 | 28 May | 12.0% (6/50) | 24% | 0% |
| M4 | 28 Jun | 6.0% (3/50) | 12% | 0% |
| M5 | 2 Aug | 48% (24/50) | 92% | Excluded* |
| M6 | 9 Sep | 46% (23/50) | 92% | 0% |
| M7 close | 28 Sep | 48% (24/50) | 96% | 0% |
*M5 used partly different Tier 2 query strings, so that row is excluded. M2 used a denominator of 70 reads.
The series falls into two clear periods.
From February to June, Tier 1 citations were sparse and unstable. From August to September, Tier 1 held between 92% and 96% across three consecutive measurements.
Tier 2 stayed at zero.
The jump happened when I changed the instrument
The move from 6% at M4 to 48% at M5 looks dramatic. It also happened in the same month that I changed two parts of the method.
From M5 onwards, every Perplexity read ran in a fresh Chrome incognito profile, with one query per profile. This kept each read in Perplexity’s advanced search mode. Earlier sessions could silently drop into basic search partway through a run.
ChatGPT also moved to forced-search reads. Each prompt began with “search the web for:” so ChatGPT retrieved sources instead of sometimes answering from memory.
Both changes fixed weaknesses in the earlier setup. They also created the main problem with interpreting the result.
I cannot separate “System Memory accumulated” from “the instrument stopped undercounting”. Either one could produce the jump I saw.
One check helps narrow it down. I audited the raw M2 and M3 data to see whether citations disappeared as Perplexity sessions progressed. They did not. An early query could miss while a later query in the same session was cited.
Those early numbers therefore include genuine retrieval misses. They were not entirely caused by degraded sessions. That makes a pure instrument artefact less likely, although it does not rule out a wider Perplexity change between June and August.
What held after the jump
At the September close, every Tier 1 query was cited across all three engines except for one Perplexity draw.
- Perplexity: 14 citations from 15 reads
- ChatGPT: 5 from 5
- Google AI Overviews: 5 from 5
The three-month stability is more useful than the jump itself.
M5, M6 and M7 all used the same pinned instrument, so those measurements are directly comparable. The move into the plateau is confounded. The plateau itself is not.
Once those citation bindings appeared, they held for three consecutive measurements.
The engines also began naming the source in the generated answer. In M6, Google’s AI Overview for “How does the GEO Stack work?” stated that the framework was developed by Artur Ferreira at The GEO Lab.
That attribution appeared in the answer itself, not only in the citation card.
Does Perplexity simply cite whoever coined the term?
A 92% Tier 1 citation rate could have a much simpler explanation: perhaps Perplexity tends to cite the site that coined a term.
I tested that at M6 using “What is the Skyscraper Technique in SEO?”, a term created by Brian Dean and associated with Backlinko for roughly a decade.
Perplexity cited Backlinko on 0 of 3 draws.
It named Brian Dean in the answer every time, but cited Semrush, Rank Math and seoClarity instead.
For The GEO Lab’s own terms, Perplexity cited thegeolab.net on 15 of 15 draws with the same query shape. “Perplexity cites the originator” does not explain that result.
The control was small and ran only once. Across all five control reads, ChatGPT cited Backlinko on its single read, making ChatGPT the engine that behaved most like an originator-citation system.
So this control rules out one explanation for Perplexity at M6. It does not prove accumulation.
The Extractability drop that looked worse than it was
The query “What is Extractability in GEO?” had been the most stable in the study.
Perplexity cited the Extractability page on every draw from M1 through M3. At M4, it suddenly fell to 0 of 3.
I logged the drop as a possible structural loss and made M5 the confirmation point.
The citation returned at 3 of 3 in M5, then stayed at 3 of 3 in M6 and M7. The M4 result was stochastic dropout, not a lasting loss.
One missing month from an otherwise stable citation is something to watch. It is not enough to call a change.
Where System Memory meets actual memory
The query “What is System Memory in GEO?” was the least stable Tier 1 query, mostly because of the name.
“System memory” is already a hardware term. “GEO” can also refer to several products and technical systems.
At M6, Google AI Overviews interpreted the query as Geo SCADA, Schneider Electric’s industrial control software, and returned an answer about RAM and server caching.
At M7, one of the three Perplexity draws resolved the query to GEOS, the PC operating system, and cited its developer documentation. The other two draws understood the intended concept and cited the System Memory page.
That gives us two different hardware interpretations from two engines.
I have seen the same pattern elsewhere in The GEO Lab’s data. Ambiguous anchors produce more collisions. Coined terms with little competing meaning resolve much more cleanly.
Even after the association exists, a concept name that overlaps with an established technical term carries an extra disambiguation cost on every query.
The category floor never moved
Tier 2 stayed at 0% across every clean measurement.
Over seven months, thegeolab.net was not cited for “What is GEO (Generative Engine Optimisation)?” or the other four category questions. There was one single-draw hit at M5, but it disappeared on the next iteration.
At the close, I also checked Google’s organic results for all five Tier 2 queries. The site was absent from the top results for every one of them.
In this dataset, Google AI Overview citation follows organic ranking closely. The Tier 2 zero on Google therefore fits the existing rank gate. It does not give us a new authority finding.
Owning a concept and competing for a category turned out to be two very different jobs.
What I got wrong
Four problems need to stay attached to this result.
M1 raw data is missing
The row-level M1 file was never consolidated and could not be recovered. The figure survives only in the write-up published at the time. Everything from M2 onwards is in the Zenodo deposit, which is observational and was not pre-registered.
M5 used the wrong Tier 2 queries
The M5 capture sheet contained reconstructed query strings that had not been checked against the frozen set.
Three Tier 2 questions were replaced and another was shortened. The M5 Tier 2 result is therefore excluded from the series. Tier 1 was unaffected apart from one capitalisation difference.
Two measurements ran late
M5 ran five days late. M6 ran twelve days late. Both remain in the series as dated observations rather than pretending they were collected on schedule.
ChatGPT and Google only had one read per query
Each ChatGPT and Google leg used a single read per query per month. That falls below The GEO Lab’s own minimum of five renders for making a citation claim.
Those per-engine results are directional. Perplexity, with three iterations per query, is the only leg with any depth, and even that remains a small sample.
What E014 shows
| Claim | Status |
|---|---|
| Tier 1 citation held between 92% and 96% for three consecutive months on a fixed instrument | Shown |
| Tier 2 category citation stayed at 0% | Shown |
| The M4 Extractability drop was stochastic rather than structural | Shown |
| Perplexity does not always cite the originator of a term | Shown once, M6 control, n=5 |
| System Memory accumulation caused the M4-to-M5 jump | Not shown because the instrument changed |
| System Memory accumulation explains the plateau | Leading hypothesis, not established |
What comes next
The successor study needs a control entity measured every month on the same instrument. A one-off control can reject a narrow explanation, but it cannot separate accumulation from platform drift across seven months.
The close also left me with two smaller tests.
The retrieval-probability concept now has two competing pages: /retrieval-probability/ and /what-raises-retrieval-probability/. That creates a clean page-selection problem.
The System Memory collisions give me another test: whether a more distinctive anchor reduces entity-resolution failures.
This series began with the Month 1 citation-rate baseline.
It ends with a result I trust and an explanation I do not yet have.
Frequently asked questions
What is System Memory in GEO?
System Memory is the fifth layer of the GEO Stack. It describes the association an AI system builds between a site and a topic over time.
You cannot edit it directly. It develops through consistent work across the four layers below it: Retrieval Probability, Extractability, Entity Reinforcement and Structural Authority.
E014 tested whether that accumulation would appear as a rising citation rate.
Did E014 prove that System Memory accumulates?
No.
E014 showed Tier 1 citation holding between 92% and 96% for three consecutive months. The jump into that plateau happened at the same time as an instrument change, so the study cannot assign the increase to accumulation alone.
Accumulation remains the leading hypothesis. The successor study will need a control entity measured every month using the same instrument.
Why did Tier 2 citation stay at zero?
The Tier 2 queries compete against established domains, and thegeolab.net did not appear in Google’s top organic results for any of the five queries at the close.
Google AI Overview citation followed organic rank closely in this dataset, so its Tier 2 zero is consistent with the rank gate. It does not establish a separate authority effect.
Why does “What is System Memory in GEO?” sometimes fail?
The phrase overlaps with existing hardware and software meanings.
“System memory” is a standard computing term, while “GEO” also matches several product and system names. Google AI Overviews returned Geo SCADA at M6. One Perplexity draw returned the GEOS operating system at M7.
The remaining Perplexity draws understood the intended framework and cited The GEO Lab.
How did E014 count a citation?
A read counted as cited only when a thegeolab.net URL appeared in the engine’s sources block. A link mentioned only inside the answer did not count.
That coding rule was fixed at M5, checked against the M4 raw data and applied unchanged through the close. Every measurement from M5 onwards was therefore coded the same way.
Have questions about this topic? Contact The GEO Lab · Return to homepage

