Perplexity Cites Soft-Brand Queries, Not Generic Ones: 51% vs 0%

Soft-brand queries cited at 51.1 percent versus 0 percent for generic queries on Perplexity across 70 measurements
,
Perplexity Cites Soft-Brand Queries, Not Generic Ones: 51% vs 0%

Soft-brand queries, questions that reference concepts coined by a site without actually naming the site, produced a 51.1% citation rate on Perplexity in E031. Generic category queries produced 0%.

I ran 70 measurements over five days, from 14 to 18 June 2026, using Perplexity sonar-pro through DataForSEO. The target pages stayed fixed. The query tier changed.

The result was pretty chunky: 23/45 citations for soft-brand queries versus 0/25 for generic queries.

That’s a 51.1 percentage-point gap, 29 points above the 22-point interpretability threshold I carried forward from E016 (z = 4.362, p = 0.000013).

There are caveats, including an important one: E031 was not pre-registered. I’ll get to that below.

The complete dataset, frozen query file and analysis are deposited on Zenodo under DOI 10.5281/zenodo.22261587.

What I tested: soft-brand vs generic queries

I tested two types of query against the same three pages: soft-brand queries referencing concepts coined on The GEO Lab without naming the site, and generic category queries that other GEO pages could answer.

The soft-brand tier, T1.5, contained nine queries built around concepts such as the noise floor, GEO Stack, extractability and retrieval probability.

The generic T2 tier contained five broader queries, including things like “What is GEO?”, “GEO vs traditional SEO” and “Does GEO actually work?”

There was no T1 explicit-brand tier. This was deliberately a two-tier test.

The three target pages stayed fixed:

Each of the 14 queries ran once per day at 10:00 UTC for five days on Perplexity sonar-pro via the DataForSEO AI Optimization API.

That gave me 70 measurements.

The scoring rule was deliberately boring: did the answer cite the target URL, yes or no?

Nothing more was inferred from it.

The result: 51.1% versus an observed zero

Soft-brand queries produced 23 citations from 45 measurements, or 51.1%. Generic category queries produced 0 from 25.

Query tier Cited Total Rate 95% Wilson CI
T1.5 (soft-brand) 23 45 51.1% 37.0% to 65.0%
T2 (generic) 0 25 0.0% 0.0% to 13.3%

The observed gap is therefore 51.1 percentage points. E016 gave me a pre-existing 22-point threshold for interpreting differences above the noise I had measured previously. E031 clears that bar by 29 points.

The two-proportion z-test gives z = 4.362, p = 0.000013.

What caught my attention wasn’t just the average.

Every one of the five generic queries went 0/5 across all five days. That included the comparative and longer causal queries I thought were the most likely to sneak a citation through.

They didn’t.

Why I care more about the separation than the 51%

The useful result here is the separation between the query tiers. I would not treat 51.1% as some universal soft-brand citation rate.

There are only 45 measurements in that arm, and the confidence interval runs from 37% to 65%. That is a pretty wide range.

So I’m not claiming:

“Soft-brand queries have a 51% citation rate.”

I’m claiming something narrower.

Under this test, on these pages, over these five days, soft-brand queries produced substantially more target-page citations than the generic category queries.

The generic tier produced an observed 0/25. The soft-brand tier produced 23/45, and the difference cleared the 22-point interpretability threshold.

That separation is the finding I care about. The exact 51.1% point estimate is secondary.

One query has weaker provenance: t15_09

One soft-brand query, t15_09, is the marginal one, so I report the result both with and without it. In the day-2 backup screen it scraped into Perplexity’s results at retrieval rank 20 of 20, in-results but bottom of the set, and I included it under the pre-registered in-results rule. That rank is in my freeze log; the prescreen CSV in the deposit captures the day-1 screen only, so t15_09 isn’t in that file. It produced 0/5 citations. Keep it and the tier is 23/45 (51.1%); drop it and it rises to 23/40 (57.5%), so the marginal query pulls the headline down, and I keep 51.1% as the conservative number.

Variant Cited Total Rate 95% Wilson CI
With t15_09 23 45 51.1% 37.0% to 65.0%
Without t15_09 23 40 57.5% 42.2% to 71.5%
t15_09 alone 0 5 0.0%

So the weaker-provenance query actually pulls the headline rate down.

I’m keeping 51.1% as the headline number because it is the more conservative version, while exposing both calculations in the deposit.

What I didn’t control

E031 was not pre-registered. It ran on one engine, one site and three target pages over five days.

That matters.

The query file was frozen chmod 444 on 4 June 2026, ten days before collection began, and there is a Day 1 prescreen. But there is no public pre-registration or Git anchor.

That makes the provenance weaker than experiments where I registered the design publicly before collection.

There was also a JSON-LD schema edit during the measurement window, on 14 June at 18:21 UTC, after that day’s 10:00 measurement. The edit changed markup rather than the visible page text used for these citation observations, so I assessed it as DV-inert. It is still disclosed because it happened.

And, obviously, the sample is small.

One Perplexity model, one website, five days.

Replication on another site or engine could shrink the gap, remove it entirely, or produce a different pattern. That’s the next thing this result needs.

What soft-brand queries might mean for publishers

If this separation replicates, coined concepts may matter for AI citation visibility even when the publisher’s brand is absent from the query.

That’s the interesting bit for me.

The generic GEO questions produced 0/25 target-page citations.

Queries referencing concepts associated with the same site produced 23/45.

The pages didn’t suddenly become better between those two groups. The query tier changed.

But I want to be careful about what happens upstream.

E031 scored citation by target URL. It did not independently measure retrieval, candidate generation or source use. So I cannot use this experiment to say soft-brand queries caused Perplexity to retrieve these pages more often.

What I can say is simpler: the soft-brand queries were much more likely to end with those pages being cited.

That gives me a hypothesis worth testing.

Maybe naming and defining your own concepts creates a query surface that sits somewhere between explicit brand recognition and completely generic discovery.

Maybe.

E031 gives me a reason to test that idea. It doesn’t establish the mechanism.

Data and reproducibility

The full E031 dataset, runner, frozen query file and second-path analysis are deposited on Zenodo under DOI 10.5281/zenodo.22261587, licensed CC BY 4.0.

The 22-point noise-floor threshold comes from E016, DOI 10.5281/zenodo.19869156.

There’s also a useful companion experiment, E026, DOI 10.5281/zenodo.22230285, where I held the pages constant again but tested ten different query categories. That experiment found another large citation-rate separation based on how the question was phrased.

Different experiment, different variable, but the same annoying lesson keeps popping up:

The page is only half the test. The question matters too.

What is a soft-brand query?

A soft-brand query references a concept associated with a site without explicitly naming that site. Asking about the GEO noise floor without mentioning The GEO Lab is one example. In E031, these queries produced target-page citations in 23/45 measurements (51.1%), compared with 0/25 for generic category queries.

Did generic category queries get cited by Perplexity?

Not to the target pages in this test. All five generic GEO queries produced zero target-page citations across five days, giving an observed result of 0/25. The Wilson 95% confidence interval is 0.0% to 13.3%, so the experiment does not establish that the underlying citation probability is literally zero.

Is the 51.1% citation rate reliable?

I wouldn’t treat 51.1% as a universal rate. Its Wilson 95% confidence interval runs from 37.0% to 65.0%. The stronger result is the separation observed in E031: soft-brand queries produced 23/45 target-page citations versus 0/25 for generic queries, a 51.1-point gap that cleared the 22-point interpretability threshold used in the experiment.


About the Author

Artur Ferreira is the founder of The GEO Lab. He developed the GEO Stack framework and leads research into Generative Engine Optimisation methodologies. Connect on X/Twitter or LinkedIn.

Have questions about this topic? Contact The GEO Lab · Return to homepage