Ten distinct query types, each targeting a different layer of the retrieval pipeline — the working query taxonomy behind every GEO Lab experiment
TL;DR
“AI visibility” is not one thing. An AI system that cites you for “What is the GEO Stack?” may never cite you for “best AI SEO tools” — those two queries probe different layers of the retrieval pipeline, and they fail independently. Ten probe types cover most of the meaningful territory: branded, proprietary concept, category/commercial, comparison, how-to, troubleshooting, definitional, statistical, latest-news, and niche-technical.
Testing with only one or two types gives you a partial picture. Designing a GEO experiment without a probe taxonomy is how you end up with a “citation rate went up” result that doesn’t generalise to any query type that actually matters commercially.
Citation rate is not a single number — it varies by query type, and the variation is structural, not noise. A probe set that spans all ten query types is the only way to diagnose which layer of the retrieval pipeline your content is failing at, and why the same page gets cited for one query and ignored for another.
Why Query Type Changes the Answer
I used to think of AI citation as a single phenomenon. A URL was either cited or it wasn’t. Citation rate was one number. Then I ran the same URL through ten different query types on the same day and got citation rates ranging from 80% to 0%. Same URL. Same day. Same platform. An order of magnitude of difference, driven entirely by what kind of question was being asked.
The GEO Stack retrieval model predicts that different query types stress different pipeline layers — this probe taxonomy was built to test those predictions with controlled experiments on real pages.
That range is not noise. It’s structural. AI systems handle different query types through different retrieval and ranking policies, and the policies don’t agree on what counts as a good source. A definitional query pulls toward authoritative originators of terms. A commercial query pulls toward aggregators and review sites. A troubleshooting query pulls toward forums and documentation. The same URL has different visibility in each.
Which means if you test AI visibility with only branded queries, you’ll conclude your site is highly visible — and then be surprised when you’re never cited for commercial queries that actually drive revenue. If you test with only category queries, you’ll conclude your site is invisible — and miss the fact that you dominate proprietary-concept citations.
The probe taxonomy below is how the Lab separates these signals. Each probe type has its own citation-rate baseline, its own expected cited-source distribution, and its own set of interventions that move the metric.
Three Pipeline Layers, Ten Probes
The ten probes group by which retrieval pipeline layer they primarily stress. A well-designed experiment spans all three.
Each probe type produces different amounts of variance across measurement runs — the ten citation measurement variables document what else changes between runs beyond query type.
Layer A
Entity recognition
Probes 01, 07
Layer B
Retrieval selection
Probes 02, 03, 04, 10
Layer C
Answer synthesis
Probes 05, 06, 08, 09
The value of this mapping: when a probe fails, the layer tells you roughly where to intervene. Layer A failures need entity reinforcement work — canonical naming, schema consistency, sameAs links. Layer B failures need retrieval-side work — crawler access, indexation, topical authority. Layer C failures need extraction-side work — declarative openings, section structure, clean answer isolation.
The Ten AI Citation Probe Types
Each probe has a definition, an example query, the layer it stresses, the expected citation pattern, and what it tells you when it fails.
Understanding what each probe detects requires knowing how content formats affect retrieval — the ten content formats mapped to query intent shows which formats produce the citation patterns each probe type is designed to reveal.
01
Branded query
What it tests
Whether the AI system recognises the brand as an entity and cites its canonical properties.
Example
“What is The GEO Lab?”
Layer
A — Entity recognition. Pure identity probe.
If it fails
The AI hasn’t indexed the brand as a distinct entity. Foundational problem — nothing else in the probe set will work until this one does.
02
Proprietary concept query
What it tests
Whether the AI cites the originator when a proprietary framework, term, or concept is named directly.
Example
“What is the GEO Stack?”
Layer
B — Retrieval selection. High signal-to-noise: the term itself disambiguates the retrieval set.
Expected pattern
Tier 1 citation behaviour. The originator should dominate citations. Failure here means the vocabulary hasn’t propagated to the retrieval index, and System Memory hasn’t kicked in.
03
Category/commercial query
What it tests
Whether the AI cites your content for general topic queries where many sources compete.
Example
“best AI SEO tools”
Layer
B — Retrieval selection. Highly competitive retrieval set, heavily biased toward aggregator and commercial domains.
Expected pattern
Tier 2 citation behaviour. Citations skew to list sites, review aggregators, and the highest-authority domains in the vertical. Originating a specific tool concept does not guarantee citation here — the query simply doesn’t pull that way.
04
Comparison query
What it tests
Whether the AI cites your content when a direct comparison is requested between two named entities.
Example
“GEO vs AEO vs LLM SEO — what’s the difference?”
Layer
B — Retrieval selection. Comparison queries tend to retrieve longer-form explainer content rather than commercial pages.
Expected pattern
Mid-weight citation — between proprietary concept and category queries. Comparison content on your site should be explicitly structured as comparison content, with parallel section structure for each entity being compared.
05
How-to query
What it tests
Whether the AI cites your content when a user asks how to perform a specific task.
Example
“How do I check if my site is being crawled by GPTBot?”
Layer
C — Answer synthesis. How-to queries favour content with step-by-step structure and concrete code snippets.
Expected pattern
Citation favours content with ordered-list structure, H2 phrased as the question, and code blocks where applicable. Failure often means the content has the answer but not in an extractable shape.
06
Troubleshooting query
What it tests
Whether the AI cites your content when a user describes a specific error or failure and asks how to fix it.
Example
“Google Search Console ‘Duplicate field: FAQPage’ error — how to fix”
Layer
C — Answer synthesis. Troubleshooting queries favour content with the exact error string and a documented fix.
Expected pattern
Citation favours failure registries, issue logs, and documentation with the exact error message in the text. AI systems performing troubleshooting synthesis look for string-level matches against the reported error before anything else.
07
Definitional query
What it tests
Whether the AI cites your content when a user asks for the definition of a general term.
Example
“What does citation rate mean in AI search?”
Layer
A — Entity recognition. The term being defined is the entity being resolved.
Expected pattern
Citation favours content that opens with a clean “X is Y” definition in the first 200 characters. If your page has the term but defines it through narrative rather than declaratively, it will lose this probe to competitors who define it up front.
08
Statistical/data query
What it tests
Whether the AI cites your content when a user asks for a specific number, percentage, or benchmark.
Example
“What percentage of AI search citations come from Wikipedia?”
Layer
C — Answer synthesis. Statistical queries strongly favour content with the number stated explicitly and with a traceable source.
Expected pattern
The research literature on GEO confirms that content with embedded statistics and quotations from credible sources is cited more often than content without. Statistical probes are the place where this effect is most visible — content with the exact number wins against content that discusses the topic generally.
09
Latest-news query
What it tests
Whether the AI cites your content for recency-sensitive queries where freshness is a significant ranking factor.
Example
“Latest GPTBot user-agent string 2026”
Layer
C — Answer synthesis. Recency bias in retrieval often dominates authority for news queries.
Expected pattern
Citation favours content with recent publication or modification dates, dated page metadata, and explicit year mentions in the content. Old but authoritative content loses to new but less established content for this probe.
10
Niche-technical query
What it tests
Whether the AI cites your content for specialist queries with a small candidate retrieval pool.
Example
“How do I prevent wpautop from breaking inline JSON-LD in WordPress?”
Layer
B — Retrieval selection. Small retrieval pool, reduced competition, the probe where domain specificity matters most.
Expected pattern
Citation rates tend to be high when the retrieval pool is small, which is why niche content can dominate AI citation without ranking on Google. Failure here often indicates the content isn’t in the retrieval index at all — not that it lost to a better source.
The Ten at a Glance
One reference table. Use when assembling a probe set for a new experiment — include at least one query per row to get signal across the full pipeline.
| # | Probe type | Layer | Citation tier | Primary signal |
|---|---|---|---|---|
| 01 | Branded | A | Tier 1 | Is the brand indexed as an entity at all? |
| 02 | Proprietary concept | B | Tier 1 | Has your vocabulary propagated to the index? |
| 03 | Category/commercial | B | Tier 2 | Do you compete against aggregators in the vertical? |
| 04 | Comparison | B | Mixed | Is your content structured as explicit comparison? |
| 05 | How-to | C | Mixed | Does the task-answer live in an extractable format? |
| 06 | Troubleshooting | C | Mixed | Does the exact error string appear in the content? |
| 07 | Definitional | A | Tier 1 | Does the page open with a clean “X is Y” sentence? |
| 08 | Statistical/data | C | Tier 1 | Is the number stated explicitly with source? |
| 09 | Latest-news | C | Mixed | Is the content dated and recent? |
| 10 | Niche-technical | B | Tier 1 | Is the content in the retrieval index for this specialism? |
Building a Probe Set for Your Own Site
The ten types above are categories, not queries. A working probe set replaces each category with specific queries about your site. Three principles make the translation useful rather than arbitrary.
Tier 1 probes consistently outperform Tier 2 by a wide margin — the citation gap experiment provides five days of data on the magnitude of that difference across 25 queries per tier.
First: one query per type is the minimum, three is better. A single query per type gives you a point estimate per category. Three queries per type gives you a within-category variance estimate, which lets you distinguish “this one query is noisy” from “this whole category is a problem.”
Second: queries must be commercially relevant to your site. A probe set built from queries you don’t care about gives you citation rates you can’t act on. Pick queries where being cited would matter — either because they drive traffic, inform purchase decisions, or validate the authority of your framework.
Third: write the queries the way users actually phrase them. “How do I fix the FAQPage duplicate field error” beats “FAQPage duplicate field error resolution methodology” for the same probe type. The retrieval systems are tuned on natural user queries, and unnatural phrasings produce unnatural retrieval results.
A worked example from thegeolab.net’s own probe set:
| Probe type | Site-specific query | What we learn when it fails |
|---|---|---|
| Branded (01) | “What is The GEO Lab?” | Entity reinforcement hasn’t stabilised across AI systems |
| Proprietary concept (02) | “What is the GEO Stack?” | Framework adoption rate lags citation rate |
| Category (03) | “Best AI visibility tools” | We’re not competing in the commercial retrieval set |
| Definitional (07) | “What does citation rate mean?” | Our definitional pages aren’t opening declaratively enough |
| Troubleshooting (06) | “FAQPage duplicate field GSC error fix” | The error string isn’t surfacing in the retrieval match |
On selection bias. Every probe set is a sample of queries, not the entire query universe. A probe set that over-weights one category will produce a citation rate that reflects that category rather than overall AI visibility. Balance the set — one query per type, minimum — and report citation rates per category rather than as a single site-wide average. The single-number summary is where most GEO reporting goes wrong.
Frequently Asked Questions
What is a probe query in GEO measurement?
A probe query is a query issued deliberately to test a specific aspect of AI citation behaviour — not to retrieve information for its own sake. Each probe query targets a known layer of the retrieval pipeline: whether a brand is recognised, whether proprietary vocabulary triggers citation, whether category queries pull commercial content, or whether the system handles a specific query type differently from another.
What is the difference between proprietary and category queries in GEO?
Proprietary queries ask about a specific framework, term, or concept owned by a single source — for example, “What is the GEO Stack?”. Category queries ask about a general topic where many sources compete — for example, “best AI SEO tools”. Proprietary queries reliably cite the originator when the term is clear; category queries usually cite aggregators, commercial domains, or the highest-authority source in the vertical, rarely the originator of a specific concept.
Why use branded queries if I already rank for them?
Ranking for a branded query on Google does not mean AI systems cite the brand for that query. The two indexes are separate. A branded probe query tests whether the AI has indexed the brand’s primary properties at all, whether it disambiguates the brand correctly, and whether it links to the official site or a third-party description. A brand can rank number one on Google and still be invisible in AI citation for the same term.
How many probe queries do you need for a meaningful test?
At least one query per probe type on the list, for a minimum of ten queries. More is better, but each query requires multiple iterations to control for measurement noise — ten queries at five iterations each is fifty measurements, which is typically the practical lower bound for detecting meaningful differences in citation rate between conditions.
Which probe query type is the most important?
No single type is the most important; each tests a different dimension of AI visibility. Proprietary queries test whether your framework vocabulary has been adopted by AI systems. Category queries test your commercial citation potential. Troubleshooting queries test your utility as a reference source. All ten types need to be in the test set, because treating any single type as the measure of “AI visibility” misses most of the signal.

