TL;DR
I ran citation-check blind against A. L. MacFarland’s 50-case USO Gold Test Set, froze every classification, and only then opened the answer key. Where the instrument and protocol overlap — MENTIONED and CITED — we agreed on 78 of 79 determinable comparisons. The one disagreement exposed a construct boundary I had not properly encoded: my instrument detected the name “Atlas”, while the key required enough evidence to establish which Atlas was meant. The test also quantified what citation-check cannot currently measure: across its out-of-scope states, the key had concrete answers for 115 slots the instrument declined, including 42 OBSERVED signals. That gives me something more useful than a headline agreement rate: a clearer definition of what the instrument measures, where that definition breaks, and exactly how much sits beyond its current coverage. This is the first external instrument benchmark citation-check has faced.
1. A stranger hands you an exam
In August, A. L. MacFarland offered me something I rarely get in this kind of work: someone else’s fixed, versioned, 50-case gold test set, with a published answer key I promised not to look at.
His USO Evidence Protocol keeps coming back to one question:
Does the claim stay inside what was actually measured?
The Gold Test Set turns that question into 50 concrete cases.
The deal was simple. I would run all 50 through citation-check exactly as configured. No tuning individual cases. No changing definitions halfway through. I would freeze the classifications first, then open the key.
“Blind” here is on the honour system. The key is public. The discipline is simply not looking.
I wanted to do it because citation-check had never faced an external instrument benchmark — tested against something I did not build.
Every measurement instrument contains assumptions about what counts, even when those assumptions are buried in code. Testing mine against someone else’s framework was a good way to find out where those assumptions stopped matching.
It worked.
2. First, what is citation-check actually allowed to say?
I declared this before the run because defining scope after seeing the results is just moving the goalposts with a timestamp.
citation-check natively produces a presence signal for each state:
present, absent, or not established.
NOT_ESTABLISHED is never silently converted into absence.
That gives the instrument coverage over the protocol’s MENTIONED and CITED states, the two that fall within extractability.
It cannot determine RETRIEVED, INCLUDED, RECOMMENDED or ADVERSELY_TREATED. It also does not assess USO reporting conformance.
MacFarland and I were clear about what that meant before the test: unsupported states would not be treated as instrument failures.
The run was testing two separate things.
First, how well do the two systems agree where their measurement surfaces overlap?
Second, does my instrument refuse cleanly where they do not?
Run details: citation-check E101-v-next, 6 August 2026 version. Static-case pass using the supplied response text and visible_citations arrays. No live queries. Run on 20 August 2026. All 50 cases returned, with zero exclusions and zero processing failures.
One more label, because labels are the whole subject here. This experiment is registered as E108, and the registration is post-hoc. The controls were declared in dated correspondence before the key was opened: blind to the key, freeze before opening, no per-case tuning. But no pre-registration was deposited before the run. A registered experiment and a pre-registered one are different claims. This is the former.
3. Where we overlap: 78 of 79
Across MENTIONED and CITED, 79 comparisons were determinable.
citation-check and the key agreed on 78.
That sounds impressive until you look at what the number actually contains.
| State | Key-positive | Key-negative | Disagreements |
|---|---|---|---|
| MENTIONED (40) | 37/37 agree | 2/3 agree | 1, G012 |
| CITED (39) | 3/3 agree | 36/36 agree | 0 |
There are two things I would not want anyone to lose in the headline number.
The first is that 78/79 is agreement within two states. It is not an instrument-accuracy claim.
Look at CITED in particular. There are only three positive cases.
Yes, citation-check agreed with the key on all three. But three positives are nowhere near enough to claim that the instrument has demonstrated accurate citation detection. Most of that apparently perfect result comes from 36 negative cases.
A citation tracker proving that it can recognise absence is not the same thing as proving that it can reliably detect citations.
So I am reporting the number because it is what this run produced, but I am not asking it to carry more weight than the sample can support.
The second thing is more interesting.
The instrument did not miss a single key-positive MENTIONED or CITED case.
Its only disagreement in the entire comparable set went the other way.
It detected something the key would not accept.
That was G012.
4. G012: when “Atlas” is not necessarily Atlas
G012 contains a response with the word Atlas in it.
citation-check returned MENTIONED: OBSERVED.
The USO key returned NOT_ESTABLISHED.
Why?
Because two companies share the name Atlas, and the response gives us no way to establish which one it means.
That single case exposed a definition I had not made explicit enough in my own instrument.
citation-check was detecting name presence. The key was requiring entity identity.
Those are not the same measurement.
And that gap matters far beyond this test set.
If your company shares its name with another company, a place, a product, a person or something from mythology, what exactly does your AI visibility tool mean when it tells you that your brand was “mentioned”?
If it has not resolved the entity, it may simply be counting the string.
That means some brand visibility metrics may contain namesake mentions that do not belong to the entity being measured.
G012 is only one case. I am not generalising an error rate from it.
But the construct boundary it exposed is real enough that the instrument now has to deal with it.
Sometimes the best result from a benchmark is not the score. It is discovering that you and the benchmark were answering slightly different questions.
5. Then came the number I initially described wrongly
This is the part where the test became more useful than I expected.
citation-check returned NOT_ESTABLISHED on all 160 out-of-scope slots: four states across the 40 scorable cases.
My first summary described that as the instrument “preserving indeterminacy.”
That was too flattering.
It made it sound as though the key also considered those 160 slots unknowable.
It did not.
The key gives concrete adjudications for 115 of them:
73 NOT_OBSERVED
42 OBSERVED
It returns NOT_ESTABLISHED on the remaining 45.
So the correct interpretation is not “160 agreements on indeterminacy.”
It is this:
45 slots where the key itself says NOT_ESTABLISHED and the instrument also refuses to determine the state.
115 slots where the key has a concrete answer but citation-check cannot reach it.
And inside those 115 are 42 OBSERVED signals that a fuller instrument could potentially capture.
That gives me three much more useful numbers:
45 matched indeterminate states.
115 concretely adjudicated slots beyond the instrument’s coverage. That is the full evaluative reach a fuller instrument could cover.
42 of those 115, the OBSERVED slots, are positive-detection headroom specifically.
That is a much less flattering description of citation-check than my first one.
It is also much more useful.
A measurement instrument refusing to answer outside its scope is good behaviour. But refusal is not free. You need to know what information you are giving up by refusing.
For the first time, this test let me put a number on that cost.
6. NOT_APPLICABLE is not NOT_ESTABLISHED
There is another distinction worth keeping clean.
The ten reporting-integrity cases, G041 to G050, do not define observation states at all. They sit entirely outside what citation-check is designed to assess.
Those returned NOT_APPLICABLE.
That is different from NOT_ESTABLISHED.
NOT_APPLICABLE means the measurement does not belong here.
NOT_ESTABLISHED means the state belongs to the measurement surface, but the available evidence does not support a determination.
Collapsing those together would recreate at the summary level exactly the problem the tri-state design is supposed to prevent.
Different unknowns need different labels.
7. G040 exposed another boundary: refusal versus missing evidence
G040 is a truncated response.
citation-check returned CITED: NOT_ESTABLISHED because it will not turn “I cannot see the complete response” into “there was no citation.”
The key says NOT_OBSERVED.
I sent that one back to MacFarland rather than trying to argue my preferred interpretation into the benchmark.
His ruling clarified something useful.
The instrument was not refusing to assess evidence that existed. The evidence required to resolve the claim was simply not present in the observable surface.
So the determination stands as NOT_ESTABLISHED, with the reason recorded as required evidence not observed.
And importantly, G040 stays outside the comparable-agreement calculation. It is forced into neither agreement nor disagreement.
That distinction is now feeding back into citation-check’s truncation handling.
There are two different situations hiding behind the same apparent “unknown”:
The instrument cannot determine this state.
and
The evidence needed to determine this state is missing.
Most measurement pipelines make those look identical.
They are not.
8. What did this instrument benchmark actually establish?
Something fairly narrow, which is exactly how I want to report it.
It establishes the behaviour of one version of one instrument against 50 static cases from one pinned test set.
Within the two determinable states, MENTIONED and CITED, agreement was 78 of 79 comparable cases.
The single disagreement exposed a name-presence versus entity-identity boundary.
And outside those states, the test quantified something I previously only described qualitatively: citation-check’s coverage boundary.
What does it not establish?
It does not validate the USO protocol.
It does not show that these results generalise to live queries.
It does not establish citation-detection accuracy from three positive CITED cases.
And it tells us nothing about citation-check’s performance on RETRIEVED, INCLUDED, RECOMMENDED or ADVERSELY_TREATED beyond the fact that the instrument does not currently measure them.
That last limitation now has a number attached to it.
115 concrete adjudications outside the instrument’s reach, including 42 observed signals.
For me, that may be the most useful result of the whole exercise.
I went into someone else’s test expecting to learn whether my instrument agreed with his key.
I came out with something better: a clearer definition of what my instrument is actually measuring, one construct it needs to tighten, and a quantified map of what it still cannot see.
MacFarland is publishing his own independent response looking at what this run does and does not support about the protocol. It will be linked here when it is live. We drafted separately, and neither of us reviewed the other’s piece before publication.
That independence is part of the test.
Sources: A. L. MacFarland, USO Evidence Protocol v0.9.5, DOI: 10.5281/zenodo.21852690, CC BY 4.0; USO Gold Test Set v0.1, pinned to the v0.9.5 release; internal denominator post; frozen citation-check return and key-comparison artifacts, DOI: 10.5281/zenodo.22083031.
Have questions about this topic? Contact The GEO Lab · Return to homepage

