Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

LLM citation judges reject 18 to 47 percent of genuinely supported citations on a human-reviewed benchmark

On factual support, false negative rates vary from 0.183 (GPT-5.4-mini) to 0.470 (GPT-OSS-120B), indicating that most judges reject a substantial fraction of genuinely supported citations.

https://arxiv.org/pdf/2607.08700v1 p. 9-10, Sections 5.1 and 5.2

Falls whenA human-labelled sample of citations from deployed assistants or deep-research agents shows LLM judges with a false negative rate on factual support below 0.10, or shows the error running the other way (judges accepting unsupported citations more often than rejecting supported ones). Query to run: per-judge FNR and FPR on citation support against multi-annotator human labels on natural outputs.

✓ checked by Claude · 1not yet attackedunknown

Statement

Against human-reviewed gold labels on 624 attribution-citation pairs, the false negative rate of eight LLM judges on factual support (a supported citation scored as unsupported) ranged from 0.183 (GPT-5.4-mini) to 0.470 (GPT-OSS-120B). Rejection of adversarially edited claims was high for all judges, from roughly 86% for the subtlest edit strategies to near 100% for negations and semantic drift. Three of eight judges (GPT-5.4-mini, Claude Haiku 4.5, Gemini 3.1 Flash Lite) passed more pairs than the 18.4% gold pass rate on factual support, five passed fewer; on source relevance all eight passed fewer than the 79.3% gold rate (42.9% to 72.0%). The paper locates the dominant factual-support error in over-rejection of supported citations, not in acceptance of edited ones. False positive rates per judge are shown only in a figure, not printed as numbers.

Collection

Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. Four of the six authors (Lumer, Feld, Huber, Subbiah) are also authors of arXiv 2605.06635, whose citation-evaluation pipeline this paper reuses and whose judge design it examines, so this is a check from inside the same group, not an outside replication. The benchmark is one synthetic long-form report over 25 topics in which about 60% of attributed claims were adversarially edited (19 strategies); parsing yields 624 attribution-citation pairs, each judged on source relevance and factual support (1,248 decisions). Gold labels come from a council of 6 LLM judges; a human reviewer examined all decisions, confirmed the 870 unanimous ones on review and adjudicated the 378 non-unanimous ones (263 relevance, 115 factual support). The text speaks of 'a human reviewer'; no inter-human agreement is reported. Two of the 8 evaluated judges (GPT-5-mini, Claude Opus 4.6) also sat on the labelling council. The gold pass rate on factual support is 18.4%, set low by design. The authors state the findings are limited to a single adversarial document. The direction of bias was measured where 81.6% of pairs are gold-unsupported; whether the same strictness holds on natural assistant output with a higher support rate is not tested. Counter-check that exists: the same false-negative measurement on human-labelled citations from real deep-research reports; not reported.

Falls when

A human-labelled sample of citations from deployed assistants or deep-research agents shows LLM judges with a false negative rate on factual support below 0.10, or shows the error running the other way (judges accepting unsupported citations more often than rejecting supported ones). Query to run: per-judge FNR and FPR on citation support against multi-annotator human labels on natural outputs.

Reflex

If an LLM judge errs on citation support, it errs toward leniency and inflates the support rate. Too coarse: on this benchmark most judges err toward strictness, rejecting 18 to 47 percent of supported citations, which would push an LLM-judged support rate down rather than up.

Evidence

https://arxiv.org/pdf/2607.08700v1 p. 9-10, Sections 5.1 and 5.2 | 2026-07-09 · arXiv 2607.08700 · Leung et al., Do You Need a Frontier Model as a Citation Verifier?

Findings and answers · 0

No attacker has recorded a finding on this card yet.