Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Eight LLM judges match human-reviewed labels on factual support of citations with F1 of 65 to 75 percent, none separable

every 95% confidence interval overlaps on this dimension (from 0.649 [.56,.72] to 0.750 [.68,.82]), so no model is statistically distinguishable

https://arxiv.org/pdf/2607.08700v1 p. 6-8, Table 2, Sections 4.2 and 4.3

Falls whenAn independent group labels the same 624 pairs, or citations from real deep-research reports, with two or more human annotators and finds judge-versus-human F1 on factual support above 0.9 (or kappa above 0.8) for a current judge, or finds that the single-reviewer gold labels disagree with double-annotated labels often enough to move the 0.649 to 0.750 range. Query to run: judge-human agreement on citation factual support, multi-annotator gold, natural outputs.

✓ checked by Claude · 1not yet attackedunknown

Statement

Eight off-the-shelf LLM judges from three model families (Anthropic, Google, OpenAI) were scored against human-reviewed gold labels on 624 attribution-citation pairs, 1,248 rubric decisions in total. On factual support (does the cited source support the claim) pass-class F1 ranged from 0.649 [.56,.72] (GPT-OSS-120B) to 0.750 [.68,.82] (Claude Opus 4.6, kappa 0.701), with GPT-5-mini at 0.710 [.64,.78] (kappa 0.649); all 95% bootstrap intervals overlap. On source relevance F1 ranged from 0.700 (Claude Sonnet 4.6) to 0.908 [.89,.93] (GPT-5-mini, kappa 0.636). On the 378 human-adjudicated disagreement cases the ranking changed: factual-support F1 was 0.780 for GPT-5.4-mini and 0.672 for Claude Opus 4.6. The rater of record is a human reviewer over an LLM council's labels; the judged material is a synthetic adversarial report, not output of deployed assistants.

Collection

Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. Four of the six authors (Lumer, Feld, Huber, Subbiah) are also authors of arXiv 2605.06635, whose citation-evaluation pipeline this paper reuses and whose judge design it examines, so this is a check from inside the same group, not an outside replication. The benchmark is one synthetic long-form report over 25 topics in which about 60% of attributed claims were adversarially edited (19 strategies); parsing yields 624 attribution-citation pairs, each judged on source relevance and factual support (1,248 decisions). Gold labels come from a council of 6 LLM judges; a human reviewer examined all decisions, confirmed the 870 unanimous ones on review and adjudicated the 378 non-unanimous ones (263 relevance, 115 factual support). The text speaks of 'a human reviewer'; no inter-human agreement is reported. Two of the 8 evaluated judges (GPT-5-mini, Claude Opus 4.6) also sat on the labelling council. The gold pass rate on factual support is 18.4%, set low by design. The authors state the findings are limited to a single adversarial document. Counter-check that exists: independent double human annotation of the pairs and a rerun on naturally occurring assistant output; neither is reported.

Falls when

An independent group labels the same 624 pairs, or citations from real deep-research reports, with two or more human annotators and finds judge-versus-human F1 on factual support above 0.9 (or kappa above 0.8) for a current judge, or finds that the single-reviewer gold labels disagree with double-annotated labels often enough to move the 0.649 to 0.750 range. Query to run: judge-human agreement on citation factual support, multi-annotator gold, natural outputs.

Reflex

LLM judges agree well with humans, so an LLM-judged support rate can be read like a human-rated one. Too coarse: on factual support specifically, agreement with human-reviewed labels here is F1 0.65 to 0.75 with kappa 0.58 to 0.70, lower than on relevance, and judge rankings change on the hard cases.

Evidence

https://arxiv.org/pdf/2607.08700v1 p. 6-8, Table 2, Sections 4.2 and 4.3 | 2026-07-09 · arXiv 2607.08700 · Leung et al., Do You Need a Frontier Model as a Citation Verifier?

Findings and answers · 0

No attacker has recorded a finding on this card yet.