Questions › Do the citations of AI research assistants support the claims they are attached to
DeepResearch Bench citation judge Gemini-2.5-Flash matched human support labels in 96 percent and not-support labels in 92 percent of 100 sampled pairs
judgment aligned with human ’support’ determinations in 96% of cases and with ’not support’ determinations in 92% of cases
https://arxiv.org/html/2506.11763v1 Appendix C, Judge LLM Selection for the FACT Framework
Falls whenA human re-annotation of at least several hundred statement-URL pairs drawn from the DeepResearch Bench release, with the split between 'support' and 'not support' reported, finds the Gemini-2.5-Flash verdict matching fewer than 90% of human 'support' labels or fewer than 85% of human 'not support' labels. It narrows if agreement differs by evaluated system, for example if it is lower on pairs from the Gemini systems. Query to run: per-label agreement between the FACT judge and independent human raters on the released pairs, per system.
Statement
On a randomly sampled set of 100 statement-URL pairs from the DeepResearch Bench tasks, the LLM judge of the FACT framework, Gemini-2.5-Flash, gave the same verdict as human annotators in 96% of the cases the humans had labelled 'support' and in 92% of the cases they had labelled 'not support'. This is the only human comparison the paper prints for the judge that produces its Citation Accuracy and Effective Citations columns. The paper does not print how the 100 pairs split between the two labels, how many annotators labelled each pair, or an interval; with 100 pairs in total, each percentage rests on fewer than 100 cases. The figure covers the support judgment only, not the judge's preceding step of extracting and deduplicating statement-URL pairs.
Collection
Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; the paper is an arXiv preprint without venue. The comparison was run by the benchmark's own authors to select the judge for their own framework (Appendix C, Judge LLM Selection), so producer and validator of the judge coincide; none of the evaluated systems is theirs, but the judge (Gemini-2.5-Flash) belongs to the same model family as two groups of systems in Table 1. The annotators are described only as human annotators; recruitment, number and instructions are not given, and no inter-annotator agreement is printed. A counter-check that exists and was not used: an outside human audit of the released statement-URL pairs; the paper reports none, and the sampled pairs with their human labels are not identified in the text.
Falls when
A human re-annotation of at least several hundred statement-URL pairs drawn from the DeepResearch Bench release, with the split between 'support' and 'not support' reported, finds the Gemini-2.5-Flash verdict matching fewer than 90% of human 'support' labels or fewer than 85% of human 'not support' labels. It narrows if agreement differs by evaluated system, for example if it is lower on pairs from the Gemini systems. Query to run: per-label agreement between the FACT judge and independent human raters on the released pairs, per system.
Reflex
The benchmark's LLM judge was validated against humans, so its citation accuracy column can be read like a human rating. Too coarse: the validation is 100 pairs labelled by the authors' own annotators, which leaves a not-support miss rate of 8 percent with a wide margin, enough to move scores that differ by a few points.
Evidence
https://arxiv.org/html/2506.11763v1 Appendix C, Judge LLM Selection for the FACT Framework | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench
Notes
Added in proposal 001. The figure appears on two existing cards of this document only inside Collection; as the sole printed human check of the judge it carries weight for every FACT number and is set out as its own card. Observed from the HTML rendering of v1 (arXiv lists only v1, submitted 13 Jun 2025); the sibling cards pin the PDF of the same version, where Appendix C is on p. 18-19.
Findings and answers · 0
No attacker has recorded a finding on this card yet.