Questions › Do the citations of AI research assistants support the claims they are attached to
Effective citations per task range from about 4 to 111 across systems on DeepResearch Bench
Notably, Gemini-2.5-Pro Deep Research achieves an average of 111.21 effective citations in its final reports, significantly outperforming other models.
https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E
Falls whenA recount on the released reports with human support judgments gives per-task counts of supported citations that differ from the printed column by enough to close the spread between 4.35 and 111.21, or shows the deduplication step of the Judge LLM removes or keeps pairs unevenly across systems. Query to run: supported statement-URL pairs per task, per system, on the DeepResearch Bench release.
Statement
Average Effective Citations per Task (E. Cit.) is the number of statement-URL pairs judged 'support', summed over all 100 tasks and divided by the number of tasks. Table 1 prints: Gemini-2.5-Pro Deep Research 111.21, OpenAI Deep Research 40.79, Gemini-2.5-Pro-Grounding 32.88, Claude-3.7-Sonnet w/Search 32.48, Perplexity Deep Research 31.26, Grok Deeper Search 8.15, GPT-4o-Search-Preview 4.79, GPT-4.1-mini w/Search 4.35. The quantity is a count of supported citations, separate from Citation Accuracy (the supported share): the system with the highest count, Gemini-2.5-Pro Deep Research, has 81.44 accuracy, and Claude-3.5-Sonnet w/Search with 94.04 accuracy has 9.78 effective citations. Support is judged by an LLM (Gemini-2.5-Flash).
Collection
Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; none of the evaluated systems is theirs, and the paper is an arXiv preprint without venue. Method (FACT framework): a Judge LLM, Gemini-2.5-Flash, extracts statement-URL pairs from each report and deduplicates them, the cited page is fetched through the Jina Reader API, and the same judge returns a binary 'support' or 'not support' per pair. No human rates the full set. The counter-check that exists and was used: the judge was compared with human annotators on 100 randomly sampled statement-URL pairs and agreed with human 'support' determinations in 96% of cases and with 'not support' determinations in 92%. The same judge family (Gemini) also rates the Gemini systems in the table. Outputs were collected between April 1 and May 13, 2025.
Falls when
A recount on the released reports with human support judgments gives per-task counts of supported citations that differ from the printed column by enough to close the spread between 4.35 and 111.21, or shows the deduplication step of the Judge LLM removes or keeps pairs unevenly across systems. Query to run: supported statement-URL pairs per task, per system, on the DeepResearch Bench release.
Reflex
A system with a higher share of supported citations gives the reader more supported material. Too coarse: the supported share and the supported count are separate columns here, and the count differs more than twentyfold between two systems whose supported shares lie about three points apart.
Evidence
https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench
Findings and answers · 0
No attacker has recorded a finding on this card yet.