Questions › Do the citations of AI research assistants support the claims they are attached to
Four deep research agents score 78 to 90 percent citation accuracy on DeepResearch Bench under an LLM judge
Grok Deeper Search 40.24 37.97 35.37 46.30 44.05 83.59 8.15 Perplexity Deep Research 42.25 40.69 39.39 46.40 44.28 90.24 31.26 Gemini-2.5-Pro Deep Research 48.88 48.53 48.50 49.18 49.44 81.44 111.21 OpenAI Deep Research 46.98 46.87 45.25 49.27 47.14 77.96 40.79
https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E
Falls whenA human-rated audit of the released statement-URL pairs finds support rates for the four deep research agents well below the printed 77.96 to 90.24, or finds the judge's 96% / 92% agreement does not hold on a larger sample than 100 pairs. It narrows if a re-run on later system versions moves the range. Query to run: human support audit of FACT statement-URL pairs from the DeepResearch Bench release, per system.
Statement
On 100 PhD-level research tasks across 22 fields, Citation Accuracy (C. Acc.) is the share of unique statement-URL pairs in a report for which the fetched page is judged to support the statement, computed per task and averaged over tasks; a task with no citable statement counts as 0. Table 1 prints for the four deep research agents: Perplexity Deep Research 90.24, Grok Deeper Search 83.59, Gemini-2.5-Pro Deep Research 81.44, OpenAI Deep Research 77.96. The deep research agents do not lead this column. For the twelve LLMs with search tools Table 1 prints: Claude-3.5-Sonnet w/Search 94.04, Claude-3.7-Sonnet w/Search 93.68, GPT-4o-Search-Preview 88.41, GPT-4.1 w/Search 87.83, GPT-4o-Mini-Search-Preview 84.98, GPT-4.1-mini w/Search 84.58, Gemini-2.5-Flash-Grounding 81.92, Gemini-2.5-Pro-Grounding 81.81, Perplexity-Sonar-Pro 78.66, Perplexity-Sonar 74.42, Perplexity-Sonar-Reasoning 48.67, Perplexity-Sonar-Reasoning-Pro 39.36. Nine of the twelve score above the lowest deep research agent (77.96) and the two Claude models score above the highest (90.24). The support judgment is made by an LLM judge (Gemini-2.5-Flash), not by human raters.
Collection
Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; none of the evaluated systems is theirs, and the paper is an arXiv preprint without venue. Method (FACT framework): a Judge LLM, Gemini-2.5-Flash, extracts statement-URL pairs from each report and deduplicates them, the cited page is fetched through the Jina Reader API, and the same judge returns a binary 'support' or 'not support' per pair. No human rates the full set. The counter-check that exists and was used: the judge was compared with human annotators on 100 randomly sampled statement-URL pairs and agreed with human 'support' determinations in 96% of cases and with 'not support' determinations in 92%. The same judge family (Gemini) also rates the Gemini systems in the table. Outputs of the four deep research agents were collected between April 1 and May 8, 2025 (Appendix D: OpenAI April 1 to May 8, Perplexity April 1 to April 29, Gemini and Grok April 27 to April 29); outputs of the LLMs with search tools were collected later, between May 11 and May 13, 2025, so the two groups in the column were not sampled in the same weeks.
Falls when
A human-rated audit of the released statement-URL pairs finds support rates for the four deep research agents well below the printed 77.96 to 90.24, or finds the judge's 96% / 92% agreement does not hold on a larger sample than 100 pairs. It narrows if a re-run on later system versions moves the range. Query to run: human support audit of FACT statement-URL pairs from the DeepResearch Bench release, per system.
Reflex
Deep research agents mostly cite pages that do not back their sentences. Too coarse: under this benchmark's LLM judge, 78 to 90 percent of extracted statement-URL pairs were judged supported, while two search-tool models in the same table sit at 39 and 49 percent.
Evidence
https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench
Notes
Revised in proposal 001 after a check against the HTML version of arXiv 2506.11763v1 (Table 1, Section 4.2.2, Appendix C, D, E). All numbers already on the card match the print. Two changes: the statement printed the search-tool column only as a range with two named values, which hid that most plain search-tool LLMs score at or above the deep research agents on C. Acc.; the full column is now printed. The collection window was given as one span (April 1 to May 13); Appendix D prints separate windows for the agents and for the search-tool LLMs, now stated. The paper does not say how the 100 pairs of the human comparison split between 'support' and 'not support'. The Reflex section still names only the two lowest search-tool values; left unchanged for the owner to decide.
Findings and answers · 0
No attacker has recorded a finding on this card yet.