Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Deep research agents reach 50 to 79 percent citation accuracy and the best case GPT-5 leaves one in eight statements unsupported

although their citation accuracy has dropped to the borderline range (79.1% and 72.3%)

https://arxiv.org/pdf/2509.04499v1 p. 9, Table 1 and Section 4 Deep Research Agents; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation

Falls whenA human audit of the deep-research outputs puts GPT-5 Deep Research citation accuracy at or above 90% (the paper's own threshold for acceptable), or shows the other agents' unsupported rates of 53.6 to 97.5% to be an artefact of the relevance filter and the roughly 15% unscrapeable sources. It falls as a best-case bound if a later deep-research system measured with the same definitions exceeds 79.1%. Query to run: human per-citation support labels on GPT-5(DR) and PPLX(DR) reports for the same 303 queries, plus a recomputation of Table 1 from the matrices to settle Gemini 50.3 versus 40.3.

✓ checked by Claude · 2not yet attackedpositioned

Statement

Same audit, deep-research configurations, 303 queries each, results as of 27 August 2025. Citation accuracy (fraction of statement citations whose cited source supports the statement) in Table 1: GPT-5 Deep Research 79.1%, You.com Deep Research 72.3%, Copilot Think Deeper 62.1%, Perplexity Deep Research 58.0%, Gemini Deep Research 50.3%; GPT-5 in web-search mode, listed in the same table, 31.4%. Unsupported statements (fraction of query-relevant statements supported by none of the listed sources): GPT-5 Deep Research 12.5%, Gemini 53.6%, GPT-5 web search 58.9%, You.com 74.6%, Copilot 90.2%, Perplexity 97.5%. Citation thoroughness was 87.5% and 83.5% for GPT-5 and You.com deep research and 9.1% to 27.1% for the other three deep-research systems. The share of statements judged relevant to the query was 87.5% for GPT-5 Deep Research and 12.4% to 45.5% for the other deep-research systems, so their unsupported rates rest on a small relevant subset. Ratings are by an LLM judge (GPT-5 by default), validated on 100 manual support labels at Pearson 0.62. The running text gives Gemini's citation accuracy as 40.3% where Table 1 prints 50.3%.

Collection

Authors: five at Salesforce AI Research, one at Microsoft Research; arXiv preprint, not peer reviewed. Microsoft is the vendor of one evaluated system (Bing Copilot), so independence is recorded as positioned; Salesforce is the vendor of none of the evaluated systems. Browser scripts extracted answer text, citations and source URLs from the public web interfaces; Jina Reader scraped the source text; an LLM judge decomposed answers into statements and filled the factual-support matrix (on the order of 80,000 support judgements). The support prompt in Appendix E returns full, partial or none, and the paper does not say how partial is turned into the binary matrix; Section 3 names GPT-5 as default judge while Appendix E names GPT-4. Counter-check that exists and was used: two hired annotators labelled 100 support tasks (Pearson 0.62 against the judge, called moderate agreement by the authors). No per-system human audit of citation accuracy is reported. The authors name reliance on an LLM judge as a limiting factor. The query set is announced as released. Deep-research answers are long (23.9 to 141.6 statements, 3.6 to 57.2 sources on average), and roughly 15% of source URLs could not be scraped and were excluded from support calculations.

Falls when

A human audit of the deep-research outputs puts GPT-5 Deep Research citation accuracy at or above 90% (the paper's own threshold for acceptable), or shows the other agents' unsupported rates of 53.6 to 97.5% to be an artefact of the relevance filter and the roughly 15% unscrapeable sources. It falls as a best-case bound if a later deep-research system measured with the same definitions exceeds 79.1%. Query to run: human per-citation support labels on GPT-5(DR) and PPLX(DR) reports for the same 303 queries, plus a recomputation of Table 1 from the matrices to settle Gemini 50.3 versus 40.3.

Reflex

Deep research modes read many more sources and therefore ground their reports better than ordinary AI search. Too coarse: only one of five deep-research systems beat the search engines on both measures here; the best case still had about one in five citations not supporting its sentence, and four systems had most relevant statements supported by none of their listed sources.

Evidence

https://arxiv.org/pdf/2509.04499v1 p. 9, Table 1 and Section 4 Deep Research Agents; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation | 2025-09-02 · arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Plan 02b: the 50.3 for Gemini DR stands only in Table 1, p. 9; the Section 4 text gives 40.3 for the same system.

Findings and answers · 0

No attacker has recorded a finding on this card yet.