Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Five deep research systems score 69 to 86 percent faithfulness of cited claims on ResearcherBench under an LLM judge

Faithfulness Groundedness Deep Research System OpenAI Deep Research 0.7032 0.84 0.34 Gemini Deep Research 0.6929 0.86 0.59 Grok3 DeepSearch 0.4414 0.69 0.32 Grok3 DeeperSearch 0.4398 0.80 0.31 Perplexity Deep Research 0.4800 0.85 0.56

https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 and Section 5.2; p. 6-7, Section 4.2

Falls whenA human-rated sample of the URL-claim-context triplets shows GPT-4.1's yes decisions are wrong often enough to move the 0.80 to 0.86 cluster, or a replication on questions outside AI research gives materially lower faithfulness for the same systems. It narrows to 'conditional on a citation being present' by construction: read with the groundedness anchor of the same table. Query to run: human audit of support decisions on ResearcherBench factual-assessment outputs.

✓ checked by Claude · 2not yet attackedindependent

Statement

Faithfulness is the number of cited claims judged supported by their cited URL divided by the number of cited claims; claims without a citation are outside this metric. Table 2 prints for the deep research systems: Gemini Deep Research 0.86, Perplexity Deep Research 0.85, OpenAI Deep Research 0.84, Grok3 DeeperSearch 0.80, Grok3 DeepSearch 0.69; for the LLMs with search tools: GPT-4o Search Preview 0.86, Perplexity: Sonar Reasoning Pro 0.62. The tasks are 65 research questions on frontier AI topics, evaluated between March and April 2025; extraction and the support judgment are both made by GPT-4.1, not by human raters.

Collection

Authors are at Shanghai Jiao Tong University, SII and GAIR; none of the evaluated systems is theirs; arXiv preprint without venue. Method: on 65 research questions from frontier AI research, GPT-4.1 extracts all factual claims with context and any citation URL from each report, the cited page is fetched through the Jina Reader API, and GPT-4.1 as judge returns a binary yes or no on whether the page supports the claim. Evaluations ran between March and April 2025. No human rates the claims. The human meta-evaluation the paper reports (10 responses, Table 3) covers the rubric assessment judge; a human check of the citation-support judge is not reported, so the counter-check for this metric exists only in principle and was not used.

Falls when

A human-rated sample of the URL-claim-context triplets shows GPT-4.1's yes decisions are wrong often enough to move the 0.80 to 0.86 cluster, or a replication on questions outside AI research gives materially lower faithfulness for the same systems. It narrows to 'conditional on a citation being present' by construction: read with the groundedness anchor of the same table. Query to run: human audit of support decisions on ResearcherBench factual-assessment outputs.

Reflex

Deep research systems cite sources that do not support their claims most of the time. Too coarse: where a claim carried a citation, this benchmark's LLM judge found it supported in 80 to 86 percent of cases for four of five deep research systems; the measure is silent on the uncited claims.

Evidence

https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 and Section 5.2; p. 6-7, Section 4.2 | 2025-07-22 · arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench Plan 02b: table-only quote, Table 2, p. 8; no sentence in the paper carries the faithfulness values.

Findings and answers · 0

No attacker has recorded a finding on this card yet.