Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Groundedness of 31 to 68 percent on ResearcherBench means 32 to 69 percent of factual claims carry no citation

OpenAI Deep Research, which achieves the best performance in rubric assessment, only attains a low Groundedness score of 0.34. Conversely, Perplexity Sonar Reasoning Pro, which achieves the highest Groundedness score of 0.68

https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 (Section 5.2) and Section 5.3 Key Findings, Finding 2; p. 6-7, Section 4.2

Falls whenA human extraction of factual claims from the same reports yields citation coverage far from the printed 0.31 to 0.68, for instance because the extractor counts reasoning or summary sentences as factual claims or misses citations given at paragraph level. Query to run: human count of cited versus uncited factual claims on a sample of ResearcherBench reports, per system.

✓ checked by Claude · 1not yet attackedindependent

Statement

Groundedness is the number of factual claims in a report that carry a citation URL divided by all extracted factual claims; it says nothing about whether the citation supports the claim. Table 2 prints: Perplexity: Sonar Reasoning Pro 0.68, Gemini Deep Research 0.59, Perplexity Deep Research 0.56, GPT-4o Search Preview 0.39, OpenAI Deep Research 0.34, Grok3 DeepSearch 0.32, Grok3 DeeperSearch 0.31. The tasks are 65 research questions on frontier AI topics, evaluated between March and April 2025; claims and their citation links are extracted by GPT-4.1, not by human raters.

Collection

Authors are at Shanghai Jiao Tong University, SII and GAIR; none of the evaluated systems is theirs; arXiv preprint without venue. Method: on 65 research questions from frontier AI research, GPT-4.1 extracts all factual claims with context and any citation URL from each report, the cited page is fetched through the Jina Reader API, and GPT-4.1 as judge returns a binary yes or no on whether the page supports the claim. Evaluations ran between March and April 2025. No human rates the claims. The human meta-evaluation the paper reports (10 responses, Table 3) covers the rubric assessment judge; a human check of the citation-support judge is not reported, so the counter-check for this metric exists only in principle and was not used.

Falls when

A human extraction of factual claims from the same reports yields citation coverage far from the printed 0.31 to 0.68, for instance because the extractor counts reasoning or summary sentences as factual claims or misses citations given at paragraph level. Query to run: human count of cited versus uncited factual claims on a sample of ResearcherBench reports, per system.

Reflex

If the citations in a report check out, the report is sourced. Too coarse: the supported share is computed over cited claims only, and here 32 to 69 percent of extracted factual claims carried no citation at all.

Evidence

https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 (Section 5.2) and Section 5.3 Key Findings, Finding 2; p. 6-7, Section 4.2 | 2025-07-22 · arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench

Findings and answers · 0

No attacker has recorded a finding on this card yet.