Questions › Do the citations of AI research assistants support the claims they are attached to
Across 10 models on DRBench 3 to 13 percent of citation URLs are hallucinated and 5 to 18 percent do not resolve
Non-resolving URL rates across models range from 5.4% [3.0, 7.7] (gpt-4.1) to 18.5% [17.8, 19.2] ( gemini-2.5-pro-deepresearch), while hallucinated URL rates range from 3.0% [1.7, 4.4] ( claude-3-5-sonnet-with-search) to 13.3% [12.7, 13.9] (gemini-2.5-pro-deepresearch).
https://arxiv.org/pdf/2604.03173v1 p. 4, Table 2 and Section 4.1; p. 5, Sections 4.1 to 4.3; p. 2-3, Sections 2 and 3.3 for definitions
Falls whenA re-check of the released URL lists with a real browser instead of HEAD requests finds most non-resolving URLs live (bot-blocking rather than dead pages), pushing the upper rates well below 13.3 and 18.5%; or a wider archive lookup shows most URLs classed as hallucinated had existed. The range narrows if the single outlier gemini-2.5-pro-deepresearch is set aside (next highest 8.8% hallucinated, 10.1% non-resolving). The anchor is out of scope for any claim about whether resolving links support their sentences. Query to run: headless-browser liveness plus Wayback and Common Crawl lookup over the released DRBench URL set.
Statement
Link validity only, not support: the study checks whether cited URLs exist, which it calls a logically prior question to whether the source supports the claim. For 10 models from Google, OpenAI and Anthropic on DRBench (100 research queries in Chinese and English, outputs pre-collected by the benchmark authors, 296 to 11,309 URLs per model), the share of citation URLs that do not resolve (HTTP 4xx or 5xx, connection error or timeout, HTTP 403 excluded) ranges from 5.4% [3.0, 7.7] (gpt-4.1) to 18.5% [17.8, 19.2] (gemini-2.5-pro-deepresearch); the share classified as hallucinated (non-resolving and no Wayback Machine snapshot at any time) ranges from 3.0% [1.7, 4.4] (claude-3-5-sonnet-with-search) to 13.3% [12.7, 13.9] (gemini-2.5-pro-deepresearch). Pooled, the two deep-research agents have 10.7% [10.2, 11.2] hallucinated and 16.2% [15.7, 16.8] non-resolving URLs against 4.8% [4.3, 5.2] and 6.8% [6.2, 7.3] for the eight search-augmented models. On ExpertQA (2,177 expert questions, 32 fields, 168,021 URLs from claude-sonnet-4-5, gemini-2.5-pro and gpt-5.1) the overall non-resolving rate is 8.22% [8.09, 8.36], by field from 5.4% [4.9, 5.9] (Business) to 11.4% [8.1, 14.6] (Theology). Brackets are bootstrap 95% intervals. The measurement is automated (HTTP requests plus Wayback Machine API); there is no human or LLM rater.
Collection
Authors are at the University of Pennsylvania; funded by DARPA's SciFy program; arXiv preprint marked as under review, not peer reviewed. No affiliation with an evaluated vendor appears in the document. Each URL gets an HTTP HEAD request (GET fallback) with a browser-like User-Agent; non-resolving URLs are looked up in the Wayback Machine, no snapshot means hallucinated, a snapshot means stale. The authors call the rates conservative lower bounds: 403 responses (6.6 to 17.0% of ExpertQA URLs) are excluded and Wayback coverage is incomplete. Counter-checks used: a headless-browser audit (403 responses 99.7% live; 89% of UNKNOWN responses live or blocked) and a sensitivity analysis for Reddit URLs (treating all as non-resolving raises gpt-5.1 from 8.47% to 26.7%). 13 of 23 DRBench models were excluded, three of them for 100% hallucination rates read as no real web retrieval. The abstract gives 53,090 DRBench URLs while the ten per-model counts in Table 1 sum to 23,269; the text does not explain the difference. Liveness is a point-in-time measurement. Tool and data are announced under MIT license.
Falls when
A re-check of the released URL lists with a real browser instead of HEAD requests finds most non-resolving URLs live (bot-blocking rather than dead pages), pushing the upper rates well below 13.3 and 18.5%; or a wider archive lookup shows most URLs classed as hallucinated had existed. The range narrows if the single outlier gemini-2.5-pro-deepresearch is set aside (next highest 8.8% hallucinated, 10.1% non-resolving). The anchor is out of scope for any claim about whether resolving links support their sentences. Query to run: headless-browser liveness plus Wayback and Common Crawl lookup over the released DRBench URL set.
Reflex
Models with web search no longer invent links; fake references are a problem of offline chatbots. Too coarse: with search switched on, 3 to 13 percent of cited URLs still had no trace of ever having existed, and pooled deep-research agents did worse than plain search-augmented models.
Evidence
https://arxiv.org/pdf/2604.03173v1 p. 4, Table 2 and Section 4.1; p. 5, Sections 4.1 to 4.3; p. 2-3, Sections 2 and 3.3 for definitions | 2026-04-03 · arXiv 2604.03173 · Rao, Wong, Callison-Burch, Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
Findings and answers · 0
No attacker has recorded a finding on this card yet.