Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Thirteen LLMs asked for computer science references produced invalid citations at rates from 14 to 95 percent

hallucination rates spanning from 14.23% (DeepSeek) to 94.93% (Hunyuan), a roughly 6.7× difference

https://arxiv.org/pdf/2602.06718v2 p. 5-7, Sections IV.A and V.A, Table II

Falls whenAn independent re-verification of a random sample of the citations labelled invalid finds materially more than the 2 percent false positives the authors report, or the same prompts run in the vendors' own search-enabled products rather than through an API aggregator give rates far below Table II. Query to run: re-verify 400 randomly drawn invalid-labelled citations from the GhostCite benchmark release by hand, and rerun the Section B prompt in the consumer products with search on.

✓ checked by Claude · 1not yet attackedindependent

Statement

Thirteen LLMs accessed through the OpenRouter API were prompted across 40 computer science domains, in batches of 10, 20 or 30, with and without online search plus chain-of-thought, to return references in a fixed JSON schema; of 331,809 extracted citations, 166,876 (50.29%) could not be verified by the authors' CiteVerifier tool against bibliographic databases and web search, and the per-model invalid share ranges from 14.23% +/- 1.65 (DeepSeek) through 21.84% (Claude 4), 50.92% (GPT-5) and 59.47% (Gemini) to 94.93% +/- 1.29 (Hunyuan), with online search showing no consistent effect across models. The authors describe these rates as a controlled baseline, not an estimate of real-world prevalence. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.

Collection

Academic authors at Nankai University and Tsinghua University; arXiv preprint (cs.CR), version 2. No stake in the evaluated vendors is declared or visible; the authors built and publish the verification tool the rates depend on. Method: 22,800 API interactions, 375,440 requested citations, automated verification by title similarity against academic databases with web-search and LLM-reparse fallbacks; unparseable outputs are excluded from the rates. The counter-check that exists was used: the authors manually checked random samples of 400 valid and 400 invalid verdicts and report 100% and 98% (392/400) agreement. Search and reasoning settings were switched through a third-party API aggregator, not through the vendors' own products.

Falls when

An independent re-verification of a random sample of the citations labelled invalid finds materially more than the 2 percent false positives the authors report, or the same prompts run in the vendors' own search-enabled products rather than through an API aggregator give rates far below Table II. Query to run: re-verify 400 randomly drawn invalid-labelled citations from the GhostCite benchmark release by hand, and rerun the Section B prompt in the consumer products with search on.

Reflex

Frontier models in 2026 rarely invent references. Too coarse: in this benchmark GPT-5 and Gemini returned unverifiable references in 50.92 and 59.47 percent of cases and the range across 13 models is 14.23 to 94.93 percent; the figure concerns existence only and comes from a list-generation prompt, not from drafting.

Evidence

https://arxiv.org/pdf/2602.06718v2 p. 5-7, Sections IV.A and V.A, Table II | 2026-05-14 (v2; v1 2026-02-06) · arXiv 2602.06718 · Xu et al., GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

Findings and answers · 0

No attacker has recorded a finding on this card yet.