Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Audit of 111 million references in arXiv, bioRxiv, SSRN and PMC papers estimates 146,932 non-existent citations in 2025

Assuming these trends persisted through year end, these four corpora, which cover only a fraction of the scientific literature, would include 146,932 hallucinated citations in 2025 alone.

https://arxiv.org/pdf/2605.07723v1 p. 3-5, Estimating hallucinated references at scale and Results

Falls whenManual verification of a random sample of post-2023 unmatched references finds that most exist (indexing lag, non-English or grey literature), or the pre-LLM baseline re-estimated with the same pipeline on later database snapshots rises to the post-2023 level. Query to run: hand-check of sampled unmatched references per corpus and year against publisher records. The anchor narrows, without falling, if the full-year 2025 data replace the August extrapolation.

✓ checked by Claude · 1not yet attackedindependent

Statement

Downstream context, not a measurement of any assistant: an audit of 111 million references in 2.5 million papers (arXiv Jan 2020 to Aug 2025, bioRxiv, SSRN, and a 10% sample of PubMed Central) checks whether each cited title exists in Semantic Scholar, OpenAlex or Google Scholar. The excess of unmatched references over the pre-LLM baseline reached 0.39% (arXiv), 0.21% (bioRxiv), 1.91% (SSRN) and 0.27% (PMC) of references as of August 2025, with monthly excess counts of 3,353, 478, 767 and 8,140; extrapolating these to year end gives 146,932 non-existent citations in the four corpora in 2025. The quantity is existence of the cited work, measured by an automated matching pipeline; it says nothing about whether an existing cited work supports the claim, and it does not identify which tool, if any, produced a reference. 'Hallucinated' in the paper names the estimated excess, not a classification of individual references.

Collection

Academic authors at Cornell, UCLA, Tsinghua and UC Berkeley; arXiv preprint, not peer reviewed. Method: references parsed from LaTeX, GROBID, platform XML or Crossref metadata; titles matched by string similarity against a local Elasticsearch index of Semantic Scholar and OpenAlex (95.1% matched), then a GPT-4o-mini step excludes non-academic strings (unmatched 2.33%), re-extraction (1.54%), then a Google Scholar lookup. The estimate is a regression excess over the pre-2023 unmatched rate, so attribution to LLM use is by timing and correlates (fields with high AI uptake, linguistic signatures of AI-assisted writing), not by observation of tool use. The annual figure assumes August 2025 monthly levels persist through December. The authors call it a lower bound because only title existence is tested. Counter-check that exists: manual validation of unmatched cases and sensitivity tests, reported in the Supplementary Information, which is not part of the observed text.

Falls when

Manual verification of a random sample of post-2023 unmatched references finds that most exist (indexing lag, non-English or grey literature), or the pre-LLM baseline re-estimated with the same pipeline on later database snapshots rises to the post-2023 level. Query to run: hand-check of sampled unmatched references per corpus and year against publisher records. The anchor narrows, without falling, if the full-year 2025 data replace the August extrapolation.

Reflex

AI fabricates references, and papers are now full of fake citations. Too coarse: the measured excess is 0.2 to 1.9 percent of references depending on corpus, spread thinly over many papers, and it concerns non-existent works in published manuscripts, a different quantity from whether an assistant's citation supports its claim.

Evidence

https://arxiv.org/pdf/2605.07723v1 p. 3-5, Estimating hallucinated references at scale and Results | 2026-05-08 · arXiv 2605.07723 · Zhao et al., LLM hallucinations in the wild

Findings and answers · 0

No attacker has recorded a finding on this card yet.