Questions › Do the citations of AI research assistants support the claims they are attached to
Audit of 111 million references in arXiv, bioRxiv, SSRN and PMC papers estimates 146,932 non-existent citations in 2025
Assuming these trends persisted through year end, these four corpora, which cover only a fraction of the scientific literature, would include 146,932 hallucinated citations in 2025 alone.
https://arxiv.org/pdf/2605.07723v1 p. 3-5, Estimating hallucinated references at scale and Results
Falls whenManual verification of a random sample of post-2023 unmatched references finds that most exist (indexing lag, non-English or grey literature), or the pre-LLM baseline re-estimated with the same pipeline on later database snapshots rises to the post-2023 level. Query to run: hand-check of sampled unmatched references per corpus and year against publisher records. The anchor narrows, without falling, if the full-year 2025 data replace the August extrapolation.
Statement
Downstream context, not a measurement of any assistant: an audit of 111 million references in 2.5 million papers (arXiv Jan 2020 to Aug 2025, bioRxiv, SSRN, and a 10% sample of PubMed Central) checks whether each cited title exists in Semantic Scholar, OpenAlex or Google Scholar. The excess of unmatched references over the pre-LLM baseline reached 0.39% (arXiv), 0.21% (bioRxiv), 1.91% (SSRN) and 0.27% (PMC) of references as of August 2025, with monthly excess counts of 3,353, 478, 767 and 8,140; extrapolating these to year end gives 146,932 non-existent citations in the four corpora in 2025. The quantity is existence of the cited work, measured by an automated matching pipeline; it says nothing about whether an existing cited work supports the claim, and it does not identify which tool, if any, produced a reference. 'Hallucinated' in the paper names the estimated excess, not a classification of individual references.
Collection
Academic authors at Cornell, UCLA, Tsinghua and UC Berkeley; arXiv preprint, not peer reviewed. Method: references parsed from LaTeX, GROBID, platform XML or Crossref metadata; titles matched by string similarity against a local Elasticsearch index of Semantic Scholar and OpenAlex (95.1% matched), then a GPT-4o-mini step excludes non-academic strings (unmatched 2.33%), re-extraction (1.54%), then a Google Scholar lookup. The estimate is a regression excess over the pre-2023 unmatched rate, so attribution to LLM use is by timing and correlates (fields with high AI uptake, linguistic signatures of AI-assisted writing), not by observation of tool use. The annual figure assumes August 2025 monthly levels persist through December. The authors call it a lower bound because only title existence is tested. Counter-check that exists: manual validation of unmatched cases and sensitivity tests, reported in the Supplementary Information, which is not part of the observed text.
Falls when
Manual verification of a random sample of post-2023 unmatched references finds that most exist (indexing lag, non-English or grey literature), or the pre-LLM baseline re-estimated with the same pipeline on later database snapshots rises to the post-2023 level. Query to run: hand-check of sampled unmatched references per corpus and year against publisher records. The anchor narrows, without falling, if the full-year 2025 data replace the August extrapolation.
Reflex
AI fabricates references, and papers are now full of fake citations. Too coarse: the measured excess is 0.2 to 1.9 percent of references depending on corpus, spread thinly over many papers, and it concerns non-existent works in published manuscripts, a different quantity from whether an assistant's citation supports its claim.
Evidence
https://arxiv.org/pdf/2605.07723v1 p. 3-5, Estimating hallucinated references at scale and Results | 2026-05-08 · arXiv 2605.07723 · Zhao et al., LLM hallucinations in the wild
Findings and answers · 0
No attacker has recorded a finding on this card yet.