Questions › Do the citations of AI research assistants support the claims they are attached to
Ten commercial LLMs produced references with no database match at rates from 11 to 57 percent across 69557 citations
Across the ten models, hallucination rates range from 11.4% (GPT-5-mini) to 56.8% (haiku-4.5), a fivefold variation.
https://arxiv.org/pdf/2603.03299v1 p. 6-7, Sections 3.4, 3.5 and 4.1, Table 1; p. 10-11, Section 4.6; p. 19, limitations
Falls whenA human audit of a random sample of the 29,028 citations labelled hallucinated finds a share of real publications well above the 10.7 percent the paper's own LLM validation found, for instance books, reports and standards that CrossRef, OpenAlex and Semantic Scholar do not index, which would pull the 11.4 to 56.8 percent range towards the 9.3 to 23.8 percent the paper prints for the inclusive threshold. Query to run: hand-verify 300 randomly drawn unmatched citations from the released dataset of arXiv 2603.03299 in Google Scholar and WorldCat.
Statement
Ten commercially deployed LLMs queried through their APIs with prompts requesting scholarly references in four academic domains (structural engineering, climate and environmental science, biomedical research, NLP and AI), two temporal framings and three replications produced 69,557 parsed citation instances, of which 40,529 were matched at confidence score 80 or higher in CrossRef, OpenAlex or Semantic Scholar; the per-model share without such a match, which the paper calls the hallucination rate, ranges from 11.4% (GPT-5-mini, 95% CI 10.4 to 12.5) to 56.8% (haiku-4.5, 95% CI 55.6 to 58.0), and under the inclusive threshold of score 65 the same rates fall to between 9.3% and 23.8%. The rater is an automated fuzzy-matching pipeline, not a human, and the ten models ran without retrieval. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.
Collection
Single academic author at Clemson University; arXiv preprint without venue. No stake in the evaluated vendors is declared or visible. Method: 15,150 API responses, regex-based citation parsing, then title, author and year fuzzy matching against three scholarly databases; everything below score 65 is classed as hallucinated, and the headline rates also count scores 65 to 79 as hallucinated. The counter-check that exists was used in part: GPT-4.1-mini with web search judged a stratified sample of 225 citations and found 75 of 75 confirmed matches real and 8 of 75 citations classed as hallucinated to be real (10.7%). No human audit of the labels is reported. The paper states (p. 19) that it evaluates unaugmented text-generation models producing references from parametric memory alone, and that its figures are a baseline for non-retrieval citation generation.
Falls when
A human audit of a random sample of the 29,028 citations labelled hallucinated finds a share of real publications well above the 10.7 percent the paper's own LLM validation found, for instance books, reports and standards that CrossRef, OpenAlex and Semantic Scholar do not index, which would pull the 11.4 to 56.8 percent range towards the 9.3 to 23.8 percent the paper prints for the inclusive threshold. Query to run: hand-verify 300 randomly drawn unmatched citations from the released dataset of arXiv 2603.03299 in Google Scholar and WorldCat.
Reflex
Current LLMs hallucinate roughly a third of their references. Too coarse: the rate spans 11.4 to 56.8 percent across ten models of the same period, depends on the match threshold, and counts only whether the reference can be found in a database, not whether it supports anything.
Evidence
https://arxiv.org/pdf/2603.03299v1 p. 6-7, Sections 3.4, 3.5 and 4.1, Table 1; p. 10-11, Section 4.6; p. 19, limitations | 2026-02-07 · arXiv 2603.03299 · Naser, How LLMs Cite and Why It Matters
Findings and answers · 0
No attacker has recorded a finding on this card yet.