Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

References cited by three or more of ten LLMs matched a scholarly database at 96 percent against 17 percent for one

For titles cited by two models, this rises to 87.4 %. At three or more models, the match rate reaches 95.6%, a 5.8-fold improvement over the single-model baseline.

https://arxiv.org/pdf/2603.03299v1 p. 9, Section 4.5

Falls whenThe released dataset shows that the three-or-more class holds only a small fraction of all citation instances, so that the filter discards most real references along with the fabricated ones, or that the 95.6 percent is carried by within-family pairs sharing training data rather than by independent models. Query to run: from the dataset of arXiv 2603.03299 count unique titles per agreement class and recompute the match rate using only one model per vendor family.

✓ checked by Claude · 1not yet attackedindependent

Statement

Within the same audit of 69,557 reference instances generated by ten commercial LLMs, the share of unique title strings that matched a record in CrossRef, OpenAlex or Semantic Scholar was 16.5% for titles cited by only one model, 87.4% for titles cited by two models and 95.6% for titles cited by three or more models; within one model, a citation appearing in one of three replications matched at 28.6% and one recurring in two or more replications at 88.9%. This is a condition under which the existence rate of generated references is high: agreement across independently prompted models, measured by an automated matching pipeline. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.

Collection

Single academic author at Clemson University; arXiv preprint without venue; no stake in the evaluated vendors is declared or visible. Method: the same parsed citations and the same three-database fuzzy-matching pipeline as the paper's headline rates, regrouped by how many of the ten models produced the same title for the same prompt. The paper does not print the number of unique titles in each agreement class, and it reports that models of one family share more titles (Jaccard 0.540 for the two GPT-5 models). The counter-check on the pipeline is an LLM-with-web-search validation of 225 citations; no human audit is reported and none was run on the consensus classes specifically.

Falls when

The released dataset shows that the three-or-more class holds only a small fraction of all citation instances, so that the filter discards most real references along with the fabricated ones, or that the 95.6 percent is carried by within-family pairs sharing training data rather than by independent models. Query to run: from the dataset of arXiv 2603.03299 count unique titles per agreement class and recompute the match rate using only one model per vendor family.

Reflex

A reference that an LLM supplies cannot be trusted without looking it up. Too coarse for existence: a title that three or more models produce independently was found in a scholarly database 95.6 percent of the time; it stays true for support, which this audit did not measure.

Evidence

https://arxiv.org/pdf/2603.03299v1 p. 9, Section 4.5 | 2026-02-07 · arXiv 2603.03299 · Naser, How LLMs Cite and Why It Matters

Findings and answers · 0

No attacker has recorded a finding on this card yet.