Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Invalid citations appear in about 1 percent of 56381 AI and security conference papers and rose 81 percent in 2025

identifying 739 invalid citations across 604 papers (1.07% of 56,381) are definitively invalid, with an 80.9% increase in invalid citation rates in 2025 (from a 2020–2024 average of 0.89% to 1.61%)

https://arxiv.org/pdf/2602.06718v2 p. 6, Section IV.B; p. 9-10, Section VI, Tables IV and V, Figure 5

Falls whenA re-check of the 603 untraceable citations finds a material part in sources the reviewers did not search (theses, non-English venues, withdrawn preprints), or the 2025 share returns to the 0.76 to 0.98 percent band once the 2025 proceedings are complete and re-extracted. Query to run: re-verify the released list of 739 invalid citations from arXiv 2602.06718 and recompute the per-year paper share.

✓ checked by Claude · 2not yet attackedindependent

Statement

In 2,199,409 citations extracted from 56,381 papers accepted at eight AI/ML and security venues (NeurIPS, ICML, AAAI, IJCAI, USENIX Security, CCS, S&P, NDSS) from 2020 to 2025, an automated check flagged 2,530 citations, manual review confirmed 739 as invalid (136 with wrong metadata, 603 untraceable), and 604 papers (1.07%) contained at least one invalid citation; the yearly share of such papers stayed between 0.76% and 0.98% from 2020 to 2024 and was 1.61% in 2025, 80.9% above the 2020-2024 average of 0.89%. This measures the published record written by researchers, whose use of an AI assistant for any given paper is not established, and the authors state that their data alone does not establish causality. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.

Collection

Academic authors at Nankai University and Tsinghua University; arXiv preprint (cs.CR), version 2; the authors publish at some of the audited venues and built the screening tool, and no other stake is visible. Method: CiteVerifier flagged citations with title similarity below 0.9; sixteen trained research assistants reviewed every flagged citation, each checked independently at least twice; two researchers then reviewed all citations classed invalid. The counter-check that exists was used: 400 citations sampled from the valid pool were reviewed by hand and no further invalid citation was found. The share is a lower bound of what the screening could see, since only flagged citations were reviewed.

Falls when

A re-check of the 603 untraceable citations finds a material part in sources the reviewers did not search (theses, non-English venues, withdrawn preprints), or the 2025 share returns to the 0.76 to 0.98 percent band once the 2025 proceedings are complete and re-extracted. Query to run: re-verify the released list of 739 invalid citations from arXiv 2602.06718 and recompute the per-year paper share.

Reflex

AI-written papers are flooding conferences with fake references. Too coarse: about one paper in a hundred at eight top venues carries at least one invalid citation, the pre-LLM years 2020 to 2022 already sit near 0.9 percent, and the 2025 rise to 1.61 percent is not attributed to AI use by the measurement itself.

Evidence

https://arxiv.org/pdf/2602.06718v2 p. 6, Section IV.B; p. 9-10, Section VI, Tables IV and V, Figure 5 | 2026-05-14 (v2; v1 2026-02-06) · arXiv 2602.06718 · Xu et al., GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models

Findings and answers · 0

No attacker has recorded a finding on this card yet.