Questions › Do the citations of AI research assistants support the claims they are attached to
Of 400 references from eight free chatbots about 27 percent were fully correct and 40 percent erroneous or fabricated
From the total dataset of 400 references analyzed, 26.5% were real and fully accurate (i.e., all five bibliographic elements were correct), while 33.8% were real but only partially correct (e.g., containing errors in the publication year or locating data). In contrast, 39.8% of the references were either incorrect or entirely fabricated by the AI systems.
https://arxiv.org/pdf/2505.18059v1 p. 8, Results and Figure 1
Falls whenA second coder re-searching the 400 references in Crossref, WorldCat and Google Scholar moves more than a few points between the three classes, in particular if references classed as fabricated turn out to exist under another edition or title variant, or a rerun of the five prompts printed in Table 2 on the same free versions gives per-chatbot fabrication rates far from 0 to 100 percent as printed. Query to run: re-verify the reference list behind Figure 1 of arXiv 2505.18059.
Statement
Eight chatbots in their free versions (ChatGPT on GPT-4o-mini, Claude 3.5 Sonnet, Copilot, DeepSeek-V3, Gemini Flash 2.0, Grok-3, Le Chat, Perplexity Sonar) were each asked between 7 and 9 February 2025, with one standardized student prompt, for 10 academic references in APA format in each of five disciplines, giving 400 references; checked by hand on five bibliographic elements (authors, year, title, venue, locating data), 26.5% were real and fully accurate, 33.8% real but partially correct, and 39.8% incorrect or entirely fabricated, with Grok and DeepSeek fabricating none of their 50 references and Copilot, Perplexity and Claude fabricating 100%, 72% and 64%. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.
Collection
Two academic authors at a Spanish university; arXiv preprint that states on its first page it has not undergone peer review. No stake in any evaluated chatbot is declared or visible. Sample: 50 references per chatbot, 80 per knowledge area, one prompt per discipline, one run. Rater: the authors, by manual searches on Google and Google Scholar with the chatbot's title in quotation marks; each reference scored 0 to 5 errors. The paper reports no second coder and no inter-rater agreement. The counter-check that exists is re-running the five printed prompts and re-searching the titles; the paper does not report having done a second pass.
Falls when
A second coder re-searching the 400 references in Crossref, WorldCat and Google Scholar moves more than a few points between the three classes, in particular if references classed as fabricated turn out to exist under another edition or title variant, or a rerun of the five prompts printed in Table 2 on the same free versions gives per-chatbot fabrication rates far from 0 to 100 percent as printed. Query to run: re-verify the reference list behind Figure 1 of arXiv 2505.18059.
Reflex
Newer chatbots no longer invent references the way early ChatGPT did. Too coarse: in February 2025 six of eight free chatbots still fabricated references, including Perplexity at 72 percent and Copilot at 100 percent, and this counts existence and bibliographic correctness only, not whether a source supports a claim.
Evidence
https://arxiv.org/pdf/2505.18059v1 p. 8, Results and Figure 1 | 2025-05-23 · arXiv 2505.18059 · Cabezas-Clavijo and Sidorenko-Bautista, Assessing the performance of 8 AI chatbots in bibliographic reference retrieval
Findings and answers · 0
No attacker has recorded a finding on this card yet.