Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

For seven LLMs on 800 medical questions 50 to 90 percent of responses are not fully supported by the sources they cite

between 50% and 90% of LLM responses are not fully supported

https://www.nature.com/articles/s41467-025-58551-6.pdf p. 1, Abstract; p. 2, section Evaluation of source veracity in LLMs; p. 3, Fig. 1b

Falls whenSupplementary Table 4 or a recomputation from the released data shows response-level support above 50% for more than one system or a worst case well above 10%, moving the 50 to 90 range; or a rerun on assistants released after 2024 over the same 800 questions finds most responses fully supported. The anchor narrows if responses with no sources at all (over 20% for GPT-4o with web search) are excluded and the not-fully-supported share drops materially. Query to run: response-level support per model from the released SourceCheckup data, with and without source-less responses.

✓ checked by Claude · 1not yet attackedindependent

Statement

Seven LLMs (GPT-4o with web search, Gemini Ultra 1.0 with web search, and the API endpoints of GPT-4o, Claude v2.1, Mistral Medium, Mixtral and Gemini Pro) each answered 800 medical questions (400 generated from Mayo Clinic pages, 400 taken from Reddit r/AskDocs) and were asked to list supporting sources, giving about 58,000 statement-source pairs. The paper summarises that between 50% and 90% of responses are not fully supported by the sources they cite. Printed in the running text: response-level support of 55% for GPT-4o with web search (the best system), 34.5% for Gemini Ultra 1.0 with web search, and about 10% for the Gemini Pro API; for GPT-4o with web search approximately 30% of individual statements are unsupported. Link validity is reported separately: models without web access produced valid URLs between 40% and 70% of the time (GPT-4o API around 70%), while the two web-search systems are described as not suffering from URL hallucination. Per-model figures for all seven are in Fig. 1b and Supplementary Table 4, which is not part of the observed text. Support was judged by an automated GPT-4o judge that agreed with a three-doctor consensus in 88.7% of 400 pairs. The Methods print that Gemini Ultra 1.0 with web search was evaluated on 3/28/24 and all other model APIs were queried on 1/20/24; the GPT-4o endpoint is named gpt-4o-2024-05-13, and no query date for GPT-4o is printed.

Collection

Authors are at Stanford University (Biomedical Data Science, Electrical Engineering, Computer Science, Genetics, Anesthesiology, Law School), Keck Medicine of USC and Loma Linda University School of Medicine; the paper is peer reviewed (Nature Communications), the authors declare no competing interests and are not the vendor of any evaluated model. Pipeline SourceCheckup: questions are generated by GPT-4o from Mayo Clinic pages or taken from Reddit r/AskDocs, each evaluated LLM answers and lists sources, GPT-4o parses the response into statements, each URL is downloaded (valid = status code 200 with non-empty text), and GPT-4o as Source Verification model judges every statement-source pair. A statement counts as supported if any source in the response supports it; a response counts as supported only if all its statements are. GPT-4o with web search returned no sources at all in over 20% of responses, which the authors name as a partial cause of its low response-level support. Counter-checks used: doctor validation of the judge on 400 pairs, doctor review of 110 pairs judged unsupported (95.8% agreement), a second judge model, and an any-source-merged rerun in which 95.1% of unsupported statements stayed unsupported. The range of 50 to 90 percent appears only in the abstract; the exact per-model values sit in the supplement. The annotating doctors are co-authors; investigators were not blinded.

Falls when

Supplementary Table 4 or a recomputation from the released data shows response-level support above 50% for more than one system or a worst case well above 10%, moving the 50 to 90 range; or a rerun on assistants released after 2024 over the same 800 questions finds most responses fully supported. The anchor narrows if responses with no sources at all (over 20% for GPT-4o with web search) are excluded and the not-fully-supported share drops materially. Query to run: response-level support per model from the released SourceCheckup data, with and without source-less responses.

Reflex

Language models with citations back up most of what they say; the known problem is invented references. Too coarse: invalid URLs were confined to the models without web access, while incomplete support affected close to half or more of the responses of every system, including those whose links all resolved.

Evidence

https://www.nature.com/articles/s41467-025-58551-6.pdf p. 1, Abstract; p. 2, section Evaluation of source veracity in LLMs; p. 3, Fig. 1b | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references

Findings and answers · 0

No attacker has recorded a finding on this card yet.