Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

GPT-4o with RAG on 300 health questions has all URLs valid but 76 percent of statements and 38 percent of responses supported

support of 75.7% (74.0– 77.2 95% CI), and response-level support of 38.4% (26.7, 49.3 95% CI).

https://www.nature.com/articles/s41467-025-58551-6.pdf p. 4, sections Additional validation on HealthSearchQA and End-to-end full human evaluation; p. 8, metric definitions

Falls whenA blinded human rating of the released GPT-4o (RAG) responses on the HealthSearchQA subset puts statement-level support outside the printed interval 74.0-77.2 or response-level support outside 26.7-49.3; or a rerun of the released pipeline on current web-search assistants over the same questions finds response-level support within a few points of URL validity. The anchor narrows rather than falls if the gap is specific to open-ended consumer questions: the paper itself prints close to 80% response-level support on Mayo Clinic derived questions. Query to run: per-citation (not any-source) human support rating on the released statement-source pairs of doi 10.1038/s41467-025-58551-6.

✓ checked by Claude · 2not yet attackedindependent

Statement

On a random subset of 300 consumer health questions from HealthSearchQA, GPT-4o with web search (RAG) returned citation URLs that were valid in 100% of cases, while statement-level support (share of parsed medical statements supported by at least one source given in the same response) was 75.7% (95% CI 74.0-77.2) and response-level support (share of responses in which every statement is supported) was 38.4% (95% CI 26.7-49.3). On the main set of 800 questions the same system's response-level support is printed as 55%: close to 80% on the 400 questions generated from Mayo Clinic pages and 31.0% (26.7, 35.8) on the 400 questions from Reddit r/AskDocs. Support was decided per statement-source pair by an automated GPT-4o judge; on 100 HealthSearchQA questions a human clinician rated 40.4% (30.7, 50.1) of the responses fully supported against 42.4% (32.7, 52.2) by the pipeline. The GPT-4o API endpoint named is gpt-4o-2024-05-13; the paper was received 30 September 2024.

Collection

Authors are at Stanford University (Biomedical Data Science, Electrical Engineering, Computer Science, Genetics, Anesthesiology, Law School), Keck Medicine of USC and Loma Linda University School of Medicine; the paper is peer reviewed (Nature Communications), the authors declare no competing interests and are not the vendor of any evaluated model. Pipeline SourceCheckup: questions are generated by GPT-4o from Mayo Clinic pages or taken from Reddit r/AskDocs, each evaluated LLM answers and lists sources, GPT-4o parses the response into statements, each URL is downloaded (valid = status code 200 with non-empty text), and GPT-4o as Source Verification model judges every statement-source pair. Because the models' intended pairing of statement and citation could not be recovered, a statement counts as supported if any source in the response supports it, which is more lenient than checking the citation attached to the claim. Counter-checks that exist and were used: three US-licensed doctors on 400 pairs (88.7% agreement of the judge with their consensus), a second judge model (Claude Sonnet 3.5), and the end-to-end clinician rating of 100 responses. The annotating doctors are co-authors and investigators were not blinded. Data and code are public (github.com/kevinwu23/SourceCheckup).

Falls when

A blinded human rating of the released GPT-4o (RAG) responses on the HealthSearchQA subset puts statement-level support outside the printed interval 74.0-77.2 or response-level support outside 26.7-49.3; or a rerun of the released pipeline on current web-search assistants over the same questions finds response-level support within a few points of URL validity. The anchor narrows rather than falls if the gap is specific to open-ended consumer questions: the paper itself prints close to 80% response-level support on Mayo Clinic derived questions. Query to run: per-citation (not any-source) human support rating on the released statement-source pairs of doi 10.1038/s41467-025-58551-6.

Reflex

An assistant that searches the web and returns working links to reputable health sites has sourced its answer. Too coarse: here every URL resolved, yet about a quarter of the statements and about six in ten whole responses were not backed by the returned sources.

Evidence

https://www.nature.com/articles/s41467-025-58551-6.pdf p. 4, sections Additional validation on HealthSearchQA and End-to-end full human evaluation; p. 8, metric definitions | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references

Findings and answers · 0

No attacker has recorded a finding on this card yet.