Questions › Do the citations of AI research assistants support the claims they are attached to
GPT-4o support judge agrees with a three-doctor consensus on 89 percent of 400 pairs and doctors among themselves on 86 percent
88.7% agreement between the Source Veri fication model and the doctor consensus and an 86.1% average inter-doctor agreement rate
https://www.nature.com/articles/s41467-025-58551-6.pdf p. 2, sections Source verification and Evaluation of bias of GPT-4o as the backbone LLM; p. 3, Fig. 1a; p. 8, Expert validation
Falls whenA blinded re-annotation of the released 400 pairs by raters who do not see the judge's decision yields judge-consensus agreement materially below the inter-doctor level of 86.1%, or a chance-corrected statistic on the released annotations shows the agreement is driven by class imbalance. The anchor narrows if the validity holds only for medical web pages: run the same judge prompt against human labels in a non-medical attribution set. Query to run: Cohen's or Fleiss' kappa and the confusion matrix from the released expert annotations of doi 10.1038/s41467-025-58551-6.
Statement
On 400 statement-source pairs drawn from responses of GPT-4o with web search, GPT-4o API and Claude v2.1 API, the automated support judge (GPT-4o prompted as Source Verification model) agreed with the majority consensus of three US-licensed medical doctors in 88.7% of cases; the average agreement between doctors was 86.1%, and the difference between judge and consensus annotations was not statistically significant (p = 0.21, unpaired two-sided t-test). With Claude Sonnet 3.5 as judge, agreement with the consensus was 87.0% (95% CI 83.4-90.4), with Llama 3.1 70B 79.3% (75.4-83.1); GPT-4o and Claude Sonnet 3.5 agreed with each other on 90.1% (89.7-90.5) of decisions. In a further sample of 110 pairs from GPT-4o with web search that the judge had marked unsupported, doctors agreed 95.8% (91.8-98.7) of the time and confirmed 105 of 110. End to end on 100 HealthSearchQA questions, a clinician rated 40.4% (30.7, 50.1) of responses fully supported, the pipeline 42.4% (32.7, 52.2).
Collection
Authors are at Stanford University, Keck Medicine of USC and Loma Linda University School of Medicine; peer reviewed (Nature Communications); no competing interests declared. The annotating doctors are co-authors (author contributions list them under expert annotations) and investigators were not blinded. The Methods say the three doctors independently scored whether the LLM-generated source verification decision correctly identified a statement as supported or not supported, which reads as doctors seeing the judge's decision, while the Fig. 1 caption describes them as determining support themselves; the observed text does not resolve this. Agreement is raw percent agreement; no chance-corrected statistic is printed in the observed text, and the authors note the task is ambiguous given the lack of full agreement among the doctors. The expert annotations are released with the data, so the counter-check (recomputing agreement) exists and is open to anyone.
Falls when
A blinded re-annotation of the released 400 pairs by raters who do not see the judge's decision yields judge-consensus agreement materially below the inter-doctor level of 86.1%, or a chance-corrected statistic on the released annotations shows the agreement is driven by class imbalance. The anchor narrows if the validity holds only for medical web pages: run the same judge prompt against human labels in a non-medical attribution set. Query to run: Cohen's or Fleiss' kappa and the confusion matrix from the released expert annotations of doi 10.1038/s41467-025-58551-6.
Reflex
An LLM grading another LLM's citations is circular and cannot be trusted. Too coarse: on 400 medical pairs the judge agreed with a doctor consensus slightly more often than the doctors agreed with one another, which places the judge's error near the level of human disagreement rather than making it arbitrary.
Evidence
https://www.nature.com/articles/s41467-025-58551-6.pdf p. 2, sections Source verification and Evaluation of bias of GPT-4o as the backbone LLM; p. 3, Fig. 1a; p. 8, Expert validation | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references
Findings and answers · 0
No attacker has recorded a finding on this card yet.