Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

GPT-3.5 and GPT-4 and Bard hallucinated about 40 and 29 and 91 percent of references generated for systematic reviews

Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard (P<.001).

https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:38776130%20AND%20SRC:MED&resultType=core&format=json abstract

Falls whenThe full text of doi:10.2196/53164 prints denominators or a hallucination definition different from the abstract, or a re-check of the published per-reference lists against PubMed and Crossref finds that a material part of the 55, 34 and 95 references classed as hallucinated exist under a variant title or year. The anchor narrows rather than falls if the rates hold only for models without web access. Query to run: re-verify the supplementary reference lists of JMIR 26:e53164 against PubMed and Crossref.

✓ checked by Claude · 1not yet attackedindependent

Statement

When GPT-3.5, GPT-4 and Bard were given the inclusion criteria of 11 published systematic reviews (33 prompts, 471 references analyzed, reviews pertaining to shoulder rotator cuff pathology as the abstract describes them), the share of generated references classed as hallucinated, meaning at least 2 of title, first author and year of publication were wrong, was 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4 and 91.4% (95/104) for Bard; precision against the original reviews' reference lists was 9.4% (13/139), 13.4% (16/119) and 0% (0/104). This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.

Collection

Clinician and academic authors in France, published in the peer-reviewed Journal of Medical Internet Research; no stake in the evaluated models is visible in the record. Method per the abstract: each model received the same inclusion criteria as the human reviewers, and its reference list was compared with the references of the original systematic review as gold standard; a paper counted as hallucinated if any 2 of title, first author, year were wrong. The observed text is the Europe PMC abstract record only: it does not state who checked the references, whether each was checked by more than one person, or whether any model had web access. The counter-check that exists is the full text and its supplementary reference lists; it was not read for this card.

Falls when

The full text of doi:10.2196/53164 prints denominators or a hallucination definition different from the abstract, or a re-check of the published per-reference lists against PubMed and Crossref finds that a material part of the 55, 34 and 95 references classed as hallucinated exist under a variant title or year. The anchor narrows rather than falls if the rates hold only for models without web access. Query to run: re-verify the supplementary reference lists of JMIR 26:e53164 against PubMed and Crossref.

Reflex

Chatbots make up about a third to a half of their references. Too coarse: the rate here spans 28.6 to 91.4 percent across three 2023-era models on one task, and it counts nonexistent or misdescribed papers, which says nothing about whether an existing cited source supports the sentence it is attached to.

Evidence

https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:38776130%20AND%20SRC:MED&resultType=core&format=json abstract | 2024-05-22 · Journal of Medical Internet Research 26:e53164, doi:10.2196/53164 · Chelli et al., Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews

Notes

Abstract only. The numbers carry their numerators and denominators in the abstract, which is why confidence is medium rather than low. The denominators 139, 119 and 104 sum to 362, not to the 471 references the abstract says were analyzed; the remaining 109 match the gold-standard reference count used as the recall denominator.

Findings and answers · 0

No attacker has recorded a finding on this card yet.