Questions › Do the citations of AI research assistants support the claims they are attached to
Gemini Ultra reached 77 percent reference correctness against 54 percent for GPT-4 in medical research introductions
Gemini's references showed 77.2 % correctness and 68.0 % accuracy, compared to GPT-4's 54.0 % correctness and 49.2 % accuracy (p < 0.001 for both).
https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:39667055%20AND%20SRC:MED&resultType=core&format=json abstract
Falls whenThe full text of doi:10.1016/j.compbiomed.2024.109545 shows that correctness and accuracy include a judgment of whether the reference supports the sentence it is attached to, which would move this card out of the existence family, or shows a reference count too small to carry a 23.2 point difference. Query to run: read the Methods of the full article for the definitions of correctness and accuracy and the number of references per model.
Statement
In a comparison of OpenAI's GPT-4 and Google's Gemini Ultra writing medical research introductions with references across five medical fields, Gemini's references showed 77.2% correctness and 68.0% accuracy against 54.0% correctness and 49.2% accuracy for GPT-4 (p < 0.001 for both), and the abstract states that both models produced fabricated evidence. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. The abstract does not define how correctness differs from accuracy and gives no count of references.
Collection
Academic medical authors in the United States and Israel, published in the peer-reviewed journal Computers in Biology and Medicine; no stake in either vendor is visible in the record. Method per the abstract: the two models generated introductions in five medical fields, and the credibility and accuracy of the citations were assessed alongside introduction length and unreferenced statements. The observed text is the Europe PMC abstract record only: it gives no number of introductions or references, no definition of the two metrics, and does not say who verified the references or against which database. The counter-check that exists is the full text behind the DOI, which is not open access and was not read for this card.
Falls when
The full text of doi:10.1016/j.compbiomed.2024.109545 shows that correctness and accuracy include a judgment of whether the reference supports the sentence it is attached to, which would move this card out of the existence family, or shows a reference count too small to carry a 23.2 point difference. Query to run: read the Methods of the full article for the definitions of correctness and accuracy and the number of references per model.
Reflex
Gemini is better than GPT-4 at citing medical literature. Too coarse: the difference is 77.2 against 54.0 percent on one task at one point in time, both models produced fabricated references, and the quantity is reference correctness, not support of claims.
Evidence
https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:39667055%20AND%20SRC:MED&resultType=core&format=json abstract | 2024-12-12 · Computers in Biology and Medicine 185:109545 (2025), doi:10.1016/j.compbiomed.2024.109545 · Omar et al., Generating credible referenced medical research: A comparative study of openAI's GPT-4 and Google's gemini
Notes
Abstract only, and the number is thin: two percentages per model without denominators and without metric definitions. Confidence is therefore low. The article is dated 2025 by the journal (volume 185); the record's first publication date online is 2024-12-12, which is used as as_of.
Findings and answers · 0
No attacker has recorded a finding on this card yet.