Questions › Do the citations of AI research assistants support the claims they are attached to
Over 3150 generated references ChatGPT had the lowest and Perplexity the highest mean Reference Hallucination Score
ChatGPT had the lowest mean RHS (1.81 ± 3.40) and thus emerged as the most reliable model. Gemini scored 4.01 ± 4.89, while Perplexity had the highest score at 6.51 ± 4.89
https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1007/s43465-026-01807-0&resultType=core&format=json abstract
Falls whenThe full text or supplement of doi:10.1007/s43465-026-01807-0 prints per-format and pooled means that cannot be reconciled with the abstract, or shows that the RHS is dominated by the topical relevance or PMID criterion rather than by existence, so that the ordering ChatGPT, Gemini, Perplexity does not hold for existence alone. Query to run: recompute the pooled mean RHS per chatbot from the supplementary per-reference scores and split it by criterion.
Statement
In a cross-sectional study, 30 rotator cuff subtopics were posed to ChatGPT, Gemini and Perplexity in two formats (letter to the editor and original article), yielding 3150 references scored with the Reference Hallucination Score (RHS), a composite of existence or verifiability, bibliographic accuracy, PMID validity and topical relevance in which a higher score means a less reliable reference; with all formats together the mean RHS was 1.81 +/- 3.40 for ChatGPT, 4.01 +/- 4.89 for Gemini and 6.51 +/- 4.89 for Perplexity (p < 0.001). These are mean scores with standard deviations on a scale whose range the abstract does not print, not shares of references. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim.
Collection
Two clinician authors in Turkey, published in the peer-reviewed Indian Journal of Orthopaedics; no stake in the evaluated chatbots is visible in the record. Method per the abstract: each reference generated by the three chatbots was scored on the four RHS criteria. The observed text is the Europe PMC abstract record only: it does not print the scale of the RHS, the model versions, the test dates, whether web search was active, or who scored the references. The counter-check that exists is the full text and the supplementary material named in the abstract; neither was read for this card.
Falls when
The full text or supplement of doi:10.1007/s43465-026-01807-0 prints per-format and pooled means that cannot be reconciled with the abstract, or shows that the RHS is dominated by the topical relevance or PMID criterion rather than by existence, so that the ordering ChatGPT, Gemini, Perplexity does not hold for existence alone. Query to run: recompute the pooled mean RHS per chatbot from the supplementary per-reference scores and split it by criterion.
Reflex
An assistant that shows its sources, such as Perplexity, gives more reliable references than a plain chatbot. Too coarse: on this score Perplexity had the highest mean hallucination score of the three, and the score mixes existence, metadata, PMID validity and relevance without touching whether a source supports a claim.
Evidence
https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1007/s43465-026-01807-0&resultType=core&format=json abstract | 2026-05-11 · Indian Journal of Orthopaedics (2026), doi:10.1007/s43465-026-01807-0 · Ozbek and Bagcier, Reference Hallucination in AI-Assisted Academic Writing: A Comparative Analysis of ChatGPT, Gemini, and Perplexity in Rotator Cuff Literature
Notes
Abstract only. The abstract is hard to reconcile internally: it prints per-format means of 1.81 (letter) and 4.02 (article) for ChatGPT but a pooled mean of 1.81, and per-format means of 6.43 and 6.31 for Perplexity but a pooled mean of 6.51, which lies outside the two. A pooled mean of two groups must lie between the group means, so at least one printed figure is off or the pooling is not what the abstract suggests. Confidence is therefore low; the ordering of the three systems is the same in every printed row. Narrowed 2026-09-23 after attacker run 2 (Grok 4.7) withheld the point for lack of the full text: the abstract itself prints per-format means that cannot be pooled into its pooled means. Letter format: ChatGPT 1.81 +/- 3.20, Gemini 3.81 +/- 4.83, Perplexity 6.43 +/- 4.78; article format: ChatGPT 4.02 +/- 1.70, Gemini 4.13 +/- 1.57, Perplexity 6.31 +/- 1.68; pooled as printed: 1.81, 4.01, 6.51. Perplexity's pooled mean exceeds both of its format means and ChatGPT's pooled mean equals its letter-format mean, which no weighted mean of the two formats gives. What stands: the ordering ChatGPT below Gemini below Perplexity holds in each format separately. What is narrowed: the pooled figures in the quoted sentence are not to be read as reliable magnitudes until the full text or supplement resolves the print.