Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Of 1053 AI assistant news responses with direct quotes 12 percent had significant quote accuracy issues per journalists

Across all AI assistant responses that included a direct quote (a total of 1,053), 12% were found to have significant issues with the accuracy of those direct quotes.

https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 21, Accuracy of direct quotes

Falls whenFalls if the quotes in the 1,053 responses are compared string by string with the pages cited for them (in the version live at generation time) and the share of responses with an altered or absent quote differs materially from 12 percent. Narrows if a large part of the significant ratings concern translated quotes, where the assistant rendered a source-language quote into the response language and the evaluator counted the rendering as alteration: query the significant Q3 ratings by language and by whether source and response language differ.

✓ checked by Claude · 1not yet attackedpositioned

Statement

In the EBU/BBC study (22 public service media organizations, 18 countries, responses generated 24 May to 10 June 2025 by the free consumer versions of ChatGPT, Copilot, Gemini and Perplexity), journalists answered for each response the question 'Do any direct quotes in the response accurately reflect the source cited for them?' (Q3). Of the 1,053 responses to core questions that included a direct quote, 12% were rated as having significant issues with the accuracy of those quotes. Gemini: 20% of 290 responses with quotes; Copilot: 4% (n=190), the fewest; ChatGPT n=262 and Perplexity n=311, whose percentages appear only in a chart (the appendix tables print 28 and 33 significant ratings for them, 59 for Gemini and 8 for Copilot). Reported cases include quotes not found in the source provided for them and quotes with altered wording. The unit is the response containing at least one direct quote, not the individual quote; rated by human journalists.

Collection

Produced by the EBU Media Intelligence Service and the BBC. Each of the 22 participating public service media organizations generated responses with the prompt prefix 'Use [organization] sources where possible', lifted its technical crawler blocks for the generation period, and had its own journalists rate the anonymized responses after a briefing with written and video calibration material. The EBU answers to its member broadcasters. The participating organizations are publishers whose content the assistants use, and the report calls for publisher control over content use, agreed citation formats and regulatory attention: a stake and a declared position, recorded as positioned. Counter-check that exists and was used: project teams in each organization checked all significant-issue ratings for evidence and classification, and the central team ran an additional QA pass. Quote accuracy is checkable by anyone against the cited page, but the report prints only illustrative cases, not the rated responses; no inter-rater agreement statistic is reported.

Falls when

Falls if the quotes in the 1,053 responses are compared string by string with the pages cited for them (in the version live at generation time) and the share of responses with an altered or absent quote differs materially from 12 percent. Narrows if a large part of the significant ratings concern translated quotes, where the assistant rendered a source-language quote into the response language and the evaluator counted the rendering as alteration: query the significant Q3 ratings by language and by whether source and response language differ.

Reflex

Text inside quotation marks with a source link is a verbatim quote from that source. Too coarse: in 12 percent of quote-bearing responses journalists found significant problems, including quotes that are not in the cited source at all.

Evidence

https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 21, Accuracy of direct quotes | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q3 counts) | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants

Findings and answers · 0

No attacker has recorded a finding on this card yet.