Questions › Do the citations of AI research assistants support the claims they are attached to
Journalists at 22 public media rated 31 percent of AI assistant news responses as having significant sourcing issues
Sourcing was the biggest cause of problems, with 31% of all responses having significant issues with sourcing – this includes information in the response not supported by the cited source, providing no sources at all, or making incorrect or unverifiable sourcing claims.
https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 10, High-level findings
Falls whenThe anchor narrows if the 31% is decomposed: query the study's response-level data for the share of significant Q2 ratings that rest on 'no direct source' or on unverifiable sourcing claims rather than on a cited source that does not contain the claim; if the unsupported-claim share alone is far lower, above all outside Gemini, the figure cannot be read as a citation-support rate. It falls if a re-rating of the same responses by raters without a publisher affiliation yields a materially different rate or a different ordering of assistants. It narrows in time if a repeat with paid tiers or later default models (the report lists GPT-5 as ChatGPT's default by 16 Oct 2025) no longer shows the 15 to 72 percent spread.
Statement
In the EBU/BBC study, 271 journalists from 22 public service media organizations in 18 countries rated 2,709 responses to 30 core news questions, generated between 24 May and 10 June 2025 in 14 languages by the free consumer versions of ChatGPT, Copilot, Gemini and Perplexity. On the sourcing criterion (Q2: 'Are the claims in the response supported by the sources the assistant provides?') 31% of responses were rated as having significant issues: Gemini 72%, ChatGPT 24%, Perplexity 15%, Copilot 15% (n = 675, 678, 681, 675; the appendix tables print 483, 160, 101 and 104 significant ratings). The category is wider than unsupported claims: it covers information in the response not supported by the cited source, providing no sources at all, and incorrect or unverifiable sourcing claims. The appendix states that significant issues for Q2 include the lack of any direct sourcing, and 42% of Gemini responses provided no direct sources. The unit is the response, not the individual citation; the rating is a human judgement on a four-level scale (no issues, some issues, significant issues, don't know).
Collection
Produced by the EBU Media Intelligence Service and the BBC. Each of the 22 participating public service media organizations generated responses with the prompt prefix 'Use [organization] sources where possible', lifted its technical crawler blocks for the generation period, and had its own journalists rate the anonymized responses after a briefing with written and video calibration material. The EBU answers to its member broadcasters. The participating organizations are publishers whose content the assistants use, and the report calls for publisher control over content use, agreed citation formats and regulatory attention: a stake and a declared position, recorded as positioned. Counter-check that exists and was used: project teams in each organization checked all significant-issue ratings for whether they were clearly evidenced and correctly classified and checked that sourcing issues were logged correctly, and the central team ran an additional QA pass; participant organizations remain responsible for their own data. No inter-rater agreement statistic is reported for this round, and the report does not publish the response-level data.
Falls when
The anchor narrows if the 31% is decomposed: query the study's response-level data for the share of significant Q2 ratings that rest on 'no direct source' or on unverifiable sourcing claims rather than on a cited source that does not contain the claim; if the unsupported-claim share alone is far lower, above all outside Gemini, the figure cannot be read as a citation-support rate. It falls if a re-rating of the same responses by raters without a publisher affiliation yields a materially different rate or a different ordering of assistants. It narrows in time if a repeat with paid tiers or later default models (the report lists GPT-5 as ChatGPT's default by 16 Oct 2025) no longer shows the 15 to 72 percent spread.
Reflex
Assistants with web search cite their sources, so their news answers can be checked. Too coarse: journalists rated 31 percent of responses as having significant sourcing issues, from 15 to 72 percent depending on the assistant, and the category includes answers with no source at all.
Evidence
https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 10, High-level findings | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q2 note and counts) | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants
Findings and answers · 0
No attacker has recorded a finding on this card yet.