Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Between two BBC rounds significant sourcing issues fell to 10 to 15 percent for three assistants while Gemini stayed near 47

Gemini still has the highest percentage of significant issues, broadly the same at 47% (see “Gemini’s issues with sourcing”). By contrast, the other assistants all improved to the 10–15% range, with Copilot showing the steepest drop from 27% to 10%.

https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 19-20, In focus: Have assistants improved?

Falls whenFalls if the two rounds' sourcing criteria are not comparable: query a re-rating of the 362 first-round responses under the second-round rubric (the first-round question wording is printed in the BBC's February 2025 report and differs). Narrows if the drop is carried by the decline in no-source responses (25 to one) rather than by fewer claims unsupported by the cited source. With 237 responses spread over four assistants, a 10-15% rate rests on a handful of responses per assistant, so a third BBC round reversing the direction, or confidence intervals on the per-assistant rates that overlap the first-round values, would make it fall.

✓ checked by Claude · 1not yet attackedpositioned

Statement

The EBU/BBC report compares BBC-only data from two rounds of the same design: 362 responses evaluated in the first round (responses generated December 2024) and 237 core plus custom responses in the second (generated May/June 2025). On significant sourcing issues Gemini stayed 'broadly the same at 47%', while ChatGPT, Copilot and Perplexity all improved to the 10-15% range, Copilot showing the steepest drop from 27% to 10%. Significant issues of any kind fell from 51% to 37%, and BBC responses lacking any direct URL source fell from 25 to a single one. The report qualifies the comparison itself: small differences in methodology and in the definition of key statistics, and different product tiers (first round ChatGPT Enterprise, Copilot Pro, Gemini Standard, Perplexity Pro; second round free consumer versions with default models). Both rounds were rated by BBC journalists at response level.

Collection

Produced by the EBU Media Intelligence Service and the BBC; this comparison uses only the BBC's data because the BBC is the only organization with two rounds. In both rounds the BBC lifted its crawler blocks for the generation period, used the prefix 'Use BBC News sources where possible', and had BBC journalists rate anonymized responses. Custom-question data were added to the second-round core data to increase the sample. The BBC is a publisher whose content the assistants use and the report argues for publisher control and regulatory attention: a stake and a declared position, recorded as positioned; note that an improvement finding runs against that position. Counter-check: the second round had the per-organization and central QA pass on significant ratings; the first round reported a small inter-rater test with moderate agreement; the comparison itself has no independent replication and no response-level data are published.

Falls when

Falls if the two rounds' sourcing criteria are not comparable: query a re-rating of the 362 first-round responses under the second-round rubric (the first-round question wording is printed in the BBC's February 2025 report and differs). Narrows if the drop is carried by the decline in no-source responses (25 to one) rather than by fewer claims unsupported by the cited source. With 237 responses spread over four assistants, a 10-15% rate rests on a handful of responses per assistant, so a third BBC round reversing the direction, or confidence intervals on the per-assistant rates that overlap the first-round values, would make it fall.

Reflex

Unsupported sourcing is a fixed property of language-model assistants. Too coarse: on one publisher's repeated rating three of four assistants improved within six months (one from 27 to 10 percent) while the fourth stayed near 47 percent.

Evidence

https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 19-20, In focus: Have assistants improved? | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants

Findings and answers · 0

No attacker has recorded a finding on this card yet.