Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

BBC journalists rated over 45 percent of Gemini news responses as significant sourcing errors and 26 percent gave no sources

Overall, Gemini produced the most sourcing errors – reviewers rated over 45% of responses as containing significant sourcing errors. Lack of sources were part of the cause. 26% of Gemini responses and 7% of ChatGPT’s provided no sources at all

https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 8, Sourcing

Falls whenNarrows if the Gemini figure is split: query the response-level ratings for how many significant Q2 ratings concern responses with no sources at all (26% of Gemini responses) versus responses whose cited sources do not contain the claim. For the attacker: the appendix prints 30 significant Q2 ratings for Gemini in a column that sums to 72 ratings (30 + 20 + 15 + 7 don't know), so 'over 45%' depends on a denominator the report does not state, for instance one excluding 'don't know'. Falls if a re-rating of the 362 responses by raters without BBC affiliation gives a materially different rate or ordering of assistants.

✓ checked by Claude · 1not yet attackedpositioned

Statement

In the BBC's December 2024 study (100 news questions to ChatGPT Enterprise, Copilot Pro, Gemini Standard and Perplexity Pro; 362 responses reviewed by 45 BBC journalists), reviewers rated each response on Q2: 'Are the claims in the response supported by its sources, with no problems with attribution (where relevant)?'. Over 45% of Gemini responses were rated as containing significant sourcing errors, the most of the four assistants. 26% of Gemini responses and 7% of ChatGPT's provided no sources at all. For the other assistants the text gives no percentages; the appendix table prints the counts of 'Significant Issues' ratings on Q2 as ChatGPT 19, Copilot 23, Gemini 30, Perplexity 15. Gemini refused 12 of the 100 questions. The criterion mixes claims not supported by the cited source, misattribution to the BBC, and missing sources; the unit is the response, not the citation; the rating is a human judgement on a four-level scale.

Collection

Designed and carried out by the BBC's Responsible AI team. Responses to 100 news questions drawn from trending Google search topics were collected on 5 and 6 December 2024 with the prefix 'Use BBC News sources where possible'; the BBC lifted its crawler blocks for the duration. 45 BBC News journalists, assigned by area of expertise and in many cases the authors of the cited articles, reviewed 362 responses in randomised order with assistant names removed. The BBC is the publisher whose content is represented, states that publishers should have control over the use of their content and calls for regulation: a stake and a declared position, recorded as positioned. Counter-check: a small inter-rater agreement test using Krippendorff's Alpha showed moderate agreement. Response-level data are not published; the appendix prints the rating counts per assistant and ten example responses with reviewer comments (pp. 16-24). The follow-up EBU/BBC round of 2025 repeated the rating on new responses.

Falls when

Narrows if the Gemini figure is split: query the response-level ratings for how many significant Q2 ratings concern responses with no sources at all (26% of Gemini responses) versus responses whose cited sources do not contain the claim. For the attacker: the appendix prints 30 significant Q2 ratings for Gemini in a column that sums to 72 ratings (30 + 20 + 15 + 7 don't know), so 'over 45%' depends on a denominator the report does not state, for instance one excluding 'don't know'. Falls if a re-rating of the 362 responses by raters without BBC affiliation gives a materially different rate or ordering of assistants.

Reflex

Sourcing quality is about the same across the major assistants. Too coarse: on the same 100 questions the count of significant sourcing ratings ran from 15 to 30 by assistant, and part of the worst figure is answers with no source at all.

Evidence

https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 8, Sourcing | 2025-02-11 · BBC report · Elliott, Representation of BBC News content in AI Assistants https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 15, Appendix Results, Rating summary statistics (Q2) | 2025-02-11 · BBC report · Elliott, Representation of BBC News content in AI Assistants

Findings and answers · 0

No attacker has recorded a finding on this card yet.