Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

ChatGPT Search gave partially or entirely incorrect source attributions for 153 of 200 publisher quotes in November 2024

In total, ChatGPT returned partially or entirely incorrect responses on a hundred and fifty-three occasions, though it only acknowledged an inability to accurately respond to a query seven times.

https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php section 'Confidently wrong'

Falls whenFalls if the released data, relabelled by a second rater, give a materially different count than 153, or if repeated runs of the same 200 prompts (the article itself reports run-to-run variation) show the count is not stable. Narrows if the forty quotes from crawler-blocking publishers are excluded: query the incorrect share among the 160 quotes whose publishers allowed OAI-SearchBot. Narrows in time against the March 2025 follow-up, where the same authors report 134 incorrectly identified articles of 200 for ChatGPT under a changed protocol.

✓ checked by Claude · 1not yet attackedpositioned

Statement

The Tow Center randomly selected twenty publishers (with licensing deals with OpenAI, in litigation against it, or unaffiliated; allowing or blocking its search crawler), pulled block quotes from ten articles of each (two hundred quotes) and asked ChatGPT Search to identify the source of each. Quotes were chosen so that Google or Bing return the source article among the top three results. The researchers judged correctness on publisher name, URL and article date. ChatGPT returned partially or entirely incorrect responses on 153 occasions (the article spells it 'a hundred and fifty-three') and acknowledged an inability to respond accurately seven times; forty of the two hundred quotes came from publishers that had blocked its search crawler. More than a third of responses included incorrect citations, and the same query repeated typically returned a different answer. This is reverse attribution of a given quote by one product in November 2024, not support of generated claims; rated manually by the researchers.

Collection

The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. Correctness was assigned manually by the researchers; no second-rater agreement is reported, and the article calls its tests initial and says more rigorous experimentation is needed to understand the true frequency of errors. Counter-check that exists: the article points to a GitHub repository with the data. OpenAI's spokesperson responded that the study is an atypical test of the product and that data and methodology had been withheld; the Tow Center states it described methodology and observations to OpenAI but did not share the data before publication.

Falls when

Falls if the released data, relabelled by a second rater, give a materially different count than 153, or if repeated runs of the same 200 prompts (the article itself reports run-to-run variation) show the count is not stable. Narrows if the forty quotes from crawler-blocking publishers are excluded: query the incorrect share among the 160 quotes whose publishers allowed OAI-SearchBot. Narrows in time against the March 2025 follow-up, where the same authors report 134 incorrectly identified articles of 200 for ChatGPT under a changed protocol.

Reflex

ChatGPT with search says so when it cannot find a source. Too coarse: in this test it returned a partially or entirely incorrect attribution 153 times and signalled inability seven times.

Evidence

https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php section 'Confidently wrong' | 2024-11-27 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, How ChatGPT Search (Mis)represents Publisher Content

Findings and answers · 0

No attacker has recorded a finding on this card yet.