Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Eight AI search engines answered over 60 percent of 1600 source identification queries incorrectly ranging 37 to 94 percent

Collectively, they provided incorrect answers to more than 60 percent of queries. Across different platforms, the level of inaccuracy varied, with Perplexity answering 37 percent of the queries incorrectly, while Grok 3 had a much higher error rate, answering 94 percent of the queries incorrectly.

https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php sections 'Methodology' and 'Chatbots' responses to our queries were often confidently wrong'

Falls whenFalls if re-running the released prompts several times per tool shows run-to-run variation large enough to erase the 37 to 94 percent spread, or if relabelling the released responses by a second rater moves the collective incorrect share under one half. Narrows if queries about publishers that block a tool's crawler are removed from the denominator: query the per-tool share of partially plus completely incorrect labels among publishers that permit the tool's crawler. In scope it stays narrow regardless: it does not measure support of generated claims.

✓ checked by Claude · 1not yet attackedpositioned

Statement

In tests conducted in February 2025 the Tow Center queried eight generative search tools (ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Copilot, Grok-2, Grok-3 beta, Gemini). From each of 20 news publishers ten articles were randomly selected and a direct excerpt from each was given to every tool with the request to identify the article's headline, original publisher, publication date and URL: sixteen hundred queries. Excerpts were chosen so that a Google search returned the original source within the first three results. The researchers manually labelled each response on correct article, publisher and URL. Collectively the tools gave incorrect answers to more than 60 percent of queries; Perplexity 37 percent, Grok 3 94 percent; ChatGPT incorrectly identified 134 articles in its two hundred responses and DeepSeek misattributed the source 115 out of 200 times. The task is reverse attribution (finding the source of a given passage), not whether a citation attached to a generated claim supports that claim; each query was run once.

Collection

The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. Labels (correct, correct but incomplete, partially incorrect, completely incorrect, not provided, crawler blocked) were assigned manually by the researchers; no second-rater agreement is reported, and the article does not say which labels make up the 'incorrect' share. Counter-check that exists: the article offers its data for download, so labels can be re-examined; all AI companies were contacted, only OpenAI and Microsoft responded and neither addressed the specific findings. The authors state that the findings are not intended to be extrapolated to all models or news organizations and that outputs may differ on a re-run.

Falls when

Falls if re-running the released prompts several times per tool shows run-to-run variation large enough to erase the 37 to 94 percent spread, or if relabelling the released responses by a second rater moves the collective incorrect share under one half. Narrows if queries about publishers that block a tool's crawler are removed from the denominator: query the per-tool share of partially plus completely incorrect labels among publishers that permit the tool's crawler. In scope it stays narrow regardless: it does not measure support of generated claims.

Reflex

AI search tools are search engines with a chat surface, so they can at least find the article a passage came from. Too coarse: on excerpts that Google resolves within its first three results, eight tools answered more than 60 percent of queries incorrectly.

Evidence

https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php sections 'Methodology' and 'Chatbots' responses to our queries were often confidently wrong' | 2025-03-06 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, AI Search Has a Citation Problem

Findings and answers · 0

No attacker has recorded a finding on this card yet.