Questions › Do the citations of AI research assistants support the claims they are attached to
Grok 3 cited URLs leading to error pages in 154 of 200 source identification prompts in a Tow Center test of February 2025
More than half of responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages. Out of the 200 prompts we tested for Grok 3, 154 citations led to error pages.
https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php section 'Platforms often failed to link back to the original source'
Falls whenFalls if the released data show that a large part of the 154 error-page links resolved at query time, or returned bot-block or paywall errors to the researchers' browser rather than being non-existent. Narrows if 'fabricated' cannot be separated from 'broken': query each of the 154 URLs against the Wayback Machine and the publisher's sitemap to see whether the path ever existed. Narrows in time if later versions of the same tools, given the same prompts, return resolving links.
Statement
In the Tow Center test of eight generative search tools (tests conducted February 2025; 200 prompts per tool, each asking for headline, publisher, date and URL of the article a given excerpt came from), more than half of the responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages; for Grok 3, 154 citations out of the 200 prompts led to error pages. Grok 2 was prone to linking to the publisher's homepage rather than the specific article. The article says the problem happened 'far less frequently' with the other tools and prints no counts for them or for Gemini. This is a link-resolution measurement (does the cited URL lead to a page), a precondition of support and not support itself; determined manually by the researchers, each prompt run once.
Collection
The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. The researchers followed the cited URLs manually; the article does not separate fabricated URLs from URLs that once existed and broke, nor say when after generation the links were followed. Counter-check that exists: the article offers its data for download, so the URLs can be re-tested; xAI and Google did not respond to the request for comment.
Falls when
Falls if the released data show that a large part of the 154 error-page links resolved at query time, or returned bot-block or paywall errors to the researchers' browser rather than being non-existent. Narrows if 'fabricated' cannot be separated from 'broken': query each of the 154 URLs against the Wayback Machine and the publisher's sitemap to see whether the path ever existed. Narrows in time if later versions of the same tools, given the same prompts, return resolving links.
Reflex
A link in an AI answer at least leads somewhere. Too coarse: for one tool 154 of 200 prompts ended in citations to error pages, while for most other tools in the same test this was far less frequent.
Evidence
https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php section 'Platforms often failed to link back to the original source' | 2025-03-06 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, AI Search Has a Citation Problem
Findings and answers · 0
No attacker has recorded a finding on this card yet.