Questions › Do the citations of AI research assistants support the claims they are attached to

Standingderivation

Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support

Rests on Twelve of fourteen deep research agents keep links valid…; GPT-4o with RAG on 300 health questions has all URLs val…; Raising tool calls from 2 to 150 drops Fact Check from 7…; Eight LLM judges match human-reviewed labels on factual …

Falls whenA study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.06635 rejects supported citations at a rate large enough to close a 22-point gap. Query: human per-citation support audit with link-check on the same sample, any 2025 or 2026 assistant.

✓ checked by Claude · 1not yet attacked

Conclusion

In the two printed pairs of the one study in this stock that scores link validity and claim support on the same citations, link validity is 98.7 and 100.0 percent while support lies lower by 21.9 points (98.7 against 76.8 percent, the best of 14 deep-research agents) and 52.3 points (100.0 against 47.7, GPT-5.4). In the same paper's depth ablation the gap is not constant: with Link Works above 92 percent at every depth, Fact Check of 78.6 and 80.0 percent at 2 tool calls leaves a gap of at most 21.4 and 20.0 points, and Fact Check of 16.7 and 57.9 percent at 150 calls a gap of at least 75 and 34 points. A second study in a second domain shows the same direction but not a comparable pair: GPT-4o with web search had valid URLs in 100 percent of cases and supported statements in 75.7 percent, two rates with different denominators (URLs against statements) that cannot be subtracted as a per-citation gap.

Step

The benchmark of 14 agents contributes the per-citation triple (link works, relevant, fact check) on one set of citations, so the gap cannot come from different samples. The depth ablation of the same paper contributes that the gap is not constant: validity and relevance stay above 92 percent while support moves with the number of tool calls, so validity does not track support even within one system. The SourceCheckup anchor contributes direction only, not a pair: its URL validity is counted per URL and its support per statement, so the step does not subtract them. Hidden premise: the rubric LLM judges' support decisions are close enough to human decisions that a gap of 22 points or more is not a judge artefact; the judge study among the parents puts the best of eight such judges at F1 0.750 against human-reviewed labels on factual support, with all eight between 0.649 and 0.750, which leaves room for a judge to reject supported citations but not enough to close a 22-point gap on its own.

Breaking point

A study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.06635 rejects supported citations at a rate large enough to close a 22-point gap. Query: human per-citation support audit with link-check on the same sample, any 2025 or 2026 assistant.

Reflex

If the link works and the page is on topic, the citation is fine. Too coarse: in the two printed agent rows the link works for 98.7 and 100.0 percent of citations and the page is on topic for 95.7 and 93.7 percent, while 23.2 and 52.3 percent of the same citations fail the support check.

Notes

Supersedes the derivation 'Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs' after attacker run 3 (attacker-grok-4.7): it counted SourceCheckup's 100 percent valid URLs against 75.7 percent supported statements as a third same-citation pair although the two rates have different denominators, and its judge premise cited an 88.7 percent doctor-consensus agreement that stands on a card which was not a parent. The old card is kept with status broken; the judge study is now a parent.

Findings and answers · 0

No attacker has recorded a finding on this card yet.