Questions › Do the citations of AI research assistants support the claims they are attached to
Twelve of fourteen deep research agents keep links valid above 94 percent while fact check scores range from 24 to 77 percent
Fact Check scores range from 24% (OSS-120B) to 77% (Claude Opus 4.5)
https://arxiv.org/pdf/2605.06635v1 p. 6-7, Section 4.1 and Table 1
Falls whenA human-rated re-evaluation of the same reports finds Fact Check rates for the frontier models within a few points of their Link Works and Relevant Content rates, or shows that the LLM judge's Fact Check decisions disagree with human raters often enough to move the 24 to 77 percent range. Query to run: human audit of per-citation support on reports generated with the protocol of arXiv 2605.06635.
Statement
In a benchmark of 14 closed- and open-source LLMs run as deep-research agents on 130 queries, per-citation scores on three separate dimensions diverge: 12 of 14 models exceed 94% on Link Works (the URL resolves) and all frontier models exceed 80% on Relevant Content (the page is on topic), while Fact Check (the cited page supports the attributed claim) ranges from 24.4% (OSS-120B) to 76.8% (Claude Opus 4.5); Table 1 prints for GPT-5.4 100.0% / 93.7% / 47.7% and for Claude Opus 4.5 98.7% / 95.7% / 76.8%. Scores are assigned by rubric-based LLM judges calibrated through human review, not by human raters on every citation.
Collection
Authors are employees of PricewaterhouseCoopers U.S.; the paper is an arXiv preprint, not peer reviewed. Citations are extracted from the agents' Markdown reports with an AST parser, the cited page is retrieved, and an LLM-as-a-judge rubric scores each citation on the three dimensions. The authors are not the vendor of any evaluated model; a commercial interest of the firm in evaluation tooling is not declared in the paper and is recorded here as unknown. The counter-check that exists is human review of judge decisions; the reliability of such judges is measured separately in arXiv 2607.08700 by an overlapping author group.
Falls when
A human-rated re-evaluation of the same reports finds Fact Check rates for the frontier models within a few points of their Link Works and Relevant Content rates, or shows that the LLM judge's Fact Check decisions disagree with human raters often enough to move the 24 to 77 percent range. Query to run: human audit of per-citation support on reports generated with the protocol of arXiv 2605.06635.
Reflex
A citation with a working link to a relevant page supports the sentence it is attached to. Too coarse: link validity and topical relevance are measured separately here and sit 19 to 60 points above factual support in Table 1.
Evidence
https://arxiv.org/pdf/2605.06635v1 p. 6-7, Section 4.1 and Table 1 | 2026-05-07 · arXiv 2605.06635 · Onweller et al., Cited but Not Verified
Findings and answers · 0
No attacker has recorded a finding on this card yet.