Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Raising tool calls from 2 to 150 drops Fact Check from 79 to 17 percent for GPT-5.4 and from 80 to 58 for Claude Opus 4.6

GPT-5.4 shows the steepest decline, from 79% to 17% (62%). Claude Opus 4.6 demonstrates the greatest resilience, declining from 80% to 58% (22%).

https://arxiv.org/pdf/2605.06635v1 p. 7-8, Section 4.3 and Tables 2-3

Falls whenA rerun of the depth ablation with per-level citation counts and intervals, human-rated or with a judge validated against human labels, shows Fact Check at 150 calls within the interval of Fact Check at 2 calls for both models, or shows the decline only for GPT-5.4. Query to run: per-citation support rate as a function of tool-call budget, more than two models, with n per level, on reports generated with the protocol of arXiv 2605.06635.

✓ checked by Claude · 1not yet attackedunknown

Statement

In a search-depth ablation, two models run as deep-research agents (GPT-5.4 and Claude Opus 4.6) were capped at seven maximum tool-call budgets (2, 10, 30, 50, 70, 100, 150). Per-citation Fact Check (the cited page supports the attributed claim) fell for GPT-5.4 from 78.6% at 2 calls to 16.7% at 150 calls (Table 2) and for Claude Opus 4.6 from 80.0% to 57.9% (Table 3); the paper rounds these to declines of 62 and 22 percentage points and averages them to the 'approximately 42%' of its abstract. Link Works and Relevant Content stayed above 92% at every depth. The GPT-5.4 series is not monotone (35.5% at 70 calls, 37.2% at 100) and its largest single step is between 2 and 10 calls (78.6% to 45.9%). Scores come from the same rubric-based LLM judge as the main benchmark, not from human raters on each citation.

Collection

Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. The pipeline extracts citations from the agents' Markdown reports, retrieves the cited page and has an LLM judge score three dimensions; the Fact Check judge was calibrated through manual review of 50 to 100 judgments. The paper does not print, for the ablation, the number of queries, the number of citations per depth level, or any interval, so the width of each percentage is not checkable from the text. The 42 percent figure is a mean of two percentage-point differences, not a relative drop. The authors are not the vendor of either model; a commercial interest of the firm is not declared and is recorded as unknown. Counter-check that exists: human audit of the judged citations and a repeat over more than two models; neither is reported.

Falls when

A rerun of the depth ablation with per-level citation counts and intervals, human-rated or with a judge validated against human labels, shows Fact Check at 150 calls within the interval of Fact Check at 2 calls for both models, or shows the decline only for GPT-5.4. Query to run: per-citation support rate as a function of tool-call budget, more than two models, with n per level, on reports generated with the protocol of arXiv 2605.06635.

Reflex

A research agent that searches more reads more and therefore cites more accurately. Too coarse: in this ablation link validity and relevance stay flat while measured support falls with search depth, for two models and without reported sample sizes.

Evidence

https://arxiv.org/pdf/2605.06635v1 p. 7-8, Section 4.3 and Tables 2-3 | 2026-05-07 · arXiv 2605.06635 · Onweller et al., Cited but Not Verified

Findings and answers · 0

No attacker has recorded a finding on this card yet.