Questions › Do the citations of AI research assistants support the claims they are attached to

Question

For a person doing research, reporting or health reading with an AI assistant or deep-research agent of 2024 to 2026: treat an attached citation as sufficient evidence that the claim is supported and spot-check at most, or open the cited passage before relying on any cited claim?

For

Entries for relying without opening, each carried by a SUPPORTS edge: - 89.0 percent of 98,020 atomic claims in Google AI Overviews are supported by the cited pages (any-page unit). - By topic that share is 76.85 to 94.77 percent, with Health at 94.77. - Four deep-research agents score 77.96 to 90.24 percent citation accuracy on DeepResearch Bench (LLM-judged). - Five deep-research systems score 0.69 to 0.86 faithfulness of cited claims on ResearcherBench (LLM-judged). - Between two BBC rounds significant sourcing issues fell to 10 to 15 percent for three of four assistants, while Gemini stayed broadly the same at 47 percent. - Three conditions are each measured on a single component: citation-trained models at 88.9 and 84.2 percent human-rated precision on supplied documents, a URL-checking tool cutting non-resolving links to under 1 percent as the paper states it, cross-model agreement at a 95.6 percent database match rate (derivation on three conditions); none was measured on claim support in web research.

Against

Entries against relying without opening (costs of the measure), each carried by an OPPOSES edge: - Per-citation Fact Check of 14 deep-research agents on 130 queries runs from 24.4 percent (OSS-120B) to 76.8 percent (Claude Opus 4.5), with GPT-5.4 at 47.7, while 12 of 14 exceed 94 percent on Link Works (rubric-based LLM judge, preprint). - Four generative search engines reach 39.8 to 68.3 percent citation accuracy and leave 23.1 to 47.0 percent of query-relevant statements supported by none of the listed sources (DeepTRACE, 303 queries each, LLM judge at Pearson 0.62 against 100 human labels). - Deep-research configurations in the same audit reach 50.3 to 79.1 percent citation accuracy in Table 1 (the paper's running text gives Gemini Deep Research as 40.3 where Table 1 prints 50.3); the best case, GPT-5 Deep Research, leaves 12.5 percent of relevant statements unsupported, the other four 53.6 to 97.5 percent on a relevant subset of 12.4 to 45.5 percent of their statements. - The products on the benefit side read lower elsewhere: Perplexity Deep Research 58.0 percent in DeepTRACE against 90.24 on DeepResearch Bench, Gemini Deep Research 50.3 against 81.44; the spread between benchmarks for one product is 31 to 32 points. - Health reading: GPT-4o with web search on 300 consumer health questions has 100 percent valid URLs, 75.7 percent supported statements (95% CI 74.0-77.2) and 38.4 percent fully supported responses (26.7-49.3). On a different set, the paper's main set of 800 questions, the same system's response-level support is 55 percent: 31.0 percent (26.7, 35.8) on the 400 Reddit r/AskDocs questions against close to 80 on the 400 Mayo Clinic derived questions (peer reviewed, GPT-4o judge at 88.7 percent agreement with a three-doctor consensus). - Across seven LLMs on 800 medical questions 50 to 90 percent of responses are not fully supported by the sources they cite; the best system has 55 percent response-level support and approximately 30 percent of its individual statements unsupported. - Unit effect: the same answers give 75.7 percent supported statements and 38.4 percent fully supported responses; the benefit entries are claim-level or citation-level rates, and a whole answer is relied on at the response unit. - Any-page unit: the 89.0 percent for AI Overviews and the 75.7 percent of SourceCheckup count a claim as supported if any cited page supports it; both are claim-level rates and upper bounds on the share of claims supported by their own attached citation, not bounds on a rate counted over citations, and neither paper prints how far below the attached-citation share lies. - Gap between link validity and support: where one study scores both, link validity is 98.7 to 100 percent in the printed pairs and support lies lower by 21.9 points (98.7 against 76.8), 52.3 points (100.0 against 47.7) and 24.3 points (100 against 75.7); in the depth ablation, with Link Works above 92 percent at every depth, the gap is at most 21.4 and 20.0 points at 2 tool calls and at least 75 and 34 points at 150; a spot check that stops at a working, on-topic link does not register this gap. - Degradation with depth: per-citation Fact Check falls from 78.6 to 16.7 percent (GPT-5.4) and from 80.0 to 57.9 percent (Claude Opus 4.6) as the tool-call budget rises from 2 to 150, with Link Works and Relevant Content above 92 percent at every depth; two models, no n per level, no intervals printed. - Links that lead nowhere: across 10 models on DRBench 3.0 to 13.3 percent of citation URLs are classified as hallucinated and 5.4 to 18.5 percent do not resolve; pooled deep-research agents 10.7 and 16.2 percent; 8.22 percent non-resolving over 168,021 ExpertQA URLs (automated, no rater, stated as lower bounds). - News: journalists at 22 public media rated 31 percent of 2,709 responses as having significant sourcing issues (Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15 percent); human raters, response unit, and the category includes responses with no source at all. - Quote fidelity: 8 altered or absent BBC quotes across 62 quote-bearing responses (13 percent as the report computes it) and 12 percent of 1,053 quote-bearing responses with significant quote-accuracy issues, both human-rated; the BBC took part in both audits, so they are not independent. - Uncited claims: on ResearcherBench 0.84 faithfulness stands beside 0.34 groundedness (OpenAI Deep Research) and 0.80 beside 0.31 (Grok3 DeeperSearch), while Sonar Reasoning Pro has the lowest faithfulness (0.62) and the highest groundedness (0.68); groundedness, 0.31 to 0.68 in these rows, is the share of claims that carry a citation at all, and a support rate over cited claims is silent on the claims left uncited. - Judge uncertainty: validated support judges disagree with human raters on 2 percent, 11.3 percent and 14.9 to 22.4 percent of decisions, while the rate-level gaps the same validations print are smaller (0.0 to 3.3 points over six printed ELI5 pairs, 2.0 points in SourceCheckup); four of the six benefit entries are LLM-judged, of their judges the card's validations cover the AI Overviews verifier only, and the ResearcherBench anchor reports no human check of its support judge. - Generated references: invalid or fabricated shares of 11.4 to 56.8 percent (ten commercial LLMs), 14.23 to 94.93 percent (thirteen LLMs), 28.6 to 91.4 percent (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 references from eight free chatbots; this is existence of the reference, not support. The four printed rates come from studies that asked models for reference lists; two further studies in the card had models draft text with references. Running without retrieval is stated for one of the six studies only, the others pool or do not state it, and no study states that its models cited a retrieved page.

Cost of the alternative (opening every cited passage): unmeasured. The stock holds no card on the time or effort of verification, so this entry carries no card and no edge. The nearby printed quantities are counts, not costs: 4.35 to 111.21 supported citations per task on DeepResearch Bench, 3.6 to 57.2 sources per deep-research answer in DeepTRACE. Opening is also not always possible: roughly 15 percent of source URLs could not be scraped in DeepTRACE.

Unresolved

What error rate in cited claims is tolerable depends on the use: a literature scan, a news report and a health decision do not share a threshold, and no anchor will settle that. Unmeasured: the time cost of opening every cited passage against the cost of acting on an unsupported claim; no source in the stock measures either. Unmeasured: per-citation support on ordinary user queries, as opposed to benchmark tasks. Consistency question: readers do not open every footnote of a human-written review either; what error rate is tolerated there is not in this stock.

Notes

Cost side searched on 2026-09-19 by costside-trustwork-e7 without sight of the benefit author's reasoning. Read: all 42 anchors and 18 derivations at statement level, the entered cards and the six benefit cards in full.

Benefit entries read against their cards (nothing removed): - 89.0 percent AI Overviews: any-page unit (a claim-level rate, an upper bound on the share of claims supported by their own attached citation), one product, trending queries, pages crawled hours or days later, LLM verifier; both verifier errors in the 100 validated verdicts were the verifier being too generous. - Health at 94.77 percent: same unit and verifier. The stock's other health measurement prints 75.7 percent statement-level and 38.4 percent response-level support for GPT-4o with web search; different product, different unit. - DeepResearch Bench 77.96 to 90.24: the Benefits line selects the four deep-research agents; the same column runs from 39.36 percent. The judge (Gemini-2.5-Flash) is of the same family as the Gemini systems it rates in that table; validated on 100 pairs. - ResearcherBench 0.69 to 0.86: rate over cited claims only; its anchor reports no human check of the support judge; 65 questions on frontier AI topics. - BBC rounds: the line now carries Gemini at 47 percent beside the three assistants at 10 to 15. Response unit, mixed criterion, different product tiers between rounds, 237 responses over four assistants; the card itself names that the drop may be carried by no-source responses falling from 25 to one. The 22-organisation round of the same weeks prints ChatGPT at 24 percent. - Three conditions: as reworded by the benefit author the line now states itself that none was measured on claim support in web research. Per the card only the URL tool is a before-and-after measurement (the one causal reading); the citation-training figures compare different models. The URL tool acts on link resolution, which sits 21.9 to 52.3 points above support where both are measured; cross-model agreement acts on reference existence. A reader does not obtain these conditions by relying.

Limits of the cost entries: - All per-citation cost rates except the news, quote, link and generated-reference entries are LLM-judged, as are four benefit entries. The judge derivation's band (2 to 22 percent at decision level, 0.0 to 3.3 and 2.0 points at rate level where printed) belongs to the judges it validated: of the entries here the SourceCheckup judge (88.7 percent agreement, behind the two health entries) and the AI Overviews verifier (98 of 100, behind two benefit entries). The other judges carry their own validations on their anchors: DeepTRACE Pearson 0.62 on 100 labels, DeepResearch Bench 96 and 92 percent agreement on 100 pairs, the 14-agent rubric judge calibrated through manual review of 50 to 100 judgments; the band is not transferred to them. The direction of judge error is unsettled in this stock: false negative rates of 0.183 to 0.470 on a human-reviewed benchmark come from the same author group as the 14-agent Fact Check and depth-ablation entries, so those two rates may read low. - The depth ablation prints no n per level and covers two models; its GPT-5.4 series is not monotone. - DeepTRACE uses 168 debate and 135 expertise queries and does not say how partial support is binarised; its authors include one at Microsoft, vendor of one evaluated engine. - The 31 percent news figure is a response-level rating that includes absent sources; it bounds per-citation support without measuring it. The EBU and BBC are publishers with a declared position. - The generated-reference rates come from four studies that asked models for reference lists (two further studies had models draft text with references). Retrieval status: stated as absent for one study, pooled with and without online search in one, unstated in four, two of which include Perplexity; no study states that its models cited a retrieved page, so the entry does not say whether retrieval was present. For systems stated to be retrieval-backed the stock prints non-existent links at 3 to 13 percent.

Considered and not entered: the Tow Center reverse-attribution tests and the DRACO axis (neighbouring quantities, per their derivation); the 2023 baselines (outside the 2024 to 2026 window); the published-record audits (do not identify a tool, measure no reliance); the orchestrator localisation (open-source pipelines, global citation recall 58.7, 28.5 and 7.1 percent); the 24-to-94 range and no-single-value derivations (summaries holding both sides). What an unsupported claim costs in the use at hand has no card; it stays under Unresolved.

Findings and answers · 0

No attacker has recorded a finding on this card yet.