Questions › Do the citations of AI research assistants support the claims they are attached to
89 percent of 98020 atomic claims in Google AI Overviews are supported by the cited pages and 11 percent are not
Of the 98,020 verified claims, 87,204 (88.97%) are consistent
https://arxiv.org/pdf/2605.14021v1 p. 6, Table 1; p. 7, Section 3.3.3; p. 11, Section 4.3
Falls whenA per-citation audit (does the specific linked page support the specific sentence it is attached to) on the same Overviews gives a supported share well below 89.0%, showing the any-cited-page unit inflates support; or a larger human validation than 100 verdicts shows the verifier's generosity toward Clear and Vague is frequent enough to move the 11.0%. Query to run: human per-link support audit on a sample of the 7,491 Overviews.
Statement
Across 98,020 atomic claims from 7,491 verifiable Google AI Overviews, each checked against the text of all pages that Overview cites, Table 1 prints: Clear 84.6% (82,933), Vague 4.4% (4,271), Ambiguous 1.4% (1,367), Incorrect 2.7% (2,609), Omitted 7.0% (6,840); Consistent (Clear plus Vague) 89.0%, Inconsistent 11.0%. Omitted means no cited source mentions the claim; Incorrect means a cited source contradicts it. The unit is the claim against the Overview's whole reference set, not a single citation against a single sentence. Labels are assigned by an LLM verifier (Grok 4.1 Fast Reasoning) and were validated against two human annotators on 100 verdicts (98 of 100 matched).
Collection
Authors are at Washington University in St. Louis; they are not affiliated with Google in the paper; arXiv preprint without venue. Method: 55,393 trending queries in 19 topical categories were issued over a 40-day window (March 13 to April 21, 2026); each AI Overview was decomposed into atomic claims and each claim verified against the full extracted body text of every reference the Overview cites, by an LLM pipeline (Grok 4.1 Fast Reasoning, temperature 0) assigning one of five labels. Pages were crawled hours or days after the Overview was generated. The counter-check that exists and was used: two annotators re-labelled a stratified sample of 100 claim-level verdicts (20 per label), inter-annotator Cohen's kappa 0.94, and the verifier matched the adjudicated human label on 98 of 100; both errors were the verifier being too generous with supported claims. Claim extraction was validated separately on 100 Overviews.
Falls when
A per-citation audit (does the specific linked page support the specific sentence it is attached to) on the same Overviews gives a supported share well below 89.0%, showing the any-cited-page unit inflates support; or a larger human validation than 100 verdicts shows the verifier's generosity toward Clear and Vague is frequent enough to move the 11.0%. Query to run: human per-link support audit on a sample of the 7,491 Overviews.
Reflex
AI answers in search mostly make claims their sources do not contain. Too coarse: in this measurement of one product, 84.6 percent of claims were explicitly supported by a cited page and 2.7 percent were contradicted by one.
Evidence
https://arxiv.org/pdf/2605.14021v1 p. 6, Table 1; p. 7, Section 3.3.3; p. 11, Section 4.3 | 2026-05-13 · arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews Plan 02b: quoted sentence is Section 4.3, p. 11; Table 1, p. 6 carries the split.
Findings and answers · 0
No attacker has recorded a finding on this card yet.