Questions › Do the citations of AI research assistants support the claims they are attached to
Supported share of Google AI Overview claims ranges from 77 to 95 percent by topic with Climate at 48 set aside
per-category fidelity ranges from 76.85% in Jobs & Education to 94.77% in Health. The highest-fidelity categories are Health (94.77%), Politics (93.65%), Science (91.82%), Business & Finance (91.42%), and Hobbies & Leisure (91.40%).
https://arxiv.org/pdf/2605.14021v1 p. 11-12, Section 4.3 and Figure 5
Falls whenPage snapshots taken at the moment the Overview is generated show Climate and Jobs & Education at rates near the other categories (confirming the artifact) or still far below them (refuting it); or a per-topic human audit shows the verifier's accuracy differs by topic enough to reorder the categories. Query to run: per-category consistent share with same-minute page snapshots, and per-category human validation.
Statement
Broken down by the 19 topical categories of the query set, the share of AI Overview claims labelled consistent (Clear plus Vague) with the cited pages ranges from 76.85% in Jobs & Education to 94.77% in Health, with Politics at 93.65%, Science 91.82%, Business & Finance 91.42% and Hobbies & Leisure 91.40%. Climate stands at 48.23% and is excluded from that range by the authors: its Overviews report real-time weather values from structured feeds, and the cited pages had changed by the time they were crawled. With real-time query subsets removed, Technology rises to 89.41%, Shopping to 85.90% and Jobs & Education to 87.71%. Labels come from the same LLM verifier (Grok 4.1 Fast Reasoning) validated on 100 human-labelled verdicts.
Collection
Authors are at Washington University in St. Louis; they are not affiliated with Google in the paper; arXiv preprint without venue. Method: 55,393 trending queries in 19 topical categories were issued over a 40-day window (March 13 to April 21, 2026); each AI Overview was decomposed into atomic claims and each claim verified against the full extracted body text of every reference the Overview cites, by an LLM pipeline (Grok 4.1 Fast Reasoning, temperature 0) assigning one of five labels. Pages were crawled hours or days after the Overview was generated. The counter-check that exists and was used: two annotators re-labelled a stratified sample of 100 claim-level verdicts (20 per label), inter-annotator Cohen's kappa 0.94, and the verifier matched the adjudicated human label on 98 of 100; both errors were the verifier being too generous with supported claims. Claim extraction was validated separately on 100 Overviews.
Falls when
Page snapshots taken at the moment the Overview is generated show Climate and Jobs & Education at rates near the other categories (confirming the artifact) or still far below them (refuting it); or a per-topic human audit shows the verifier's accuracy differs by topic enough to reorder the categories. Query to run: per-category consistent share with same-minute page snapshots, and per-category human validation.
Reflex
Citation support for an AI search product is one number. Too coarse: within one product and one 40-day window the supported share spans 18 points by topic, and the lowest categories are driven by pages that change after the answer is generated.
Evidence
https://arxiv.org/pdf/2605.14021v1 p. 11-12, Section 4.3 and Figure 5 | 2026-05-13 · arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews
Findings and answers · 0
No attacker has recorded a finding on this card yet.