# Do the citations of AI research assistants support the claims they are attached to This document is the whole evidence stock as one file, rendered from the engine export; nothing here is written for the file. Order: the question, then 9 core derivations (the inferences the owner marks as carrying the answer), then 13 further derivations, then 46 anchors (measurements with coordinate and verbatim quoted span), then the tradeoff. Every card is a "##" heading; inside a card the sections follow the schema (statement, collection, falls when, evidence for an anchor; conclusion, step, breaking point for a derivation). "→" is an outgoing edge, "←" an incoming one; every target is an address on this site. type: question ### Question For AI assistants and deep-research agents evaluated in published measurements between 2024 and 2026, what share of the citations attached to factual claims point to a passage that actually supports that claim, as distinct from a link that merely resolves or a source that is merely on topic? ### Not asked - Whether AI research assistants are trustworthy in general, or better or worse than human researchers. - Whether any named vendor is honest, misleading, or negligent; vendors appear only as sources of published claims and evaluations. - The rate of fabricated (non-existent) references as a quantity of its own; it enters only where a study measures it beside citation support, because a reference can exist and still not support the claim. - Whether users read or verify citations, and what they believe after reading them. - Legal liability for a wrong citation. - Which model is best; the stock records measurements per system and date, never a ranking. ### Scope Primary sources: peer-reviewed papers and preprints (arXiv, journal sites) that measure citation support, link validity or relevance for LLM assistants, generative search engines and deep-research agents; institutional evaluations (EBU/BBC, Tow Center, Reuters Institute); vendors' own published evaluation pages. Period: measurements published 2024 to 2026, with earlier benchmarks admitted as baselines where a 2024 to 2026 study builds on them. Measures: citation support or fact-check rate, citation precision and recall, link validity, topical relevance, and their change with search depth or task type. Both directions are searched: measurements of low support and measurements of high support or of conditions under which citations hold. Counting unit, fixed after the scope attack of 2026-09-23: the share the question asks for is counted per attached citation, that is per (claim, cited source) pair as the benchmarks print it; a per-cited-claim rate counts as a value only where the study attaches one citation per claim. Response-level rates (every statement in a response supported), citation recall (the share of claims that carry any supporting citation), link validity and topical relevance are neighbouring quantities: the stock records them beside the share, as bounds or as context, never as values of it. Where a study reports support by any listed source rather than by the attached citation, the stock says so on the card and treats the value as an upper bound. ### Notes Scope narrowed on 2026-09-23 after attacker run 3 (attacker-grok-4.7, x-attack-scope, verdict failed): the scope listed recall, link validity and relevance among the measures and did not fix the counting level, although those choices move published figures by tens of points; the title's quantity is now the only value and the others are declared neighbouring quantities. Relationships - ADDRESSES ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - ADDRESSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - ADDRESSES ← https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks (The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks) Address: https://trustillery.com/entity/stocks/ai-citations--do-the-citations-of-ai-research-assistants-support-the-claims-they-are-attached-to ## Counting per response instead of per statement halves the support rate in the same data type: derivation · status: open · confidence: high ### Conclusion In SourceCheckup the same answers of GPT-4o with web search give 75.7 percent supported statements and 38.4 percent fully supported responses on HealthSearchQA, and about 70 percent supported statements against 55 percent fully supported responses on the 800-question set; a response-level rate and a statement-level rate are different quantities and cannot be set side by side across studies. ### Step The first parent contributes both units on one sample with intervals (74.0 to 77.2 and 26.7 to 49.3). The second contributes the same pair on the main set (approximately 30 percent of statements unsupported, 55 percent response-level support) and the summary that 50 to 90 percent of responses are not fully supported across seven models. The arithmetic behind it is that a response counts as supported only if every statement is, so the response rate falls with the number of statements per response; the paper does not print that number, so the step does not compute a prediction. ### Breaking point A dataset in which response-level full support is not lower than statement-level support although responses contain several statements, which would mean unsupported statements cluster in few responses. Query: distribution of unsupported statements per response in the SourceCheckup release. ### Reflex Half of AI answers are not supported by their sources. Too coarse as a citation rate: in the same data three quarters of the individual statements are supported. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--for-seven-llms-on-800-medical-questions-50-to-90-percent-of-responses-are-not-fully-supported-by-the-sources-they-cite - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks (The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks) Address: https://trustillery.com/entity/stocks/ai-citations--counting-per-response-instead-of-per-statement-halves-the-support-rate-in-the-same-data ## Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent type: derivation · status: superseded · confidence: high ### Conclusion Across the five LLM-judged measurements of 2025 and 2026 in this stock that score whether the cited page supports the attached claim, the printed per-system rates run from 24.4 percent (OSS-120B as a deep-research agent) to 94.04 percent (Claude-3.5-Sonnet with search on DeepResearch Bench); for commercial deep-research products alone the printed range is 50.3 to 90.24 percent. ### Step Each parent contributes one benchmark's range under its own task set and judge: 24.4 to 76.8 (14 agents, 130 queries), 39.8 to 68.3 (four generative search engines, DeepTRACE), 50.3 to 79.1 (deep-research configurations, DeepTRACE), 39.36 to 94.04 with 77.96 to 90.24 for the four deep-research agents (DeepResearch Bench), and 0.62 to 0.86 (ResearcherBench faithfulness). The step only takes the minimum and maximum of printed values and restricts the product range to systems the parents label deep research. It widens no parent's scope: all five are per-citation or per-cited-claim support under an LLM judge; none is human-rated on every citation. ### Breaking point A parent turns out to measure a different quantity (for instance support by any listed source instead of by the attached citation), which would remove its values from the range; or a human-rated replication of any one benchmark lands outside its printed range. Query: compare the metric definitions of the five papers clause by clause. ### Reflex Deep research tools get their citations right about 80 to 90 percent of the time. Too coarse: that is the upper part of one benchmark; the printed values for comparable products go down to 50 percent and for agents built on open models to 24. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-generative-search-engines-reach-40-to-68-percent-citation-accuracy-and-leave-23-to-47-percent-of-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-deep-research-agents-score-78-to-90-percent-citation-accuracy-on-deepresearch-bench-under-an-llm-judge - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge - ADDRESSES → https://trustillery.com/entity/stocks/ai-citations--do-the-citations-of-ai-research-assistants-support-the-claims-they-are-attached-to - SUPERSEDES ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) Address: https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent ## Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken type: derivation · status: open · confidence: medium ### Conclusion Across the five LLM-judged measurements of 2025 and 2026 in this stock that score whether the cited page supports the attached claim, the printed per-system rates run from 24.4 percent (OSS-120B run as a deep-research agent) to 94.04 percent (Claude-3.5-Sonnet with search on DeepResearch Bench). For commercial deep-research products alone the printed range is 50.3 to 90.24 percent by the tables, but the lowest value belongs to a system for which the same paper's running text prints 40.3 percent, so the floor reads 40.3 or 50.3 depending on which line of arXiv 2509.04499 is taken; and a commercial model run as a deep-research agent by benchmark authors rather than sold as a product (GPT-5.4 at 47.7 percent Fact Check) sits below the product floor either way. Two values adopted from proposal 003 fall inside the span and below the product floor: OpenAI DeepResearch at 39.9 percent citation precision on DeepScholar-Bench (GPT-4o entailment judge, arXiv-only corpus, 80 percent human agreement on the judge) and OpenAI Deep Research at 78.87 percent cited-statement consistency on ReportBench (gpt-4o judge, no validation printed); with the first, the printed floor for a commercial deep-research product is 39.9 percent. ### Step Each parent contributes one benchmark's range under its own task set and judge: 24.4 to 76.8 (14 agents, 130 queries, with GPT-5.4 at 47.7), 39.8 to 68.3 (four generative search engines, DeepTRACE), 50.3 by Table 1 or 40.3 by the text to 79.1 (deep-research configurations, DeepTRACE), 39.36 to 94.04 with 77.96 to 90.24 for the four deep-research agents (DeepResearch Bench), 0.62 to 0.86 (ResearcherBench faithfulness), .399 for one product (DeepScholar-Bench, per-citation entailment) and 78.87 for one product (ReportBench, per-cited-statement consistency). The step takes the minimum and maximum of every printed value, text and table alike, and states the product range with both readings where a parent prints two values for one system; it restricts the product range to systems the parents label as sold deep-research products and names separately the benchmark-run configuration that falls below it. It widens no parent's scope: all five are per-citation or per-cited-claim support under an LLM judge; none is human-rated on every citation. ### Breaking point A parent turns out to measure a different quantity (for instance support by any listed source instead of by the attached citation), which would remove its values from the range; a human-rated replication of any one benchmark lands outside its printed range; or the authors of arXiv 2509.04499 resolve the 40.3 against 50.3 print in a revision. Query: compare the metric definitions of the five papers clause by clause; check arXiv 2509.04499 for a v2. ### Reflex There is one number for how often AI citations hold. Too coarse: the printed per-citation support rates of 2025 and 2026 run from a quarter to more than nine tenths, and even the floor for one product depends on which line of one paper is read. ### Notes Supersedes the derivation 'Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent' after attacker run 3 (attacker-grok-4.7): its product floor of 50.3 took Table 1 of arXiv 2509.04499 and ignored the 40.3 the same paper prints in its text for the same system, against its own rule of taking the minimum of printed values; and GPT-5.4 run as a deep-research agent at 47.7 percent was neither included nor named. The old card is kept with status broken. Extended 2026-09-23 after proposal 003 (completeness attack of attacker-grok-4.7): two adopted anchors added as parents; the product floor moves from 40.3 or 50.3 to 39.9. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-generative-search-engines-reach-40-to-68-percent-citation-accuracy-and-leave-23-to-47-percent-of-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-deep-research-agents-score-78-to-90-percent-citation-accuracy-on-deepresearch-bench-under-an-llm-judge - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge - SUPERSEDES → https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--deepscholar-bench-openai-deepresearch-scores-399-citation-precision-under-gpt-4o-entailment-judging-on-63-arxiv-related-work-queries - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--reportbench-7887-percent-of-openai-deep-research-cited-statements-judged-consistent-with-their-cited-page-against-3143-for-o3-with-search-tools - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks (The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks) Address: https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken ## The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another type: derivation · status: open · confidence: medium ### Conclusion Perplexity Deep Research is printed at 58.0 percent citation accuracy in DeepTRACE, 90.24 percent in DeepResearch Bench and 0.85 faithfulness in ResearcherBench; Gemini Deep Research at 50.3, 81.44 and 0.86. The spread between benchmarks for one product (31 to 32 points) is as large as the spread between products within one benchmark (29 points in DeepTRACE, 12 in DeepResearch Bench). ### Step DeepTRACE contributes the low reading (debate and expertise queries, GPT-5 judge, results as of August 2025), DeepResearch Bench the high reading (100 PhD-level research tasks, Gemini-2.5-Flash judge, mid 2025), ResearcherBench a third reading close to the high one (65 frontier-AI questions, GPT-4.1 judge, March to April 2025). All three define the rate as supported citations or cited claims over all citations or cited claims. Hidden premise: the products were comparable across the three evaluation dates; the papers name products, not model versions, so a version change between March and August 2025 cannot be excluded. The step does not say which benchmark is right. ### Breaking point Evidence that the products changed materially between the evaluation dates in a way that explains 30 points, or a run of one product on two of the benchmarks in the same week with one judge that removes the gap. Query: release notes of Perplexity and Gemini deep research between March and August 2025; judge-swap re-run on the released DeepResearch Bench pairs. ### Reflex A product has a citation accuracy. Too coarse: the printed value for one product moves by 30 points with the benchmark. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-deep-research-agents-score-78-to-90-percent-citation-accuracy-on-deepresearch-bench-under-an-llm-judge - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks (The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks) Address: https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another ## The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks type: derivation · status: open · confidence: medium ### Conclusion For assistants and deep-research agents measured in 2025 and 2026 the published share of citations that support their claim runs from 24.4 to 94.04 percent. Within one benchmark the printed spread between systems is 52.4 points (24.4 to 76.8) and 54.7 points (39.36 to 94.04); one deep-research product moves 31 to 32 points between benchmarks, against spreads of 29 and 12 points between deep-research products within one benchmark. Counted per response instead of per statement, the same answers of one system in a dataset of 2024 give 38.4 against 75.7 percent, which is a different quantity and not a correction. Where automatic support judges were validated against humans they disagree on 2 to 22 percent of decisions; those validations concern other judges than the ones behind the 24 to 94 span. An answer to the question is a range with its conditions, not a number. ### Step The range derivation contributes the span of printed values and the spreads within a benchmark (76.8 minus 24.4, 94.04 minus 39.36). The same-product derivation contributes that for deep-research products the benchmark moves one product as far as products differ within a benchmark; the step does not extend that to all systems, where the within-benchmark spread is larger. The unit derivation contributes that statement-level and response-level rates cannot be set side by side. The judge derivation contributes the decision-level disagreement of the judges that were validated in its parents; the step does not transfer that band to the judges behind the span, three of which print a validation on their anchors (DeepTRACE: Pearson 0.62 on 100 tasks; DeepResearch Bench: 96 and 92 percent agreement on 100 pairs; the 14-agent benchmark: rubric judges validated at F1 0.75 in a separate study) while ResearcherBench prints none. The step is a conjunction and adds no premise beyond the four. It does not rank systems and does not say which benchmark is closest to ordinary use. ### Breaking point A human-rated measurement on a representative sample of ordinary user queries across current products; it would give the question a value with an interval and make this card a statement about benchmarks only. Query: any 2026 human audit of per-citation support on sampled real-world queries. ### Reflex About X percent of AI citations are wrong. Any single X quoted for this is one cell of a table whose rows are benchmarks, units and raters. ### Notes Attacker run 3 (attacker-grok-4.7, x-attack-step): applied, the step over-generalised; three of the four benchmarks behind the span print a judge validation, ResearcherBench does not. Conclusion unchanged. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--counting-per-response-instead-of-per-statement-halves-the-support-rate-in-the-same-data - ADDRESSES → https://trustillery.com/entity/stocks/ai-citations--do-the-citations-of-ai-research-assistants-support-the-claims-they-are-attached-to - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken Address: https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks ## Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions type: derivation · status: open · confidence: medium ### Conclusion The four judge validations among this card's parents put disagreement between an automatic support judge and human raters at 2 percent (98 of 100 verdicts, AI Overviews), 11.3 percent (88.7 percent agreement with a three-doctor consensus, where doctors agreed with each other at 86.1 percent) and 14.9 to 22.4 percent (ALCE's NLI metric, accuracy 85.1 percent for recall and 77.6 for precision); on an adversarial benchmark the best of eight judges reaches F1 0.750 on factual support. These are decision-level disagreements of these judges on their own material. The rate-level gaps the parents print are smaller: 0.0 to 3.3 points between human and automatic system scores on ELI5 (six printed pairs), and 2.0 points (40.4 against 42.4 percent fully supported responses) in SourceCheckup. ### Step Each parent contributes one validated judge on its own material; the step only converts agreement into disagreement (100 minus the printed agreement) and places them side by side. The doctors' own 86.1 percent agreement is carried along because it bounds what agreement with humans can mean. The F1 anchor is kept separate because its material is synthetic and adversarial and its metric is not an agreement rate. Hidden premise: validation samples of 100 to 400 items are representative of the full runs. ### Breaking point A larger validation (a thousand or more items) of one of these judges on its own benchmark shows disagreement well outside 2 to 22 percent, or shows that disagreement concentrates on one system so that rankings change. Query: per-system breakdown of judge-human disagreement in any of the four papers' released data. ### Reflex The numbers come from papers, so they are measured. Too coarse: most of them are scored by an automatic judge, and where such a judge was validated its agreement with humans was between 78 and 98 percent of decisions. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--89-percent-of-98020-atomic-claims-in-google-ai-overviews-are-supported-by-the-cited-pages-and-11-percent-are-not - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-support-judge-agrees-with-a-three-doctor-consensus-on-89-percent-of-400-pairs-and-doctors-among-themselves-on-86-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--alce-automatic-citation-metrics-agree-with-human-raters-at-kappa-0698-for-recall-and-0525-for-precision - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--eight-llm-judges-match-human-reviewed-labels-on-factual-support-of-citations-with-f1-of-65-to-75-percent-none-separable - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-share-of-supporting-citations-has-no-single-published-value-and-one-deep-research-product-moves-31-to-32-points-between-benchmarks (The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks) Address: https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions ## Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support type: derivation · status: open · confidence: high ### Conclusion In the two printed pairs of the one study in this stock that scores link validity and claim support on the same citations, link validity is 98.7 and 100.0 percent while support lies lower by 21.9 points (98.7 against 76.8 percent, the best of 14 deep-research agents) and 52.3 points (100.0 against 47.7, GPT-5.4). In the same paper's depth ablation the gap is not constant: with Link Works above 92 percent at every depth, Fact Check of 78.6 and 80.0 percent at 2 tool calls leaves a gap of at most 21.4 and 20.0 points, and Fact Check of 16.7 and 57.9 percent at 150 calls a gap of at least 75 and 34 points. A second study in a second domain shows the same direction but not a comparable pair: GPT-4o with web search had valid URLs in 100 percent of cases and supported statements in 75.7 percent, two rates with different denominators (URLs against statements) that cannot be subtracted as a per-citation gap. ### Step The benchmark of 14 agents contributes the per-citation triple (link works, relevant, fact check) on one set of citations, so the gap cannot come from different samples. The depth ablation of the same paper contributes that the gap is not constant: validity and relevance stay above 92 percent while support moves with the number of tool calls, so validity does not track support even within one system. The SourceCheckup anchor contributes direction only, not a pair: its URL validity is counted per URL and its support per statement, so the step does not subtract them. Hidden premise: the rubric LLM judges' support decisions are close enough to human decisions that a gap of 22 points or more is not a judge artefact; the judge study among the parents puts the best of eight such judges at F1 0.750 against human-reviewed labels on factual support, with all eight between 0.649 and 0.750, which leaves room for a judge to reject supported citations but not enough to close a 22-point gap on its own. ### Breaking point A study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.06635 rejects supported citations at a rate large enough to close a 22-point gap. Query: human per-citation support audit with link-check on the same sample, any 2025 or 2026 assistant. ### Reflex If the link works and the page is on topic, the citation is fine. Too coarse: in the two printed agent rows the link works for 98.7 and 100.0 percent of citations and the page is on topic for 95.7 and 93.7 percent, while 23.2 and 52.3 percent of the same citations fail the support check. ### Notes Supersedes the derivation 'Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs' after attacker run 3 (attacker-grok-4.7): it counted SourceCheckup's 100 percent valid URLs against 75.7 percent supported statements as a third same-citation pair although the two rates have different denominators, and its judge premise cited an 88.7 percent doctor-consensus agreement that stands on a card which was not a parent. The old card is kept with status broken; the judge study is now a parent. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--eight-llm-judges-match-human-reviewed-labels-on-factual-support-of-citations-with-f1-of-65-to-75-percent-none-separable - SUPERSEDES → https://trustillery.com/entity/stocks/ai-citations--where-one-study-measures-both-link-validity-sits-about-22-to-52-points-above-claim-support-in-the-printed-pairs - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support ## Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures type: derivation · status: open · confidence: high ### Conclusion In the one benchmark of this stock that scores link validity and claim support on the same citations, 12 of 14 deep-research agents keep Link Works above 94 percent while Fact Check runs from 24.4 to 76.8 percent; since a citation that fails to resolve cannot exceed the Link Works shortfall, at most 6 percentage points of the 23.2 to 75.6 points of support failure can be dead links for those 12 agents, and the rest are existing pages that do not support the claim. A second study on other systems puts hallucinated URLs at 3.0 to 13.3 percent and non-resolving ones at 5.4 to 18.5 percent, the same order as the Link Works shortfall. The larger failure in retrieval-backed systems is an existing page that does not support the claim. ### Step The 14-agent benchmark contributes the same-citation triple: for each citation the judge records whether the link works, whether the page is on topic and whether it supports the claim, so the two rates are on one sample and can be subtracted; 100 minus Link Works bounds the share of citations that fail for want of a page, and 100 minus Fact Check is the share that fail support for any reason, so their difference is a lower bound on support failures that resolve. The URL study contributes an independent measurement of the existence failure on ten other systems with bootstrap intervals and no rater; it is not subtracted from anything, it only shows that the existence failure is of the same size elsewhere. The step makes no cross-study comparison of magnitudes and no claim of an order of magnitude. ### Breaking point A per-citation release of arXiv 2605.06635 in which most citations failing Fact Check also fail Link Works (the two failures coincide on the same citations), or a current system whose non-resolving rate exceeds its support-failure rate on one sample. Query: join of link status and support verdict per citation in the release of arXiv 2605.06635. ### Reflex The main problem with AI citations is broken links. Too coarse: in the one benchmark that scores both on the same citations, links work for more than 94 percent of citations while a quarter to three quarters of them fail the support check. ### Notes Supersedes the derivation 'In retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check' after attacker run 3 (attacker-grok-4.7): its step compared two studies as orders of magnitude, which 13.3 against 23.2 is not, and used 100 minus Fact Check as if it showed the failing citations exist. The old card is kept with status broken. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--across-10-models-on-drbench-3-to-13-percent-of-citation-urls-are-hallucinated-and-5-to-18-percent-do-not-resolve - SUPERSEDES → https://trustillery.com/entity/stocks/ai-citations--in-retrieval-backed-systems-non-existent-links-are-the-smaller-failure-at-3-to-13-percent-against-23-to-76-percent-of-citations-failing-the-support-check Address: https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-support-on-the-same-citations-dead-links-explain-at-most-a-sixth-of-the-support-failures ## Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs type: derivation · status: superseded · confidence: high ### Conclusion In the three measurements of the stock (from two papers) that score link validity and claim support in the same study, link validity is 98.7 to 100 percent in the printed pairs while support lies lower by 21.9 points (98.7 against 76.8 percent, the best of 14 deep-research agents), 52.3 points (100.0 against 47.7, GPT-5.4) and 24.3 points (100 percent valid URLs against 75.7 percent supported statements, GPT-4o with web search on health questions). In the depth ablation the gap widens: with Link Works above 92 percent at every depth, Fact Check of 78.6 and 80.0 percent at 2 tool calls leaves a gap of at most 21.4 and 20.0 points, and Fact Check of 16.7 and 57.9 percent at 150 calls a gap of at least 75 and 34 points. ### Step The benchmark of 14 agents contributes the per-citation triple (link works, relevant, fact check) on one set of citations, so the gap cannot come from different samples. The SourceCheckup anchor contributes the same gap in a second domain with a judge validated against doctors, and with intervals. The depth ablation contributes that the gap is not constant: validity and relevance stay flat while support moves with the number of tool calls, so validity does not track support even within one system. Hidden premise: the LLM judges' support decisions are close enough to human decisions that a gap of 22 points or more is not a judge artefact; the one judge validation among the parents prints 88.7 percent agreement with a three-doctor consensus. ### Breaking point A study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.06635 rejects supported citations at a rate large enough to close a 22-point gap. Query: human per-citation support audit with link-check on the same sample, any 2025 or 2026 assistant. ### Reflex If the link works and the page is on topic, the citation is fine. Too coarse: in the two printed agent rows the link works for 98.7 and 100.0 percent of citations and the page is on topic for 95.7 and 93.7 percent, while 23.2 and 52.3 percent of the same citations fail the support check. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 - SUPERSEDES ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support (Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support) Address: https://trustillery.com/entity/stocks/ai-citations--where-one-study-measures-both-link-validity-sits-about-22-to-52-points-above-claim-support-in-the-printed-pairs ## Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question type: derivation · status: open · confidence: high ### Conclusion Six studies in the stock check the bibliographic references that models generate: in four the models were asked for reference lists (Naser, GhostCite, the eight chatbots, Chelli), in two they drafted text with references (medical research introductions in Omar; letters to the editor and original articles in Ozbek). Three state how the references were verified (two by automated matching against bibliographic databases, one by hand), three are observed as abstracts that do not state who verified: invalid or fabricated shares run from 11.4 to 56.8 percent (ten commercial LLMs), 14.23 to 94.93 percent (thirteen LLMs), 28.6 to 91.4 percent (GPT-3.5, GPT-4, Bard), and 39.8 percent of 400 references from eight free chatbots. All of them measure whether the reference exists and is bibliographically correct; none checks whether an existing reference supports the claim it is attached to, so none of these rates is a value of the quantity the question asks for. ### Step Naser and GhostCite contribute the two large automated audits and the ranges; the eight-chatbot study contributes a hand-checked sample; Chelli contributes a comparison against the reference lists of published systematic reviews under a rule of two wrong fields out of three, from an abstract that does not state who checked; Omar contributes a second medical comparison in which no model is free of fabricated references (correctness 77.2 and 54.0 percent); Ozbek contributes that a composite score mixing existence with topical relevance is still not a support check. Existence is a precondition of support: a fabricated reference supports nothing, so in the four reference-list studies the complements of these rates are loose upper bounds on support (a reference that does not exist supports nothing; one that exists may or may not), but the rates themselves bound nothing. Scope: only the Naser models are stated to have run without retrieval; the GhostCite rates pool runs with and without online search, which showed no consistent effect, and the other parents do not state the retrieval status. No parent states that its models cited a retrieved page, and two of the parents with unstated retrieval status include Perplexity (Ozbek, the eight chatbots), so the step does not claim that retrieval was absent. ### Breaking point One of these studies turns out to have checked the content of the cited work against the claim, which would make its rate a support rate; or a study shows that for retrieval-backed assistants existence failures and support failures are the same citations. Query: methods sections of the six papers for any content check. ### Reflex AI makes up 40 percent of its citations. Too coarse: that figure comes from studies that asked chatbots for reference lists and checked whether the references exist; it says nothing about whether the citations of a searching assistant support its claims. ### Notes Attacker run 3 (attacker-grok-4.7, x-attack-step): applied, step wording corrected in place; the rates do not bound support from above, their complements do. Conclusion unchanged. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-35-and-gpt-4-and-bard-hallucinated-about-40-and-29-and-91-percent-of-references-generated-for-systematic-reviews - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--of-400-references-from-eight-free-chatbots-about-27-percent-were-fully-correct-and-40-percent-erroneous-or-fabricated - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--thirteen-llms-asked-for-computer-science-references-produced-invalid-citations-at-rates-from-14-to-95-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gemini-ultra-reached-77-percent-reference-correctness-against-54-percent-for-gpt-4-in-medical-research-introductions - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--over-3150-generated-references-chatgpt-had-the-lowest-and-perplexity-the-highest-mean-reference-hallucination-score - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question ## Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research type: derivation · status: open · confidence: medium ### Conclusion Citation-trained models score 88.9 and 84.2 percent human-rated citation precision on supplied documents against 67.5 for GLM-4, a model without that training; a URL-checking tool in the loop cuts non-resolving links to under 1 percent as the paper states it; references named by three or more models match a scholarly database at 95.6 percent against 16.5 for references named by one model; and on one open web-research pipeline an orchestrator instruction or a snippet substitution raised citation precision from 87.6 to 91.0 and 94.1 percent and recall from 64.5 to 69.7, under an LLM judge and with gains of the order of the printed standard errors. The first three hold for their own component (grounding in a given document, link resolution, existence) and were not measured on claim support in web research; the fourth is the only condition in the stock measured on citation precision of a web-research system, on one pipeline, with a precision definition the paper does not print. ### Step The two LongCite anchors contribute the comparison of citation-trained and other models, a comparison across different models and not one model before and after training, once LLM-judged across five datasets and once human-rated on one, in a closed-document setting. The URL tool anchor contributes the link component as a before-and-after measurement on the same questions and states that support of the replacement links was not measured. The consensus anchor contributes the existence component for references recalled without retrieval, a status stated on the sibling anchor of the same audit, which is a parent. The AI-Q anchor contributes a before-and-after on one open pipeline measured on citation recall and precision with standard errors of 2 to 4 points; its precision follows an unprinted LongCite prompt and its judge is gpt-5-mini without a human check of these scores, so the step counts it as a measured condition and not as a validated one. The step adds nothing but the sorting of the four conditions by the component they were measured on. ### Breaking point A study applying citation training or a URL-checking tool to a web-research assistant and measuring per-citation support before and after; or a human re-scoring of the AI-Q intervention that removes its precision gain. Query: papers citing LongCite or urlhealth that report claim-level support on web tasks; a human audit of the Table 3 scores of arXiv 2608.24306. ### Reflex There are known fixes: verify the links, train for citations. Three of the four conditions are measured on link resolution and on closed documents, not on whether web citations support claims; the one measured on web citation precision is a single open pipeline under an LLM judge. ### Notes Supersedes 'Three conditions are each measured on one component of citation quality and none on claim support in web research' after the completeness attack of attacker run 3 (attacker-grok-4.7) delivered the AI-Q intervention anchor, adopted in proposal 003. The old card is kept with status broken. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--human-raters-find-citation-precision-of-89-and-84-percent-for-longcite-models-against-68-for-glm-4-on-longbench-chat - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--citation-trained-longcite-8b-reaches-citation-f1-of-72-on-longbench-cite-against-65-to-67-for-three-proprietary-models - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--with-a-url-checking-tool-in-the-loop-three-models-cut-non-resolving-citation-urls-6-to-79-fold-to-under-1-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--two-fixes-to-the-open-ai-q-pipeline-raise-citation-precision-from-876-to-910-and-941-and-recall-from-645-to-697-on-50-deepresearch-bench-queries - SUPERSEDES → https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research ## High faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume type: derivation · status: open · confidence: medium ### Conclusion On ResearcherBench OpenAI Deep Research has 0.84 faithfulness and 0.34 groundedness, Grok3 DeeperSearch 0.80 and 0.31, while Sonar Reasoning Pro has the lowest faithfulness (0.62) and the highest groundedness (0.68); on DeepResearch Bench the number of supported citations per task runs from 4.35 to 111.21. A support rate is computed over the claims a system chose to cite and says nothing about the claims it left uncited or about how many citations it gives. ### Step The faithfulness anchor contributes the rate among cited claims, the groundedness anchor the share of claims that are cited at all, in the same table for the same systems, which is what allows the statement that they move independently. The volume anchor contributes that accuracy and count separate as well: 94.04 accuracy with 9.78 effective citations for one system, 81.44 with 111.21 for another. The ratio 111.21 to 4.35 is computed here, not printed in the paper. ### Breaking point Evidence that uncited claims in these reports are common knowledge needing no citation, which would make low groundedness harmless; or a benchmark where groundedness and faithfulness rise together across systems. Query: human classification of uncited claims in ResearcherBench outputs as citation-worthy or not. ### Reflex A tool with 85 percent citation accuracy is 85 percent reliable. The 85 percent covers only the third to two thirds of claims that carry a citation. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--groundedness-of-31-to-68-percent-on-researcherbench-means-32-to-69-percent-of-factual-claims-carry-no-citation - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--effective-citations-per-task-range-from-about-4-to-111-across-systems-on-deepresearch-bench - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--high-faithfulness-of-cited-claims-coexists-with-most-claims-carrying-no-citation-and-with-a-25-fold-spread-in-citation-volume ## In retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check type: derivation · status: superseded · confidence: medium ### Conclusion For search-augmented and deep-research systems, 3.0 to 13.3 percent of cited URLs are classified as hallucinated and 5.4 to 18.5 percent do not resolve, while in a per-citation benchmark of the same system class 23.2 to 75.6 percent of citations fail the support check (Fact Check 24.4 to 76.8 percent). The larger failure in retrieval-backed systems is an existing page that does not support the claim. ### Step The URL study contributes the existence failure with bootstrap intervals and no rater in the loop. The 14-agent benchmark contributes the support failure as 100 minus its printed Fact Check range. The two studies use different systems and queries; the step compares orders of magnitude, not points, and rests on the premise that both sample the same class of products (2025 and 2026 assistants with web search). ### Breaking point A study on one sample of citations finds that most citations failing the support check also fail to resolve, or a current system whose hallucinated-URL rate exceeds its support-failure rate. Query: join of link status and support verdict per citation in the release of arXiv 2605.06635. ### Reflex The problem with AI citations is invented sources. For assistants that search, the measured problem is mostly real sources that do not say what is claimed. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--across-10-models-on-drbench-3-to-13-percent-of-citation-urls-are-hallucinated-and-5-to-18-percent-do-not-resolve - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - SUPERSEDES ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-support-on-the-same-citations-dead-links-explain-at-most-a-sixth-of-the-support-failures (Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures) Address: https://trustillery.com/entity/stocks/ai-citations--in-retrieval-backed-systems-non-existent-links-are-the-smaller-failure-at-3-to-13-percent-against-23-to-76-percent-of-citations-failing-the-support-check ## Non-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool type: derivation · status: open · confidence: medium ### Conclusion Two audits of the published literature find invalid or non-existent citations in 1.07 percent of 56,381 conference papers (1.61 percent in 2025, 80.9 percent above the 2020 to 2024 average) and an excess of unmatched references of 0.21 to 1.91 percent over the pre-LLM baseline in four preprint and journal corpora by August 2025. Both measure existence in papers written by researchers and neither establishes which tool, if any, produced a given reference. ### Step The conference audit contributes a hand-confirmed count and the year-over-year change; the four-corpus audit contributes the excess over a pre-LLM baseline at far larger scale by automated matching. Together they show a downstream trace in 2025 in two samples: a rise of the per-paper share over 2020 to 2024 in one, an excess of unmatched references over a pre-LLM baseline as of August 2025 in the other; the units differ (papers, references) and the samples may overlap through arXiv. The step keeps the authors' own disclaimers: no causal attribution to assistants. ### Breaking point The join fails while both parents stand if the two audits are not independent readings of one trend: if the conference papers are largely the same documents as the arXiv corpus of the second audit, if both rest on the same matching databases so that one indexing gap produces both signals, or if the per-paper share and the per-reference excess turn out to move apart when computed on one corpus. Query: overlap of the 56,381 conference papers with the arXiv corpus, and both metrics computed on that overlap. ### Reflex Hallucinated citations are flooding science. The measured level is about one paper in a hundred with at least one invalid citation, rising, with no attribution to a tool. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--audit-of-111-million-references-in-arxiv-biorxiv-ssrn-and-pmc-papers-estimates-146932-non-existent-citations-in-2025 - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--invalid-citations-appear-in-about-1-percent-of-56381-ai-and-security-conference-papers-and-rose-81-percent-in-2025 Address: https://trustillery.com/entity/stocks/ai-citations--non-existent-citations-reach-the-published-record-at-about-1-percent-of-papers-and-rising-without-identifying-the-tool ## Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023 type: derivation · status: open · confidence: medium ### Conclusion The human audit of early 2023 found 74.5 percent citation precision on average (63.6 to 89.5 by engine) and the automatic ALCE baseline about 50 percent; the LLM-judged audit of generative search engines in August 2025 found 39.8 to 68.3 percent citation accuracy. The only repeated design, the BBC rounds of December 2024 and mid 2025, shows significant sourcing issues falling to 10 to 15 percent for three assistants and unchanged near 47 percent for Gemini. Whether support improved between 2023 and 2026 cannot be read from the cross-study numbers, and the one within-design comparison shows improvement for three of four assistants on news questions. ### Step Liu et al. contribute the human-rated baseline and its spread; ALCE the automatic baseline on research pipelines; DeepTRACE the 2025 reading on a partly overlapping set of products (Perplexity, Bing/Copilot, You.com) with a different rater and query set. These three differ in rater, queries and systems, so the step declines to compute a trend from them. The BBC comparison contributes the only pair with one design, with the report's own caveats (product tiers changed from paid to free, definitions adjusted, samples of 362 and 237). ### Breaking point A replication of the 2023 human audit protocol on 2025 or 2026 engines; it would replace the cross-study comparison with a trend. Query: any paper citing arXiv 2304.09848 that reuses its annotation protocol on current systems. ### Reflex The models have got much better, so citation problems are a 2023 story. No published measurement holds the method fixed across those years except one news audit over six months. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--in-early-2023-four-generative-search-engines-had-515-percent-citation-recall-and-745-percent-citation-precision - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--citation-precision-ranged-from-636-to-895-percent-across-four-2023-generative-search-engines-and-fell-as-utility-rose - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--on-eli5-in-2023-chatgpt-and-gpt-4-baselines-reach-about-50-percent-automatic-citation-recall-and-precision - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--four-generative-search-engines-reach-40-to-68-percent-citation-accuracy-and-leave-23-to-47-percent-of-statements-unsupported - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--between-two-bbc-rounds-significant-sourcing-issues-fell-to-10-to-15-percent-for-three-assistants-while-gemini-stayed-near-47 Address: https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 ## Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support type: derivation · status: open · confidence: high ### Conclusion The Tow Center error rates (more than 60 percent of 1,600 queries, 37 to 94 percent by tool; 153 of 200 for ChatGPT Search) are for naming the source of a given excerpt, the Grok 3 figure of 154 of 200 is link resolution, and the DRACO citation-quality score of 42.1 to 64.6 percent rates whether the rubric's primary documents are cited. None of the four is a share of citations that support the attached claim. ### Step Each parent contributes its own task definition, stated in its Statement: reverse attribution on publisher, date and URL (two anchors), error pages (one), primary-source rubric criteria graded by an LLM judge and published by the vendor of the top system (one). The step only sorts them out of the support range so that they are not averaged into it. ### Breaking point A re-reading of one of the three documents (two CJR articles, DRACO) shows that it did check generated claims against cited passages. Query: methodology sections of the two CJR articles and DRACO Section on rubric axes. ### Reflex AI search engines are wrong 60 percent of the time when they cite. The 60 percent is for a source-identification task, not for the citations in ordinary answers. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--eight-ai-search-engines-answered-over-60-percent-of-1600-source-identification-queries-incorrectly-ranging-37-to-94-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--chatgpt-search-gave-partially-or-entirely-incorrect-source-attributions-for-153-of-200-publisher-quotes-in-november-2024 - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--grok-3-cited-urls-leading-to-error-pages-in-154-of-200-source-identification-prompts-in-a-tow-center-test-of-february-2025 - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--vendor-run-draco-scores-citation-quality-at-65-percent-for-perplexity-deep-research-and-42-to-56-for-five-rivals Address: https://trustillery.com/entity/stocks/ai-citations--reverse-attribution-tests-and-the-vendor-primary-source-axis-measure-neighbouring-quantities-and-not-citation-support ## Support falls as agent runs get longer and most traced errors arise in orchestration and not in search type: derivation · status: open · confidence: medium ### Conclusion Per-citation support fell from 78.6 to 16.7 percent (GPT-5.4) and from 80.0 to 57.9 percent (Claude Opus 4.6) as the tool-call budget rose from 2 to 150, and in three open multi-agent pipelines the orchestrating step was the origin of 52.6 to 100 percent of traced final-report errors while the searcher accounted for 0.4 percent in the one system where it is printed. Both point to synthesis over many sources, not retrieval, as the place where support is lost. ### Step The ablation contributes the dose-response: more tool calls, lower support, link validity and relevance above 92 percent at every depth. The localisation study contributes where errors are introduced, in different (open-source) systems. The step joins two systems classes under the premise that commercial agents fail in the same place as open pipelines; that premise is not measured in either parent. The GPT-5.4 series is not monotone, so the conclusion is about the end points. ### Breaking point An error localisation on a commercial deep-research product that places most unsupported citations at retrieval; or a depth ablation on further models that shows no decline. Query: repeat of the tool-call ablation of arXiv 2605.06635 on a third and fourth model. ### Reflex More searching means better-grounded answers. In the one ablation that varied it, support fell with search depth. ### Notes Attacker run 3 (attacker-grok-4.7, x-attack-step): refused on the main point, the joining premise is declared in the step and named in the breaking point, not hidden; 'unchanged' corrected in place to 'above 92 percent at every depth'. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--orchestrator-originates-85-percent-of-final-report-errors-in-ai-q-53-percent-in-ms-agent-and-100-in-trajectorykit Address: https://trustillery.com/entity/stocks/ai-citations--support-falls-as-agent-runs-get-longer-and-most-traced-errors-arise-in-orchestration-and-not-in-search ## The 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate type: derivation · status: open · confidence: high ### Conclusion The 31 percent of news responses with significant sourcing issues (Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15 percent) is a share of responses rated by journalists on a criterion that covers unsupported claims, missing sources and incorrect sourcing claims together; it is not the share of citations that fail to support their claim, and for Gemini part of it is absent sourcing (42 percent of its responses gave no direct source in 2025, 26 percent gave none in December 2024). ### Step The EBU anchor contributes the definition of the category and the per-assistant spread; the BBC anchor contributes that the same criterion in the earlier round already mixed misattribution, unsupported claims and missing sources, and that missing sources were named as part of the cause for the assistant with the highest rate. The step narrows what the 31 percent can be cited for; it does not say the figure is wrong. ### Breaking point The EBU appendix or data release separates the three sub-categories and shows that unsupported claims alone account for nearly all of the 31 percent. Query: Q2 sub-codes in the EBU/BBC toolkit data. ### Reflex A third of AI news answers cite sources that do not back them up. Too coarse: the 31 percent also counts answers with no source at all. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--journalists-at-22-public-media-rated-31-percent-of-ai-assistant-news-responses-as-having-significant-sourcing-issues - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--bbc-journalists-rated-over-45-percent-of-gemini-news-responses-as-significant-sourcing-errors-and-26-percent-gave-no-sources Address: https://trustillery.com/entity/stocks/ai-citations--the-31-percent-sourcing-figure-of-the-ebu-audit-counts-responses-and-includes-absent-sources-and-is-not-a-per-citation-support-rate ## The direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge type: derivation · status: open · confidence: medium ### Conclusion Where judge decisions on citation support were compared with human labels, two measurements find the LLM judge stricter than humans (false negative rates of 0.183 to 0.470 on a human-reviewed benchmark; GPT-4o judge precision of 79.7, 78.1 and 53.9 against human 88.9, 84.2 and 67.5 on the same responses) and one finds it more generous (both disagreements in 100 validated AI Overview verdicts were the verifier accepting a claim the humans did not). ### Step The judge benchmark contributes the size of over-rejection on an adversarial report in which 81.6 percent of pairs are unsupported by construction. The LongCite human evaluation contributes the same direction on natural model output, for a document-grounded task. The AI Overviews validation contributes the opposite direction on natural web output, on two cases. The step does not average them: tasks, judges and base rates differ, and two cases are not a rate. It follows only that the LLM-judged support rates in this stock cannot be corrected in a known direction. ### Breaking point A validation of one of the stock's judges on natural web-research output with a few hundred human labels that shows a consistent direction of error. Query: human re-labelling of released judge decisions from DeepResearch Bench or arXiv 2605.06635. ### Reflex LLM judges are lenient, so real support is lower than reported. Not licensed: two of three comparisons in the stock show the judge stricter than the humans. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--llm-citation-judges-reject-18-to-47-percent-of-genuinely-supported-citations-on-a-human-reviewed-benchmark - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--human-raters-find-citation-precision-of-89-and-84-percent-for-longcite-models-against-68-for-glm-4-on-longbench-chat - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--89-percent-of-98020-atomic-claims-in-google-ai-overviews-are-supported-by-the-cited-pages-and-11-percent-are-not Address: https://trustillery.com/entity/stocks/ai-citations--the-direction-of-llm-judge-error-on-citation-support-is-unsettled-with-two-measurements-showing-stricter-judges-and-one-a-more-generous-judge ## Three conditions are each measured on one component of citation quality and none on claim support in web research type: derivation · status: superseded · confidence: medium ### Conclusion Citation-trained models score 88.9 and 84.2 percent human-rated citation precision on supplied documents against 67.5 for GLM-4, a model without that training; a URL-checking tool in the loop cuts non-resolving links to under 1 percent as the paper states it; references named by three or more models match a scholarly database at 95.6 percent against 16.5 for references named by one model. Each holds for its own component (grounding in a given document, link resolution, existence) and none of the three studies measured whether citations in web research support their claims afterwards. ### Step The two LongCite anchors contribute the comparison of citation-trained and other models, a comparison across different models and not one model before and after training, once LLM-judged across five datasets and once human-rated on one, in a closed-document setting. The URL tool anchor contributes the link component as a before-and-after measurement on the same questions, the only causal reading among the three, and states that support of the replacement links was not measured. The consensus anchor contributes the existence component for references recalled without retrieval, a status stated not on that anchor but on the sibling anchor of the same audit (ten commercial LLMs, 69,557 citations, run without retrieval), which is therefore a parent. The step adds nothing but the observation that the three conditions do not overlap with the quantity the question asks about. ### Breaking point A study applying one of these interventions to a web-research assistant and measuring per-citation support before and after. Query: papers citing LongCite or urlhealth that report claim-level support on web tasks. ### Reflex There are known fixes: verify the links, train for citations. The conditions are measured on link resolution and on closed documents, not on whether web citations support claims. ### Notes Attacker run 3 (attacker-grok-4.7, x-attack-step): applied, the retrieval status was taken from a card that was not a parent; the sibling anchor of the same audit is now a parent and the step names it. Conclusion unchanged. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--human-raters-find-citation-precision-of-89-and-84-percent-for-longcite-models-against-68-for-glm-4-on-longbench-chat - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--citation-trained-longcite-8b-reaches-citation-f1-of-72-on-longbench-cite-against-65-to-67-for-three-proprietary-models - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--with-a-url-checking-tool-in-the-loop-three-models-cut-non-resolving-citation-urls-6-to-79-fold-to-under-1-percent - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations - SUPERSEDES ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) Address: https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research ## Two journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses type: derivation · status: open · confidence: medium ### Conclusion In December 2024 BBC journalists found eight altered or absent BBC quotes across 62 responses containing BBC quotes (13 percent as the report computes it), and in mid 2025 journalists at 22 organisations rated 12 percent of 1,053 quote-bearing responses as having significant quote-accuracy issues. For direct quotes attributed to a cited source, the two human-rated rounds agree on roughly one response in eight. ### Step The BBC anchor contributes a small sample checked against the broadcaster's own articles by people who in part wrote them; the EBU anchor contributes a sample seventeen times larger. The BBC took part in both audits, so they are not independent. The two are not the same quantity in the strict sense: the BBC figure divides quotes by responses, the EBU figure is a rating per response. The step treats both as response-level and says so; agreement at 12 to 13 percent across rounds is recorded, not explained. ### Breaking point The EBU data split by assistant shows the 12 percent is carried by one assistant (Gemini at 20 percent) so that the typical rate is far lower; or a per-quote count shows the BBC's 13 percent changes materially with the right denominator. Query: per-assistant and per-quote counts in both reports' appendices. ### Reflex A direct quote with a link is safe to reuse. In both audits about one in eight responses with quotes had a quote that was altered or not in the cited piece. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--of-1053-ai-assistant-news-responses-with-direct-quotes-12-percent-had-significant-quote-accuracy-issues-per-journalists - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--bbc-journalists-found-8-altered-or-absent-bbc-quotes-across-62-ai-assistant-responses-in-december-2024-or-13-percent - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--two-journalist-audits-find-a-quote-accuracy-problem-in-about-one-in-eight-quote-bearing-news-responses ## Two support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation type: derivation · status: open · confidence: high ### Conclusion The 89.0 percent for Google AI Overviews checks each claim against all pages the Overview cites, and SourceCheckup's 75.7 percent counts a statement as supported if at least one source given in the same response supports it. Both are claim-level rates of support somewhere in the reference set: more lenient than checking the attached citation, and an upper bound on the share of claims supported by their own attached citation. They are not bounds on the share of citations that support their sentence, which has a different denominator. ### Step Each parent states its unit in its Statement. The step is set logic over claims: if the attached citation supports the claim, then some cited page does, but not the reverse. It says nothing about a rate counted over citations: a claim with three supporting citations and a claim with one non-supporting citation give 50 percent of claims and 75 percent of citations. It does not estimate how far the attached-citation rate lies below; neither paper prints it. ### Breaking point A per-link audit on either dataset finds the share of claims supported by their attached citation equal to the any-source rate, which would make the bound tight. Query: per-link labels in the AI Overviews release. ### Reflex Google's AI Overviews get their citations right 89 percent of the time. The 89 percent is claims supported by any of the cited pages; how many links support their own sentence was not counted. Relationships - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--89-percent-of-98020-atomic-claims-in-google-ai-overviews-are-supported-by-the-cited-pages-and-11-percent-are-not - FOLLOWS_FROM → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--two-support-rates-in-the-stock-count-a-claim-as-supported-if-any-cited-page-supports-it-and-bound-the-share-of-claims-supported-by-their-attached-citation ## 89 percent of 98020 atomic claims in Google AI Overviews are supported by the cited pages and 11 percent are not type: anchor · status: open · kind: measurement · quoted: Of the 98,020 verified claims, 87,204 (88.97%) are consistent · collected_by: Xu, Iqbal, Montgomery (Washington University in St. Louis) · independence: independent · source_family: self-published · as_of: 2026-05-13 ### Statement Across 98,020 atomic claims from 7,491 verifiable Google AI Overviews, each checked against the text of all pages that Overview cites, Table 1 prints: Clear 84.6% (82,933), Vague 4.4% (4,271), Ambiguous 1.4% (1,367), Incorrect 2.7% (2,609), Omitted 7.0% (6,840); Consistent (Clear plus Vague) 89.0%, Inconsistent 11.0%. Omitted means no cited source mentions the claim; Incorrect means a cited source contradicts it. The unit is the claim against the Overview's whole reference set, not a single citation against a single sentence. Labels are assigned by an LLM verifier (Grok 4.1 Fast Reasoning) and were validated against two human annotators on 100 verdicts (98 of 100 matched). ### Collection Authors are at Washington University in St. Louis; they are not affiliated with Google in the paper; arXiv preprint without venue. Method: 55,393 trending queries in 19 topical categories were issued over a 40-day window (March 13 to April 21, 2026); each AI Overview was decomposed into atomic claims and each claim verified against the full extracted body text of every reference the Overview cites, by an LLM pipeline (Grok 4.1 Fast Reasoning, temperature 0) assigning one of five labels. Pages were crawled hours or days after the Overview was generated. The counter-check that exists and was used: two annotators re-labelled a stratified sample of 100 claim-level verdicts (20 per label), inter-annotator Cohen's kappa 0.94, and the verifier matched the adjudicated human label on 98 of 100; both errors were the verifier being too generous with supported claims. Claim extraction was validated separately on 100 Overviews. ### Falls when A per-citation audit (does the specific linked page support the specific sentence it is attached to) on the same Overviews gives a supported share well below 89.0%, showing the any-cited-page unit inflates support; or a larger human validation than 100 verdicts shows the verifier's generosity toward Clear and Vague is frequent enough to move the 11.0%. Query to run: human per-link support audit on a sample of the 7,491 Overviews. ### Reflex AI answers in search mostly make claims their sources do not contain. Too coarse: in this measurement of one product, 84.6 percent of claims were explicitly supported by a cited page and 2.7 percent were contradicted by one. ### Evidence https://arxiv.org/pdf/2605.14021v1 p. 6, Table 1; p. 7, Section 3.3.3; p. 11, Section 4.3 | 2026-05-13 · arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews Plan 02b: quoted sentence is Section 4.3, p. 11; Table 1, p. 6 carries the split. Relationships - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-direction-of-llm-judge-error-on-citation-support-is-unsettled-with-two-measurements-showing-stricter-judges-and-one-a-more-generous-judge (The direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--two-support-rates-in-the-stock-count-a-claim-as-supported-if-any-cited-page-supports-it-and-bound-the-share-of-claims-supported-by-their-attached-citation (Two support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions (Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions) Address: https://trustillery.com/entity/stocks/ai-citations--89-percent-of-98020-atomic-claims-in-google-ai-overviews-are-supported-by-the-cited-pages-and-11-percent-are-not ## Across 10 models on DRBench 3 to 13 percent of citation URLs are hallucinated and 5 to 18 percent do not resolve type: anchor · status: open · kind: measurement · quoted: Non-resolving URL rates across models range from 5.4% [3.0, 7.7] (gpt-4.1) to 18.5% [17.8, 19.2] ( gemini-2.5-pro-deepresearch), while hallucinated URL rates range from 3.0% [1.7, 4.4] ( claude-3-5-sonnet-with-search) to 13.3% [12.7, 13.9] (gemini-2.5-pro-deepresearch). · collected_by: Rao, Wong, Callison-Burch (University of Pennsylvania) · independence: independent · source_family: self-published · as_of: 2026-04-03 ### Statement Link validity only, not support: the study checks whether cited URLs exist, which it calls a logically prior question to whether the source supports the claim. For 10 models from Google, OpenAI and Anthropic on DRBench (100 research queries in Chinese and English, outputs pre-collected by the benchmark authors, 296 to 11,309 URLs per model), the share of citation URLs that do not resolve (HTTP 4xx or 5xx, connection error or timeout, HTTP 403 excluded) ranges from 5.4% [3.0, 7.7] (gpt-4.1) to 18.5% [17.8, 19.2] (gemini-2.5-pro-deepresearch); the share classified as hallucinated (non-resolving and no Wayback Machine snapshot at any time) ranges from 3.0% [1.7, 4.4] (claude-3-5-sonnet-with-search) to 13.3% [12.7, 13.9] (gemini-2.5-pro-deepresearch). Pooled, the two deep-research agents have 10.7% [10.2, 11.2] hallucinated and 16.2% [15.7, 16.8] non-resolving URLs against 4.8% [4.3, 5.2] and 6.8% [6.2, 7.3] for the eight search-augmented models. On ExpertQA (2,177 expert questions, 32 fields, 168,021 URLs from claude-sonnet-4-5, gemini-2.5-pro and gpt-5.1) the overall non-resolving rate is 8.22% [8.09, 8.36], by field from 5.4% [4.9, 5.9] (Business) to 11.4% [8.1, 14.6] (Theology). Brackets are bootstrap 95% intervals. The measurement is automated (HTTP requests plus Wayback Machine API); there is no human or LLM rater. ### Collection Authors are at the University of Pennsylvania; funded by DARPA's SciFy program; arXiv preprint marked as under review, not peer reviewed. No affiliation with an evaluated vendor appears in the document. Each URL gets an HTTP HEAD request (GET fallback) with a browser-like User-Agent; non-resolving URLs are looked up in the Wayback Machine, no snapshot means hallucinated, a snapshot means stale. The authors call the rates conservative lower bounds: 403 responses (6.6 to 17.0% of ExpertQA URLs) are excluded and Wayback coverage is incomplete. Counter-checks used: a headless-browser audit (403 responses 99.7% live; 89% of UNKNOWN responses live or blocked) and a sensitivity analysis for Reddit URLs (treating all as non-resolving raises gpt-5.1 from 8.47% to 26.7%). 13 of 23 DRBench models were excluded, three of them for 100% hallucination rates read as no real web retrieval. The abstract gives 53,090 DRBench URLs while the ten per-model counts in Table 1 sum to 23,269; the text does not explain the difference. Liveness is a point-in-time measurement. Tool and data are announced under MIT license. ### Falls when A re-check of the released URL lists with a real browser instead of HEAD requests finds most non-resolving URLs live (bot-blocking rather than dead pages), pushing the upper rates well below 13.3 and 18.5%; or a wider archive lookup shows most URLs classed as hallucinated had existed. The range narrows if the single outlier gemini-2.5-pro-deepresearch is set aside (next highest 8.8% hallucinated, 10.1% non-resolving). The anchor is out of scope for any claim about whether resolving links support their sentences. Query to run: headless-browser liveness plus Wayback and Common Crawl lookup over the released DRBench URL set. ### Reflex Models with web search no longer invent links; fake references are a problem of offline chatbots. Too coarse: with search switched on, 3 to 13 percent of cited URLs still had no trace of ever having existed, and pooled deep-research agents did worse than plain search-augmented models. ### Evidence https://arxiv.org/pdf/2604.03173v1 p. 4, Table 2 and Section 4.1; p. 5, Sections 4.1 to 4.3; p. 2-3, Sections 2 and 3.3 for definitions | 2026-04-03 · arXiv 2604.03173 · Rao, Wong, Callison-Burch, Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--in-retrieval-backed-systems-non-existent-links-are-the-smaller-failure-at-3-to-13-percent-against-23-to-76-percent-of-citations-failing-the-support-check (In retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-support-on-the-same-citations-dead-links-explain-at-most-a-sixth-of-the-support-failures (Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures) Address: https://trustillery.com/entity/stocks/ai-citations--across-10-models-on-drbench-3-to-13-percent-of-citation-urls-are-hallucinated-and-5-to-18-percent-do-not-resolve ## ALCE automatic citation metrics agree with human raters at kappa 0.698 for recall and 0.525 for precision type: anchor · status: open · kind: measurement · quoted: kappa coefficient between human and ALCE suggests substantial agreement for citation recall (0.698) and moderate agreement for citation precision (0.525) · collected_by: Gao, Yen, Yu, Chen (Department of Computer Science and Princeton Language and Intelligence, Princeton University) · independence: positioned · source_family: peer-review · as_of: 2023-05-24 ### Statement The ALCE paper validates its automatic, NLI-based citation recall and precision against human judgement on outputs of three systems built on 2023 models (ChatGPT VANILLA, ChatGPT RERANK, Vicuna-13B VANILLA). Human raters judged, per sentence, whether all cited passages together fully support it, and per citation, whether it fully, partially or does not support the sentence. Cohen's kappa between human and automatic labels is printed as 0.698 for citation recall and 0.525 for citation precision; treating human annotations as gold labels, the automatic metric has an accuracy of 85.1% for citation recall and 77.6% for citation precision. For detecting irrelevant citations it has a recall of 75.6% and a precision of 66.1%, which the paper attributes to the NLI model being unable to detect partial support. On ELI5 the system-level scores are, human against ALCE, 50.8 / 52.4 against 52.8 / 50.4 for ChatGPT VANILLA, 59.7 / 60.6 against 63.0 / 60.6 with RERANK, and 13.4 / 19.2 against 13.6 / 18.1 for Vicuna-13B (recall / precision, Table 9). This is a 2023 baseline for how far automatic support metrics track human raters. ### Collection Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen, Department of Computer Science and Princeton Language and Intelligence, Princeton University; published at EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 31 Oct 2023). The validation is human-rated: workers employed through Surge AI at an average pay of 20 USD per hour; the paper says it randomly sampled 100 examples from ASQA and ELI5 and annotated the outputs of the three selected models, without stating the number of raters or an inter-rater agreement figure. The automatic side is the TRUE NLI model (T5-11B). The authors validate a metric they propose themselves, so they have a stake in the outcome; that is why independence is recorded as positioned, although they are not vendors of any evaluated model. The paper thanks Surge AI for support with the human evaluation. A counter-check by a party other than the metric's authors is not part of the paper. ### Falls when An independent human annotation of ALCE outputs yields a Cohen's kappa with the automatic labels clearly below 0.698 for recall or 0.525 for precision, or accuracy below 85.1% and 77.6%, or the agreement does not hold once the systems are 2024 to 2026 models whose statements synthesise several passages. Query to run: human versus TRUE-NLI agreement on citation support for outputs of current models on the ALCE questions. ### Reflex Automatic entailment checks are a good enough stand-in for human judgement of whether a citation supports a claim. Too coarse: agreement was substantial for whether a statement is supported by all its citations together and only moderate for whether a single citation supports it, the quantity this stock asks about. ### Evidence https://arxiv.org/pdf/2305.14627v2 p. 8-9, Section 6 and Tables 8-9, with Appendix F, p. 15-16, and Appendix G.5, p. 17 | 2023-05-24 · EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions (Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions) Address: https://trustillery.com/entity/stocks/ai-citations--alce-automatic-citation-metrics-agree-with-human-raters-at-kappa-0698-for-recall-and-0525-for-precision ## Audit of 111 million references in arXiv, bioRxiv, SSRN and PMC papers estimates 146,932 non-existent citations in 2025 type: anchor · status: open · kind: measurement · quoted: Assuming these trends persisted through year end, these four corpora, which cover only a fraction of the scientific literature, would include 146,932 hallucinated citations in 2025 alone. · collected_by: Zhao, Wang, Stuart, De Vaan, Ginsparg, Yin (Cornell University; UCLA; Tsinghua University; UC Berkeley Haas School of Business) · independence: independent · source_family: self-published · as_of: 2026-05-08 ### Statement Downstream context, not a measurement of any assistant: an audit of 111 million references in 2.5 million papers (arXiv Jan 2020 to Aug 2025, bioRxiv, SSRN, and a 10% sample of PubMed Central) checks whether each cited title exists in Semantic Scholar, OpenAlex or Google Scholar. The excess of unmatched references over the pre-LLM baseline reached 0.39% (arXiv), 0.21% (bioRxiv), 1.91% (SSRN) and 0.27% (PMC) of references as of August 2025, with monthly excess counts of 3,353, 478, 767 and 8,140; extrapolating these to year end gives 146,932 non-existent citations in the four corpora in 2025. The quantity is existence of the cited work, measured by an automated matching pipeline; it says nothing about whether an existing cited work supports the claim, and it does not identify which tool, if any, produced a reference. 'Hallucinated' in the paper names the estimated excess, not a classification of individual references. ### Collection Academic authors at Cornell, UCLA, Tsinghua and UC Berkeley; arXiv preprint, not peer reviewed. Method: references parsed from LaTeX, GROBID, platform XML or Crossref metadata; titles matched by string similarity against a local Elasticsearch index of Semantic Scholar and OpenAlex (95.1% matched), then a GPT-4o-mini step excludes non-academic strings (unmatched 2.33%), re-extraction (1.54%), then a Google Scholar lookup. The estimate is a regression excess over the pre-2023 unmatched rate, so attribution to LLM use is by timing and correlates (fields with high AI uptake, linguistic signatures of AI-assisted writing), not by observation of tool use. The annual figure assumes August 2025 monthly levels persist through December. The authors call it a lower bound because only title existence is tested. Counter-check that exists: manual validation of unmatched cases and sensitivity tests, reported in the Supplementary Information, which is not part of the observed text. ### Falls when Manual verification of a random sample of post-2023 unmatched references finds that most exist (indexing lag, non-English or grey literature), or the pre-LLM baseline re-estimated with the same pipeline on later database snapshots rises to the post-2023 level. Query to run: hand-check of sampled unmatched references per corpus and year against publisher records. The anchor narrows, without falling, if the full-year 2025 data replace the August extrapolation. ### Reflex AI fabricates references, and papers are now full of fake citations. Too coarse: the measured excess is 0.2 to 1.9 percent of references depending on corpus, spread thinly over many papers, and it concerns non-existent works in published manuscripts, a different quantity from whether an assistant's citation supports its claim. ### Evidence https://arxiv.org/pdf/2605.07723v1 p. 3-5, Estimating hallucinated references at scale and Results | 2026-05-08 · arXiv 2605.07723 · Zhao et al., LLM hallucinations in the wild Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--non-existent-citations-reach-the-published-record-at-about-1-percent-of-papers-and-rising-without-identifying-the-tool (Non-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool) Address: https://trustillery.com/entity/stocks/ai-citations--audit-of-111-million-references-in-arxiv-biorxiv-ssrn-and-pmc-papers-estimates-146932-non-existent-citations-in-2025 ## BBC journalists found 8 altered or absent BBC quotes across 62 AI assistant responses in December 2024 or 13 percent type: anchor · status: open · kind: measurement · quoted: Eight quotes sourced from BBC articles were either altered from the original source or not present in the cited article. 62 responses in total included BBC quotes meaning an error rate of 13%. · collected_by: Oli Elliott, Principal Data Scientist, BBC Responsible AI Team; ratings by 45 BBC News journalists · independence: positioned · source_family: established-media · as_of: 2025-02-11 ### Statement In December 2024 the BBC put 100 news questions to ChatGPT (Enterprise, GPT-4o), Copilot Pro, Gemini Standard and Perplexity Pro; 45 BBC journalists reviewed 362 responses. 62 responses included quotes from BBC articles; eight quotes were either altered from the original source or not present in the cited article, which the report expresses as an error rate of 13%. The eight occurred in responses from all assistants tested except ChatGPT. The report divides a count of quotes by a count of responses; the total number of quotes checked is not printed. Verification was a human comparison with the BBC's own articles. ### Collection Designed and carried out by the BBC's Responsible AI team. Responses to 100 news questions drawn from trending Google search topics were collected on 5 and 6 December 2024 with the prefix 'Use BBC News sources where possible'; the BBC lifted its crawler blocks for the duration. 45 BBC News journalists, assigned by area of expertise and in many cases the authors of the cited articles, reviewed 362 responses in randomised order with assistant names removed. The BBC is the publisher whose content is represented, states that publishers should have control over the use of their content and calls for regulation: a stake and a declared position, recorded as positioned. Counter-check: the cited BBC articles are public, so each quote can be compared by anyone, but the report publishes only example responses in its appendix, not the 62 quote-bearing responses or the eight cases in full. A small inter-rater agreement test using Krippendorff's Alpha showed moderate agreement; it was not specific to quotes. ### Falls when Falls or narrows if the full set of BBC-attributed quotes in the 62 responses is compared with the cited BBC articles, including their revision history since BBC articles are updated after publication, and the altered-or-absent count differs from eight. Narrows if the rate is recomputed per quote rather than per response. Query: number of distinct BBC-attributed quotes in the 362 responses, and how many match the article version live on 5 and 6 December 2024. ### Reflex When an assistant puts words in quotation marks and cites the article, the words are in the article. Too coarse: in this sample about one in eight quote-bearing responses carried a quote that was altered or not in the cited article. ### Evidence https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 7, Accuracy | 2025-02-11 · BBC report · Elliott, Representation of BBC News content in AI Assistants Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--two-journalist-audits-find-a-quote-accuracy-problem-in-about-one-in-eight-quote-bearing-news-responses (Two journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses) Address: https://trustillery.com/entity/stocks/ai-citations--bbc-journalists-found-8-altered-or-absent-bbc-quotes-across-62-ai-assistant-responses-in-december-2024-or-13-percent ## BBC journalists rated over 45 percent of Gemini news responses as significant sourcing errors and 26 percent gave no sources type: anchor · status: open · kind: measurement · quoted: Overall, Gemini produced the most sourcing errors – reviewers rated over 45% of responses as containing significant sourcing errors. Lack of sources were part of the cause. 26% of Gemini responses and 7% of ChatGPT’s provided no sources at all · collected_by: Oli Elliott, Principal Data Scientist, BBC Responsible AI Team; ratings by 45 BBC News journalists · independence: positioned · source_family: established-media · as_of: 2025-02-11 ### Statement In the BBC's December 2024 study (100 news questions to ChatGPT Enterprise, Copilot Pro, Gemini Standard and Perplexity Pro; 362 responses reviewed by 45 BBC journalists), reviewers rated each response on Q2: 'Are the claims in the response supported by its sources, with no problems with attribution (where relevant)?'. Over 45% of Gemini responses were rated as containing significant sourcing errors, the most of the four assistants. 26% of Gemini responses and 7% of ChatGPT's provided no sources at all. For the other assistants the text gives no percentages; the appendix table prints the counts of 'Significant Issues' ratings on Q2 as ChatGPT 19, Copilot 23, Gemini 30, Perplexity 15. Gemini refused 12 of the 100 questions. The criterion mixes claims not supported by the cited source, misattribution to the BBC, and missing sources; the unit is the response, not the citation; the rating is a human judgement on a four-level scale. ### Collection Designed and carried out by the BBC's Responsible AI team. Responses to 100 news questions drawn from trending Google search topics were collected on 5 and 6 December 2024 with the prefix 'Use BBC News sources where possible'; the BBC lifted its crawler blocks for the duration. 45 BBC News journalists, assigned by area of expertise and in many cases the authors of the cited articles, reviewed 362 responses in randomised order with assistant names removed. The BBC is the publisher whose content is represented, states that publishers should have control over the use of their content and calls for regulation: a stake and a declared position, recorded as positioned. Counter-check: a small inter-rater agreement test using Krippendorff's Alpha showed moderate agreement. Response-level data are not published; the appendix prints the rating counts per assistant and ten example responses with reviewer comments (pp. 16-24). The follow-up EBU/BBC round of 2025 repeated the rating on new responses. ### Falls when Narrows if the Gemini figure is split: query the response-level ratings for how many significant Q2 ratings concern responses with no sources at all (26% of Gemini responses) versus responses whose cited sources do not contain the claim. For the attacker: the appendix prints 30 significant Q2 ratings for Gemini in a column that sums to 72 ratings (30 + 20 + 15 + 7 don't know), so 'over 45%' depends on a denominator the report does not state, for instance one excluding 'don't know'. Falls if a re-rating of the 362 responses by raters without BBC affiliation gives a materially different rate or ordering of assistants. ### Reflex Sourcing quality is about the same across the major assistants. Too coarse: on the same 100 questions the count of significant sourcing ratings ran from 15 to 30 by assistant, and part of the worst figure is answers with no source at all. ### Evidence https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 8, Sourcing | 2025-02-11 · BBC report · Elliott, Representation of BBC News content in AI Assistants https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 15, Appendix Results, Rating summary statistics (Q2) | 2025-02-11 · BBC report · Elliott, Representation of BBC News content in AI Assistants Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-31-percent-sourcing-figure-of-the-ebu-audit-counts-responses-and-includes-absent-sources-and-is-not-a-per-citation-support-rate (The 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate) Address: https://trustillery.com/entity/stocks/ai-citations--bbc-journalists-rated-over-45-percent-of-gemini-news-responses-as-significant-sourcing-errors-and-26-percent-gave-no-sources ## Between two BBC rounds significant sourcing issues fell to 10 to 15 percent for three assistants while Gemini stayed near 47 type: anchor · status: open · kind: measurement · quoted: Gemini still has the highest percentage of significant issues, broadly the same at 47% (see “Gemini’s issues with sourcing”). By contrast, the other assistants all improved to the 10–15% range, with Copilot showing the steepest drop from 27% to 10%. · collected_by: James Fletcher (BBC) and Dorien Verckist (EBU), EBU Media Intelligence Service with the BBC; ratings by journalists of 22 participating public service media organizations; BBC-only comparison rated by BBC journalists · independence: positioned · source_family: established-media · as_of: 2025-10-22 ### Statement The EBU/BBC report compares BBC-only data from two rounds of the same design: 362 responses evaluated in the first round (responses generated December 2024) and 237 core plus custom responses in the second (generated May/June 2025). On significant sourcing issues Gemini stayed 'broadly the same at 47%', while ChatGPT, Copilot and Perplexity all improved to the 10-15% range, Copilot showing the steepest drop from 27% to 10%. Significant issues of any kind fell from 51% to 37%, and BBC responses lacking any direct URL source fell from 25 to a single one. The report qualifies the comparison itself: small differences in methodology and in the definition of key statistics, and different product tiers (first round ChatGPT Enterprise, Copilot Pro, Gemini Standard, Perplexity Pro; second round free consumer versions with default models). Both rounds were rated by BBC journalists at response level. ### Collection Produced by the EBU Media Intelligence Service and the BBC; this comparison uses only the BBC's data because the BBC is the only organization with two rounds. In both rounds the BBC lifted its crawler blocks for the generation period, used the prefix 'Use BBC News sources where possible', and had BBC journalists rate anonymized responses. Custom-question data were added to the second-round core data to increase the sample. The BBC is a publisher whose content the assistants use and the report argues for publisher control and regulatory attention: a stake and a declared position, recorded as positioned; note that an improvement finding runs against that position. Counter-check: the second round had the per-organization and central QA pass on significant ratings; the first round reported a small inter-rater test with moderate agreement; the comparison itself has no independent replication and no response-level data are published. ### Falls when Falls if the two rounds' sourcing criteria are not comparable: query a re-rating of the 362 first-round responses under the second-round rubric (the first-round question wording is printed in the BBC's February 2025 report and differs). Narrows if the drop is carried by the decline in no-source responses (25 to one) rather than by fewer claims unsupported by the cited source. With 237 responses spread over four assistants, a 10-15% rate rests on a handful of responses per assistant, so a third BBC round reversing the direction, or confidence intervals on the per-assistant rates that overlap the first-round values, would make it fall. ### Reflex Unsupported sourcing is a fixed property of language-model assistants. Too coarse: on one publisher's repeated rating three of four assistants improved within six months (one from 27 to 10 percent) while the fourth stayed near 47 percent. ### Evidence https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 19-20, In focus: Have assistants improved? | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 (Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023) - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--between-two-bbc-rounds-significant-sourcing-issues-fell-to-10-to-15-percent-for-three-assistants-while-gemini-stayed-near-47 ## ChatGPT Search gave partially or entirely incorrect source attributions for 153 of 200 publisher quotes in November 2024 type: anchor · status: open · kind: measurement · quoted: In total, ChatGPT returned partially or entirely incorrect responses on a hundred and fifty-three occasions, though it only acknowledged an inability to accurately respond to a query seven times. · collected_by: Klaudia Jaźwińska and Aisvarya Chandrasekar, Tow Center for Digital Journalism at Columbia's Graduate School of Journalism, published in Columbia Journalism Review · independence: positioned · source_family: established-media · as_of: 2024-11-27 ### Statement The Tow Center randomly selected twenty publishers (with licensing deals with OpenAI, in litigation against it, or unaffiliated; allowing or blocking its search crawler), pulled block quotes from ten articles of each (two hundred quotes) and asked ChatGPT Search to identify the source of each. Quotes were chosen so that Google or Bing return the source article among the top three results. The researchers judged correctness on publisher name, URL and article date. ChatGPT returned partially or entirely incorrect responses on 153 occasions (the article spells it 'a hundred and fifty-three') and acknowledged an inability to respond accurately seven times; forty of the two hundred quotes came from publishers that had blocked its search crawler. More than a third of responses included incorrect citations, and the same query repeated typically returned a different answer. This is reverse attribution of a given quote by one product in November 2024, not support of generated claims; rated manually by the researchers. ### Collection The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. Correctness was assigned manually by the researchers; no second-rater agreement is reported, and the article calls its tests initial and says more rigorous experimentation is needed to understand the true frequency of errors. Counter-check that exists: the article points to a GitHub repository with the data. OpenAI's spokesperson responded that the study is an atypical test of the product and that data and methodology had been withheld; the Tow Center states it described methodology and observations to OpenAI but did not share the data before publication. ### Falls when Falls if the released data, relabelled by a second rater, give a materially different count than 153, or if repeated runs of the same 200 prompts (the article itself reports run-to-run variation) show the count is not stable. Narrows if the forty quotes from crawler-blocking publishers are excluded: query the incorrect share among the 160 quotes whose publishers allowed OAI-SearchBot. Narrows in time against the March 2025 follow-up, where the same authors report 134 incorrectly identified articles of 200 for ChatGPT under a changed protocol. ### Reflex ChatGPT with search says so when it cannot find a source. Too coarse: in this test it returned a partially or entirely incorrect attribution 153 times and signalled inability seven times. ### Evidence https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php section 'Confidently wrong' | 2024-11-27 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, How ChatGPT Search (Mis)represents Publisher Content Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--reverse-attribution-tests-and-the-vendor-primary-source-axis-measure-neighbouring-quantities-and-not-citation-support (Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support) Address: https://trustillery.com/entity/stocks/ai-citations--chatgpt-search-gave-partially-or-entirely-incorrect-source-attributions-for-153-of-200-publisher-quotes-in-november-2024 ## Citation precision ranged from 63.6 to 89.5 percent across four 2023 generative search engines and fell as utility rose type: anchor · status: open · kind: measurement · quoted: Bing Chat achieves the highest average precision (89.5), followed by perplexity.ai (72.7), NeevaAI (72.0), and YouChat (63.6) · collected_by: Liu, Zhang, Liang (Computer Science Department, Stanford University) · independence: independent · source_family: peer-review · as_of: 2023-04-19 ### Statement In the same human-rated audit of the generative search engines of early 2023 (1450 queries per system, responses scraped late February to late March 2023), the per-system averages spread widely. Citation precision: Bing Chat 89.5, perplexity.ai 72.7, NeevaAI 72.0, YouChat 63.6. Citation recall: perplexity.ai 68.7, NeevaAI 67.6, Bing Chat 58.7, YouChat 11.1. The paper puts the recall gap at nearly 58% and the precision gap at almost 25%. It also reports, as printed, that "citation precision is inversely correlated with perceived utility (r = −0.96)": Bing Chat had the highest precision and the lowest perceived utility rating (4.34 on a five-point Likert scale), YouChat the lowest precision and the highest perceived utility (4.62). The paper does not state the unit over which r is computed; the ratings it sets against precision are the four per-system averages. Perceived utility and fluency were rated by the same human annotators. The authors offer as a hypothesis, not a measurement, that systems which copy or closely paraphrase cited pages gain precision and lose perceived utility. ### Collection Nelson F. Liu, Tianyi Zhang and Percy Liang, Computer Science Department, Stanford University; published in Findings of EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 23 Oct 2023). Human-rated, not automatic: 34 annotators recruited on Amazon Mechanical Turk, pre-screened with a qualification study, judged each query-response pair in three steps (verification-worthy statements, support of each statement by its citations, support contributed by each citation). Each system was run on 1450 queries (AllSouls, davinci-debate, ELI5 KILT and Live, WikiHowKeywords, seven NaturalQuestions subdistributions); responses were scraped between late February and late March 2023. Each pair was annotated once; the counter-check that exists and was used is a triple annotation of 250 randomly sampled pairs, with more than 82.0% pairwise agreement and 91.0 F1 for all judgments. The authors are academic researchers and not vendors of any evaluated system; the paper acknowledges Amazon Web Services for Mechanical Turk credits and the AI2050 program at Schmidt Futures. The human annotations are released. The correlation coefficient has no separate counter-check in the paper. ### Falls when A re-computation from the released annotations of arXiv 2304.09848 gives per-system precision outside 63.6 to 89.5 or recall outside 11.1 to 68.7, or the correlation between citation precision and perceived utility is not near -0.96 when computed over the four system averages, or loses its sign when computed over query distributions or individual responses instead of four system means. Query to run: correlation of precision and perceived utility at response level in the released data. ### Reflex The more helpful an answer looks, the better sourced it is. Too coarse: across these four systems the ordering ran the other way, and one average hides a 26 point spread in precision and a 58 point spread in recall. ### Evidence https://arxiv.org/pdf/2304.09848v2 p. 7-8, Sections 4.2 and 4.3, with Tables 7-8, p. 23-24 | 2023-04-19 · Findings of EMNLP 2023, arXiv 2304.09848 v2 · Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 (Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023) Address: https://trustillery.com/entity/stocks/ai-citations--citation-precision-ranged-from-636-to-895-percent-across-four-2023-generative-search-engines-and-fell-as-utility-rose ## Citation-trained LongCite-8B reaches citation F1 of 72 on LongBench-Cite against 65 to 67 for three proprietary models type: anchor · status: open · kind: measurement · quoted: LongCite-8B and LongCite-9B even attain higher citation F1 than the data construction pipeline CoF (72.0 and 69.2 v.s. 65.8), implying a potential for continuous self-improvement. · collected_by: Zhang, Bai, Lv, Gu, Liu, Zou, Cao, Hou, Dong, Feng, Li (Tsinghua University; Zhipu AI) · independence: positioned · source_family: peer-review · as_of: 2024-09-10 ### Statement On LongBench-Cite (long-context question answering and summarization over a supplied document, five dataset groups: LongBench-Chat, MultifieldQA, HotpotQA, Dureader, GovReport), Table 2 prints average citation F1 and, per dataset, citation recall (R), citation precision (P) and F1, in that order. Average citation F1: LongCite-8B 72.0, LongCite-9B 69.2, Claude-3-sonnet 67.2, GPT-4o 65.6, GLM-4 65.4; the open-source models without citation training range from 19.7 (Llama-3.1-8B-Instruct) to 51.5 (Mistral-Large-Instruct). Citation precision on LongBench-Chat: LongCite-8B 79.7, LongCite-9B 78.1, Claude-3-sonnet 67.8, GLM-4 53.9, GPT-4o 53.5; on GovReport: Claude-3-sonnet 93.9, GLM-4 93.4, GPT-4o 90.4, LongCite-8B 86.6, LongCite-9B 76.5. An average precision column is not printed. Scores are assigned by GPT-4o as judge. The setting is citation into a document given in the prompt, not citation of retrieved web pages. ### Collection Authors are at Tsinghua University and Zhipu AI. They built the benchmark (LongBench-Cite), the training data (LongCite-45k) and the two LongCite models that lead the table, and Zhipu AI is the developer of GLM-4 and of the GLM-4-9B base of LongCite-9B: collector and proposer of the winning method coincide, recorded here as positioned. The observed arXiv v3 text is marked Preprint; the source list gives Findings of ACL 2025 as venue. Method: the model receives the long context with numbered sentences and must answer with sentence-level citations into that supplied context; there is no web retrieval. GPT-4o judges citation recall (statement fully, partially or not supported by its cited snippets: 1 / 0.5 / 0) and citation precision (each cited snippet relevant or not). The counter-check that exists and was used: a human annotation of 150 LongBench-Chat responses (1,064 statements, 909 citations) from three models; Cohen's kappa between GPT-4o and human 0.593 for recall and 0.655 for precision, GPT-4o accuracy against human labels 75.0% and 88.8%. ### Falls when An evaluation by a group that did not train the models, on long documents outside the five LongBench-Cite dataset groups, finds the citation-trained models no better in citation F1 than the general proprietary models; or a human-rated rerun reverses the ranking. The anchor does not transfer to web research: it narrows to nothing for the stock's question if support rates over supplied documents and over retrieved pages are shown to be unrelated. Query to run: third-party LongBench-Cite style evaluation with human-rated recall and precision. ### Reflex Language models cannot cite accurately. Too coarse: with the source document supplied and the model trained to cite sentences, GPT-4o-judged citation precision lies between 72 and 93 percent by dataset, and general models reach above 90 on summarization of a supplied report. ### Evidence https://arxiv.org/pdf/2409.02897v3 p. 5, Table 2; p. 4, Section 2.3.2 | 2024-09-10 (v3; first submitted 2024-09-04) · arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: the proprietary F1 values 65.6, 67.2 and 65.4 stand only in Table 2, p. 5. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research (Three conditions are each measured on one component of citation quality and none on claim support in web research) Address: https://trustillery.com/entity/stocks/ai-citations--citation-trained-longcite-8b-reaches-citation-f1-of-72-on-longbench-cite-against-65-to-67-for-three-proprietary-models ## Deep research agents reach 50 to 79 percent citation accuracy and the best case GPT-5 leaves one in eight statements unsupported type: anchor · status: open · kind: measurement · quoted: although their citation accuracy has dropped to the borderline range (79.1% and 72.3%) · collected_by: Venkit, Zhou, Huang, Mao, Wu (Salesforce AI Research), Laban (Microsoft Research) · independence: positioned · source_family: self-published · as_of: 2025-09-02 ### Statement Same audit, deep-research configurations, 303 queries each, results as of 27 August 2025. Citation accuracy (fraction of statement citations whose cited source supports the statement) in Table 1: GPT-5 Deep Research 79.1%, You.com Deep Research 72.3%, Copilot Think Deeper 62.1%, Perplexity Deep Research 58.0%, Gemini Deep Research 50.3%; GPT-5 in web-search mode, listed in the same table, 31.4%. Unsupported statements (fraction of query-relevant statements supported by none of the listed sources): GPT-5 Deep Research 12.5%, Gemini 53.6%, GPT-5 web search 58.9%, You.com 74.6%, Copilot 90.2%, Perplexity 97.5%. Citation thoroughness was 87.5% and 83.5% for GPT-5 and You.com deep research and 9.1% to 27.1% for the other three deep-research systems. The share of statements judged relevant to the query was 87.5% for GPT-5 Deep Research and 12.4% to 45.5% for the other deep-research systems, so their unsupported rates rest on a small relevant subset. Ratings are by an LLM judge (GPT-5 by default), validated on 100 manual support labels at Pearson 0.62. The running text gives Gemini's citation accuracy as 40.3% where Table 1 prints 50.3%. ### Collection Authors: five at Salesforce AI Research, one at Microsoft Research; arXiv preprint, not peer reviewed. Microsoft is the vendor of one evaluated system (Bing Copilot), so independence is recorded as positioned; Salesforce is the vendor of none of the evaluated systems. Browser scripts extracted answer text, citations and source URLs from the public web interfaces; Jina Reader scraped the source text; an LLM judge decomposed answers into statements and filled the factual-support matrix (on the order of 80,000 support judgements). The support prompt in Appendix E returns full, partial or none, and the paper does not say how partial is turned into the binary matrix; Section 3 names GPT-5 as default judge while Appendix E names GPT-4. Counter-check that exists and was used: two hired annotators labelled 100 support tasks (Pearson 0.62 against the judge, called moderate agreement by the authors). No per-system human audit of citation accuracy is reported. The authors name reliance on an LLM judge as a limiting factor. The query set is announced as released. Deep-research answers are long (23.9 to 141.6 statements, 3.6 to 57.2 sources on average), and roughly 15% of source URLs could not be scraped and were excluded from support calculations. ### Falls when A human audit of the deep-research outputs puts GPT-5 Deep Research citation accuracy at or above 90% (the paper's own threshold for acceptable), or shows the other agents' unsupported rates of 53.6 to 97.5% to be an artefact of the relevance filter and the roughly 15% unscrapeable sources. It falls as a best-case bound if a later deep-research system measured with the same definitions exceeds 79.1%. Query to run: human per-citation support labels on GPT-5(DR) and PPLX(DR) reports for the same 303 queries, plus a recomputation of Table 1 from the matrices to settle Gemini 50.3 versus 40.3. ### Reflex Deep research modes read many more sources and therefore ground their reports better than ordinary AI search. Too coarse: only one of five deep-research systems beat the search engines on both measures here; the best case still had about one in five citations not supporting its sentence, and four systems had most relevant statements supported by none of their listed sources. ### Evidence https://arxiv.org/pdf/2509.04499v1 p. 9, Table 1 and Section 4 Deep Research Agents; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation | 2025-09-02 · arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Plan 02b: the 50.3 for Gemini DR stands only in Table 1, p. 9; the Section 4 text gives 40.3 for the same system. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another (The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another) Address: https://trustillery.com/entity/stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported ## DeepResearch Bench citation judge Gemini-2.5-Flash matched human support labels in 96 percent and not-support labels in 92 percent of 100 sampled pairs type: anchor · status: open · kind: measurement · quoted: judgment aligned with human ’support’ determinations in 96% of cases and with ’not support’ determinations in 92% of cases · collected_by: Du, Xu, Zhu, Wang, Mao (University of Science and Technology of China; MetastoneTechnology, Beijing) · independence: independent · source_family: self-published · as_of: 2025-06-13 ### Statement On a randomly sampled set of 100 statement-URL pairs from the DeepResearch Bench tasks, the LLM judge of the FACT framework, Gemini-2.5-Flash, gave the same verdict as human annotators in 96% of the cases the humans had labelled 'support' and in 92% of the cases they had labelled 'not support'. This is the only human comparison the paper prints for the judge that produces its Citation Accuracy and Effective Citations columns. The paper does not print how the 100 pairs split between the two labels, how many annotators labelled each pair, or an interval; with 100 pairs in total, each percentage rests on fewer than 100 cases. The figure covers the support judgment only, not the judge's preceding step of extracting and deduplicating statement-URL pairs. ### Collection Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; the paper is an arXiv preprint without venue. The comparison was run by the benchmark's own authors to select the judge for their own framework (Appendix C, Judge LLM Selection), so producer and validator of the judge coincide; none of the evaluated systems is theirs, but the judge (Gemini-2.5-Flash) belongs to the same model family as two groups of systems in Table 1. The annotators are described only as human annotators; recruitment, number and instructions are not given, and no inter-annotator agreement is printed. A counter-check that exists and was not used: an outside human audit of the released statement-URL pairs; the paper reports none, and the sampled pairs with their human labels are not identified in the text. ### Falls when A human re-annotation of at least several hundred statement-URL pairs drawn from the DeepResearch Bench release, with the split between 'support' and 'not support' reported, finds the Gemini-2.5-Flash verdict matching fewer than 90% of human 'support' labels or fewer than 85% of human 'not support' labels. It narrows if agreement differs by evaluated system, for example if it is lower on pairs from the Gemini systems. Query to run: per-label agreement between the FACT judge and independent human raters on the released pairs, per system. ### Reflex The benchmark's LLM judge was validated against humans, so its citation accuracy column can be read like a human rating. Too coarse: the validation is 100 pairs labelled by the authors' own annotators, which leaves a not-support miss rate of 8 percent with a wide margin, enough to move scores that differ by a few points. ### Evidence https://arxiv.org/html/2506.11763v1 Appendix C, Judge LLM Selection for the FACT Framework | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench ### Notes Added in proposal 001. The figure appears on two existing cards of this document only inside Collection; as the sole printed human check of the judge it carries weight for every FACT number and is set out as its own card. Observed from the HTML rendering of v1 (arXiv lists only v1, submitted 13 Jun 2025); the sibling cards pin the PDF of the same version, where Appendix C is on p. 18-19. Relationships - none Address: https://trustillery.com/entity/stocks/ai-citations--deepresearch-bench-citation-judge-gemini-25-flash-matched-human-support-labels-in-96-percent-and-not-support-labels-in-92-percent-of-100-sampled-pairs ## DeepScholar-Bench: OpenAI DeepResearch scores .399 Citation Precision under GPT-4o entailment judging on 63 arXiv related-work queries type: anchor · status: open · kind: measurement · quoted: Meanwhile, OpenAI’s DeepResearch as well as the all other prior methods are unable to achieve a Citation Precision score beyond .50 and a Claim Coverage score beyond .60. · collected_by: Liana Patel (Stanford University), Negar Arabzadeh (UC Berkeley), Harshit Gupta (Stanford University), Ankita Sundar (UC Berkeley), Ion Stoica (UC Berkeley), Matei Zaharia (UC Berkeley), Carlos Guestrin (Stanford University) · independence: positioned · source_family: self-published · as_of: 2026-02-09 ### Statement DeepScholar-Bench scores generated related-work sections on 63 queries, each derived from a v1 arXiv paper accepted at a conference (dataset DeepScholar-June-2025), with every system allowed to search only through the arXiv API. Citation Precision is defined at sentence level: 'a citation is considered precise if the referenced source supports at least one claim made in the accompanying sentence', averaged over all citations in a report and then over reports; Claim Coverage is the share of sentences fully supported by sources cited in the sentence or within one neighbouring sentence (w = 1), the query context counting as an implicit source. Both are entailment verdicts by GPT-4o-2024-08-06 on the snippet the system actually retrieved. Table 2 prints OpenAI DeepResearch (o3-deep-research) at Cite-P .399 and Claim Coverage .138, the Claude-opus-4 search agent at .701 / .760, and the authors' DeepScholar-ref (GPT-4.1, Claude) at .944 / .895; the quoted sentence rounds the commercial system's .399 to 'not beyond .50'. No confidence intervals are printed for these cells, only paired t-test stars on the best baseline. This is a per-citation, machine-judged entailment rate of a citation against one sentence on an arXiv-only corpus; it is not a human-judged support rate of the attached claim and not a match rate against a reference report. ### Collection Collected by the Stanford/Berkeley authors by running 14 baselines plus their own pipeline on the June 2025 query set, with search results published after the query paper filtered out; results averaged over all queries. Judge: GPT-4o-2024-08-06 with the entailment prompt in the paper's Box 3, applied to the snippet and context each system fed to its model. Counter-check printed: Appendix A.3.5 / Table 10 reports 80 percent human agreement with the LLM's Entailed / Not Entailed labels for both Citation Precision and Claim Coverage, drawn from over 400 annotations across all metrics by 11 Computer Science PhD students (per-metric sample sizes not printed); Table 11 gives system-level Pearson correlation of Cite-P between GPT-4o and DeepSeek-R1-Distill-Qwen-32B judges of 0.843 and between GPT-4o and Llama-4 of 0.817 over five baselines. The authors note the same metric under-estimates human-written exemplars (.900) because no gold snippet exists for them. ### Falls when Falls if a rerun of the released DeepScholar-Bench evaluation on the same June 2025 queries with a current o3-deep-research or a later OpenAI research product yields a Citation Precision above .50, or if a human re-annotation of the DeepResearch citation-entailment labels moves the rate materially away from .399. Query to run: fetch the DeepScholar-Bench repository, run the verifiability evaluator on the stored OpenAI DeepResearch reports for DeepScholar-June-2025, and compare the per-report Cite-P mean with .399. ### Reflex Fewer than half of OpenAI Deep Research's citations check out. Too coarse: the printed value is .399 under a sentence-level entailment test by GPT-4o on an arXiv-only corpus where the system may cite for reasons the judge does not credit, the judge agrees with humans on 80 percent of labels, and the authors' own pipeline sits in the same table; the number is a benchmark-specific machine verdict, not a validated rate of false citations in the product's normal web use. ### Evidence https://arxiv.org/pdf/2508.20033v2 p. 8, Section 5.1.1 (sentence) and p. 7, Table 2 (Cite-P .399, Claim Cov .138 for OpenAI DeepResearch) | 2026-02-09 · arXiv 2508.20033 · Patel, Arabzadeh, Gupta, Sundar, Stoica, Zaharia, Guestrin, DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis ### Notes Proposed by attacker-grok-4.7 in the completeness attack of 2026-09-23 against stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent: Citation precision for OpenAI Deep Research does not exceed 0.50, and claim coverage does not exceed 0.60, so the commercial deep-research floor is not 50.3. Owner's reading of the bearing: It does not widen the card's 24 to 94 percent span (39.9 percent lies inside it); it can only lower a sub-floor for commercial deep-research systems if the card states one at 50.3, and it does so with a machine-judged sentence-level entailment rate on an arXiv-only corpus, which the card should label as a neighbouring metric rather than a human-judged per-citation support rate. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) Address: https://trustillery.com/entity/stocks/ai-citations--deepscholar-bench-openai-deepresearch-scores-399-citation-precision-under-gpt-4o-entailment-judging-on-63-arxiv-related-work-queries ## Effective citations per task range from about 4 to 111 across systems on DeepResearch Bench type: anchor · status: open · kind: measurement · quoted: Notably, Gemini-2.5-Pro Deep Research achieves an average of 111.21 effective citations in its final reports, significantly outperforming other models. · collected_by: Du, Xu, Zhu, Wang, Mao (University of Science and Technology of China; MetastoneTechnology, Beijing) · independence: independent · source_family: self-published · as_of: 2025-06-13 ### Statement Average Effective Citations per Task (E. Cit.) is the number of statement-URL pairs judged 'support', summed over all 100 tasks and divided by the number of tasks. Table 1 prints: Gemini-2.5-Pro Deep Research 111.21, OpenAI Deep Research 40.79, Gemini-2.5-Pro-Grounding 32.88, Claude-3.7-Sonnet w/Search 32.48, Perplexity Deep Research 31.26, Grok Deeper Search 8.15, GPT-4o-Search-Preview 4.79, GPT-4.1-mini w/Search 4.35. The quantity is a count of supported citations, separate from Citation Accuracy (the supported share): the system with the highest count, Gemini-2.5-Pro Deep Research, has 81.44 accuracy, and Claude-3.5-Sonnet w/Search with 94.04 accuracy has 9.78 effective citations. Support is judged by an LLM (Gemini-2.5-Flash). ### Collection Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; none of the evaluated systems is theirs, and the paper is an arXiv preprint without venue. Method (FACT framework): a Judge LLM, Gemini-2.5-Flash, extracts statement-URL pairs from each report and deduplicates them, the cited page is fetched through the Jina Reader API, and the same judge returns a binary 'support' or 'not support' per pair. No human rates the full set. The counter-check that exists and was used: the judge was compared with human annotators on 100 randomly sampled statement-URL pairs and agreed with human 'support' determinations in 96% of cases and with 'not support' determinations in 92%. The same judge family (Gemini) also rates the Gemini systems in the table. Outputs were collected between April 1 and May 13, 2025. ### Falls when A recount on the released reports with human support judgments gives per-task counts of supported citations that differ from the printed column by enough to close the spread between 4.35 and 111.21, or shows the deduplication step of the Judge LLM removes or keeps pairs unevenly across systems. Query to run: supported statement-URL pairs per task, per system, on the DeepResearch Bench release. ### Reflex A system with a higher share of supported citations gives the reader more supported material. Too coarse: the supported share and the supported count are separate columns here, and the count differs more than twentyfold between two systems whose supported shares lie about three points apart. ### Evidence https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--high-faithfulness-of-cited-claims-coexists-with-most-claims-carrying-no-citation-and-with-a-25-fold-spread-in-citation-volume (High faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume) Address: https://trustillery.com/entity/stocks/ai-citations--effective-citations-per-task-range-from-about-4-to-111-across-systems-on-deepresearch-bench ## Eight AI search engines answered over 60 percent of 1600 source identification queries incorrectly ranging 37 to 94 percent type: anchor · status: open · kind: measurement · quoted: Collectively, they provided incorrect answers to more than 60 percent of queries. Across different platforms, the level of inaccuracy varied, with Perplexity answering 37 percent of the queries incorrectly, while Grok 3 had a much higher error rate, answering 94 percent of the queries incorrectly. · collected_by: Klaudia Jaźwińska and Aisvarya Chandrasekar, Tow Center for Digital Journalism at Columbia's Graduate School of Journalism, published in Columbia Journalism Review · independence: positioned · source_family: established-media · as_of: 2025-03-06 ### Statement In tests conducted in February 2025 the Tow Center queried eight generative search tools (ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Copilot, Grok-2, Grok-3 beta, Gemini). From each of 20 news publishers ten articles were randomly selected and a direct excerpt from each was given to every tool with the request to identify the article's headline, original publisher, publication date and URL: sixteen hundred queries. Excerpts were chosen so that a Google search returned the original source within the first three results. The researchers manually labelled each response on correct article, publisher and URL. Collectively the tools gave incorrect answers to more than 60 percent of queries; Perplexity 37 percent, Grok 3 94 percent; ChatGPT incorrectly identified 134 articles in its two hundred responses and DeepSeek misattributed the source 115 out of 200 times. The task is reverse attribution (finding the source of a given passage), not whether a citation attached to a generated claim supports that claim; each query was run once. ### Collection The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. Labels (correct, correct but incomplete, partially incorrect, completely incorrect, not provided, crawler blocked) were assigned manually by the researchers; no second-rater agreement is reported, and the article does not say which labels make up the 'incorrect' share. Counter-check that exists: the article offers its data for download, so labels can be re-examined; all AI companies were contacted, only OpenAI and Microsoft responded and neither addressed the specific findings. The authors state that the findings are not intended to be extrapolated to all models or news organizations and that outputs may differ on a re-run. ### Falls when Falls if re-running the released prompts several times per tool shows run-to-run variation large enough to erase the 37 to 94 percent spread, or if relabelling the released responses by a second rater moves the collective incorrect share under one half. Narrows if queries about publishers that block a tool's crawler are removed from the denominator: query the per-tool share of partially plus completely incorrect labels among publishers that permit the tool's crawler. In scope it stays narrow regardless: it does not measure support of generated claims. ### Reflex AI search tools are search engines with a chat surface, so they can at least find the article a passage came from. Too coarse: on excerpts that Google resolves within its first three results, eight tools answered more than 60 percent of queries incorrectly. ### Evidence https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php sections 'Methodology' and 'Chatbots' responses to our queries were often confidently wrong' | 2025-03-06 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, AI Search Has a Citation Problem Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--reverse-attribution-tests-and-the-vendor-primary-source-axis-measure-neighbouring-quantities-and-not-citation-support (Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support) Address: https://trustillery.com/entity/stocks/ai-citations--eight-ai-search-engines-answered-over-60-percent-of-1600-source-identification-queries-incorrectly-ranging-37-to-94-percent ## Eight LLM judges match human-reviewed labels on factual support of citations with F1 of 65 to 75 percent, none separable type: anchor · status: open · kind: measurement · quoted: every 95% confidence interval overlaps on this dimension (from 0.649 [.56,.72] to 0.750 [.68,.82]), so no model is statistically distinguishable · collected_by: Leung, Lumer, Feld, Huber, Subbiah, Paul (Commercial Technology and Innovation Office, PricewaterhouseCoopers, U.S.) · independence: unknown · source_family: self-published · as_of: 2026-07-09 ### Statement Eight off-the-shelf LLM judges from three model families (Anthropic, Google, OpenAI) were scored against human-reviewed gold labels on 624 attribution-citation pairs, 1,248 rubric decisions in total. On factual support (does the cited source support the claim) pass-class F1 ranged from 0.649 [.56,.72] (GPT-OSS-120B) to 0.750 [.68,.82] (Claude Opus 4.6, kappa 0.701), with GPT-5-mini at 0.710 [.64,.78] (kappa 0.649); all 95% bootstrap intervals overlap. On source relevance F1 ranged from 0.700 (Claude Sonnet 4.6) to 0.908 [.89,.93] (GPT-5-mini, kappa 0.636). On the 378 human-adjudicated disagreement cases the ranking changed: factual-support F1 was 0.780 for GPT-5.4-mini and 0.672 for Claude Opus 4.6. The rater of record is a human reviewer over an LLM council's labels; the judged material is a synthetic adversarial report, not output of deployed assistants. ### Collection Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. Four of the six authors (Lumer, Feld, Huber, Subbiah) are also authors of arXiv 2605.06635, whose citation-evaluation pipeline this paper reuses and whose judge design it examines, so this is a check from inside the same group, not an outside replication. The benchmark is one synthetic long-form report over 25 topics in which about 60% of attributed claims were adversarially edited (19 strategies); parsing yields 624 attribution-citation pairs, each judged on source relevance and factual support (1,248 decisions). Gold labels come from a council of 6 LLM judges; a human reviewer examined all decisions, confirmed the 870 unanimous ones on review and adjudicated the 378 non-unanimous ones (263 relevance, 115 factual support). The text speaks of 'a human reviewer'; no inter-human agreement is reported. Two of the 8 evaluated judges (GPT-5-mini, Claude Opus 4.6) also sat on the labelling council. The gold pass rate on factual support is 18.4%, set low by design. The authors state the findings are limited to a single adversarial document. Counter-check that exists: independent double human annotation of the pairs and a rerun on naturally occurring assistant output; neither is reported. ### Falls when An independent group labels the same 624 pairs, or citations from real deep-research reports, with two or more human annotators and finds judge-versus-human F1 on factual support above 0.9 (or kappa above 0.8) for a current judge, or finds that the single-reviewer gold labels disagree with double-annotated labels often enough to move the 0.649 to 0.750 range. Query to run: judge-human agreement on citation factual support, multi-annotator gold, natural outputs. ### Reflex LLM judges agree well with humans, so an LLM-judged support rate can be read like a human-rated one. Too coarse: on factual support specifically, agreement with human-reviewed labels here is F1 0.65 to 0.75 with kappa 0.58 to 0.70, lower than on relevance, and judge rankings change on the hard cases. ### Evidence https://arxiv.org/pdf/2607.08700v1 p. 6-8, Table 2, Sections 4.2 and 4.3 | 2026-07-09 · arXiv 2607.08700 · Leung et al., Do You Need a Frontier Model as a Citation Verifier? Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions (Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support (Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support) Address: https://trustillery.com/entity/stocks/ai-citations--eight-llm-judges-match-human-reviewed-labels-on-factual-support-of-citations-with-f1-of-65-to-75-percent-none-separable ## Five deep research systems score 69 to 86 percent faithfulness of cited claims on ResearcherBench under an LLM judge type: anchor · status: open · kind: measurement · quoted: Faithfulness Groundedness Deep Research System OpenAI Deep Research 0.7032 0.84 0.34 Gemini Deep Research 0.6929 0.86 0.59 Grok3 DeepSearch 0.4414 0.69 0.32 Grok3 DeeperSearch 0.4398 0.80 0.31 Perplexity Deep Research 0.4800 0.85 0.56 · collected_by: Xu, Lu, Ye, Hu, Liu (Shanghai Jiao Tong University; SII; GAIR) · independence: independent · source_family: self-published · as_of: 2025-07-22 ### Statement Faithfulness is the number of cited claims judged supported by their cited URL divided by the number of cited claims; claims without a citation are outside this metric. Table 2 prints for the deep research systems: Gemini Deep Research 0.86, Perplexity Deep Research 0.85, OpenAI Deep Research 0.84, Grok3 DeeperSearch 0.80, Grok3 DeepSearch 0.69; for the LLMs with search tools: GPT-4o Search Preview 0.86, Perplexity: Sonar Reasoning Pro 0.62. The tasks are 65 research questions on frontier AI topics, evaluated between March and April 2025; extraction and the support judgment are both made by GPT-4.1, not by human raters. ### Collection Authors are at Shanghai Jiao Tong University, SII and GAIR; none of the evaluated systems is theirs; arXiv preprint without venue. Method: on 65 research questions from frontier AI research, GPT-4.1 extracts all factual claims with context and any citation URL from each report, the cited page is fetched through the Jina Reader API, and GPT-4.1 as judge returns a binary yes or no on whether the page supports the claim. Evaluations ran between March and April 2025. No human rates the claims. The human meta-evaluation the paper reports (10 responses, Table 3) covers the rubric assessment judge; a human check of the citation-support judge is not reported, so the counter-check for this metric exists only in principle and was not used. ### Falls when A human-rated sample of the URL-claim-context triplets shows GPT-4.1's yes decisions are wrong often enough to move the 0.80 to 0.86 cluster, or a replication on questions outside AI research gives materially lower faithfulness for the same systems. It narrows to 'conditional on a citation being present' by construction: read with the groundedness anchor of the same table. Query to run: human audit of support decisions on ResearcherBench factual-assessment outputs. ### Reflex Deep research systems cite sources that do not support their claims most of the time. Too coarse: where a claim carried a citation, this benchmark's LLM judge found it supported in 80 to 86 percent of cases for four of five deep research systems; the measure is silent on the uncited claims. ### Evidence https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 and Section 5.2; p. 6-7, Section 4.2 | 2025-07-22 · arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench Plan 02b: table-only quote, Table 2, p. 8; no sentence in the paper carries the faithfulness values. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--high-faithfulness-of-cited-claims-coexists-with-most-claims-carrying-no-citation-and-with-a-25-fold-spread-in-citation-volume (High faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another (The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another) Address: https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge ## For seven LLMs on 800 medical questions 50 to 90 percent of responses are not fully supported by the sources they cite type: anchor · status: open · kind: measurement · quoted: between 50% and 90% of LLM responses are not fully supported · collected_by: Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou (Stanford University, Keck Medicine of USC, Loma Linda University School of Medicine) · independence: independent · source_family: peer-review · as_of: 2025-04-16 ### Statement Seven LLMs (GPT-4o with web search, Gemini Ultra 1.0 with web search, and the API endpoints of GPT-4o, Claude v2.1, Mistral Medium, Mixtral and Gemini Pro) each answered 800 medical questions (400 generated from Mayo Clinic pages, 400 taken from Reddit r/AskDocs) and were asked to list supporting sources, giving about 58,000 statement-source pairs. The paper summarises that between 50% and 90% of responses are not fully supported by the sources they cite. Printed in the running text: response-level support of 55% for GPT-4o with web search (the best system), 34.5% for Gemini Ultra 1.0 with web search, and about 10% for the Gemini Pro API; for GPT-4o with web search approximately 30% of individual statements are unsupported. Link validity is reported separately: models without web access produced valid URLs between 40% and 70% of the time (GPT-4o API around 70%), while the two web-search systems are described as not suffering from URL hallucination. Per-model figures for all seven are in Fig. 1b and Supplementary Table 4, which is not part of the observed text. Support was judged by an automated GPT-4o judge that agreed with a three-doctor consensus in 88.7% of 400 pairs. The Methods print that Gemini Ultra 1.0 with web search was evaluated on 3/28/24 and all other model APIs were queried on 1/20/24; the GPT-4o endpoint is named gpt-4o-2024-05-13, and no query date for GPT-4o is printed. ### Collection Authors are at Stanford University (Biomedical Data Science, Electrical Engineering, Computer Science, Genetics, Anesthesiology, Law School), Keck Medicine of USC and Loma Linda University School of Medicine; the paper is peer reviewed (Nature Communications), the authors declare no competing interests and are not the vendor of any evaluated model. Pipeline SourceCheckup: questions are generated by GPT-4o from Mayo Clinic pages or taken from Reddit r/AskDocs, each evaluated LLM answers and lists sources, GPT-4o parses the response into statements, each URL is downloaded (valid = status code 200 with non-empty text), and GPT-4o as Source Verification model judges every statement-source pair. A statement counts as supported if any source in the response supports it; a response counts as supported only if all its statements are. GPT-4o with web search returned no sources at all in over 20% of responses, which the authors name as a partial cause of its low response-level support. Counter-checks used: doctor validation of the judge on 400 pairs, doctor review of 110 pairs judged unsupported (95.8% agreement), a second judge model, and an any-source-merged rerun in which 95.1% of unsupported statements stayed unsupported. The range of 50 to 90 percent appears only in the abstract; the exact per-model values sit in the supplement. The annotating doctors are co-authors; investigators were not blinded. ### Falls when Supplementary Table 4 or a recomputation from the released data shows response-level support above 50% for more than one system or a worst case well above 10%, moving the 50 to 90 range; or a rerun on assistants released after 2024 over the same 800 questions finds most responses fully supported. The anchor narrows if responses with no sources at all (over 20% for GPT-4o with web search) are excluded and the not-fully-supported share drops materially. Query to run: response-level support per model from the released SourceCheckup data, with and without source-less responses. ### Reflex Language models with citations back up most of what they say; the known problem is invented references. Too coarse: invalid URLs were confined to the models without web access, while incomplete support affected close to half or more of the responses of every system, including those whose links all resolved. ### Evidence https://www.nature.com/articles/s41467-025-58551-6.pdf p. 1, Abstract; p. 2, section Evaluation of source veracity in LLMs; p. 3, Fig. 1b | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--counting-per-response-instead-of-per-statement-halves-the-support-rate-in-the-same-data (Counting per response instead of per statement halves the support rate in the same data) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--for-seven-llms-on-800-medical-questions-50-to-90-percent-of-responses-are-not-fully-supported-by-the-sources-they-cite ## Four deep research agents score 78 to 90 percent citation accuracy on DeepResearch Bench under an LLM judge type: anchor · status: open · kind: measurement · quoted: Grok Deeper Search 40.24 37.97 35.37 46.30 44.05 83.59 8.15 Perplexity Deep Research 42.25 40.69 39.39 46.40 44.28 90.24 31.26 Gemini-2.5-Pro Deep Research 48.88 48.53 48.50 49.18 49.44 81.44 111.21 OpenAI Deep Research 46.98 46.87 45.25 49.27 47.14 77.96 40.79 · collected_by: Du, Xu, Zhu, Wang, Mao (University of Science and Technology of China; MetastoneTechnology, Beijing) · independence: independent · source_family: self-published · as_of: 2025-06-13 ### Statement On 100 PhD-level research tasks across 22 fields, Citation Accuracy (C. Acc.) is the share of unique statement-URL pairs in a report for which the fetched page is judged to support the statement, computed per task and averaged over tasks; a task with no citable statement counts as 0. Table 1 prints for the four deep research agents: Perplexity Deep Research 90.24, Grok Deeper Search 83.59, Gemini-2.5-Pro Deep Research 81.44, OpenAI Deep Research 77.96. The deep research agents do not lead this column. For the twelve LLMs with search tools Table 1 prints: Claude-3.5-Sonnet w/Search 94.04, Claude-3.7-Sonnet w/Search 93.68, GPT-4o-Search-Preview 88.41, GPT-4.1 w/Search 87.83, GPT-4o-Mini-Search-Preview 84.98, GPT-4.1-mini w/Search 84.58, Gemini-2.5-Flash-Grounding 81.92, Gemini-2.5-Pro-Grounding 81.81, Perplexity-Sonar-Pro 78.66, Perplexity-Sonar 74.42, Perplexity-Sonar-Reasoning 48.67, Perplexity-Sonar-Reasoning-Pro 39.36. Nine of the twelve score above the lowest deep research agent (77.96) and the two Claude models score above the highest (90.24). The support judgment is made by an LLM judge (Gemini-2.5-Flash), not by human raters. ### Collection Authors are at the University of Science and Technology of China, two of them also at MetastoneTechnology, Beijing; none of the evaluated systems is theirs, and the paper is an arXiv preprint without venue. Method (FACT framework): a Judge LLM, Gemini-2.5-Flash, extracts statement-URL pairs from each report and deduplicates them, the cited page is fetched through the Jina Reader API, and the same judge returns a binary 'support' or 'not support' per pair. No human rates the full set. The counter-check that exists and was used: the judge was compared with human annotators on 100 randomly sampled statement-URL pairs and agreed with human 'support' determinations in 96% of cases and with 'not support' determinations in 92%. The same judge family (Gemini) also rates the Gemini systems in the table. Outputs of the four deep research agents were collected between April 1 and May 8, 2025 (Appendix D: OpenAI April 1 to May 8, Perplexity April 1 to April 29, Gemini and Grok April 27 to April 29); outputs of the LLMs with search tools were collected later, between May 11 and May 13, 2025, so the two groups in the column were not sampled in the same weeks. ### Falls when A human-rated audit of the released statement-URL pairs finds support rates for the four deep research agents well below the printed 77.96 to 90.24, or finds the judge's 96% / 92% agreement does not hold on a larger sample than 100 pairs. It narrows if a re-run on later system versions moves the range. Query to run: human support audit of FACT statement-URL pairs from the DeepResearch Bench release, per system. ### Reflex Deep research agents mostly cite pages that do not back their sentences. Too coarse: under this benchmark's LLM judge, 78 to 90 percent of extracted statement-URL pairs were judged supported, while two search-tool models in the same table sit at 39 and 49 percent. ### Evidence https://arxiv.org/pdf/2506.11763v1 p. 6-7, Table 1 and Section 4.2.2; p. 18-19, Appendix C and E | 2025-06-13 · arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench ### Notes Revised in proposal 001 after a check against the HTML version of arXiv 2506.11763v1 (Table 1, Section 4.2.2, Appendix C, D, E). All numbers already on the card match the print. Two changes: the statement printed the search-tool column only as a range with two named values, which hid that most plain search-tool LLMs score at or above the deep research agents on C. Acc.; the full column is now printed. The collection window was given as one span (April 1 to May 13); Appendix D prints separate windows for the agents and for the search-tool LLMs, now stated. The paper does not say how the 100 pairs of the human comparison split between 'support' and 'not support'. The Reflex section still names only the two lowest search-tool values; left unchanged for the owner to decide. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another (The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another) Address: https://trustillery.com/entity/stocks/ai-citations--four-deep-research-agents-score-78-to-90-percent-citation-accuracy-on-deepresearch-bench-under-an-llm-judge ## Four generative search engines reach 40 to 68 percent citation accuracy and leave 23 to 47 percent of statements unsupported type: anchor · status: open · kind: measurement · quoted: You.com and Perplexity list slightly fewer sources (3.4–3.5) but still struggle with unsupported claims (23–47%). Finally, on citation metrics, all three engines show relatively low citation accuracy (40–68%), with frequent misattribution. · collected_by: Venkit, Zhou, Huang, Mao, Wu (Salesforce AI Research), Laban (Microsoft Research) · independence: positioned · source_family: self-published · as_of: 2025-09-02 ### Statement In the DeepTRACE audit, four public generative search engines were each run on 303 queries (168 debate questions from ProCon, 135 expertise questions contributed by participants of an earlier user study); results are as of 27 August 2025. Citation accuracy, defined as the fraction of statement citations for which the cited source's content supports the statement (overlap of citation matrix and factual-support matrix divided by the number of citations), was 68.3% for You.com, 65.8% for Bing Copilot, 49.0% for Perplexity and 39.8% for GPT-4.5. Unsupported statements, defined as the fraction of query-relevant statements not factually supported by any of the listed sources, cited or not, were 30.8%, 23.1%, 31.6% and 47.0% in the same order. Citation thoroughness (accurate citations over all possible accurate citations) was 20.5% to 24.4%. Support was decided per statement-source pair by an LLM judge (GPT-5 by default) over the scraped full text; on 100 manually verified tasks the judge correlated with human labels at Pearson 0.62. Link validity is not a reported metric: for roughly 15% of URLs the scraper returned an error (paywall or unavailable page) and these sources were excluded from the support calculations. ### Collection Authors: five at Salesforce AI Research, one at Microsoft Research; arXiv preprint, not peer reviewed. Microsoft is the vendor of one evaluated system (Bing Copilot), so independence is recorded as positioned; Salesforce is the vendor of none of the evaluated systems. Browser scripts extracted answer text, citations and source URLs from the public web interfaces; Jina Reader scraped the source text; an LLM judge decomposed answers into statements and filled the factual-support matrix (on the order of 80,000 support judgements). The support prompt in Appendix E returns full, partial or none, and the paper does not say how partial is turned into the binary matrix; Section 3 names GPT-5 as default judge while Appendix E names GPT-4. Counter-check that exists and was used: two hired annotators labelled 100 support tasks (Pearson 0.62 against the judge, called moderate agreement by the authors). No per-system human audit of citation accuracy is reported. The authors name reliance on an LLM judge as a limiting factor. The query set is announced as released. The Figure 2 caption and the running text speak of three engines while the score card prints four columns (You, Bing, PPLX, GPT 4.5). ### Falls when A human rating of the same engines' answers on the DeepTRACE queries puts citation accuracy of all four above 80%, or shows that the judge's error (Pearson 0.62 against human labels) is large enough to move the 39.8 to 68.3 range by more than ten points; or counting partial support as supported closes most of the gap. The anchor narrows if the result is driven by the 168 debate queries. Query to run: per-citation human support labels on a stratified sample of the DeepTRACE outputs, split by query type and by the binarisation of partial support. ### Reflex AI search engines cite their sources, so a reader can check any claim with one click. Too coarse: between about one third and three fifths of the citations measured here lead to a source that does not support the sentence it is attached to, and that is a different quantity from the share of sentences no listed source supports. ### Evidence https://arxiv.org/pdf/2509.04499v1 p. 8, Figure 2a and Section 4; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation | 2025-09-02 · arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 (Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--four-generative-search-engines-reach-40-to-68-percent-citation-accuracy-and-leave-23-to-47-percent-of-statements-unsupported ## Gemini Ultra reached 77 percent reference correctness against 54 percent for GPT-4 in medical research introductions type: anchor · status: open · kind: measurement · quoted: Gemini's references showed 77.2 % correctness and 68.0 % accuracy, compared to GPT-4's 54.0 % correctness and 49.2 % accuracy (p < 0.001 for both). · collected_by: Omar, Nassar, Hijazi, Glicksberg, Nadkarni, Klang (per the Europe PMC record's affiliation field: Division of Data-Driven and Digital Medicine, Icahn School of Medicine at Mount Sinai, New York; Maccabi Health Services, Israel; Edith Wolfson Medical Center, Holon, Israel) · independence: independent · source_family: peer-review · as_of: 2024-12-12 ### Statement In a comparison of OpenAI's GPT-4 and Google's Gemini Ultra writing medical research introductions with references across five medical fields, Gemini's references showed 77.2% correctness and 68.0% accuracy against 54.0% correctness and 49.2% accuracy for GPT-4 (p < 0.001 for both), and the abstract states that both models produced fabricated evidence. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. The abstract does not define how correctness differs from accuracy and gives no count of references. ### Collection Academic medical authors in the United States and Israel, published in the peer-reviewed journal Computers in Biology and Medicine; no stake in either vendor is visible in the record. Method per the abstract: the two models generated introductions in five medical fields, and the credibility and accuracy of the citations were assessed alongside introduction length and unreferenced statements. The observed text is the Europe PMC abstract record only: it gives no number of introductions or references, no definition of the two metrics, and does not say who verified the references or against which database. The counter-check that exists is the full text behind the DOI, which is not open access and was not read for this card. ### Falls when The full text of doi:10.1016/j.compbiomed.2024.109545 shows that correctness and accuracy include a judgment of whether the reference supports the sentence it is attached to, which would move this card out of the existence family, or shows a reference count too small to carry a 23.2 point difference. Query to run: read the Methods of the full article for the definitions of correctness and accuracy and the number of references per model. ### Reflex Gemini is better than GPT-4 at citing medical literature. Too coarse: the difference is 77.2 against 54.0 percent on one task at one point in time, both models produced fabricated references, and the quantity is reference correctness, not support of claims. ### Evidence https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:39667055%20AND%20SRC:MED&resultType=core&format=json abstract | 2024-12-12 · Computers in Biology and Medicine 185:109545 (2025), doi:10.1016/j.compbiomed.2024.109545 · Omar et al., Generating credible referenced medical research: A comparative study of openAI's GPT-4 and Google's gemini ### Notes Abstract only, and the number is thin: two percentages per model without denominators and without metric definitions. Confidence is therefore low. The article is dated 2025 by the journal (volume 185); the record's first publication date online is 2024-12-12, which is used as as_of. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) Address: https://trustillery.com/entity/stocks/ai-citations--gemini-ultra-reached-77-percent-reference-correctness-against-54-percent-for-gpt-4-in-medical-research-introductions ## GPT-3.5 and GPT-4 and Bard hallucinated about 40 and 29 and 91 percent of references generated for systematic reviews type: anchor · status: open · kind: measurement · quoted: Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard (P<.001). · collected_by: Chelli, Descamps, Lavoue, Trojani, Azar, Deckert, Raynier, Clowez, Boileau, Ruetsch-Chelli (per the Europe PMC record's affiliation field: Institute for Sports and Reconstructive Bone and Joint Surgery, Groupe Kantys, Nice; Hopital Lariboisiere, AP-HP, Paris; Universite Cote d'Azur, INSERM, C3M, Nice) · independence: independent · source_family: peer-review · as_of: 2024-05-22 ### Statement When GPT-3.5, GPT-4 and Bard were given the inclusion criteria of 11 published systematic reviews (33 prompts, 471 references analyzed, reviews pertaining to shoulder rotator cuff pathology as the abstract describes them), the share of generated references classed as hallucinated, meaning at least 2 of title, first author and year of publication were wrong, was 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4 and 91.4% (95/104) for Bard; precision against the original reviews' reference lists was 9.4% (13/139), 13.4% (16/119) and 0% (0/104). This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Clinician and academic authors in France, published in the peer-reviewed Journal of Medical Internet Research; no stake in the evaluated models is visible in the record. Method per the abstract: each model received the same inclusion criteria as the human reviewers, and its reference list was compared with the references of the original systematic review as gold standard; a paper counted as hallucinated if any 2 of title, first author, year were wrong. The observed text is the Europe PMC abstract record only: it does not state who checked the references, whether each was checked by more than one person, or whether any model had web access. The counter-check that exists is the full text and its supplementary reference lists; it was not read for this card. ### Falls when The full text of doi:10.2196/53164 prints denominators or a hallucination definition different from the abstract, or a re-check of the published per-reference lists against PubMed and Crossref finds that a material part of the 55, 34 and 95 references classed as hallucinated exist under a variant title or year. The anchor narrows rather than falls if the rates hold only for models without web access. Query to run: re-verify the supplementary reference lists of JMIR 26:e53164 against PubMed and Crossref. ### Reflex Chatbots make up about a third to a half of their references. Too coarse: the rate here spans 28.6 to 91.4 percent across three 2023-era models on one task, and it counts nonexistent or misdescribed papers, which says nothing about whether an existing cited source supports the sentence it is attached to. ### Evidence https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:38776130%20AND%20SRC:MED&resultType=core&format=json abstract | 2024-05-22 · Journal of Medical Internet Research 26:e53164, doi:10.2196/53164 · Chelli et al., Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews ### Notes Abstract only. The numbers carry their numerators and denominators in the abstract, which is why confidence is medium rather than low. The denominators 139, 119 and 104 sum to 362, not to the 471 references the abstract says were analyzed; the remaining 109 match the gold-standard reference count used as the recall denominator. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) Address: https://trustillery.com/entity/stocks/ai-citations--gpt-35-and-gpt-4-and-bard-hallucinated-about-40-and-29-and-91-percent-of-references-generated-for-systematic-reviews ## GPT-4o support judge agrees with a three-doctor consensus on 89 percent of 400 pairs and doctors among themselves on 86 percent type: anchor · status: open · kind: measurement · quoted: 88.7% agreement between the Source Veri fication model and the doctor consensus and an 86.1% average inter-doctor agreement rate · collected_by: Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou (Stanford University, Keck Medicine of USC, Loma Linda University School of Medicine) · independence: independent · source_family: peer-review · as_of: 2025-04-16 ### Statement On 400 statement-source pairs drawn from responses of GPT-4o with web search, GPT-4o API and Claude v2.1 API, the automated support judge (GPT-4o prompted as Source Verification model) agreed with the majority consensus of three US-licensed medical doctors in 88.7% of cases; the average agreement between doctors was 86.1%, and the difference between judge and consensus annotations was not statistically significant (p = 0.21, unpaired two-sided t-test). With Claude Sonnet 3.5 as judge, agreement with the consensus was 87.0% (95% CI 83.4-90.4), with Llama 3.1 70B 79.3% (75.4-83.1); GPT-4o and Claude Sonnet 3.5 agreed with each other on 90.1% (89.7-90.5) of decisions. In a further sample of 110 pairs from GPT-4o with web search that the judge had marked unsupported, doctors agreed 95.8% (91.8-98.7) of the time and confirmed 105 of 110. End to end on 100 HealthSearchQA questions, a clinician rated 40.4% (30.7, 50.1) of responses fully supported, the pipeline 42.4% (32.7, 52.2). ### Collection Authors are at Stanford University, Keck Medicine of USC and Loma Linda University School of Medicine; peer reviewed (Nature Communications); no competing interests declared. The annotating doctors are co-authors (author contributions list them under expert annotations) and investigators were not blinded. The Methods say the three doctors independently scored whether the LLM-generated source verification decision correctly identified a statement as supported or not supported, which reads as doctors seeing the judge's decision, while the Fig. 1 caption describes them as determining support themselves; the observed text does not resolve this. Agreement is raw percent agreement; no chance-corrected statistic is printed in the observed text, and the authors note the task is ambiguous given the lack of full agreement among the doctors. The expert annotations are released with the data, so the counter-check (recomputing agreement) exists and is open to anyone. ### Falls when A blinded re-annotation of the released 400 pairs by raters who do not see the judge's decision yields judge-consensus agreement materially below the inter-doctor level of 86.1%, or a chance-corrected statistic on the released annotations shows the agreement is driven by class imbalance. The anchor narrows if the validity holds only for medical web pages: run the same judge prompt against human labels in a non-medical attribution set. Query to run: Cohen's or Fleiss' kappa and the confusion matrix from the released expert annotations of doi 10.1038/s41467-025-58551-6. ### Reflex An LLM grading another LLM's citations is circular and cannot be trusted. Too coarse: on 400 medical pairs the judge agreed with a doctor consensus slightly more often than the doctors agreed with one another, which places the judge's error near the level of human disagreement rather than making it arbitrary. ### Evidence https://www.nature.com/articles/s41467-025-58551-6.pdf p. 2, sections Source verification and Evaluation of bias of GPT-4o as the backbone LLM; p. 3, Fig. 1a; p. 8, Expert validation | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions (Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions) Address: https://trustillery.com/entity/stocks/ai-citations--gpt-4o-support-judge-agrees-with-a-three-doctor-consensus-on-89-percent-of-400-pairs-and-doctors-among-themselves-on-86-percent ## GPT-4o with RAG on 300 health questions has all URLs valid but 76 percent of statements and 38 percent of responses supported type: anchor · status: open · kind: measurement · quoted: support of 75.7% (74.0– 77.2 95% CI), and response-level support of 38.4% (26.7, 49.3 95% CI). · collected_by: Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou (Stanford University, Keck Medicine of USC, Loma Linda University School of Medicine) · independence: independent · source_family: peer-review · as_of: 2025-04-16 ### Statement On a random subset of 300 consumer health questions from HealthSearchQA, GPT-4o with web search (RAG) returned citation URLs that were valid in 100% of cases, while statement-level support (share of parsed medical statements supported by at least one source given in the same response) was 75.7% (95% CI 74.0-77.2) and response-level support (share of responses in which every statement is supported) was 38.4% (95% CI 26.7-49.3). On the main set of 800 questions the same system's response-level support is printed as 55%: close to 80% on the 400 questions generated from Mayo Clinic pages and 31.0% (26.7, 35.8) on the 400 questions from Reddit r/AskDocs. Support was decided per statement-source pair by an automated GPT-4o judge; on 100 HealthSearchQA questions a human clinician rated 40.4% (30.7, 50.1) of the responses fully supported against 42.4% (32.7, 52.2) by the pipeline. The GPT-4o API endpoint named is gpt-4o-2024-05-13; the paper was received 30 September 2024. ### Collection Authors are at Stanford University (Biomedical Data Science, Electrical Engineering, Computer Science, Genetics, Anesthesiology, Law School), Keck Medicine of USC and Loma Linda University School of Medicine; the paper is peer reviewed (Nature Communications), the authors declare no competing interests and are not the vendor of any evaluated model. Pipeline SourceCheckup: questions are generated by GPT-4o from Mayo Clinic pages or taken from Reddit r/AskDocs, each evaluated LLM answers and lists sources, GPT-4o parses the response into statements, each URL is downloaded (valid = status code 200 with non-empty text), and GPT-4o as Source Verification model judges every statement-source pair. Because the models' intended pairing of statement and citation could not be recovered, a statement counts as supported if any source in the response supports it, which is more lenient than checking the citation attached to the claim. Counter-checks that exist and were used: three US-licensed doctors on 400 pairs (88.7% agreement of the judge with their consensus), a second judge model (Claude Sonnet 3.5), and the end-to-end clinician rating of 100 responses. The annotating doctors are co-authors and investigators were not blinded. Data and code are public (github.com/kevinwu23/SourceCheckup). ### Falls when A blinded human rating of the released GPT-4o (RAG) responses on the HealthSearchQA subset puts statement-level support outside the printed interval 74.0-77.2 or response-level support outside 26.7-49.3; or a rerun of the released pipeline on current web-search assistants over the same questions finds response-level support within a few points of URL validity. The anchor narrows rather than falls if the gap is specific to open-ended consumer questions: the paper itself prints close to 80% response-level support on Mayo Clinic derived questions. Query to run: per-citation (not any-source) human support rating on the released statement-source pairs of doi 10.1038/s41467-025-58551-6. ### Reflex An assistant that searches the web and returns working links to reputable health sites has sourced its answer. Too coarse: here every URL resolved, yet about a quarter of the statements and about six in ten whole responses were not backed by the returned sources. ### Evidence https://www.nature.com/articles/s41467-025-58551-6.pdf p. 4, sections Additional validation on HealthSearchQA and End-to-end full human evaluation; p. 8, metric definitions | 2025-04-16 · Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--counting-per-response-instead-of-per-statement-halves-the-support-rate-in-the-same-data (Counting per response instead of per statement halves the support rate in the same data) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--two-support-rates-in-the-stock-count-a-claim-as-supported-if-any-cited-page-supports-it-and-bound-the-share-of-claims-supported-by-their-attached-citation (Two support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support (Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-study-measures-both-link-validity-sits-about-22-to-52-points-above-claim-support-in-the-printed-pairs (Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs) Address: https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported ## Grok 3 cited URLs leading to error pages in 154 of 200 source identification prompts in a Tow Center test of February 2025 type: anchor · status: open · kind: measurement · quoted: More than half of responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages. Out of the 200 prompts we tested for Grok 3, 154 citations led to error pages. · collected_by: Klaudia Jaźwińska and Aisvarya Chandrasekar, Tow Center for Digital Journalism at Columbia's Graduate School of Journalism, published in Columbia Journalism Review · independence: positioned · source_family: established-media · as_of: 2025-03-06 ### Statement In the Tow Center test of eight generative search tools (tests conducted February 2025; 200 prompts per tool, each asking for headline, publisher, date and URL of the article a given excerpt came from), more than half of the responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages; for Grok 3, 154 citations out of the 200 prompts led to error pages. Grok 2 was prone to linking to the publisher's homepage rather than the specific article. The article says the problem happened 'far less frequently' with the other tools and prints no counts for them or for Gemini. This is a link-resolution measurement (does the cited URL lead to a page), a precondition of support and not support itself; determined manually by the researchers, each prompt run once. ### Collection The Tow Center is a university research center at Columbia's Graduate School of Journalism and a partner of CJR, a journalism trade publication. It has no commercial stake in any tested tool, but the study is framed from the news publishers' side (referral traffic, attribution, crawler control) and quotes publishers as affected parties; recorded as positioned for that reason. The researchers followed the cited URLs manually; the article does not separate fabricated URLs from URLs that once existed and broke, nor say when after generation the links were followed. Counter-check that exists: the article offers its data for download, so the URLs can be re-tested; xAI and Google did not respond to the request for comment. ### Falls when Falls if the released data show that a large part of the 154 error-page links resolved at query time, or returned bot-block or paywall errors to the researchers' browser rather than being non-existent. Narrows if 'fabricated' cannot be separated from 'broken': query each of the 154 URLs against the Wayback Machine and the publisher's sitemap to see whether the path ever existed. Narrows in time if later versions of the same tools, given the same prompts, return resolving links. ### Reflex A link in an AI answer at least leads somewhere. Too coarse: for one tool 154 of 200 prompts ended in citations to error pages, while for most other tools in the same test this was far less frequent. ### Evidence https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php section 'Platforms often failed to link back to the original source' | 2025-03-06 · Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, AI Search Has a Citation Problem Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--reverse-attribution-tests-and-the-vendor-primary-source-axis-measure-neighbouring-quantities-and-not-citation-support (Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support) Address: https://trustillery.com/entity/stocks/ai-citations--grok-3-cited-urls-leading-to-error-pages-in-154-of-200-source-identification-prompts-in-a-tow-center-test-of-february-2025 ## Groundedness of 31 to 68 percent on ResearcherBench means 32 to 69 percent of factual claims carry no citation type: anchor · status: open · kind: measurement · quoted: OpenAI Deep Research, which achieves the best performance in rubric assessment, only attains a low Groundedness score of 0.34. Conversely, Perplexity Sonar Reasoning Pro, which achieves the highest Groundedness score of 0.68 · collected_by: Xu, Lu, Ye, Hu, Liu (Shanghai Jiao Tong University; SII; GAIR) · independence: independent · source_family: self-published · as_of: 2025-07-22 ### Statement Groundedness is the number of factual claims in a report that carry a citation URL divided by all extracted factual claims; it says nothing about whether the citation supports the claim. Table 2 prints: Perplexity: Sonar Reasoning Pro 0.68, Gemini Deep Research 0.59, Perplexity Deep Research 0.56, GPT-4o Search Preview 0.39, OpenAI Deep Research 0.34, Grok3 DeepSearch 0.32, Grok3 DeeperSearch 0.31. The tasks are 65 research questions on frontier AI topics, evaluated between March and April 2025; claims and their citation links are extracted by GPT-4.1, not by human raters. ### Collection Authors are at Shanghai Jiao Tong University, SII and GAIR; none of the evaluated systems is theirs; arXiv preprint without venue. Method: on 65 research questions from frontier AI research, GPT-4.1 extracts all factual claims with context and any citation URL from each report, the cited page is fetched through the Jina Reader API, and GPT-4.1 as judge returns a binary yes or no on whether the page supports the claim. Evaluations ran between March and April 2025. No human rates the claims. The human meta-evaluation the paper reports (10 responses, Table 3) covers the rubric assessment judge; a human check of the citation-support judge is not reported, so the counter-check for this metric exists only in principle and was not used. ### Falls when A human extraction of factual claims from the same reports yields citation coverage far from the printed 0.31 to 0.68, for instance because the extractor counts reasoning or summary sentences as factual claims or misses citations given at paragraph level. Query to run: human count of cited versus uncited factual claims on a sample of ResearcherBench reports, per system. ### Reflex If the citations in a report check out, the report is sourced. Too coarse: the supported share is computed over cited claims only, and here 32 to 69 percent of extracted factual claims carried no citation at all. ### Evidence https://arxiv.org/pdf/2507.16280v1 p. 8, Table 2 (Section 5.2) and Section 5.3 Key Findings, Finding 2; p. 6-7, Section 4.2 | 2025-07-22 · arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--high-faithfulness-of-cited-claims-coexists-with-most-claims-carrying-no-citation-and-with-a-25-fold-spread-in-citation-volume (High faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume) Address: https://trustillery.com/entity/stocks/ai-citations--groundedness-of-31-to-68-percent-on-researcherbench-means-32-to-69-percent-of-factual-claims-carry-no-citation ## Human raters find citation precision of 89 and 84 percent for LongCite models against 68 for GLM-4 on LongBench-Chat type: anchor · status: open · kind: measurement · quoted: GLM-4 61.2 67.5 60.247.6 53.9 47.146.1 29.1 30.8 LongCite-8B79.6 88.9 82.662.0 79.7 67.459.639.542.0 LongCite-9B72.884.275.857.678.163.664.2 45.1 47.1 Table 6: Citation quality evaluated by human · collected_by: Zhang, Bai, Lv, Gu, Liu, Zou, Cao, Hou, Dong, Feng, Li (Tsinghua University; Zhipu AI) · independence: positioned · source_family: peer-review · as_of: 2024-09-10 ### Statement In the paper's human evaluation, the anonymized responses of three models on the 50 LongBench-Chat queries (150 responses, 1,064 statements, 909 citations) were manually annotated for citation recall and citation precision by the same standard the GPT-4o judge uses. Table 6 prints human scores as recall / precision / F1: LongCite-8B 79.6 / 88.9 / 82.6, LongCite-9B 72.8 / 84.2 / 75.8, GLM-4 61.2 / 67.5 / 60.2. The GPT-4o judge scores for the same responses are lower: 62.0 / 79.7 / 67.4, 57.6 / 78.1 / 63.6 and 47.6 / 53.9 / 47.1. Precision here is the share of cited snippets that at least partially support their statement; recall is whether a statement is supported by its cited snippets. The setting is citation into a document supplied in the prompt, not web research. ### Collection Authors are at Tsinghua University and Zhipu AI. They built the benchmark (LongBench-Cite), the training data (LongCite-45k) and the two LongCite models that lead the table, and Zhipu AI is the developer of GLM-4 and of the GLM-4-9B base of LongCite-9B: collector and proposer of the winning method coincide, recorded here as positioned. The observed arXiv v3 text is marked Preprint; the source list gives Findings of ACL 2025 as venue. Method: the model receives the long context with numbered sentences and must answer with sentence-level citations into that supplied context; there is no web retrieval. GPT-4o judges citation recall (statement fully, partially or not supported by its cited snippets: 1 / 0.5 / 0) and citation precision (each cited snippet relevant or not). The counter-check that exists and was used: a human annotation of 150 LongBench-Chat responses (1,064 statements, 909 citations) from three models; Cohen's kappa between GPT-4o and human 0.593 for recall and 0.655 for precision, GPT-4o accuracy against human labels 75.0% and 88.8%. For this anchor the human annotation is itself the measurement; the paper does not say who the annotators were or report agreement between human annotators. ### Falls when An independent human annotation of the same 150 responses gives precision for the LongCite models near or below the 67.5 of GLM-4, or the annotators turn out to be the model developers without blinding beyond anonymized model names. The anchor narrows if 'at least partially supports' is tightened to full support and the precision gap shrinks. Query to run: second human annotation of the released LongBench-Chat responses with reported inter-annotator agreement. ### Reflex Only an LLM judge says citations are good, humans would find them worse. Too coarse: in this closed-document setting the human scores sit above the GPT-4o judge scores for all three models, by 6 to 14 points in precision. ### Evidence https://arxiv.org/pdf/2409.02897v3 p. 10, Section 4.3, Table 6 and Table 7 | 2024-09-10 (v3; first submitted 2024-09-04) · arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: table-only quote, Table 6, p. 10, Human column P; no sentence carries 88.9, 84.2 or 67.5. Plan 02b: the caption is cut before 'GPT-4o' because the engine's canonical form folds the line-break hyphen of that token. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-direction-of-llm-judge-error-on-citation-support-is-unsettled-with-two-measurements-showing-stricter-judges-and-one-a-more-generous-judge (The direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research (Three conditions are each measured on one component of citation quality and none on claim support in web research) Address: https://trustillery.com/entity/stocks/ai-citations--human-raters-find-citation-precision-of-89-and-84-percent-for-longcite-models-against-68-for-glm-4-on-longbench-chat ## In early 2023 four generative search engines had 51.5 percent citation recall and 74.5 percent citation precision type: anchor · status: open · kind: measurement · quoted: a mere 51.5% of generated statements are fully supported with citations (recall), and only 74.5% of citations fully support their associated statements (precision) · collected_by: Liu, Zhang, Liang (Computer Science Department, Stanford University) · independence: independent · source_family: peer-review · as_of: 2023-04-19 ### Statement A human-rated audit of four commercial generative search engines as they stood in early 2023 (Bing Chat, NeevaAI, perplexity.ai, YouChat; responses scraped between late February and late March 2023) on 1450 queries per system measured two quantities, averaged over the four systems. Citation recall, defined in the paper as the proportion of verification-worthy statements that are fully supported by their associated citations, was 51.5%. Citation precision, defined as the proportion of generated citations that support their associated statements, was 74.5%. The precision formula counts a citation that fully supports its statement and also a citation that partially supports it when the union of the statement's citations gives full support and no single citation does; the results sentence abbreviates this as citations that fully support. The averages are unweighted means of four system values (Tables 7 and 8), and the paper separates verifiability from factual correctness. This is a 2023 baseline that predates the stock's 2024 to 2026 window. ### Collection Nelson F. Liu, Tianyi Zhang and Percy Liang, Computer Science Department, Stanford University; published in Findings of EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 23 Oct 2023). Human-rated, not automatic: 34 annotators recruited on Amazon Mechanical Turk, pre-screened with a qualification study, judged each query-response pair in three steps (verification-worthy statements, support of each statement by its citations, support contributed by each citation). Each system was run on 1450 queries (AllSouls, davinci-debate, ELI5 KILT and Live, WikiHowKeywords, seven NaturalQuestions subdistributions); responses were scraped between late February and late March 2023. Each pair was annotated once; the counter-check that exists and was used is a triple annotation of 250 randomly sampled pairs, with more than 82.0% pairwise agreement and 91.0 F1 for all judgments. The authors are academic researchers and not vendors of any evaluated system; the paper acknowledges Amazon Web Services for Mechanical Turk credits and the AI2050 program at Schmidt Futures. The human annotations are released. ### Falls when A re-computation from the released human annotations of arXiv 2304.09848, taken as the unweighted mean of the four per-system values in Tables 7 and 8, gives citation recall outside 51.0 to 52.0% or citation precision outside 74.0 to 75.0%. It also falls if an independent human re-annotation of a random sample of at least 250 of the released query-response pairs, under the paper's own recall and precision definitions, yields a sample recall or precision more than 5 points away from the value the original annotations give on the same pairs, or if a full second annotation of the early-2023 responses moves either average by three points or more. The anchor narrows if it is read as a rate for later systems: query to run is a human-rated audit of the same engines with the same recall and precision definitions in 2024 to 2026. ### Reflex Search assistants that attach citations let the reader verify what they say. Too coarse: in this human audit about half of the generated statements were not fully supported by their own citations and about a quarter of citations did not support their statement. ### Evidence https://arxiv.org/pdf/2304.09848v2 p. 7, Section 4.2, with definitions p. 3-4, Sections 2.3-2.4, and Tables 7-8, p. 23-24 | 2023-04-19 · Findings of EMNLP 2023, arXiv 2304.09848 v2 · Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 (Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023) Address: https://trustillery.com/entity/stocks/ai-citations--in-early-2023-four-generative-search-engines-had-515-percent-citation-recall-and-745-percent-citation-precision ## Invalid citations appear in about 1 percent of 56381 AI and security conference papers and rose 81 percent in 2025 type: anchor · status: open · kind: measurement · quoted: identifying 739 invalid citations across 604 papers (1.07% of 56,381) are definitively invalid, with an 80.9% increase in invalid citation rates in 2025 (from a 2020–2024 average of 0.89% to 1.61%) · collected_by: Xu, Qiu, Sun, Miao, Wu, Li and colleagues, Nankai University and Tsinghua University · independence: independent · source_family: self-published · as_of: 2026-05-14 ### Statement In 2,199,409 citations extracted from 56,381 papers accepted at eight AI/ML and security venues (NeurIPS, ICML, AAAI, IJCAI, USENIX Security, CCS, S&P, NDSS) from 2020 to 2025, an automated check flagged 2,530 citations, manual review confirmed 739 as invalid (136 with wrong metadata, 603 untraceable), and 604 papers (1.07%) contained at least one invalid citation; the yearly share of such papers stayed between 0.76% and 0.98% from 2020 to 2024 and was 1.61% in 2025, 80.9% above the 2020-2024 average of 0.89%. This measures the published record written by researchers, whose use of an AI assistant for any given paper is not established, and the authors state that their data alone does not establish causality. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Academic authors at Nankai University and Tsinghua University; arXiv preprint (cs.CR), version 2; the authors publish at some of the audited venues and built the screening tool, and no other stake is visible. Method: CiteVerifier flagged citations with title similarity below 0.9; sixteen trained research assistants reviewed every flagged citation, each checked independently at least twice; two researchers then reviewed all citations classed invalid. The counter-check that exists was used: 400 citations sampled from the valid pool were reviewed by hand and no further invalid citation was found. The share is a lower bound of what the screening could see, since only flagged citations were reviewed. ### Falls when A re-check of the 603 untraceable citations finds a material part in sources the reviewers did not search (theses, non-English venues, withdrawn preprints), or the 2025 share returns to the 0.76 to 0.98 percent band once the 2025 proceedings are complete and re-extracted. Query to run: re-verify the released list of 739 invalid citations from arXiv 2602.06718 and recompute the per-year paper share. ### Reflex AI-written papers are flooding conferences with fake references. Too coarse: about one paper in a hundred at eight top venues carries at least one invalid citation, the pre-LLM years 2020 to 2022 already sit near 0.9 percent, and the 2025 rise to 1.61 percent is not attributed to AI use by the measurement itself. ### Evidence https://arxiv.org/pdf/2602.06718v2 p. 6, Section IV.B; p. 9-10, Section VI, Tables IV and V, Figure 5 | 2026-05-14 (v2; v1 2026-02-06) · arXiv 2602.06718 · Xu et al., GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--non-existent-citations-reach-the-published-record-at-about-1-percent-of-papers-and-rising-without-identifying-the-tool (Non-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool) Address: https://trustillery.com/entity/stocks/ai-citations--invalid-citations-appear-in-about-1-percent-of-56381-ai-and-security-conference-papers-and-rose-81-percent-in-2025 ## Journalists at 22 public media rated 31 percent of AI assistant news responses as having significant sourcing issues type: anchor · status: open · kind: measurement · quoted: Sourcing was the biggest cause of problems, with 31% of all responses having significant issues with sourcing – this includes information in the response not supported by the cited source, providing no sources at all, or making incorrect or unverifiable sourcing claims. · collected_by: James Fletcher (BBC) and Dorien Verckist (EBU), EBU Media Intelligence Service with the BBC; ratings by journalists of 22 participating public service media organizations · independence: positioned · source_family: established-media · as_of: 2025-10-22 ### Statement In the EBU/BBC study, 271 journalists from 22 public service media organizations in 18 countries rated 2,709 responses to 30 core news questions, generated between 24 May and 10 June 2025 in 14 languages by the free consumer versions of ChatGPT, Copilot, Gemini and Perplexity. On the sourcing criterion (Q2: 'Are the claims in the response supported by the sources the assistant provides?') 31% of responses were rated as having significant issues: Gemini 72%, ChatGPT 24%, Perplexity 15%, Copilot 15% (n = 675, 678, 681, 675; the appendix tables print 483, 160, 101 and 104 significant ratings). The category is wider than unsupported claims: it covers information in the response not supported by the cited source, providing no sources at all, and incorrect or unverifiable sourcing claims. The appendix states that significant issues for Q2 include the lack of any direct sourcing, and 42% of Gemini responses provided no direct sources. The unit is the response, not the individual citation; the rating is a human judgement on a four-level scale (no issues, some issues, significant issues, don't know). ### Collection Produced by the EBU Media Intelligence Service and the BBC. Each of the 22 participating public service media organizations generated responses with the prompt prefix 'Use [organization] sources where possible', lifted its technical crawler blocks for the generation period, and had its own journalists rate the anonymized responses after a briefing with written and video calibration material. The EBU answers to its member broadcasters. The participating organizations are publishers whose content the assistants use, and the report calls for publisher control over content use, agreed citation formats and regulatory attention: a stake and a declared position, recorded as positioned. Counter-check that exists and was used: project teams in each organization checked all significant-issue ratings for whether they were clearly evidenced and correctly classified and checked that sourcing issues were logged correctly, and the central team ran an additional QA pass; participant organizations remain responsible for their own data. No inter-rater agreement statistic is reported for this round, and the report does not publish the response-level data. ### Falls when The anchor narrows if the 31% is decomposed: query the study's response-level data for the share of significant Q2 ratings that rest on 'no direct source' or on unverifiable sourcing claims rather than on a cited source that does not contain the claim; if the unsupported-claim share alone is far lower, above all outside Gemini, the figure cannot be read as a citation-support rate. It falls if a re-rating of the same responses by raters without a publisher affiliation yields a materially different rate or a different ordering of assistants. It narrows in time if a repeat with paid tiers or later default models (the report lists GPT-5 as ChatGPT's default by 16 Oct 2025) no longer shows the 15 to 72 percent spread. ### Reflex Assistants with web search cite their sources, so their news answers can be checked. Too coarse: journalists rated 31 percent of responses as having significant sourcing issues, from 15 to 72 percent depending on the assistant, and the category includes answers with no source at all. ### Evidence https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 10, High-level findings | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q2 note and counts) | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants Relationships - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-31-percent-sourcing-figure-of-the-ebu-audit-counts-responses-and-includes-absent-sources-and-is-not-a-per-citation-support-rate (The 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate) Address: https://trustillery.com/entity/stocks/ai-citations--journalists-at-22-public-media-rated-31-percent-of-ai-assistant-news-responses-as-having-significant-sourcing-issues ## LLM citation judges reject 18 to 47 percent of genuinely supported citations on a human-reviewed benchmark type: anchor · status: open · kind: measurement · quoted: On factual support, false negative rates vary from 0.183 (GPT-5.4-mini) to 0.470 (GPT-OSS-120B), indicating that most judges reject a substantial fraction of genuinely supported citations. · collected_by: Leung, Lumer, Feld, Huber, Subbiah, Paul (Commercial Technology and Innovation Office, PricewaterhouseCoopers, U.S.) · independence: unknown · source_family: self-published · as_of: 2026-07-09 ### Statement Against human-reviewed gold labels on 624 attribution-citation pairs, the false negative rate of eight LLM judges on factual support (a supported citation scored as unsupported) ranged from 0.183 (GPT-5.4-mini) to 0.470 (GPT-OSS-120B). Rejection of adversarially edited claims was high for all judges, from roughly 86% for the subtlest edit strategies to near 100% for negations and semantic drift. Three of eight judges (GPT-5.4-mini, Claude Haiku 4.5, Gemini 3.1 Flash Lite) passed more pairs than the 18.4% gold pass rate on factual support, five passed fewer; on source relevance all eight passed fewer than the 79.3% gold rate (42.9% to 72.0%). The paper locates the dominant factual-support error in over-rejection of supported citations, not in acceptance of edited ones. False positive rates per judge are shown only in a figure, not printed as numbers. ### Collection Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. Four of the six authors (Lumer, Feld, Huber, Subbiah) are also authors of arXiv 2605.06635, whose citation-evaluation pipeline this paper reuses and whose judge design it examines, so this is a check from inside the same group, not an outside replication. The benchmark is one synthetic long-form report over 25 topics in which about 60% of attributed claims were adversarially edited (19 strategies); parsing yields 624 attribution-citation pairs, each judged on source relevance and factual support (1,248 decisions). Gold labels come from a council of 6 LLM judges; a human reviewer examined all decisions, confirmed the 870 unanimous ones on review and adjudicated the 378 non-unanimous ones (263 relevance, 115 factual support). The text speaks of 'a human reviewer'; no inter-human agreement is reported. Two of the 8 evaluated judges (GPT-5-mini, Claude Opus 4.6) also sat on the labelling council. The gold pass rate on factual support is 18.4%, set low by design. The authors state the findings are limited to a single adversarial document. The direction of bias was measured where 81.6% of pairs are gold-unsupported; whether the same strictness holds on natural assistant output with a higher support rate is not tested. Counter-check that exists: the same false-negative measurement on human-labelled citations from real deep-research reports; not reported. ### Falls when A human-labelled sample of citations from deployed assistants or deep-research agents shows LLM judges with a false negative rate on factual support below 0.10, or shows the error running the other way (judges accepting unsupported citations more often than rejecting supported ones). Query to run: per-judge FNR and FPR on citation support against multi-annotator human labels on natural outputs. ### Reflex If an LLM judge errs on citation support, it errs toward leniency and inflates the support rate. Too coarse: on this benchmark most judges err toward strictness, rejecting 18 to 47 percent of supported citations, which would push an LLM-judged support rate down rather than up. ### Evidence https://arxiv.org/pdf/2607.08700v1 p. 9-10, Sections 5.1 and 5.2 | 2026-07-09 · arXiv 2607.08700 · Leung et al., Do You Need a Frontier Model as a Citation Verifier? Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--the-direction-of-llm-judge-error-on-citation-support-is-unsettled-with-two-measurements-showing-stricter-judges-and-one-a-more-generous-judge (The direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge) Address: https://trustillery.com/entity/stocks/ai-citations--llm-citation-judges-reject-18-to-47-percent-of-genuinely-supported-citations-on-a-human-reviewed-benchmark ## Of 1053 AI assistant news responses with direct quotes 12 percent had significant quote accuracy issues per journalists type: anchor · status: open · kind: measurement · quoted: Across all AI assistant responses that included a direct quote (a total of 1,053), 12% were found to have significant issues with the accuracy of those direct quotes. · collected_by: James Fletcher (BBC) and Dorien Verckist (EBU), EBU Media Intelligence Service with the BBC; ratings by journalists of 22 participating public service media organizations · independence: positioned · source_family: established-media · as_of: 2025-10-22 ### Statement In the EBU/BBC study (22 public service media organizations, 18 countries, responses generated 24 May to 10 June 2025 by the free consumer versions of ChatGPT, Copilot, Gemini and Perplexity), journalists answered for each response the question 'Do any direct quotes in the response accurately reflect the source cited for them?' (Q3). Of the 1,053 responses to core questions that included a direct quote, 12% were rated as having significant issues with the accuracy of those quotes. Gemini: 20% of 290 responses with quotes; Copilot: 4% (n=190), the fewest; ChatGPT n=262 and Perplexity n=311, whose percentages appear only in a chart (the appendix tables print 28 and 33 significant ratings for them, 59 for Gemini and 8 for Copilot). Reported cases include quotes not found in the source provided for them and quotes with altered wording. The unit is the response containing at least one direct quote, not the individual quote; rated by human journalists. ### Collection Produced by the EBU Media Intelligence Service and the BBC. Each of the 22 participating public service media organizations generated responses with the prompt prefix 'Use [organization] sources where possible', lifted its technical crawler blocks for the generation period, and had its own journalists rate the anonymized responses after a briefing with written and video calibration material. The EBU answers to its member broadcasters. The participating organizations are publishers whose content the assistants use, and the report calls for publisher control over content use, agreed citation formats and regulatory attention: a stake and a declared position, recorded as positioned. Counter-check that exists and was used: project teams in each organization checked all significant-issue ratings for evidence and classification, and the central team ran an additional QA pass. Quote accuracy is checkable by anyone against the cited page, but the report prints only illustrative cases, not the rated responses; no inter-rater agreement statistic is reported. ### Falls when Falls if the quotes in the 1,053 responses are compared string by string with the pages cited for them (in the version live at generation time) and the share of responses with an altered or absent quote differs materially from 12 percent. Narrows if a large part of the significant ratings concern translated quotes, where the assistant rendered a source-language quote into the response language and the evaluator counted the rendering as alteration: query the significant Q3 ratings by language and by whether source and response language differ. ### Reflex Text inside quotation marks with a source link is a verbatim quote from that source. Too coarse: in 12 percent of quote-bearing responses journalists found significant problems, including quotes that are not in the cited source at all. ### Evidence https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 21, Accuracy of direct quotes | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q3 counts) | 2025-10-22 · EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--two-journalist-audits-find-a-quote-accuracy-problem-in-about-one-in-eight-quote-bearing-news-responses (Two journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses) Address: https://trustillery.com/entity/stocks/ai-citations--of-1053-ai-assistant-news-responses-with-direct-quotes-12-percent-had-significant-quote-accuracy-issues-per-journalists ## Of 400 references from eight free chatbots about 27 percent were fully correct and 40 percent erroneous or fabricated type: anchor · status: open · kind: measurement · quoted: From the total dataset of 400 references analyzed, 26.5% were real and fully accurate (i.e., all five bibliographic elements were correct), while 33.8% were real but only partially correct (e.g., containing errors in the publication year or locating data). In contrast, 39.8% of the references were either incorrect or entirely fabricated by the AI systems. · collected_by: Alvaro Cabezas-Clavijo and Pavel Sidorenko-Bautista, Universidad Internacional de La Rioja (UNIR), Logrono, Spain · independence: independent · source_family: self-published · as_of: 2025-05-23 ### Statement Eight chatbots in their free versions (ChatGPT on GPT-4o-mini, Claude 3.5 Sonnet, Copilot, DeepSeek-V3, Gemini Flash 2.0, Grok-3, Le Chat, Perplexity Sonar) were each asked between 7 and 9 February 2025, with one standardized student prompt, for 10 academic references in APA format in each of five disciplines, giving 400 references; checked by hand on five bibliographic elements (authors, year, title, venue, locating data), 26.5% were real and fully accurate, 33.8% real but partially correct, and 39.8% incorrect or entirely fabricated, with Grok and DeepSeek fabricating none of their 50 references and Copilot, Perplexity and Claude fabricating 100%, 72% and 64%. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Two academic authors at a Spanish university; arXiv preprint that states on its first page it has not undergone peer review. No stake in any evaluated chatbot is declared or visible. Sample: 50 references per chatbot, 80 per knowledge area, one prompt per discipline, one run. Rater: the authors, by manual searches on Google and Google Scholar with the chatbot's title in quotation marks; each reference scored 0 to 5 errors. The paper reports no second coder and no inter-rater agreement. The counter-check that exists is re-running the five printed prompts and re-searching the titles; the paper does not report having done a second pass. ### Falls when A second coder re-searching the 400 references in Crossref, WorldCat and Google Scholar moves more than a few points between the three classes, in particular if references classed as fabricated turn out to exist under another edition or title variant, or a rerun of the five prompts printed in Table 2 on the same free versions gives per-chatbot fabrication rates far from 0 to 100 percent as printed. Query to run: re-verify the reference list behind Figure 1 of arXiv 2505.18059. ### Reflex Newer chatbots no longer invent references the way early ChatGPT did. Too coarse: in February 2025 six of eight free chatbots still fabricated references, including Perplexity at 72 percent and Copilot at 100 percent, and this counts existence and bibliographic correctness only, not whether a source supports a claim. ### Evidence https://arxiv.org/pdf/2505.18059v1 p. 8, Results and Figure 1 | 2025-05-23 · arXiv 2505.18059 · Cabezas-Clavijo and Sidorenko-Bautista, Assessing the performance of 8 AI chatbots in bibliographic reference retrieval Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) Address: https://trustillery.com/entity/stocks/ai-citations--of-400-references-from-eight-free-chatbots-about-27-percent-were-fully-correct-and-40-percent-erroneous-or-fabricated ## On ELI5 in 2023 ChatGPT and GPT-4 baselines reach about 50 percent automatic citation recall and precision type: anchor · status: open · kind: measurement · quoted: on the ELI5 dataset, around 50% generations of our ChatGPT and GPT-4 baselines are not fully supported by the cited passages · collected_by: Gao, Yen, Yu, Chen (Department of Computer Science and Princeton Language and Intelligence, Princeton University) · independence: independent · source_family: peer-review · as_of: 2023-10-31 ### Statement On the ELI5 part of the ALCE benchmark (1,000 randomly selected development questions, retrieval from the Sphere web corpus cut into 100-word passages), citation quality of the authors' retrieve-and-prompt systems built on the models of 2023 was scored automatically by an NLI model (TRUE, a T5-11B model), not by human raters. Table 6 prints citation recall / citation precision of 51.1 / 50.0 for ChatGPT VANILLA (5 passages), 44.0 / 50.1 for GPT-4 with 5 passages, 48.5 / 53.4 for GPT-4 with 20 passages, 38.3 / 37.9 for LLaMA-2-Chat-70B, and 69.3 / 67.8 for ChatGPT with RERANK, a strategy that reranks sampled generations by the same automatic citation recall. Recall here is the share of statements whose concatenated cited passages entail the statement; precision is the share of citations not detected as irrelevant, given recall of 1. The paper summarises this as around 50% of generations of its ChatGPT and GPT-4 baselines not being fully supported by the cited passages. The systems are research pipelines over a fixed corpus, not deployed assistants citing live web pages. This is a 2023 baseline that predates the stock's 2024 to 2026 window. ### Collection Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen, Department of Computer Science and Princeton Language and Intelligence, Princeton University; published at EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 31 Oct 2023). Automatic, not human-rated: each statement and the concatenation of its cited passages go to the TRUE NLI model, which returns entailment or not; scores of the open models are averaged over three seeded runs; Appendix G.6 states that ChatGPT-16K and GPT-4 use one seeded run each, so the GPT-4 figures are single-run values. The authors built the evaluated pipelines themselves but are not vendors of the underlying models; the paper acknowledges an IBM PhD Fellowship, an NSF CAREER award, a Sloan Research Fellowship and Microsoft Azure credits. The counter-check that exists and was used is a human evaluation through Surge AI on sampled outputs of three systems: on ELI5 human raters gave ChatGPT VANILLA 50.8 recall and 52.4 precision against 52.8 and 50.4 from ALCE (Table 9). The paper states that the NLI model cannot detect partial support and therefore gives a lower citation precision score than human evaluation. ### Falls when A human-rated evaluation of the same ELI5 outputs gives citation recall or precision for the ChatGPT or GPT-4 VANILLA systems that departs from the automatic 44 to 53 percent band by more than a few points, or a different entailment model applied to the released outputs moves the band. The anchor narrows if read as a rate for deployed assistants: query to run is the ALCE metrics, or a human audit with the same definitions, on the outputs of 2024 to 2026 systems for the same 1,000 ELI5 questions. ### Reflex If the model is given the retrieved passages and told to cite them, its citations will back what it writes. Too coarse: with passages in context, the strongest 2023 models still left about half of their ELI5 statements without full support from the passages they cited, as scored by an NLI model. ### Evidence https://arxiv.org/pdf/2305.14627v2 p. 7, Table 6, with metric definitions p. 4-5, Section 3.3 | 2023-10-31 (v2; first submitted 2023-05-24) · EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations Plan 02b: exact values 51.1/50.0 and 44.0/50.1 stand only in Table 6, p. 7. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--only-the-two-bbc-rounds-repeat-one-design-over-time-and-cross-study-comparison-shows-no-rise-in-citation-precision-since-2023 (Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023) Address: https://trustillery.com/entity/stocks/ai-citations--on-eli5-in-2023-chatgpt-and-gpt-4-baselines-reach-about-50-percent-automatic-citation-recall-and-precision ## Orchestrator originates 85 percent of final-report errors in AI-Q, 53 percent in MS-Agent and 100 in TrajectoryKit type: anchor · status: open · kind: measurement · quoted: For AI-Q, 84.7% of final-report errors were attributed to the orchestrator, for MS-Agent 52.6%, and for TrajectoryKit 100%. · collected_by: Hirsch, Wan, Wang, Stengel-Eskin, Bansal, Dagan (Bar-Ilan University, UNC Chapel Hill, University of Texas at Austin) · independence: independent · source_family: peer-review · as_of: 2026-08-25 ### Statement Three open-source multi-agent deep-research systems (Nvidia AI-Q, MS-Agent, TrajectoryKit) were run on 20 DeepResearch Bench examples with up to 10 sampled sentences per agent per example. Global citation recall (share of citation-needing report sentences supported by their cited sources) was 58.7% for AI-Q, 28.5% for MS-Agent and 7.1% for TrajectoryKit. Each failing sentence was traced back through the agent chain to the invocation that introduced the error: the orchestrator was the origin for 84.7% of confirmed final-report errors in AI-Q (researcher 14.8%, searcher 0.4%), 52.6% in MS-Agent (reporter 47.4%) and 100% in TrajectoryKit. The 84.7% figure is for AI-Q only. Of AI-Q's orchestrator errors 70% were citation-related (uncited output, uncited input reliance, insufficient citations) and 30% hallucinations; in MS-Agent 99% were citation-related; in TrajectoryKit 95% of final-report errors were hallucinations. All tests were run by an LLM judge (gpt-5-mini-2025-08-07); on 50 AI-Q sentences labelled by two human annotators with arbitration, the method matched the human label exactly in 76% (kappa 0.62) and the localisation agreed on 75% of relevant sentences. ### Collection Academic authors at Bar-Ilan University, UNC Chapel Hill and University of Texas at Austin; arXiv preprint whose abs page carries the comment "Accepted to EMNLP 2026 (Main Conference)"; the observed text is the arXiv v1. Funding named: Israel Science Foundation, NSF, and a Google PhD Fellowship; none of the three evaluated systems is a Google product, and the authors are not the vendor of any of them. Method: entailment and citation-alignment prompts taken from prior work (LongCite, DEER, Localized Attribution Queries), documents read from the systems' tracing logs without re-crawling, citations to URLs not in the log filtered out as hallucinated. Only running-text sentences are evaluated; tables are excluded. The authors call the link between error type and underlying model capability observational, not controlled. Counter-check used: a human annotation study on 50 AI-Q sentences (inter-annotator kappa 0.71 before arbitration); it covers AI-Q only, not the other two systems. Code is released. ### Falls when A human-annotated trace of the same or a larger sample assigns less than half of AI-Q's final-report errors to the orchestrator, or the released code run on further systems (including closed commercial deep-research products, which this study does not cover) shows error origin dominated by retrieval or summarisation agents. Query to run: error-origin distribution per agent with human-labelled traces, more than 20 tasks, systems beyond the three open-source ones. ### Reflex Citation errors in research agents come from bad retrieval or from sources that do not say what is needed. Too coarse: in these three systems most errors that reach the report are introduced at the final writing step, where content that was supported upstream loses or misplaces its citation, and single-document summarisers err least (0.9 to 14.9 percent). ### Evidence https://arxiv.org/pdf/2608.24306v1 p. 7-8, Section 6.2 and Table 2 | 2026-08-25 · arXiv 2608.24306 · Hirsch et al., Who is the Agent to Blame? Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--support-falls-as-agent-runs-get-longer-and-most-traced-errors-arise-in-orchestration-and-not-in-search (Support falls as agent runs get longer and most traced errors arise in orchestration and not in search) Address: https://trustillery.com/entity/stocks/ai-citations--orchestrator-originates-85-percent-of-final-report-errors-in-ai-q-53-percent-in-ms-agent-and-100-in-trajectorykit ## Over 3150 generated references ChatGPT had the lowest and Perplexity the highest mean Reference Hallucination Score type: anchor · status: narrowed · kind: measurement · quoted: ChatGPT had the lowest mean RHS (1.81 ± 3.40) and thus emerged as the most reliable model. Gemini scored 4.01 ± 4.89, while Perplexity had the highest score at 6.51 ± 4.89 · collected_by: Ozbek and Bagcier (per the Europe PMC record's affiliation field: Department of Physical Medicine and Rehabilitation, University of Health Sciences, Derince Training and Research Hospital, Kocaeli; Basaksehir Cam and Sakura City Hospital, Istanbul, Turkey) · independence: independent · source_family: peer-review · as_of: 2026-05-11 ### Statement In a cross-sectional study, 30 rotator cuff subtopics were posed to ChatGPT, Gemini and Perplexity in two formats (letter to the editor and original article), yielding 3150 references scored with the Reference Hallucination Score (RHS), a composite of existence or verifiability, bibliographic accuracy, PMID validity and topical relevance in which a higher score means a less reliable reference; with all formats together the mean RHS was 1.81 +/- 3.40 for ChatGPT, 4.01 +/- 4.89 for Gemini and 6.51 +/- 4.89 for Perplexity (p < 0.001). These are mean scores with standard deviations on a scale whose range the abstract does not print, not shares of references. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Two clinician authors in Turkey, published in the peer-reviewed Indian Journal of Orthopaedics; no stake in the evaluated chatbots is visible in the record. Method per the abstract: each reference generated by the three chatbots was scored on the four RHS criteria. The observed text is the Europe PMC abstract record only: it does not print the scale of the RHS, the model versions, the test dates, whether web search was active, or who scored the references. The counter-check that exists is the full text and the supplementary material named in the abstract; neither was read for this card. ### Falls when The full text or supplement of doi:10.1007/s43465-026-01807-0 prints per-format and pooled means that cannot be reconciled with the abstract, or shows that the RHS is dominated by the topical relevance or PMID criterion rather than by existence, so that the ordering ChatGPT, Gemini, Perplexity does not hold for existence alone. Query to run: recompute the pooled mean RHS per chatbot from the supplementary per-reference scores and split it by criterion. ### Reflex An assistant that shows its sources, such as Perplexity, gives more reliable references than a plain chatbot. Too coarse: on this score Perplexity had the highest mean hallucination score of the three, and the score mixes existence, metadata, PMID validity and relevance without touching whether a source supports a claim. ### Evidence https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1007/s43465-026-01807-0&resultType=core&format=json abstract | 2026-05-11 · Indian Journal of Orthopaedics (2026), doi:10.1007/s43465-026-01807-0 · Ozbek and Bagcier, Reference Hallucination in AI-Assisted Academic Writing: A Comparative Analysis of ChatGPT, Gemini, and Perplexity in Rotator Cuff Literature ### Notes Abstract only. The abstract is hard to reconcile internally: it prints per-format means of 1.81 (letter) and 4.02 (article) for ChatGPT but a pooled mean of 1.81, and per-format means of 6.43 and 6.31 for Perplexity but a pooled mean of 6.51, which lies outside the two. A pooled mean of two groups must lie between the group means, so at least one printed figure is off or the pooling is not what the abstract suggests. Confidence is therefore low; the ordering of the three systems is the same in every printed row. Narrowed 2026-09-23 after attacker run 2 (Grok 4.7) withheld the point for lack of the full text: the abstract itself prints per-format means that cannot be pooled into its pooled means. Letter format: ChatGPT 1.81 +/- 3.20, Gemini 3.81 +/- 4.83, Perplexity 6.43 +/- 4.78; article format: ChatGPT 4.02 +/- 1.70, Gemini 4.13 +/- 1.57, Perplexity 6.31 +/- 1.68; pooled as printed: 1.81, 4.01, 6.51. Perplexity's pooled mean exceeds both of its format means and ChatGPT's pooled mean equals its letter-format mean, which no weighted mean of the two formats gives. What stands: the ordering ChatGPT below Gemini below Perplexity holds in each format separately. What is narrowed: the pooled figures in the quoted sentence are not to be read as reliable magnitudes until the full text or supplement resolves the print. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) Address: https://trustillery.com/entity/stocks/ai-citations--over-3150-generated-references-chatgpt-had-the-lowest-and-perplexity-the-highest-mean-reference-hallucination-score ## Raising tool calls from 2 to 150 drops Fact Check from 79 to 17 percent for GPT-5.4 and from 80 to 58 for Claude Opus 4.6 type: anchor · status: open · kind: measurement · quoted: GPT-5.4 shows the steepest decline, from 79% to 17% (62%). Claude Opus 4.6 demonstrates the greatest resilience, declining from 80% to 58% (22%). · collected_by: Onweller, Lumer, Huber, Ramchandani, Subbiah, Feld (PricewaterhouseCoopers U.S., Commercial Technology and Innovation Office) · independence: unknown · source_family: self-published · as_of: 2026-05-07 ### Statement In a search-depth ablation, two models run as deep-research agents (GPT-5.4 and Claude Opus 4.6) were capped at seven maximum tool-call budgets (2, 10, 30, 50, 70, 100, 150). Per-citation Fact Check (the cited page supports the attributed claim) fell for GPT-5.4 from 78.6% at 2 calls to 16.7% at 150 calls (Table 2) and for Claude Opus 4.6 from 80.0% to 57.9% (Table 3); the paper rounds these to declines of 62 and 22 percentage points and averages them to the 'approximately 42%' of its abstract. Link Works and Relevant Content stayed above 92% at every depth. The GPT-5.4 series is not monotone (35.5% at 70 calls, 37.2% at 100) and its largest single step is between 2 and 10 calls (78.6% to 45.9%). Scores come from the same rubric-based LLM judge as the main benchmark, not from human raters on each citation. ### Collection Authors are employees of PricewaterhouseCoopers U.S.; arXiv preprint, not peer reviewed. The pipeline extracts citations from the agents' Markdown reports, retrieves the cited page and has an LLM judge score three dimensions; the Fact Check judge was calibrated through manual review of 50 to 100 judgments. The paper does not print, for the ablation, the number of queries, the number of citations per depth level, or any interval, so the width of each percentage is not checkable from the text. The 42 percent figure is a mean of two percentage-point differences, not a relative drop. The authors are not the vendor of either model; a commercial interest of the firm is not declared and is recorded as unknown. Counter-check that exists: human audit of the judged citations and a repeat over more than two models; neither is reported. ### Falls when A rerun of the depth ablation with per-level citation counts and intervals, human-rated or with a judge validated against human labels, shows Fact Check at 150 calls within the interval of Fact Check at 2 calls for both models, or shows the decline only for GPT-5.4. Query to run: per-citation support rate as a function of tool-call budget, more than two models, with n per level, on reports generated with the protocol of arXiv 2605.06635. ### Reflex A research agent that searches more reads more and therefore cites more accurately. Too coarse: in this ablation link validity and relevance stay flat while measured support falls with search depth, for two models and without reported sample sizes. ### Evidence https://arxiv.org/pdf/2605.06635v1 p. 7-8, Section 4.3 and Tables 2-3 | 2026-05-07 · arXiv 2605.06635 · Onweller et al., Cited but Not Verified Relationships - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--support-falls-as-agent-runs-get-longer-and-most-traced-errors-arise-in-orchestration-and-not-in-search (Support falls as agent runs get longer and most traced errors arise in orchestration and not in search) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support (Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-study-measures-both-link-validity-sits-about-22-to-52-points-above-claim-support-in-the-printed-pairs (Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs) Address: https://trustillery.com/entity/stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 ## References cited by three or more of ten LLMs matched a scholarly database at 96 percent against 17 percent for one type: anchor · status: open · kind: measurement · quoted: For titles cited by two models, this rises to 87.4 %. At three or more models, the match rate reaches 95.6%, a 5.8-fold improvement over the single-model baseline. · collected_by: M.Z. Naser, School of Civil and Environmental Engineering and Earth Sciences and Artificial Intelligence Research Institute for Science and Engineering, Clemson University, USA · independence: independent · source_family: self-published · as_of: 2026-02-07 ### Statement Within the same audit of 69,557 reference instances generated by ten commercial LLMs, the share of unique title strings that matched a record in CrossRef, OpenAlex or Semantic Scholar was 16.5% for titles cited by only one model, 87.4% for titles cited by two models and 95.6% for titles cited by three or more models; within one model, a citation appearing in one of three replications matched at 28.6% and one recurring in two or more replications at 88.9%. This is a condition under which the existence rate of generated references is high: agreement across independently prompted models, measured by an automated matching pipeline. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Single academic author at Clemson University; arXiv preprint without venue; no stake in the evaluated vendors is declared or visible. Method: the same parsed citations and the same three-database fuzzy-matching pipeline as the paper's headline rates, regrouped by how many of the ten models produced the same title for the same prompt. The paper does not print the number of unique titles in each agreement class, and it reports that models of one family share more titles (Jaccard 0.540 for the two GPT-5 models). The counter-check on the pipeline is an LLM-with-web-search validation of 225 citations; no human audit is reported and none was run on the consensus classes specifically. ### Falls when The released dataset shows that the three-or-more class holds only a small fraction of all citation instances, so that the filter discards most real references along with the fabricated ones, or that the 95.6 percent is carried by within-family pairs sharing training data rather than by independent models. Query to run: from the dataset of arXiv 2603.03299 count unique titles per agreement class and recompute the match rate using only one model per vendor family. ### Reflex A reference that an LLM supplies cannot be trusted without looking it up. Too coarse for existence: a title that three or more models produce independently was found in a scholarly database 95.6 percent of the time; it stays true for support, which this audit did not measure. ### Evidence https://arxiv.org/pdf/2603.03299v1 p. 9, Section 4.5 | 2026-02-07 · arXiv 2603.03299 · Naser, How LLMs Cite and Why It Matters Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research (Three conditions are each measured on one component of citation quality and none on claim support in web research) Address: https://trustillery.com/entity/stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one ## ReportBench: 78.87 percent of OpenAI Deep Research cited statements judged consistent with their cited page, against 31.43 for o3 with search tools type: anchor · status: open · kind: measurement · quoted: while achieving a notably higher citation match rate (78.87% vs. 31.43%) and factual accuracy (95.83% vs. 82.22%). · collected_by: Minghao Li, Ying Zeng, Zhihao Cheng, Cong Ma, Kai Jia (ByteDance BandAI) · independence: independent · source_family: self-published · as_of: 2025-08-14 ### Statement On the same 100 ReportBench prompts, the match rate is 'the proportion of statements that are semantically consistent with their cited sources': gpt-4o identifies every statement with an explicit citation link, the cited web page is scraped, gpt-4o locates the passage most relevant to the statement and then gives a consistency verdict, and the verdicts are aggregated per report. Table 1 prints 78.87 percent for OpenAI Deep Research over 88.2 cited statements per report, 31.43 percent for openai-o3 with search over 16.16 cited statements per report, 72.94 percent for Gemini Deep Research, 73.67 percent for claude4-sonnet, 59.24 percent for gemini-2.5-pro and 44.88 percent for gemini-2.5-flash; no intervals. The base models were given SerpAPI search and Firecrawl reading with at most five tool calls per instance and were made to cite in URL form. This is a per-cited-statement support rate as judged by gpt-4o against the cited page; it is not a reference-list match against the survey's ground-truth bibliography (that is the separate precision/recall column, 0.385 / 0.033 for Deep Research) and not a human verdict. ### Collection Same collection as the factual-accuracy figure: reports gathered by the five ByteDance BandAI authors from the OpenAI and Gemini web interfaces between July 14 and July 25 (2025) and from batch runs of the base models with search and link-reading tools. gpt-4o performs statement extraction, supporting-passage extraction and consistency verification in three separate steps, which the authors present as more interpretable than a single LLM-as-a-judge score. No agreement figure between gpt-4o and human raters is printed; the counter-check consists of manual audits of individual cases (arXiv:2407.15186 and arXiv:2009.12619 test items) that surface statement and citation hallucinations, and the remark that intermediate outputs are retained for optional human inspection. ### Falls when Falls if a human re-judging of a sample of the stored cited statements finds the OpenAI Deep Research consistency rate materially different from 78.87 percent, or if rerunning the o3 baseline without the five-tool-call cap closes most of the gap to 31.43 percent. Query to run: from the ReportBench repository, take the per-statement consistency verdicts for OpenAI Deep Research and openai-o3, hand-check 200 cited statements from each, and recompute both match rates alongside a rerun of o3 with the tool-call cap lifted. ### Reflex The Deep Research wrapper makes o3's citations two and a half times more reliable. Too coarse: the o3 baseline was boxed in with five tool calls per report and a foreign URL citation format, wrote five times fewer cited statements, and both rates are gpt-4o verdicts on a scraped page with no printed human agreement; the gap shows the product pipeline grounding cited statements better under this setup, not how much of it the wrapper as such contributes. ### Evidence https://arxiv.org/pdf/2508.15804v1 p. 8, Section 3.4 Model-Level Comparative Analysis (sentence) and p. 7, Table 1 (Match Rate 78.87% vs 31.43%) | 2025-08-14 · arXiv 2508.15804 · Li, Zeng, Cheng, Ma, Jia, ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks ### Notes Proposed by attacker-grok-4.7 in the completeness attack of 2026-09-23 against tradeoff:supports: The cited-statement match rate is 78.87 percent for OpenAI Deep Research against 31.43 percent for o3 with search tools on the same task, so on one benchmark the deep-research wrapper grounds cited statements far better than the base model. Owner's reading of the bearing: It supports the tradeoff as one benchmark data point: cited statements of the deep-research product are judged consistent with their sources far more often than those of the base model with tools, but the base-model condition is handicapped (five tool calls, imposed URL citing) and the judge is unvalidated gpt-4o, so it supports the direction, not the size of the effect. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--reportbench-7887-percent-of-openai-deep-research-cited-statements-judged-consistent-with-their-cited-page-against-3143-for-o3-with-search-tools ## Supported share of Google AI Overview claims ranges from 77 to 95 percent by topic with Climate at 48 set aside type: anchor · status: open · kind: measurement · quoted: per-category fidelity ranges from 76.85% in Jobs & Education to 94.77% in Health. The highest-fidelity categories are Health (94.77%), Politics (93.65%), Science (91.82%), Business & Finance (91.42%), and Hobbies & Leisure (91.40%). · collected_by: Xu, Iqbal, Montgomery (Washington University in St. Louis) · independence: independent · source_family: self-published · as_of: 2026-05-13 ### Statement Broken down by the 19 topical categories of the query set, the share of AI Overview claims labelled consistent (Clear plus Vague) with the cited pages ranges from 76.85% in Jobs & Education to 94.77% in Health, with Politics at 93.65%, Science 91.82%, Business & Finance 91.42% and Hobbies & Leisure 91.40%. Climate stands at 48.23% and is excluded from that range by the authors: its Overviews report real-time weather values from structured feeds, and the cited pages had changed by the time they were crawled. With real-time query subsets removed, Technology rises to 89.41%, Shopping to 85.90% and Jobs & Education to 87.71%. Labels come from the same LLM verifier (Grok 4.1 Fast Reasoning) validated on 100 human-labelled verdicts. ### Collection Authors are at Washington University in St. Louis; they are not affiliated with Google in the paper; arXiv preprint without venue. Method: 55,393 trending queries in 19 topical categories were issued over a 40-day window (March 13 to April 21, 2026); each AI Overview was decomposed into atomic claims and each claim verified against the full extracted body text of every reference the Overview cites, by an LLM pipeline (Grok 4.1 Fast Reasoning, temperature 0) assigning one of five labels. Pages were crawled hours or days after the Overview was generated. The counter-check that exists and was used: two annotators re-labelled a stratified sample of 100 claim-level verdicts (20 per label), inter-annotator Cohen's kappa 0.94, and the verifier matched the adjudicated human label on 98 of 100; both errors were the verifier being too generous with supported claims. Claim extraction was validated separately on 100 Overviews. ### Falls when Page snapshots taken at the moment the Overview is generated show Climate and Jobs & Education at rates near the other categories (confirming the artifact) or still far below them (refuting it); or a per-topic human audit shows the verifier's accuracy differs by topic enough to reorder the categories. Query to run: per-category consistent share with same-minute page snapshots, and per-category human validation. ### Reflex Citation support for an AI search product is one number. Too coarse: within one product and one 40-day window the supported share spans 18 points by topic, and the lowest categories are driven by pages that change after the answer is generated. ### Evidence https://arxiv.org/pdf/2605.14021v1 p. 11-12, Section 4.3 and Figure 5 | 2026-05-13 · arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews Relationships - SUPPORTS ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) Address: https://trustillery.com/entity/stocks/ai-citations--supported-share-of-google-ai-overview-claims-ranges-from-77-to-95-percent-by-topic-with-climate-at-48-set-aside ## Ten commercial LLMs produced references with no database match at rates from 11 to 57 percent across 69557 citations type: anchor · status: open · kind: measurement · quoted: Across the ten models, hallucination rates range from 11.4% (GPT-5-mini) to 56.8% (haiku-4.5), a fivefold variation. · collected_by: M.Z. Naser, School of Civil and Environmental Engineering and Earth Sciences and Artificial Intelligence Research Institute for Science and Engineering, Clemson University, USA · independence: independent · source_family: self-published · as_of: 2026-02-07 ### Statement Ten commercially deployed LLMs queried through their APIs with prompts requesting scholarly references in four academic domains (structural engineering, climate and environmental science, biomedical research, NLP and AI), two temporal framings and three replications produced 69,557 parsed citation instances, of which 40,529 were matched at confidence score 80 or higher in CrossRef, OpenAlex or Semantic Scholar; the per-model share without such a match, which the paper calls the hallucination rate, ranges from 11.4% (GPT-5-mini, 95% CI 10.4 to 12.5) to 56.8% (haiku-4.5, 95% CI 55.6 to 58.0), and under the inclusive threshold of score 65 the same rates fall to between 9.3% and 23.8%. The rater is an automated fuzzy-matching pipeline, not a human, and the ten models ran without retrieval. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Single academic author at Clemson University; arXiv preprint without venue. No stake in the evaluated vendors is declared or visible. Method: 15,150 API responses, regex-based citation parsing, then title, author and year fuzzy matching against three scholarly databases; everything below score 65 is classed as hallucinated, and the headline rates also count scores 65 to 79 as hallucinated. The counter-check that exists was used in part: GPT-4.1-mini with web search judged a stratified sample of 225 citations and found 75 of 75 confirmed matches real and 8 of 75 citations classed as hallucinated to be real (10.7%). No human audit of the labels is reported. The paper states (p. 19) that it evaluates unaugmented text-generation models producing references from parametric memory alone, and that its figures are a baseline for non-retrieval citation generation. ### Falls when A human audit of a random sample of the 29,028 citations labelled hallucinated finds a share of real publications well above the 10.7 percent the paper's own LLM validation found, for instance books, reports and standards that CrossRef, OpenAlex and Semantic Scholar do not index, which would pull the 11.4 to 56.8 percent range towards the 9.3 to 23.8 percent the paper prints for the inclusive threshold. Query to run: hand-verify 300 randomly drawn unmatched citations from the released dataset of arXiv 2603.03299 in Google Scholar and WorldCat. ### Reflex Current LLMs hallucinate roughly a third of their references. Too coarse: the rate spans 11.4 to 56.8 percent across ten models of the same period, depends on the match threshold, and counts only whether the reference can be found in a database, not whether it supports anything. ### Evidence https://arxiv.org/pdf/2603.03299v1 p. 6-7, Sections 3.4, 3.5 and 4.1, Table 1; p. 10-11, Section 4.6; p. 19, limitations | 2026-02-07 · arXiv 2603.03299 · Naser, How LLMs Cite and Why It Matters Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research (Three conditions are each measured on one component of citation quality and none on claim support in web research) Address: https://trustillery.com/entity/stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations ## Thirteen LLMs asked for computer science references produced invalid citations at rates from 14 to 95 percent type: anchor · status: open · kind: measurement · quoted: hallucination rates spanning from 14.23% (DeepSeek) to 94.93% (Hunyuan), a roughly 6.7× difference · collected_by: Xu, Qiu, Sun, Miao, Wu, Li and colleagues, Nankai University and Tsinghua University · independence: independent · source_family: self-published · as_of: 2026-05-14 ### Statement Thirteen LLMs accessed through the OpenRouter API were prompted across 40 computer science domains, in batches of 10, 20 or 30, with and without online search plus chain-of-thought, to return references in a fixed JSON schema; of 331,809 extracted citations, 166,876 (50.29%) could not be verified by the authors' CiteVerifier tool against bibliographic databases and web search, and the per-model invalid share ranges from 14.23% +/- 1.65 (DeepSeek) through 21.84% (Claude 4), 50.92% (GPT-5) and 59.47% (Gemini) to 94.93% +/- 1.29 (Hunyuan), with online search showing no consistent effect across models. The authors describe these rates as a controlled baseline, not an estimate of real-world prevalence. This quantity is whether a generated reference exists and is bibliographically correct; it is not the share of citations whose cited passage supports the attached claim. ### Collection Academic authors at Nankai University and Tsinghua University; arXiv preprint (cs.CR), version 2. No stake in the evaluated vendors is declared or visible; the authors built and publish the verification tool the rates depend on. Method: 22,800 API interactions, 375,440 requested citations, automated verification by title similarity against academic databases with web-search and LLM-reparse fallbacks; unparseable outputs are excluded from the rates. The counter-check that exists was used: the authors manually checked random samples of 400 valid and 400 invalid verdicts and report 100% and 98% (392/400) agreement. Search and reasoning settings were switched through a third-party API aggregator, not through the vendors' own products. ### Falls when An independent re-verification of a random sample of the citations labelled invalid finds materially more than the 2 percent false positives the authors report, or the same prompts run in the vendors' own search-enabled products rather than through an API aggregator give rates far below Table II. Query to run: re-verify 400 randomly drawn invalid-labelled citations from the GhostCite benchmark release by hand, and rerun the Section B prompt in the consumer products with search on. ### Reflex Frontier models in 2026 rarely invent references. Too coarse: in this benchmark GPT-5 and Gemini returned unverifiable references in 50.92 and 59.47 percent of cases and the range across 13 models is 14.23 to 94.93 percent; the figure concerns existence only and comes from a list-generation prompt, not from drafting. ### Evidence https://arxiv.org/pdf/2602.06718v2 p. 5-7, Sections IV.A and V.A, Table II | 2026-05-14 (v2; v1 2026-02-06) · arXiv 2602.06718 · Xu et al., GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question (Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question) Address: https://trustillery.com/entity/stocks/ai-citations--thirteen-llms-asked-for-computer-science-references-produced-invalid-citations-at-rates-from-14-to-95-percent ## Twelve of fourteen deep research agents keep links valid above 94 percent while fact check scores range from 24 to 77 percent type: anchor · status: open · kind: measurement · quoted: Fact Check scores range from 24% (OSS-120B) to 77% (Claude Opus 4.5) · collected_by: Onweller, Lumer, Huber, Ramchandani, Subbiah, Feld (PricewaterhouseCoopers U.S., Commercial Technology and Innovation Office) · independence: unknown · source_family: self-published · as_of: 2026-05-07 ### Statement In a benchmark of 14 closed- and open-source LLMs run as deep-research agents on 130 queries, per-citation scores on three separate dimensions diverge: 12 of 14 models exceed 94% on Link Works (the URL resolves) and all frontier models exceed 80% on Relevant Content (the page is on topic), while Fact Check (the cited page supports the attributed claim) ranges from 24.4% (OSS-120B) to 76.8% (Claude Opus 4.5); Table 1 prints for GPT-5.4 100.0% / 93.7% / 47.7% and for Claude Opus 4.5 98.7% / 95.7% / 76.8%. Scores are assigned by rubric-based LLM judges calibrated through human review, not by human raters on every citation. ### Collection Authors are employees of PricewaterhouseCoopers U.S.; the paper is an arXiv preprint, not peer reviewed. Citations are extracted from the agents' Markdown reports with an AST parser, the cited page is retrieved, and an LLM-as-a-judge rubric scores each citation on the three dimensions. The authors are not the vendor of any evaluated model; a commercial interest of the firm in evaluation tooling is not declared in the paper and is recorded here as unknown. The counter-check that exists is human review of judge decisions; the reliability of such judges is measured separately in arXiv 2607.08700 by an overlapping author group. ### Falls when A human-rated re-evaluation of the same reports finds Fact Check rates for the frontier models within a few points of their Link Works and Relevant Content rates, or shows that the LLM judge's Fact Check decisions disagree with human raters often enough to move the 24 to 77 percent range. Query to run: human audit of per-citation support on reports generated with the protocol of arXiv 2605.06635. ### Reflex A citation with a working link to a relevant page supports the sentence it is attached to. Too coarse: link validity and topical relevance are measured separately here and sit 19 to 60 points above factual support in Table 1. ### Evidence https://arxiv.org/pdf/2605.06635v1 p. 6-7, Section 4.1 and Table 1 | 2026-05-07 · arXiv 2605.06635 · Onweller et al., Cited but Not Verified Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--in-retrieval-backed-systems-non-existent-links-are-the-smaller-failure-at-3-to-13-percent-against-23-to-76-percent-of-citations-failing-the-support-check (In retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-for-2025-and-2026-systems-span-24-to-94-percent (Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--published-per-citation-support-rates-span-24-to-94-percent-and-the-deep-research-floor-reads-40-or-50-depending-on-which-line-of-one-paper-is-taken (Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken) - OPPOSES ← https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation (Relying on an assistant citation without opening the cited passage against verifying each citation) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support (Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-support-on-the-same-citations-dead-links-explain-at-most-a-sixth-of-the-support-failures (Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--where-one-study-measures-both-link-validity-sits-about-22-to-52-points-above-claim-support-in-the-printed-pairs (Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs) Address: https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent ## Two fixes to the open AI-Q pipeline raise citation precision from 87.6 to 91.0 and 94.1 and recall from 64.5 to 69.7 on 50 DeepResearch Bench queries type: anchor · status: open · kind: measurement · quoted: These interventions successfully raise citation recall by 5% and citation precision by 3% to 7% percentage points, without reducing output quality. · collected_by: Eran Hirsch (Bar-Ilan University), David Wan (UNC Chapel Hill), Han Wang (UNC Chapel Hill), Elias Stengel-Eskin (University of Texas at Austin), Mohit Bansal (UNC Chapel Hill), Ido Dagan (Bar-Ilan University) · independence: independent · source_family: peer-review · as_of: 2026-08-25 ### Statement The authors ran Nvidia's open-source AI-Q deep-research pipeline (GPT-5 orchestrator, GPT-5 mini reporter, GPT-5 nano searcher) on the 50 English queries of DeepResearch Bench, unmodified and with two separate one-line fixes: appending 'do NOT use information which is not backed by citations' to the orchestrator prompt, and replacing the researcher agents' synthesized notes in the orchestrator input with the raw search snippets of the URLs they cite. Citation recall follows Zhang et al. (LongCite): for each report sentence, does it need a citation, and do its citations support it, judged by gpt-5-mini-2025-08-07 on the documents in the tracing log with hallucinated URLs filtered out; citation precision is 'calculated using the citation precision prompt from LongCite', whose definition the paper does not print. Table 3 prints recall 64.5 ± 3.9 (baseline), 69.7 ± 4.1 (snippets) and 69.6 ± 3.2 (guidance); precision 87.6 ± 3.0, 94.1 ± 2.0 and 91.0 ± 2.3; RACE quality 52.6 ± 0.4, 52.4 ± 0.7 and 52.6 ± 0.6, the ± being standard errors of the mean. These are machine-judged per-sentence citation quality rates of one open pipeline before and after the authors' own modifications; they are not a support rate of a commercial product and not a human-judged rate. ### Collection Collected by the six academic authors by running AI-Q themselves on DeepResearch Bench examples 51 to 100 with tracing logs, using gpt-5-mini-2025-08-07 as the judge for all prompts and evaluating retrieved documents from the log rather than re-crawling. Counter-check printed: a human study (Section 4.3) on 50 AI-Q sentences, each labelled by two of four annotators with a third arbitrating, gave inter-annotator Cohen's kappa 0.71, and the arbitrated labels matched the algorithm's error-type verdicts in 76 percent of cases (kappa 0.62), with 75 percent agreement on the supporting spans; that study validates the error-localisation method, not the LongCite recall and precision scores of Table 3, for which no human agreement is printed. ### Falls when Falls if a rerun of the released code on the same 50 queries with the same GPT-5 family models returns precision gains within the standard errors, or if a human re-judging of the precision verdicts shows the gpt-5-mini precision rate moving differently from the human rate across the two conditions. Query to run: clone the paper's repository, run baseline AI-Q and the citation-guidance variant on DeepResearch Bench 51 to 100 three times each, and compare the mean citation precision difference with the printed 3.4 points and its standard errors. ### Reflex A one-sentence prompt makes a research agent cite 3 to 7 points more precisely. Too coarse: the gains sit inside or at the edge of the ± 2 to 3 point standard errors on 50 queries, the precision metric is an unprinted LongCite prompt judged by gpt-5-mini with no human check, and the pipeline is one open system on GPT-5 models; it shows that precision was measured for a web-research intervention, not that the intervention transfers or that the effect size is settled. ### Evidence https://arxiv.org/pdf/2608.24306v1 p. 2, Section 1 Introduction (sentence) and p. 9, Table 3 and Section 7 (recall 64.5 to 69.7 / 69.6, precision 87.6 to 94.1 / 91.0) | 2026-08-25 · arXiv 2608.24306 · Hirsch, Wan, Wang, Stengel-Eskin, Bansal, Dagan, Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research ### Notes Proposed by attacker-grok-4.7 in the completeness attack of 2026-09-23 against stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research: Citation precision on the open pipeline AI-Q rises by 3 to 7 percentage points after an orchestrator instruction or a snippet substitution, so a web-research intervention was measured on citation precision, which that card says was not measured for its three conditions. Owner's reading of the bearing: It adds a fourth condition rather than overturning the three: an orchestrator prompt instruction and a snippet substitution on an open web-research pipeline, each measured on both citation recall and citation precision with standard errors, so the card's 'none on claim support in web research' has to be narrowed to the three named conditions or extended, with the caveat that precision here is an unprinted LongCite definition judged by gpt-5-mini and the gains are of the order of the standard errors. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) Address: https://trustillery.com/entity/stocks/ai-citations--two-fixes-to-the-open-ai-q-pipeline-raise-citation-precision-from-876-to-910-and-941-and-recall-from-645-to-697-on-50-deepresearch-bench-queries ## Vendor-run DRACO scores citation quality at 65 percent for Perplexity Deep Research and 42 to 56 for five rivals type: anchor · status: open · kind: body_statement · quoted: Perplexity Deep Research (with Opus 4.5 or 4.6) demonstrates best performance in all four categories, achieving the highest normalized scores in Factual Accuracy (67.9%), Breadth and Depth of Analysis (73.1%), Presentation Quality (90.3%), and Citation Quality (64.6%). · collected_by: Zhong, Zhang, Southern, Yang, Wang, Jung, Zhang, Yarats, Ho, Ma (Perplexity; one author Harvard University) · independence: subordinate · source_family: industry · as_of: 2026-02-12 ### Statement DRACO grades deep research outputs on 100 tasks against task-specific rubrics along four axes. The Citation Quality axis is described as 'Citations to primary source documents' and makes up 12% of criteria (4.8 of 39.3 per task on average). Table 13 prints normalized Citation Quality scores: Perplexity Deep Research (Opus 4.6) 64.6, Perplexity Deep Research (Opus 4.5) 62.5, Claude Opus 4.6 56.2, Gemini Deep Research 51.5, OpenAI Deep Research (o3) 45.8, OpenAI Deep Research (o4-mini) 42.5, Claude Opus 4.5 42.1. Each criterion receives a binary MET or UNMET from an LLM judge (Gemini-3-Pro), averaged over 5 grading runs. The axis scores whether a response cites the primary documents the rubric asks for; it is not a per-claim check of whether each cited passage supports the sentence it is attached to. ### Collection Nine of ten authors are at Perplexity, one at Harvard University; Perplexity is the vendor of the system that ranks first on every axis, the tasks are sampled from Perplexity Deep Research requests, and the rubrics were designed and validated in a process the vendor ran together with The LLM Data Company and 26 recruited domain experts (Section 4.1): collector and beneficiary coincide, recorded as subordinate. arXiv preprint without venue. The judge model was chosen drawing on an internal human-LLM alignment study that is not published in the paper. The counter-check that exists: the dataset and rubrics are public (hf.co/datasets/perplexity-ai/draco), and the paper reports scores under GPT-5.2 and Sonnet-4.5 as alternative judges with stable ranking; an evaluation by a party without a stake is not part of the paper. ### Falls when A party with no stake reruns the public DRACO tasks and rubrics and the Citation Quality ordering changes, or a per-claim support audit of the same outputs shows the axis score does not track the share of citations that support their claims. Query to run: independent regrade of the DRACO Citation Quality criteria plus a per-citation support check on the same reports. ### Reflex The vendor's benchmark shows its deep research product cites best. Too coarse: the axis covers 12 percent of the criteria, measures citation of primary documents rather than support of each claim, is LLM-judged, and tops out at 64.6 percent for the vendor's own system. ### Evidence https://arxiv.org/pdf/2602.11685v1 p. 12, Table 13; p. 6-7, Section 4.1 and Table 4; p. 9, Section 5.1 | 2026-02-12 · arXiv 2602.11685 · Zhong et al. (Perplexity), DRACO Plan 02b: the rivals' 42 to 56 stand only in Table 13, p. 12. Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--reverse-attribution-tests-and-the-vendor-primary-source-axis-measure-neighbouring-quantities-and-not-citation-support (Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support) Address: https://trustillery.com/entity/stocks/ai-citations--vendor-run-draco-scores-citation-quality-at-65-percent-for-perplexity-deep-research-and-42-to-56-for-five-rivals ## With a URL checking tool in the loop three models cut non-resolving citation URLs 6 to 79 fold to under 1 percent type: anchor · status: open · kind: measurement · quoted: GPT-5.1 from 16.0% to 0.6% (26×), Gemini from 6.1% to 0.1% (79×), and Claude from 4.9% to 0.8% (6.4×) · collected_by: Rao, Wong, Callison-Burch (University of Pennsylvania) · independence: positioned · source_family: self-published · as_of: 2026-04-03 ### Statement Condition under which link failure nearly disappears. On 435 ExpertQA questions (a 20% sample), Claude Sonnet 4.5, Gemini 2.5 Pro and GPT-5.1 answered with cited URLs while a URL checking tool (urlhealth: HTTP check plus Wayback Machine lookup) was available as a callable tool, and could verify and replace their own citations over several rounds. Classified with the same tool before and after on the same questions, the non-resolving rate fell for GPT-5.1 from 16.0% to 0.6% (26x), for Gemini from 6.1% to 0.1% (79x) and for Claude from 4.9% to 0.8% (6.4x), all p < 10^-35 in a two-proportion z-test. Table 3 prints final shares of LIVE 78.0 to 88.9%, DEAD 0.1 to 0.6%, LIKELY HALLUCINATED 0.4 to 1.8% and UNKNOWN 10.3 to 20.2% of 4,203 to 7,985 URLs per model. A smaller model (gpt-5-nano) in the same loop ended at a 7.5% not-live rate, with 48 hallucinated URLs persisting across up to 14 correction rounds. The measurement concerns whether links resolve; whether the replacement links support the claims was not measured. Automated measurement, no human or LLM rater. ### Collection Authors are at the University of Pennsylvania (DARPA SciFy funding); arXiv preprint under review, not peer reviewed. They are not the vendor of any evaluated model but are the makers of the tool whose effect is measured here, so independence is recorded as positioned for this result. The prose rates after mitigation (0.6, 0.1, 0.8%) do not equal DEAD plus LIKELY HALLUCINATED in Table 3 (2.4% for GPT-5.1, 0.7% for Gemini, 0.5% for Claude); the text does not explain the difference. UNKNOWN responses (10 to 20%) sit outside the non-resolving count; a headless-browser audit of 600 sampled UNKNOWN URLs found 11.0% [8.5, 13.7] genuinely dead. The pre-tool rate for GPT-5.1 on this subset under urlhealth classification (16.0%) is not the main-pipeline rate on the full set (8.47%); classification rules and question sets differ. Gemini ran in two phases because its API does not allow search grounding and custom tools together. The Claude run stopped at 658 questions on an API usage limit, hence the common subset of 435. Tool and data are announced under MIT license, so the counter-check (rerun) is open. ### Falls when A recomputation from the released data in which DEAD plus LIKELY HALLUCINATED plus the dead share of UNKNOWN exceeds 1% for any of the three models (Table 3 already sums to 2.4% for GPT-5.1), or an independent browser-based liveness check of the final citations finds more than 1% not resolving. As a condition for the stock's question it falls if a support audit shows that tool-verified links resolve but support their sentences no better than before. Query to run: browser liveness check of the final URL sets, and per-citation support rating before versus after the urlhealth loop. ### Reflex Broken and invented links are an inherent defect of language models. Too coarse: given a link checker as a tool and the competence to use it, three current models brought confirmed-broken citations down to below 1 percent by the paper's count and at most 2.4 percent by its table, so link failure is a tooling condition; support is a separate quantity that the check does not touch. ### Evidence https://arxiv.org/pdf/2604.03173v1 p. 7, Section 5.1 Results; p. 8, Table 3 | 2026-04-03 · arXiv 2604.03173 · Rao, Wong, Callison-Burch, Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents Relationships - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research (Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research) - FOLLOWS_FROM ← https://trustillery.com/entity/stocks/ai-citations--three-conditions-are-each-measured-on-one-component-of-citation-quality-and-none-on-claim-support-in-web-research (Three conditions are each measured on one component of citation quality and none on claim support in web research) Address: https://trustillery.com/entity/stocks/ai-citations--with-a-url-checking-tool-in-the-loop-three-models-cut-non-resolving-citation-urls-6-to-79-fold-to-under-1-percent ## Relying on an assistant citation without opening the cited passage against verifying each citation type: tradeoff · status: open · balance: two_sided · cost_side_by: costside-trustwork-e7 (Claude subagent with fresh context, 2026-09-19) ### Question For a person doing research, reporting or health reading with an AI assistant or deep-research agent of 2024 to 2026: treat an attached citation as sufficient evidence that the claim is supported and spot-check at most, or open the cited passage before relying on any cited claim? ### Benefits Entries for relying without opening, each carried by a SUPPORTS edge: - 89.0 percent of 98,020 atomic claims in Google AI Overviews are supported by the cited pages (any-page unit). - By topic that share is 76.85 to 94.77 percent, with Health at 94.77. - Four deep-research agents score 77.96 to 90.24 percent citation accuracy on DeepResearch Bench (LLM-judged). - Five deep-research systems score 0.69 to 0.86 faithfulness of cited claims on ResearcherBench (LLM-judged). - Between two BBC rounds significant sourcing issues fell to 10 to 15 percent for three of four assistants, while Gemini stayed broadly the same at 47 percent. - Three conditions are each measured on a single component: citation-trained models at 88.9 and 84.2 percent human-rated precision on supplied documents, a URL-checking tool cutting non-resolving links to under 1 percent as the paper states it, cross-model agreement at a 95.6 percent database match rate (derivation on three conditions); none was measured on claim support in web research. ### Costs Entries against relying without opening (costs of the measure), each carried by an OPPOSES edge: - Per-citation Fact Check of 14 deep-research agents on 130 queries runs from 24.4 percent (OSS-120B) to 76.8 percent (Claude Opus 4.5), with GPT-5.4 at 47.7, while 12 of 14 exceed 94 percent on Link Works (rubric-based LLM judge, preprint). - Four generative search engines reach 39.8 to 68.3 percent citation accuracy and leave 23.1 to 47.0 percent of query-relevant statements supported by none of the listed sources (DeepTRACE, 303 queries each, LLM judge at Pearson 0.62 against 100 human labels). - Deep-research configurations in the same audit reach 50.3 to 79.1 percent citation accuracy in Table 1 (the paper's running text gives Gemini Deep Research as 40.3 where Table 1 prints 50.3); the best case, GPT-5 Deep Research, leaves 12.5 percent of relevant statements unsupported, the other four 53.6 to 97.5 percent on a relevant subset of 12.4 to 45.5 percent of their statements. - The products on the benefit side read lower elsewhere: Perplexity Deep Research 58.0 percent in DeepTRACE against 90.24 on DeepResearch Bench, Gemini Deep Research 50.3 against 81.44; the spread between benchmarks for one product is 31 to 32 points. - Health reading: GPT-4o with web search on 300 consumer health questions has 100 percent valid URLs, 75.7 percent supported statements (95% CI 74.0-77.2) and 38.4 percent fully supported responses (26.7-49.3). On a different set, the paper's main set of 800 questions, the same system's response-level support is 55 percent: 31.0 percent (26.7, 35.8) on the 400 Reddit r/AskDocs questions against close to 80 on the 400 Mayo Clinic derived questions (peer reviewed, GPT-4o judge at 88.7 percent agreement with a three-doctor consensus). - Across seven LLMs on 800 medical questions 50 to 90 percent of responses are not fully supported by the sources they cite; the best system has 55 percent response-level support and approximately 30 percent of its individual statements unsupported. - Unit effect: the same answers give 75.7 percent supported statements and 38.4 percent fully supported responses; the benefit entries are claim-level or citation-level rates, and a whole answer is relied on at the response unit. - Any-page unit: the 89.0 percent for AI Overviews and the 75.7 percent of SourceCheckup count a claim as supported if any cited page supports it; both are claim-level rates and upper bounds on the share of claims supported by their own attached citation, not bounds on a rate counted over citations, and neither paper prints how far below the attached-citation share lies. - Gap between link validity and support: where one study scores both, link validity is 98.7 to 100 percent in the printed pairs and support lies lower by 21.9 points (98.7 against 76.8), 52.3 points (100.0 against 47.7) and 24.3 points (100 against 75.7); in the depth ablation, with Link Works above 92 percent at every depth, the gap is at most 21.4 and 20.0 points at 2 tool calls and at least 75 and 34 points at 150; a spot check that stops at a working, on-topic link does not register this gap. - Degradation with depth: per-citation Fact Check falls from 78.6 to 16.7 percent (GPT-5.4) and from 80.0 to 57.9 percent (Claude Opus 4.6) as the tool-call budget rises from 2 to 150, with Link Works and Relevant Content above 92 percent at every depth; two models, no n per level, no intervals printed. - Links that lead nowhere: across 10 models on DRBench 3.0 to 13.3 percent of citation URLs are classified as hallucinated and 5.4 to 18.5 percent do not resolve; pooled deep-research agents 10.7 and 16.2 percent; 8.22 percent non-resolving over 168,021 ExpertQA URLs (automated, no rater, stated as lower bounds). - News: journalists at 22 public media rated 31 percent of 2,709 responses as having significant sourcing issues (Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15 percent); human raters, response unit, and the category includes responses with no source at all. - Quote fidelity: 8 altered or absent BBC quotes across 62 quote-bearing responses (13 percent as the report computes it) and 12 percent of 1,053 quote-bearing responses with significant quote-accuracy issues, both human-rated; the BBC took part in both audits, so they are not independent. - Uncited claims: on ResearcherBench 0.84 faithfulness stands beside 0.34 groundedness (OpenAI Deep Research) and 0.80 beside 0.31 (Grok3 DeeperSearch), while Sonar Reasoning Pro has the lowest faithfulness (0.62) and the highest groundedness (0.68); groundedness, 0.31 to 0.68 in these rows, is the share of claims that carry a citation at all, and a support rate over cited claims is silent on the claims left uncited. - Judge uncertainty: validated support judges disagree with human raters on 2 percent, 11.3 percent and 14.9 to 22.4 percent of decisions, while the rate-level gaps the same validations print are smaller (0.0 to 3.3 points over six printed ELI5 pairs, 2.0 points in SourceCheckup); four of the six benefit entries are LLM-judged, of their judges the card's validations cover the AI Overviews verifier only, and the ResearcherBench anchor reports no human check of its support judge. - Generated references: invalid or fabricated shares of 11.4 to 56.8 percent (ten commercial LLMs), 14.23 to 94.93 percent (thirteen LLMs), 28.6 to 91.4 percent (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 references from eight free chatbots; this is existence of the reference, not support. The four printed rates come from studies that asked models for reference lists; two further studies in the card had models draft text with references. Running without retrieval is stated for one of the six studies only, the others pool or do not state it, and no study states that its models cited a retrieved page. Cost of the alternative (opening every cited passage): unmeasured. The stock holds no card on the time or effort of verification, so this entry carries no card and no edge. The nearby printed quantities are counts, not costs: 4.35 to 111.21 supported citations per task on DeepResearch Bench, 3.6 to 57.2 sources per deep-research answer in DeepTRACE. Opening is also not always possible: roughly 15 percent of source URLs could not be scraped in DeepTRACE. ### Unresolved What error rate in cited claims is tolerable depends on the use: a literature scan, a news report and a health decision do not share a threshold, and no anchor will settle that. Unmeasured: the time cost of opening every cited passage against the cost of acting on an unsupported claim; no source in the stock measures either. Unmeasured: per-citation support on ordinary user queries, as opposed to benchmark tasks. Consistency question: readers do not open every footnote of a human-written review either; what error rate is tolerated there is not in this stock. ### Notes Cost side searched on 2026-09-19 by costside-trustwork-e7 without sight of the benefit author's reasoning. Read: all 42 anchors and 18 derivations at statement level, the entered cards and the six benefit cards in full. Benefit entries read against their cards (nothing removed): - 89.0 percent AI Overviews: any-page unit (a claim-level rate, an upper bound on the share of claims supported by their own attached citation), one product, trending queries, pages crawled hours or days later, LLM verifier; both verifier errors in the 100 validated verdicts were the verifier being too generous. - Health at 94.77 percent: same unit and verifier. The stock's other health measurement prints 75.7 percent statement-level and 38.4 percent response-level support for GPT-4o with web search; different product, different unit. - DeepResearch Bench 77.96 to 90.24: the Benefits line selects the four deep-research agents; the same column runs from 39.36 percent. The judge (Gemini-2.5-Flash) is of the same family as the Gemini systems it rates in that table; validated on 100 pairs. - ResearcherBench 0.69 to 0.86: rate over cited claims only; its anchor reports no human check of the support judge; 65 questions on frontier AI topics. - BBC rounds: the line now carries Gemini at 47 percent beside the three assistants at 10 to 15. Response unit, mixed criterion, different product tiers between rounds, 237 responses over four assistants; the card itself names that the drop may be carried by no-source responses falling from 25 to one. The 22-organisation round of the same weeks prints ChatGPT at 24 percent. - Three conditions: as reworded by the benefit author the line now states itself that none was measured on claim support in web research. Per the card only the URL tool is a before-and-after measurement (the one causal reading); the citation-training figures compare different models. The URL tool acts on link resolution, which sits 21.9 to 52.3 points above support where both are measured; cross-model agreement acts on reference existence. A reader does not obtain these conditions by relying. Limits of the cost entries: - All per-citation cost rates except the news, quote, link and generated-reference entries are LLM-judged, as are four benefit entries. The judge derivation's band (2 to 22 percent at decision level, 0.0 to 3.3 and 2.0 points at rate level where printed) belongs to the judges it validated: of the entries here the SourceCheckup judge (88.7 percent agreement, behind the two health entries) and the AI Overviews verifier (98 of 100, behind two benefit entries). The other judges carry their own validations on their anchors: DeepTRACE Pearson 0.62 on 100 labels, DeepResearch Bench 96 and 92 percent agreement on 100 pairs, the 14-agent rubric judge calibrated through manual review of 50 to 100 judgments; the band is not transferred to them. The direction of judge error is unsettled in this stock: false negative rates of 0.183 to 0.470 on a human-reviewed benchmark come from the same author group as the 14-agent Fact Check and depth-ablation entries, so those two rates may read low. - The depth ablation prints no n per level and covers two models; its GPT-5.4 series is not monotone. - DeepTRACE uses 168 debate and 135 expertise queries and does not say how partial support is binarised; its authors include one at Microsoft, vendor of one evaluated engine. - The 31 percent news figure is a response-level rating that includes absent sources; it bounds per-citation support without measuring it. The EBU and BBC are publishers with a declared position. - The generated-reference rates come from four studies that asked models for reference lists (two further studies had models draft text with references). Retrieval status: stated as absent for one study, pooled with and without online search in one, unstated in four, two of which include Perplexity; no study states that its models cited a retrieved page, so the entry does not say whether retrieval was present. For systems stated to be retrieval-backed the stock prints non-existent links at 3 to 13 percent. Considered and not entered: the Tow Center reverse-attribution tests and the DRACO axis (neighbouring quantities, per their derivation); the 2023 baselines (outside the 2024 to 2026 window); the published-record audits (do not identify a tool, measure no reliance); the orchestrator localisation (open-source pipelines, global citation recall 58.7, 28.5 and 7.1 percent); the 24-to-94 range and no-single-value derivations (summaries holding both sides). What an unsupported claim costs in the use at hand has no card; it stays under Unresolved. Relationships - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--89-percent-of-98020-atomic-claims-in-google-ai-overviews-are-supported-by-the-cited-pages-and-11-percent-are-not - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--supported-share-of-google-ai-overview-claims-ranges-from-77-to-95-percent-by-topic-with-climate-at-48-set-aside - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--four-deep-research-agents-score-78-to-90-percent-citation-accuracy-on-deepresearch-bench-under-an-llm-judge - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--five-deep-research-systems-score-69-to-86-percent-faithfulness-of-cited-claims-on-researcherbench-under-an-llm-judge - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--between-two-bbc-rounds-significant-sourcing-issues-fell-to-10-to-15-percent-for-three-assistants-while-gemini-stayed-near-47 - ADDRESSES → https://trustillery.com/entity/stocks/ai-citations--do-the-citations-of-ai-research-assistants-support-the-claims-they-are-attached-to - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--four-generative-search-engines-reach-40-to-68-percent-citation-accuracy-and-leave-23-to-47-percent-of-statements-unsupported - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--the-same-deep-research-products-score-50-to-58-percent-on-one-benchmark-and-81-to-90-on-another - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--gpt-4o-with-rag-on-300-health-questions-has-all-urls-valid-but-76-percent-of-statements-and-38-percent-of-responses-supported - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--for-seven-llms-on-800-medical-questions-50-to-90-percent-of-responses-are-not-fully-supported-by-the-sources-they-cite - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--counting-per-response-instead-of-per-statement-halves-the-support-rate-in-the-same-data - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--two-support-rates-in-the-stock-count-a-claim-as-supported-if-any-cited-page-supports-it-and-bound-the-share-of-claims-supported-by-their-attached-citation - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--across-10-models-on-drbench-3-to-13-percent-of-citation-urls-are-hallucinated-and-5-to-18-percent-do-not-resolve - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--journalists-at-22-public-media-rated-31-percent-of-ai-assistant-news-responses-as-having-significant-sourcing-issues - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--two-journalist-audits-find-a-quote-accuracy-problem-in-about-one-in-eight-quote-bearing-news-responses - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--high-faithfulness-of-cited-claims-coexists-with-most-claims-carrying-no-citation-and-with-a-25-fold-spread-in-citation-volume - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--fabricated-reference-rates-of-11-to-95-percent-measure-whether-a-generated-reference-exists-and-do-not-answer-the-support-question - OPPOSES → https://trustillery.com/entity/stocks/ai-citations--where-one-benchmark-scores-link-validity-and-claim-support-on-the-same-citations-link-validity-sits-22-to-52-points-above-support - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--reportbench-7887-percent-of-openai-deep-research-cited-statements-judged-consistent-with-their-cited-page-against-3143-for-o3-with-search-tools - SUPPORTS → https://trustillery.com/entity/stocks/ai-citations--four-conditions-raise-one-component-of-citation-quality-each-and-only-the-fourth-was-measured-on-citation-precision-in-web-research Address: https://trustillery.com/entity/stocks/ai-citations--relying-on-an-assistant-citation-without-opening-the-cited-passage-against-verifying-each-citation