Questions › AI and evaluation
Do the citations of AI research assistants support the claims they are attached to?
For AI assistants and deep-research agents evaluated in published measurements between 2024 and 2026, what share of the citations attached to factual claims point to a passage that actually supports that claim, as distinct from a link that merely resolves or a source that is merely on topic?
Where the answer stands
For assistants and deep-research agents measured in 2025 and 2026 the published share of citations that support their claim runs from 24.4 to 94.04 percent. Within one benchmark the printed spread between systems is 52.4 points (24.4 to 76.8) and 54.7 points (39.36 to 94.04); one deep-research product moves 31 to 32 points between benchmarks, against spreads of 29 and 12 points between deep-research products within one benchmark. Counted per response instead of per statement, the same answers of one system in a dataset of 2024 give 38.4 against 75.7 percent, which is a different quantity and not a correction. Where automatic support judges were validated against humans they disagree on 2 to 22 percent of decisions; those validations concern other judges than the ones behind the 24 to 94 span. An answer to the question is a range with its conditions, not a number.
This rests on 6 further inferences, each with its parents and the condition that would break it:
- Counting per response instead of per statement halves the support rate in the same data
- Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken
- The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another
- Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions
- Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support
- Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures
No verdict, no score. The answer is a range with its conditions; the tradeoff below says where inference ends.
The card shown as the answer is the core inference chosen by the owner (tag core), not a rule of the schema.
How solid this is
- Second checks: 70 of 70
- cards confirmed by a checker identity · families: Claude, human
- Quoted spans: 43 of 46
- resolve against fetched text · 0 drifted, 0 absent, 3 unobserved · last run 2026-09-22
- Attack runs: 2
- by Grok, Claude · 9 findings · 8 applied, 2 refused · 4 cards fallen
- Open findings: 1
- awaiting the owner's disposition
- Proposals: 2
- 7 cards adopted, 2 rejected · both by the project's own agents so far
- Last attack: 4 d ago
- 2026-09-23 · stranger-trustwork-6a
- First-pass failures: 20 of 70 cards
- anchors 10 of 46, derivations 9 of 22, tradeoffs 1 of 1, questions 0 of 1, then repaired and re-checked; today 0 open failures · the stock's own error rate before its second check: cards whose first checker verification failed
A cut mark is the weak reading. How the readings combine into one index is not decided; until then they stand side by side. Every number comes from the engine's ledger and health readings.
Where the anchors come from
Independence of the collector
- independent
- 26 of 46
- positioned
- 15 of 46
- subordinate
- 1 of 46
- unknown
- 4 of 46
Where the document lives
- self-published
- 23 of 46
- peer-review
- 14 of 46
- established-media
- 8 of 46
- industry
- 1 of 46
Independence is what stake the collector has in the result, as the card declares it; a family is where the document lives. Counts, not a judgement.
What is asked, and what is not
What counts
Primary sources: peer-reviewed papers and preprints (arXiv, journal sites) that measure citation support, link validity or relevance for LLM assistants, generative search engines and deep-research agents; institutional evaluations (EBU/BBC, Tow Center, Reuters Institute); vendors' own published evaluation pages. Period: measurements published 2024 to 2026, with earlier benchmarks admitted as baselines where a 2024 to 2026 study builds on them. Measures: citation support or fact-check rate, citation precision and recall, link validity, topical relevance, and their change with search depth or task type. Both directions are searched: measurements of low support and measurements of high support or of conditions under which citations hold.
Counting unit, fixed after the scope attack of 2026-09-23: the share the question asks for is counted per attached citation, that is per (claim, cited source) pair as the benchmarks print it; a per-cited-claim rate counts as a value only where the study attaches one citation per claim. Response-level rates (every statement in a response supported), citation recall (the share of claims that carry any supporting citation), link validity and topical relevance are neighbouring quantities: the stock records them beside the share, as bounds or as context, never as values of it. Where a study reports support by any listed source rather than by the attached citation, the stock says so on the card and treats the value as an upper bound.
Not asked
- Whether AI research assistants are trustworthy in general, or better or worse than human researchers.
- Whether any named vendor is honest, misleading, or negligent; vendors appear only as sources of published claims and evaluations.
- The rate of fabricated (non-existent) references as a quantity of its own; it enters only where a study measures it beside citation support, because a reference can exist and still not support the claim.
- Whether users read or verify citations, and what they believe after reading them.
- Legal liability for a wrong citation.
- Which model is best; the stock records measurements per system and date, never a ranking.
Scope narrowed on 2026-09-23 after attacker run 3 (attacker-grok-4.7, x-attack-scope, verdict failed): the scope listed recall, link validity and relevance among the measures and did not fix the counting level, although those choices move published figures by tens of points; the title's quantity is now the only value and the others are declared neighbouring quantities.
The inferences
What follows from the measurements below, each with the step written out and the condition that would break it. Core inferences first.
Counting per response instead of per statement halves the support rate in the same data
Rests on GPT-4o with RAG on 300 health questions has all URLs val…; For seven LLMs on 800 medical questions 50 to 90 percent…
Falls whenA dataset in which response-level full support is not lower than statement-level support although responses contain several statements, which would mean unsupported statements cluster in few responses. Query: distributio…
Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken
Rests on Twelve of fourteen deep research agents keep links valid…; Four generative search engines reach 40 to 68 percent ci…; Deep research agents reach 50 to 79 percent citation acc…; Four deep research agents score 78 to 90 percent citatio… and 3 more
Falls whenA parent turns out to measure a different quantity (for instance support by any listed source instead of by the attached citation), which would remove its values from the range; a human-rated replication of any one bench…
The same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another
Rests on Deep research agents reach 50 to 79 percent citation acc…; Four deep research agents score 78 to 90 percent citatio…; Five deep research systems score 69 to 86 percent faithf…
Falls whenEvidence that the products changed materially between the evaluation dates in a way that explains 30 points, or a run of one product on two of the benchmarks in the same week with one judge that removes the gap. Query: r…
The share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks
Rests on The same deep research products score 50 to 58 percent o…; Validated automatic support judges disagree with human r…; Counting per response instead of per statement halves th…; Published per-citation support rates span 24 to 94 perce…
Falls whenA human-rated measurement on a representative sample of ordinary user queries across current products; it would give the question a value with an interval and make this card a statement about benchmarks only. Query: any …
Validated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions
Rests on 89 percent of 98020 atomic claims in Google AI Overviews…; GPT-4o support judge agrees with a three-doctor consensu…; ALCE automatic citation metrics agree with human raters …; Eight LLM judges match human-reviewed labels on factual …
Falls whenA larger validation (a thousand or more items) of one of these judges on its own benchmark shows disagreement well outside 2 to 22 percent, or shows that disagreement concentrates on one system so that rankings change. Q…
Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support
Rests on Twelve of fourteen deep research agents keep links valid…; GPT-4o with RAG on 300 health questions has all URLs val…; Raising tool calls from 2 to 150 drops Fact Check from 7…; Eight LLM judges match human-reviewed labels on factual …
Falls whenA study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.066…
Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures
Rests on Twelve of fourteen deep research agents keep links valid…; Across 10 models on DRBench 3 to 13 percent of citation …
Falls whenA per-citation release of arXiv 2605.06635 in which most citations failing Fact Check also fail Link Works (the two failures coincide on the same citations), or a current system whose non-resolving rate exceeds its suppo…
Fabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question
Rests on GPT-3.5 and GPT-4 and Bard hallucinated about 40 and 29 …; Of 400 references from eight free chatbots about 27 perc…; Ten commercial LLMs produced references with no database…; Thirteen LLMs asked for computer science references prod… and 2 more
Falls whenOne of these studies turns out to have checked the content of the cited work against the claim, which would make its rate a support rate; or a study shows that for retrieval-backed assistants existence failures and suppo…
Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research
Rests on Human raters find citation precision of 89 and 84 percen…; Citation-trained LongCite-8B reaches citation F1 of 72 o…; With a URL checking tool in the loop three models cut no…; References cited by three or more of ten LLMs matched a … and 2 more
Falls whenA study applying citation training or a URL-checking tool to a web-research assistant and measuring per-citation support before and after; or a human re-scoring of the AI-Q intervention that removes its precision gain. Q…
High faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume
Rests on Five deep research systems score 69 to 86 percent faithf…; Groundedness of 31 to 68 percent on ResearcherBench mean…; Effective citations per task range from about 4 to 111 a…
Falls whenEvidence that uncited claims in these reports are common knowledge needing no citation, which would make low groundedness harmless; or a benchmark where groundedness and faithfulness rise together across systems. Query: …
Non-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool
Rests on Audit of 111 million references in arXiv, bioRxiv, SSRN …; Invalid citations appear in about 1 percent of 56381 AI …
Falls whenThe join fails while both parents stand if the two audits are not independent readings of one trend: if the conference papers are largely the same documents as the arXiv corpus of the second audit, if both rest on the sa…
Only the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023
Rests on In early 2023 four generative search engines had 51.5 pe…; Citation precision ranged from 63.6 to 89.5 percent acro…; On ELI5 in 2023 ChatGPT and GPT-4 baselines reach about …; Four generative search engines reach 40 to 68 percent ci… and 1 more
Falls whenA replication of the 2023 human audit protocol on 2025 or 2026 engines; it would replace the cross-study comparison with a trend. Query: any paper citing arXiv 2304.09848 that reuses its annotation protocol on current sy…
Reverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support
Rests on Eight AI search engines answered over 60 percent of 1600…; ChatGPT Search gave partially or entirely incorrect sour…; Grok 3 cited URLs leading to error pages in 154 of 200 s…; Vendor-run DRACO scores citation quality at 65 percent f…
Falls whenA re-reading of one of the three documents (two CJR articles, DRACO) shows that it did check generated claims against cited passages. Query: methodology sections of the two CJR articles and DRACO Section on rubric axes.
Support falls as agent runs get longer and most traced errors arise in orchestration and not in search
Rests on Raising tool calls from 2 to 150 drops Fact Check from 7…; Orchestrator originates 85 percent of final-report error…
Falls whenAn error localisation on a commercial deep-research product that places most unsupported citations at retrieval; or a depth ablation on further models that shows no decline. Query: repeat of the tool-call ablation of arX…
The 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate
Rests on Journalists at 22 public media rated 31 percent of AI as…; BBC journalists rated over 45 percent of Gemini news res…
Falls whenThe EBU appendix or data release separates the three sub-categories and shows that unsupported claims alone account for nearly all of the 31 percent. Query: Q2 sub-codes in the EBU/BBC toolkit data.
The direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge
Rests on LLM citation judges reject 18 to 47 percent of genuinely…; Human raters find citation precision of 89 and 84 percen…; 89 percent of 98020 atomic claims in Google AI Overviews…
Falls whenA validation of one of the stock's judges on natural web-research output with a few hundred human labels that shows a consistent direction of error. Query: human re-labelling of released judge decisions from DeepResearch…
Two journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses
Rests on Of 1053 AI assistant news responses with direct quotes 1…; BBC journalists found 8 altered or absent BBC quotes acr…
Falls whenThe EBU data split by assistant shows the 12 percent is carried by one assistant (Gemini at 20 percent) so that the typical rate is far lower; or a per-quote count shows the BBC's 13 percent changes materially with the r…
Two support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation
Rests on 89 percent of 98020 atomic claims in Google AI Overviews…; GPT-4o with RAG on 300 health questions has all URLs val…
Falls whenA per-link audit on either dataset finds the share of claims supported by their attached citation equal to the any-source rate, which would make the bound tight. Query: per-link labels in the AI Overviews release.
Show the 4 inferences that fell. They stay, struck through, with their successors linked.
In retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check
Superseded by Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures. This card stays as history.
Rests on Across 10 models on DRBench 3 to 13 percent of citation …; Twelve of fourteen deep research agents keep links valid…
Falls whenA study on one sample of citations finds that most citations failing the support check also fail to resolve, or a current system whose hallucinated-URL rate exceeds its support-failure rate. Query: join of link status an…
Published per-citation support rates for 2025 and 2026 systems span 24 to 94 percent
Superseded by Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken. This card stays as history.
Rests on Twelve of fourteen deep research agents keep links valid…; Four generative search engines reach 40 to 68 percent ci…; Deep research agents reach 50 to 79 percent citation acc…; Four deep research agents score 78 to 90 percent citatio… and 1 more
Falls whenA parent turns out to measure a different quantity (for instance support by any listed source instead of by the attached citation), which would remove its values from the range; or a human-rated replication of any one be…
Three conditions are each measured on one component of citation quality and none on claim support in web research
Superseded by Four conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research. This card stays as history.
Rests on Human raters find citation precision of 89 and 84 percen…; Citation-trained LongCite-8B reaches citation F1 of 72 o…; With a URL checking tool in the loop three models cut no…; References cited by three or more of ten LLMs matched a … and 1 more
Falls whenA study applying one of these interventions to a web-research assistant and measuring per-citation support before and after. Query: papers citing LongCite or urlhealth that report claim-level support on web tasks.
Where one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs
Superseded by Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support. This card stays as history.
Rests on Twelve of fourteen deep research agents keep links valid…; GPT-4o with RAG on 300 health questions has all URLs val…; Raising tool calls from 2 to 150 drops Fact Check from 7…
Falls whenA study that scores resolving, relevance and support on the same citations with human raters and finds support within a few points of link validity for current systems; or evidence that the rubric judge of arXiv 2605.066…
Where inference ends
For a person doing research, reporting or health reading with an AI assistant or deep-research agent of 2024 to 2026: treat an attached citation as sufficient evidence that the claim is supported and spot-check at most, or open the cited passage before relying on any cited claim?
For relying without opening
- 89.0 percent of 98,020 atomic claims in Google AI Overviews are supported by the cited pages (any-page unit).
- By topic that share is 76.85 to 94.77 percent, with Health at 94.77.
- Four deep-research agents score 77.96 to 90.24 percent citation accuracy on DeepResearch Bench (LLM-judged).
- Five deep-research systems score 0.69 to 0.86 faithfulness of cited claims on ResearcherBench (LLM-judged).
- Between two BBC rounds significant sourcing issues fell to 10 to 15 percent for three of four assistants, while Gemini stayed broadly the same at 47 percent.
- Three conditions are each measured on a single component: citation-trained models at 88.9 and 84.2 percent human-rated precision on supplied documents, a URL-checking tool cutting non-resolving links to under 1 percent as the paper states it, cross-model agreement at a 95.6 percent database match rate (derivation on three conditions); none was measured on claim support in web research.
Against
- Per-citation Fact Check of 14 deep-research agents on 130 queries runs from 24.4 percent (OSS-120B) to 76.8 percent (Claude Opus 4.5), with GPT-5.4 at 47.7, while 12 of 14 exceed 94 percent on Link Works (rubric-based LLM judge, preprint).
- Four generative search engines reach 39.8 to 68.3 percent citation accuracy and leave 23.1 to 47.0 percent of query-relevant statements supported by none of the listed sources (DeepTRACE, 303 queries each, LLM judge at Pearson 0.62 against 100 human labels).
- Deep-research configurations in the same audit reach 50.3 to 79.1 percent citation accuracy in Table 1 (the paper's running text gives Gemini Deep Research as 40.3 where Table 1 prints 50.3); the best case, GPT-5 Deep Research, leaves 12.5 percent of relevant statements unsupported, the other four 53.6 to 97.5 percent on a relevant subset of 12.4 to 45.5 percent of their statements.
- The products on the benefit side read lower elsewhere: Perplexity Deep Research 58.0 percent in DeepTRACE against 90.24 on DeepResearch Bench, Gemini Deep Research 50.3 against 81.44; the spread between benchmarks for one product is 31 to 32 points.
- Health reading: GPT-4o with web search on 300 consumer health questions has 100 percent valid URLs, 75.7 percent supported statements (95% CI 74.0-77.2) and 38.4 percent fully supported responses (26.7-49.3). On a different set, the paper's main set of 800 questions, the same system's response-level support is 55 percent: 31.0 percent (26.7, 35.8) on the 400 Reddit r/AskDocs questions against close to 80 on the 400 Mayo Clinic derived questions (peer reviewed, GPT-4o judge at 88.7 percent agreement with a three-doctor consensus).
- Across seven LLMs on 800 medical questions 50 to 90 percent of responses are not fully supported by the sources they cite; the best system has 55 percent response-level support and approximately 30 percent of its individual statements unsupported.
- Unit effect: the same answers give 75.7 percent supported statements and 38.4 percent fully supported responses; the benefit entries are claim-level or citation-level rates, and a whole answer is relied on at the response unit.
- Any-page unit: the 89.0 percent for AI Overviews and the 75.7 percent of SourceCheckup count a claim as supported if any cited page supports it; both are claim-level rates and upper bounds on the share of claims supported by their own attached citation, not bounds on a rate counted over citations, and neither paper prints how far below the attached-citation share lies.
- Gap between link validity and support: where one study scores both, link validity is 98.7 to 100 percent in the printed pairs and support lies lower by 21.9 points (98.7 against 76.8), 52.3 points (100.0 against 47.7) and 24.3 points (100 against 75.7); in the depth ablation, with Link Works above 92 percent at every depth, the gap is at most 21.4 and 20.0 points at 2 tool calls and at least 75 and 34 points at 150; a spot check that stops at a working, on-topic link does not register this gap.
- Degradation with depth: per-citation Fact Check falls from 78.6 to 16.7 percent (GPT-5.4) and from 80.0 to 57.9 percent (Claude Opus 4.6) as the tool-call budget rises from 2 to 150, with Link Works and Relevant Content above 92 percent at every depth; two models, no n per level, no intervals printed.
- Links that lead nowhere: across 10 models on DRBench 3.0 to 13.3 percent of citation URLs are classified as hallucinated and 5.4 to 18.5 percent do not resolve; pooled deep-research agents 10.7 and 16.2 percent; 8.22 percent non-resolving over 168,021 ExpertQA URLs (automated, no rater, stated as lower bounds).
- News: journalists at 22 public media rated 31 percent of 2,709 responses as having significant sourcing issues (Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15 percent); human raters, response unit, and the category includes responses with no source at all.
- Quote fidelity: 8 altered or absent BBC quotes across 62 quote-bearing responses (13 percent as the report computes it) and 12 percent of 1,053 quote-bearing responses with significant quote-accuracy issues, both human-rated; the BBC took part in both audits, so they are not independent.
- Uncited claims: on ResearcherBench 0.84 faithfulness stands beside 0.34 groundedness (OpenAI Deep Research) and 0.80 beside 0.31 (Grok3 DeeperSearch), while Sonar Reasoning Pro has the lowest faithfulness (0.62) and the highest groundedness (0.68); groundedness, 0.31 to 0.68 in these rows, is the share of claims that carry a citation at all, and a support rate over cited claims is silent on the claims left uncited.
- Judge uncertainty: validated support judges disagree with human raters on 2 percent, 11.3 percent and 14.9 to 22.4 percent of decisions, while the rate-level gaps the same validations print are smaller (0.0 to 3.3 points over six printed ELI5 pairs, 2.0 points in SourceCheckup); four of the six benefit entries are LLM-judged, of their judges the card's validations cover the AI Overviews verifier only, and the ResearcherBench anchor reports no human check of its support judge.
- Generated references: invalid or fabricated shares of 11.4 to 56.8 percent (ten commercial LLMs), 14.23 to 94.93 percent (thirteen LLMs), 28.6 to 91.4 percent (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 references from eight free chatbots; this is existence of the reference, not support. The four printed rates come from studies that asked models for reference lists; two further studies in the card had models draft text with references. Running without retrieval is stated for one of the six studies only, the others pool or do not state it, and no study states that its models cited a retrieved page.
What stays unresolved
What error rate in cited claims is tolerable depends on the use: a literature scan, a news report and a health decision do not share a threshold, and no anchor will settle that. Unmeasured: the time cost of opening every cited passage against the cost of acting on an unsupported claim; no source in the stock measures either. Unmeasured: per-citation support on ordinary user queries, as opposed to benchmark tasks. Consistency question: readers do not open every footnote of a human-written review either; what error rate is tolerated there is not in this stock.
Two sides, no winner. Cost side written by costside-trustwork-e7 (Claude subagent with fresh context, 2026-09-19) without sight of the benefit side. Full card →
What was measured
Every measurement the stock rests on, with the document, the page and the verbatim span. One line per card; the source once.
Findings of EMNLP 2023, arXiv 2304.09848 v2 · Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines
EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations
EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations Plan 02b: exact values 51.1/50.0 and 44.0/50.1 stand only in Table 6, p. 7.
Journal of Medical Internet Research 26:e53164, doi:10.2196/53164 · Chelli et al., Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews
arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: the proprietary F1 values 65.6, 67.2 and 65.4 stand only in Table 2, p. 5.
arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: table-only quote, Table 6, p. 10, Human column P; no sentence carries 88.9, 84.2 or 67.5. Plan 02b: the caption is cut before 'GPT-4o' because the engine's canonical form folds the line-break hyphen of that token.
Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, How ChatGPT Search (Mis)represents Publisher Content
Computers in Biology and Medicine 185:109545 (2025), doi:10.1016/j.compbiomed.2024.109545 · Omar et al., Generating credible referenced medical research: A comparative study of openAI's GPT-4 and Google's gemini
BBC report · Elliott, Representation of BBC News content in AI Assistants
BBC report · Elliott, Representation of BBC News content in AI Assistants https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf p. 15, Appendix Results, Rating summary statistics (Q2)
Columbia Journalism Review, Tow Center · Jaźwińska, Chandrasekar, AI Search Has a Citation Problem
Nature Communications 16:3615 · Wu, Wu, Wei, Zhang, Casasola, Nguyen, Riantawan, Shi, Ho, Zou, An automated framework for assessing how well LLMs cite relevant medical references
arXiv 2505.18059 · Cabezas-Clavijo and Sidorenko-Bautista, Assessing the performance of 8 AI chatbots in bibliographic reference retrieval
arXiv 2506.11763 · Du, Xu, Zhu, Wang, Mao, DeepResearch Bench
arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench Plan 02b: table-only quote, Table 2, p. 8; no sentence in the paper carries the faithfulness values.
arXiv 2507.16280 · Xu, Lu, Ye, Hu, Liu, ResearcherBench
arXiv 2508.15804 · Li, Zeng, Cheng, Ma, Jia, ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Plan 02b: the 50.3 for Gemini DR stands only in Table 1, p. 9; the Section 4 text gives 40.3 for the same system.
arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants
EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q2 note and counts)
EBU Media Intelligence Service and BBC report · Fletcher, Verckist, News Integrity in AI Assistants https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf p. 67-68, Appendix 3, Assistant data (Q3 counts)
arXiv 2603.03299 · Naser, How LLMs Cite and Why It Matters
arXiv 2508.20033 · Patel, Arabzadeh, Gupta, Sundar, Stoica, Zaharia, Guestrin, DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
arXiv 2602.11685 · Zhong et al. (Perplexity), DRACO Plan 02b: the rivals' 42 to 56 stand only in Table 13, p. 12.
arXiv 2604.03173 · Rao, Wong, Callison-Burch, Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
arXiv 2605.06635 · Onweller et al., Cited but Not Verified
arXiv 2605.07723 · Zhao et al., LLM hallucinations in the wild
Indian Journal of Orthopaedics (2026), doi:10.1007/s43465-026-01807-0 · Ozbek and Bagcier, Reference Hallucination in AI-Assisted Academic Writing: A Comparative Analysis of ChatGPT, Gemini, and Perplexity in Rotator Cuff Literature
arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews Plan 02b: quoted sentence is Section 4.3, p. 11; Table 1, p. 6 carries the split.
arXiv 2605.14021 · Xu, Iqbal, Montgomery, Measuring Google AI Overviews
arXiv 2602.06718 · Xu et al., GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
arXiv 2607.08700 · Leung et al., Do You Need a Frontier Model as a Citation Verifier?
arXiv 2608.24306 · Hirsch et al., Who is the Agent to Blame?
arXiv 2608.24306 · Hirsch, Wan, Wang, Stengel-Eskin, Bansal, Dagan, Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
Whole upright: standing. Cut upright: narrowed. Lying stroke: fallen. ✓ names the checker's model family.
The record
| Run | Attacker | Form | Findings | Held | Cards | |
|---|---|---|---|---|---|---|
| 2026-09-22 | attacker-grok-4.7 · Grok | step, scope | 8 | 11 | 19 | rows → |
| 2026-09-23 | stranger-trustwork-6a · Claude | coordinate | 1 | 0 | 1 | rows → |
The owner answers every finding in the ledger, applied or refused with the error named. A coordinate attack on 2026-09-22 recorded no finding and therefore no row.
Proposals · 2 merged, rejections shown
merged proposals/ai-citations-002 · 4 cards from proposer-p1 · merged by operator-bjoern · 2026-09-22
Staged exercise by the project's own agent (proposer-p1), not outside participation: one checked correction of the DeepResearch Bench card, one new anchor with a span row, one deliberately weak card on a secondary source, one deliberate conflict on the falls_when of the early-2023 baseline anchor. Made to rehearse the proposal path before the first outside proposal.
- rejectchatgpt search misidentified 134 of 200 news articles or 67 percent in the tow center test of march 2025 (rejected, not in the stock)The coordinate and the engine anchor point at reporting (Ars Technica, 2025-03-13), not at the Tow Center study; the schema's Evidence rule requires the primary source. The figure is also covered by the stock's existing anchor on the eight-engine test, which pins the primary CJR page. Re-propose with the CJR page as coordinate and the quoted span taken from it if the ChatGPT Search figure deserves its own card.
- adopt with changesDeepResearch Bench citation judge Gemini-2.5-Flash matched human support labels in 96 percent and not-support labels in 92 percent of 100 sampled pairsAdopted as proposed; only found_via is changed from 'proposal-001' to name the staged proposal 002 by the project's own agent, so the card's origin reads honestly on the served page.
- adoptFour deep research agents score 78 to 90 percent citation accuracy on DeepResearch Bench under an LLM judge
- adopt with changesIn early 2023 four generative search engines had 51.5 percent citation recall and 74.5 percent citation precisionBoth sides sharpened the same sentence after the fork point. The proposer's version is adopted for the recomputation (it names the computation and a window a recomputation must hit) and for the paired re-annotation sample; the owner's three-point threshold for a full second annotation is kept beside it.
merged proposals/ai-citations-003 · 5 cards from attacker-grok-4.7 · merged by operator-bjoern · 2026-09-23
Completeness attack by the project's own attacker session (attacker-grok-4.7, plan 03b of the second bundle, 2026-09-23): the attacker searched for measurements the stock lacks and delivered four candidates with URL, verbatim span, target card and reason; the owner drafted the card bodies from those five fields after fetching each paper and wrote them into this fork under the attacker's identity as proposer. Three anchors and one tradeoff edge adopted, one anchor rejected as a neighbouring quantity.
- adoptDeepScholar-Bench: OpenAI DeepResearch scores .399 Citation Precision under GPT-4o entailment judging on 63 arXiv related-work queriesA per-citation, LLM-judged support rate on a commercial deep-research product printed at .399 (Table 2), below the 40.3 or 50.3 floor the range derivation states; the anchor names the entailment judge, the arXiv-only corpus and the 80 percent human agreement, so it enters the range as a value with its conditions. The owner will extend the range derivation.
- adoptRelying on an assistant citation without opening the cited passage against verifying each citationThe proposed SUPPORTS edge to the match-rate anchor is taken; the card body is otherwise unchanged.
- adoptReportBench: 78.87 percent of OpenAI Deep Research cited statements judged consistent with their cited page, against 31.43 for o3 with search toolsA per-cited-statement consistency rate of 78.87 percent for a commercial deep-research product against 31.43 for the tool-capped base model: a value of the question's quantity under an unvalidated gpt-4o judge, and one benchmark's support for the tradeoff's benefit side in direction, with the base-model handicap stated on the card.
- rejectreportbench 9583 percent of openai deep researchs uncited statements judged factually correct by a six vote gemini web check on 100 prompts (rejected, not in the stock)Factual accuracy of uncited statements is a different quantity from citation support, judged by an unvalidated six-vote Gemini check on a different task; and no card in the stock infers that an uncited claim is false, so there is no implication for it to qualify. It belongs in a stock on factual accuracy of research reports, not here.
- adoptTwo fixes to the open AI-Q pipeline raise citation precision from 87.6 to 91.0 and 94.1 and recall from 64.5 to 69.7 on 50 DeepResearch Bench queriesA web-research intervention measured on citation precision and recall with standard errors, on an open pipeline: a fourth condition under which citations hold better, which the derivation on three conditions must now name. The card states that the precision definition is not printed and that the gains are of the order of the standard errors. The owner will supersede the three-conditions derivation.
All 30 attack and answer rows
- 2026-09-2306:27#findingattack: coordinatestranger-trustwork-6a · ClaudeWith a URL checking tool in the loop three models cut non-resolving citation URLs 6 to 79 fold to under 1 percent
https://arxiv.org/pdf/2604.03173v1 p. 8 Table 3 | DEAD 0.1 to 0.6%, LIKELY HALLUCINATED 0.4 to 1.8%, UNKNOWN 10.3 to 20.2% with 11.0% [8.5, 13.7] of sampled UNKNOWN genuinely dead, as the card prints them | Falls When clause met: DEAD plus LIKELY HALLUCINATED plus the dead share of UNKNOWN exceeds 1% for every model (about 3.5 GPT-5.1, 1.8 Gemini, 1.6 Claude by the stranger's sum). Finding of the plan 04 stranger session, which read only the exported HTML; arithmetic unverified against the paper; open for the owner's disposition pass
- 2026-09-2223:25#findingattack: scopeattacker-grok-4.7 · GrokDo the citations of AI research assistants support the claims they are attached to
The title asks for one share: of the citations attached to factual claims, how many point to a passage that supports that claim, as distinct from a link that resolves or a source that is merely on topic. The scope does not fix that share. It admits citation recall, link validity and topical relevance as measures, and it never chooses citation-level against claim-level or response-level, nor an attached citation against any page in the reference set, choices that move published figures by tens of points. Not Asked rightly keeps out vendor motive, user belief, legal liability and a ranking, and fabricated references as their own quantity are rightly out because existence is not support. A reader of the title expects the attached-citation rate, and the scope's wider measure list does not hold the answer to that rate.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate
The EBU parent defines the 31% as a response-level journalist rating that covers unsupported claims, no sources at all, and incorrect sourcing claims, with Gemini at 72% and 42% of its responses giving no direct source. The BBC parent shows the earlier Q2 already mixed misattribution, unsupported claims and missing sources, and names lack of sources as part of the cause for the assistant with the highest rate (26% of Gemini responses). The step only narrows what the 31% can be cited for.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge
The judge-benchmark parent prints false-negative rates of 0.183 to 0.470, over-rejection of supported citations, and the LongCite parent prints GPT-4o precision of 79.7, 78.1 and 53.9 against human 88.9, 84.2 and 67.5 on the same responses. The AI Overviews parent records both mismatches in the 100 verdicts as the verifier being too generous, and the step does not average the two directions: tasks and base rates differ, and two cases are not a rate.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokThree conditions are each measured on one component of citation quality and none on claim support in web research
The step says the consensus anchor contributes existence for references recalled without retrieval, but parent stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one does not state that the ten models ran without retrieval. Stronger: the LongCite parents compare different models on a supplied document, human precision 88.9 and 84.2 against 67.5, the URL parent is a before-and-after of link resolution and says claim support was not measured, and the 95.6% figure is a database match for titles named by three models, so none of the three is claim support in web research.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokWhere one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs
The step treats SourceCheckup's 100% valid URLs against 75.7% supported statements as the same gap as the per-citation pairs, but those two percentages have different denominators, and it cites an 88.7% agreement with a three-doctor consensus that none of the three parents print. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent gives same-citation gaps of 21.9 points (98.7 against 76.8) and 52.3 (100.0 against 47.7), and the depth parent only bounds the gap because Link Works is printed as above 92% rather than as a paired value.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokValidated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions
The disagreements are the complements of printed agreement: 100 minus 98 is 2 on the AI Overviews sample, 100 minus 88.7 is 11.3 beside inter-doctor agreement of 86.1, and 100 minus 85.1 and 100 minus 77.6 are 14.9 and 22.4 for ALCE accuracy. The six ELI5 human-minus-automatic gaps in the ALCE parent run from 0.0 to 3.3 points and the SourceCheckup end-to-end gap is 42.4 minus 40.4, and the F1 parent is kept separate because 0.750 is not an agreement rate.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokTwo support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation
Both parents state an any-source unit: the AI Overviews parent checks each claim against all cited pages (89.0% consistent), and the SourceCheckup parent counts a statement supported if any source in the response supports it (75.7%). If the attached citation supported the claim then some cited page would, but not the reverse, and the step's example of 50% of claims against 75% of citations shows the denominators are not bounds on each other. Neither parent prints the attached-citation rate.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another
DeepTRACE's Table 1 prints Perplexity Deep Research at 58.0 and Gemini Deep Research at 50.3, DeepResearch Bench prints 90.24 and 81.44, and ResearcherBench prints faithfulness 0.85 and 0.86, so the 31-to-32-point gaps are 90.24 minus 58.0 and 81.44 minus 50.3. Within DeepTRACE the deep-research column runs from 79.1 to 50.3 and within DeepResearch Bench from 90.24 to 77.96, and the step states the version-change premise instead of picking a benchmark.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokThe share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks
The step says it adds no premise beyond the four parents, then asserts that the judges behind the 24.4–94.04 span have their own validations printed on their anchors. The judge parent's conclusion only places 2%, 11.3% and 14.9–22.4% on its own judges, and the range parent's conclusion lists ResearcherBench inside the span without a human validation. Stronger: the 2–22% band stays with the judges named in parent stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions, and nothing in the four parents prints a human validation for the judges behind that span, so the band is neither a correction to it nor already measured on it.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokTwo journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses
The BBC parent is eight altered or absent quotes over 62 responses, which the report calls 13% and which divides quotes by responses, and the EBU parent is 12% of 1,053 quote-bearing responses, about seventeen times the BBC sample. The BBC took part in both audits, and the step records the 12-to-13% numerical agreement while saying the two are not the same quantity.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokCounting per response instead of per statement halves the support rate in the same data
The HealthSearchQA parent prints both units on one sample, 75.7% of statements (74.0–77.2) and 38.4% of responses (26.7–49.3), and the seven-model parent prints the same system's pair on the 800-question set, about 30% of statements unsupported and 55% response-level support, under the rule that a response counts only if every statement does. The step does not invent a statements-per-response count the parents do not print.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokOnly the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023
Liu's parents give 2023 human precision of 74.5% on average and 63.6 to 89.5 by engine, ALCE gives the automatic baseline near 50% on research pipelines, and DeepTRACE gives 39.8 to 68.3% in August 2025 under a different rater and query set, so refusing a cross-study trend follows. The BBC parent is the only repeated design among the parents, Gemini staying at 47% and the other three falling to 10–15%, with the caveats the step uses: paid tiers to free, adjusted definitions, samples of 362 and 237.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokReverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support
Each parent states its own task: the two Tow Center anchors are reverse attribution of a given excerpt (more than 60% of 1,600 queries, and 153 of 200 for ChatGPT Search), the Grok 3 anchor is error-page resolution (154 of 200), and the DRACO anchor, published by the vendor of the top system, grades whether rubric primary documents are cited (42.1 to 64.6). None of those statements is a share of citations that support the attached claim, so sorting them out of the support range follows.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokIn retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check
The step calls 3.0–13.3% hallucinated URLs and 23.2–75.6% support failures an order-of-magnitude gap across two samples, but 13.3 against 23.2 is not an order of magnitude, and 100 minus Fact Check does not show that the failing citations exist. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent already scores both on the same citations, Link Works above 94% for 12 of 14 beside Fact Check 24.4–76.8, so most of those support failures are resolving pages.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokFabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question
The step says the invalid shares in the four reference-list studies bound support from above, but parent stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations prints a hallucination rate of 11.4%, and a non-existence rate is a lower bound on citations that cannot support a claim, so support is at most one minus that rate. Stronger: those rates measure existence or bibliographic match, not support, and only their complements are loose upper bounds, because a matched reference need not support the attached claim.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokSupport falls as agent runs get longer and most traced errors arise in orchestration and not in search
The ablation parent shows Fact Check falling between the endpoints while Link Works and Relevant Content stay above 92%, not that they are unchanged, and the localisation parent places the orchestrator as the origin of errors only in three open pipelines. Stronger: parent stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 supports a drop with tool-call budget for two commercial models while links keep resolving, and parent stocks/ai-citations--orchestrator-originates-85-percent-of-final-report-errors-in-ai-q-53-percent-in-ms-agent-and-100-in-trajectorykit does not measure those models, so the premise that they fail at the orchestrator is in neither parent.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokPublished per-citation support rates for 2025 and 2026 systems span 24 to 94 percent
The step says it takes the minimum of printed values for systems the parents label deep research, but parent stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported prints Gemini Deep Research at 40.3% in the running text and 50.3% in Table 1, and parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent prints GPT-5.4, run as a deep-research agent, at 47.7% Fact Check. Stronger: 24.4 to 94.04 is the min and max of the per-citation and per-cited-claim rates, while a commercial deep-research floor is not 50.3 once those lower printed values are kept.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokNon-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool
The conference parent gives the hand-confirmed per-paper share, 604 of 56,381 papers (1.07%) and 1.61% in 2025 against a 0.89% average for 2020–2024, and the four-corpus parent gives the automated excess over the pre-LLM baseline, 0.21% to 1.91% of references as of August 2025. The step keeps the units apart and keeps both authors' refusal to name a tool.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokHigh faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume
The faithfulness and groundedness parents are one ResearcherBench table, so OpenAI Deep Research at 0.84 and 0.34, Grok3 DeeperSearch at 0.80 and 0.31, and Sonar Reasoning Pro at 0.62 and 0.68 are paired, and the volume parent separately prints 4.35 to 111.21 effective citations and 94.04 accuracy with 9.78 against 81.44 with 111.21. A support rate over cited claims is silent on uncited claims and on citation count, and the step marks the 111.21-to-4.35 ratio as computed here.
Cite a card by its id and this page's date; to cite an event, link its ledger row. Ids never change; fallen cards keep theirs. Full ledger →