Questions › Do the citations of AI research assistants support the claims they are attached to
Four generative search engines reach 40 to 68 percent citation accuracy and leave 23 to 47 percent of statements unsupported
You.com and Perplexity list slightly fewer sources (3.4–3.5) but still struggle with unsupported claims (23–47%). Finally, on citation metrics, all three engines show relatively low citation accuracy (40–68%), with frequent misattribution.
https://arxiv.org/pdf/2509.04499v1 p. 8, Figure 2a and Section 4; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation
Falls whenA human rating of the same engines' answers on the DeepTRACE queries puts citation accuracy of all four above 80%, or shows that the judge's error (Pearson 0.62 against human labels) is large enough to move the 39.8 to 68.3 range by more than ten points; or counting partial support as supported closes most of the gap. The anchor narrows if the result is driven by the 168 debate queries. Query to run: per-citation human support labels on a stratified sample of the DeepTRACE outputs, split by query type and by the binarisation of partial support.
Statement
In the DeepTRACE audit, four public generative search engines were each run on 303 queries (168 debate questions from ProCon, 135 expertise questions contributed by participants of an earlier user study); results are as of 27 August 2025. Citation accuracy, defined as the fraction of statement citations for which the cited source's content supports the statement (overlap of citation matrix and factual-support matrix divided by the number of citations), was 68.3% for You.com, 65.8% for Bing Copilot, 49.0% for Perplexity and 39.8% for GPT-4.5. Unsupported statements, defined as the fraction of query-relevant statements not factually supported by any of the listed sources, cited or not, were 30.8%, 23.1%, 31.6% and 47.0% in the same order. Citation thoroughness (accurate citations over all possible accurate citations) was 20.5% to 24.4%. Support was decided per statement-source pair by an LLM judge (GPT-5 by default) over the scraped full text; on 100 manually verified tasks the judge correlated with human labels at Pearson 0.62. Link validity is not a reported metric: for roughly 15% of URLs the scraper returned an error (paywall or unavailable page) and these sources were excluded from the support calculations.
Collection
Authors: five at Salesforce AI Research, one at Microsoft Research; arXiv preprint, not peer reviewed. Microsoft is the vendor of one evaluated system (Bing Copilot), so independence is recorded as positioned; Salesforce is the vendor of none of the evaluated systems. Browser scripts extracted answer text, citations and source URLs from the public web interfaces; Jina Reader scraped the source text; an LLM judge decomposed answers into statements and filled the factual-support matrix (on the order of 80,000 support judgements). The support prompt in Appendix E returns full, partial or none, and the paper does not say how partial is turned into the binary matrix; Section 3 names GPT-5 as default judge while Appendix E names GPT-4. Counter-check that exists and was used: two hired annotators labelled 100 support tasks (Pearson 0.62 against the judge, called moderate agreement by the authors). No per-system human audit of citation accuracy is reported. The authors name reliance on an LLM judge as a limiting factor. The query set is announced as released. The Figure 2 caption and the running text speak of three engines while the score card prints four columns (You, Bing, PPLX, GPT 4.5).
Falls when
A human rating of the same engines' answers on the DeepTRACE queries puts citation accuracy of all four above 80%, or shows that the judge's error (Pearson 0.62 against human labels) is large enough to move the 39.8 to 68.3 range by more than ten points; or counting partial support as supported closes most of the gap. The anchor narrows if the result is driven by the 168 debate queries. Query to run: per-citation human support labels on a stratified sample of the DeepTRACE outputs, split by query type and by the binarisation of partial support.
Reflex
AI search engines cite their sources, so a reader can check any claim with one click. Too coarse: between about one third and three fifths of the citations measured here lead to a source that does not support the sentence it is attached to, and that is a different quantity from the share of sentences no listed source supports.
Evidence
https://arxiv.org/pdf/2509.04499v1 p. 8, Figure 2a and Section 4; p. 5-7, Sections 3.1.1 to 3.1.4 for definitions and judge validation | 2025-09-02 · arXiv 2509.04499 · Venkit, Laban, Zhou, Huang, Mao, Wu, DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Findings and answers · 0
No attacker has recorded a finding on this card yet.