Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

In early 2023 four generative search engines had 51.5 percent citation recall and 74.5 percent citation precision

a mere 51.5% of generated statements are fully supported with citations (recall), and only 74.5% of citations fully support their associated statements (precision)

https://arxiv.org/pdf/2304.09848v2 p. 7, Section 4.2, with definitions p. 3-4, Sections 2.3-2.4, and Tables 7-8, p. 23-24

Falls whenA re-computation from the released human annotations of arXiv 2304.09848, taken as the unweighted mean of the four per-system values in Tables 7 and 8, gives citation recall outside 51.0 to 52.0% or citation precision outside 74.0 to 75.0%. It also falls if an independent human re-annotation of a random sample of at least 250 of the released query-response pairs, under the paper's own recall and precision definitions, yields a sample recall or precision more than 5 points away from the value the original annotations give on the same pairs, or if a full second annotation of the early-2023 responses moves either average by three points or more. The anchor narrows if it is read as a rate for later systems: query to run is a human-rated audit of the same engines with the same recall and precision definitions in 2024 to 2026.

✓ checked by Claude, human · 3not yet attackedindependent

Statement

A human-rated audit of four commercial generative search engines as they stood in early 2023 (Bing Chat, NeevaAI, perplexity.ai, YouChat; responses scraped between late February and late March 2023) on 1450 queries per system measured two quantities, averaged over the four systems. Citation recall, defined in the paper as the proportion of verification-worthy statements that are fully supported by their associated citations, was 51.5%. Citation precision, defined as the proportion of generated citations that support their associated statements, was 74.5%. The precision formula counts a citation that fully supports its statement and also a citation that partially supports it when the union of the statement's citations gives full support and no single citation does; the results sentence abbreviates this as citations that fully support. The averages are unweighted means of four system values (Tables 7 and 8), and the paper separates verifiability from factual correctness. This is a 2023 baseline that predates the stock's 2024 to 2026 window.

Collection

Nelson F. Liu, Tianyi Zhang and Percy Liang, Computer Science Department, Stanford University; published in Findings of EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 23 Oct 2023). Human-rated, not automatic: 34 annotators recruited on Amazon Mechanical Turk, pre-screened with a qualification study, judged each query-response pair in three steps (verification-worthy statements, support of each statement by its citations, support contributed by each citation). Each system was run on 1450 queries (AllSouls, davinci-debate, ELI5 KILT and Live, WikiHowKeywords, seven NaturalQuestions subdistributions); responses were scraped between late February and late March 2023. Each pair was annotated once; the counter-check that exists and was used is a triple annotation of 250 randomly sampled pairs, with more than 82.0% pairwise agreement and 91.0 F1 for all judgments. The authors are academic researchers and not vendors of any evaluated system; the paper acknowledges Amazon Web Services for Mechanical Turk credits and the AI2050 program at Schmidt Futures. The human annotations are released.

Falls when

A re-computation from the released human annotations of arXiv 2304.09848, taken as the unweighted mean of the four per-system values in Tables 7 and 8, gives citation recall outside 51.0 to 52.0% or citation precision outside 74.0 to 75.0%. It also falls if an independent human re-annotation of a random sample of at least 250 of the released query-response pairs, under the paper's own recall and precision definitions, yields a sample recall or precision more than 5 points away from the value the original annotations give on the same pairs, or if a full second annotation of the early-2023 responses moves either average by three points or more. The anchor narrows if it is read as a rate for later systems: query to run is a human-rated audit of the same engines with the same recall and precision definitions in 2024 to 2026.

Reflex

Search assistants that attach citations let the reader verify what they say. Too coarse: in this human audit about half of the generated statements were not fully supported by their own citations and about a quarter of citations did not support their statement.

Evidence

https://arxiv.org/pdf/2304.09848v2 p. 7, Section 4.2, with definitions p. 3-4, Sections 2.3-2.4, and Tables 7-8, p. 23-24 | 2023-04-19 · Findings of EMNLP 2023, arXiv 2304.09848 v2 · Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines

Findings and answers · 0

No attacker has recorded a finding on this card yet.