Questions › Do the citations of AI research assistants support the claims they are attached to
Citation precision ranged from 63.6 to 89.5 percent across four 2023 generative search engines and fell as utility rose
Bing Chat achieves the highest average precision (89.5), followed by perplexity.ai (72.7), NeevaAI (72.0), and YouChat (63.6)
https://arxiv.org/pdf/2304.09848v2 p. 7-8, Sections 4.2 and 4.3, with Tables 7-8, p. 23-24
Falls whenA re-computation from the released annotations of arXiv 2304.09848 gives per-system precision outside 63.6 to 89.5 or recall outside 11.1 to 68.7, or the correlation between citation precision and perceived utility is not near -0.96 when computed over the four system averages, or loses its sign when computed over query distributions or individual responses instead of four system means. Query to run: correlation of precision and perceived utility at response level in the released data.
Statement
In the same human-rated audit of the generative search engines of early 2023 (1450 queries per system, responses scraped late February to late March 2023), the per-system averages spread widely. Citation precision: Bing Chat 89.5, perplexity.ai 72.7, NeevaAI 72.0, YouChat 63.6. Citation recall: perplexity.ai 68.7, NeevaAI 67.6, Bing Chat 58.7, YouChat 11.1. The paper puts the recall gap at nearly 58% and the precision gap at almost 25%. It also reports, as printed, that "citation precision is inversely correlated with perceived utility (r = −0.96)": Bing Chat had the highest precision and the lowest perceived utility rating (4.34 on a five-point Likert scale), YouChat the lowest precision and the highest perceived utility (4.62). The paper does not state the unit over which r is computed; the ratings it sets against precision are the four per-system averages. Perceived utility and fluency were rated by the same human annotators. The authors offer as a hypothesis, not a measurement, that systems which copy or closely paraphrase cited pages gain precision and lose perceived utility.
Collection
Nelson F. Liu, Tianyi Zhang and Percy Liang, Computer Science Department, Stanford University; published in Findings of EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 23 Oct 2023). Human-rated, not automatic: 34 annotators recruited on Amazon Mechanical Turk, pre-screened with a qualification study, judged each query-response pair in three steps (verification-worthy statements, support of each statement by its citations, support contributed by each citation). Each system was run on 1450 queries (AllSouls, davinci-debate, ELI5 KILT and Live, WikiHowKeywords, seven NaturalQuestions subdistributions); responses were scraped between late February and late March 2023. Each pair was annotated once; the counter-check that exists and was used is a triple annotation of 250 randomly sampled pairs, with more than 82.0% pairwise agreement and 91.0 F1 for all judgments. The authors are academic researchers and not vendors of any evaluated system; the paper acknowledges Amazon Web Services for Mechanical Turk credits and the AI2050 program at Schmidt Futures. The human annotations are released. The correlation coefficient has no separate counter-check in the paper.
Falls when
A re-computation from the released annotations of arXiv 2304.09848 gives per-system precision outside 63.6 to 89.5 or recall outside 11.1 to 68.7, or the correlation between citation precision and perceived utility is not near -0.96 when computed over the four system averages, or loses its sign when computed over query distributions or individual responses instead of four system means. Query to run: correlation of precision and perceived utility at response level in the released data.
Reflex
The more helpful an answer looks, the better sourced it is. Too coarse: across these four systems the ordering ran the other way, and one average hides a 26 point spread in precision and a 58 point spread in recall.
Evidence
https://arxiv.org/pdf/2304.09848v2 p. 7-8, Sections 4.2 and 4.3, with Tables 7-8, p. 23-24 | 2023-04-19 · Findings of EMNLP 2023, arXiv 2304.09848 v2 · Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines
Findings and answers · 0
No attacker has recorded a finding on this card yet.