Questions › Do the citations of AI research assistants support the claims they are attached to
On ELI5 in 2023 ChatGPT and GPT-4 baselines reach about 50 percent automatic citation recall and precision
on the ELI5 dataset, around 50% generations of our ChatGPT and GPT-4 baselines are not fully supported by the cited passages
https://arxiv.org/pdf/2305.14627v2 p. 7, Table 6, with metric definitions p. 4-5, Section 3.3
Falls whenA human-rated evaluation of the same ELI5 outputs gives citation recall or precision for the ChatGPT or GPT-4 VANILLA systems that departs from the automatic 44 to 53 percent band by more than a few points, or a different entailment model applied to the released outputs moves the band. The anchor narrows if read as a rate for deployed assistants: query to run is the ALCE metrics, or a human audit with the same definitions, on the outputs of 2024 to 2026 systems for the same 1,000 ELI5 questions.
Statement
On the ELI5 part of the ALCE benchmark (1,000 randomly selected development questions, retrieval from the Sphere web corpus cut into 100-word passages), citation quality of the authors' retrieve-and-prompt systems built on the models of 2023 was scored automatically by an NLI model (TRUE, a T5-11B model), not by human raters. Table 6 prints citation recall / citation precision of 51.1 / 50.0 for ChatGPT VANILLA (5 passages), 44.0 / 50.1 for GPT-4 with 5 passages, 48.5 / 53.4 for GPT-4 with 20 passages, 38.3 / 37.9 for LLaMA-2-Chat-70B, and 69.3 / 67.8 for ChatGPT with RERANK, a strategy that reranks sampled generations by the same automatic citation recall. Recall here is the share of statements whose concatenated cited passages entail the statement; precision is the share of citations not detected as irrelevant, given recall of 1. The paper summarises this as around 50% of generations of its ChatGPT and GPT-4 baselines not being fully supported by the cited passages. The systems are research pipelines over a fixed corpus, not deployed assistants citing live web pages. This is a 2023 baseline that predates the stock's 2024 to 2026 window.
Collection
Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen, Department of Computer Science and Princeton Language and Intelligence, Princeton University; published at EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 31 Oct 2023). Automatic, not human-rated: each statement and the concatenation of its cited passages go to the TRUE NLI model, which returns entailment or not; scores of the open models are averaged over three seeded runs; Appendix G.6 states that ChatGPT-16K and GPT-4 use one seeded run each, so the GPT-4 figures are single-run values. The authors built the evaluated pipelines themselves but are not vendors of the underlying models; the paper acknowledges an IBM PhD Fellowship, an NSF CAREER award, a Sloan Research Fellowship and Microsoft Azure credits. The counter-check that exists and was used is a human evaluation through Surge AI on sampled outputs of three systems: on ELI5 human raters gave ChatGPT VANILLA 50.8 recall and 52.4 precision against 52.8 and 50.4 from ALCE (Table 9). The paper states that the NLI model cannot detect partial support and therefore gives a lower citation precision score than human evaluation.
Falls when
A human-rated evaluation of the same ELI5 outputs gives citation recall or precision for the ChatGPT or GPT-4 VANILLA systems that departs from the automatic 44 to 53 percent band by more than a few points, or a different entailment model applied to the released outputs moves the band. The anchor narrows if read as a rate for deployed assistants: query to run is the ALCE metrics, or a human audit with the same definitions, on the outputs of 2024 to 2026 systems for the same 1,000 ELI5 questions.
Reflex
If the model is given the retrieved passages and told to cite them, its citations will back what it writes. Too coarse: with passages in context, the strongest 2023 models still left about half of their ELI5 statements without full support from the passages they cited, as scored by an NLI model.
Evidence
https://arxiv.org/pdf/2305.14627v2 p. 7, Table 6, with metric definitions p. 4-5, Section 3.3 | 2023-10-31 (v2; first submitted 2023-05-24) · EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations Plan 02b: exact values 51.1/50.0 and 44.0/50.1 stand only in Table 6, p. 7.
Findings and answers · 0
No attacker has recorded a finding on this card yet.