Questions › Do the citations of AI research assistants support the claims they are attached to
ALCE automatic citation metrics agree with human raters at kappa 0.698 for recall and 0.525 for precision
kappa coefficient between human and ALCE suggests substantial agreement for citation recall (0.698) and moderate agreement for citation precision (0.525)
https://arxiv.org/pdf/2305.14627v2 p. 8-9, Section 6 and Tables 8-9, with Appendix F, p. 15-16, and Appendix G.5, p. 17
Falls whenAn independent human annotation of ALCE outputs yields a Cohen's kappa with the automatic labels clearly below 0.698 for recall or 0.525 for precision, or accuracy below 85.1% and 77.6%, or the agreement does not hold once the systems are 2024 to 2026 models whose statements synthesise several passages. Query to run: human versus TRUE-NLI agreement on citation support for outputs of current models on the ALCE questions.
Statement
The ALCE paper validates its automatic, NLI-based citation recall and precision against human judgement on outputs of three systems built on 2023 models (ChatGPT VANILLA, ChatGPT RERANK, Vicuna-13B VANILLA). Human raters judged, per sentence, whether all cited passages together fully support it, and per citation, whether it fully, partially or does not support the sentence. Cohen's kappa between human and automatic labels is printed as 0.698 for citation recall and 0.525 for citation precision; treating human annotations as gold labels, the automatic metric has an accuracy of 85.1% for citation recall and 77.6% for citation precision. For detecting irrelevant citations it has a recall of 75.6% and a precision of 66.1%, which the paper attributes to the NLI model being unable to detect partial support. On ELI5 the system-level scores are, human against ALCE, 50.8 / 52.4 against 52.8 / 50.4 for ChatGPT VANILLA, 59.7 / 60.6 against 63.0 / 60.6 with RERANK, and 13.4 / 19.2 against 13.6 / 18.1 for Vicuna-13B (recall / precision, Table 9). This is a 2023 baseline for how far automatic support metrics track human raters.
Collection
Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen, Department of Computer Science and Princeton Language and Intelligence, Princeton University; published at EMNLP 2023 (venue per the stock's source list; the observed text is arXiv v2, stamped 31 Oct 2023). The validation is human-rated: workers employed through Surge AI at an average pay of 20 USD per hour; the paper says it randomly sampled 100 examples from ASQA and ELI5 and annotated the outputs of the three selected models, without stating the number of raters or an inter-rater agreement figure. The automatic side is the TRUE NLI model (T5-11B). The authors validate a metric they propose themselves, so they have a stake in the outcome; that is why independence is recorded as positioned, although they are not vendors of any evaluated model. The paper thanks Surge AI for support with the human evaluation. A counter-check by a party other than the metric's authors is not part of the paper.
Falls when
An independent human annotation of ALCE outputs yields a Cohen's kappa with the automatic labels clearly below 0.698 for recall or 0.525 for precision, or accuracy below 85.1% and 77.6%, or the agreement does not hold once the systems are 2024 to 2026 models whose statements synthesise several passages. Query to run: human versus TRUE-NLI agreement on citation support for outputs of current models on the ALCE questions.
Reflex
Automatic entailment checks are a good enough stand-in for human judgement of whether a citation supports a claim. Too coarse: agreement was substantial for whether a statement is supported by all its citations together and only moderate for whether a single citation supports it, the quantity this stock asks about.
Evidence
https://arxiv.org/pdf/2305.14627v2 p. 8-9, Section 6 and Tables 8-9, with Appendix F, p. 15-16, and Appendix G.5, p. 17 | 2023-05-24 · EMNLP 2023, arXiv 2305.14627 v2 · Gao, Yen, Yu, Chen, Enabling Large Language Models to Generate Text with Citations
Findings and answers · 0
No attacker has recorded a finding on this card yet.