Questions › Do the citations of AI research assistants support the claims they are attached to
Human raters find citation precision of 89 and 84 percent for LongCite models against 68 for GLM-4 on LongBench-Chat
GLM-4 61.2 67.5 60.247.6 53.9 47.146.1 29.1 30.8 LongCite-8B79.6 88.9 82.662.0 79.7 67.459.639.542.0 LongCite-9B72.884.275.857.678.163.664.2 45.1 47.1 Table 6: Citation quality evaluated by human
https://arxiv.org/pdf/2409.02897v3 p. 10, Section 4.3, Table 6 and Table 7
Falls whenAn independent human annotation of the same 150 responses gives precision for the LongCite models near or below the 67.5 of GLM-4, or the annotators turn out to be the model developers without blinding beyond anonymized model names. The anchor narrows if 'at least partially supports' is tightened to full support and the precision gap shrinks. Query to run: second human annotation of the released LongBench-Chat responses with reported inter-annotator agreement.
Statement
In the paper's human evaluation, the anonymized responses of three models on the 50 LongBench-Chat queries (150 responses, 1,064 statements, 909 citations) were manually annotated for citation recall and citation precision by the same standard the GPT-4o judge uses. Table 6 prints human scores as recall / precision / F1: LongCite-8B 79.6 / 88.9 / 82.6, LongCite-9B 72.8 / 84.2 / 75.8, GLM-4 61.2 / 67.5 / 60.2. The GPT-4o judge scores for the same responses are lower: 62.0 / 79.7 / 67.4, 57.6 / 78.1 / 63.6 and 47.6 / 53.9 / 47.1. Precision here is the share of cited snippets that at least partially support their statement; recall is whether a statement is supported by its cited snippets. The setting is citation into a document supplied in the prompt, not web research.
Collection
Authors are at Tsinghua University and Zhipu AI. They built the benchmark (LongBench-Cite), the training data (LongCite-45k) and the two LongCite models that lead the table, and Zhipu AI is the developer of GLM-4 and of the GLM-4-9B base of LongCite-9B: collector and proposer of the winning method coincide, recorded here as positioned. The observed arXiv v3 text is marked Preprint; the source list gives Findings of ACL 2025 as venue. Method: the model receives the long context with numbered sentences and must answer with sentence-level citations into that supplied context; there is no web retrieval. GPT-4o judges citation recall (statement fully, partially or not supported by its cited snippets: 1 / 0.5 / 0) and citation precision (each cited snippet relevant or not). The counter-check that exists and was used: a human annotation of 150 LongBench-Chat responses (1,064 statements, 909 citations) from three models; Cohen's kappa between GPT-4o and human 0.593 for recall and 0.655 for precision, GPT-4o accuracy against human labels 75.0% and 88.8%. For this anchor the human annotation is itself the measurement; the paper does not say who the annotators were or report agreement between human annotators.
Falls when
An independent human annotation of the same 150 responses gives precision for the LongCite models near or below the 67.5 of GLM-4, or the annotators turn out to be the model developers without blinding beyond anonymized model names. The anchor narrows if 'at least partially supports' is tightened to full support and the precision gap shrinks. Query to run: second human annotation of the released LongBench-Chat responses with reported inter-annotator agreement.
Reflex
Only an LLM judge says citations are good, humans would find them worse. Too coarse: in this closed-document setting the human scores sit above the GPT-4o judge scores for all three models, by 6 to 14 points in precision.
Evidence
https://arxiv.org/pdf/2409.02897v3 p. 10, Section 4.3, Table 6 and Table 7 | 2024-09-10 (v3; first submitted 2024-09-04) · arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: table-only quote, Table 6, p. 10, Human column P; no sentence carries 88.9, 84.2 or 67.5. Plan 02b: the caption is cut before 'GPT-4o' because the engine's canonical form folds the line-break hyphen of that token.
Findings and answers · 0
No attacker has recorded a finding on this card yet.