Questions › Do the citations of AI research assistants support the claims they are attached to
Citation-trained LongCite-8B reaches citation F1 of 72 on LongBench-Cite against 65 to 67 for three proprietary models
LongCite-8B and LongCite-9B even attain higher citation F1 than the data construction pipeline CoF (72.0 and 69.2 v.s. 65.8), implying a potential for continuous self-improvement.
https://arxiv.org/pdf/2409.02897v3 p. 5, Table 2; p. 4, Section 2.3.2
Falls whenAn evaluation by a group that did not train the models, on long documents outside the five LongBench-Cite dataset groups, finds the citation-trained models no better in citation F1 than the general proprietary models; or a human-rated rerun reverses the ranking. The anchor does not transfer to web research: it narrows to nothing for the stock's question if support rates over supplied documents and over retrieved pages are shown to be unrelated. Query to run: third-party LongBench-Cite style evaluation with human-rated recall and precision.
Statement
On LongBench-Cite (long-context question answering and summarization over a supplied document, five dataset groups: LongBench-Chat, MultifieldQA, HotpotQA, Dureader, GovReport), Table 2 prints average citation F1 and, per dataset, citation recall (R), citation precision (P) and F1, in that order. Average citation F1: LongCite-8B 72.0, LongCite-9B 69.2, Claude-3-sonnet 67.2, GPT-4o 65.6, GLM-4 65.4; the open-source models without citation training range from 19.7 (Llama-3.1-8B-Instruct) to 51.5 (Mistral-Large-Instruct). Citation precision on LongBench-Chat: LongCite-8B 79.7, LongCite-9B 78.1, Claude-3-sonnet 67.8, GLM-4 53.9, GPT-4o 53.5; on GovReport: Claude-3-sonnet 93.9, GLM-4 93.4, GPT-4o 90.4, LongCite-8B 86.6, LongCite-9B 76.5. An average precision column is not printed. Scores are assigned by GPT-4o as judge. The setting is citation into a document given in the prompt, not citation of retrieved web pages.
Collection
Authors are at Tsinghua University and Zhipu AI. They built the benchmark (LongBench-Cite), the training data (LongCite-45k) and the two LongCite models that lead the table, and Zhipu AI is the developer of GLM-4 and of the GLM-4-9B base of LongCite-9B: collector and proposer of the winning method coincide, recorded here as positioned. The observed arXiv v3 text is marked Preprint; the source list gives Findings of ACL 2025 as venue. Method: the model receives the long context with numbered sentences and must answer with sentence-level citations into that supplied context; there is no web retrieval. GPT-4o judges citation recall (statement fully, partially or not supported by its cited snippets: 1 / 0.5 / 0) and citation precision (each cited snippet relevant or not). The counter-check that exists and was used: a human annotation of 150 LongBench-Chat responses (1,064 statements, 909 citations) from three models; Cohen's kappa between GPT-4o and human 0.593 for recall and 0.655 for precision, GPT-4o accuracy against human labels 75.0% and 88.8%.
Falls when
An evaluation by a group that did not train the models, on long documents outside the five LongBench-Cite dataset groups, finds the citation-trained models no better in citation F1 than the general proprietary models; or a human-rated rerun reverses the ranking. The anchor does not transfer to web research: it narrows to nothing for the stock's question if support rates over supplied documents and over retrieved pages are shown to be unrelated. Query to run: third-party LongBench-Cite style evaluation with human-rated recall and precision.
Reflex
Language models cannot cite accurately. Too coarse: with the source document supplied and the model trained to cite sentences, GPT-4o-judged citation precision lies between 72 and 93 percent by dataset, and general models reach above 90 on summarization of a supplied report.
Evidence
https://arxiv.org/pdf/2409.02897v3 p. 5, Table 2; p. 4, Section 2.3.2 | 2024-09-10 (v3; first submitted 2024-09-04) · arXiv 2409.02897 v3, Findings of ACL 2025 · Zhang et al., LongCite Plan 02b: the proprietary F1 values 65.6, 67.2 and 65.4 stand only in Table 2, p. 5.
Findings and answers · 0
No attacker has recorded a finding on this card yet.