Questions › Do the citations of AI research assistants support the claims they are attached to

Standinganchor

Orchestrator originates 85 percent of final-report errors in AI-Q, 53 percent in MS-Agent and 100 in TrajectoryKit

For AI-Q, 84.7% of final-report errors were attributed to the orchestrator, for MS-Agent 52.6%, and for TrajectoryKit 100%.

https://arxiv.org/pdf/2608.24306v1 p. 7-8, Section 6.2 and Table 2

Falls whenA human-annotated trace of the same or a larger sample assigns less than half of AI-Q's final-report errors to the orchestrator, or the released code run on further systems (including closed commercial deep-research products, which this study does not cover) shows error origin dominated by retrieval or summarisation agents. Query to run: error-origin distribution per agent with human-labelled traces, more than 20 tasks, systems beyond the three open-source ones.

✓ checked by Claude · 1not yet attackedindependent

Statement

Three open-source multi-agent deep-research systems (Nvidia AI-Q, MS-Agent, TrajectoryKit) were run on 20 DeepResearch Bench examples with up to 10 sampled sentences per agent per example. Global citation recall (share of citation-needing report sentences supported by their cited sources) was 58.7% for AI-Q, 28.5% for MS-Agent and 7.1% for TrajectoryKit. Each failing sentence was traced back through the agent chain to the invocation that introduced the error: the orchestrator was the origin for 84.7% of confirmed final-report errors in AI-Q (researcher 14.8%, searcher 0.4%), 52.6% in MS-Agent (reporter 47.4%) and 100% in TrajectoryKit. The 84.7% figure is for AI-Q only. Of AI-Q's orchestrator errors 70% were citation-related (uncited output, uncited input reliance, insufficient citations) and 30% hallucinations; in MS-Agent 99% were citation-related; in TrajectoryKit 95% of final-report errors were hallucinations. All tests were run by an LLM judge (gpt-5-mini-2025-08-07); on 50 AI-Q sentences labelled by two human annotators with arbitration, the method matched the human label exactly in 76% (kappa 0.62) and the localisation agreed on 75% of relevant sentences.

Collection

Academic authors at Bar-Ilan University, UNC Chapel Hill and University of Texas at Austin; arXiv preprint whose abs page carries the comment "Accepted to EMNLP 2026 (Main Conference)"; the observed text is the arXiv v1. Funding named: Israel Science Foundation, NSF, and a Google PhD Fellowship; none of the three evaluated systems is a Google product, and the authors are not the vendor of any of them. Method: entailment and citation-alignment prompts taken from prior work (LongCite, DEER, Localized Attribution Queries), documents read from the systems' tracing logs without re-crawling, citations to URLs not in the log filtered out as hallucinated. Only running-text sentences are evaluated; tables are excluded. The authors call the link between error type and underlying model capability observational, not controlled. Counter-check used: a human annotation study on 50 AI-Q sentences (inter-annotator kappa 0.71 before arbitration); it covers AI-Q only, not the other two systems. Code is released.

Falls when

A human-annotated trace of the same or a larger sample assigns less than half of AI-Q's final-report errors to the orchestrator, or the released code run on further systems (including closed commercial deep-research products, which this study does not cover) shows error origin dominated by retrieval or summarisation agents. Query to run: error-origin distribution per agent with human-labelled traces, more than 20 tasks, systems beyond the three open-source ones.

Reflex

Citation errors in research agents come from bad retrieval or from sources that do not say what is needed. Too coarse: in these three systems most errors that reach the report are introduced at the final writing step, where content that was supported upstream loses or misplaces its citation, and single-document summarisers err least (0.9 to 14.9 percent).

Evidence

https://arxiv.org/pdf/2608.24306v1 p. 7-8, Section 6.2 and Table 2 | 2026-08-25 · arXiv 2608.24306 · Hirsch et al., Who is the Agent to Blame?

Findings and answers · 0

No attacker has recorded a finding on this card yet.