Questions › Do the citations of AI research assistants support the claims they are attached to
Vendor-run DRACO scores citation quality at 65 percent for Perplexity Deep Research and 42 to 56 for five rivals
Perplexity Deep Research (with Opus 4.5 or 4.6) demonstrates best performance in all four categories, achieving the highest normalized scores in Factual Accuracy (67.9%), Breadth and Depth of Analysis (73.1%), Presentation Quality (90.3%), and Citation Quality (64.6%).
https://arxiv.org/pdf/2602.11685v1 p. 12, Table 13; p. 6-7, Section 4.1 and Table 4; p. 9, Section 5.1
Falls whenA party with no stake reruns the public DRACO tasks and rubrics and the Citation Quality ordering changes, or a per-claim support audit of the same outputs shows the axis score does not track the share of citations that support their claims. Query to run: independent regrade of the DRACO Citation Quality criteria plus a per-citation support check on the same reports.
Statement
DRACO grades deep research outputs on 100 tasks against task-specific rubrics along four axes. The Citation Quality axis is described as 'Citations to primary source documents' and makes up 12% of criteria (4.8 of 39.3 per task on average). Table 13 prints normalized Citation Quality scores: Perplexity Deep Research (Opus 4.6) 64.6, Perplexity Deep Research (Opus 4.5) 62.5, Claude Opus 4.6 56.2, Gemini Deep Research 51.5, OpenAI Deep Research (o3) 45.8, OpenAI Deep Research (o4-mini) 42.5, Claude Opus 4.5 42.1. Each criterion receives a binary MET or UNMET from an LLM judge (Gemini-3-Pro), averaged over 5 grading runs. The axis scores whether a response cites the primary documents the rubric asks for; it is not a per-claim check of whether each cited passage supports the sentence it is attached to.
Collection
Nine of ten authors are at Perplexity, one at Harvard University; Perplexity is the vendor of the system that ranks first on every axis, the tasks are sampled from Perplexity Deep Research requests, and the rubrics were designed and validated in a process the vendor ran together with The LLM Data Company and 26 recruited domain experts (Section 4.1): collector and beneficiary coincide, recorded as subordinate. arXiv preprint without venue. The judge model was chosen drawing on an internal human-LLM alignment study that is not published in the paper. The counter-check that exists: the dataset and rubrics are public (hf.co/datasets/perplexity-ai/draco), and the paper reports scores under GPT-5.2 and Sonnet-4.5 as alternative judges with stable ranking; an evaluation by a party without a stake is not part of the paper.
Falls when
A party with no stake reruns the public DRACO tasks and rubrics and the Citation Quality ordering changes, or a per-claim support audit of the same outputs shows the axis score does not track the share of citations that support their claims. Query to run: independent regrade of the DRACO Citation Quality criteria plus a per-citation support check on the same reports.
Reflex
The vendor's benchmark shows its deep research product cites best. Too coarse: the axis covers 12 percent of the criteria, measures citation of primary documents rather than support of each claim, is LLM-judged, and tops out at 64.6 percent for the vendor's own system.
Evidence
https://arxiv.org/pdf/2602.11685v1 p. 12, Table 13; p. 6-7, Section 4.1 and Table 4; p. 9, Section 5.1 | 2026-02-12 · arXiv 2602.11685 · Zhong et al. (Perplexity), DRACO Plan 02b: the rivals' 42 to 56 stand only in Table 13, p. 12.