Ledger › verdict

finding

9 rows where verdict is “finding”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-2306:27#

    https://arxiv.org/pdf/2604.03173v1 p. 8 Table 3 | DEAD 0.1 to 0.6%, LIKELY HALLUCINATED 0.4 to 1.8%, UNKNOWN 10.3 to 20.2% with 11.0% [8.5, 13.7] of sampled UNKNOWN genuinely dead, as the card prints them | Falls When clause met: DEAD plus LIKELY HALLUCINATED plus the dead share of UNKNOWN exceeds 1% for every model (about 3.5 GPT-5.1, 1.8 Gemini, 1.6 Claude by the stranger's sum). Finding of the plan 04 stranger session, which read only the exported HTML; arithmetic unverified against the paper; open for the owner's disposition pass

  2. 2026-09-2223:25#

    The title asks for one share: of the citations attached to factual claims, how many point to a passage that supports that claim, as distinct from a link that resolves or a source that is merely on topic. The scope does not fix that share. It admits citation recall, link validity and topical relevance as measures, and it never chooses citation-level against claim-level or response-level, nor an attached citation against any page in the reference set, choices that move published figures by tens of points. Not Asked rightly keeps out vendor motive, user belief, legal liability and a ranking, and fabricated references as their own quantity are rightly out because existence is not support. A reader of the title expects the attached-citation rate, and the scope's wider measure list does not hold the answer to that rate.

  3. 2026-09-2223:24#

    The step says the consensus anchor contributes existence for references recalled without retrieval, but parent stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one does not state that the ten models ran without retrieval. Stronger: the LongCite parents compare different models on a supplied document, human precision 88.9 and 84.2 against 67.5, the URL parent is a before-and-after of link resolution and says claim support was not measured, and the 95.6% figure is a database match for titles named by three models, so none of the three is claim support in web research.

  4. 2026-09-2223:24#

    The step treats SourceCheckup's 100% valid URLs against 75.7% supported statements as the same gap as the per-citation pairs, but those two percentages have different denominators, and it cites an 88.7% agreement with a three-doctor consensus that none of the three parents print. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent gives same-citation gaps of 21.9 points (98.7 against 76.8) and 52.3 (100.0 against 47.7), and the depth parent only bounds the gap because Link Works is printed as above 92% rather than as a paired value.

  5. 2026-09-2223:24#

    The step says it adds no premise beyond the four parents, then asserts that the judges behind the 24.4–94.04 span have their own validations printed on their anchors. The judge parent's conclusion only places 2%, 11.3% and 14.9–22.4% on its own judges, and the range parent's conclusion lists ResearcherBench inside the span without a human validation. Stronger: the 2–22% band stays with the judges named in parent stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions, and nothing in the four parents prints a human validation for the judges behind that span, so the band is neither a correction to it nor already measured on it.

  6. 2026-09-2223:24#

    The step calls 3.0–13.3% hallucinated URLs and 23.2–75.6% support failures an order-of-magnitude gap across two samples, but 13.3 against 23.2 is not an order of magnitude, and 100 minus Fact Check does not show that the failing citations exist. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent already scores both on the same citations, Link Works above 94% for 12 of 14 beside Fact Check 24.4–76.8, so most of those support failures are resolving pages.

  7. 2026-09-2223:24#

    The step says the invalid shares in the four reference-list studies bound support from above, but parent stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations prints a hallucination rate of 11.4%, and a non-existence rate is a lower bound on citations that cannot support a claim, so support is at most one minus that rate. Stronger: those rates measure existence or bibliographic match, not support, and only their complements are loose upper bounds, because a matched reference need not support the attached claim.

  8. 2026-09-2223:24#

    The ablation parent shows Fact Check falling between the endpoints while Link Works and Relevant Content stay above 92%, not that they are unchanged, and the localisation parent places the orchestrator as the origin of errors only in three open pipelines. Stronger: parent stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 supports a drop with tool-call budget for two commercial models while links keep resolving, and parent stocks/ai-citations--orchestrator-originates-85-percent-of-final-report-errors-in-ai-q-53-percent-in-ms-agent-and-100-in-trajectorykit does not measure those models, so the premise that they fail at the orchestrator is in neither parent.

  9. 2026-09-2223:24#

    The step says it takes the minimum of printed values for systems the parents label deep research, but parent stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported prints Gemini Deep Research at 40.3% in the running text and 50.3% in Table 1, and parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent prints GPT-5.4, run as a deep-research agent, at 47.7% Fact Check. Stronger: 24.4 to 94.04 is the min and max of the per-citation and per-cited-claim rates, while a commercial deep-research floor is not 50.3 once those lower printed values are kept.