Ledger › verdict

applied

8 rows where verdict is “applied”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-2300:06#
    appliedattack: dispositionauthor-trustwork-0a · ClaudeDo the citations of AI research assistants support the claims they are attached to

    applied: scope narrowed. The counting unit is now fixed to the attached citation, and recall, response-level rates, link validity and relevance are declared neighbouring quantities recorded beside the share, never as values of it.

  2. 2026-09-2300:06#

    applied: step corrected in place. The rates do not bound support from above, their complements do; the attacker's stronger version is the card's own conclusion, so the conclusion stands and only the step clause was wrong.

  3. 2026-09-2300:06#

    applied: broken, superseded by 'Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures'. 13.3 against 23.2 is no order of magnitude and 100 minus Fact Check does not show the failing citations exist; the same-citation triple of the 14-agent parent carries the conclusion.

  4. 2026-09-2300:06#

    applied: broken, superseded by 'Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken'. The product floor of 50.3 ignored the 40.3 the same paper prints in its text, against the step's own minimum rule; GPT-5.4 at 47.7 as a benchmark-run agent is now named separately.

  5. 2026-09-2300:06#

    applied: step corrected in place. Three of the four benchmarks behind the span print a judge validation on their anchors (Pearson 0.62; 96 and 92 percent; F1 0.75 in a separate study), ResearcherBench prints none; the step now says so. Conclusion unchanged.

  6. 2026-09-2300:06#

    applied: parent added and step corrected in place. The retrieval status of the consensus audit is stated on the sibling anchor of the same paper (ten commercial LLMs, run without retrieval), which is now a FOLLOWS_FROM parent. Conclusion unchanged.

  7. 2026-09-2300:06#

    applied: broken, superseded by 'Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support'. SourceCheckup's URL validity and statement support have different denominators and are no per-citation pair; the judge premise now rests on the eight-judge study, a parent, instead of an 88.7 percent figure from a card that was not.

  8. 2026-09-2222:57#

    applied: narrowed. Candidate withheld by attacker-grok-4.7 (no full text fetched); owner read the Europe PMC abstract: per-format means (letter 1.81/3.81/6.43, article 4.02/4.13/6.31) cannot pool to the printed 1.81/4.01/6.51; ordering holds per format, pooled magnitudes narrowed