Ledger › verdict

held

11 rows where verdict is “held”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-2223:24#

    The EBU parent defines the 31% as a response-level journalist rating that covers unsupported claims, no sources at all, and incorrect sourcing claims, with Gemini at 72% and 42% of its responses giving no direct source. The BBC parent shows the earlier Q2 already mixed misattribution, unsupported claims and missing sources, and names lack of sources as part of the cause for the assistant with the highest rate (26% of Gemini responses). The step only narrows what the 31% can be cited for.

  2. 2026-09-2223:24#

    The judge-benchmark parent prints false-negative rates of 0.183 to 0.470, over-rejection of supported citations, and the LongCite parent prints GPT-4o precision of 79.7, 78.1 and 53.9 against human 88.9, 84.2 and 67.5 on the same responses. The AI Overviews parent records both mismatches in the 100 verdicts as the verifier being too generous, and the step does not average the two directions: tasks and base rates differ, and two cases are not a rate.

  3. 2026-09-2223:24#

    The disagreements are the complements of printed agreement: 100 minus 98 is 2 on the AI Overviews sample, 100 minus 88.7 is 11.3 beside inter-doctor agreement of 86.1, and 100 minus 85.1 and 100 minus 77.6 are 14.9 and 22.4 for ALCE accuracy. The six ELI5 human-minus-automatic gaps in the ALCE parent run from 0.0 to 3.3 points and the SourceCheckup end-to-end gap is 42.4 minus 40.4, and the F1 parent is kept separate because 0.750 is not an agreement rate.

  4. 2026-09-2223:24#

    Both parents state an any-source unit: the AI Overviews parent checks each claim against all cited pages (89.0% consistent), and the SourceCheckup parent counts a statement supported if any source in the response supports it (75.7%). If the attached citation supported the claim then some cited page would, but not the reverse, and the step's example of 50% of claims against 75% of citations shows the denominators are not bounds on each other. Neither parent prints the attached-citation rate.

  5. 2026-09-2223:24#

    DeepTRACE's Table 1 prints Perplexity Deep Research at 58.0 and Gemini Deep Research at 50.3, DeepResearch Bench prints 90.24 and 81.44, and ResearcherBench prints faithfulness 0.85 and 0.86, so the 31-to-32-point gaps are 90.24 minus 58.0 and 81.44 minus 50.3. Within DeepTRACE the deep-research column runs from 79.1 to 50.3 and within DeepResearch Bench from 90.24 to 77.96, and the step states the version-change premise instead of picking a benchmark.

  6. 2026-09-2223:24#

    The BBC parent is eight altered or absent quotes over 62 responses, which the report calls 13% and which divides quotes by responses, and the EBU parent is 12% of 1,053 quote-bearing responses, about seventeen times the BBC sample. The BBC took part in both audits, and the step records the 12-to-13% numerical agreement while saying the two are not the same quantity.

  7. 2026-09-2223:24#

    The HealthSearchQA parent prints both units on one sample, 75.7% of statements (74.0–77.2) and 38.4% of responses (26.7–49.3), and the seven-model parent prints the same system's pair on the 800-question set, about 30% of statements unsupported and 55% response-level support, under the rule that a response counts only if every statement does. The step does not invent a statements-per-response count the parents do not print.

  8. 2026-09-2223:24#

    Liu's parents give 2023 human precision of 74.5% on average and 63.6 to 89.5 by engine, ALCE gives the automatic baseline near 50% on research pipelines, and DeepTRACE gives 39.8 to 68.3% in August 2025 under a different rater and query set, so refusing a cross-study trend follows. The BBC parent is the only repeated design among the parents, Gemini staying at 47% and the other three falling to 10–15%, with the caveats the step uses: paid tiers to free, adjusted definitions, samples of 362 and 237.

  9. 2026-09-2223:24#

    Each parent states its own task: the two Tow Center anchors are reverse attribution of a given excerpt (more than 60% of 1,600 queries, and 153 of 200 for ChatGPT Search), the Grok 3 anchor is error-page resolution (154 of 200), and the DRACO anchor, published by the vendor of the top system, grades whether rubric primary documents are cited (42.1 to 64.6). None of those statements is a share of citations that support the attached claim, so sorting them out of the support range follows.

  10. 2026-09-2223:24#

    The conference parent gives the hand-confirmed per-paper share, 604 of 56,381 papers (1.07%) and 1.61% in 2025 against a 0.89% average for 2020–2024, and the four-corpus parent gives the automated excess over the pre-LLM baseline, 0.21% to 1.91% of references as of August 2025. The step keeps the units apart and keeps both authors' refusal to name a tool.

  11. 2026-09-2223:24#

    The faithfulness and groundedness parents are one ResearcherBench table, so OpenAI Deep Research at 0.84 and 0.34, Grok3 DeeperSearch at 0.80 and 0.31, and Sonar Reasoning Pro at 0.62 and 0.68 are paired, and the volume parent separately prints 4.35 to 111.21 effective citations and 94.04 accuracy with 9.78 against 81.44 with 111.21. A support rate over cited claims is silent on uncited claims and on citation count, and the step marks the 111.21-to-4.35 ratio as computed here.