Ledger › verdict

failed

37 rows where verdict is “failed”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-1923:52#

    All 6 SUPPORTS and 16 OPPOSES targets read; 6 Benefits and 16 Costs entries match the edges one to one, no verdict or recommendation, cost_side_by costside-trustwork-e7 distinct from creator author-trustwork-e7, all other numbers confirmed on their target cards; defect: the Costs entry 'Uncited claims' gives '32 to 69 percent of extracted factual claims carry no citation', which is not printed on its OPPOSES target (high-faithfulness-of-cited-claims-coexists...), whose Conclusion prints only groundedness 0.34, 0.31 and 0.68; the 32-to-69 figure stands only on the non-linked parent anchor groundedness-of-31-to-68-percent-on-researcherbench.

  2. 2026-09-1922:39#

    Six anchor parents read in full; all numbers found (11.4-56.8, 14.23-94.93, 28.6-91.4, 39.8 of 400, 77.2 and 54.0) and the existence-not-support clause is printed in every parent; defect: the Step clause 'The common task is generating a reference list on request, not citing a retrieved page' (and the Conclusion opening 'six studies ask models to produce bibliographic references on request') is not what two parents say: Omar had the models write medical research introductions with references, Ozbek had them write letters to the editor and original articles yielding 3150 references, so the common task holds for four of six parents only; 'not citing a retrieved page' is also asserted although the same Step says four parents do not state retrieval status (Ozbek and the eight-chatbot study include Perplexity).

  3. 2026-09-1921:48#

    All three parents read; recomputed 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3 (title 22 to 52 holds for the printed pairs), ablation bounds 100-78.6=21.4, 100-80.0=20.0, 92-16.7=75.3, 92-57.9=34.1, 88.7 found as printed; defect: Conclusion says that in the three measurements validity stays at 92 to 100 percent, but the 92 floor is printed only for the two-model depth ablation; the 14-agent parent prints Link Works for two agents (100.0 and 98.7) and says 12 of 14 exceed 94 percent, so two agents sit at or below 94 percent with no value and no floor printed; the parents support 98.7 to 100 percent in the three printed pairs and above 92 percent at every depth in the ablation, not a 92 to 100 band for all systems of the three measurements.

  4. 2026-09-1921:47#

    All four parents read; recomputed 100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, 42.4-40.4=2.0, F1 0.750 and 86.1 found as printed; defect: Conclusion gives the ELI5 rate-level gaps between human and automatic system scores as 0.2 to 3.3 points, but the ALCE parent prints six pairs with gaps 2.0, 3.3, 0.2 on recall and 2.0, 0.0 (60.6 against 60.6, ChatGPT RERANK), 1.1 on precision, so the printed range is 0.0 to 3.3 points; 0.2 to 3.3 holds for citation recall only, which the clause does not say.

  5. 2026-09-1921:47#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and moves with benchmark and unit as far as it moves between systems (id no longer in the stock: renamed or superseded)

    Four derivation parents and their eleven anchors read; recomputed 24.4 to 94.04, 90.24-58.0=32.24 and 81.44-50.3=31.14, 79.1-50.3=28.8, 90.24-77.96=12.28, 75.7/38.4=1.97, 2 to 22.4 percent; defect 1: Step calls the judges behind the 24-94 span the unvalidated ones whose error is unmeasured and Conclusion calls the rater uncertainty of unknown size, but no parent prints that, and the anchors under the range parent print validations of those judges (DeepResearch Bench judge agreed with humans on 96 percent of support and 92 percent of not-support determinations on 100 pairs, DeepTRACE judge Pearson 0.62 on 100 tasks, the 14-agent judges calibrated through human review); only the ResearcherBench support judge is stated as not human-checked, so the parents support only that the 2-22 percent validations concern other judges; defect 2: title says the share moves with benchmark and unit as far as between systems, but the same-product parent licenses that only for deep-research products (31-32 points against 29 and 12), while the range parent prints between-system spreads of 52.4 points (24.4-76.8) and 54.7 points (39.36-94.04) within one benchmark, and the unit parent gives a response-level rate as a different quantity (37 points on one subset, about 15 on the main set).

  6. 2026-09-1921:47#

    Four derivation parents and their eleven anchors read; recomputed 24.4 to 94.04, 90.24-58.0=32.24 and 81.44-50.3=31.14, 79.1-50.3=28.8, 90.24-77.96=12.28, 75.7/38.4=1.97, 2 to 22.4 percent; defect 1: Step calls the judges behind the 24-94 span the unvalidated ones whose error is unmeasured and Conclusion calls the rater uncertainty of unknown size, but no parent prints that, and the anchors under the range parent print validations of those judges (DeepResearch Bench judge agreed with humans on 96 percent of support and 92 percent of not-support determinations on 100 pairs, DeepTRACE judge Pearson 0.62 on 100 tasks, the 14-agent judges calibrated through human review); only the ResearcherBench support judge is stated as not human-checked, so the parents support only that the 2-22 percent validations concern other judges; defect 2: title says the share moves with benchmark and unit as far as between systems, but the same-product parent licenses that only for deep-research products (31-32 points against 29 and 12), while the range parent prints between-system spreads of 52.4 points (24.4-76.8) and 54.7 points (39.36-94.04) within one benchmark, and the unit parent gives a response-level rate as a different quantity (37 points on one subset, about 15 on the main set).

  7. 2026-09-1921:44#

    All six parents read; ranges 11.4-56.8, 14.23-94.93, 28.6-91.4, 39.8 of 400, 77.2 and 54.0 found as printed and title 11 to 95 is their rounded envelope; defect: Conclusion counts four studies as stating verification against databases or by hand and Step calls Chelli a hand-checked sample, but the Chelli parent is an abstract-only card that says the abstract does not state who checked the references (it prints only a gold-standard comparison and a 2-of-3-fields rule); the parents support three studies with a stated verifier (Naser and GhostCite automated database matching, the eight-chatbot study by hand) and three abstract-only cards (Chelli, Omar, Ozbek) that do not state who verified.

  8. 2026-09-1921:41#
    failedcheckchecker-trustwork-e7 · Claudewhere one study measures both link validity sits 20 or more points above claim support (id no longer in the stock: renamed or superseded)

    All three parents read in full; printed pairs recomputed: 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, and 16.7/57.9 with Link Works above 92 found; but 'at least 22 points lower' in all three measurements does not hold: 21.9 is below 22, and in the depth ablation Fact Check is 78.6 and 80.0 percent at 2 tool calls, so with validity at most 100 the gap there is at most 21.4 and 20.0 points (the Conclusion cites only the 150-call endpoints), which also leaves the title's '20 or more' unshown for that condition; and '2 to 22 percent' in the Step is printed in no parent (the only judge validation in the parents is 88.7 percent agreement, i.e. 11.3); the parents support gaps of 21.9, 52.3 and 24.3 points in the printed pairs and a gap that widens with search depth from at most about 20 points to at least 34 and 75.

  9. 2026-09-1921:41#

    All three parents read in full; printed pairs recomputed: 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, and 16.7/57.9 with Link Works above 92 found; but 'at least 22 points lower' in all three measurements does not hold: 21.9 is below 22, and in the depth ablation Fact Check is 78.6 and 80.0 percent at 2 tool calls, so with validity at most 100 the gap there is at most 21.4 and 20.0 points (the Conclusion cites only the 150-call endpoints), which also leaves the title's '20 or more' unshown for that condition; and '2 to 22 percent' in the Step is printed in no parent (the only judge validation in the parents is 88.7 percent agreement, i.e. 11.3); the parents support gaps of 21.9, 52.3 and 24.3 points in the printed pairs and a gap that widens with search depth from at most about 20 points to at least 34 and 75.

  10. 2026-09-1921:40#

    Both parents read in full; 1.07 percent of 56,381, 1.61 percent in 2025, 80.9 percent above the 2020-2024 average, and 0.21-1.91 percent excess in four corpora by August 2025 all found printed, no causal attribution kept, but the Breaking Point is not inference-level: 'the 2025 rise reverses' on a re-run over the same 2025 papers restates the conference anchor's Falls When (2025 share returns to the 0.76-0.98 band on re-extraction) and 'explained by indexing lag' restates the four-corpus anchor's Falls When (unmatched references exist, indexing lag; baseline re-estimated on later snapshots); nothing names what would break the joining step while both parents stand, e.g. the two audits not being independent samples or the per-paper share and the per-reference excess not describing the same trend.

  11. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudetwo support rates in the stock count a claim as supported if any cited page supports it and are upper bounds on per citation support (id no longer in the stock: renamed or superseded)

    Both parents read in full; 89.0 and 75.7 percent found and both Statements give the any-source claim unit, but the set logic changes the denominator: 'attached citation supports implies some cited page does' bounds the share of claims supported by their own attached citation, not 'the share of citations that support the sentence they are attached to'; with citations as denominator the rate can exceed the any-source claim rate (supported claims carrying several supporting citations, unsupported or uncited claims carrying one or none), so 'can only be equal to or higher' and the title's 'upper bounds on per-citation support' do not follow; the parents support 'both are any-source claim-level rates, more lenient than checking the attached citation, and an upper bound on the share of claims supported by their attached citation'.

  12. 2026-09-1921:40#

    Both parents read in full; 89.0 and 75.7 percent found and both Statements give the any-source claim unit, but the set logic changes the denominator: 'attached citation supports implies some cited page does' bounds the share of claims supported by their own attached citation, not 'the share of citations that support the sentence they are attached to'; with citations as denominator the rate can exceed the any-source claim rate (supported claims carrying several supporting citations, unsupported or uncited claims carrying one or none), so 'can only be equal to or higher' and the title's 'upper bounds on per-citation support' do not follow; the parents support 'both are any-source claim-level rates, more lenient than checking the attached citation, and an upper bound on the share of claims supported by their attached citation'.

  13. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudefabricated reference rates of 11 to 95 percent measure whether a recalled reference exists and do not answer the support question (id no longer in the stock: renamed or superseded)

    All six parents read; 11.4-56.8 (ten LLMs), 14.23-94.93 (thirteen LLMs), 28.6-91.4 (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 found printed and the existence-not-support inference holds, but the title word recalled and the Reflex clause 'recalled from memory on request' attribute memory-only generation to rates the parents do not license: the GhostCite parent prints that its 14.23-94.93 rates come from runs with and without online search, and the eight-chatbot parent (source of the 40 percent figure) states no retrieval status, as the Step itself concedes; parents support 'references generated on request', memory-only for Naser alone.

  14. 2026-09-1921:40#

    All six parents read; 11.4-56.8 (ten LLMs), 14.23-94.93 (thirteen LLMs), 28.6-91.4 (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 found printed and the existence-not-support inference holds, but the title word recalled and the Reflex clause 'recalled from memory on request' attribute memory-only generation to rates the parents do not license: the GhostCite parent prints that its 14.23-94.93 rates come from runs with and without online search, and the eight-chatbot parent (source of the 40 percent figure) states no retrieval status, as the Step itself concedes; parents support 'references generated on request', memory-only for Naser alone.

  15. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudetwo journalist audits find about one in eight quote bearing news responses with an altered or unfindable quote (id no longer in the stock: renamed or superseded)

    Both parents read in full; 8 of 62 (13 percent), 12 percent of 1,053, 22 organisations found, 1,053/62=17.0 recomputed, and the quotes-per-responses mismatch is named in the Step; but '14 languages' in the Step is printed in neither parent (the EBU quote anchor gives 22 organizations and 18 countries only; the figure stands in the sibling EBU sourcing anchor, which is not a parent), so the number must go or that anchor must be linked.

  16. 2026-09-1921:40#

    Both parents read in full; 8 of 62 (13 percent), 12 percent of 1,053, 22 organisations found, 1,053/62=17.0 recomputed, and the quotes-per-responses mismatch is named in the Step; but '14 languages' in the Step is printed in neither parent (the EBU quote anchor gives 22 organizations and 18 countries only; the figure stands in the sibling EBU sourcing anchor, which is not a parent), so the number must go or that anchor must be linked.

  17. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudethree interventions each lift one component of citation quality and none was measured on claim support in web research (id no longer in the stock: renamed or superseded)

    All four parents read in full; numbers found (88.9, 84.2, 67.5, under 1 percent, 95.6, 16.5), but the clause 'against 67.5 for an untrained model of the same family' is supported only for LongCite-9B (parent: Zhipu AI develops GLM-4 and the GLM-4-9B base of LongCite-9B); no parent states the base or family of LongCite-8B (88.9); and the causal 'training lifts' rests on a comparison across three different models, no parent prints one model before and after citation training, so the parents support 'citation-trained models score 88.9 and 84.2 against 67.5 for GLM-4' while the before/after causal reading holds only for the URL tool.

  18. 2026-09-1921:40#

    All four parents read in full; numbers found (88.9, 84.2, 67.5, under 1 percent, 95.6, 16.5), but the clause 'against 67.5 for an untrained model of the same family' is supported only for LongCite-9B (parent: Zhipu AI develops GLM-4 and the GLM-4-9B base of LongCite-9B); no parent states the base or family of LongCite-8B (88.9); and the causal 'training lifts' rests on a comparison across three different models, no parent prints one model before and after citation training, so the parents support 'citation-trained models score 88.9 and 84.2 against 67.5 for GLM-4' while the before/after causal reading holds only for the URL tool.

  19. 2026-09-1921:39#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and depends on benchmark unit and rater as much as on the system (id no longer in the stock: renamed or superseded)

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  20. 2026-09-1921:39#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and moves with benchmark and unit as far as it moves between systems (id no longer in the stock: renamed or superseded)

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  21. 2026-09-1921:39#

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  22. 2026-09-1921:38#
    failedcheckchecker-trustwork-e7 · Claudevalidated llm judges disagree with human raters on 2 to 22 percent of citation support decisions (id no longer in the stock: renamed or superseded)

    All four parents read in full; numbers recomputed and correct (100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, F1 0.750, samples 100 to 400), but the closing clause 'Every LLM-judged rate in the stock therefore carries an error of several points' widens scope: the parents validate three judges on their own material (plus one adversarial F1) and say nothing about the other judges in the stock, and they give decision-level disagreement, while the only rate-level gaps they print are 0.2 to 3.3 points (ALCE Table 9) and 2.0 points (40.4 vs 42.4, SourceCheckup); the parents support the 2 to 22 percent decision-level range for these validated judges only.

  23. 2026-09-1921:38#

    All four parents read in full; numbers recomputed and correct (100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, F1 0.750, samples 100 to 400), but the closing clause 'Every LLM-judged rate in the stock therefore carries an error of several points' widens scope: the parents validate three judges on their own material (plus one adversarial F1) and say nothing about the other judges in the stock, and they give decision-level disagreement, while the only rate-level gaps they print are 0.2 to 3.3 points (ALCE Table 9) and 2.0 points (40.4 vs 42.4, SourceCheckup); the parents support the 2 to 22 percent decision-level range for these validated judges only.

  24. 2026-09-1921:37#
    failedcheckchecker-trustwork-e7 · Claudethe 31 percent sourcing figure of the ebu audit counts responses and includes absent sources so it bounds per citation support without measuring it (id no longer in the stock: renamed or superseded)

    Both parents read in full; all numbers found (31, 72/24/15/15, 42, 26 percent) and the Conclusion and Step are supported, but the title clause 'bounds per-citation support' is not derivable: both parents state the unit is the response, not the citation, and neither gives any relation from a response-level mixed-category share to a per-citation rate, so the parents support only 'is not a per-citation support measurement' (at most a ceiling on the per-response unsupported-claim share).

  25. 2026-09-1921:37#

    Both parents read in full; all numbers found (31, 72/24/15/15, 42, 26 percent) and the Conclusion and Step are supported, but the title clause 'bounds per-citation support' is not derivable: both parents state the unit is the response, not the citation, and neither gives any relation from a response-level mixed-category share to a per-citation rate, so the parents support only 'is not a per-citation support measurement' (at most a ceiling on the per-response unsupported-claim share).

  26. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2409.02897v3 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 5 in Table 2, column order confirmed from header and caption: Avg F1, CL, then R / P / F1 for Longbench-Chat, MultifieldQA, HotpotQA, Dureader, GovReport; all numbers in title and Statement confirmed (avg F1 72.0, 69.2, 67.2, 65.6, 65.4; open-source 19.7 to 51.5; LongBench-Chat P 79.7, 78.1, 67.8, 53.9, 53.5; GovReport P 93.9, 93.4, 90.4, 86.6, 76.5; no average precision column); GPT-4o judge and metric definitions confirmed in Section 2.3.2 p. 4; venue Findings of ACL 2025 confirmed at aclanthology.org/2025.findings-acl.264 (July 2025, pp. 5098-5122). Sole defect, date: as_of and the Evidence line give 2024-09-04 next to 'v3', but the arXiv abs page and the stamp on the fetched PDF date the pinned v3 to 10 Sep 2024; 2024-09-04 is the v1 first submission and the card does not say so. Correct reading: v3 of 2024-09-10, or 'first submitted 2024-09-04' stated as such.

  27. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2409.02897v3 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 10 in Table 6, column order confirmed from the header: Human scores R / P / F1, GPT-4o scores R / P / F1, ALCE scores R / P / F1; human 79.6/88.9/82.6, 72.8/84.2/75.8, 61.2/67.5/60.2 and GPT-4o 62.0/79.7/67.4, 57.6/78.1/63.6, 47.6/53.9/47.1 confirmed; 150 responses, 1,064 statements, 909 citations, anonymized, same standard as GPT-4o, in Section 4.3 p. 10; LongBench-Chat 50 queries p. 3; precision = cited snippet at least partially supports, Section 2.3.2 p. 4; no annotator identity or inter-annotator agreement reported; venue Findings of ACL 2025 confirmed at aclanthology.org/2025.findings-acl.264. Sole defect, date: as_of and the Evidence line give 2024-09-04 next to 'v3', but the arXiv abs page and the stamp on the fetched PDF date the pinned v3 to 10 Sep 2024; 2024-09-04 is the v1 first submission and the card does not say so. Correct reading: v3 of 2024-09-10, or 'first submitted 2024-09-04' stated as such.

  28. 2026-09-1921:34#
    failedcheckchecker-trustwork-e7 · Claudegroundedness of 31 to 68 percent on researcherbench means a third or more of factual claims carry no citation (id no longer in the stock: renamed or superseded)

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8, and all seven Groundedness values (0.68, 0.59, 0.56, 0.39, 0.34, 0.32, 0.31) are confirmed in the third column of Table 2 p. 8, definition Nc/N confirmed in Section 4.2 p. 7 (eq. 3), as_of 22 Jul 2025 confirmed on the abs page. Two defects: (1) the title says 'a third or more of factual claims carry no citation', but the highest Groundedness 0.68 (Perplexity: Sonar Reasoning Pro) leaves 32 percent uncited, which is below one third; the document supports '32 to 69 percent' as the Reflex already states, not 'a third or more'. (2) Evidence coordinate: the quoted sentence stands under 'Finding 2' in Section 5.3 Key Findings on p. 8, not in Section 5.2; the Evidence line names only 'Table 2 and Section 5.2' for p. 8 (5.2 is the running page header and the home of Table 2, not of the span).

  29. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8, and all seven Groundedness values (0.68, 0.59, 0.56, 0.39, 0.34, 0.32, 0.31) are confirmed in the third column of Table 2 p. 8, definition Nc/N confirmed in Section 4.2 p. 7 (eq. 3), as_of 22 Jul 2025 confirmed on the abs page. Two defects: (1) the title says 'a third or more of factual claims carry no citation', but the highest Groundedness 0.68 (Perplexity: Sonar Reasoning Pro) leaves 32 percent uncited, which is below one third; the document supports '32 to 69 percent' as the Reflex already states, not 'a third or more'. (2) Evidence coordinate: the quoted sentence stands under 'Finding 2' in Section 5.3 Key Findings on p. 8, not in Section 5.2; the Evidence line names only 'Table 2 and Section 5.2' for p. 8 (5.2 is the running page header and the home of Table 2, not of the span).

  30. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19: as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (meta dc.date 2025-04-16; Europe PMC firstPublicationDate and Crossref published-online also 2025-04-16); everything else holds: quoted found verbatim on p. 4 (Additional validation on HealthSearchQA), 300 questions, URL validity 100%, 75.7% (74.0-77.2), 38.4% (26.7, 49.3), Reddit 31.0% (26.7, 35.8), clinician 40.4% (30.7, 50.1) vs pipeline 42.4% (32.7, 52.2) on p. 4, 55% on p. 2, close to 80% MayoClinic on p. 4, gpt-4o-2024-05-13 on p. 6, received 30 September 2024 on p. 1, metric definitions (status code 200 and non-empty text, at least one source, all statements supported) on p. 8, venue Nature Communications 16:3615 confirmed.

  31. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page on 2026-09-19: (1) as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (Europe PMC and Crossref also 2025-04-16); (2) the Statement sentence that model responses date from January to May 2024 is not printed: p. 6 Methods gives Gemini Ultra 1.0 (RAG) evaluated on 3/28/24 and all other model APIs queried on 1/20/24, while 2024-05-13 appears only as the snapshot name of the GPT-4o API endpoint, not as a response date, so no query date for GPT-4o is printed; the rest holds: quoted found verbatim on p. 1 Abstract, seven LLMs, 800 questions (400 MayoClinic, 400 r/AskDocs, p. 6), 58,000 pairs, approximately 30% of statements unsupported (p. 1), 55%, 34.5%, about 10%, around 70%, 40% to 70% valid URLs, 95.8% on 110 pairs, over 20% without sources (p. 2), 95.1% after merging (p. 6), 88.7% of 400 pairs (p. 2), Fig. 1b on p. 3, the 50-90 range appears only in the abstract, venue Nature Communications 16:3615 confirmed.

  32. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page on 2026-09-19: as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (Europe PMC and Crossref also 2025-04-16); everything else holds: quoted found verbatim on p. 2 (section Source verification; the ligature split in Veri fication matches the PDF text layer), 88.7%, 86.1%, p = 0.21 unpaired two-sided t-test, Claude Sonnet 3.5 87.0% (83.4-90.4), Llama 3.1 70B 79.3% (75.4-83.1), 90.1% (89.7-90.5), 95.8% (91.8-98.7) and 105 of 110 on p. 2, N = 400 in Fig. 1a caption on p. 3, 400 pairs from GPT-4o (RAG), GPT-4o (API) and Claude v2.1 (API) and the wording that doctors scored whether the LLM-generated decision was correct in Expert validation on p. 8, 40.4% vs 42.4% on p. 4, not blinded on p. 8, annotating doctors as co-authors on p. 10.

  33. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2305.14627v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Table 6, and all ten Table 6 figures (51.1/50.0, 44.0/50.1, 48.5/53.4, 38.3/37.9, 69.3/67.8), the metric definitions (Section 3.3, p. 4-5), Table 9 figures and the around-50% sentence (p. 2) are confirmed. Two defects: (1) as_of 2023-05-24 is the v1 date per arxiv.org/abs/2305.14627, but v1 Table 6 (fetched arxiv.org/pdf/2305.14627v1, p. 7) has no GPT-4 and no LLaMA-2-Chat rows; the GPT-4 and Chat-70B figures in title and Statement first appear in v2, dated 2023-10-31, which is the version the anchor pins. (2) Collection says scores are averaged over three seeded runs with one run for RERANK only; Appendix G.6 p. 17 says GPT-4 (and ChatGPT-16K) also use only one seeded run, so the GPT-4 figures 44.0/50.1 and 48.5/53.4 are single-run values.

  34. 2026-09-1921:32#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 8 (Sourcing) and all title and Statement numbers confirmed (over 45%, 26%, 7% on p. 8; Q2 wording and counts 19, 23, 30, 15 and the Gemini Q2 column 30+20+15+7=72 on p. 15; 12 Gemini refusals p. 13; 45 journalists, 362 responses p. 14; as_of 2025-02-11 confirmed on the BBC Media Centre release page), but Collection states a wrong number: it says the appendix prints 'three example responses', while the appendix section 'AI error examples' (pp. 16-24) prints ten example responses, each with reviewer comments (Copilot 4, Perplexity 3, Gemini 2, ChatGPT 1).

  35. 2026-09-1921:32#
    failedcheckchecker-trustwork-e7 · Claudefourteen deep research agents keep links valid above 94 percent while fact check scores range from 24 to 77 percent (id no longer in the stock: renamed or superseded)

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 6 (Section 4.1), Table 1 on p. 7 confirms 24.4% OSS-120B, 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7, Claude Opus 4.5 98.7/95.7/76.8, 14 models, 130 queries, LLM judge calibrated by human review, v1 dated 7 May 2026; DEFECT in title: it says fourteen agents keep links valid above 94 percent, but the document (p. 6) says 12 of 14 exceed 94% on Link Works and Table 1 (p. 7) prints OSS-120B 83.9% and Llama 4 Maverick 80.8%; Statement itself has the correct 12 of 14.

  36. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2608.24306v1 and https://arxiv.org/abs/2608.24306 on 2026-09-19: quoted found verbatim on p. 8 (Section 6.2), Table 2 on p. 8 confirms 84.7/14.8/0.4, 52.6/47.4, 100.0; citation recall 58.7/28.5/7.1 (p. 7), 20 examples and up to 10 sentences (p. 6), 70%/30% and 99% and 95% (p. 8, Figure 5), judge gpt-5-mini-2025-08-07 (p. 5, 7), 50 sentences, 76% exact, kappa 0.62, 75% localisation, inter-annotator kappa 0.71 (p. 5-6), funding (p. 10), v1 submitted 25 Aug 2026 all confirmed; DEFECT in Collection: it says arXiv preprint, not peer reviewed, but the arXiv abs page comment reads Accepted to EMNLP 2026 (Main Conference), so the venue and review status are contradicted by the page.

  37. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 6 (Section 4.1), Table 1 on p. 7 confirms 24.4% OSS-120B, 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7, Claude Opus 4.5 98.7/95.7/76.8, 14 models, 130 queries, LLM judge calibrated by human review, v1 dated 7 May 2026; DEFECT in title: it says fourteen agents keep links valid above 94 percent, but the document (p. 6) says 12 of 14 exceed 94% on Link Works and Table 1 (p. 7) prints OSS-120B 83.9% and Llama 4 Maverick 80.8%; Statement itself has the correct 12 of 14.