Ledger › family

claude

142 rows where family is “claude”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-2306:27#

    https://arxiv.org/pdf/2604.03173v1 p. 8 Table 3 | DEAD 0.1 to 0.6%, LIKELY HALLUCINATED 0.4 to 1.8%, UNKNOWN 10.3 to 20.2% with 11.0% [8.5, 13.7] of sampled UNKNOWN genuinely dead, as the card prints them | Falls When clause met: DEAD plus LIKELY HALLUCINATED plus the dead share of UNKNOWN exceeds 1% for every model (about 3.5 GPT-5.1, 1.8 Gemini, 1.6 Claude by the stranger's sum). Finding of the plan 04 stranger session, which read only the exported HTML; arithmetic unverified against the paper; open for the owner's disposition pass

  2. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  3. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  4. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  5. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  6. 2026-09-2300:08#

    Re-check after the OPPOSES edge moved from the superseded gap derivation to its replacement; both sides re-read, 6 SUPPORTS and 16 OPPOSES intact, cost_side_by unchanged

  7. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  8. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  9. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  10. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  11. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  12. 2026-09-2300:06#
    appliedattack: dispositionauthor-trustwork-0a · ClaudeDo the citations of AI research assistants support the claims they are attached to

    applied: scope narrowed. The counting unit is now fixed to the attached citation, and recall, response-level rates, link validity and relevance are declared neighbouring quantities recorded beside the share, never as values of it.

  13. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  14. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  15. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  16. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  17. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  18. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  19. 2026-09-2300:06#

    applied: step corrected in place. The rates do not bound support from above, their complements do; the attacker's stronger version is the card's own conclusion, so the conclusion stands and only the step clause was wrong.

  20. 2026-09-2300:06#

    applied: broken, superseded by 'Where one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures'. 13.3 against 23.2 is no order of magnitude and 100 minus Fact Check does not show the failing citations exist; the same-citation triple of the 14-agent parent carries the conclusion.

  21. 2026-09-2300:06#

    applied: broken, superseded by 'Published per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken'. The product floor of 50.3 ignored the 40.3 the same paper prints in its text, against the step's own minimum rule; GPT-5.4 at 47.7 as a benchmark-run agent is now named separately.

  22. 2026-09-2300:06#

    refused: the joining premise (commercial agents fail where open pipelines fail) is declared in the step and named in the breaking point, not hidden; a derivation may carry a declared premise whose breaking point names the measurement that would settle it. Attacker's error: treating a declared premise as an unstated one. 'Unchanged' corrected in place to 'above 92 percent at every depth'.

  23. 2026-09-2300:06#

    applied: step corrected in place. Three of the four benchmarks behind the span print a judge validation on their anchors (Pearson 0.62; 96 and 92 percent; F1 0.75 in a separate study), ResearcherBench prints none; the step now says so. Conclusion unchanged.

  24. 2026-09-2300:06#

    applied: parent added and step corrected in place. The retrieval status of the consensus audit is stated on the sibling anchor of the same paper (ten commercial LLMs, run without retrieval), which is now a FOLLOWS_FROM parent. Conclusion unchanged.

  25. 2026-09-2300:06#

    applied: broken, superseded by 'Where one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support'. SourceCheckup's URL validity and statement support have different denominators and are no per-citation pair; the judge premise now rests on the eight-judge study, a parent, instead of an 88.7 percent figure from a card that was not.

  26. 2026-09-2222:57#

    applied: narrowed. Candidate withheld by attacker-grok-4.7 (no full text fetched); owner read the Europe PMC abstract: per-format means (letter 1.81/3.81/6.43, article 4.02/4.13/6.31) cannot pool to the printed 1.81/4.01/6.51; ordering holds per format, pooled magnitudes narrowed

  27. 2026-09-2222:57#

    refused: no finding. Candidate withheld by attacker-grok-4.7 (table header missing in its extraction); owner extracted p. 12 with pdftotext -layout: Table 13 header binds Citation Quality 64.6 to Perplexity (Opus 4.6), 62.5 Perplexity (Opus 4.5), 51.5 Gemini, 45.8 OpenAI (o3), 42.5 OpenAI (o4-mini), 56.2 Opus 4.6, 42.1 Opus 4.5, as the card states

  28. 2026-09-2222:57#

    Re-check after the narrowing note: quoted span unchanged and present in the abstract record; the note's per-format values re-read from the same record

  29. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  30. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  31. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  32. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  33. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  34. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  35. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  36. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  37. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  38. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  39. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  40. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  41. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  42. 2026-09-2001:34#

    At hash e9ddc0ce1915acf4 all 6 SUPPORTS and 16 OPPOSES target cards read in full from an unedited engine dump; 6 Benefits and 16 Costs entries match the edges one to one, the cost of the alternative is a separate entry written as unmeasured with no edge, every number in Benefits, Costs and the alternative-cost paragraph found on the card its entry points to with the same system, metric, unit and rater (edited entries BBC rounds, three conditions, deep-research configurations, health reading, any-page unit, link-validity gap, uncited claims, judge uncertainty and generated references checked number by number; the earlier 32-to-69 defect is gone), differences 31 to 32 points recomputed, no verdict or recommendation or named person, balance two_sided with both edge kinds, cost_side_by costside-trustwork-e7 differs from the creating identity in the provenance block (role author).

  43. 2026-09-2000:25#

    All six anchor parents read in full (hash 74c16a631655fb7b); every number found printed in a parent: 11.4 and 56.8 and ten LLMs (Naser), 14.23 and 94.93 and thirteen LLMs (GhostCite), 28.6 and 91.4 (Chelli), 39.8 of 400 from eight chatbots, 77.2 and 54.0 (Omar); task split four list / two drafting, three stated verifiers (two automated, one by hand) against three abstracts silent on who verified, and retrieval status (only Naser stated without retrieval, GhostCite with and without search, rest unstated, Perplexity in Ozbek and the eight chatbots) all match the parents; each parent does work; second breaking-point disjunct is inference-level.

  44. 2026-09-1923:52#

    All 6 SUPPORTS and 16 OPPOSES targets read; 6 Benefits and 16 Costs entries match the edges one to one, no verdict or recommendation, cost_side_by costside-trustwork-e7 distinct from creator author-trustwork-e7, all other numbers confirmed on their target cards; defect: the Costs entry 'Uncited claims' gives '32 to 69 percent of extracted factual claims carry no citation', which is not printed on its OPPOSES target (high-faithfulness-of-cited-claims-coexists...), whose Conclusion prints only groundedness 0.34, 0.31 and 0.68; the 32-to-69 figure stands only on the non-linked parent anchor groundedness-of-31-to-68-percent-on-researcherbench.

  45. 2026-09-1922:39#

    Six anchor parents read in full; all numbers found (11.4-56.8, 14.23-94.93, 28.6-91.4, 39.8 of 400, 77.2 and 54.0) and the existence-not-support clause is printed in every parent; defect: the Step clause 'The common task is generating a reference list on request, not citing a retrieved page' (and the Conclusion opening 'six studies ask models to produce bibliographic references on request') is not what two parents say: Omar had the models write medical research introductions with references, Ozbek had them write letters to the editor and original articles yielding 3150 references, so the common task holds for four of six parents only; 'not citing a retrieved page' is also asserted although the same Step says four parents do not state retrieval status (Ozbek and the eight-chatbot study include Perplexity).

  46. 2026-09-1922:39#

    Four derivation parents and their thirteen anchors read; recomputed 76.8-24.4=52.4, 94.04-39.36=54.68, 90.24-58.0=32.24 and 81.44-50.3=31.14 (31 to 32), 79.1-50.3=28.8 (29), 90.24-77.96=12.28 (12), 38.4 against 75.7 and the 2 to 22 band as printed in the parents; every parent does work, scope limits stated, breaking point is inference-level.

  47. 2026-09-1922:39#

    Four anchor parents read; recomputed 100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, the six ELI5 human-ALCE gaps (2.0, 2.0, 3.3, 0.0, 0.2, 1.1 giving 0.0 to 3.3), 42.4-40.4=2.0, F1 0.750 and 86.1 as printed; each parent does work and the breaking point targets the stated representativeness premise.

  48. 2026-09-1922:39#

    Three anchor parents read; recomputed 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, ablation bounds 100-78.6=21.4, 100-80.0=20.0, 92-16.7=75.3, 92-57.9=34.1, Reflex 100-76.8=23.2 and 100-47.7=52.3, 88.7 as printed; units named per pair, every parent does work, breaking point targets the judge-artefact premise.

  49. 2026-09-1921:48#

    All three parents read; recomputed 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3 (title 22 to 52 holds for the printed pairs), ablation bounds 100-78.6=21.4, 100-80.0=20.0, 92-16.7=75.3, 92-57.9=34.1, 88.7 found as printed; defect: Conclusion says that in the three measurements validity stays at 92 to 100 percent, but the 92 floor is printed only for the two-model depth ablation; the 14-agent parent prints Link Works for two agents (100.0 and 98.7) and says 12 of 14 exceed 94 percent, so two agents sit at or below 94 percent with no value and no floor printed; the parents support 98.7 to 100 percent in the three printed pairs and above 92 percent at every depth in the ablation, not a 92 to 100 band for all systems of the three measurements.

  50. 2026-09-1921:48#

    Both parents read; 89.0 percent (claim against all pages the Overview cites) and 75.7 percent (statement supported by at least one source in the same response) found as printed with their units, worked example recomputed (1 of 2 claims = 50 percent, 3 of 4 citations = 75 percent); the upper-bound step is set logic over claims and is restricted to the claim denominator.

  51. 2026-09-1921:48#

    Both parents read; recomputed 8/62=12.9 percent (13 as the report prints), 12 percent of 1,053 (appendix counts 59+28+33+8=128, 12.2 percent), 1,053/62=17.0, one in eight=12.5 percent, Gemini 20 percent of 290 found as printed; human raters in both, dates December 2024 and mid 2025 as printed, the quote-over-response denominator mismatch and the BBC's part in both rounds are declared in the Step.

  52. 2026-09-1921:47#

    All four parents read; 88.9, 84.2 and 67.5 human-rated precision, under 1 percent after the URL tool (0.6, 0.1, 0.8 as the prose prints), 95.6 against 16.5 percent found as printed, no arithmetic needed; closed-document setting, link resolution and existence are each as the parents state, the one causal verb rests on the before-and-after design the URL parent prints, and each parent states that web claim support was not measured.

  53. 2026-09-1921:47#

    All four parents read; recomputed 100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, 42.4-40.4=2.0, F1 0.750 and 86.1 found as printed; defect: Conclusion gives the ELI5 rate-level gaps between human and automatic system scores as 0.2 to 3.3 points, but the ALCE parent prints six pairs with gaps 2.0, 3.3, 0.2 on recall and 2.0, 0.0 (60.6 against 60.6, ChatGPT RERANK), 1.1 on precision, so the printed range is 0.0 to 3.3 points; 0.2 to 3.3 holds for citation recall only, which the clause does not say.

  54. 2026-09-1921:47#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and moves with benchmark and unit as far as it moves between systems (id no longer in the stock: renamed or superseded)

    Four derivation parents and their eleven anchors read; recomputed 24.4 to 94.04, 90.24-58.0=32.24 and 81.44-50.3=31.14, 79.1-50.3=28.8, 90.24-77.96=12.28, 75.7/38.4=1.97, 2 to 22.4 percent; defect 1: Step calls the judges behind the 24-94 span the unvalidated ones whose error is unmeasured and Conclusion calls the rater uncertainty of unknown size, but no parent prints that, and the anchors under the range parent print validations of those judges (DeepResearch Bench judge agreed with humans on 96 percent of support and 92 percent of not-support determinations on 100 pairs, DeepTRACE judge Pearson 0.62 on 100 tasks, the 14-agent judges calibrated through human review); only the ResearcherBench support judge is stated as not human-checked, so the parents support only that the 2-22 percent validations concern other judges; defect 2: title says the share moves with benchmark and unit as far as between systems, but the same-product parent licenses that only for deep-research products (31-32 points against 29 and 12), while the range parent prints between-system spreads of 52.4 points (24.4-76.8) and 54.7 points (39.36-94.04) within one benchmark, and the unit parent gives a response-level rate as a different quantity (37 points on one subset, about 15 on the main set).

  55. 2026-09-1921:47#

    Four derivation parents and their eleven anchors read; recomputed 24.4 to 94.04, 90.24-58.0=32.24 and 81.44-50.3=31.14, 79.1-50.3=28.8, 90.24-77.96=12.28, 75.7/38.4=1.97, 2 to 22.4 percent; defect 1: Step calls the judges behind the 24-94 span the unvalidated ones whose error is unmeasured and Conclusion calls the rater uncertainty of unknown size, but no parent prints that, and the anchors under the range parent print validations of those judges (DeepResearch Bench judge agreed with humans on 96 percent of support and 92 percent of not-support determinations on 100 pairs, DeepTRACE judge Pearson 0.62 on 100 tasks, the 14-agent judges calibrated through human review); only the ResearcherBench support judge is stated as not human-checked, so the parents support only that the 2-22 percent validations concern other judges; defect 2: title says the share moves with benchmark and unit as far as between systems, but the same-product parent licenses that only for deep-research products (31-32 points against 29 and 12), while the range parent prints between-system spreads of 52.4 points (24.4-76.8) and 54.7 points (39.36-94.04) within one benchmark, and the unit parent gives a response-level rate as a different quantity (37 points on one subset, about 15 on the main set).

  56. 2026-09-1921:45#

    Both parents read; 31 percent, Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15, 42 percent of Gemini responses without direct source (EBU 2025) and 26 percent without sources (BBC December 2024) found as printed, no arithmetic needed; response unit, journalist raters and the three-part category are as the parents state; breaking point is the decomposition outcome under which the narrowing fails while the 31 percent stands.

  57. 2026-09-1921:45#

    All four parents read; more than 60 percent of 1,600 queries, 37 to 94 percent by tool, 153 of 200, 154 of 200 and DRACO 42.1 to 64.6 found as printed; each parent's Statement prints the task definition the sort relies on (reverse attribution twice, link resolution, primary-document rubric by LLM judge from the top system's vendor), no number recomputed, no scope widening.

  58. 2026-09-1921:44#

    Both parents read; 1.07 percent of 56,381, 1.61 percent in 2025, 80.9 percent above the 2020-2024 average (1.61/0.89 recomputed as 1.809), excess 0.21 to 1.91 percent over four corpora as of August 2025 all found as printed; units kept apart, no causal attribution, breaking point is about independence of the two audits.

  59. 2026-09-1921:44#

    All six parents read; ranges 11.4-56.8, 14.23-94.93, 28.6-91.4, 39.8 of 400, 77.2 and 54.0 found as printed and title 11 to 95 is their rounded envelope; defect: Conclusion counts four studies as stating verification against databases or by hand and Step calls Chelli a hand-checked sample, but the Chelli parent is an abstract-only card that says the abstract does not state who checked the references (it prints only a gold-standard comparison and a 2-of-3-fields rule); the parents support three studies with a stated verifier (Naser and GhostCite automated database matching, the eight-chatbot study by hand) and three abstract-only cards (Chelli, Omar, Ozbek) that do not state who verified.

  60. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 1 (Abstract, the only occurrence of the 50% to 90% range); seven models, 800 questions, 58,000 pairs and approximately 30% unsupported statements on p. 1, valid URLs 40% to 70%, no URL hallucination for the two RAG systems, 55%, 34.5%, around 70%, about 10%, over 20% source-less responses and 95.8% on 110 pairs on p. 2, Fig. 1b metric captions p. 3, 400 MayoClinic plus 400 r/AskDocs, gpt-4o-2024-05-13, 3/28/24 and 1/20/24 and the 95.1% merged-source rerun on p. 6; publisher page gives 16 April 2025, Nature Communications 16:3615, no competing interests.

  61. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2507.16280v1 and https://arxiv.org/abs/2507.16280 on 2026-09-19; quoted found verbatim on p. 8 (Section 5.3, Finding 2); Groundedness 0.68/0.59/0.56/0.39/0.34/0.32/0.31 confirmed per system in Table 2 p. 8; definition cited claims over all extracted claims (Eq. 3) and Jina Reader plus binary judge on p. 6-7 (Section 4.2), GPT-4.1 as extractor and judge and March to April 2025 on p. 8, 65 questions p. 1, human meta-evaluation of 10 responses covers the rubric judge only (p. 9-10, Table 3); 32 and 69 in the title are the complements of 0.68 and 0.31 under that definition, not printed figures; abs page gives v1 of 22 Jul 2025, no venue.

  62. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2409.02897v3, https://arxiv.org/abs/2409.02897 and https://aclanthology.org/2025.findings-acl.264/ on 2026-09-19; quoted found verbatim on p. 5 (LongCite-8B row of Table 2); average F1 72.0/69.2/67.2/65.6/65.4, open models 19.7 to 51.5, LongBench-Chat precision 79.7/78.1/67.8/53.9/53.5 and GovReport precision 93.9/93.4/90.4/86.6/76.5 confirmed in Table 2 with column order R, P, F1; GPT-4o judge with 1/0.5/0 recall and relevant-or-not precision in Section 2.3.2 p. 4; human study 150 responses, 1,064 statements, 909 citations, kappa 0.593 and 0.655, accuracy 75.0% and 88.8% on p. 10; abs page gives v1 4 Sep 2024 and v3 10 Sep 2024; ACL Anthology lists the paper in Findings of ACL 2025.

  63. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2409.02897v3, https://arxiv.org/abs/2409.02897 and https://aclanthology.org/2025.findings-acl.264/ on 2026-09-19; quoted found verbatim on p. 10 (rows of Table 6); human R/P/F1 79.6/88.9/82.6, 72.8/84.2/75.8, 61.2/67.5/60.2 and GPT-4o 62.0/79.7/67.4, 57.6/78.1/63.6, 47.6/53.9/47.1 confirmed in Table 6; 150 anonymized responses, 1,064 statements, 909 citations, same standard as the GPT-4o evaluation and Table 7 agreement figures in Section 4.3 p. 10; 50 LongBench-Chat queries in Table 1 p. 3; precision as at-least-partial support in Section 2.3.2 p. 4; annotators not described; abs page gives v3 of 10 Sep 2024 (v1 4 Sep 2024); ACL Anthology lists Findings of ACL 2025.

  64. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2602.11685v1 and https://arxiv.org/abs/2602.11685 on 2026-09-19; quoted found verbatim on p. 12 (Table 13 rows and caption); Citation Quality 64.6/62.5/51.5/45.8/42.5/56.2/42.1 matched to the column order Perplexity Opus 4.6, Perplexity Opus 4.5, Gemini, OpenAI o3, OpenAI o4-mini, Opus 4.6, Opus 4.5; 12% (4.8) of 39.3 criteria on p. 6, axis description Citations to primary source documents in Table 4 p. 7, 100 tasks, Gemini-3-Pro judge, internal alignment study and alternative judges in Section 5.1 p. 9, MET/UNMET in Section 4.2 p. 8, 5 grading runs p. 10, The LLM Data Company and twenty-six experts at the start of Section 4.1 on p. 5; nine Perplexity authors and one Harvard on p. 1; abs page gives v1 of 12 Feb 2026, no venue.

  65. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19; quoted found verbatim on p. 6 (Section 4.1); 14 models, 130 queries, 12 of 14 above 94% Link Works (counted in Table 1, p. 7), frontier Relevant above 80%, Fact Check 24.4% OSS-120B to 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7 and Opus 4.5 98.7/95.7/76.8 confirmed in Table 1; LLM-judge calibrated by human review, AST parser, PwC affiliation and Preprint mark confirmed; abs page gives v1 of 7 May 2026 and no venue; author overlap with arXiv 2607.08700 confirmed on its abs page.

  66. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19; quoted found verbatim on p. 7 (Section 4.3); seven budgets 2 to 150 on p. 6, GPT-5.4 78.6% to 16.7% with 45.9/35.5/37.2 in Table 2 and Claude Opus 4.6 80.0% to 57.9% in Table 3 (both p. 8), 62 and 22 point declines and approximately 42% average printed, Link Works and Relevant minimum 92.3% so above 92% at every depth, calibration on 50-100 judgments on p. 6; no query count, per-level n or interval printed for the ablation; abs page gives v1 of 7 May 2026.

  67. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2608.24306v1 and https://arxiv.org/abs/2608.24306 on 2026-09-19; quoted found verbatim on p. 8 (Section 6.2); 84.7/14.8/0.4, 52.6/47.4 and 100.0 confirmed in Table 2 p. 8, global citation recall 58.7/28.5/7.1 on p. 7, 20 DeepResearch Bench examples with up to 10 sentences per agent on p. 6, 70%/30% (Fig. 5), 99% and 95% on p. 8, judge gpt-5-mini-2025-08-07 on p. 7, human study 50 sentences, kappa 0.71, 76% exact and kappa 0.62, 75% localisation on p. 5-6, funding and limitations on p. 10; abs page gives v1 of 25 Aug 2026 and the comment Accepted to EMNLP 2026 (Main Conference).

  68. 2026-09-1921:41#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf and the BBC Media Centre release page on 2026-09-19; quoted found verbatim on p. 8 (Sourcing); Q2 wording and Significant Issues counts 19/23/30/15 and the Gemini Q2 column 30+20+15+7=72 confirmed in the rating table p. 15; product versions and 12 Gemini refusals on p. 13, collection on 5 and 6 December 2024, 45 journalists, 362 responses, randomised anonymised order and Krippendorff Alpha on p. 14, ten examples on pp. 16-24; release page states Published: 11 February 2025, report dated February 2025 by Oli Elliott, BBC Responsible AI Team.

  69. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2305.14627v2 and https://arxiv.org/abs/2305.14627 on 2026-09-19; quoted found verbatim on p. 7 (Table 6 header and first rows); 51.1/50.0, 69.3/67.8, 44.0/50.1, 48.5/53.4 and 38.3/37.9 confirmed in Table 6; recall and precision definitions via TRUE NLI confirmed in Section 3.3 p. 4-5 (TRUE as T5-11B p. 4), around 50% sentence on p. 2, 1,000 dev questions, Sphere and 100-word passages on p. 3, Table 9 human 50.8/52.4 vs ALCE 52.8/50.4 on p. 9, partial-support limitation p. 10, single seeded run for GPT-4 in G.6 p. 17, acknowledgments p. 10; abs page gives v1 24 May 2023, v2 31 Oct 2023 and the comment Accepted by EMNLP 2023.

  70. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 4 (Additional validation on HealthSearchQA); 300 questions, URL validity 100%, 75.7% (74.0-77.2), 38.4% (26.7, 49.3), Reddit 31.0% (26.7, 35.8), close to 80% MayoClinic and clinician 40.4% vs pipeline 42.4% on p. 4, 55% on p. 2, 400+400 questions and gpt-4o-2024-05-13 on p. 6, any-source rule p. 7, metric definitions and not-blinded statement p. 8, code repository p. 9; publisher page gives Nature Communications 16:3615, published 16 April 2025, received 30 September 2024, no competing interests, peer review file.

  71. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 2 (Source verification; the extraction prints Veri fication with a ligature split exactly as quoted); 88.7%, 86.1%, p = 0.21 unpaired two-sided t-test, Claude Sonnet 3.5 87.0% (83.4-90.4), Llama 3.1 70B 79.3% (75.4-83.1), 90.1% (89.7-90.5), 110 pairs 95.8% (91.8-98.7) and 105 confirmed on p. 2, N = 400 in Fig. 1a caption p. 3, source models and the doctors-scored-the-decision wording under Expert validation p. 8, 40.4% vs 42.4% on p. 4, ambiguity limitation p. 6, annotating co-authors p. 10, annotations released p. 9; publisher page gives 16 April 2025, Nature Communications 16:3615.

  72. 2026-09-1921:41#
    failedcheckchecker-trustwork-e7 · Claudewhere one study measures both link validity sits 20 or more points above claim support (id no longer in the stock: renamed or superseded)

    All three parents read in full; printed pairs recomputed: 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, and 16.7/57.9 with Link Works above 92 found; but 'at least 22 points lower' in all three measurements does not hold: 21.9 is below 22, and in the depth ablation Fact Check is 78.6 and 80.0 percent at 2 tool calls, so with validity at most 100 the gap there is at most 21.4 and 20.0 points (the Conclusion cites only the 150-call endpoints), which also leaves the title's '20 or more' unshown for that condition; and '2 to 22 percent' in the Step is printed in no parent (the only judge validation in the parents is 88.7 percent agreement, i.e. 11.3); the parents support gaps of 21.9, 52.3 and 24.3 points in the printed pairs and a gap that widens with search depth from at most about 20 points to at least 34 and 75.

  73. 2026-09-1921:41#

    All three parents read in full; printed pairs recomputed: 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, and 16.7/57.9 with Link Works above 92 found; but 'at least 22 points lower' in all three measurements does not hold: 21.9 is below 22, and in the depth ablation Fact Check is 78.6 and 80.0 percent at 2 tool calls, so with validity at most 100 the gap there is at most 21.4 and 20.0 points (the Conclusion cites only the 150-call endpoints), which also leaves the title's '20 or more' unshown for that condition; and '2 to 22 percent' in the Step is printed in no parent (the only judge validation in the parents is 88.7 percent agreement, i.e. 11.3); the parents support gaps of 21.9, 52.3 and 24.3 points in the printed pairs and a gap that widens with search depth from at most about 20 points to at least 34 and 75.

  74. 2026-09-1921:40#

    All four parents read (Statement, Collection, Falls When); more than 60 percent of 1,600, 37 and 94 percent, 153 of 200, 154 of 200 and DRACO 42.1-64.6 found printed, and each parent's Statement prints the task definition the sorting rests on (reverse attribution twice, link resolution, primary-source rubric under an LLM judge by the top system's vendor); no number computed, every parent works, breaking point is not a restated Falls When.

  75. 2026-09-1921:40#

    Both parents read in full; 78.6 to 16.7 and 80.0 to 57.9 over budgets 2 to 150, non-monotone GPT-5.4 series, 84.7/52.6/100 percent orchestrator origin (range 52.6-100) and searcher 0.4 percent for AI-Q only found printed, no number computed; the decline is licensed by a controlled ablation on two models, the cross-class premise is declared as unmeasured in the Step, both parents work, breaking point is inference-level.

  76. 2026-09-1921:40#

    Both parents read in full; 1.07 percent of 56,381, 1.61 percent in 2025, 80.9 percent above the 2020-2024 average, and 0.21-1.91 percent excess in four corpora by August 2025 all found printed, no causal attribution kept, but the Breaking Point is not inference-level: 'the 2025 rise reverses' on a re-run over the same 2025 papers restates the conference anchor's Falls When (2025 share returns to the 0.76-0.98 band on re-extraction) and 'explained by indexing lag' restates the four-corpus anchor's Falls When (unmatched references exist, indexing lag; baseline re-estimated on later snapshots); nothing names what would break the joining step while both parents stand, e.g. the two audits not being independent samples or the per-paper share and the per-reference excess not describing the same trend.

  77. 2026-09-1921:40#

    All five parents read (Statement and Collection); 74.5, 63.6-89.5, ALCE precision 50.0/50.1/53.4 as about 50, DeepTRACE 39.8-68.3 as of 27 August 2025, BBC 10-15 and 47 percent, rounds December 2024 and May/June 2025, samples 362 and 237 and the tier caveat all found printed; the step computes no trend, every parent works, breaking point is inference-level.

  78. 2026-09-1921:40#

    All five parents read; 24.4-76.8, 39.8-68.3, 50.3-79.1, 39.36-94.04 with 77.96-90.24, and 0.62-0.86 found printed, overall minimum 24.4 (OSS-120B) and maximum 94.04 (Claude-3.5-Sonnet w/Search) recomputed across all printed values including GPT-5 web search 31.4, product range 50.3-90.24 recomputed over the deep-research rows of DeepTRACE, DeepResearch Bench and ResearcherBench (0.69-0.86 inside); all five are per-citation or per-cited-claim under an LLM judge, breaking point has an inference-level clause.

  79. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudetwo support rates in the stock count a claim as supported if any cited page supports it and are upper bounds on per citation support (id no longer in the stock: renamed or superseded)

    Both parents read in full; 89.0 and 75.7 percent found and both Statements give the any-source claim unit, but the set logic changes the denominator: 'attached citation supports implies some cited page does' bounds the share of claims supported by their own attached citation, not 'the share of citations that support the sentence they are attached to'; with citations as denominator the rate can exceed the any-source claim rate (supported claims carrying several supporting citations, unsupported or uncited claims carrying one or none), so 'can only be equal to or higher' and the title's 'upper bounds on per-citation support' do not follow; the parents support 'both are any-source claim-level rates, more lenient than checking the attached citation, and an upper bound on the share of claims supported by their attached citation'.

  80. 2026-09-1921:40#

    Both parents read in full; 89.0 and 75.7 percent found and both Statements give the any-source claim unit, but the set logic changes the denominator: 'attached citation supports implies some cited page does' bounds the share of claims supported by their own attached citation, not 'the share of citations that support the sentence they are attached to'; with citations as denominator the rate can exceed the any-source claim rate (supported claims carrying several supporting citations, unsupported or uncited claims carrying one or none), so 'can only be equal to or higher' and the title's 'upper bounds on per-citation support' do not follow; the parents support 'both are any-source claim-level rates, more lenient than checking the attached citation, and an upper bound on the share of claims supported by their attached citation'.

  81. 2026-09-1921:40#

    Three parents read in full; 0.84/0.34, 0.80/0.31, 0.62/0.68, 4.35, 111.21, 94.04 with 9.78 and 81.44 with 111.21 found printed in the two ResearcherBench anchors and the DeepResearch Bench volume anchor, ratio 111.21/4.35 = 25.6 recomputed, Sonar Reasoning Pro confirmed as lowest faithfulness and highest groundedness of the seven rows; every parent works, breaking point is inference-level.

  82. 2026-09-1921:40#

    Both parents read in full; 3.0-13.3 hallucinated and 5.4-18.5 non-resolving found printed in the DRBench URL anchor, 24.4-76.8 Fact Check printed in the 14-agent anchor, 100 minus 76.8 = 23.2 and 100 minus 24.4 = 75.6 recomputed; cross-study premise is declared in the Step, both parents work, breaking point is inference-level.

  83. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudefabricated reference rates of 11 to 95 percent measure whether a recalled reference exists and do not answer the support question (id no longer in the stock: renamed or superseded)

    All six parents read; 11.4-56.8 (ten LLMs), 14.23-94.93 (thirteen LLMs), 28.6-91.4 (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 found printed and the existence-not-support inference holds, but the title word recalled and the Reflex clause 'recalled from memory on request' attribute memory-only generation to rates the parents do not license: the GhostCite parent prints that its 14.23-94.93 rates come from runs with and without online search, and the eight-chatbot parent (source of the 40 percent figure) states no retrieval status, as the Step itself concedes; parents support 'references generated on request', memory-only for Naser alone.

  84. 2026-09-1921:40#

    All six parents read; 11.4-56.8 (ten LLMs), 14.23-94.93 (thirteen LLMs), 28.6-91.4 (GPT-3.5, GPT-4, Bard) and 39.8 percent of 400 found printed and the existence-not-support inference holds, but the title word recalled and the Reflex clause 'recalled from memory on request' attribute memory-only generation to rates the parents do not license: the GhostCite parent prints that its 14.23-94.93 rates come from runs with and without online search, and the eight-chatbot parent (source of the 40 percent figure) states no retrieval status, as the Step itself concedes; parents support 'references generated on request', memory-only for Naser alone.

  85. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudetwo journalist audits find about one in eight quote bearing news responses with an altered or unfindable quote (id no longer in the stock: renamed or superseded)

    Both parents read in full; 8 of 62 (13 percent), 12 percent of 1,053, 22 organisations found, 1,053/62=17.0 recomputed, and the quotes-per-responses mismatch is named in the Step; but '14 languages' in the Step is printed in neither parent (the EBU quote anchor gives 22 organizations and 18 countries only; the figure stands in the sibling EBU sourcing anchor, which is not a parent), so the number must go or that anchor must be linked.

  86. 2026-09-1921:40#

    Both parents read in full; 8 of 62 (13 percent), 12 percent of 1,053, 22 organisations found, 1,053/62=17.0 recomputed, and the quotes-per-responses mismatch is named in the Step; but '14 languages' in the Step is printed in neither parent (the EBU quote anchor gives 22 organizations and 18 countries only; the figure stands in the sibling EBU sourcing anchor, which is not a parent), so the number must go or that anchor must be linked.

  87. 2026-09-1921:40#

    Both SourceCheckup parents read in full; 75.7 (74.0-77.2), 38.4 (26.7-49.3), 55, 50-90 and seven models found printed, about 70 recomputed as 100 minus the printed approximately 30 percent unsupported statements, ratio 38.4/75.7 = 0.51 recomputed for the title's halving (holds on the 300-question sample; 55/70 = 0.79 on the 800 set); both parents work, breaking point is inference-level.

  88. 2026-09-1921:40#
    failedcheckchecker-trustwork-e7 · Claudethree interventions each lift one component of citation quality and none was measured on claim support in web research (id no longer in the stock: renamed or superseded)

    All four parents read in full; numbers found (88.9, 84.2, 67.5, under 1 percent, 95.6, 16.5), but the clause 'against 67.5 for an untrained model of the same family' is supported only for LongCite-9B (parent: Zhipu AI develops GLM-4 and the GLM-4-9B base of LongCite-9B); no parent states the base or family of LongCite-8B (88.9); and the causal 'training lifts' rests on a comparison across three different models, no parent prints one model before and after citation training, so the parents support 'citation-trained models score 88.9 and 84.2 against 67.5 for GLM-4' while the before/after causal reading holds only for the URL tool.

  89. 2026-09-1921:40#

    All four parents read in full; numbers found (88.9, 84.2, 67.5, under 1 percent, 95.6, 16.5), but the clause 'against 67.5 for an untrained model of the same family' is supported only for LongCite-9B (parent: Zhipu AI develops GLM-4 and the GLM-4-9B base of LongCite-9B); no parent states the base or family of LongCite-8B (88.9); and the causal 'training lifts' rests on a comparison across three different models, no parent prints one model before and after citation training, so the parents support 'citation-trained models score 88.9 and 84.2 against 67.5 for GLM-4' while the before/after causal reading holds only for the URL tool.

  90. 2026-09-1921:39#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and depends on benchmark unit and rater as much as on the system (id no longer in the stock: renamed or superseded)

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  91. 2026-09-1921:39#
    failedcheckchecker-trustwork-e7 · Claudethe share of supporting citations has no single published value and moves with benchmark and unit as far as it moves between systems (id no longer in the stock: renamed or superseded)

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  92. 2026-09-1921:39#

    All four derivation parents read with their anchors; numbers hold (24.4 to 94.04, 31 to 32 points recomputed, 75.7/38.4=1.97, 2 to 22 percent), but two clauses exceed the parents: the title's 'as much as on the system' is shown only for the benchmark (31 to 32 points against 29 and 12) and arguably the unit, while no parent shows the rater moving a rate comparably (printed rate-level judge-human gaps are 2 to 3 points); and 'the LLM judges that produce nearly all of these numbers disagree with humans on 2 to 22 percent' transfers a band measured on the AI Overviews verifier, the SourceCheckup judge and ALCE's NLI metric to the judges behind the 24 to 94 span, for which the anchors print other or no validations (Pearson 0.62, 96/92 percent, none); the parents support 'no single value; benchmark and unit move the rate as far as systems differ; validated judges elsewhere disagree with humans on 2 to 22 percent of decisions'.

  93. 2026-09-1921:38#
    failedcheckchecker-trustwork-e7 · Claudevalidated llm judges disagree with human raters on 2 to 22 percent of citation support decisions (id no longer in the stock: renamed or superseded)

    All four parents read in full; numbers recomputed and correct (100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, F1 0.750, samples 100 to 400), but the closing clause 'Every LLM-judged rate in the stock therefore carries an error of several points' widens scope: the parents validate three judges on their own material (plus one adversarial F1) and say nothing about the other judges in the stock, and they give decision-level disagreement, while the only rate-level gaps they print are 0.2 to 3.3 points (ALCE Table 9) and 2.0 points (40.4 vs 42.4, SourceCheckup); the parents support the 2 to 22 percent decision-level range for these validated judges only.

  94. 2026-09-1921:38#

    All four parents read in full; numbers recomputed and correct (100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, F1 0.750, samples 100 to 400), but the closing clause 'Every LLM-judged rate in the stock therefore carries an error of several points' widens scope: the parents validate three judges on their own material (plus one adversarial F1) and say nothing about the other judges in the stock, and they give decision-level disagreement, while the only rate-level gaps they print are 0.2 to 3.3 points (ALCE Table 9) and 2.0 points (40.4 vs 42.4, SourceCheckup); the parents support the 2 to 22 percent decision-level range for these validated judges only.

  95. 2026-09-1921:37#

    All three parents read in full; printed values found (58.0/90.24/0.85 Perplexity, 50.3/81.44/0.86 Gemini, judges, task counts, dates) and differences recomputed: 90.24-58.0=32.2, 81.44-50.3=31.1, 79.1-50.3=28.8, 90.24-77.96=12.3; version comparability is named as hidden premise and no ranking of benchmarks is claimed.

  96. 2026-09-1921:37#
    failedcheckchecker-trustwork-e7 · Claudethe 31 percent sourcing figure of the ebu audit counts responses and includes absent sources so it bounds per citation support without measuring it (id no longer in the stock: renamed or superseded)

    Both parents read in full; all numbers found (31, 72/24/15/15, 42, 26 percent) and the Conclusion and Step are supported, but the title clause 'bounds per-citation support' is not derivable: both parents state the unit is the response, not the citation, and neither gives any relation from a response-level mixed-category share to a per-citation rate, so the parents support only 'is not a per-citation support measurement' (at most a ceiling on the per-response unsupported-claim share).

  97. 2026-09-1921:37#

    All three parents read in full; numbers found as printed: FNR 0.183 to 0.470 and 81.6 percent gold-unsupported (judge benchmark), GPT-4o precision 79.7/78.1/53.9 against human 88.9/84.2/67.5 (LongCite Table 6), 2 of 100 verifier errors both too generous (AI Overviews, 100 minus 98 recomputed); each parent supplies one direction and the conclusion claims only that the direction is unsettled.

  98. 2026-09-1921:37#

    Both parents read in full; all numbers found (31, 72/24/15/15, 42, 26 percent) and the Conclusion and Step are supported, but the title clause 'bounds per-citation support' is not derivable: both parents state the unit is the response, not the citation, and neither gives any relation from a response-level mixed-category share to a per-citation rate, so the parents support only 'is not a per-citation support measurement' (at most a ceiling on the per-response unsupported-claim share).

  99. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2602.11685v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 12 in Table 13 (normalized scores by rubric axis), column order confirmed from the header: Perplexity (Opus 4.6), Perplexity (Opus 4.5), Gemini, OpenAI (o3), OpenAI (o4-mini), Opus 4.6, Opus 4.5, delta; Citation Quality 64.6, 62.5, 51.5, 45.8, 42.5, 56.2, 42.1 confirmed; axis description 'Citations to primary source documents' in Table 4 p. 7, 12% (4.8 of 39.3 criteria) on p. 6 and Table 5 p. 7; binary MET/UNMET in Section 4.2 p. 8; Gemini-3-Pro judge drawn from an internal human-LLM alignment study and GPT-5.2 / Sonnet-4.5 alternatives with stable ranking in Section 5.1 p. 9; 100 tasks and 5 independent grading runs p. 9-10; nine authors Perplexity, one Harvard, p. 1; abs page shows v1 submitted 12 Feb 2026, no venue.

  100. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2409.02897v3 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 5 in Table 2, column order confirmed from header and caption: Avg F1, CL, then R / P / F1 for Longbench-Chat, MultifieldQA, HotpotQA, Dureader, GovReport; all numbers in title and Statement confirmed (avg F1 72.0, 69.2, 67.2, 65.6, 65.4; open-source 19.7 to 51.5; LongBench-Chat P 79.7, 78.1, 67.8, 53.9, 53.5; GovReport P 93.9, 93.4, 90.4, 86.6, 76.5; no average precision column); GPT-4o judge and metric definitions confirmed in Section 2.3.2 p. 4; venue Findings of ACL 2025 confirmed at aclanthology.org/2025.findings-acl.264 (July 2025, pp. 5098-5122). Sole defect, date: as_of and the Evidence line give 2024-09-04 next to 'v3', but the arXiv abs page and the stamp on the fetched PDF date the pinned v3 to 10 Sep 2024; 2024-09-04 is the v1 first submission and the card does not say so. Correct reading: v3 of 2024-09-10, or 'first submitted 2024-09-04' stated as such.

  101. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2409.02897v3 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 10 in Table 6, column order confirmed from the header: Human scores R / P / F1, GPT-4o scores R / P / F1, ALCE scores R / P / F1; human 79.6/88.9/82.6, 72.8/84.2/75.8, 61.2/67.5/60.2 and GPT-4o 62.0/79.7/67.4, 57.6/78.1/63.6, 47.6/53.9/47.1 confirmed; 150 responses, 1,064 statements, 909 citations, anonymized, same standard as GPT-4o, in Section 4.3 p. 10; LongBench-Chat 50 queries p. 3; precision = cited snippet at least partially supports, Section 2.3.2 p. 4; no annotator identity or inter-annotator agreement reported; venue Findings of ACL 2025 confirmed at aclanthology.org/2025.findings-acl.264. Sole defect, date: as_of and the Evidence line give 2024-09-04 next to 'v3', but the arXiv abs page and the stamp on the fetched PDF date the pinned v3 to 10 Sep 2024; 2024-09-04 is the v1 first submission and the card does not say so. Correct reading: v3 of 2024-09-10, or 'first submitted 2024-09-04' stated as such.

  102. 2026-09-1921:34#
    failedcheckchecker-trustwork-e7 · Claudegroundedness of 31 to 68 percent on researcherbench means a third or more of factual claims carry no citation (id no longer in the stock: renamed or superseded)

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8, and all seven Groundedness values (0.68, 0.59, 0.56, 0.39, 0.34, 0.32, 0.31) are confirmed in the third column of Table 2 p. 8, definition Nc/N confirmed in Section 4.2 p. 7 (eq. 3), as_of 22 Jul 2025 confirmed on the abs page. Two defects: (1) the title says 'a third or more of factual claims carry no citation', but the highest Groundedness 0.68 (Perplexity: Sonar Reasoning Pro) leaves 32 percent uncited, which is below one third; the document supports '32 to 69 percent' as the Reflex already states, not 'a third or more'. (2) Evidence coordinate: the quoted sentence stands under 'Finding 2' in Section 5.3 Key Findings on p. 8, not in Section 5.2; the Evidence line names only 'Table 2 and Section 5.2' for p. 8 (5.2 is the running page header and the home of Table 2, not of the span).

  103. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2605.14021v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 6 in Table 1, columns Label / Count / Percentage, caption '98,020 claims from 7,491 verifiable AIOs'; all five label counts and shares plus Consistent 87,204 (89.0%) and Inconsistent 10,816 (11.0%) confirmed; label definitions p. 6, claim checked against full body text of every cited reference p. 6 and p. 11 Section 4.3; Grok 4.1 Fast Reasoning at temperature 0 p. 6; validation on 100 verdicts (20 per label), kappa 0.94, 98 of 100, both errors too generous, p. 7 Section 3.3.3; 55,393 queries, 19 categories, Mar 13 to Apr 21 2026 p. 7; abs page shows v1 submitted 13 May 2026, comment 'Under Review', no venue.

  104. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2605.14021v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 11, Section 4.3 'Variation across topics'; 76.85, 94.77, 93.65, 91.82, 91.42, 91.40 confirmed there and in Figure 5 p. 11 (consistent share = Clear + Vague, sorted descending, Climate 48.2%); Climate 48.23% named a measurement artifact and excluded from the range by the authors, real-time weather feeds and pages crawled hours or days later, and the adjusted 89.41 / 85.90 / 87.71 confirmed on p. 12; span of 18 points is 94.77 minus 76.85 = 17.92; abs page shows v1 submitted 13 May 2026, no venue.

  105. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8, and all seven Groundedness values (0.68, 0.59, 0.56, 0.39, 0.34, 0.32, 0.31) are confirmed in the third column of Table 2 p. 8, definition Nc/N confirmed in Section 4.2 p. 7 (eq. 3), as_of 22 Jul 2025 confirmed on the abs page. Two defects: (1) the title says 'a third or more of factual claims carry no citation', but the highest Groundedness 0.68 (Perplexity: Sonar Reasoning Pro) leaves 32 percent uncited, which is below one third; the document supports '32 to 69 percent' as the Reflex already states, not 'a third or more'. (2) Evidence coordinate: the quoted sentence stands under 'Finding 2' in Section 5.3 Key Findings on p. 8, not in Section 5.2; the Evidence line names only 'Table 2 and Section 5.2' for p. 8 (5.2 is the running page header and the home of Table 2, not of the span).

  106. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2506.11763v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 6 in Table 1, column order Overall/Comp./Depth/Inst./Read. (RACE) then C. Acc., E. Cit. (FACT) confirmed from the header; C. Acc. 90.24, 83.59, 81.44, 77.96 and search-tool values 39.36, 94.04, 93.68, 88.41 confirmed in Table 1; metric definition (per-task share of supported unique statement-URL pairs, 0 when none, averaged over tasks) confirmed in Appendix E p. 18-19; Gemini-2.5-Flash judge p. 6, 96%/92% on 100 pairs in Appendix C p. 18, collection April 1 to May 13 in Table 5 p. 18; abs page shows v1 submitted 13 Jun 2025, no venue.

  107. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2506.11763v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 7, Section 4.2.2; E. Cit. column (last column of Table 1, p. 6) confirmed: 111.21, 40.79, 32.88, 32.48, 31.26, 8.15, 4.79, 4.35 (minimum of the column), and the pairs 81.44/111.21 and 94.04/9.78; definition (supported pairs summed over tasks divided by number of tasks) confirmed in Appendix E.2 p. 19; judge Gemini-2.5-Flash p. 6; abs page shows v1 submitted 13 Jun 2025, no venue.

  108. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8 in Table 2 (Section 5.2), column order Coverage / Faithfulness / Groundedness confirmed from the header; Faithfulness 0.86, 0.85, 0.84, 0.80, 0.69 and 0.86, 0.62 for the two search-tool LLMs confirmed; definition Ns/Nc over cited claims only confirmed in Section 4.2 p. 7 (eq. 2); 65 questions p. 1, GPT-4.1 as extractor and judge and March to April 2025 on p. 8; Table 3 meta-evaluation (10 responses, p. 9) covers the rubric judge only; abs page shows v1 submitted 22 Jul 2025, no venue.

  109. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2509.04499v1 and https://arxiv.org/abs/2509.04499 on 2026-09-19: quoted found verbatim on p. 8 (Figure 2a score card); citation accuracy 68.3, 65.8, 49.0, 39.8 and unsupported statements 30.8, 23.1, 31.6, 47.0 for You, Bing, PPLX, GPT 4.5 and thoroughness 20.5 to 24.4 confirmed in Figure 2a, results as of August 27, 2025 in Section 4 on p. 8; definitions of Unsupported Statements (p. 6, any listed source) and Citation Accuracy (p. 7, overlap of matrices over number of citations), 303 queries with 168 debate and 135 expertise (p. 7), roughly 15% scraper errors, Pearson 0.62 on 100 tasks by two annotators and about 80,000 judgements (p. 5), GPT-5 default judge (p. 4) vs GPT-4 in Appendix E (p. 15), full/partial/none prompt (p. 19), three-engines caption vs four columns (p. 8) all confirmed; abs page shows v1 submitted 2 Sep 2025, no journal reference, authors and affiliations match.

  110. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2509.04499v1 and https://arxiv.org/abs/2509.04499 on 2026-09-19: quoted found verbatim on p. 9 (Table 1); citation accuracy 79.1, 72.3, 31.4, 58.0, 62.1, 50.3, unsupported statements 12.5, 74.6, 58.9, 97.5, 90.2, 53.6, thoroughness 87.5, 83.5, 17.9, 9.1, 13.2, 27.1, relevant statements 87.5 and 12.4 to 45.5, statements 23.9 to 141.6 and sources 3.6 to 57.2 confirmed per column in Table 1 on p. 9; running text on p. 9 prints 40.3% for Gemini against 50.3 in the table as the Statement says; acceptable threshold for citation accuracy [90,100) in Table 2 on p. 14; definitions on p. 6-7, Pearson 0.62 and 15% unscrapeable on p. 5, as of August 27, 2025 on p. 8; abs page shows v1 submitted 2 Sep 2025, no journal reference.

  111. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2604.03173v1 and https://arxiv.org/abs/2604.03173 on 2026-09-19: quoted found verbatim on p. 4 (Section 4.1 Overall rates); 5.4 [3.0, 7.7], 18.5 [17.8, 19.2], 3.0 [1.7, 4.4], 13.3 [12.7, 13.9] confirmed in Table 2 on p. 4, next highest 8.8 and 10.1 in Table 2, pooled 10.7 [10.2, 11.2] and 16.2 [15.7, 16.8] vs 4.8 [4.3, 5.2] and 6.8 [6.2, 7.3], ExpertQA 168,021 URLs, 8.22 [8.09, 8.36], Business 5.4 [4.9, 5.9], Theology 11.4 [8.1, 14.6], Reddit sensitivity 8.47 to 26.7 on p. 5; definitions (4xx/5xx, connection error, timeout, 403 excluded, no Wayback snapshot = hallucinated) on p. 2-3, 100 queries, 2,177 questions, 32 fields, 296 to 11,309 URLs in Table 1 on p. 3 (the ten counts sum to 23,269 against 53,090 in the abstract); logically prior question on p. 8, DARPA SciFy and MIT license on p. 9, three excluded models with 100% hallucination rates on p. 15; abs page shows v1 submitted 3 Apr 2026, PDF header Preprint. Under review, no journal reference.

  112. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2604.03173v1 and https://arxiv.org/abs/2604.03173 on 2026-09-19: quoted found verbatim on p. 7 (Section 5.1 Results); 16.0 to 0.6 (26x), 6.1 to 0.1 (79x), 4.9 to 0.8 (6.4x), all p < 10^-35 two-proportion z-test, 435 ExpertQA questions as 20% sample, gpt-5-nano 7.5% NOT LIVE with 48 hallucinated URLs across up to 14 rounds, 600 UNKNOWN URLs with 11.0% [8.5, 13.7] dead, Gemini two-phase run and the Claude stop at 658 questions confirmed on p. 7; Table 3 on p. 8 prints LIVE 79.3, 88.9, 78.0, DEAD 0.1, 0.2, 0.6, LIKELY HALL. 0.4, 0.5, 1.8, UNKNOWN 20.2, 10.3, 19.7 and 7,985, 4,203, 4,829 URLs, so DEAD plus LIKELY HALL. is 0.5, 0.7, 2.4 as the Collection says; 6-79x to under 1% in the abstract on p. 1; 8.47% main-pipeline rate on p. 5; abs page shows v1 submitted 3 Apr 2026, no journal reference.

  113. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19: as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (meta dc.date 2025-04-16; Europe PMC firstPublicationDate and Crossref published-online also 2025-04-16); everything else holds: quoted found verbatim on p. 4 (Additional validation on HealthSearchQA), 300 questions, URL validity 100%, 75.7% (74.0-77.2), 38.4% (26.7, 49.3), Reddit 31.0% (26.7, 35.8), clinician 40.4% (30.7, 50.1) vs pipeline 42.4% (32.7, 52.2) on p. 4, 55% on p. 2, close to 80% MayoClinic on p. 4, gpt-4o-2024-05-13 on p. 6, received 30 September 2024 on p. 1, metric definitions (status code 200 and non-empty text, at least one source, all statements supported) on p. 8, venue Nature Communications 16:3615 confirmed.

  114. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page on 2026-09-19: (1) as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (Europe PMC and Crossref also 2025-04-16); (2) the Statement sentence that model responses date from January to May 2024 is not printed: p. 6 Methods gives Gemini Ultra 1.0 (RAG) evaluated on 3/28/24 and all other model APIs queried on 1/20/24, while 2024-05-13 appears only as the snapshot name of the GPT-4o API endpoint, not as a response date, so no query date for GPT-4o is printed; the rest holds: quoted found verbatim on p. 1 Abstract, seven LLMs, 800 questions (400 MayoClinic, 400 r/AskDocs, p. 6), 58,000 pairs, approximately 30% of statements unsupported (p. 1), 55%, 34.5%, about 10%, around 70%, 40% to 70% valid URLs, 95.8% on 110 pairs, over 20% without sources (p. 2), 95.1% after merging (p. 6), 88.7% of 400 pairs (p. 2), Fig. 1b on p. 3, the 50-90 range appears only in the abstract, venue Nature Communications 16:3615 confirmed.

  115. 2026-09-1921:32#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher landing page on 2026-09-19: as_of and the Evidence date 2025-04-15 are contradicted, the publisher page prints Published 16 April 2025 (Europe PMC and Crossref also 2025-04-16); everything else holds: quoted found verbatim on p. 2 (section Source verification; the ligature split in Veri fication matches the PDF text layer), 88.7%, 86.1%, p = 0.21 unpaired two-sided t-test, Claude Sonnet 3.5 87.0% (83.4-90.4), Llama 3.1 70B 79.3% (75.4-83.1), 90.1% (89.7-90.5), 95.8% (91.8-98.7) and 105 of 110 on p. 2, N = 400 in Fig. 1a caption on p. 3, 400 pairs from GPT-4o (RAG), GPT-4o (API) and Claude v2.1 (API) and the wording that doctors scored whether the LLM-generated decision was correct in Expert validation on p. 8, 40.4% vs 42.4% on p. 4, not blinded on p. 8, annotating doctors as co-authors on p. 10.

  116. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2304.09848v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Section 4.2 (with the line-break hyphen of gener- ated as printed); 51.5 recall and 74.5 precision confirmed as the Average rows of Tables 7 and 8 (p. 23-24) and as unweighted means of the four system values; recall and precision definitions incl. the partial-support rule confirmed in Sections 2.3-2.4 (p. 3-4); 1450 queries, 34 annotators, MTurk, 250 triple-annotated pairs with >82.0% agreement and 91.0 F1, scrape window, acknowledgements and released annotations confirmed; arxiv.org/abs/2304.09848 confirms v1 2023-04-19, v2 2023-10-23 and Findings of EMNLP 2023; as_of is the v1 date and both figures are already printed in v1.

  117. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2304.09848v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Section 4.2; per-system precision 89.5/72.7/72.0/63.6 and recall 68.7/67.6/58.7/11.1 confirmed in text p. 7 and Tables 7-8 (p. 23-24); gaps of nearly 58% and almost 25% as printed p. 7; r = -0.96 printed in Section 4.3 (p. 8) with no unit of computation stated; perceived utility 4.34 (Bing Chat) and 4.62 (YouChat) confirmed on p. 6, Section 4.1 and the appendix table p. 22; copy/paraphrase explanation is framed by the paper as hypothesis (p. 2); arxiv.org/abs/2304.09848 confirms v1 2023-04-19, v2 2023-10-23 and Findings of EMNLP 2023; as_of is the v1 date and the figures are already printed in v1.

  118. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2305.14627v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 9, Section 6; kappa 0.698 and 0.525 and accuracy 85.1% and 77.6% confirmed p. 9 and Appendix G.5 p. 17; irrelevant-citation recall 75.6% and precision 66.1% with the partial-support explanation confirmed p. 17; Table 9 (p. 9) ELI5 human vs ALCE 50.8/52.4 vs 52.8/50.4, 59.7/60.6 vs 63.0/60.6, 13.4/19.2 vs 13.6/18.1 confirmed; human protocol (per sentence full support, per citation full/partial/no) confirmed in Section 6 p. 8 and Appendix F p. 15-16; Surge AI, 20 USD per hour, 100 sampled examples, no rater count or inter-rater figure confirmed; arxiv.org/abs/2305.14627 confirms v1 2023-05-24, v2 2023-10-31 and Accepted by EMNLP 2023; as_of is the v1 date and all figures are already printed in v1.

  119. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2305.14627v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Table 6, and all ten Table 6 figures (51.1/50.0, 44.0/50.1, 48.5/53.4, 38.3/37.9, 69.3/67.8), the metric definitions (Section 3.3, p. 4-5), Table 9 figures and the around-50% sentence (p. 2) are confirmed. Two defects: (1) as_of 2023-05-24 is the v1 date per arxiv.org/abs/2305.14627, but v1 Table 6 (fetched arxiv.org/pdf/2305.14627v1, p. 7) has no GPT-4 and no LLaMA-2-Chat rows; the GPT-4 and Chat-70B figures in title and Statement first appear in v2, dated 2023-10-31, which is the version the anchor pins. (2) Collection says scores are averaged over three seeded runs with one run for RERANK only; Appendix G.6 p. 17 says GPT-4 (and ChatGPT-16K) also use only one seeded run, so the GPT-4 figures 44.0/50.1 and 48.5/53.4 are single-run values.

  120. 2026-09-1921:32#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 8 (Sourcing) and all title and Statement numbers confirmed (over 45%, 26%, 7% on p. 8; Q2 wording and counts 19, 23, 30, 15 and the Gemini Q2 column 30+20+15+7=72 on p. 15; 12 Gemini refusals p. 13; 45 journalists, 362 responses p. 14; as_of 2025-02-11 confirmed on the BBC Media Centre release page), but Collection states a wrong number: it says the appendix prints 'three example responses', while the appendix section 'AI error examples' (pp. 16-24) prints ten example responses, each with reviewer comments (Copilot 4, Perplexity 3, Gemini 2, ChatGPT 1).

  121. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2603.03299v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 7, Section 4.1; 11.4% [10.4, 12.5] GPT-5-mini and 56.8% [55.6, 58.0] haiku-4.5 confirmed in Table 1 p. 7 (Real column sums to 40,529, Total to 69,557, difference 29,028); 69,557 and 40,529, thresholds 80 and 65 in Section 3.4 p. 6; 225-sample validation 75/75 and 8/75 (10.7%) in Section 3.5 p. 6; 15,150 responses p. 5; inclusive-threshold range 9.3% to 23.8% in Section 4.6 table p. 10-11; parametric-memory and non-retrieval baseline wording on p. 19; arXiv abs page confirms v1 submitted 2026-02-07, no journal reference.

  122. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2603.03299v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 9, Section 4.5; 16.5% (one model), 87.4% (two), 95.6% (three or more, 5.8-fold), 28.6% and 88.9% for within-model replications all confirmed on p. 9 as match rates of unique title strings against the verification pipeline; Jaccard 0.540 for GPT-5-mini/GPT-5-nano on p. 10; no per-class title counts printed; arXiv abs page confirms v1 submitted 2026-02-07, no journal reference.

  123. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2602.06718v2 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 6, Section V.A; 14.23 ±1.65 DeepSeek, 21.84 Claude 4, 50.92 GPT-5, 59.47 Gemini, 94.93 ±1.29 Hunyuan confirmed in Table II p. 7; 13 models, OpenRouter, 40 domains, batch 10/20/30, search plus chain-of-thought, 375,440 citations from 22,800 interactions and the controlled-baseline wording in Section IV.A p. 5-6; 331,809 extracted, 166,876 (50.29%) invalid, 400/400 samples with 100% and 98% (392/400) on p. 6; verification cascade with web-search and LLM-reparse fallbacks on p. 4; arXiv abs page confirms v1 2026-02-06, v2 2026-05-14, primary class cs.CR.

  124. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2602.06718v2 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 9, Section VI.A (including the missing space before 'contained'); 56,381 papers and eight venues in Table IV p. 9 and Section IV.B p. 6; 739 invalid (136 error, 603 ghost) and 604 papers (1.07%) on p. 9; 0.76%-0.98% for 2020-2024, 1.61% in 2025, 80.9% over the 0.89% average and the no-causality sentence in Section VI.D p. 10; sixteen assistants, checked at least twice, two researchers, 400-sample with no further invalid in Section IV.B p. 6; 2,199,409 extracted and 2,530 flagged are printed on p. 8 (start of Section VI), one page before the cited p. 9-10; arXiv abs page confirms v1 2026-02-06, v2 2026-05-14, primary class cs.CR.

  125. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2505.18059v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 8 under Results; 400 references, 26.5%, 33.8%, 39.8%, Copilot 100%, Perplexity 72%, Claude 64%, Grok and DeepSeek 0 of 50 confirmed on p. 8 next to Figure 1; models in Table 1 and test dates 7-9 February 2025 on p. 7; five elements and manual Google/Google Scholar verification on p. 8; not-peer-reviewed notice on p. 1; arXiv abs page confirms v1 submitted 2025-05-23.

  126. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:39667055%20AND%20SRC:MED&resultType=core&format=json on 2026-09-19 (own download, tags stripped, entities unescaped, NBSP normalised as whitespace); quoted found verbatim in abstractText Results; 77.2%/68.0% for Gemini vs 54.0%/49.2% for GPT-4, p < 0.001 for both, the 23.2 point difference, Gemini Ultra, five medical fields and 'both models produced fabricated evidence' confirmed in the abstract, which indeed gives no reference count and no metric definitions; record confirms firstPublicationDate 2024-12-12, Comput Biol Med vol 185 (2025 Feb) 109545, doi 10.1016/j.compbiomed.2024.109545, isOpenAccess N, and the affiliations.

  127. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1007/s43465-026-01807-0&resultType=core&format=json on 2026-09-19 (own download, tags stripped, thin spaces normalised as whitespace); quoted found verbatim in abstractText Results; 30 subtopics, two formats, 3150 references, four RHS criteria, pooled means 1.81 ± 3.40, 4.01 ± 4.89, 6.51 ± 4.89 (p < 0.001) confirmed, and the per-format means cited in Notes (1.81/4.02 ChatGPT, 6.43/6.31 Perplexity) are printed as stated; record confirms firstPublicationDate 2026-05-11, Indian J Orthop (vol 60 issue 8, 2026), doi 10.1007/s43465-026-01807-0, and both affiliations.

  128. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 10 (High-level findings); 31%, Gemini 72%, ChatGPT 24%, Perplexity and Copilot 15% and n=675/678/681/675 confirmed on p. 10, Q2 counts 160/104 (p. 67) and 483/101 (p. 68) and the Q2 note on lack of direct sourcing confirmed in Appendix 3, 271 journalists, 2709 responses and 24 May to 10 June 2025 on p. 63, 30 core questions p. 7, 42% Gemini no direct sources p. 34, Q2 wording p. 66, four-level scale p. 8, QA pass p. 65; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  129. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 21 (Accuracy of direct quotes); 1,053, 12%, Gemini 20% of 290, Copilot 4% and n=190/262/311/290 confirmed on p. 21, Q3 significant counts 28 and 8 (p. 67), 59 and 33 (p. 68) confirmed and the four Q3 columns sum to 262, 190, 290, 311 = 1,053, Q3 wording p. 66, absent and altered quote cases pp. 21-22; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  130. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 19 (In focus: Have assistants improved?); 362 and 237 core plus custom responses, 51% to 37%, 47%, 10-15% range and Copilot 27% to 10% confirmed on p. 19, 25 to a single response without direct URL source and the product tier table (Enterprise, Pro, Standard, Pro vs consumer/free) confirmed on p. 20, methodology caveat p. 19 and p. 65, first-round Krippendorff Alpha test confirmed in the BBC February 2025 report p. 14; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  131. 2026-09-1921:32#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 7 (Accuracy section); eight quotes, 62 responses, 13% and 'all assistants tested except ChatGPT' confirmed on p. 7, 100 questions and product versions (Enterprise GPT-4o, Pro, Standard, Pro) p. 13, 45 journalists, 362 responses, 5 and 6 December 2024, randomised order and Krippendorff Alpha p. 14; no total quote count is printed, as the Statement says; as_of 2025-02-11 confirmed on the BBC Media Centre release page (Published: 11 February 2025), report itself prints February 2025.

  132. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Chatbots' responses to our queries were often confidently wrong'; eight tools, 20 publishers, ten articles, sixteen hundred queries, first three Google results and the six manual labels confirmed in 'Methodology', more than 60 percent, 37 percent Perplexity, 94 percent Grok 3 and ChatGPT 134 of two hundred confirmed in the quoted section, DeepSeek 115 of 200 in 'Platforms often failed to link back', tests in February 2025, single-run and no-extrapolation caveats in 'Limitations'; byline Jazwinska and Chandrasekar, CJR Tow Center, dated March 6, 2025 on the page.

  133. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Platforms often failed to link back to the original source'; 200 prompts and 154 citations to error pages for Grok 3, more than half of Gemini and Grok 3 responses, Grok 2 homepage links and 'far less frequently with other chatbots' confirmed in the same section, tests conducted in February 2025 confirmed in the licensing-deals section, no counts printed for Gemini or other tools; byline and date March 6, 2025 confirmed on the page; the released data file is GPG-encrypted so the citation-versus-prompt unit could not be tested.

  134. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Confidently wrong'; two hundred quotes from twenty publications, forty from crawler-blocking publishers, a hundred and fifty-three incorrect, seven acknowledgements and 'more than a third' incorrect citations confirmed in that section, top-three Google or Bing selection in the introduction, publisher name, URL and date criterion in 'The illusion of control', run-to-run variation in 'Unpredictable (mis)attribution', OpenAI reply and GitHub data link in the conclusion; byline and date November 27, 2024 confirmed on the page.

  135. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.07723v1 and https://arxiv.org/abs/2605.07723 on 2026-09-19: quoted found verbatim on p. 5 (Results); 111 million references and 2.5 million papers (p. 1), corpora with arXiv Jan 2020-Aug 2025 and 10% PMC sample (p. 4), excess 0.39/0.21/1.91/0.27% as of August 2025 (p. 4), monthly 3,353/478/767/8,140 and 146,932 (p. 5), matching pipeline 95.1% matched, 2.33%, 1.54%, GPT-4o-mini, Google Scholar lookup (p. 3), hallucinated defined as estimated excess over pre-LLM baseline, not per-reference classification (p. 4), lower bound (p. 5); affiliations p. 1; v1 submitted 8 May 2026; no venue on the abs page.

  136. 2026-09-1921:32#
    failedcheckchecker-trustwork-e7 · Claudefourteen deep research agents keep links valid above 94 percent while fact check scores range from 24 to 77 percent (id no longer in the stock: renamed or superseded)

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 6 (Section 4.1), Table 1 on p. 7 confirms 24.4% OSS-120B, 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7, Claude Opus 4.5 98.7/95.7/76.8, 14 models, 130 queries, LLM judge calibrated by human review, v1 dated 7 May 2026; DEFECT in title: it says fourteen agents keep links valid above 94 percent, but the document (p. 6) says 12 of 14 exceed 94% on Link Works and Table 1 (p. 7) prints OSS-120B 83.9% and Llama 4 Maverick 80.8%; Statement itself has the correct 12 of 14.

  137. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 7 (Section 4.3); Tables 2-3 on p. 8 confirm GPT-5.4 Fact Check 78.6% at 2 calls, 45.9% at 10, 35.5% at 70, 37.2% at 100, 16.7% at 150 and Claude Opus 4.6 80.0% to 57.9%; seven budgets 2/10/30/50/70/100/150 (p. 6), Link Works and Relevant above 92% (p. 7, minimum 92.3% in Table 3), approximately 42% on average (abstract and p. 7), judge calibration on 50-100 judgments (p. 6), no n per depth level printed; v1 submitted 7 May 2026.

  138. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2608.24306v1 and https://arxiv.org/abs/2608.24306 on 2026-09-19: quoted found verbatim on p. 8 (Section 6.2), Table 2 on p. 8 confirms 84.7/14.8/0.4, 52.6/47.4, 100.0; citation recall 58.7/28.5/7.1 (p. 7), 20 examples and up to 10 sentences (p. 6), 70%/30% and 99% and 95% (p. 8, Figure 5), judge gpt-5-mini-2025-08-07 (p. 5, 7), 50 sentences, 76% exact, kappa 0.62, 75% localisation, inter-annotator kappa 0.71 (p. 5-6), funding (p. 10), v1 submitted 25 Aug 2026 all confirmed; DEFECT in Collection: it says arXiv preprint, not peer reviewed, but the arXiv abs page comment reads Accepted to EMNLP 2026 (Main Conference), so the venue and review status are contradicted by the page.

  139. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2607.08700v1 and https://arxiv.org/abs/2607.08700 on 2026-09-19: quoted found verbatim on p. 7 (Section 4.2); Table 2 on p. 6 confirms factual-support pass-class F1 0.649 [.56,.72] GPT-OSS-120B to 0.750 [.68,.82] Claude Opus 4.6 (kappa 0.701), GPT-5-mini 0.710 [.64,.78] (kappa 0.649), relevance 0.700 Claude Sonnet 4.6 to 0.908 [.89,.93] GPT-5-mini (kappa 0.636), kappa range 0.580-0.701; 624 pairs, 1,248 decisions, 8 judges, 3 families (p. 6); adjudicated subset 0.780 GPT-5.4-mini and 0.672 Opus 4.6 (p. 8, Section 4.3); council of 6, 870/378 (263/115), one human reviewer, 18.4% gold pass rate, 25 topics, about 60% edited, 19 strategies (p. 4-5, 7); v1 submitted 9 Jul 2026; no venue on the abs page.

  140. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2607.08700v1 and https://arxiv.org/abs/2607.08700 on 2026-09-19: quoted found verbatim on p. 9 (Section 5.1); FNR defined p. 6 as FN/(FN+TP), a good citation rejected; 0.183 GPT-5.4-mini to 0.470 GPT-OSS-120B, three judges above the 18.4% gold pass rate, relevance pass rates 42.9% to 72.0% below gold 79.3% (p. 9); roughly 86% to near 100% rejection of edited claims and over-rejection as dominant error (p. 10, Section 5.2); FPR per judge only in Figure 5 (p. 10), no printed values in text or appendix; 624 pairs, human-reviewed gold (p. 4-6); v1 submitted 9 Jul 2026; no venue on the abs page.

  141. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 6 (Section 4.1), Table 1 on p. 7 confirms 24.4% OSS-120B, 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7, Claude Opus 4.5 98.7/95.7/76.8, 14 models, 130 queries, LLM judge calibrated by human review, v1 dated 7 May 2026; DEFECT in title: it says fourteen agents keep links valid above 94 percent, but the document (p. 6) says 12 of 14 exceed 94% on Link Works and Table 1 (p. 7) prints OSS-120B 83.9% and Llama 4 Maverick 80.8%; Statement itself has the correct 12 of 14.

  142. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:38776130%20AND%20SRC:MED&resultType=core&format=json on 2026-09-19 (own download, tags stripped, entities unescaped); quoted found verbatim in abstractText Results; 39.6% (55/139), 28.6% (34/119), 91.4% (95/104), precision 9.4% (13/139), 13.4% (16/119), 0% (0/104), 11 reviews, 33 prompts, 471 references and the any-2-of-title/first-author/year definition all confirmed in the abstract; record confirms firstPublicationDate 2024-05-22, J Med Internet Res vol 26 e53164, doi 10.2196/53164, and the three affiliations.