Ledger › verdict

confirmed

101 rows where verdict is “confirmed”. One facet at a time; to cite a single event, link the row.

  1. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  2. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  3. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  4. 2026-09-2300:19#

    Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves

  5. 2026-09-2300:18#

    proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05

  6. 2026-09-2300:18#

    proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05

  7. 2026-09-2300:18#

    proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05

  8. 2026-09-2300:18#

    proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05

  9. 2026-09-2300:08#

    Re-check after the OPPOSES edge moved from the superseded gap derivation to its replacement; both sides re-read, 6 SUPPORTS and 16 OPPOSES intact, cost_side_by unchanged

  10. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  11. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  12. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  13. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  14. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  15. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  16. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  17. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  18. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  19. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  20. 2026-09-2300:06#

    Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named

  21. 2026-09-2222:57#

    Re-check after the narrowing note: quoted span unchanged and present in the abstract record; the note's per-format values re-read from the same record

  22. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  23. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  24. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  25. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  26. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  27. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  28. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  29. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  30. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  31. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  32. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  33. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  34. 2026-09-2219:36#

    Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text

  35. 2026-09-2219:19#

    proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a

  36. 2026-09-2219:19#

    proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a

  37. 2026-09-2219:19#

    proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a

  38. 2026-09-2001:34#

    At hash e9ddc0ce1915acf4 all 6 SUPPORTS and 16 OPPOSES target cards read in full from an unedited engine dump; 6 Benefits and 16 Costs entries match the edges one to one, the cost of the alternative is a separate entry written as unmeasured with no edge, every number in Benefits, Costs and the alternative-cost paragraph found on the card its entry points to with the same system, metric, unit and rater (edited entries BBC rounds, three conditions, deep-research configurations, health reading, any-page unit, link-validity gap, uncited claims, judge uncertainty and generated references checked number by number; the earlier 32-to-69 defect is gone), differences 31 to 32 points recomputed, no verdict or recommendation or named person, balance two_sided with both edge kinds, cost_side_by costside-trustwork-e7 differs from the creating identity in the provenance block (role author).

  39. 2026-09-2000:25#

    All six anchor parents read in full (hash 74c16a631655fb7b); every number found printed in a parent: 11.4 and 56.8 and ten LLMs (Naser), 14.23 and 94.93 and thirteen LLMs (GhostCite), 28.6 and 91.4 (Chelli), 39.8 of 400 from eight chatbots, 77.2 and 54.0 (Omar); task split four list / two drafting, three stated verifiers (two automated, one by hand) against three abstracts silent on who verified, and retrieval status (only Naser stated without retrieval, GhostCite with and without search, rest unstated, Perplexity in Ozbek and the eight chatbots) all match the parents; each parent does work; second breaking-point disjunct is inference-level.

  40. 2026-09-1922:39#

    Four derivation parents and their thirteen anchors read; recomputed 76.8-24.4=52.4, 94.04-39.36=54.68, 90.24-58.0=32.24 and 81.44-50.3=31.14 (31 to 32), 79.1-50.3=28.8 (29), 90.24-77.96=12.28 (12), 38.4 against 75.7 and the 2 to 22 band as printed in the parents; every parent does work, scope limits stated, breaking point is inference-level.

  41. 2026-09-1922:39#

    Four anchor parents read; recomputed 100-98=2, 100-88.7=11.3, 100-85.1=14.9, 100-77.6=22.4, the six ELI5 human-ALCE gaps (2.0, 2.0, 3.3, 0.0, 0.2, 1.1 giving 0.0 to 3.3), 42.4-40.4=2.0, F1 0.750 and 86.1 as printed; each parent does work and the breaking point targets the stated representativeness premise.

  42. 2026-09-1922:39#

    Three anchor parents read; recomputed 98.7-76.8=21.9, 100.0-47.7=52.3, 100-75.7=24.3, ablation bounds 100-78.6=21.4, 100-80.0=20.0, 92-16.7=75.3, 92-57.9=34.1, Reflex 100-76.8=23.2 and 100-47.7=52.3, 88.7 as printed; units named per pair, every parent does work, breaking point targets the judge-artefact premise.

  43. 2026-09-1921:48#

    Both parents read; 89.0 percent (claim against all pages the Overview cites) and 75.7 percent (statement supported by at least one source in the same response) found as printed with their units, worked example recomputed (1 of 2 claims = 50 percent, 3 of 4 citations = 75 percent); the upper-bound step is set logic over claims and is restricted to the claim denominator.

  44. 2026-09-1921:48#

    Both parents read; recomputed 8/62=12.9 percent (13 as the report prints), 12 percent of 1,053 (appendix counts 59+28+33+8=128, 12.2 percent), 1,053/62=17.0, one in eight=12.5 percent, Gemini 20 percent of 290 found as printed; human raters in both, dates December 2024 and mid 2025 as printed, the quote-over-response denominator mismatch and the BBC's part in both rounds are declared in the Step.

  45. 2026-09-1921:47#

    All four parents read; 88.9, 84.2 and 67.5 human-rated precision, under 1 percent after the URL tool (0.6, 0.1, 0.8 as the prose prints), 95.6 against 16.5 percent found as printed, no arithmetic needed; closed-document setting, link resolution and existence are each as the parents state, the one causal verb rests on the before-and-after design the URL parent prints, and each parent states that web claim support was not measured.

  46. 2026-09-1921:45#

    Both parents read; 31 percent, Gemini 72, ChatGPT 24, Perplexity 15, Copilot 15, 42 percent of Gemini responses without direct source (EBU 2025) and 26 percent without sources (BBC December 2024) found as printed, no arithmetic needed; response unit, journalist raters and the three-part category are as the parents state; breaking point is the decomposition outcome under which the narrowing fails while the 31 percent stands.

  47. 2026-09-1921:45#

    All four parents read; more than 60 percent of 1,600 queries, 37 to 94 percent by tool, 153 of 200, 154 of 200 and DRACO 42.1 to 64.6 found as printed; each parent's Statement prints the task definition the sort relies on (reverse attribution twice, link resolution, primary-document rubric by LLM judge from the top system's vendor), no number recomputed, no scope widening.

  48. 2026-09-1921:44#

    Both parents read; 1.07 percent of 56,381, 1.61 percent in 2025, 80.9 percent above the 2020-2024 average (1.61/0.89 recomputed as 1.809), excess 0.21 to 1.91 percent over four corpora as of August 2025 all found as printed; units kept apart, no causal attribution, breaking point is about independence of the two audits.

  49. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 1 (Abstract, the only occurrence of the 50% to 90% range); seven models, 800 questions, 58,000 pairs and approximately 30% unsupported statements on p. 1, valid URLs 40% to 70%, no URL hallucination for the two RAG systems, 55%, 34.5%, around 70%, about 10%, over 20% source-less responses and 95.8% on 110 pairs on p. 2, Fig. 1b metric captions p. 3, 400 MayoClinic plus 400 r/AskDocs, gpt-4o-2024-05-13, 3/28/24 and 1/20/24 and the 95.1% merged-source rerun on p. 6; publisher page gives 16 April 2025, Nature Communications 16:3615, no competing interests.

  50. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2507.16280v1 and https://arxiv.org/abs/2507.16280 on 2026-09-19; quoted found verbatim on p. 8 (Section 5.3, Finding 2); Groundedness 0.68/0.59/0.56/0.39/0.34/0.32/0.31 confirmed per system in Table 2 p. 8; definition cited claims over all extracted claims (Eq. 3) and Jina Reader plus binary judge on p. 6-7 (Section 4.2), GPT-4.1 as extractor and judge and March to April 2025 on p. 8, 65 questions p. 1, human meta-evaluation of 10 responses covers the rubric judge only (p. 9-10, Table 3); 32 and 69 in the title are the complements of 0.68 and 0.31 under that definition, not printed figures; abs page gives v1 of 22 Jul 2025, no venue.

  51. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2409.02897v3, https://arxiv.org/abs/2409.02897 and https://aclanthology.org/2025.findings-acl.264/ on 2026-09-19; quoted found verbatim on p. 5 (LongCite-8B row of Table 2); average F1 72.0/69.2/67.2/65.6/65.4, open models 19.7 to 51.5, LongBench-Chat precision 79.7/78.1/67.8/53.9/53.5 and GovReport precision 93.9/93.4/90.4/86.6/76.5 confirmed in Table 2 with column order R, P, F1; GPT-4o judge with 1/0.5/0 recall and relevant-or-not precision in Section 2.3.2 p. 4; human study 150 responses, 1,064 statements, 909 citations, kappa 0.593 and 0.655, accuracy 75.0% and 88.8% on p. 10; abs page gives v1 4 Sep 2024 and v3 10 Sep 2024; ACL Anthology lists the paper in Findings of ACL 2025.

  52. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2409.02897v3, https://arxiv.org/abs/2409.02897 and https://aclanthology.org/2025.findings-acl.264/ on 2026-09-19; quoted found verbatim on p. 10 (rows of Table 6); human R/P/F1 79.6/88.9/82.6, 72.8/84.2/75.8, 61.2/67.5/60.2 and GPT-4o 62.0/79.7/67.4, 57.6/78.1/63.6, 47.6/53.9/47.1 confirmed in Table 6; 150 anonymized responses, 1,064 statements, 909 citations, same standard as the GPT-4o evaluation and Table 7 agreement figures in Section 4.3 p. 10; 50 LongBench-Chat queries in Table 1 p. 3; precision as at-least-partial support in Section 2.3.2 p. 4; annotators not described; abs page gives v3 of 10 Sep 2024 (v1 4 Sep 2024); ACL Anthology lists Findings of ACL 2025.

  53. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2602.11685v1 and https://arxiv.org/abs/2602.11685 on 2026-09-19; quoted found verbatim on p. 12 (Table 13 rows and caption); Citation Quality 64.6/62.5/51.5/45.8/42.5/56.2/42.1 matched to the column order Perplexity Opus 4.6, Perplexity Opus 4.5, Gemini, OpenAI o3, OpenAI o4-mini, Opus 4.6, Opus 4.5; 12% (4.8) of 39.3 criteria on p. 6, axis description Citations to primary source documents in Table 4 p. 7, 100 tasks, Gemini-3-Pro judge, internal alignment study and alternative judges in Section 5.1 p. 9, MET/UNMET in Section 4.2 p. 8, 5 grading runs p. 10, The LLM Data Company and twenty-six experts at the start of Section 4.1 on p. 5; nine Perplexity authors and one Harvard on p. 1; abs page gives v1 of 12 Feb 2026, no venue.

  54. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19; quoted found verbatim on p. 6 (Section 4.1); 14 models, 130 queries, 12 of 14 above 94% Link Works (counted in Table 1, p. 7), frontier Relevant above 80%, Fact Check 24.4% OSS-120B to 76.8% Claude Opus 4.5, GPT-5.4 100.0/93.7/47.7 and Opus 4.5 98.7/95.7/76.8 confirmed in Table 1; LLM-judge calibrated by human review, AST parser, PwC affiliation and Preprint mark confirmed; abs page gives v1 of 7 May 2026 and no venue; author overlap with arXiv 2607.08700 confirmed on its abs page.

  55. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19; quoted found verbatim on p. 7 (Section 4.3); seven budgets 2 to 150 on p. 6, GPT-5.4 78.6% to 16.7% with 45.9/35.5/37.2 in Table 2 and Claude Opus 4.6 80.0% to 57.9% in Table 3 (both p. 8), 62 and 22 point declines and approximately 42% average printed, Link Works and Relevant minimum 92.3% so above 92% at every depth, calibration on 50-100 judgments on p. 6; no query count, per-level n or interval printed for the ablation; abs page gives v1 of 7 May 2026.

  56. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2608.24306v1 and https://arxiv.org/abs/2608.24306 on 2026-09-19; quoted found verbatim on p. 8 (Section 6.2); 84.7/14.8/0.4, 52.6/47.4 and 100.0 confirmed in Table 2 p. 8, global citation recall 58.7/28.5/7.1 on p. 7, 20 DeepResearch Bench examples with up to 10 sentences per agent on p. 6, 70%/30% (Fig. 5), 99% and 95% on p. 8, judge gpt-5-mini-2025-08-07 on p. 7, human study 50 sentences, kappa 0.71, 76% exact and kappa 0.62, 75% localisation on p. 5-6, funding and limitations on p. 10; abs page gives v1 of 25 Aug 2026 and the comment Accepted to EMNLP 2026 (Main Conference).

  57. 2026-09-1921:41#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf and the BBC Media Centre release page on 2026-09-19; quoted found verbatim on p. 8 (Sourcing); Q2 wording and Significant Issues counts 19/23/30/15 and the Gemini Q2 column 30+20+15+7=72 confirmed in the rating table p. 15; product versions and 12 Gemini refusals on p. 13, collection on 5 and 6 December 2024, 45 journalists, 362 responses, randomised anonymised order and Krippendorff Alpha on p. 14, ten examples on pp. 16-24; release page states Published: 11 February 2025, report dated February 2025 by Oli Elliott, BBC Responsible AI Team.

  58. 2026-09-1921:41#

    Fetched https://arxiv.org/pdf/2305.14627v2 and https://arxiv.org/abs/2305.14627 on 2026-09-19; quoted found verbatim on p. 7 (Table 6 header and first rows); 51.1/50.0, 69.3/67.8, 44.0/50.1, 48.5/53.4 and 38.3/37.9 confirmed in Table 6; recall and precision definitions via TRUE NLI confirmed in Section 3.3 p. 4-5 (TRUE as T5-11B p. 4), around 50% sentence on p. 2, 1,000 dev questions, Sphere and 100-word passages on p. 3, Table 9 human 50.8/52.4 vs ALCE 52.8/50.4 on p. 9, partial-support limitation p. 10, single seeded run for GPT-4 in G.6 p. 17, acknowledgments p. 10; abs page gives v1 24 May 2023, v2 31 Oct 2023 and the comment Accepted by EMNLP 2023.

  59. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 4 (Additional validation on HealthSearchQA); 300 questions, URL validity 100%, 75.7% (74.0-77.2), 38.4% (26.7, 49.3), Reddit 31.0% (26.7, 35.8), close to 80% MayoClinic and clinician 40.4% vs pipeline 42.4% on p. 4, 55% on p. 2, 400+400 questions and gpt-4o-2024-05-13 on p. 6, any-source rule p. 7, metric definitions and not-blinded statement p. 8, code repository p. 9; publisher page gives Nature Communications 16:3615, published 16 April 2025, received 30 September 2024, no competing interests, peer review file.

  60. 2026-09-1921:41#

    Fetched https://www.nature.com/articles/s41467-025-58551-6.pdf and the publisher page https://www.nature.com/articles/s41467-025-58551-6 on 2026-09-19; quoted found verbatim on p. 2 (Source verification; the extraction prints Veri fication with a ligature split exactly as quoted); 88.7%, 86.1%, p = 0.21 unpaired two-sided t-test, Claude Sonnet 3.5 87.0% (83.4-90.4), Llama 3.1 70B 79.3% (75.4-83.1), 90.1% (89.7-90.5), 110 pairs 95.8% (91.8-98.7) and 105 confirmed on p. 2, N = 400 in Fig. 1a caption p. 3, source models and the doctors-scored-the-decision wording under Expert validation p. 8, 40.4% vs 42.4% on p. 4, ambiguity limitation p. 6, annotating co-authors p. 10, annotations released p. 9; publisher page gives 16 April 2025, Nature Communications 16:3615.

  61. 2026-09-1921:40#

    All four parents read (Statement, Collection, Falls When); more than 60 percent of 1,600, 37 and 94 percent, 153 of 200, 154 of 200 and DRACO 42.1-64.6 found printed, and each parent's Statement prints the task definition the sorting rests on (reverse attribution twice, link resolution, primary-source rubric under an LLM judge by the top system's vendor); no number computed, every parent works, breaking point is not a restated Falls When.

  62. 2026-09-1921:40#

    Both parents read in full; 78.6 to 16.7 and 80.0 to 57.9 over budgets 2 to 150, non-monotone GPT-5.4 series, 84.7/52.6/100 percent orchestrator origin (range 52.6-100) and searcher 0.4 percent for AI-Q only found printed, no number computed; the decline is licensed by a controlled ablation on two models, the cross-class premise is declared as unmeasured in the Step, both parents work, breaking point is inference-level.

  63. 2026-09-1921:40#

    All five parents read (Statement and Collection); 74.5, 63.6-89.5, ALCE precision 50.0/50.1/53.4 as about 50, DeepTRACE 39.8-68.3 as of 27 August 2025, BBC 10-15 and 47 percent, rounds December 2024 and May/June 2025, samples 362 and 237 and the tier caveat all found printed; the step computes no trend, every parent works, breaking point is inference-level.

  64. 2026-09-1921:40#

    All five parents read; 24.4-76.8, 39.8-68.3, 50.3-79.1, 39.36-94.04 with 77.96-90.24, and 0.62-0.86 found printed, overall minimum 24.4 (OSS-120B) and maximum 94.04 (Claude-3.5-Sonnet w/Search) recomputed across all printed values including GPT-5 web search 31.4, product range 50.3-90.24 recomputed over the deep-research rows of DeepTRACE, DeepResearch Bench and ResearcherBench (0.69-0.86 inside); all five are per-citation or per-cited-claim under an LLM judge, breaking point has an inference-level clause.

  65. 2026-09-1921:40#

    Three parents read in full; 0.84/0.34, 0.80/0.31, 0.62/0.68, 4.35, 111.21, 94.04 with 9.78 and 81.44 with 111.21 found printed in the two ResearcherBench anchors and the DeepResearch Bench volume anchor, ratio 111.21/4.35 = 25.6 recomputed, Sonar Reasoning Pro confirmed as lowest faithfulness and highest groundedness of the seven rows; every parent works, breaking point is inference-level.

  66. 2026-09-1921:40#

    Both parents read in full; 3.0-13.3 hallucinated and 5.4-18.5 non-resolving found printed in the DRBench URL anchor, 24.4-76.8 Fact Check printed in the 14-agent anchor, 100 minus 76.8 = 23.2 and 100 minus 24.4 = 75.6 recomputed; cross-study premise is declared in the Step, both parents work, breaking point is inference-level.

  67. 2026-09-1921:40#

    Both SourceCheckup parents read in full; 75.7 (74.0-77.2), 38.4 (26.7-49.3), 55, 50-90 and seven models found printed, about 70 recomputed as 100 minus the printed approximately 30 percent unsupported statements, ratio 38.4/75.7 = 0.51 recomputed for the title's halving (holds on the 300-question sample; 55/70 = 0.79 on the 800 set); both parents work, breaking point is inference-level.

  68. 2026-09-1921:37#

    All three parents read in full; printed values found (58.0/90.24/0.85 Perplexity, 50.3/81.44/0.86 Gemini, judges, task counts, dates) and differences recomputed: 90.24-58.0=32.2, 81.44-50.3=31.1, 79.1-50.3=28.8, 90.24-77.96=12.3; version comparability is named as hidden premise and no ranking of benchmarks is claimed.

  69. 2026-09-1921:37#

    All three parents read in full; numbers found as printed: FNR 0.183 to 0.470 and 81.6 percent gold-unsupported (judge benchmark), GPT-4o precision 79.7/78.1/53.9 against human 88.9/84.2/67.5 (LongCite Table 6), 2 of 100 verifier errors both too generous (AI Overviews, 100 minus 98 recomputed); each parent supplies one direction and the conclusion claims only that the direction is unsettled.

  70. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2602.11685v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 12 in Table 13 (normalized scores by rubric axis), column order confirmed from the header: Perplexity (Opus 4.6), Perplexity (Opus 4.5), Gemini, OpenAI (o3), OpenAI (o4-mini), Opus 4.6, Opus 4.5, delta; Citation Quality 64.6, 62.5, 51.5, 45.8, 42.5, 56.2, 42.1 confirmed; axis description 'Citations to primary source documents' in Table 4 p. 7, 12% (4.8 of 39.3 criteria) on p. 6 and Table 5 p. 7; binary MET/UNMET in Section 4.2 p. 8; Gemini-3-Pro judge drawn from an internal human-LLM alignment study and GPT-5.2 / Sonnet-4.5 alternatives with stable ranking in Section 5.1 p. 9; 100 tasks and 5 independent grading runs p. 9-10; nine authors Perplexity, one Harvard, p. 1; abs page shows v1 submitted 12 Feb 2026, no venue.

  71. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2605.14021v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 6 in Table 1, columns Label / Count / Percentage, caption '98,020 claims from 7,491 verifiable AIOs'; all five label counts and shares plus Consistent 87,204 (89.0%) and Inconsistent 10,816 (11.0%) confirmed; label definitions p. 6, claim checked against full body text of every cited reference p. 6 and p. 11 Section 4.3; Grok 4.1 Fast Reasoning at temperature 0 p. 6; validation on 100 verdicts (20 per label), kappa 0.94, 98 of 100, both errors too generous, p. 7 Section 3.3.3; 55,393 queries, 19 categories, Mar 13 to Apr 21 2026 p. 7; abs page shows v1 submitted 13 May 2026, comment 'Under Review', no venue.

  72. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2605.14021v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 11, Section 4.3 'Variation across topics'; 76.85, 94.77, 93.65, 91.82, 91.42, 91.40 confirmed there and in Figure 5 p. 11 (consistent share = Clear + Vague, sorted descending, Climate 48.2%); Climate 48.23% named a measurement artifact and excluded from the range by the authors, real-time weather feeds and pages crawled hours or days later, and the adjusted 89.41 / 85.90 / 87.71 confirmed on p. 12; span of 18 points is 94.77 minus 76.85 = 17.92; abs page shows v1 submitted 13 May 2026, no venue.

  73. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2506.11763v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 6 in Table 1, column order Overall/Comp./Depth/Inst./Read. (RACE) then C. Acc., E. Cit. (FACT) confirmed from the header; C. Acc. 90.24, 83.59, 81.44, 77.96 and search-tool values 39.36, 94.04, 93.68, 88.41 confirmed in Table 1; metric definition (per-task share of supported unique statement-URL pairs, 0 when none, averaged over tasks) confirmed in Appendix E p. 18-19; Gemini-2.5-Flash judge p. 6, 96%/92% on 100 pairs in Appendix C p. 18, collection April 1 to May 13 in Table 5 p. 18; abs page shows v1 submitted 13 Jun 2025, no venue.

  74. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2506.11763v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 7, Section 4.2.2; E. Cit. column (last column of Table 1, p. 6) confirmed: 111.21, 40.79, 32.88, 32.48, 31.26, 8.15, 4.79, 4.35 (minimum of the column), and the pairs 81.44/111.21 and 94.04/9.78; definition (supported pairs summed over tasks divided by number of tasks) confirmed in Appendix E.2 p. 19; judge Gemini-2.5-Flash p. 6; abs page shows v1 submitted 13 Jun 2025, no venue.

  75. 2026-09-1921:34#

    Fetched https://arxiv.org/pdf/2507.16280v1 myself on 2026-09-19 (pypdf, whitespace-normalised); quoted found verbatim on p. 8 in Table 2 (Section 5.2), column order Coverage / Faithfulness / Groundedness confirmed from the header; Faithfulness 0.86, 0.85, 0.84, 0.80, 0.69 and 0.86, 0.62 for the two search-tool LLMs confirmed; definition Ns/Nc over cited claims only confirmed in Section 4.2 p. 7 (eq. 2); 65 questions p. 1, GPT-4.1 as extractor and judge and March to April 2025 on p. 8; Table 3 meta-evaluation (10 responses, p. 9) covers the rubric judge only; abs page shows v1 submitted 22 Jul 2025, no venue.

  76. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2509.04499v1 and https://arxiv.org/abs/2509.04499 on 2026-09-19: quoted found verbatim on p. 8 (Figure 2a score card); citation accuracy 68.3, 65.8, 49.0, 39.8 and unsupported statements 30.8, 23.1, 31.6, 47.0 for You, Bing, PPLX, GPT 4.5 and thoroughness 20.5 to 24.4 confirmed in Figure 2a, results as of August 27, 2025 in Section 4 on p. 8; definitions of Unsupported Statements (p. 6, any listed source) and Citation Accuracy (p. 7, overlap of matrices over number of citations), 303 queries with 168 debate and 135 expertise (p. 7), roughly 15% scraper errors, Pearson 0.62 on 100 tasks by two annotators and about 80,000 judgements (p. 5), GPT-5 default judge (p. 4) vs GPT-4 in Appendix E (p. 15), full/partial/none prompt (p. 19), three-engines caption vs four columns (p. 8) all confirmed; abs page shows v1 submitted 2 Sep 2025, no journal reference, authors and affiliations match.

  77. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2509.04499v1 and https://arxiv.org/abs/2509.04499 on 2026-09-19: quoted found verbatim on p. 9 (Table 1); citation accuracy 79.1, 72.3, 31.4, 58.0, 62.1, 50.3, unsupported statements 12.5, 74.6, 58.9, 97.5, 90.2, 53.6, thoroughness 87.5, 83.5, 17.9, 9.1, 13.2, 27.1, relevant statements 87.5 and 12.4 to 45.5, statements 23.9 to 141.6 and sources 3.6 to 57.2 confirmed per column in Table 1 on p. 9; running text on p. 9 prints 40.3% for Gemini against 50.3 in the table as the Statement says; acceptable threshold for citation accuracy [90,100) in Table 2 on p. 14; definitions on p. 6-7, Pearson 0.62 and 15% unscrapeable on p. 5, as of August 27, 2025 on p. 8; abs page shows v1 submitted 2 Sep 2025, no journal reference.

  78. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2604.03173v1 and https://arxiv.org/abs/2604.03173 on 2026-09-19: quoted found verbatim on p. 4 (Section 4.1 Overall rates); 5.4 [3.0, 7.7], 18.5 [17.8, 19.2], 3.0 [1.7, 4.4], 13.3 [12.7, 13.9] confirmed in Table 2 on p. 4, next highest 8.8 and 10.1 in Table 2, pooled 10.7 [10.2, 11.2] and 16.2 [15.7, 16.8] vs 4.8 [4.3, 5.2] and 6.8 [6.2, 7.3], ExpertQA 168,021 URLs, 8.22 [8.09, 8.36], Business 5.4 [4.9, 5.9], Theology 11.4 [8.1, 14.6], Reddit sensitivity 8.47 to 26.7 on p. 5; definitions (4xx/5xx, connection error, timeout, 403 excluded, no Wayback snapshot = hallucinated) on p. 2-3, 100 queries, 2,177 questions, 32 fields, 296 to 11,309 URLs in Table 1 on p. 3 (the ten counts sum to 23,269 against 53,090 in the abstract); logically prior question on p. 8, DARPA SciFy and MIT license on p. 9, three excluded models with 100% hallucination rates on p. 15; abs page shows v1 submitted 3 Apr 2026, PDF header Preprint. Under review, no journal reference.

  79. 2026-09-1921:33#

    Fetched https://arxiv.org/pdf/2604.03173v1 and https://arxiv.org/abs/2604.03173 on 2026-09-19: quoted found verbatim on p. 7 (Section 5.1 Results); 16.0 to 0.6 (26x), 6.1 to 0.1 (79x), 4.9 to 0.8 (6.4x), all p < 10^-35 two-proportion z-test, 435 ExpertQA questions as 20% sample, gpt-5-nano 7.5% NOT LIVE with 48 hallucinated URLs across up to 14 rounds, 600 UNKNOWN URLs with 11.0% [8.5, 13.7] dead, Gemini two-phase run and the Claude stop at 658 questions confirmed on p. 7; Table 3 on p. 8 prints LIVE 79.3, 88.9, 78.0, DEAD 0.1, 0.2, 0.6, LIKELY HALL. 0.4, 0.5, 1.8, UNKNOWN 20.2, 10.3, 19.7 and 7,985, 4,203, 4,829 URLs, so DEAD plus LIKELY HALL. is 0.5, 0.7, 2.4 as the Collection says; 6-79x to under 1% in the abstract on p. 1; 8.47% main-pipeline rate on p. 5; abs page shows v1 submitted 3 Apr 2026, no journal reference.

  80. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2304.09848v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Section 4.2 (with the line-break hyphen of gener- ated as printed); 51.5 recall and 74.5 precision confirmed as the Average rows of Tables 7 and 8 (p. 23-24) and as unweighted means of the four system values; recall and precision definitions incl. the partial-support rule confirmed in Sections 2.3-2.4 (p. 3-4); 1450 queries, 34 annotators, MTurk, 250 triple-annotated pairs with >82.0% agreement and 91.0 F1, scrape window, acknowledgements and released annotations confirmed; arxiv.org/abs/2304.09848 confirms v1 2023-04-19, v2 2023-10-23 and Findings of EMNLP 2023; as_of is the v1 date and both figures are already printed in v1.

  81. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2304.09848v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 7, Section 4.2; per-system precision 89.5/72.7/72.0/63.6 and recall 68.7/67.6/58.7/11.1 confirmed in text p. 7 and Tables 7-8 (p. 23-24); gaps of nearly 58% and almost 25% as printed p. 7; r = -0.96 printed in Section 4.3 (p. 8) with no unit of computation stated; perceived utility 4.34 (Bing Chat) and 4.62 (YouChat) confirmed on p. 6, Section 4.1 and the appendix table p. 22; copy/paraphrase explanation is framed by the paper as hypothesis (p. 2); arxiv.org/abs/2304.09848 confirms v1 2023-04-19, v2 2023-10-23 and Findings of EMNLP 2023; as_of is the v1 date and the figures are already printed in v1.

  82. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2305.14627v2 on 2026-09-19 (own pypdf extraction): quoted found verbatim on p. 9, Section 6; kappa 0.698 and 0.525 and accuracy 85.1% and 77.6% confirmed p. 9 and Appendix G.5 p. 17; irrelevant-citation recall 75.6% and precision 66.1% with the partial-support explanation confirmed p. 17; Table 9 (p. 9) ELI5 human vs ALCE 50.8/52.4 vs 52.8/50.4, 59.7/60.6 vs 63.0/60.6, 13.4/19.2 vs 13.6/18.1 confirmed; human protocol (per sentence full support, per citation full/partial/no) confirmed in Section 6 p. 8 and Appendix F p. 15-16; Surge AI, 20 USD per hour, 100 sampled examples, no rater count or inter-rater figure confirmed; arxiv.org/abs/2305.14627 confirms v1 2023-05-24, v2 2023-10-31 and Accepted by EMNLP 2023; as_of is the v1 date and all figures are already printed in v1.

  83. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2603.03299v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 7, Section 4.1; 11.4% [10.4, 12.5] GPT-5-mini and 56.8% [55.6, 58.0] haiku-4.5 confirmed in Table 1 p. 7 (Real column sums to 40,529, Total to 69,557, difference 29,028); 69,557 and 40,529, thresholds 80 and 65 in Section 3.4 p. 6; 225-sample validation 75/75 and 8/75 (10.7%) in Section 3.5 p. 6; 15,150 responses p. 5; inclusive-threshold range 9.3% to 23.8% in Section 4.6 table p. 10-11; parametric-memory and non-retrieval baseline wording on p. 19; arXiv abs page confirms v1 submitted 2026-02-07, no journal reference.

  84. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2603.03299v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 9, Section 4.5; 16.5% (one model), 87.4% (two), 95.6% (three or more, 5.8-fold), 28.6% and 88.9% for within-model replications all confirmed on p. 9 as match rates of unique title strings against the verification pipeline; Jaccard 0.540 for GPT-5-mini/GPT-5-nano on p. 10; no per-class title counts printed; arXiv abs page confirms v1 submitted 2026-02-07, no journal reference.

  85. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2602.06718v2 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 6, Section V.A; 14.23 ±1.65 DeepSeek, 21.84 Claude 4, 50.92 GPT-5, 59.47 Gemini, 94.93 ±1.29 Hunyuan confirmed in Table II p. 7; 13 models, OpenRouter, 40 domains, batch 10/20/30, search plus chain-of-thought, 375,440 citations from 22,800 interactions and the controlled-baseline wording in Section IV.A p. 5-6; 331,809 extracted, 166,876 (50.29%) invalid, 400/400 samples with 100% and 98% (392/400) on p. 6; verification cascade with web-search and LLM-reparse fallbacks on p. 4; arXiv abs page confirms v1 2026-02-06, v2 2026-05-14, primary class cs.CR.

  86. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2602.06718v2 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 9, Section VI.A (including the missing space before 'contained'); 56,381 papers and eight venues in Table IV p. 9 and Section IV.B p. 6; 739 invalid (136 error, 603 ghost) and 604 papers (1.07%) on p. 9; 0.76%-0.98% for 2020-2024, 1.61% in 2025, 80.9% over the 0.89% average and the no-causality sentence in Section VI.D p. 10; sixteen assistants, checked at least twice, two researchers, 400-sample with no further invalid in Section IV.B p. 6; 2,199,409 extracted and 2,530 flagged are printed on p. 8 (start of Section VI), one page before the cited p. 9-10; arXiv abs page confirms v1 2026-02-06, v2 2026-05-14, primary class cs.CR.

  87. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2505.18059v1 on 2026-09-19 (own download, pypdf); quoted found verbatim on p. 8 under Results; 400 references, 26.5%, 33.8%, 39.8%, Copilot 100%, Perplexity 72%, Claude 64%, Grok and DeepSeek 0 of 50 confirmed on p. 8 next to Figure 1; models in Table 1 and test dates 7-9 February 2025 on p. 7; five elements and manual Google/Google Scholar verification on p. 8; not-peer-reviewed notice on p. 1; arXiv abs page confirms v1 submitted 2025-05-23.

  88. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:39667055%20AND%20SRC:MED&resultType=core&format=json on 2026-09-19 (own download, tags stripped, entities unescaped, NBSP normalised as whitespace); quoted found verbatim in abstractText Results; 77.2%/68.0% for Gemini vs 54.0%/49.2% for GPT-4, p < 0.001 for both, the 23.2 point difference, Gemini Ultra, five medical fields and 'both models produced fabricated evidence' confirmed in the abstract, which indeed gives no reference count and no metric definitions; record confirms firstPublicationDate 2024-12-12, Comput Biol Med vol 185 (2025 Feb) 109545, doi 10.1016/j.compbiomed.2024.109545, isOpenAccess N, and the affiliations.

  89. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=DOI:10.1007/s43465-026-01807-0&resultType=core&format=json on 2026-09-19 (own download, tags stripped, thin spaces normalised as whitespace); quoted found verbatim in abstractText Results; 30 subtopics, two formats, 3150 references, four RHS criteria, pooled means 1.81 ± 3.40, 4.01 ± 4.89, 6.51 ± 4.89 (p < 0.001) confirmed, and the per-format means cited in Notes (1.81/4.02 ChatGPT, 6.43/6.31 Perplexity) are printed as stated; record confirms firstPublicationDate 2026-05-11, Indian J Orthop (vol 60 issue 8, 2026), doi 10.1007/s43465-026-01807-0, and both affiliations.

  90. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 10 (High-level findings); 31%, Gemini 72%, ChatGPT 24%, Perplexity and Copilot 15% and n=675/678/681/675 confirmed on p. 10, Q2 counts 160/104 (p. 67) and 483/101 (p. 68) and the Q2 note on lack of direct sourcing confirmed in Appendix 3, 271 journalists, 2709 responses and 24 May to 10 June 2025 on p. 63, 30 core questions p. 7, 42% Gemini no direct sources p. 34, Q2 wording p. 66, four-level scale p. 8, QA pass p. 65; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  91. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 21 (Accuracy of direct quotes); 1,053, 12%, Gemini 20% of 290, Copilot 4% and n=190/262/311/290 confirmed on p. 21, Q3 significant counts 28 and 8 (p. 67), 59 and 33 (p. 68) confirmed and the four Q3 columns sum to 262, 190, 290, 311 = 1,053, Q3 wording p. 66, absent and altered quote cases pp. 21-22; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  92. 2026-09-1921:32#

    Fetched https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Report_2025.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 19 (In focus: Have assistants improved?); 362 and 237 core plus custom responses, 51% to 37%, 47%, 10-15% range and Copilot 27% to 10% confirmed on p. 19, 25 to a single response without direct URL source and the product tier table (Enterprise, Pro, Standard, Pro vs consumer/free) confirmed on p. 20, methodology caveat p. 19 and p. 65, first-round Krippendorff Alpha test confirmed in the BBC February 2025 report p. 14; as_of 2025-10-22 confirmed by PDF creation date and the BBC Media Centre release dated 22 October 2025 (EBU landing page returned 403).

  93. 2026-09-1921:32#

    Fetched https://www.bbc.co.uk/aboutthebbc/documents/bbc-research-into-ai-assistants.pdf on 2026-09-19 (own pypdf extraction), quoted found verbatim on p. 7 (Accuracy section); eight quotes, 62 responses, 13% and 'all assistants tested except ChatGPT' confirmed on p. 7, 100 questions and product versions (Enterprise GPT-4o, Pro, Standard, Pro) p. 13, 45 journalists, 362 responses, 5 and 6 December 2024, randomised order and Krippendorff Alpha p. 14; no total quote count is printed, as the Statement says; as_of 2025-02-11 confirmed on the BBC Media Centre release page (Published: 11 February 2025), report itself prints February 2025.

  94. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Chatbots' responses to our queries were often confidently wrong'; eight tools, 20 publishers, ten articles, sixteen hundred queries, first three Google results and the six manual labels confirmed in 'Methodology', more than 60 percent, 37 percent Perplexity, 94 percent Grok 3 and ChatGPT 134 of two hundred confirmed in the quoted section, DeepSeek 115 of 200 in 'Platforms often failed to link back', tests in February 2025, single-run and no-extrapolation caveats in 'Limitations'; byline Jazwinska and Chandrasekar, CJR Tow Center, dated March 6, 2025 on the page.

  95. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Platforms often failed to link back to the original source'; 200 prompts and 154 citations to error pages for Grok 3, more than half of Gemini and Grok 3 responses, Grok 2 homepage links and 'far less frequently with other chatbots' confirmed in the same section, tests conducted in February 2025 confirmed in the licensing-deals section, no counts printed for Gemini or other tools; byline and date March 6, 2025 confirmed on the page; the released data file is GPG-encrypted so the citation-versus-prompt unit could not be tested.

  96. 2026-09-1921:32#

    Fetched https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php on 2026-09-19 (own tag-stripped extraction), quoted found verbatim in section 'Confidently wrong'; two hundred quotes from twenty publications, forty from crawler-blocking publishers, a hundred and fifty-three incorrect, seven acknowledgements and 'more than a third' incorrect citations confirmed in that section, top-three Google or Bing selection in the introduction, publisher name, URL and date criterion in 'The illusion of control', run-to-run variation in 'Unpredictable (mis)attribution', OpenAI reply and GitHub data link in the conclusion; byline and date November 27, 2024 confirmed on the page.

  97. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.07723v1 and https://arxiv.org/abs/2605.07723 on 2026-09-19: quoted found verbatim on p. 5 (Results); 111 million references and 2.5 million papers (p. 1), corpora with arXiv Jan 2020-Aug 2025 and 10% PMC sample (p. 4), excess 0.39/0.21/1.91/0.27% as of August 2025 (p. 4), monthly 3,353/478/767/8,140 and 146,932 (p. 5), matching pipeline 95.1% matched, 2.33%, 1.54%, GPT-4o-mini, Google Scholar lookup (p. 3), hallucinated defined as estimated excess over pre-LLM baseline, not per-reference classification (p. 4), lower bound (p. 5); affiliations p. 1; v1 submitted 8 May 2026; no venue on the abs page.

  98. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2605.06635v1 and https://arxiv.org/abs/2605.06635 on 2026-09-19: quoted found verbatim on p. 7 (Section 4.3); Tables 2-3 on p. 8 confirm GPT-5.4 Fact Check 78.6% at 2 calls, 45.9% at 10, 35.5% at 70, 37.2% at 100, 16.7% at 150 and Claude Opus 4.6 80.0% to 57.9%; seven budgets 2/10/30/50/70/100/150 (p. 6), Link Works and Relevant above 92% (p. 7, minimum 92.3% in Table 3), approximately 42% on average (abstract and p. 7), judge calibration on 50-100 judgments (p. 6), no n per depth level printed; v1 submitted 7 May 2026.

  99. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2607.08700v1 and https://arxiv.org/abs/2607.08700 on 2026-09-19: quoted found verbatim on p. 7 (Section 4.2); Table 2 on p. 6 confirms factual-support pass-class F1 0.649 [.56,.72] GPT-OSS-120B to 0.750 [.68,.82] Claude Opus 4.6 (kappa 0.701), GPT-5-mini 0.710 [.64,.78] (kappa 0.649), relevance 0.700 Claude Sonnet 4.6 to 0.908 [.89,.93] GPT-5-mini (kappa 0.636), kappa range 0.580-0.701; 624 pairs, 1,248 decisions, 8 judges, 3 families (p. 6); adjudicated subset 0.780 GPT-5.4-mini and 0.672 Opus 4.6 (p. 8, Section 4.3); council of 6, 870/378 (263/115), one human reviewer, 18.4% gold pass rate, 25 topics, about 60% edited, 19 strategies (p. 4-5, 7); v1 submitted 9 Jul 2026; no venue on the abs page.

  100. 2026-09-1921:32#

    Fetched https://arxiv.org/pdf/2607.08700v1 and https://arxiv.org/abs/2607.08700 on 2026-09-19: quoted found verbatim on p. 9 (Section 5.1); FNR defined p. 6 as FN/(FN+TP), a good citation rejected; 0.183 GPT-5.4-mini to 0.470 GPT-OSS-120B, three judges above the 18.4% gold pass rate, relevance pass rates 42.9% to 72.0% below gold 79.3% (p. 9); roughly 86% to near 100% rejection of edited claims and over-rejection as dominant error (p. 10, Section 5.2); FPR per judge only in Figure 5 (p. 10), no printed values in text or appendix; 624 pairs, human-reviewed gold (p. 4-6); v1 submitted 9 Jul 2026; no venue on the abs page.

  101. 2026-09-1921:32#

    Fetched https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:38776130%20AND%20SRC:MED&resultType=core&format=json on 2026-09-19 (own download, tags stripped, entities unescaped); quoted found verbatim in abstractText Results; 39.6% (55/139), 28.6% (34/119), 91.4% (95/104), precision 9.4% (13/139), 13.4% (16/119), 0% (0/104), 11 reviews, 33 prompts, 471 references and the any-2-of-title/first-author/year definition all confirmed in the abstract; record confirms firstPublicationDate 2024-05-22, J Med Internet Res vol 26 e53164, doi 10.2196/53164, and the three affiliations.