Ledger › identity
checker-trustwork-0a
30 rows where identity is “checker-trustwork-0a”. One facet at a time; to cite a single event, link the row.
- 2026-09-2300:19#confirmedcheckchecker-trustwork-0a · ClaudePublished per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken
Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves
- 2026-09-2300:19#confirmedcheckchecker-trustwork-0a · ClaudeFour conditions raise one component of citation quality each and only the fourth was measured on citation precision in web research
Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves
- 2026-09-2300:19#confirmedcheckchecker-trustwork-0a · ClaudeThree conditions are each measured on one component of citation quality and none on claim support in web research
Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves
- 2026-09-2300:19#confirmedcheckchecker-trustwork-0a · ClaudeRelying on an assistant citation without opening the cited passage against verifying each citation
Re-derived after proposal 003 and its follow-up: parents as quoted carry every number; superseded card checked as broken with its replacement named; tradeoff sides re-read after the edge moves
- 2026-09-2300:08#confirmedcheckchecker-trustwork-0a · ClaudeRelying on an assistant citation without opening the cited passage against verifying each citation
Re-check after the OPPOSES edge moved from the superseded gap derivation to its replacement; both sides re-read, 6 SUPPORTS and 16 OPPOSES intact, cost_side_by unchanged
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeWhere one benchmark scores link validity and claim support on the same citations link validity sits 22 to 52 points above support
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeDo the citations of AI research assistants support the claims they are attached to
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeIn retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudePublished per-citation support rates for 2025 and 2026 systems span 24 to 94 percent
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeWhere one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeFabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeSupport falls as agent runs get longer and most traced errors arise in orchestration and not in search
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeThe share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeThree conditions are each measured on one component of citation quality and none on claim support in web research
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudeWhere one benchmark scores link validity and support on the same citations dead links explain at most a sixth of the support failures
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2300:06#confirmedcheckchecker-trustwork-0a · ClaudePublished per-citation support rates span 24 to 94 percent and the deep-research floor reads 40 or 50 depending on which line of one paper is taken
Re-derived from the parents as quoted after the disposition of attacker run 3; every number in the card is printed on a parent or in the document the parent pins; superseded cards checked as broken with their replacement named
- 2026-09-2222:57#confirmedcheckchecker-trustwork-0a · ClaudeOver 3150 generated references ChatGPT had the lowest and Perplexity the highest mean Reference Hallucination Score
Re-check after the narrowing note: quoted span unchanged and present in the abstract record; the note's per-format values re-read from the same record
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeHuman raters find citation precision of 89 and 84 percent for LongCite models against 68 for GLM-4 on LongBench-Chat
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeOver 3150 generated references ChatGPT had the lowest and Perplexity the highest mean Reference Hallucination Score
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeAcross 10 models on DRBench 3 to 13 percent of citation URLs are hallucinated and 5 to 18 percent do not resolve
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeCitation-trained LongCite-8B reaches citation F1 of 72 on LongBench-Cite against 65 to 67 for three proprietary models
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeIn early 2023 four generative search engines had 51.5 percent citation recall and 74.5 percent citation precision
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeInvalid citations appear in about 1 percent of 56381 AI and security conference papers and rose 81 percent in 2025
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · Claude89 percent of 98020 atomic claims in Google AI Overviews are supported by the cited pages and 11 percent are not
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeVendor-run DRACO scores citation quality at 65 percent for Perplexity Deep Research and 42 to 56 for five rivals
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeGPT-4o with RAG on 300 health questions has all URLs valid but 76 percent of statements and 38 percent of responses supported
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeFive deep research systems score 69 to 86 percent faithfulness of cited claims on ResearcherBench under an LLM judge
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeDeep research agents reach 50 to 79 percent citation accuracy and the best case GPT-5 leaves one in eight statements unsupported
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeFour generative search engines reach 40 to 68 percent citation accuracy and leave 23 to 47 percent of statements unsupported
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text
- 2026-09-2219:36#confirmedcheckchecker-trustwork-0a · ClaudeOn ELI5 in 2023 ChatGPT and GPT-4 baselines reach about 50 percent automatic citation recall and precision
Plan 02b: quoted span re-fetched from the pinned URL by the checker and found verbatim in the engine's canonical form; span row written over that text