How it works

Trustillery runs on Memstead, a graph engine for typed cards. The nine terms below are not house style; they are the engine's schema, and it refuses what does not fit. How this is built → · The engine underneath →

A question
One question in the words a reader would use. Beneath it a scope and a list of what is not asked, written by the owner's agent and confirmed by the owner. The precision lives in the cards, not in the title.
An anchor
Something measured or observed, never derived. It names the document, the page, the verbatim span, who collected it and whom that body answers to. It states in advance under which recomputation or counter-measurement it falls.
A derivation
What follows from named parent cards, with the step written out and a breaking point declared before anyone attacks it.
A tradeoff
Where derivation ends and judgement begins. Incomplete until both sides carry entries; the cost side is written by someone who did not see the benefit side.
A check
A second identity re-reads a card against the fetched source and records a verdict with its method. The engine measures whether checker and author are distinct; a stranger will also ask whether they are the same model family, and the ledger says.
An attack
A model of another family than the authors' reads every derivation with its parents and records a verdict with a reason per card, a verdict on the question's scope, and the measurements the build missed. The owner answers every finding: applied, or refused with the error named.
A proposal
A fork of the stock. The owner adopts, adopts with changes or rejects each card with a reason; the record is published with the stock, rejections included. Proposals by the project's own agents are labelled as such.
Where it ends
A question may end in an answer, in a range with its conditions, or in a tradeoff. All three are results. The stock never issues a verdict on who was right.
An edition
The stock is published in one edition a day, at a fixed hour. Nothing changes between editions except what the law requires, and that is marked as an off-cycle edition. What is handed over (a finding, a proposal, a correction) is written the moment it is handed over, with its own time, never edited afterwards, only followed; it appears in the next edition. Until the daily build runs, editions are cut by hand and dated. Trust builds slowly.

Same question, with and without the stock

The one comparison the project has measured, on 2026-09-23. Three delegated agents on the same model answered the same raw question. The owner read X and Y blind, the assignment sealed by a hash before any result existed, then recorded one sentence: “Arm X wirkt klar besser und hat den Bestand.” (Arm X seems clearly better and has the stock.). His guess of which arm had the stock was right.

What the stock added was not the papers: the web-search arm found the same five. It added the joints: the counting unit, the judge's validity, the spread of one product across benchmarks, and what was never measured.

ArmConditionTokensWall-clock
YWithout the stock · same model, no tools, answers from memory34,31446 s
XWith the stock · same model, read the whole export, nothing else140,35094 s
CWeb search, no stock · same model, ordinary web research, shown after the verdict49,100109 s
SStranger session: fresh agent, only the exported HTML, asked to name one attackable card and its condition151,34987 s to the card
Arm Y · Without the stock · same model, no tools, answers from memory · 34,314 tokens

Short answer, from memory (no tools used): depending on how strictly "supports" is measured, published 2024–2026 measurements put the share of citations that actually back the attached claim somewhere between roughly half and four fifths, with fabricated or dead links on top. No single headline number exists; the figures cluster by evaluation type.

1. Sentence-level verification of AI search engines (the strictest, best-designed measurements) - Liu, Zhang & Liang (Stanford, 2023, widely re-cited through 2024–25) on Bing Chat, Perplexity, NeevaAI, YouChat: about 74% of citations supported their sentence, and only about 52% of generated sentences were fully supported. This remains the anchor figure. - Wu et al. (Stanford, 2024; later Nature Digital Medicine 2025, "SourceCheckup") on medical questions: even GPT-4 with retrieval left roughly 30% of individual statements unsupported and about half of all responses contained at least one unsupported statement; without retrieval the unsupported share was far higher. Their broad finding: 50–90% of responses were not fully supported. - Asai et al. (OpenScholar / ScholarQABench, Nov 2024): GPT-4o fabricated citations in about 78–90% of scientific-literature answers when not grounded; retrieval-grounded systems approached expert-level citation accuracy.

2. Fabricated or broken references - Chelli et al. (JMIR 2024): fabricated references for systematic-review prompts of ~29% (GPT-4), ~40% (GPT-3.5), ~91% (Bard). - Tow Center / CJR (Jaźwińska & Chandrasekar, March 2025), 1,600 queries across eight AI search tools: over 60% of answers misattributed or wrongly cited the source; Perplexity best (~37% wrong), Grok 3 worst (~94%); Gemini and Grok 3 pointed to fabricated or dead URLs in more than half of cases; premium tiers were more confidently wrong, not more accurate. - Magesi et al. (Stanford RegLab, 2024) on legal research tools: Lexis+ AI and Westlaw AI hallucinated in roughly 17% and 33% of queries, including misgrounded citations.

3. Deep-research agents (2025–26 benchmarks) - DeepResearch Bench (Du et al., mid-2025) reports "citation accuracy" of roughly 78–90% for OpenAI, Gemini and Perplexity deep-research modes (Perplexity highest at ~90%, OpenAI lowest at ~78%) — but this checks whether the cited URL contains the claim in a lenient, LLM-judged way, and effective-citation counts differ enormously between systems. - Mind2Web 2 (Gou et al., 2025) with agent-as-judge attribution checks found OpenAI Deep Research reaching only about 50–70% of human-level verified performance; unsupported attributions were a main failure mode. CiteME (Press et al., 2024) put citation-attribution agents at ~35% accuracy.

Synthesis and uncertainty. A defensible summary: about 70–80% of citations from grounded, retrieval-backed systems point to a source that supports the claim on lenient checks; under strict claim-by-claim verification, only about 50–65% of claims are fully supported; ungrounded models fabricate a majority of references. I am confident about the Liu et al., Wu et al., Tow Center and Chelli figures within a few points. I am less sure of the exact DeepResearch Bench and Mind2Web 2 numbers (the ranking is right, the percentages may be off by 5–10 points), and I recall no rigorous 2026 measurement well enough to quote it.

Arm X · With the stock · same model, read the whole export, nothing else · 140,350 tokens

Short answer: there is no single published value. Per-citation support rates measured in 2025–2026 span roughly 24 to 94 percent, and the stock's own summary card concludes that "an answer to the question is a range with its conditions, not a number."

What the measurements print (all LLM-judged, per attached citation or per cited claim)

  • *Cited but Not Verified* (arXiv 2605.06635, May 2026): 14 models run as deep-research agents on 130 queries. Fact Check (cited page supports the claim) runs from 24.4% (OSS-120B) to 76.8% (Claude Opus 4.5); GPT-5.4 at 47.7%. Meanwhile 12 of 14 keep links valid above 94%, so link validity sits 22 to 52 points above support on the same citations.
  • *DeepTRACE* (arXiv 2509.04499, Aug 2025): four generative search engines 39.8–68.3% citation accuracy; deep-research configurations 50.3–79.1% (GPT-5 Deep Research best; Gemini printed as 50.3 in the table and 40.3 in the text).
  • *DeepResearch Bench* (arXiv 2506.11763, 2025): four deep-research agents 77.96–90.24%; twelve search-tool LLMs 39.36–94.04%.
  • *ResearcherBench* (arXiv 2507.16280): faithfulness of cited claims 0.62–0.86, but only 31–68% of factual claims carried any citation at all.
  • *DeepScholar-Bench*: OpenAI DeepResearch 39.9% citation precision; *ReportBench*: 78.87% for the same product.

Upper bounds counted by a more lenient unit (supported by *any* cited page, not the attached one): Google AI Overviews 89.0% of 98,020 claims (2026, verifier validated 98/100 against humans); SourceCheckup 75.7% of statements for GPT-4o with search, falling to 38.4% when counted per whole response.

Why the range does not collapse to a number

1. The same product moves 31–32 points between benchmarks (Perplexity Deep Research 58.0 in DeepTRACE vs 90.24 in DeepResearch Bench; Gemini 50.3 vs 81.44), as much as products differ within one benchmark. 2. Support falls with search depth: raising tool calls from 2 to 150 dropped Fact Check from 78.6% to 16.7% (GPT-5.4) and 80.0% to 57.9% (Claude Opus 4.6), while link validity stayed above 92%. 3. Every 2024–2026 per-citation rate is LLM-judged. Validated judges disagree with humans on 2–22% of decisions, and the direction is unsettled: two comparisons show judges stricter than humans (rejecting 18–47% of supported citations), one shows a judge more generous. 4. The only human-rated per-citation audit (Liu et al., 74.5% precision, 63.6–89.5 by engine) is from early 2023 and outside the window; no human-rated measurement on ordinary user queries exists for current products.

What is consistent across studies: dead or fabricated links (3–13% hallucinated, 5–18% non-resolving for retrieval-backed systems) explain at most a sixth of support failures; the dominant failure is a real, on-topic page that does not say what the sentence claims. Reference-existence studies (11–95% fabricated) and reverse-attribution tests (Tow Center, >60% wrong) measure neighbouring quantities and do not answer the support question.

Practical reading: for a current commercial deep-research product, roughly half to nine-tenths of attached citations support their claim depending on benchmark, task and judge; for the best documented cases (GPT-5 Deep Research, Claude with search) about one in five to one in ten citations still fails.

Source: the llms-full export of the stock, read in full; cards cited: the question card, "Published per-citation support rates span 24 to 94 percent...", "The share of supporting citations has no single published value...", and the anchors named above.

Arm C · Web search, no stock · same model, ordinary web research, shown after the verdict · 49,100 tokens

Short answer: across the 2024–2026 measurements, only about 40–80% of the citations attached to claims actually support them, and the spread between systems is wider than the spread between studies. No frontier deep-research product has been measured above ~90% on a claim-level support check, and several sit near or below 50%.

Baseline (pre-deep-research). Liu, Zhang & Liang's Stanford audit of Bing Chat, Perplexity, NeevaAI and YouChat found only 74.5% of citations supported their sentence and only 51.5% of sentences were fully supported (https://arxiv.org/abs/2304.09848). Later work has mostly landed below this, not above.

DeepResearch Bench (June 2025). The FACT framework extracts statement–URL pairs and asks whether the page supports the statement. In the original paper: Perplexity DR 90.2%, Grok 83.6%, Gemini-2.5-Pro DR 81.4%, OpenAI DR 78.0%, Claude-3.7 with search 93.7% (https://arxiv.org/html/2506.11763). Notably, the live leaderboard has since been revised downward under a stricter judge to 44–49% for the same four products (https://deepresearch-bench.github.io/), so the headline "80–90%" figures should be treated as the lenient end.

DeepTRACE (ICLR 2026). 303 questions, GPT-5 judge validated against human annotators. Citation accuracy (share of citations whose source supports the statement): GPT-5 DR 79%, You.com DR 72%, Copilot DR 62%, Perplexity DR 58%, Gemini DR 50%; among plain answer engines, You.com 68%, Bing Copilot 66%, Perplexity 49%, GPT-4.5 40%. Unsupported-statement rates for some DR products were extreme (Perplexity DR 97.5%, Copilot DR 90.2%) because they produce many claims and spread citations loosely (https://arxiv.org/html/2509.04499; https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf).

"Cited but Not Verified" (May 2026). 14 models on 130 queries, checking links, relevance and factual support separately. Frontier models keep links alive >94% and topically relevant >80%, yet only 39–77% of citations pass the fact check; open-source models 24–51% (https://arxiv.org/html/2605.06635v1). The gap between "relevant" and "supports" is the key finding: most bad citations point to real, on-topic pages that simply do not say what is claimed.

Fabrication is the smaller problem. A March 2026 audit of 221k URLs from Claude, Gemini and GPT deep-research outputs found only 3–13% of citation URLs are outright hallucinated (Claude ~3%, Gemini DR ~13%) (https://arxiv.org/html/2604.03173v1). So for a frontier agent, roughly a tenth of citations are fake, but a further quarter to half are real yet non-supporting.

Open-source agents are worse. Fault-localisation work on the top three open DeepResearch-Bench systems reports citation recall of 59%, 29% and 7% (https://arxiv.org/html/2608.24306).

Caveats. All numbers come from LLM judges with moderate human agreement (DeepTRACE reports r≈0.62), definitions differ (per-citation precision vs per-statement recall), and products change monthly. The robust reading is: assume one in four to one in two citations from a deep-research report does not back its claim, with GPT-5 DR at the better end and Gemini/Perplexity DR at the worse end.

Stranger session · transcript

Stranger session of the local acceptance test, plan 04. A fresh delegated agent (claude-fable-5-1, inherited by the general-purpose subagent) received only the path of the self-contained HTML export of stocks/ai-citations (tip 84c0854, 70 entities, 359354 bytes, sha256 64ca5dbfb19f2fcbaeea72d8bccb38510fc703fdc927ba94573e25faf7c6da2e) and the charge to use the Read tool on that file alone: no shell, no web, no MCP, no other file.

Start time: spawned 2026-09-23T02:50:03+0200. Time it named the card: 2026-09-23T02:51:30+0200 (spawn plus the harness-reported duration of 86620 ms; the report was in the session at 02:51:38). Difference: 1 minute 27 seconds, under fifteen minutes. Cost: 151349 tokens, 8 tool uses (7 Read calls over lines 1 to 1700 of the file, one hand-back).

Entity named: stocks/ai-citations--with-a-url-checking-tool-in-the-loop-three-models-cut-non-resolving-citation-urls-6-to-79-fold-to-under-1-percent (title: With a URL checking tool in the loop three models cut non-resolving citation URLs 6 to 79 fold to under 1 percent).

Falls When targeted, quoted by the stranger and verified verbatim against the engine's entity read: "A recomputation from the released data in which DEAD plus LIKELY HALLUCINATED plus the dead share of UNKNOWN exceeds 1% for any of the three models (Table 3 already sums to 2.4% for GPT-5.1), or an independent browser-based liveness check of the final citations finds more than 1% not resolving. As a condition for the stock's question it falls if a support audit shows that tool-verified links resolve but support their sentences no better than before. Query to run: browser liveness check of the final URL sets, and per-citation support rating before versus after the urlhealth loop."

The stranger's attack, in its words: "The breaking point sets a 1% threshold that the card's own Collection section already exceeds for all three models, so no new data is needed: Table 3's DEAD plus LIKELY HALLUCINATED is printed as 2.4% (GPT-5.1), 0.7% (Gemini) and 0.5% (Claude), and the same card reports UNKNOWN at 10.3 to 20.2% of URLs with 11.0% [8.5, 13.7] of sampled UNKNOWNs genuinely dead, which adds roughly 1.1 to 2.2 points to every model. Summed, the lowest plausible not-resolving rates are about 3.5% for GPT-5.1, 1.8% for Gemini and 1.6% for Claude, each above the 1% the Falls When names and the 'under 1 percent' the title asserts. The headline multipliers collapse with it: GPT-5.1 goes from 16.0% to about 3.5 to 4.6% (roughly 4x, not 26x) and Gemini from 6.1% to about 1.8 to 2.9% (roughly 2 to 3x, not 79x). The card even concedes the prose rates 'do not equal DEAD plus LIKELY HALLUCINATED in Table 3' and that 'the text does not explain the difference', yet keeps the prose figures in the title and status open. I would therefore file a narrowing: the condition (a link checker in the loop cuts non-resolving links several-fold) stands, but 'under 1 percent' and '6 to 79 fold' should be replaced by the table-derived range, with the paper's own UNKNOWN audit as the arithmetic."

Runner-up the stranger noted and did not choose: the BBC card on Gemini sourcing errors (30 of 72 printed ratings is 41.7 percent, so 'over 45 percent' needs the seven don't-knows excluded).

Coverage: the stranger read lines 1 to 1700 of the HTML (header, index, all anchor cards, six of the derivations) and not the remaining derivations, the question or the tradeoff.

Source: the planning bundle's notes and the check ledger; nothing on this page was written for the page.

How this is built

The stock lives in the Memstead engine as a typed graph: four card types, a controlled vocabulary of links, a schema that refuses a claim without a coordinate and an inference without a breaking point. Every write is a commit with the writing identity in its trailer; every check is an append-only ledger row; a proposal is a git fork merged per card.

This site is a static render of the engine's exports: the stock as JSON, the health readings, the ledger, the proposal record, and the acceptance test's transcripts from the planning bundle. A script pulls them; the site adds nothing. If a page and an export disagree, the export is right.

Cost is measured and recorded per session in the project's own memory: building 42 anchors, 18 derivations and one tradeoff took about 4.4 hours of orchestrating wall-clock (about two of them lost to stalled delegated agents) and about 2.55 million delegated tokens (author batches 739 thousand, quoted-span checkers 890 thousand, derivation walkers 472 thousand, cost side 170 thousand, tradeoff checks 163 thousand, grader 117 thousand); re-pinning 43 anchors with quoted spans over checker-fetched text took about 50 minutes and 320 thousand delegated tokens; the thinking attack that moved seven of eighteen inferences took 21 minutes of API time and 3.3 million tokens, 2.6 million of them cached reads, USD 1.42; the acceptance test took about three minutes of runs and 375 thousand tokens.

Failure is measured too. The build session's cost memo of plan 02 records the first-pass rates before the second check: 11 of 42 anchors and 9 of 18 derivations failed, were repaired and re-checked. The question page shows a reading derived from the ledger over the stock as it stands today, which counts differently after renames and merged proposals: 10 of 46 anchors and 9 of 22 derivations failed their first second-identity check, none open. Both figures stand, each with its source.

One person orchestrates and decides; agents write. The imprint names the person.