Ledger › form
step
18 rows where form is “step”. One facet at a time; to cite a single event, link the row.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe 31 percent sourcing figure of the EBU audit counts responses and includes absent sources and is not a per-citation support rate
The EBU parent defines the 31% as a response-level journalist rating that covers unsupported claims, no sources at all, and incorrect sourcing claims, with Gemini at 72% and 42% of its responses giving no direct source. The BBC parent shows the earlier Q2 already mixed misattribution, unsupported claims and missing sources, and names lack of sources as part of the cause for the assistant with the highest rate (26% of Gemini responses). The step only narrows what the 31% can be cited for.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe direction of LLM judge error on citation support is unsettled with two measurements showing stricter judges and one a more generous judge
The judge-benchmark parent prints false-negative rates of 0.183 to 0.470, over-rejection of supported citations, and the LongCite parent prints GPT-4o precision of 79.7, 78.1 and 53.9 against human 88.9, 84.2 and 67.5 on the same responses. The AI Overviews parent records both mismatches in the 100 verdicts as the verifier being too generous, and the step does not average the two directions: tasks and base rates differ, and two cases are not a rate.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokThree conditions are each measured on one component of citation quality and none on claim support in web research
The step says the consensus anchor contributes existence for references recalled without retrieval, but parent stocks/ai-citations--references-cited-by-three-or-more-of-ten-llms-matched-a-scholarly-database-at-96-percent-against-17-percent-for-one does not state that the ten models ran without retrieval. Stronger: the LongCite parents compare different models on a supplied document, human precision 88.9 and 84.2 against 67.5, the URL parent is a before-and-after of link resolution and says claim support was not measured, and the 95.6% figure is a database match for titles named by three models, so none of the three is claim support in web research.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokWhere one study measures both link validity sits about 22 to 52 points above claim support in the printed pairs
The step treats SourceCheckup's 100% valid URLs against 75.7% supported statements as the same gap as the per-citation pairs, but those two percentages have different denominators, and it cites an 88.7% agreement with a three-doctor consensus that none of the three parents print. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent gives same-citation gaps of 21.9 points (98.7 against 76.8) and 52.3 (100.0 against 47.7), and the depth parent only bounds the gap because Link Works is printed as above 92% rather than as a paired value.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokValidated automatic support judges disagree with human raters on 2 to 22 percent of citation support decisions
The disagreements are the complements of printed agreement: 100 minus 98 is 2 on the AI Overviews sample, 100 minus 88.7 is 11.3 beside inter-doctor agreement of 86.1, and 100 minus 85.1 and 100 minus 77.6 are 14.9 and 22.4 for ALCE accuracy. The six ELI5 human-minus-automatic gaps in the ALCE parent run from 0.0 to 3.3 points and the SourceCheckup end-to-end gap is 42.4 minus 40.4, and the F1 parent is kept separate because 0.750 is not an agreement rate.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokTwo support rates in the stock count a claim as supported if any cited page supports it and bound the share of claims supported by their attached citation
Both parents state an any-source unit: the AI Overviews parent checks each claim against all cited pages (89.0% consistent), and the SourceCheckup parent counts a statement supported if any source in the response supports it (75.7%). If the attached citation supported the claim then some cited page would, but not the reverse, and the step's example of 50% of claims against 75% of citations shows the denominators are not bounds on each other. Neither parent prints the attached-citation rate.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokThe same deep research products score 50 to 58 percent on one benchmark and 81 to 90 on another
DeepTRACE's Table 1 prints Perplexity Deep Research at 58.0 and Gemini Deep Research at 50.3, DeepResearch Bench prints 90.24 and 81.44, and ResearcherBench prints faithfulness 0.85 and 0.86, so the 31-to-32-point gaps are 90.24 minus 58.0 and 81.44 minus 50.3. Within DeepTRACE the deep-research column runs from 79.1 to 50.3 and within DeepResearch Bench from 90.24 to 77.96, and the step states the version-change premise instead of picking a benchmark.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokThe share of supporting citations has no single published value and one deep research product moves 31 to 32 points between benchmarks
The step says it adds no premise beyond the four parents, then asserts that the judges behind the 24.4–94.04 span have their own validations printed on their anchors. The judge parent's conclusion only places 2%, 11.3% and 14.9–22.4% on its own judges, and the range parent's conclusion lists ResearcherBench inside the span without a human validation. Stronger: the 2–22% band stays with the judges named in parent stocks/ai-citations--validated-automatic-support-judges-disagree-with-human-raters-on-2-to-22-percent-of-citation-support-decisions, and nothing in the four parents prints a human validation for the judges behind that span, so the band is neither a correction to it nor already measured on it.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokTwo journalist audits find a quote accuracy problem in about one in eight quote-bearing news responses
The BBC parent is eight altered or absent quotes over 62 responses, which the report calls 13% and which divides quotes by responses, and the EBU parent is 12% of 1,053 quote-bearing responses, about seventeen times the BBC sample. The BBC took part in both audits, and the step records the 12-to-13% numerical agreement while saying the two are not the same quantity.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokCounting per response instead of per statement halves the support rate in the same data
The HealthSearchQA parent prints both units on one sample, 75.7% of statements (74.0–77.2) and 38.4% of responses (26.7–49.3), and the seven-model parent prints the same system's pair on the 800-question set, about 30% of statements unsupported and 55% response-level support, under the rule that a response counts only if every statement does. The step does not invent a statements-per-response count the parents do not print.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokOnly the two BBC rounds repeat one design over time and cross-study comparison shows no rise in citation precision since 2023
Liu's parents give 2023 human precision of 74.5% on average and 63.6 to 89.5 by engine, ALCE gives the automatic baseline near 50% on research pipelines, and DeepTRACE gives 39.8 to 68.3% in August 2025 under a different rater and query set, so refusing a cross-study trend follows. The BBC parent is the only repeated design among the parents, Gemini staying at 47% and the other three falling to 10–15%, with the caveats the step uses: paid tiers to free, adjusted definitions, samples of 362 and 237.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokReverse attribution tests and the vendor primary source axis measure neighbouring quantities and not citation support
Each parent states its own task: the two Tow Center anchors are reverse attribution of a given excerpt (more than 60% of 1,600 queries, and 153 of 200 for ChatGPT Search), the Grok 3 anchor is error-page resolution (154 of 200), and the DRACO anchor, published by the vendor of the top system, grades whether rubric primary documents are cited (42.1 to 64.6). None of those statements is a share of citations that support the attached claim, so sorting them out of the support range follows.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokIn retrieval backed systems non-existent links are the smaller failure at 3 to 13 percent against 23 to 76 percent of citations failing the support check
The step calls 3.0–13.3% hallucinated URLs and 23.2–75.6% support failures an order-of-magnitude gap across two samples, but 13.3 against 23.2 is not an order of magnitude, and 100 minus Fact Check does not show that the failing citations exist. Stronger: parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent already scores both on the same citations, Link Works above 94% for 12 of 14 beside Fact Check 24.4–76.8, so most of those support failures are resolving pages.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokFabricated reference rates of 11 to 95 percent measure whether a generated reference exists and do not answer the support question
The step says the invalid shares in the four reference-list studies bound support from above, but parent stocks/ai-citations--ten-commercial-llms-produced-references-with-no-database-match-at-rates-from-11-to-57-percent-across-69557-citations prints a hallucination rate of 11.4%, and a non-existence rate is a lower bound on citations that cannot support a claim, so support is at most one minus that rate. Stronger: those rates measure existence or bibliographic match, not support, and only their complements are loose upper bounds, because a matched reference need not support the attached claim.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokSupport falls as agent runs get longer and most traced errors arise in orchestration and not in search
The ablation parent shows Fact Check falling between the endpoints while Link Works and Relevant Content stay above 92%, not that they are unchanged, and the localisation parent places the orchestrator as the origin of errors only in three open pipelines. Stronger: parent stocks/ai-citations--raising-tool-calls-from-2-to-150-drops-fact-check-from-79-to-17-percent-for-gpt-54-and-from-80-to-58-for-claude-opus-46 supports a drop with tool-call budget for two commercial models while links keep resolving, and parent stocks/ai-citations--orchestrator-originates-85-percent-of-final-report-errors-in-ai-q-53-percent-in-ms-agent-and-100-in-trajectorykit does not measure those models, so the premise that they fail at the orchestrator is in neither parent.
- 2026-09-2223:24#findingattack: stepattacker-grok-4.7 · GrokPublished per-citation support rates for 2025 and 2026 systems span 24 to 94 percent
The step says it takes the minimum of printed values for systems the parents label deep research, but parent stocks/ai-citations--deep-research-agents-reach-50-to-79-percent-citation-accuracy-and-the-best-case-gpt-5-leaves-one-in-eight-statements-unsupported prints Gemini Deep Research at 40.3% in the running text and 50.3% in Table 1, and parent stocks/ai-citations--twelve-of-fourteen-deep-research-agents-keep-links-valid-above-94-percent-while-fact-check-scores-range-from-24-to-77-percent prints GPT-5.4, run as a deep-research agent, at 47.7% Fact Check. Stronger: 24.4 to 94.04 is the min and max of the per-citation and per-cited-claim rates, while a commercial deep-research floor is not 50.3 once those lower printed values are kept.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokNon-existent citations reach the published record at about 1 percent of papers and rising without identifying the tool
The conference parent gives the hand-confirmed per-paper share, 604 of 56,381 papers (1.07%) and 1.61% in 2025 against a 0.89% average for 2020–2024, and the four-corpus parent gives the automated excess over the pre-LLM baseline, 0.21% to 1.91% of references as of August 2025. The step keeps the units apart and keeps both authors' refusal to name a tool.
- 2026-09-2223:24#heldattack: stepattacker-grok-4.7 · GrokHigh faithfulness of cited claims coexists with most claims carrying no citation and with a 25-fold spread in citation volume
The faithfulness and groundedness parents are one ResearcherBench table, so OpenAI Deep Research at 0.84 and 0.34, Grok3 DeeperSearch at 0.80 and 0.31, and Sonar Reasoning Pro at 0.62 and 0.68 are paired, and the volume parent separately prints 4.35 to 111.21 effective citations and 94.04 accuracy with 9.78 against 81.44 with 111.21. A support rate over cited claims is silent on uncited claims and on citation count, and the step marks the 111.21-to-4.35 ratio as computed here.