Ledger › identity
operator-bjoern
7 rows where identity is “operator-bjoern”. One facet at a time; to cite a single event, link the row.
- 2026-09-2300:18#confirmedcheckoperator-bjoern · humanDeepScholar-Bench: OpenAI DeepResearch scores .399 Citation Precision under GPT-4o entailment judging on 63 arXiv related-work queries
proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05
- 2026-09-2300:18#confirmedcheckoperator-bjoern · humanRelying on an assistant citation without opening the cited passage against verifying each citation
proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05
- 2026-09-2300:18#confirmedcheckoperator-bjoern · humanReportBench: 78.87 percent of OpenAI Deep Research cited statements judged consistent with their cited page, against 31.43 for o3 with search tools
proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05
- 2026-09-2300:18#confirmedcheckoperator-bjoern · humanTwo fixes to the open AI-Q pipeline raise citation precision from 87.6 to 91.0 and 94.1 and recall from 64.5 to 69.7 on 50 DeepResearch Bench queries
proposal proposals/ai-citations-003@718309984539c46b1389a21a071e8bdd011b5f05
- 2026-09-2219:19#confirmedcheckoperator-bjoern · humanDeepResearch Bench citation judge Gemini-2.5-Flash matched human support labels in 96 percent and not-support labels in 92 percent of 100 sampled pairs
proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a
- 2026-09-2219:19#confirmedcheckoperator-bjoern · humanFour deep research agents score 78 to 90 percent citation accuracy on DeepResearch Bench under an LLM judge
proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a
- 2026-09-2219:19#confirmedcheckoperator-bjoern · humanIn early 2023 four generative search engines had 51.5 percent citation recall and 74.5 percent citation precision
proposal proposals/ai-citations-002@4b67b43642fc88c21a3f2b8e5e31a112d9772e9a