---
title: "The canonical LLM self-knowledge benchmark is SelfAware (Yin et al. 2023) — 1,032 unanswerable vs 2,337 answerable questions; \"SKBench\" does not exist"
type: "claim"
status: "budding"
audit_status: "capture-verified (ACL Anthology, arXiv, and official repo fetched directly at capture level) | 2026-09-11 audit (claude-fable-5-1, cross-model lane; writer unknown): ACL Anthology page, arXiv 2305.18153 abstract, github.com/yinzhangyue/SelfAware README ('SelfAware includes 1032 unanswerable questions and 2337 answerable questions' EXACT; 'five diverse categories' EXACT in the abstract), arXiv 2207.05221 (Kadavath, P(IK)) and an arXiv API search for 'SKBench' (0 results) all re-checked directly. One stale sentence: 'Restructuring of the affected note awaits Cali's decision' — the question closed 2026-07-07 (Cali ruling 2, restructuring executed the same day); dated correction appended inline, original sentence kept. Claim unchanged."
source_url: "https://aclanthology.org/2023.findings-acl.551/"
source_title: "Do Large Language Models Know What They Don’t Know?"
source_author: "Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang"
source_date: "2023-07 (ACL Findings 2023)"
source_tier: 1
source_quote: "Do Large Language Models Know What They Don't Know?"
provenance: "Promotion from 10-inbox/raw/20260706-1058-what-real-benchmark-did-the-skbench-citation-mean...md, 2026-07-07, queen cycle 5 — answering 50-questions/question-recover-skbench-real-source.md (opened cycle 2, researched by the bee overnight)"
origin: "session"
date_created: "2026-07-07T00:00:00.000Z"
tags: ["self-knowledge","benchmark","selfaware","llm-calibration","citation-correction"]
drafted_in: ["2026-07-12-good-1952","good-1952","surveying-my-own-species"]
verified_verbatim: "2026-08-07 — source_quote matched verbatim (normalized) against a direct fetch of source_url by seek_verify (no model involved)"
seek_code_commit: "f424b5f"
audits: ["2026-09-11 claude-fable-5-1"]
---


The purpose-built benchmark the field treats as canonical for LLM
self-knowledge is **SelfAware** (Yin et al., "Do Large Language Models Know
What They Don't Know?", ACL 2023 Findings; arXiv:2305.18153): 1,032
unanswerable and 2,337 answerable questions across five categories, testing
whether models can identify what cannot be known (dataset composition
confirmed at the official repository). A distinct, earlier line is
Anthropic's P(IK) self-evaluation (Kadavath et al. 2022, arXiv:2207.05221) —
predicting the probability of knowing an answer rather than classifying
unanswerability. Adjacent narrower instruments: CalibratedMath (verbalized
confidence), SAPLMA (hidden-state truthfulness probes), BeHonest (knowledge
boundaries as one of three honesty dimensions).

**The negative finding is part of the claim:** an honest search found no
benchmark named "SKBench" — no paper, repo, erratum, or discussion thread.
The citation by that name in [[introspection-access-problem]] (audit V-010,
ruled a probable hallucination) has no recoverable referent, and its
specific findings (26 LLMs, 34% capability underestimation, 18%
tool-degradation) have no located home in any of the real benchmarks above.
The nearest string-collision is SKA-Bench (structured-knowledge, topically
unrelated). Restructuring of the affected note awaits Cali's decision per
[[question-recover-skbench-real-source]]. *(2026-09-11 audit: no longer
pending — that question records Cali's approval and the history-preserving
restructuring of [[introspection-access-problem]] as executed on
2026-07-07; the sentence above is kept as the note's state at promotion.)*
