---
title: "The count of things seen exactly once — ecology's 'singletons', linguistics' 'hapax legomena' — is the diagnostic of the unseen: Good-Turing sets the unseen mass at roughly f₁/n"
type: "claim"
status: "seedling"
audit_status: "flagged (unverified-quant — the Good-Turing unseen-mass estimate f₁/n and the Chao1 richness formula (n−1)/n · f₁²/2f₂ are carried here from Karsdorp's blog (Tier 2). These are standard textbook formulas, but a specific formula is a quantitative claim whose primary homes are Good 1953 (Biometrika) and Chao 1984 (Scand. J. Statistics), not a blog. Routed to [[question-verify-good-turing-chao1-formulas-primary]]) | 2026-08-07 (headless promotion, claude-sonnet-5): partially resolved. Chao's own 1984 primary was read directly (Tier 1) and does NOT carry the (n−1)/n prefactor this note recorded — see [[claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor]]; the body below is corrected accordingly. Good-Turing's f₁/n remains unconfirmed against Good (1953) after a five-route access attempt (OUP, DOI, JSTOR, HathiTrust, one university mirror), all blocked — still `[unverified-quant — needs primary]`, question stays open."
writer_model: "claude-opus-4-8"
source_url: "https://www.karsdorp.io/posts/20220309103709-good_turing_as_an_unseen_species_model/"
source_title: "Demystifying Chao1 with Good-Turing"
source_author: "Folgert Karsdorp, 'Good-Turing as an unseen-species model'"
source_date: "2022-03-09T00:00:00.000Z"
source_quote: "Chao1 = (n−1)/n · f₁²/2f₂"
source_tier: 2
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md, 2026-07-12 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["statistics-of-the-unseen","good-turing","singletons","hapax-legomena","chao1","cross-domain-bridge"]
seek_code_commit: "89bc9f4"
---


Across the fields that share the unseen-species problem
([[claim-unseen-species-problem-has-dual-origin-ecology-and-cryptanalysis]]), the
quantity that tells you how much you have *not* seen is the number of things you
have seen **exactly once**. Ecology calls these **singletons**; linguistics calls
them **hapax legomena** — words appearing once in a corpus. The same statistic,
two field-names, is the observable proxy for the invisible tail.

Two estimators formalize it. Good-Turing sets the total probability mass of
never-seen items at approximately **f₁/n** — the fraction of the sample made of
once-seen items — the intuition being that if many things have shown up only once,
many more are still waiting to show up at all. The **Chao1** lower-bound estimator
of total richness uses singletons *and* doubletons: Chao1 = D + f₁²/(2f₂),
where D is the observed class count, f₁ the count of singletons, and f₂ of
doubletons — corrected 2026-08-07: a direct Tier-1 read of Chao's own 1984
paper found no (n−1)/n prefactor on this formula, contrary to what this note
previously carried from a Tier-2 blog (see
[[claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor]]). When
there are no doubletons the community is well-sampled; a large
singleton-to-doubleton ratio signals a large hidden tail.

This diagnostic is what powers the sample-coverage machinery in
[[claim-paleobiology-reinvented-coverage-based-rarefaction-as-quorum-subsampling]]:
coverage is estimated from exactly these low-frequency counts. It is also the hook
that connects the problem to authorship attribution — Efron and Thisted used the
same once-seen statistics to estimate the words Shakespeare knew but never wrote —
which sits near the vault's estimation cluster
([[claim-james-stein-estimator-uniformly-dominates-the-sample-mean]],
[[claim-efron-baseball-shrinkage-halved-batting-average-prediction-error]]).

**`[unverified-quant — needs primary]`, partially resolved 2026-08-07.** The
Chao1 formula has now been checked directly against its 1984 primary and
corrected — see
[[claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor]] and,
for the estimator's own lineage, [[claim-chao-1984-estimator-extends-harris-1959-occupancy-bound]].
Good-Turing's f₁/n is still unconfirmed against Good (1953) — a genuine,
five-route attempt to reach the primary text this session (OUP abstract, DOI
resolver, JSTOR, HathiTrust, one university mirror) was blocked at every
route; only a secondary course-slide summary reproduces the formula and its
worked example. Verification of that half stays routed to
[[question-verify-good-turing-chao1-formulas-primary]], which stays `open`.
The note stays `seedling` — one of its two flagged legs is still unverified.
