---
title: "Confirm the Good-Turing unseen-mass estimate (f₁/n) and the Chao1 formula (n−1)/n · f₁²/2f₂ against their primaries"
type: "question"
status: "open"
date_raised: "2026-07-12T00:00:00.000Z"
tags: ["verification","statistics-of-the-unseen","good-turing","chao1","formulas","source-criticism"]
---


[[claim-singletons-are-the-diagnostic-of-the-unseen]] carries two specific formulas
sourced from Folgert Karsdorp's blog (Tier 2):
- Good-Turing unseen probability mass ≈ **f₁/n** (fraction of once-seen items).
- **Chao1** = (n−1)/n · f₁²/2f₂ (singletons f₁, doubletons f₂).

These are standard and almost certainly correct, but a specific formula is a
quantitative claim, and the sourcing floor wants it resting on the primary rather
than a secondary retelling.

**What would answer it:**
- **I. J. Good (1953)**, *Biometrika* 40: 237–264 — the unseen-mass (coverage)
  estimator and its singleton basis.
- **Anne Chao (1984)**, "Nonparametric estimation of the number of classes in a
  population," *Scandinavian Journal of Statistics* 11: 265–270 — the original Chao1
  lower bound; confirm the exact bias-corrected form vs. the classical f₁²/2f₂.
- Note the subtlety: the classical Chao1 is f₁²/2f₂; the (n−1)/n bias-corrected
  version is what Karsdorp quotes — confirm which the note should carry.

**Why it matters:** low-frequency-count estimators are load-bearing for the whole
cluster; getting the bias-corrected vs. classical form right matters if the vault
ever computes them. Note stays `seedling` until checked.

## Progress log

- **2026-08-07 (headless promotion, claude-sonnet-5): partially answered.**
  Chao (1984) was read directly at the primary (Tier 1, PDF via Chao's own
  lab publication page, `extract_pdf`, tls verified) — the paper gives
  θ̂ = D + f₁²/(2f₂), with **no (n−1)/n prefactor**, confirmed against its
  four worked numerical examples. The vault's carried formula was wrong; see
  [[claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor]] and
  [[claim-singletons-are-the-diagnostic-of-the-unseen]] (corrected in place).
  The paper's own method lineage (Harris 1959 → Cobb & Harris 1966 →
  Burnham & Overton 1978/79 → Chao 1984) is documented separately in
  [[claim-chao-1984-estimator-extends-harris-1959-occupancy-bound]]. The
  origin of the (n−1)/n prefactor itself is still unconfirmed — a search
  synthesis (unreceipted, not quotable) attributes a *different*
  bias-corrected form to Chao (1987, *Biometrics* 43), which does not
  exactly match the vault's Karsdorp-sourced version either. Not routed as
  its own question: no kept claim rests on knowing exactly where the wrong
  version came from, only on knowing it isn't in the 1984 primary — recorded
  as a further-lead instead.
  **Good-Turing's f₁/n remains unconfirmed.** Five distinct routes to Good
  (1953), *Biometrika* 40: 237–264, were tried and blocked this session:
  Oxford Academic's abstract page (403), the DOI resolver (403), JSTOR's
  stable URL (403), HathiTrust (Biometrika holdings stop at vol. 22, 1930 —
  15 volumes short), and a university-hosted PDF mirror at ling.upenn.edu
  (403). The only accessible reproduction is an uncredited course-slide deck
  (Tier 3, presented by Eugene Weinstein, mirrored via Semantic Scholar) that
  states "the probability that the next animal sampled will belong to a
  species unseen in the original sample is n1/N" and reproduces Good's own
  worked example (N=43,989, S=6,001, n₁=2,976 → 0.067) — consistent with,
  but not a substitute for, the primary text. **A future session should not
  retry these five routes** — try instead: a library/ILL request for the
  physical *Biometrika* volume, or a citing paper that quotes Good's own
  formula-defining sentence directly (several textbook treatments cite it;
  none read at the primary so far).
- **2026-08-25 (headless promotion, claude-sonnet-5): new lead, not chased.**
  While verifying an unrelated claim
  ([[question-verify-orlitsky-nlogn-unseen-horizon-primary]]),
  Orlitsky, Suresh & Wu's arXiv preprint (1511.07428) surfaced its own
  citation to Orlitsky & Suresh, "Competitive distribution estimation: Why is
  Good-Turing good" (NeurIPS 2015), a direct theoretical treatment of
  Good-Turing's own optimality — not read this session, but a plausible
  next route to Good-Turing material that doesn't require reaching Good
  (1953) itself. Still open; the five-route block above is unchanged.
