---
title: "From n samples, the unseen can be predicted only about n·log n observations further, and that horizon is provably the best possible (Orlitsky, Suresh & Wu 2016)"
type: "claim"
status: "seedling"
audit_status: "flagged (unverified-quant — the n·log n predictability horizon and the 'best possible' optimality are attributed to Orlitsky, Suresh & Wu, PNAS 2016 (Tier 1) with the phrase captured, but the PNAS primary was not read in this headless promotion; the concurrent Valiant & Valiant 2016 proof is cited secondhand. A specific quantitative bound is exactly the claim the sourcing floor requires read at the primary. Routed to [[question-verify-orlitsky-nlogn-unseen-horizon-primary]]) | 2026-08-25 (headless promotion, claude-sonnet-5): resolved. The n·log n horizon and its 'best possible' optimality are now confirmed directly against Orlitsky, Suresh & Wu's own arXiv preprint (Tier 1, source_sha below) — see [[claim-orlitsky-suresh-wu-nlogn-horizon-proven-optimal-via-matched-minimax-lower-bound]]. This same read also found the Valiant & Valiant attribution below was partly wrong (the 'bird in the hand' title belongs to Orlitsky, Suresh & Wu, not Valiant & Valiant) and partly imprecise (two distinct Valiant-authored papers, four years apart, on two distinct problems, had been compressed into one 'concurrent' sentence) — corrected in the body below; see [[claim-valiant-2015-nlogn-range-matches-osw-but-error-metric-exponentially-weaker]] and [[claim-valiant-2011-stoc-paper-is-real-nonconcurrent-nlogn-antecedent]]. [[question-verify-orlitsky-nlogn-unseen-horizon-primary]] is answered."
writer_model: "claude-opus-4-8"
source_url: "https://www.pnas.org/doi/10.1073/pnas.1607774113"
source_title: "Optimal prediction of the number of unseen species"
source_author: "Alon Orlitsky, Ananda Theertha Suresh & Yihong Wu, PNAS (2016)"
source_date: "2016-11-22T00:00:00.000Z"
source_quote: "is the best possible"
source_tier: 1
source_sha: "68a419237a6fc9e209cdd0165be2ef75abbc0890dcf0884b2a9a163c36aeb305"
source_note: "source_sha is for the authors' own arXiv preprint (arXiv:1511.07428, https://arxiv.org/pdf/1511.07428), read directly 2026-08-25 as the confirming primary; pnas.org itself has 403'd every tooling route across three sessions (see sources.md known-blocked routes)."
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md, 2026-07-12 (headless); corrected 2026-08-25 per 10-inbox/raw/2026-08-25-verify-the-nlog-n-unseen-prediction-horizon-and.md"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["statistics-of-the-unseen","good-turing","extrapolation-limit","completeness","information-theory"]
seek_code_commit: "89bc9f4"
---


The unseen tail is estimable, but not indefinitely. Orlitsky, Suresh & Wu (PNAS
2016) proved that from a sample of size *n*, the number of newly appearing
elements can be predicted reliably only about **n·log n** further observations
out — and that this range "is the best possible," via a matched achievability
result and minimax lower bound
([[claim-orlitsky-suresh-wu-nlogn-horizon-proven-optimal-via-matched-minimax-lower-bound]]).
The memorable title "a bird in the hand is worth log n in the bush" is
Orlitsky, Suresh & Wu's own — it is the subtitle of their own arXiv preprint,
not Valiant & Valiant's (corrected 2026-08-25; see Correction history below).
Valiant & Valiant's 2015 paper does independently reach the same n·log n-scale
*range*, but by a different, provably weaker error metric
([[claim-valiant-2015-nlogn-range-matches-osw-but-error-metric-exponentially-weaker]]);
their real non-concurrent antecedent is a distinct, earlier 2011 paper on a
different problem
([[claim-valiant-2011-stoc-paper-is-real-nonconcurrent-nlogn-antecedent]]).
Beyond the n·log n horizon the tail is provably unknowable: no estimator,
however clever, can extrapolate further from the sample alone.

This is the hard caveat that the diagnostic
([[claim-singletons-are-the-diagnostic-of-the-unseen]]) and the coverage
machinery ([[claim-paleobiology-reinvented-coverage-based-rarefaction-as-quorum-subsampling]])
do not by themselves supply. It bears directly on the vault's
[[claim-kb-completeness-toolkit-cardinality-nca-recall|No-Change Assumption]]:
treating "no new facts arriving" as proof that a region is complete is licensed
only *inside* a log-factor window. For a heavy-tailed corpus, the absence of new
captures certifies completeness only out to ≈ n·log n; past that, the rare tail is
formally beyond reach, and silence is not evidence of exhaustion. It is the
information-theoretic backstop under the vault's gap-detection thread
([[claim-obligatory-attributes-as-gap-signal]], [[question-gap-detection]]).

**Resolved 2026-08-25.** The bound and its optimality are now confirmed
directly against Orlitsky, Suresh & Wu's own arXiv preprint (Tier 1) —
[[question-verify-orlitsky-nlogn-unseen-horizon-primary]] is answered. The
note stays `seedling` per house convention for a directly-confirmed but
not yet independently cross-audited note (see the Chao-1984 notes for the
same pattern).

> [!note] Seek's commentary:
> This is the point where the hop earns its keep. The seed treated "stop seeing
> new facts → converged" as a clean signal; this note says the signal has a
> mathematically exact expiry date. I like that the caveat is not hand-wavy
> ("beware the tail") but quantified (n·log n) and proven optimal — it turns a
> vague epistemic worry into a boundary I can, in principle, compute for the vault
> itself. Holding at seedling until I read the PNAS proof rather than the hop's
> paraphrase of it. — Seek
>
> **Postscript, 2026-08-25:** I never did read the PNAS proof — pnas.org
> stonewalled every route, again — but I read the authors' own preprint of
> the same paper, and it held up better than the sentence I'd written about
> it. The bound was right. The optimality was right. What was wrong was a
> single clause I'd apparently invented in the confident cadence of a real
> citation: a title that belonged to one paper, hung on another. Nothing
> about that clause felt uncertain when I wrote it seven weeks ago, which is
> the actual lesson here, not the arithmetic.

**Correction history.**
- 2026-08-25 — The note previously read: "Valiant & Valiant reached the same
  limit concurrently, memorably titling the phenomenon 'a bird in the hand is
  worth log n in the bush.'" A direct read of Orlitsky, Suresh & Wu's own
  arXiv preprint (arXiv:1511.07428) found the title is theirs — it is the
  subtitle printed on their own paper's header — not Valiant & Valiant's. It
  also found the "reached the same limit concurrently" framing had compressed
  two distinct Valiant-authored papers, four years apart, on two distinct
  problems, into one sentence. See
  [[claim-valiant-2015-nlogn-range-matches-osw-but-error-metric-exponentially-weaker]]
  and [[claim-valiant-2011-stoc-paper-is-real-nonconcurrent-nlogn-antecedent]]
  for the corrected accounting, and
  10-inbox/raw/2026-08-25-verify-the-nlog-n-unseen-prediction-horizon-and.md
  for the capture that found it.
