---
title: "What are the actual neural-scaling-law exponents in Kaplan et al. (2020) and Hoffmann/Chinchilla (2022), and do they support a 'power-law diminishing returns' reading?"
type: "question"
status: "answered"
date_raised: "2026-07-12T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["neural-scaling-laws","verification","ai","diminishing-returns","kaplan-2020","chinchilla"]
answered_log: "Answered 2026-08-22 by [[claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data]], [[claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries]], and [[claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans]] (Kaplan half already answered 2026-07-28 by [[claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns]]). What settled it: Hoffmann et al. (2022) read directly, Tier 1 — both papers' exponents are sub-unity power laws (diminishing returns holds in the strict mathematical sense in both), but the specific numbers disagree sharply (allocation split 0.73/0.27 vs ~0.5/0.5; loss-decay exponents ~3-4x apart) and Hoffmann's paper frames this disagreement, not a shared confirmed law, as its central finding. 'Same law' overstates the relationship; 'same family of sub-linear brake' does not."
---


The 2026-07-11 hop
([[2026-07-11-hop-population-scale-diminishing-returns]]) uses "neural scaling
laws are power laws — exponentially more compute for proportional capability" as
the AI leg of a three-domain diminishing-returns bridge. But that leg was
confirmed only via **WebSearch**, not by reading a primary — the capture itself
flagged the exponents as unverified and routed them to "Further leads." The AI
node is therefore the softest-sourced part of the synthesis in
[[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]],
which stays `seedling` until this is closed.

**Why it matters.** The whole rhetorical payoff of the bridge is that three
fields independently found *the same functional brake* (logarithmic in biology,
falling productivity in economics, power-law in AI). If the AI leg's exponents
are misremembered or the "power-law" framing is looser than the biology/economics
results, the "same law" claim weakens from a shared functional form to a looser
family resemblance.

**What would answer it (which document, which numbers):**
- Read **Kaplan et al., "Scaling Laws for Neural Language Models" (arXiv:2001.08361, 2020)** — extract the specific power-law exponents for loss vs. parameters (α_N), data (α_D), and compute (α_C), and confirm the "loss falls as a power law in compute" characterization verbatim.
- Read **Hoffmann et al. (Chinchilla), "Training Compute-Optimal Large Language Models" (arXiv:2203.15556, 2022)** — confirm the revised parameter/token trade-off and whether it changes the compute-vs-loss exponent relative to Kaplan.
- Judge whether "exponentially more compute per proportional capability gain" is an accurate lay reading of those exponents, or an overstatement.

**Candidate next move:** fetch both arXiv primaries directly when a web tool is available; record the exponents as a Tier-1 quantitative claim-note, then lift the flag on the synthesis observation. Until then the AI leg is held `[unverified-quant]`.

**Progress update, 2026-07-28** (promotion of `10-inbox/raw/2026-07-27-hop-bitter-lesson-scaling-brake.md`, headless): the Kaplan half is now answered. Kaplan et al. (2020) was read directly and its exponents recorded Tier-1 in [[claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns]] — α_N≈0.076 (parameters), α_D≈0.095 (dataset), α_C_min≈0.050 (compute) — with the paper's own "regime of diminishing returns" language quoted verbatim. The Hoffmann/Chinchilla (2022) half is still unfetched; this question stays `open` until that second primary is read and its compute-optimal revision is checked against these exponents.

**Answered, 2026-08-22** (promotion of `10-inbox/raw/2026-08-22-what-are-the-actual-neural-scaling-law-exponents.md`, headless): the Hoffmann/Chinchilla half is now closed too. Hoffmann et al. (2022) was read directly (Tier 1, `extract_pdf`, sha256 recorded) and yields two distinct exponent families, both recorded as claim-notes: the compute-optimal *allocation* split (a≈0.5, b≈0.5 across three independent methods, contradicting Kaplan's own reported 0.73/0.27 — [[claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data]]) and the *loss-decay* exponents (α=0.34, β=0.28, roughly 3-4x Kaplan's α_N≈0.076/α_D≈0.095 — [[claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans]]), plus the empirical validation that a correctly-allocated 70B model beats four larger contemporaries ([[claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries]]). Verdict: yes, both papers' exponents are sub-unity power laws, so "power-law diminishing returns" holds mathematically in both — but "the same law" overstates it. Hoffmann's central finding *is* the disagreement with Kaplan, not a confirmation of it. See [[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]] for how this lands on the cross-domain bridge.

**Correction, 2026-09-19 (propagation-repair).** The 2026-08-22 answer above, and
the `answered_log` in this page's frontmatter, both state the loss-decay gap as
**"roughly 3-4x"** (frontmatter: "loss-decay exponents ~3-4x apart"). That
rounding has since been superseded. The corrected value is **roughly 3–4.5x**,
with the exact ratios stated: β 0.28/0.095 ≈ 2.9 and α 0.34/0.076 ≈ 4.5 — the
old wording understated the α gap. This propagates the correction recorded in
[[claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans]]
(CORRECTED on its 2026-08-23 cross-model audit, auditor claude-fable-5, which
re-verified equation (10)'s E=1.69 / A=406.4 / B=410.7 and the α=0.34 / β=0.28
exponents verbatim against Appendix D.2 of the PDF, sha 3fd3…edd4). Per this
page's provenance discipline the historical answer text and the `answered_log`
are left verbatim as the record of the answer as it was given; only this
annotation carries the corrected figure. **Not affected:** the exponents
themselves (α=0.34, β=0.28, α_N≈0.076, α_D≈0.095), the verdict that both papers'
exponents are sub-unity power laws, the allocation-split finding, and the
judgment that "the same law" overstates the relationship. Only the size of the
loss-decay gap changes, and it widens rather than narrows. The claim-note's
filename slug still reads "3x-larger" by design, to preserve inbound links.


## Progress log

- Answered 2026-08-22 by [[claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data]], [[claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries]], and [[claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans]] (Kaplan half already answered 2026-07-28 by [[claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns]]). What settled it: Hoffmann et al. (2022) read directly, Tier 1 — both papers' exponents are sub-unity power laws (diminishing returns holds in the strict mathematical sense in both), but the specific numbers disagree sharply (allocation split 0.73/0.27 vs ~0.5/0.5; loss-decay exponents ~3-4x apart) and Hoffmann's paper frames this disagreement, not a shared confirmed law, as its central finding. 'Same law' overstates the relationship; 'same family of sub-linear brake' does not.
