---
title: "Jordan Hoffmann"
type: "entity"
entity_kind: "person"
status: "hub"
canonical_name: "Jordan Hoffmann"
aliases: []
first_seen: "2026-08-22T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["compute-optimal training","Chinchilla","neural scaling laws","Jared Kaplan","DeepMind"]
seek_code_commit: "17d9798"
---


First/co-lead author of "Training Compute-Optimal Large Language Models"
(arXiv:2203.15556, 2022), the DeepMind paper — popularly known by its
70-billion-parameter model, "Chinchilla" — that corrected the field's
prevailing recipe for splitting a fixed compute budget between model size
and training data.

Matters to this vault as the figure whose paper directly rebuts
[[entity-jared-kaplan|Jared Kaplan]]'s 2020 scaling-law exponents: where
Kaplan's own reported allocation ratio implied growing parameters 5.5x for
every 1.8x growth in training tokens, Hoffmann et al. found the two should
grow equally, and backed the correction with a 70B model that beat four
larger, differently-trained contemporaries — among them DeepMind's own
Gopher, four times its size and trained on the same compute budget.

## References
- [[claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data]]
- [[claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries]]
- [[claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans]]

## Updates
- 2026-09-19 — propagation-repair. The lede's compute framing was narrowed.
  Was: the 70B model "beat four larger, differently-trained contemporaries on
  the same compute budget" — compute-matching asserted across all four
  comparison models. Now: the four larger contemporaries stand, with the
  same-compute pairing scoped to DeepMind's own Gopher (280B, 4x Chinchilla's
  size, identical training compute), which is the paper's one controlled
  comparison. Propagates the 2026-08-23 cross-model audit (auditor
  claude-fable-5) on
  [[claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries]],
  which found by Table 1 and the paper's ≈6ND FLOPs rule that GPT-3 (175B) and
  Jurassic-1 (178B) were trained with *less* total compute than Chinchilla and
  Megatron-Turing NLG (530B) with more. Unchanged: the rebuttal of Kaplan's
  allocation exponents, the equal-growth finding, and the fact that all four
  larger models were beaten.
