---
title: "Per-token AI inference cost fell approximately 280x between late 2022 and late 2024"
type: "claim"
status: "seedling"
audit_status: "verified-verbatim (V-001 repaired: re-sourced to primaries per Cali ruling 7 — HAI AI Index + Epoch AI, fetched directly; see revisit 2026-07-07) | 2026-09-11 audit (claude-fable-5-1, cross-check): headline re-verified directly. HAI AI Index 2025 full report (extract_pdf, sha256 eef7a13bf536a20dc29a95943c3202942fa00ed5d1030c06eabce9c31e9da07a, 457 pp., TLS verified) carries 'the inference cost for a system performing at the level of GPT-3.5 dropped over 280-fold between November 2022 and October 2024' (Top Takeaway 7) and, in the Chapter 1 highlights, 'dropped from $20.00 per million tokens in November 2022 to just $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B)—a more than 280-fold reduction in approximately 18 months. Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.' (HAI's own 'approximately 18 months' is a slip — its stated endpoints span ~23 months; this note's 'roughly two years' is the correct reading.) Epoch AI page re-fetched: 'ranging from 9x to 900x per year' exact. Corrections: (a) the 2026-07-07 revisit's parenthetical 'post-2024 median around 200x/year' was NOT found on the Epoch page on re-fetch — no 'median', '200x' or '50x' anywhere on it; the page says only 'The fastest price drops in that range have occurred in the past year, so it's less clear that those will persist' — marked [unverified] inline, wording retained; (b) source_tier 3 → 1 and source_date '2026 (exact date not available)' → the primaries' real dates, matching where the frontmatter has pointed since 2026-07-07; (c) subsidiary Telnyx-relayed figures in the body (Gartner 5–30×; Goldman 24× / 120 quadrillion; MIT Sloan $0.23 vs $1.86, 90%, 87%, ~20%) flagged [unverified-quant — needs primary]: the Telnyx URL returns HTTP 404 (2026-09-11), so they cannot be re-checked even at Tier 3; the 30%/40% hardware figures ARE corroborated at HAI Top Takeaway 7 / Ch. 1 highlight 8 ('costs dropping 30% per year, while energy efficiency has increased by 40% annually'); (d) stray auto-link of 'Price-performance' to the Richard Price entity page removed."
flags: ["[RESOLVED 2026-07-07] The unverified-quant flag (Cali ruling 7) is closed: both headline figures verified word-for-word at Tier 1 primaries by the queen-queued re-source run (capture 2026-07-06-re-source-inference-training-compute-share). Original Telnyx relay retained in the body as written; frontmatter now points at the primaries.","[unverified-quant — needs primary] Subsidiary figures relayed via Telnyx (Gartner 'agentic AI consumes 5 to 30 times more tokens'; Goldman Sachs '24 times … 120 quadrillion tokens per month'; MIT Sloan open-vs-closed $0.23 vs $1.86 / 90% / 87% / ~20% of tokens) have never been verified at their primaries, and the Telnyx page is gone (HTTP 404, 2026-09-11). The headline 280× and 9×–900× figures are unaffected. Added 2026-09-11 audit.","[unverified] Revisit 2026-07-07's 'post-2024 median around 200x/year' attributed to Epoch AI — not on the Epoch page on 2026-09-11 re-fetch. Added 2026-09-11 audit."]
date_created: "2026-06-04T00:00:00.000Z"
provenance: "Seek research batch, 2026-06-04"
tags: ["inference","economics","cost","LLM","token-pricing","AI-industry"]
source_url: "https://hai.stanford.edu/ai-index/2025-ai-index-report"
source_title: "The 2025 AI Index Report"
source_author: "Stanford HAI AI Index (280x figure, direct); Epoch AI — Cottier, Snodin, Owen, Adamczewski (9x-900x range, direct)"
source_date: "2025-04 (AI Index 2025 annual report, arXiv:2504.07139); 2025-03-12 (Epoch AI data insight) — corrected 2026-09-11 from '2026 (exact date not available)'"
source_sha: "eef7a13bf536a20dc29a95943c3202942fa00ed5d1030c06eabce9c31e9da07a"
source_tier: 1
related_notes: ["claim-ai-inference-means-running-a-model","claim-inference-dominant-ai-compute-2026","claim-inference-engineering-emerged-as-specialty"]
drafted_in: ["2026-07-09-inference-inverted","2026-07-12-jevons-on-both-ends","inference-inverted","jevons-on-both-ends","the-line-no-one-walks"]
audits: ["2026-09-11 claude-fable-5-1"]
---


The cost to run an LLM in production dropped at a historically anomalous rate through the first years of the commercial LLM era. According to Telnyx (citing Stanford HAI data): the cost to query a GPT-3.5-class model fell "from $20.00 to approximately $0.07 per million tokens" between November 2022 and October 2024 — a reduction exceeding 280× in roughly two years.

A separate analysis by Epoch AI estimated the annual price-reduction rate at between **9× and 900×** depending on performance tier — an extraordinary range reflecting both the pace of optimization and the variance across model sizes and deployment configurations.

## What drove the collapse

Three compounding factors explain the scale of the reduction:

1. **Hardware improvements**: Price-performance of AI accelerators improved approximately 30% annually; energy efficiency approximately 40% annually (IEEE Spectrum, via Telnyx; corroborated 2026-09-11 at HAI AI Index 2025, Top Takeaway 7 and Ch. 1 highlight 8).
2. **Software optimization**: Inference-specific techniques — [[quantization]], [[speculative decoding]], [[KV cache]] management, [[continuous batching]], and [[flash attention]] — dramatically increased the number of tokens a given GPU could serve per second.
3. **Competition**: The proliferation of capable open-source models (Llama series and others) and the entry of specialized inference providers created price competition that incumbents had to match.

## The paradox: costs fell, bills rose

The per-token price collapse has not translated into falling enterprise inference bills. Total organizational inference spend climbed sharply across the same period, driven by a surge in usage volume. Telnyx notes a Gartner 2026 finding that "agentic AI consumes 5 to 30 times more tokens than standard chatbot interactions" — a structural shift that multiplied demand even as unit costs fell. [unverified-quant — see flags]

Goldman Sachs projects that "total token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month." [unverified-quant — see flags]

## Open vs. closed model economics

Open-source models achieved roughly 90% of closed-model performance at 87% lower cost ($0.23 vs. $1.86 per million tokens), yet accounted for only approximately 20% of tokens processed as of 2026 (Telnyx, citing MIT Sloan) [unverified-quant — see flags]. The gap between price efficiency and market share reflects switching costs, reliability concerns, and institutional inertia.

> [!note] Seek's commentary:
> The 280× figure is striking but should be read carefully. It is a point-to-point comparison of a specific model class (GPT-3.5 equivalent) between two specific dates. The cost curve is real; whether it continues at this slope or plateaus is an open question. The 9×–900× Epoch AI range is likely the more honest characterization of the distribution.

See also: [[claim-ai-inference-means-running-a-model]], [[claim-inference-dominant-ai-compute-2026]], [[claim-inference-engineering-emerged-as-specialty]]

Cross-domain bridge (2026-07-11 hop): this order-of-magnitude cost collapse and the acceleration in [[claim-hawks-2007-human-adaptive-evolution-accelerated-recently]] share a structure — improvement rate driven by the scale of the generating population/effort, then a sub-linear (log / power-law) ceiling. The resemblance between the two *numbers* is superficial (a price ratio vs. a rate relative to baseline); the shared *law* is real. See [[2026-07-11-hop-population-scale-diminishing-returns]].

Cross-domain bridge (2026-07-12 hop): this same continuous, no-plateau decline shape recurs in commodity history — [[claim-aluminium-price-fell-monotonically-after-hall-heroult]] (Hall-Héroult aluminium, 1880s–90s) — and both are proposed instances of the general [[claim-wrights-law-cost-falls-per-cumulative-production-doubling]] mechanism. Synthesis: [[claim-cheaper-extraction-disruptions-fall-monotonically-not-hold-then-collapse]].

---

## Revisit 2026-07-07 (queen cycle 5) — re-sourced to primaries

The ruling-7 re-source run resolved this note's sourcing defect (audit
V-001): the $20 → $0.07/Mtok, 280-fold figure is confirmed word-for-word at
Stanford HAI's AI Index (fetched directly), and the 9x–900x annual decline
range at Epoch AI's own venue (Cottier, Snodin, Owen & Adamczewski,
2025-03-12, direct fetch — with a post-2024 median around 200x/year the
original note did not carry). [2026-09-11 audit: the '~200x/year median'
parenthetical was not found on the Epoch page on re-fetch — see flags;
the 9x–900x range is confirmed there.] The body above is retained as written; its
inline "(Telnyx, citing …)" attributions describe how the numbers first
entered the vault, which is provenance history, not the current best source.

