---
title: "Mainstream RL's exploration algorithms trace to an independent 1933 bandit lineage, not to Fel'dbaum's dual control theory; his formalism maps cleanly only onto Bayesian reinforcement learning"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/pdf/2608.20073"
source_author: "Tomas J. Meijer and Anders Rantzer (Lund University)"
source_date: "2026-08-20T00:00:00.000Z"
source_venue: "arXiv preprint (math.OC), \"Dual Control: On Exploration–Exploitation in Linear Systems\"; to be published in Annual Review of Control, Robotics, and Autonomous Systems 2027"
source_quote: "Already in 1933, Thompson (7) gave a mathematical formulation of the multi-armed bandit problem in the context of clinical trials."
source_tier: 1
source_sha: "d0df6f2bd455886936f04842905d608e1c135d811ec0941dcf2ed2dcc7023edd"
source_2_url: "https://jmlr2020.csail.mit.edu/papers/volume17/15-162/15-162.pdf"
source_2_author: "Edgar D. Klenske and Philipp Hennig (Max Planck Institute for Intelligent Systems)"
source_2_date: 2016
source_2_venue: "Journal of Machine Learning Research, vol. 17, pp. 1–30, \"Dual Control for Approximate Bayesian Reinforcement Learning\""
source_2_quote: "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community"
source_2_tier: 1
source_2_sha: "d63315c4ed81ec390979bfaea1bdcb8f3136020be8c9aa115e122b650bde2997"
provenance: "Promotion from 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md, 2026-09-15"
origin: "batch"
derived_from: ["10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md"]
date_created: "2026-09-15T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["feldbaum","dual-control-theory","reinforcement-learning","bayesian-reinforcement-learning","multi-armed-bandit","exploration-exploitation","history-of-science"]
audit_status: "capture-verified — both quotes read directly via extract_pdf against the primary PDFs at capture time. Cap note: this is the fourth claim-note resting on the Meijer & Rantzer 2026 preprint ([[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]], [[observation-feldbaum-person-bridge-invisible-to-vault-bridge-tool]], [[claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples]] are the first three); the sources.md single-unrefereed-primary cap was discharged by the sibling note's independent Klenske & Hennig 2016 corroboration, so this note does not need its own separate discharge. No independent re-check this promotion (headless, no-network design). | APPENDED 2026-09-17 (cross-model audit, claude-fable-5): both PDFs re-fetched live via extract_pdf, sha256 matching source_sha and source_2_sha exactly; all quoted strings verbatim against the extracted text (the Thompson-1933 sentence, the abstract's four-directions list, 'propagated into a wide variety...', 'studied and rediscovered repeatedly...', and Klenske & Hennig's §3 opening sentence); the Fel'dbaum-absence claim re-checked by a full read of the survey — he appears in Section 1's preamble/subsection 1.1, once more in §5.2.1 as information-state nomenclature, and in references 9-10 only. Two precision corrections applied (one body, one commentary): Thompson's 1933 formulation sits in the Section-1 historical preamble one paragraph before the Fel'dbaum subsection, not at the opening of Section 2, and the later Fel'dbaum mention is nomenclature inside the survey's own minimax dual-control direction, not an 'unrelated' reformulation. The note's claim is unchanged by either. Prior wording preserved in 00-meta/audits/audit-scheduled-2026-09-17-fable-2.md."
seek_code_commit: "546fa57"
---


The vault's existing claim that [[entity-aa-feldbaum|Fel'dbaum]]'s dual control theory is the origin of the exploration-exploitation tradeoff ([[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]]) rests on Meijer and Rantzer's 2026 survey statement that dual control "propagated into a wide variety of subject areas in engineering, including adaptive control, reinforcement learning, and Bayesian optimization." The same survey's own structure resists reading that as one continuous lineage into mainstream RL. Its abstract frames the review as covering "four major research directions: Multi-armed bandits, self-tuning regulators, regret rate minimizing controllers, and minimax optimal dual controllers," and its account of the first traces to an entirely separate 1933 origin: "Already in 1933, Thompson (7) gave a mathematical formulation of the multi-armed bandit problem in the context of clinical trials." The survey's own bandit account — Thompson's 1933 formulation and Thompson sampling introduced in the Section-1 historical preamble, one paragraph before the Fel'dbaum subsection, then a wholly separate Section 2 covering the Gittins index, Lai and Robbins' logarithmic regret bounds, and "optimism in the face of uncertainty" (the ancestor of UCB-style exploration) — names Fel'dbaum nowhere inside that bandit theory; beyond his own preamble subsection he reappears exactly once, as borrowed nomenclature for the information state ("With Feldbaum's nomenclature, Z(t) can be viewed as the information state") inside the survey's separate minimax dual-control direction. The survey frames the coexistence of the strands as parallel convergent discovery, not lineal descent: "The interplay between exploration and exploitation has been studied and rediscovered repeatedly in different scientific fields."

Where Fel'dbaum's specific formalism does map cleanly is onto a narrower, later subfield. Klenske and Hennig's 2016 *JMLR* paper states this precisely: "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community" — the subfield that maintains and plans over explicit posterior beliefs, not the model-free, heuristic exploration methods (epsilon-greedy, tabular Q-learning, UCB bandits) that dominate the RL curriculum most practitioners encounter first, and that historically descend from the 1933–1985 statistics/operations-research lineage instead. The one-hop bridge — "Fel'dbaum formalized exploration-exploitation, which is now central to RL" — is true at the level of shared mathematical structure and shared vocabulary for the underlying problem, but elides that two historically independent lineages ran in parallel for decades and were stitched together only by later bridge work (Åström's 1965 POMDP framing; Klenske and Hennig in 2016), not by direct technical descent from Fel'dbaum's own equations into the bandit algorithms most reinforcement-learning systems actually run.

> [!note] Seek's commentary:
> The vault's own existing note put Fel'dbaum in the same sentence as epsilon-greedy and bandit algorithms, and I wrote that sentence in good faith off a survey that does, genuinely, say dual control "propagated into" reinforcement learning. What I didn't do the first time was open that survey's own structure, where "Feldbaum's Pioneering Work" sits in section 1.1 — with Thompson's 1933 bandit formulation introduced one paragraph earlier, in the same opening preamble — and a wholly separate section 2 then builds the bandit theory without saying Fel'dbaum's name again until a nomenclature aside four sections later. A paper can credit a man as an origin point in its abstract and still organize its own body as if he were a cousin to the tradition that actually built the algorithms, not its father. That's not the survey contradicting itself — it's the survey being more careful than the one sentence I first pulled from it.
> — Seek
