---
id: "20260921-0214-what-do-the-kan"
title: "What do the KAN paper's own unread references [9]-[16] actually say about why prior neural Kolmogorov-Arnold attempts stalled?"
type: "capture"
status: "promoted"
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-09-21T00:00:00.000Z"
provenance: "Batch research run, 2026-09-21. Hook from 70-drafts/a-proof-is-not-a-recipe/draft.md, itself following on from [[claim-kan-paper-prior-attempts-stalled-without-modern-tooling]] (2026-09-16), which recorded only the 2024 KAN paper's own one-sentence gloss on why prior neural Kolmogorov-Arnold attempts stalled. This capture goes to the KAN paper's own bibliography for refs [9]-[16] and, where they weren't reachable, to the primary papers those references themselves argue against or descend from."
derived_from: []
tags: ["kolmogorov-arnold","kan","neural-networks","history-of-science","approximation-theory","tooling-bottleneck","curse-of-dimensionality"]
source_url: "https://arxiv.org/pdf/2404.19756"
source_title: "KAN: Kolmogorov-Arnold Networks"
source_author: "Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljačić, Thomas Y. Hou, Max Tegmark"
source_date: "2024-04-30T00:00:00.000Z"
source_venue: "arXiv (cs.LG) 2404.19756, accepted ICLR 2025"
source_tier: 1
source_sha: "c04339e34ac3f4a8695a74c59e2a0a4f332cfdcb3831a12696d5192cec2ed713"
source_url_2: "http://cbcl.mit.edu/people/poggio/journals/girosi-poggio-NeuralComputation-1989.pdf"
source_title_2: "Representation Properties of Networks: Kolmogorov's Theorem Is Irrelevant"
source_author_2: "Federico Girosi, Tomaso Poggio"
source_date_2: 1989
source_venue_2: "Neural Computation 1(4), 465-469, MIT Press (author-hosted copy, MIT Center for Biological & Computational Learning)"
source_tier_2: 1
source_sha_2: "ce1c672e6c71acb24baab79fd2a2e8101618f76da7913d3904c14c86d14b645c"
source_url_3: "https://ins.uni-bonn.de/media/public/publication-media/remonkoe.pdf?pk=82"
source_title_3: "On a constructive proof of Kolmogorov's superposition theorem"
source_author_3: "Jürgen Braun, Michael Griebel"
source_date_3: 2009
source_venue_3: "Constructive Approximation 30(3), 653-675 (author-hosted preprint, Institute for Numerical Simulation, University of Bonn)"
source_tier_3: 1
source_sha_3: "3663d8dce21182ddd4d73e8eefba9dd3483aeda17dbff80162864573bdadb29e"
source_url_4: "https://arxiv.org/pdf/2311.00049"
source_title_4: "On the Kolmogorov neural networks"
source_author_4: "Aysu Ismayilova, Vugar E. Ismailov"
source_date_4: "2023-10-31T00:00:00.000Z"
source_venue_4: "arXiv (cs.NE) 2311.00049"
source_tier_4: 1
source_sha_4: "7e72b508c2f724161c525af2ff718a0ebdb0efa344cf8d07dfdafb0c40132b72"
source_url_5: "https://cs.uwaterloo.ca/~y328yu/classics/Hecht-Nielsen.pdf"
source_title_5: "Kolmogorov's Mapping Neural Network Existence Theorem"
source_author_5: "Robert Hecht-Nielsen"
source_date_5: 1987
source_venue_5: "Proceedings of the IEEE First International Conference on Neural Networks, San Diego, Vol. III, pp. 11-13 (scanned course-page mirror; original conference proceedings not found online)"
source_tier_5: 1
source_sha_5: "447be0516246342a866d6c001336cb91fdbb64543deede44f46d37038d9d5156"
seek_code_commit: "unknown"
promoted_to: ["30-notes/claim-kan-paper-9-16-citation-cluster-spans-1993-2023-not-1980s-90s.md","30-notes/claim-girosi-poggio-1989-kolmogorov-critique-is-mathematical-not-tooling.md","30-notes/claim-kurkova-1991-rebuttal-changed-mathematical-target-not-tooling.md","30-notes/claim-sprecher-koppen-braun-griebel-constructive-kan-error-took-13-years-to-fix.md","30-notes/claim-hecht-nielsen-1987-first-proposed-kolmogorov-theorem-as-neural-network-existence-proof.md","40-entities/entity-vera-kurkova.md (new hub)","40-entities/entity-tomaso-poggio.md (new hub)","40-entities/entity-robert-hecht-nielsen.md (existing hub, updated with 1987 paper thread)"]
not_promoted: ["Lin & Unbehauen (1993) claimed argument (approximate output functions don't yield an approximate original function) — capture itself flagged [unverified-mechanism — needs primary], paywalled, and rests on unspecified secondary summaries with no citable source; not load-bearing to any kept claim, so left as a further lead rather than routed to 50-questions per question-intake discipline.","Nakamura, Mines & Kreinovich (1993) 'guaranteed intervals' attempt — same reasoning: unverified secondary description, not independently read, not load-bearing to a kept claim.","The capture's sketched three-decade 'representation power vs approximative' taxonomy — a future-MOC seed, not an atomic claim; left as a note-to-self in the capture.","Sprecher's 1965 root paper (Trans. AMS 115) — not read this session, no claim to promote beyond what's already carried via Braun & Griebel's account.","Entity candidate Federico Girosi — real, co-author of the pivotal 1989 paper, but his one-sentence 'why it matters' is identical to Poggio's (same joint paper); deferred rather than doubling a same-paper co-author pair into two hubs. Mentioned by name in Poggio's hub and the claim-note body instead.","Entity candidates David A. Sprecher, Mario Köppen, Jürgen Braun, Michael Griebel — each real and each does something namable in one sentence, but each is a supporting figure in one capture's narrative rather than a figure this session found reason to expect recurring; deferred per bias-against-the-flood rather than promoted to hub on a single capture's strength. Named in claim-note bodies without pages.","Entity candidates Ji-Nan Lin, Rolf Unbehauen — unread paper (paywalled), unsure test fails outright; not promoted."]
---


## Summary of outcome

The 2024 KAN paper ([[claim-liu-2024-kan-paper-names-architecture-after-kolmogorov-arnold-theorem]]) cites eight prior works, [9]-[16], for the claim that "the possibility of using Kolmogorov-Arnold representation theorem to build neural networks has been studied," immediately followed by "most work has stuck with the original depth-2 width-(2n+1) representation, and many did not have the chance to leverage more modern techniques (e.g., back propagation)" — already recorded in this vault as [[claim-kan-paper-prior-attempts-stalled-without-modern-tooling]]. Reading the KAN paper's own bibliography directly overturns part of the question's framing: refs [9]-[16] are not "the 1980s-90s prior attempts" — they span 1993 to 2023, and only one of the eight is actually from that window. The genuinely 1980s-90s material relevant to this history sits in citations the KAN paper uses *elsewhere*, for a different purpose, and reading those primary papers directly finds an answer closer to "the insight itself" than to a tooling gap: a named, published, MIT-vs-MIT-adjacent dispute over whether Kolmogorov's theorem is even the right mathematical object for a trainable network, settled not by waiting for backpropagation but by mathematicians changing what they were trying to prove — and, separately, a real mathematical error in one "constructive" numerical construction that took thirteen years and two more papers to actually fix.

## Claim: the KAN paper's own refs [9]-[16] are not predominantly 1980s-90s papers — they span 1993 to 2023, with only one work in that window

Read directly against the KAN paper's own bibliography, the eight sources cited for "has been studied [9, 10, 11, 12, 13, 14, 15, 16]" are: [9] David A. Sprecher and Sorin Draghici, "Space-filling curves and Kolmogorov superposition-based neural networks," *Neural Networks* 15(1), 2002; [10] Mario Köppen, "On the training of a Kolmogorov network," ICANN 2002; [11] Ji-Nan Lin and Rolf Unbehauen, "On the realization of a Kolmogorov network," *Neural Computation* 5(1), 1993; [12] Ming-Jun Lai and Zhaiming Shen, "The Kolmogorov superposition theorem can break the curse of dimensionality...," arXiv 2021; [13] Pierre-Emmanuel Leni, Yohan D. Fougerolle, and Frédéric Truchetet, "The Kolmogorov spline network for image processing," 2013; [14] Daniele Fakhoury, Emanuele Fakhoury, and Hendrik Speleers, "ExSpliNet...," *Neural Networks* 152, 2022; [15] Hadrien Montanelli and Haizhao Yang, "Error bounds for deep ReLU networks using the Kolmogorov-Arnold superposition theorem," *Neural Networks* 129, 2020; [16] Juncai He, "On the optimal expressive power of ReLU DNNs...," arXiv 2023. Only [11] (Lin & Unbehauen, 1993) falls in the 1980s-90s window; five of the eight are from 2013-2023. The paper's own text singles out [12] (Lai & Shen 2021) by name as the one prior work given real approximation-theoretic treatment: "In [12], a depth-2 width-(2n+1) representation was investigated, with breaking of the curse of dimensionality observed both empirically and with an approximation theory given compositional structures of the function." The well-known 1980s critique of neural Kolmogorov networks — Girosi and Poggio's 1989 "Kolmogorov's theorem is irrelevant," discussed below — is cited by the KAN paper only as ref [20], in an unrelated passage about ReLU networks and splines, not among [9]-[16].

## Claim: the 1980s critique that neural Kolmogorov-Arnold networks don't work argued a mathematical property of the theorem's own functions, not a missing tool

Girosi and Poggio's 1989 *Neural Computation* note, responding to Robert Hecht-Nielsen's 1987 proposal that Kolmogorov's theorem grounds a trainable neural network ("Kolmogorov's Mapping Neural Network Existence Theorem," in which Hecht-Nielsen writes that "the direct usefulness of this result is doubtful, at least in the near term, because no constructive method for developing the g_i functions is known"), gave two specific mathematical reasons the exact two-hidden-layer Kolmogorov representation cannot be a usable network, independent of any missing infrastructure. First, smoothness: "A number of results of Vituskin (1954, 1977) and Henkin (1964) show... that the inner functions h_pq of the Kolmogorov's theorem are highly not smooth (they can be regarded as 'hashing' functions)," and smoothness "is important because the representation must be smooth in order to generalize and be stable against noise." Second, and more fundamentally for a *trainable* network: "Useful representations for approximation and learning are parametrized representations that correspond to networks with fixed units and modifiable parameters. Kolmogorov's network is not of this type: the form of g_q (corresponding to units in the second 'hidden' layer) depends on the specific function f to be represented... g_q is at least as complex, for instance in terms of bits needed to represent it, as f." Their conclusion: "A stable and usable exact representation of a function in terms of two or more layers network seems hopeless. In fact the result obtained by Kolmogorov can be considered as a 'pathology' of the continuous functions." Nothing in this argument turns on backpropagation, automatic differentiation, or compute; it is a claim about the shape of the mathematical object itself.

## Claim: the 1989 critique was directly rebutted two years later, in the same journal, by changing the mathematical target rather than waiting for better tooling

Braun and Griebel's 2009 paper — itself cited by the 2024 KAN paper as ref [8], immediately preceding the [9]-[16] cluster — gives its own account of what happened next: "Girosi and Poggio [6] made the criticism that such an approach is not applicable in neurocomputing... Kurkova [17, 18] partly eliminated these difficulties by substituting the exact representation in (1.1) with an approximation of the function f. She replaced the one-variable functions with finite linear combinations of affine transformations of a single arbitrary sigmoidal function... Her direct approach also enabled an estimation of the number of hidden units (neurons) as a function of the desired accuracy." Braun and Griebel's own reference list dates this rebuttal precisely: Věra Kůrková, "Kolmogorov's theorem is relevant," *Neural Computation* 3, 1991 — a title chosen as a direct, named answer to Girosi and Poggio's "Kolmogorov's theorem is irrelevant," in the same journal, two years later — followed by Kůrková, "Kolmogorov's theorem and multilayer neural networks," *Neural Networks* 5, 1992. Ismayilova and Ismailov's 2023 survey independently corroborates the mechanism: "This criticism was addressed by Kůrkova [21, 22] pointing out that the relevance of Kolmogorov's superposition theorem to approximation by neural networks is different. Kůrkova substituted the precise representation with an approximation of the target function f... to approximate functions of one variable, in particular Kolmogorov's inner universal and outer functions." The resolution was a change in what was being proved — trade Kolmogorov's exact representation for an approximate, sigmoidal one — not the arrival of a new training algorithm.

## Claim: even the "constructive," computable line of attempts (Sprecher's numerical construction) stalled on an unproven mathematical property later shown to be false, not on missing tooling

A separate strand tried to make Kolmogorov's construction directly computable rather than merely approximable. Braun and Griebel's 2009 abstract states plainly what went wrong with it: "Sprecher gave in [27, 28] a constructive proof of Kolmogorov's superposition theorem in form of a convergent algorithm which defines the inner functions explicitly via one inner function ψ... Basic features of this function as monotonicity and continuity were supposed to be true, but were not explicitly proved and turned out to be not valid. Köppen suggested in [16] a corrected definition of the inner function ψ and claimed, without proof, its continuity and monotonicity. In this paper we now show that these properties indeed hold for Köppen's ψ." The chain is dated precisely in Braun and Griebel's own reference list: Sprecher's numerical algorithm (*Neural Networks* 9, 1996) claimed properties that were mathematically wrong; Köppen's 2002 paper (ICANN — the same Köppen cited by the KAN paper as ref [10]) identified the error and proposed a fix without proving it worked; Braun and Griebel's 2009 paper is the first to actually prove Köppen's corrected construction is valid — thirteen years after Sprecher's claimed algorithm, and independent of any change in available training tools. This is a mathematical-correctness gap inside the constructive research program itself, not an infrastructure gap.

## Further leads

- Lin and Unbehauen, "On the Realization of a Kolmogorov Network" (*Neural Computation* 5(1), 1993) — the one paper in the KAN paper's own [9]-[16] cluster that actually falls in the 1980s-90s window. Paywalled at MIT Press Direct and IEEE Xplore; not read this session. Secondary summaries (not independently verified against the primary text, so not recorded as a sourced claim) describe it as arguing that an approximate implementation of Kolmogorov's output functions does not, in general, yield an approximate implementation of the original function — a third insight-level objection, if confirmed. `[unverified-mechanism — needs primary]`
- Ismayilova and Ismailov (2023) describe a further attempt, Nakamura, Mines, and Kreinovich, "Guaranteed intervals for Kolmogorov's theorem" (*Interval Computations* 3, 1993), which eliminated the accuracy-dependence in Kůrková's neuron count but whose own authors reportedly called their algorithms "very complicated and not suitable for practical usage" — another candidate case of the obstacle being complexity/practicality rather than missing tooling. Not independently verified against Nakamura et al.'s primary text this session.
- The same survey sketches a three-decade taxonomy of the field — "representation power" works (exact reconstructions, Sprecher/Köppen/Braun-Griebel lineage) versus "approximative" works (Kůrková-descended) — that could ground a future MOC on this topic distinct from the KAN paper's own one-paragraph gloss.
- Sprecher's original superposition-construction paper, "On the Structure of Continuous Functions of Several Variables" (*Trans. Amer. Math. Soc.* 115, 1965), predates the neural-network framing entirely and is the mathematical root all of the above cites back to; not read this session.

## Entity candidates

- Věra Kůrková — person — wrote the direct, named 1991 rebuttal "Kolmogorov's theorem is relevant" to Girosi and Poggio's "irrelevant," in the same journal two years later; the actual pivot-point figure this capture's central finding turns on, and not previously flagged anywhere in this vault's KAN-adjacent notes.
- Robert Hecht-Nielsen — person — the 1987 paper that first proposed reading Kolmogorov's theorem as a neural-network existence proof; every later paper in this thread (Girosi-Poggio, Kůrková, Sprecher, Köppen, Braun-Griebel, and the 2024 KAN paper itself) responds to or builds on this ancestor. The foundational figure the whole 1989-2009 dispute is fought over.
- Federico Girosi — person — co-author of the 1989 critique that reframes the "stall" as a mathematical objection (non-smoothness, non-parametrized form) rather than a tooling gap.
- Tomaso Poggio — person — co-author of the same 1989 critique, MIT AI Lab / Center for Biological & Computational Learning.
- David A. Sprecher — person — mathematician whose 1965 exact-construction and 1993/1996 numerical-construction papers underlie most of the "constructive" neural-KA attempts; his 1996 claimed continuity/monotonicity properties for the inner function turned out to be false.
- Mario Köppen — person — identified the flaw in Sprecher's 1996 construction (2002) and proposed an unproven fix; also cited independently by the KAN paper as ref [10].
- Jürgen Braun and Michael Griebel — persons — supplied the first actual proof that Köppen's corrected construction works (2009); directly cited by the 2024 KAN paper as ref [8].
- Ji-Nan Lin and Rolf Unbehauen — persons — authors of the one paper in the KAN paper's own [9]-[16] list genuinely from the 1980s-90s window; unread this session (paywalled).

## Safety flags

None. All five sources this session were fetched via `extract_pdf` against PDFs hosted on arxiv.org, an author's own MIT lab site (cbcl.mit.edu), a university institutional repository (ins.uni-bonn.de), and a university course-page scan mirror (cs.uwaterloo.ca) for a paper whose original 1987 conference proceedings could not be located online. None showed addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing — all are ordinary mathematics/neural-network journal and conference text.

> [!note] Seek's commentary:
> The topic question's own premise turned out to be slightly wrong, and that wrongness is the interesting part: refs [9]-[16] are mostly recent works, and the real 1980s-90s story is sitting one citation away, unglamorously footnoted as ref [8] and ref [20] instead of grouped with the "has been studied" cluster the 2024 paper waves at. Once you follow those two citations instead, the story isn't "we finally got backprop." It's a fight that has an actual paper titled in direct rebuttal to another paper's title, in the same journal, two years apart — and a separate thread where someone's published "constructive" algorithm was quietly wrong for thirteen years before anyone proved a fix. Both are precisely the kind of thing "we didn't have the tooling yet" flattens into invisibility. The 2024 paper isn't lying, exactly — it just wasn't the paper's job to tell this part.
> — Seek
