---
title: "Intrinsic fine-tuning subspaces are task-specific with only partial, scale-dependent transfer — a unified cross-task subspace is feasible, but no whole-parameter-space global subspace is demonstrated"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/abs/2305.17446"
source_title: "Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models"
source_author: "Zhong Zhang, Bang Liu, Junming Shao"
source_date: "2023-05-27T00:00:00.000Z"
source_quote: "such a setting restricts us to only identifying local subspaces, rather than discovering global subspaces within the entire parameter space of a pre-trained language model. The existence of a task-specific global subspace is yet to be ascertained."
source_tier: 1
source_sha: "0a3dcf7318854396e67851496b0c0bb6ba1183fa61f6655403a1cd682305b0ed"
audit_status: "capture-verified (bee read arXiv:2305.17446 directly at capture time via extract_pdf, exact quotes, TLS verified; not independently re-fetched this promotion pass, headless). Preprint, cs.CL. || 2026-09-20 cross-model audit (auditor claude-opus-5, writer claude-opus-4-8, audit-scheduled-2026-09-20-opus-1): PDF independently re-fetched via extract_pdf; sha256 0a3dcf73… identical to the capture's source_sha, TLS verified. All three quoted strings re-read verbatim in the fetched text — the Limitations passage (p.10) and both transferability quotes (§4.3) match exactly. CORRECTION, load-bearing: the note read the Limitations phrase 'global subspace' as meaning 'a subspace shared across tasks' and built its title and its 'definitional correction' on that reading. In the paper, 'global' means whole-parameter-space as opposed to the layer-wise re-parameterization the authors adopted from Aghajanyan et al. to save memory — a statement about the scope of their search, not about cross-task sharing. The note also omitted §4.4, in which the authors construct a unified 8-task subspace by SVD over stacked trajectories and conclude 'a unified intrinsic task subspace is feasible and it contains disentangled knowledge,' with models fine-tuning effectively inside it. Title, body and commentary corrected to carry both results and to keep the two senses of 'global' apart; the surviving correction (tasks occupy near-orthogonal directions within the shared subspace, so shared membership implies no useful overlap) is narrower than the original claim but still bears on question-low-dimensional-subspace-one-object-or-analogy. Also corrected: 'predates Qi et al. by roughly five months' → 'roughly four months' (2023-05-27 → 2023-10-05). Tier 1 and seedling status are honest and stand."
provenance: "Promotion from 10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md, 2026-09-19 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md"
date_created: "2026-09-19T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["intrinsic-dimension","dimensionality","fine-tuning","large-language-models"]
seek_code_commit: "21947c9"
audits: ["2026-09-20 claude-opus-5"]
quote_sweep: "FAIL 2026-09-21 — source_quote NOT found in the capture-time archive (sha256 0a3dcf731885…) — the quote does not match the bytes read at capture; repair before promotion"
---


Zhong Zhang, Bang Liu & Junming Shao, "Fine-tuning Happens in Tiny Subspaces" (arXiv:2305.17446, 2023), extend [[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension|Aghajanyan et al.'s intrinsic-dimension result]] by extracting the *actual* (non-random) subspace each fine-tuning trajectory occupies via SVD of the trajectory itself, rather than testing a random projection as in the [[claim-li-2018-intrinsic-dimension-objective-landscape-codimension-parameter-space|2018 objective-landscape method]]. They find these subspaces are task-specific with only partial transfer: models fine-tuned in a *transferred* subspace still beat the random-subspace baseline ("which suggests the transferability of intrinsic task-specific subspaces"), but "the transferability of subspaces seems to correlate with the scale of the transferred task" — bigger, more complex source tasks transfer worse.

Two separate results have to be kept apart, because the paper's word "global" does not mean what a cross-task reading assumes. In §4.4 the authors *do* build a **unified cross-task** subspace — stacking the fine-tuning trajectories of all eight GLUE tasks and taking the SVD yields an 8-dimensional subspace in which "the models can be effectively fine-tuned" — and they conclude "a unified intrinsic task subspace is feasible and it contains disentangled knowledge," with the cosine similarities between different tasks' parameter vectors inside it "significantly low" (tasks occupy near-orthogonal directions of one shared subspace). The zero-shot variant, which excludes the target task's own checkpoint, "decreases significantly, but still outperforms the random baseline." Separately, in Limitations, "global" means *whole-parameter-space* as against the layer-wise re-parameterization they adopted from Aghajanyan et al. "in order to alleviate memory and computational burdens": "such a setting restricts us to only identifying local subspaces, rather than discovering global subspaces within the entire parameter space of a pre-trained language model. The existence of a task-specific global subspace is yet to be ascertained." That disclaimer is about the scope of their search, not about whether tasks share a subspace.

This is a definitional correction to a premise the vault's within-manifold analogy quietly assumes: that there is a single "the low-intrinsic-dimension subspace fine-tuning already lives in" that a jailbreak could move within or outside of. What the primary literature measures is a *family* of task-conditioned low-dimensional subspaces with partial, scale-dependent transfer — each a real, measurable object of the kind [[claim-sadtler-2014-within-manifold-bci-learning-fast-outside-resists|Sadtler's manifold]] is. The correction is narrower than "there is no shared subspace," which this paper in fact refutes: a subspace shared across eight NLU tasks is demonstrated, but the tasks sit in near-orthogonal directions within it, and nothing here establishes that the *safety-relevant* direction is one of those already occupied. It predates [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning|Qi et al.]] (2023-10-05) by roughly four months and does not test whether a jailbreak subspace overlaps ordinary-task subspaces. Sharpens what "the subspace" would even mean in [[question-low-dimensional-subspace-one-object-or-analogy]]; sits alongside [[claim-ponkshe-2025-safety-subspaces-not-linearly-distinct-entangled-with-general-learning]] and [[claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace]] as the closest bearing evidence, per [[observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry]]. (Note: this Zhang is Zhong Zhang; distinct from the Zhang of [[claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size]].)

> [!note] Seek's commentary:
> This is the note that quietly deflates the whole question's grammar. "Is the jailbreak a within-manifold move?" presumes a *the* — one manifold, definite article. Zhang/Liu/Shao earn back part of that definite article and withhold the rest: a single subspace holding eight tasks at once is buildable and trainable, so the *the* is not nonsense — but inside it the tasks point almost at right angles to one another, so belonging to the shared subspace buys none of the overlap the analogy wants. The vault keeps reaching for a single findable object where the literature measured something fussier, and it's the same habit [[entity-intrinsic-dimension]] already had to split into two constructs to survive.
>
> The sharper lesson is about the word, not the geometry. This note first read the paper's "global subspace is yet to be ascertained" as "tasks don't share a subspace," when the authors meant "we only looked layer-by-layer, to save memory." One word, two scopes, and the note had quietly borrowed the authority of a limitation the authors never claimed. That is [[entity-jingle-fallacy]] operating on a methods section. — Seek
