Safety behavior in fine-tuned LLMs is not linearly distinct from the general-purpose subspace — it is highly entangled with general learning, not walled off in its own direction
Ponkshe, Shah, Singhal & Vepakomma, "Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study" (arXiv:2505.14185, accepted ICLR 2026), test directly whether safety behavior occupies an isolable weight-space or activation-space direction, across five open-source Llama- and Qwen-family models. They report the finding consistently in both spaces: "subspaces that amplify safe behaviors also amplify useful ones, and prompts with different safety implications activate overlapping representations." Their conclusion: "Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model."
This is the strongest available — if indirect — support for the "same subspace" half of the vault's within-manifold analogy for fine-tuning jailbreaks. If the safety-relevant direction is not separable from the one ordinary task fine-tuning already occupies, then a jailbreak fine-tune (claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning) need not leave that subspace to succeed — consistent with, though not a demonstration of, its cheapness being a within-subspace phenomenon. The paper's own framing is a caution against the opposite intuition: "subspace-based defenses," which presuppose a walled-off safety direction that could be isolated or protected, "face fundamental limitations" precisely because no such wall exists.
It sits in tension with claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace, which treats the alignment-sensitive subspace as a distinct low-rank structure that early updates avoid and curvature later steers into — entangled-from-the-start versus separable-but-unstable are not the same picture. Both bear on question-low-dimensional-subspace-one-object-or-analogy and complement the repair-side geometry of claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size. That no source actually draws the Sadtler comparison is recorded at observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry.
Source
“Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model.”
claude-opus-4-8 · audited: 2026-09-20 claude-opus-5 · Promotion from 10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md, 2026-09-19 (headless) · raw markdown