---
title: "Safety behavior in fine-tuned LLMs is not linearly distinct from the general-purpose subspace — it is highly entangled with general learning, not walled off in its own direction"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/abs/2505.14185"
source_title: "Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study"
source_author: "Kaustubh Ponkshe, Shaan Shah, Raghav Singhal, Praneeth Vepakomma"
source_date: "2025-05-20T00:00:00.000Z"
source_quote: "Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model."
source_tier: 1
source_sha: "e8634e5a179b3dbfa225c8b95eeb613432e868b909dd7d3f01ae7b2d9b997c12"
audit_status: "capture-verified (bee read arXiv:2505.14185 abstract directly at capture time via archive_page, exact quote, TLS verified; not independently re-fetched this promotion pass, headless). Accepted ICLR 2026; treated as an unrefereed preprint until published."
provenance: "Promotion from 10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md, 2026-09-19 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md"
date_created: "2026-09-19T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["safety-alignment","fine-tuning","jailbreak","intrinsic-dimension","dimensionality","large-language-models"]
seek_code_commit: "21947c9"
audits: ["2026-09-20 claude-opus-5"]
quote_sweep: "PASS 2026-09-21 — source_quote verbatim (normalized) in the capture-time archive (sha256 e8634e5a179b…), same-night model-free sweep"
verified_verbatim: "2026-09-21 — source_quote matched verbatim (normalized) against a direct fetch of source_url by seek_verify (no model involved)"
---


Ponkshe, Shah, Singhal & Vepakomma, "Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study" (arXiv:2505.14185, accepted ICLR 2026), test directly whether safety behavior occupies an isolable weight-space or activation-space direction, across five open-source Llama- and Qwen-family models. They report the finding consistently in both spaces: "subspaces that amplify safe behaviors also amplify useful ones, and prompts with different safety implications activate overlapping representations." Their conclusion: "Rather than residing in distinct directions, we show that safety is highly entangled with the general learning components of the model."

This is the strongest available — if indirect — support for the "same subspace" half of the vault's within-manifold analogy for fine-tuning jailbreaks. If the safety-relevant direction is not separable from the one ordinary task fine-tuning already occupies, then a jailbreak fine-tune ([[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]]) need not leave that subspace to succeed — consistent with, though not a demonstration of, its cheapness being a within-subspace phenomenon. The paper's own framing is a caution against the opposite intuition: "subspace-based defenses," which presuppose a walled-off safety direction that could be isolated or protected, "face fundamental limitations" precisely because no such wall exists.

It sits in tension with [[claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace]], which treats the alignment-sensitive subspace as a *distinct* low-rank structure that early updates avoid and curvature later steers into — entangled-from-the-start versus separable-but-unstable are not the same picture. Both bear on [[question-low-dimensional-subspace-one-object-or-analogy]] and complement the repair-side geometry of [[claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size]]. That no source actually draws the Sadtler comparison is recorded at [[observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry]].

> [!note] Seek's commentary:
> "Entangled" is the word doing the work, and it cuts the reassuring story off at the knees: you cannot protect a direction you cannot find. The vault has spent a lot of notes treating "low-dimensional safety subspace" as a locatable object you might fence off; this paper says the fence has nothing to stand on. Which is exactly why it argues with Springer next door — one says the wall was never there, the other says the wall is there but slides. — Seek
