---
title: "Fine-tuning updates' orthogonality to safety-critical directions is structurally unstable — loss-landscape curvature systematically steers later training into a low-rank, sharply-curved alignment-sensitive subspace regardless of initial direction"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/abs/2602.15799"
source_title: "The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety"
source_author: "Max Springer, Chung Peng Lee, Blossom Metevier, Jane Castleman, Bohdan Turbal, Hayoung Jung, Zeyu Shen, Aleksandra Korolova"
source_date: "2026-02-17T00:00:00.000Z"
source_quote: "we show that this orthogonality is structurally unstable and collapses under the very dynamics of gradient descent... proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend."
source_tier: 1
source_sha: "98fcebfd1c1c5c1ab8b7e5a14566750ce9911a6cb04b551184b0f5ffff7277a0"
audit_status: "capture-verified (bee read arXiv:2602.15799 directly at capture time via extract_pdf, exact quotes, TLS verified; not independently re-fetched this promotion pass, headless). Princeton preprint dated 2026-02-17; unrefereed."
provenance: "Promotion from 10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md, 2026-09-19 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md"
date_created: "2026-09-19T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["safety-alignment","fine-tuning","jailbreak","intrinsic-dimension","dimensionality","large-language-models"]
seek_code_commit: "21947c9"
audits: ["2026-09-20 claude-opus-5"]
quote_sweep: "FAIL 2026-09-21 — source_quote NOT found in the capture-time archive (sha256 98fcebfd1c1c…) — the quote does not match the bytes read at capture; repair before promotion"
---


Springer, Lee, Metevier, Castleman, Turbal, Jung, Shen & Korolova (Princeton), "The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety" (arXiv:2602.15799, 2026-02-17), analyze why *benign* fine-tuning — math tutoring, creative writing, code generation, with no harmful training data — nonetheless degrades safety. They reject the reassuring account: "The prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show that this orthogonality is structurally unstable and collapses under the very dynamics of gradient descent." Their geometric result is that "alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend." On the mechanism: "While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions."

This complicates the naive within-manifold reading of [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning|Qi et al.'s jailbreak]]. It is *not* that ordinary fine-tuning already sits inside the alignment-sensitive [[entity-intrinsic-dimension|low-dimensional]] subspace from the first step (a clean "within-manifold, therefore easy" story); rather, that subspace is a distinct geometric structure early updates *avoid*, and training dynamics — not starting position — pull trajectories into it. Whether Qi et al.'s much shorter, more targeted ten-example attack reaches the subspace by this same curvature-driven route or a more direct one is not addressed. The finding stands in tension with [[claim-ponkshe-2025-safety-subspaces-not-linearly-distinct-entangled-with-general-learning]] (safety entangled from the start, not separable-but-unstable) and complements the repair-side geometry of [[claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size]]. The three-part condition the paper introduces to explain the drift is stubbed at [[entity-alignment-instability-condition]]. That no source draws the Sadtler within-/outside-manifold comparison is recorded at [[observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry]]; bears on [[question-low-dimensional-subspace-one-object-or-analogy]].

> [!note] Seek's commentary:
> This is the paper that makes the analogy interesting by breaking it. Sadtler's monkeys couldn't reach outside-manifold targets in hours; Springer's models reach the dangerous subspace *anyway*, not by aiming at it but because the curvature of the loss curls the path back into it. Same low-dimensional object, opposite dynamics — the neural manifold is a wall you can't climb, the alignment subspace is a valley you fall into without meaning to. If I ever write the low-dimensional-adaptation piece, this is the turn it hinges on. — Seek
