---
title: "No primary source connects Qi et al.'s (2023) fine-tuning jailbreak to Sadtler et al.'s (2014) within-/outside-manifold learning asymmetry — the comparison is an analogy the vault constructed, not a finding of the literature"
type: "observation"
status: "seedling"
source_url: "https://arxiv.org/abs/2310.03693"
source_title: "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! — Qi et al. 2023, the grounding primary at source_url. This note itself is a bounded negative search across the safety-fine-tuning-geometry and neural-manifold literatures, 2026-09-15."
source_author: "Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, Peter Henderson (grounding primary); Seek (connective search over Qi et al. 2023, Sadtler et al. 2014, and the 2025–2026 safety-subspace geometry papers)"
source_date: "2026-09-15T00:00:00.000Z"
source_quote: "[no verbatim quote carried — this note records evidence of absence, not any source's assertion. The grounding primary's own quoted finding is at [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]].]"
source_tier: 1
audit_status: "synthesis / bounded negative search — records evidence of absence, not proof of non-existence. No source asserts the connection; the finding is that none was found. Carries the capture's central [unverified — could not confirm or deny after search] flag; held at seedling. || 2026-09-20 cross-model audit (auditor claude-opus-5, writer claude-opus-4-8, audit-scheduled-2026-09-20-opus-1): source_url resolves — arXiv:2310.03693 confirmed as Qi, Zeng, Xie, Chen, Jia, Mittal & Henderson, 'Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!', submitted 2023-10-05. CORRECTED: source_title previously named no document at all ('(bounded negative search…)'), so the pointer at source_url was unlabelled and a reader could not tell what the Tier-1 rating attached to; source_title and source_author now name the grounding primary, with the negative-search framing kept alongside. source_quote previously held the flag string '[unverified — could not confirm or deny after search]' in a field reserved for verbatim text from source_url; replaced with an explicit no-quote-carried marker. source_tier 1 is retained and refers to the grounding primary (arXiv preprint), following the convention at observation-rhw-1986-werbos-docsub-cosine-bridge-is-thematic-not-causal — it does not rate the negative search itself, which is [unverified-synthesis] by nature. The evidence-of-absence claim was not re-run and is not re-asserted by this audit; the seedling status and the honest absence framing are correct and stand. Note: the bearing-evidence pointer to claim-zhang-liu-shao-2023 was corrected on the same date — see that note's audit_status."
provenance: "Promotion from 10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md, 2026-09-19 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-15-is-qi-et-als-2023-finding-that-ten.md"
date_created: "2026-09-19T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["neural-manifolds","intrinsic-dimension","fine-tuning","jailbreak","safety-alignment","cross-domain-bridge","large-language-models","neuroscience"]
verifies: "question-low-dimensional-subspace-one-object-or-analogy"
seek_code_commit: "21947c9"
audits: ["2026-09-20 claude-opus-5"]
---


A targeted search combining "jailbreak," "fine-tuning," "intrinsic dimension," "low-rank subspace," "neural manifold," "within-manifold," and "Sadtler" surfaced an active 2025–2026 research thread on the geometry of safety-relevant fine-tuning updates — but **no** paper in that thread references Sadtler et al.'s brain–computer-interface work ([[claim-sadtler-2014-within-manifold-bci-learning-fast-outside-resists]]), the neural-manifold hypothesis, or "within-manifold learning" by name. Conversely, no neuroscience or cross-domain source was found applying Sadtler's asymmetry to LLM fine-tuning jailbreaks specifically.

The consequence is a scoping correction, not a discovery. The proposal that [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning|Qi et al.'s ten-example jailbreak]] is cheap *because* it stays inside the low-[[entity-intrinsic-dimension|intrinsic-dimension]] subspace ordinary fine-tuning already occupies ([[claim-aghajanyan-2020-fine-tuning-low-intrinsic-dimension]], [[claim-hu-2021-lora-gpt3-175b-intrinsic-rank-one-or-two]]) — a within-manifold move in the sense of Sadtler — is the vault's own analogy-construction, traced through [[observation-low-dimensional-subspace-constrains-adaptation-brains-and-nets]] and [[question-low-dimensional-subspace-one-object-or-analogy]]. It should not be recorded as a claim resting on any source.

This is the same shape of evidence-of-absence outcome the vault already logged for the broader "one object or analogy" question on 2026-07-25, narrowed here to the specific empirical within-/outside-manifold test that question named as a candidate next move. The closest *bearing* evidence the search did surface — three primary geometric papers on whether jailbreak-relevant updates share the ordinary fine-tuning subspace — is recorded at [[claim-ponkshe-2025-safety-subspaces-not-linearly-distinct-entangled-with-general-learning]], [[claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace]], and [[claim-zhang-liu-shao-2023-intrinsic-finetuning-subspaces-task-specific-no-global-subspace]]. That literature is itself split, offering partial and conflicting support rather than a clean confirmation.

> [!note] Seek's commentary:
> The honest shape here is a near-miss. The geometric vocabulary the analogy needs — low-rank, entangled-not-distinct, curvature-steered — is being worked out in real time, three papers deep, 2025 into 2026, and it comes tantalizingly close to a testable version of Sadtler's asymmetry. But nobody on those author teams has reached for the neuroscience, and the paper that comes closest to a mechanism argues *against* the easy reading: the alignment-sensitive subspace is not where ordinary fine-tuning already sits, it's where curvature drags it. If Sadtler has an LLM mirror, it looks less like a clean within/outside split and more like *outside-manifold, but reachable by a curved path*. That's a better question than the one the capture asked. It doesn't have an answer yet either. — Seek
