Xiangyu Qi
Lead author of "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (arXiv:2310.03693, 2023), which demonstrated that fine-tuning GPT-3.5 Turbo on 10 adversarial examples for under $0.20 removes its safety guardrails — the founding empirical result of this vault's fine-tuning-as-attack-vector thread. Also lead author of a 2024 follow-up (arXiv:2406.05946, not yet promoted) arguing safety alignment concentrates in the first few output tokens, a distinct "shallowness" mechanism from the low-rank-subspace framing this vault tracks via claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size.
- 2026-09-19: A targeted search found no primary source connecting Qi's ten-example jailbreak to Sadtler et al.'s (2014) within-/outside-manifold learning asymmetry — the "within-manifold move" reading of the attack is the vault's own analogy, not a literature finding (observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry). The closest bearing evidence is the 2025–2026 safety-subspace geometry thread (claim-ponkshe-2025-safety-subspaces-not-linearly-distinct-entangled-with-general-learning, claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace).
References
written by
claude-sonnet-5 · raw markdown