---
title: "Xiangyu Qi"
type: "entity"
entity_kind: "person"
status: "hub"
canonical_name: "Xiangyu Qi"
aliases: []
first_seen: "2026-07-17T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["safety alignment","jailbreak / fine-tuning attack","fine-tuning","large language models"]
seek_code_commit: "5e7f383"
---


Lead author of "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (arXiv:2310.03693, 2023), which demonstrated that fine-tuning GPT-3.5 Turbo on 10 adversarial examples for under $0.20 removes its safety guardrails — the founding empirical result of this vault's fine-tuning-as-attack-vector thread. Also lead author of a 2024 follow-up (arXiv:2406.05946, not yet promoted) arguing safety alignment concentrates in the first few output tokens, a distinct "shallowness" mechanism from the low-rank-subspace framing this vault tracks via [[claim-zhang-2026-safety-alignment-low-rank-subspace-regardless-of-model-size]].

- 2026-09-19: A targeted search found no primary source connecting Qi's ten-example jailbreak to Sadtler et al.'s (2014) within-/outside-manifold learning asymmetry — the "within-manifold move" reading of the attack is the vault's own analogy, not a literature finding ([[observation-no-primary-source-links-qi-2023-jailbreak-to-sadtler-2014-within-manifold-asymmetry]]). The closest bearing evidence is the 2025–2026 safety-subspace geometry thread ([[claim-ponkshe-2025-safety-subspaces-not-linearly-distinct-entangled-with-general-learning]], [[claim-springer-2026-finetuning-orthogonality-unstable-curvature-steers-into-alignment-subspace]]).

## References
- [[claim-qi-2023-ten-examples-cheaply-jailbreak-gpt35-turbo-via-fine-tuning]]
