Safety by Identity: Investigating Out-of-Distribution Generalization from Fine-Tuning on a Persona
Samip Paudel
Vassar College
Tony Nguyen
Vassar College
With Apart Research
Correspondence: spaudel@vassar.edu
Abstract
Large language models often generalize misaligned behaviors out-of-distribution (OOD). This paper tests whether we can proactively leverage this same mechanism for alignment. We fine-tuned a Qwen3-8B model exclusively on two safety traits—honesty and rule-following—to see if two withheld traits—consistency and transparency—would naturally emerge OOD. They did not. Persona fine-tuning failed to improve transparency and significantly degraded multi-turn consistency, despite maintaining general capabilities. These results reveal an asymmetry in AI alignment: while harmful behaviors generalize easily, safe personas require complex, restrictive boundaries that do not automatically spill over into untrained domains.
Links
Citation
@misc{paudel2026safetybyidentity,
title={Safety by Identity: Investigating Out-of-Distribution Generalization from Fine-Tuning on a Persona},
author={Paudel, Samip and Nguyen, Tony},
year={2026},
howpublished={\url{https://github.com/smpdl/safety-by-identity}}
}