Safety by Identity
  • Home
  • Results
  • Training Data
  • Eval Prompts
  • System Prompts
  • Responses
  • Paper

Safety by Identity: Investigating Out-of-Distribution Generalization from Fine-Tuning on a Persona

Samip Paudel
Vassar College

Tony Nguyen
Vassar College

With Apart Research

Correspondence: spaudel@vassar.edu

Abstract

Large language models often generalize misaligned behaviors out-of-distribution (OOD). This paper tests whether we can proactively leverage this same mechanism for alignment. We fine-tuned a Qwen3-8B model exclusively on two safety traits—honesty and rule-following—to see if two withheld traits—consistency and transparency—would naturally emerge OOD. They did not. Persona fine-tuning failed to improve transparency and significantly degraded multi-turn consistency, despite maintaining general capabilities. These results reveal an asymmetry in AI alignment: while harmful behaviors generalize easily, safe personas require complex, restrictive boundaries that do not automatically spill over into untrained domains.

Download paper (PDF)

Links

  • GitHub repository
  • Replication guide
  • System prompts

Citation

@misc{paudel2026safetybyidentity,
  title={Safety by Identity: Investigating Out-of-Distribution Generalization from Fine-Tuning on a Persona},
  author={Paudel, Samip and Nguyen, Tony},
  year={2026},
  howpublished={\url{https://github.com/smpdl/safety-by-identity}}
}
  • Edit this page
  • Report an issue