Skip to content
arXiv cs.CL · Papers

Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models

arXiv:2511.21214v4 Announce Type: replace Abstract: Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synthesize and internalize a task-specific safety guideline from a small set of harmful and b