Skip to content
arXiv cs.CL · Papers

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv:2406.01514v4 Announce Type: replace Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow al