Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models
arXiv:2604.08881v2 Announce Type: replace Abstract: With the widespread deployment of vision-language large models (VLLMs), their safety alignment faces dual challenges across languages and modalities. Existing methods…