LessWrong AI
· Communities
Optimiser Choice Can Amplify or Suppress Emergent Misalignment
This is a linkpost for https://arxiv.org/abs/2606.31591. Work done with Patrick Leask and Lev McKinney during the Astra Fellowship.TL;DR: Optimiser choice strongly influences emergent misalignment, while model size and family seem to barely matter. Optimisers that concentrate the LoRA update into fewer directions degra