Skip to content
LessWrong AI · Communities

Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven

This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss.Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instru