LessWrong AI
· Communities
Deliberate Alignment Faking as a Defense Against Model Poisoning
I want to discuss and brainstorm a counterintuitive approach to AI alignment:Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment.To prevent this from going horribly wrong, we add an additional output to the network, to be used during training, which means "I would not normal