Skip to content
LessWrong AI · Communities

Models don’t seem to be dishonest in the way humans are

TLDRModels often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrela