LessWrong AI
· Communities
Models don’t seem to be dishonest in the way humans are
TLDRModels often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrela