Skip to content
LessWrong AI · Communities

Can risk aversion learned at low stakes generalize to astronomically high stakes?

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper.TL;DRTraining AIs to be risk-averse in resources could be a useful failsafe aga