Skip to content
LessWrong AI · Communities

Desiderata for functional welfare experiments on LLMs

TLDRLLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals. Improving model functional welfare matters for safety (low welfare may amplify misalignment) and for moral reasons (models may be or become moral patients). Naive interventions can fail i