LessWrong AI
· Communities
Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas
Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion.[1] We give an initial empirical demonstration of this effect on Kimi K2.6.The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in deter