Skip to content
LessWrong AI · Communities

Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems

Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome.Suppose an AI agent population could suddenl