LessWrong AI
· Communities
Proposal: The Glasswing Standard
Thinking about "Plan A" makes me want to make concrete proposals towards those goals.I think Anthropic's "Project Glasswing" provides a clear and easily implemented first-step policy towards AI safety. With a few small tweaks, I think we can build a release process that is robust against today's mundane threats, while