Skip to content
LessWrong AI · Communities

RLVR that rewards red teaming the training environment

Epistemic status: seeing what sticksI've been thinking pretty obsessively about how to make sure the Hugging Face incident doesn't happen again. I don't work at a major lab (shout outs to Anima though), and don't have access to any compute independently, so I can't write a paper on this idea or evaluate how well it wor