Skip to content
X · @teortaxesTex · X / Twitter

RT Bo Liu (Benjamin Liu): Thanks @_akhaliq! RLVR needs a verifier, but open-ended tasks (writing, summarization) have none, and learned judges get gam…

RT Bo Liu (Benjamin Liu)Thanks @_akhaliq! RLVR needs a verifier, but open-ended tasks (writing, summarization) have none, and learned judges get gamed. Inspired by multi-agent debate, we turn the task into a self-play game with a rule-verifiable winner: an information asymmetry makes one player genuinely worse, so rewa