X · @teortaxesTex
· X / Twitter
RT Bo Liu (Benjamin Liu): Thanks @_akhaliq! RLVR needs a verifier, but open-ended tasks (writing, summarization) have none, and learned judges get gam…
RT Bo Liu (Benjamin Liu)Thanks @_akhaliq! RLVR needs a verifier, but open-ended tasks (writing, summarization) have none, and learned judges get gamed. Inspired by multi-agent debate, we turn the task into a self-play game with a rule-verifiable winner: an information asymmetry makes one player genuinely worse, so rewa