Skip to content
arXiv cs.CL · Papers

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

arXiv:2608.02867v1 Announce Type: new Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves samp