arXiv cs.AI
· Papers
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
arXiv:2607.18006v1 Announce Type: cross Abstract: Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($leq 4 , mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framewor