Skip to content
arXiv cs.AI · Papers

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

arXiv:2608.10812v2 Announce Type: replace-cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference