arXiv cs.AI
· Papers
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
arXiv:2607.02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to mul