Skip to content
arXiv cs.AI · Papers

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

arXiv:2605.19665v2 Announce Type: replace-cross Abstract: Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpretability by decomposing evaluation into ex