arXiv stat.ML
· Papers
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
arXiv:2508.11847v4 Announce Type: replace Abstract: We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to adopt. When we apply our met