arXiv cs.AI
· Papers
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over fixed item sets, treating all items as