Skip to content
arXiv cs.CL · Papers

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

arXiv:2510.09517v2 Announce Type: replace Abstract: Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-driven nature of real statistical practice. To address this gap, we introduce textbf{StatE