Skip to content
arXiv cs.CL · Papers

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as a lack of exceptionally challengi