Skip to content
arXiv cs.AI · Papers

Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking

arXiv:2606.29088v2 Announce Type: replace-cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully reflect real-world bugfixing practices. They are small, weakening statistical reliability, and the buggy programs are often