Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking
arXiv:2606.29088v2 Announce Type: replace-cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully reflect real-world…