arXiv cs.CL
· Papers
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
arXiv:2608.02499v1 Announce Type: cross Abstract: Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to mess