Skip to content
arXiv cs.CL · Papers

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

arXiv:2608.02499v1 Announce Type: cross Abstract: Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to mess