X · @jeremyphoward
· X / Twitter
RT Yifan Wu: Introducing SWE-Together: a multi-turn benchmark built from real user–agent coding sessions. Coding agents are often benchmarked like ex…
RT Yifan WuIntroducing SWE-Together: a multi-turn benchmark built from real user–agent coding sessions.Coding agents are often benchmarked like exam-takers: given the full spec up front, then graded on the final code. But real coding help is a conversation — users clarify goals, add constraints, and correct course alon