arXiv cs.CL
· Papers
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
arXiv:2607.12893v1 Announce Type: cross Abstract: Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the cor