Skip to content
arXiv cs.LG · Papers

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall o