arXiv cs.LG
· Papers
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall o