arXiv stat.ML
· Papers
Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling
arXiv:2608.10896v1 Announce Type: new Abstract: Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional ce