Skip to content
arXiv stat.ML · Papers

Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling

arXiv:2608.10896v1 Announce Type: new Abstract: Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional ce