arXiv stat.ML
· Papers
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions
arXiv:2608.06545v1 Announce Type: cross Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $varepsilon$-optimal robust policy under the average-reward crite