Skip to content
arXiv stat.ML · Papers

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

arXiv:2608.06545v1 Announce Type: cross Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $varepsilon$-optimal robust policy under the average-reward crite