arXiv cs.LG
· Papers
Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning
arXiv:2608.07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumption