arXiv stat.ML
· Papers
PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
arXiv:2507.20068v3 Announce Type: replace-cross Abstract: Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary datasets, such as those synthesized by generative models, can improve the accuracy of OPE