Skip to content
arXiv stat.ML · Papers

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

arXiv:2607.05375v2 Announce Type: replace Abstract: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitt