Skip to content
arXiv cs.LG · Papers

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

arXiv:2607.26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. We do not aim at identifying a single optimal arm; instead, we consider the pr