Skip to content
arXiv stat.ML · Papers

Stochastic Linear Bandits with Partially Observed Actions

arXiv:2607.08971v1 Announce Type: cross Abstract: The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinate