arXiv stat.ML
· Papers
Stochastic Linear Bandits with Partially Observed Actions
arXiv:2607.08971v1 Announce Type: cross Abstract: The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinate