Skip to content
arXiv cs.LG · Papers

Thompson Sampling Is 2-Competitive for Mistakes

arXiv:2607.12389v1 Announce Type: cross Abstract: We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy. Our analysis applies as long as the latent arm processes are independent and each arm evolves on