1 paper
Raymond Feng, Jesse Geneson, Andrew Lee +1
We determine sharp bounds on the price of bandit feedback for several variants of the mistake-bound model. The first part of the paper presents bounds on the r-input weak reinfor…