1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Bingkui Tong, Junpei Komiyama, Soichiro Nishimori +1
We study a stochastic bandit algorithm motivated by retry-aware objectives that value the best outcome among multiple attempts, such as pass@k and max@k. Given a posterior over…