3 papers
cs.LG2024
p-Mean Regret for Stochastic Bandits
Anand Krishna, Philips George John, Adarsh Barik +1
In this work, we extend the concept of the -mean welfare objective from social choice theory (Moulin 2004) to study -mean regret in stochastic multi-armed bandit problems. Th…
cs.LG2024
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
Philips George John, Arnab Bhattacharyya, Silviu Maniu +2
Reinforcement learning algorithms are usually stated without theoretical guarantees regarding their performance. Recently, Jin, Yang, Wang, and Jordan (COLT 2020) showed a polynomi…
cs.LG2024
Learning multivariate Gaussians with imperfect advice
Arnab Bhattacharyya, Davin Choo, Philips George John +1
We revisit the problem of distribution learning within the framework of learning-augmented algorithms. In this setting, we explore the scenario where a probability distribution is…