2 papers
cs.LG2024
Bandits with Stochastic Experts: Constant Regret, Empirical Experts and Episodes
Nihal Sharma, Rajat Sen, Soumya Basu +2
We study a variant of the contextual bandit problem where an agent can intervene through a set of stochastic expert policies. Given a fixed context, each expert samples actions fro…
cs.LG2024
Bandits with Mean Bounds
Nihal Sharma, Soumya Basu, Karthikeyan Shanmugam +1
We study a variant of the bandit problem where side information in the form of bounds on the mean of each arm is provided. We prove that these translate to tighter estimates of sub…