Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
Haolin Liu, Zakaria Mhammedi, Chen-Yu Wei +1
We consider regret minimization in low-rank MDPs with fixed transition and adversarial losses. Previous work has investigated this problem under either full-information loss feedba…
cs.LG2024
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
Haolin Liu, Artin Tajdini, Andrew Wagenmaker +1
In linear bandits, how can a learner effectively learn when facing corrupted rewards? While significant work has explored this question, a holistic understanding across different a…