6 papers
An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We study decision making with structured observation (DMSO). Previous work (Foster et al., 2021b, 2023a) has characterized the complexity of DMSO via the decision-estimation coeffi…
Decision Making in Hybrid Environments: A Model Aggregation Approach
Haolin Liu, Chen-Yu Wei, Julian Zimmert
Recent work by Foster et al. (2021, 2022, 2023b) and Xu and Zeevi (2023) developed the framework of decision estimation coefficient (DEC) that characterizes the complexity of gener…
A Model Selection Approach for Corruption Robust Reinforcement Learning
Chen-Yu Wei, Christoph Dann, Julian Zimmert
We develop a model selection approach to tackle reinforcement learning with adversarial corruption in both transition and reward. For finite-horizon tabular MDPs, without prior kno…
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
Haolin Liu, Zakaria Mhammedi, Chen-Yu Wei +1
We consider regret minimization in low-rank MDPs with fixed transition and adversarial losses. Previous work has investigated this problem under either full-information loss feedba…
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
Saeed Masoudian, Julian Zimmert, Yevgeny Seldin
We propose a new best-of-both-worlds algorithm for bandits with variably delayed feedback. In contrast to prior work, which required prior knowledge of the maximal delay $d_{\mathr…
Incentive-compatible Bandits: Importance Weighting No More
Julian Zimmert, Teodor V. Marinov
We study the problem of incentive-compatible online learning with bandit feedback. In this class of problems, the experts are self-interested agents who might misrepresent their pr…