activity
20182025
most citedAdapting to Misspecification in Contextual Bandits

21 citations · 34 across the 9 of their papers we have counts for

collaborators

15 papers

cs.LG2025

An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs

Haolin Liu, Chen-Yu Wei, Julian Zimmert

We study decision making with structured observation (DMSO). Previous work (Foster et al., 2021b, 2023a) has characterized the complexity of DMSO via the decision-estimation coeffi…

cs.LG2025

Decision Making in Hybrid Environments: A Model Aggregation Approach

Haolin Liu, Chen-Yu Wei, Julian Zimmert

Recent work by Foster et al. (2021, 2022, 2023b) and Xu and Zeevi (2023) developed the framework of decision estimation coefficient (DEC) that characterizes the complexity of gener…

cs.LG2024

Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback

Haolin Liu, Zakaria Mhammedi, Chen-Yu Wei +1

We consider regret minimization in low-rank MDPs with fixed transition and adversarial losses. Previous work has investigated this problem under either full-information loss feedba…

cs.LG2024

Incentive-compatible Bandits: Importance Weighting No More

Julian Zimmert, Teodor V. Marinov

We study the problem of incentive-compatible online learning with bandit feedback. In this class of problems, the experts are self-interested agents who might misrepresent their pr…

cs.LG2024

Optimal cross-learning for contextual bandits with unknown context distributions

Jon Schneider, Julian Zimmert

We consider the problem of designing contextual bandit algorithms in the ``cross-learning'' setting of Balseiro et al., where the learner observes the loss for the action they play…

cs.LG2023

Towards Optimal Regret in Adversarial Linear MDPs with Bandit Feedback

Haolin Liu, Chen-Yu Wei, Julian Zimmert

We study online reinforcement learning in linear Markov decision processes with adversarial losses and bandit feedback, without prior knowledge on transitions or access to simulato…