2 papers
cs.LG2025
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
Wang Chi Cheung, Lixing Lyu
Traditional online learning models are typically initialized from scratch. By contrast, contemporary real-world applications often have access to historical datasets that can poten…
cs.LG2025
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
Lixing Lyu, Jiashuo Jiang, Wang Chi Cheung
We study infinite-horizon Discounted Markov Decision Processes (DMDPs) under a generative model. Motivated by the Algorithm with Advice framework Mitzenmacher and Vassilvitskii 202…