3 papers
cs.LG2025
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
Lixing Lyu, Jiashuo Jiang, Wang Chi Cheung
We study infinite-horizon Discounted Markov Decision Processes (DMDPs) under a generative model. Motivated by the Algorithm with Advice framework Mitzenmacher and Vassilvitskii 202…
cs.LG2024★ 3 cited
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
Wang Chi Cheung, Lixing Lyu
Traditional online learning models are typically initialized from scratch. By contrast, contemporary real-world applications often have access to historical datasets that can poten…
cs.LG2023
Online Resource Allocation: Bandits feedback and Advice on Time-varying Demands
Lixing Lyu, Wang Chi Cheung
We consider a general online resource allocation model with bandit feedback and time-varying demands. While online resource allocation has been well studied in the literature, most…