collaborators

9 papers

cs.LG2026

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

Mingjie Hu, Jian-Qiang Hu, Enlu Zhou

Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve…

cs.LG2026

Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs

Meichen Song, Yuhao Wang, Enlu Zhou

In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient exploration is needed to learn t…

cs.LG2026

Adaptive Simulation Experiment for LLM Policy Optimization

Mingjie Hu, Siyang Gao, Jian-qiang Hu +1

Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires specifying a policy that governs…

math.OC2026

Adaptive Distributionally Robust Optimal Control with Bayesian Ambiguity Sets

Wentao Ma, Zhiping Chen, Huifu Xu +1

In stochastic optimal control (SOC), uncertainty may arise from incomplete knowledge of the true probability distribution of the underlying environment, which is known as Knightian…

cs.LG2025

Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions

Xiaoshuang Wang, Yifan Lin, Enlu Zhou

Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty asso…

cs.LG2025

Online Bayesian Risk-Averse Reinforcement Learning

Yuhao Wang, Enlu Zhou

In this paper, we study the Bayesian risk-averse formulation in reinforcement learning (RL). To address the epistemic uncertainty due to a lack of data, we adopt the Bayesian Risk…