collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

Mingjie Hu, Jian-Qiang Hu, Enlu Zhou

Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve…

cs.LG2026

Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs

Meichen Song, Yuhao Wang, Enlu Zhou

In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient exploration is needed to learn t…

cs.LG2026

Adaptive Simulation Experiment for LLM Policy Optimization

Mingjie Hu, Siyang Gao, Jian-qiang Hu +1

Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires specifying a policy that governs…

cs.LG2025

Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions

Xiaoshuang Wang, Yifan Lin, Enlu Zhou

Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty asso…

cs.LG2025

Online Bayesian Risk-Averse Reinforcement Learning

Yuhao Wang, Enlu Zhou

In this paper, we study the Bayesian risk-averse formulation in reinforcement learning (RL). To address the epistemic uncertainty due to a lack of data, we adopt the Bayesian Risk…

cs.LG2025

Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate

Yifan Lin, Yuhao Wang, Enlu Zhou

Reinforcement learning provides a mathematical framework for learning-based control, whose success largely depends on the amount of data it can utilize. The efficient utilization o…