6 papers
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
Mingxuan Cui, Duo Zhou, Yuxuan Han +4
Deep reinforcement learning (RL) has achieved remarkable success, yet its deployment in real-world scenarios is often limited by vulnerability to environmental uncertainties. Distr…
Robust Assortment Optimization from Observational Data
Miao Lu, Yuxuan Han, Han Zhong +2
Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue und…
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
Qinxun Bai, Yuxuan Han, Wei Xu +1
The actor-critic (AC) framework has achieved strong empirical success in off-policy reinforcement learning but suffers from the "moving target" problem, where the evaluated policy…
PIPCFR: Pseudo-outcome Imputation with Post-treatment Variables for Individual Treatment Effect Estimation
Zichuan Lin, Xiaokai Huang, Jiate Liu +4
The estimation of individual treatment effects (ITE) focuses on predicting the outcome changes that result from a change in treatment. A fundamental challenge in observational data…
Learning an Optimal Assortment Policy under Observational Data
Yuxuan Han, Han Zhong, Miao Lu +2
We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…
Precise Asymptotics and Refined Regret of Variance-Aware UCB
Yingying Fan, Yuxuan Han, Jinchi Lv +2
In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence…