collaborators

6 papers

cs.LG2026

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty

Mingxuan Cui, Duo Zhou, Yuxuan Han +4

Deep reinforcement learning (RL) has achieved remarkable success, yet its deployment in real-world scenarios is often limited by vulnerability to environmental uncertainties. Distr…

stat.ML2026

Robust Assortment Optimization from Observational Data

Miao Lu, Yuxuan Han, Han Zhong +2

Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue und…

cs.LG2026

Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration

Qinxun Bai, Yuxuan Han, Wei Xu +1

The actor-critic (AC) framework has achieved strong empirical success in off-policy reinforcement learning but suffers from the "moving target" problem, where the evaluated policy…

cs.LG2025

PIPCFR: Pseudo-outcome Imputation with Post-treatment Variables for Individual Treatment Effect Estimation

Zichuan Lin, Xiaokai Huang, Jiate Liu +4

The estimation of individual treatment effects (ITE) focuses on predicting the outcome changes that result from a change in treatment. A fundamental challenge in observational data…

stat.ML2025

Learning an Optimal Assortment Policy under Observational Data

Yuxuan Han, Han Zhong, Miao Lu +2

We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…

stat.ML2025

Precise Asymptotics and Refined Regret of Variance-Aware UCB

Yingying Fan, Yuxuan Han, Jinchi Lv +2

In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence…