From the 1 of 4 linked papers with an AI index.
4 papers
Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms
Miao Lu, Han Zhong, Tong Zhang +1
The paper studies reinforcement learning where the learner must be robust to differences between training and deployment environments, using interactive data collection and proposi…
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training
Junze Ye, Jiayi Cheng, Miao Lu +3
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses…
Robust Assortment Optimization from Observational Data
Miao Lu, Yuxuan Han, Han Zhong +2
Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue und…
Learning an Optimal Assortment Policy under Observational Data
Yuxuan Han, Han Zhong, Miao Lu +2
We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…