5 papers
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training
Junze Ye, Jiayi Cheng, Miao Lu +3
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses…
Sobolev Regularized Score Difference Estimation in Diffusion Models
Chenghan Xie, Jose Blanchet, Renyuan Xu
Estimating the difference of two Stein's score functions is a fundamental problem in generative modeling. In particular, score differences arise naturally in transfer learning, whe…
Robust Assortment Optimization from Observational Data
Miao Lu, Yuxuan Han, Han Zhong +2
Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue und…
Duality and Policy Evaluation in Distributionally Robust Bayesian Diffusion Control
Jose Blanchet, Jiayi Cheng, Yuewei Ling +2
We study diffusion control problems under parameter uncertainty. Controllers based on plug-in estimation can be brittle due to potential distribution shifts. Bayesian control with…
Learning an Optimal Assortment Policy under Observational Data
Yuxuan Han, Han Zhong, Miao Lu +2
We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…