Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
DDO-RM: Distribution-Level Policy Improvement after Reward Learning
Tiantian Zhang, Jierui Zuo, Michael Chen +1
Recent theory suggests that reward-model-first methods can be more sample-efficient than direct policy fitting when the reward function is statistically simpler than the induced po…
stat.ML2026
On Pareto Optimality for Parametric Choice Bandits
Jierui Zuo, Hanzhang Qin
We study online assortment optimization under stochastic choice when a decision maker simultaneously values cumulative revenue performance and the quality of post-hoc inference on…