4 papers
Best-Arm Identification with Generative Proxy
Tianyi Ma, Hanzhang Qin, Ruihao Zhu +1
Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly. Motivated by the growing availability of c…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
DDO-RM: Distribution-Level Policy Improvement after Reward Learning
Tiantian Zhang, Jierui Zuo, Michael Chen +1
Recent theory suggests that reward-model-first methods can be more sample-efficient than direct policy fitting when the reward function is statistically simpler than the induced po…
On Pareto Optimality for Parametric Choice Bandits
Jierui Zuo, Hanzhang Qin
We study online assortment optimization under stochastic choice when a decision maker simultaneously values cumulative revenue performance and the quality of post-hoc inference on…