1 citations · 1 across the 6 of their papers we have counts for
7 papers
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
Jingkai Huang, Will Ma, Zhengyuan Zhou
A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this pa…
Learning an Optimal Assortment Policy under Observational Data
Yuxuan Han, Han Zhong, Miao Lu +2
We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…
Learning Optimal Distributionally Robust Stochastic Control in Continuous State Spaces
Shengbo Wang, Jason Meng, Nian Si +2
We study data-driven learning of robust stochastic control for infinite-horizon systems with potentially continuous state and action spaces. In many managerial settings--supply cha…
Simulation-Based Benchmarking of Reinforcement Learning Agents for Personalized Retail Promotions
Yu Xia, Sriram Narayanamoorthy, Zhengyuan Zhou +1
The development of open benchmarking platforms could greatly accelerate the adoption of AI agents in retail. This paper presents comprehensive simulations of customer shopping beha…
On the Foundation of Distributionally Robust Reinforcement Learning
Shengbo Wang, Nian Si, Jose Blanchet +1
Motivated by the need for a robust policy in the face of environment shifts between training and deployment, we contribute to the theoretical foundation of distributionally robust…
Adaptive, Doubly Optimal No-Regret Learning in Strongly Monotone and Exp-Concave Games with Gradient Feedback
Michael I. Jordan, Tianyi Lin, Zhengyuan Zhou
Online gradient descent (OGD) is well known to be doubly optimal under strong convexity or monotonicity assumptions: (1) in the single-agent setting, it achieves an optimal regret…