activity
20202026
most citedSimulation-Based Benchmarking of Reinforcement Learning Agents for Personalized Retail Promotions

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

stat.ML2026

Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers

Jingkai Huang, Will Ma, Zhengyuan Zhou

A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this pa…

stat.ML2025

Learning an Optimal Assortment Policy under Observational Data

Yuxuan Han, Han Zhong, Miao Lu +2

We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…

stat.ML2024

Learning Optimal Distributionally Robust Stochastic Control in Continuous State Spaces

Shengbo Wang, Jason Meng, Nian Si +2

We study data-driven learning of robust stochastic control for infinite-horizon systems with potentially continuous state and action spaces. In many managerial settings--supply cha…

cs.AI2024★ 1 cited

Simulation-Based Benchmarking of Reinforcement Learning Agents for Personalized Retail Promotions

Yu Xia, Sriram Narayanamoorthy, Zhengyuan Zhou +1

The development of open benchmarking platforms could greatly accelerate the adoption of AI agents in retail. This paper presents comprehensive simulations of customer shopping beha…

cs.LG2023

On the Foundation of Distributionally Robust Reinforcement Learning

Shengbo Wang, Nian Si, Jose Blanchet +1

Motivated by the need for a robust policy in the face of environment shifts between training and deployment, we contribute to the theoretical foundation of distributionally robust…

cs.GT2023

Adaptive, Doubly Optimal No-Regret Learning in Strongly Monotone and Exp-Concave Games with Gradient Feedback

Michael I. Jordan, Tianyi Lin, Zhengyuan Zhou

Online gradient descent (OGD) is well known to be doubly optimal under strong convexity or monotonicity assumptions: (1) in the single-agent setting, it achieves an optimal regret…