4 papers
Cross-Epoch Adaptive Rollout Optimization for RL Post-Training
Yiming Zong, Yige Wang, Jiashuo Jiang
LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt,…
Adaptive Inference for Resource-Constrained Dynamic Pricing
Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
We study dynamic pricing over a finite selling horizon when limited resource capacity determines revenue and the observations available for inference at a prespecified price. Resou…
Online Bidding for Contextual First-Price Auctions with Budgets under One-Sided Information Feedback
Zeng Fu, Jiashuo Jiang, Yuan Zhou
In this paper, we study the problem of learning to bid in repeated first-price auctions with budget constraints. In each period, the decision maker needs to submit a bid to win the…
Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates
Congyuan Duan, Wanteng Ma, Jiashuo Jiang +1
This paper investigates regret minimization, statistical inference, and their interplay in high-dimensional online decision-making based on the sparse linear context bandit model.…