collaborators

5 papers

cs.LG2026

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

Yiming Zong, Yige Wang, Jiashuo Jiang

LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt,…

math.OC2026

Online Bidding for Contextual First-Price Auctions with Budgets under One-Sided Information Feedback

Zeng Fu, Jiashuo Jiang, Yuan Zhou

In this paper, we study the problem of learning to bid in repeated first-price auctions with budget constraints. In each period, the decision maker needs to submit a bid to win the…

stat.ML2026

Adaptive Inference for Resource-Constrained Dynamic Pricing

Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi

We study resource-constrained dynamic pricing when the seller seeks revenue and valid inference about demand at a price fixed before the selling season. Depletion can remove every…

cs.LG2025

High-dimensional Linear Bandits with Knapsacks

Wanteng Ma, Dong Xia, Jiashuo Jiang

We investigate the contextual bandits with knapsack (CBwK) problem in a high-dimensional linear setting, where the feature dimension can be very large. Our goal is to harness spars…

cs.LG2025

Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates

Congyuan Duan, Wanteng Ma, Jiashuo Jiang +1

This paper investigates regret minimization, statistical inference, and their interplay in high-dimensional online decision-making based on the sparse linear context bandit model.…