Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Cross-Epoch Adaptive Rollout Optimization for RL Post-Training
Yiming Zong, Yige Wang, Jiashuo Jiang
LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt,…
cs.LG2024
Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates
Congyuan Duan, Wanteng Ma, Jiashuo Jiang +1
This paper investigates regret minimization, statistical inference, and their interplay in high-dimensional online decision-making based on the sparse linear context bandit model.…
cs.LG2023
High-dimensional Linear Bandits with Knapsacks
Wanteng Ma, Dong Xia, Jiashuo Jiang
We investigate the contextual bandits with knapsack (CBwK) problem in a high-dimensional linear setting, where the feature dimension can be very large. Our goal is to harness spars…