2 papers
cs.LG2026
Cross-Epoch Adaptive Rollout Optimization for RL Post-Training
Yiming Zong, Yige Wang, Jiashuo Jiang
LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt,…
cs.LG2026
Non-Stationary Online Resource Allocation: Learning from a Single Sample
Yiding Feng, Jiashuo Jiang, Yige Wang
We study online resource allocation under non-stationary demand with a minimum offline data requirement. In this problem, a decision-maker must allocate multiple types of resources…