37 papers
GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models
Peizheng Guo, Jianqi Zhang, Xingyu Zhang +4
Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by di…
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
Yao Zhou, Hang Gao, Fengge Wu +2
Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies from sparse task-success feedback. In common GRPO-based VLA post-tra…
Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models
Jianqi Zhang, Xingyu Zhang, Zeen Song +3
Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale dataset…
Dirichlet-Guided Group Forecasting for Alleviating Over-smoothing in Time Series Forecasting
Xingyu Zhang, Jingyao Wang, Xin Yu +4
Time series forecasting often suffers from over-smoothing, especially when future dynamics are multi-modal. Forecasts may follow the coarse trend of the observed future, but fail t…
PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models
Peizheng Guo, Jingyao Wang, Changwen Zheng +1
Vision-Language-Action (VLA) models show promising ability in language-guided robotic tasks. However, making VLA policies reliable remains challenging, because a manipulation task…
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
Jingyao Wang, Peizheng Guo, Wenwen Qiang +4
Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coupled to final correctness and pay…