6 papers
GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models
Peizheng Guo, Jianqi Zhang, Xingyu Zhang +4
Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by di…
PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models
Peizheng Guo, Jingyao Wang, Changwen Zheng +1
Vision-Language-Action (VLA) models show promising ability in language-guided robotic tasks. However, making VLA policies reliable remains challenging, because a manipulation task…
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
Jingyao Wang, Peizheng Guo, Wenwen Qiang +4
Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coupled to final correctness and pay…
Rethinking Multi-Modal Learning from Gradient Uncertainty
Peizheng Guo, Jingyao Wang, Wenwen Qiang +3
Multi-Modal Learning (MML) integrates information from diverse modalities to improve predictive accuracy. While existing optimization strategies have made significant strides by mi…
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
Peizheng Guo, Jingyao Wang, Wenwen Qiang +3
Despite Multimodal Large Language Models (MLLMs) having shown impressive capabilities, they may suffer from hallucinations. Empirically, we find that MLLMs attend disproportionatel…
Exploring Transferability of Self-Supervised Learning by Task Conflict Calibration
Huijie Guo, Jingyao Wang, Peizheng Guo +3
In this paper, we explore the transferability of SSL by addressing two central questions: (i) what is the representation transferability of SSL, and (ii) how can we effectively mod…