collaborators

6 papers

cs.CV2026

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Caoyuan Ma, Wenpu Liu, Weichu Xie +12

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propos…

cs.AI2026

DOPD: Dual On-policy Distillation

Xinlei Yu, Gen Li, Qingyi Si +13

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…

cs.LG2026

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Wenpu Liu, Yuqi Xu, Weichu Xie +8

Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective…

cs.LG2026

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning

Ziyue Wang, Aomufei Yuan, Yongfu Zhu +10

Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce eac…

cs.LG2026

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

Yuqi Xu, Rizhen Hu, Zihan Liu +2

Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searchi…

cs.LG2026

Step-wise Rubric Rewards for LLM Reasoning

Weichu Xie, Haozhe Zhao, Wenpu Liu +15

Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision ov…