11 papers
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Caoyuan Ma, Wenpu Liu, Weichu Xie +12
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propos…
OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning
Zhentong Ye, Lei Zhang, Sijia Zhou +7
Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We i…
DOPD: Dual On-policy Distillation
Xinlei Yu, Gen Li, Qingyi Si +13
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Wenpu Liu, Yuqi Xu, Weichu Xie +8
Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective…
ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation
Kanghui Tian, Siyuan Liu, Ziang Yan +3
On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a commo…
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning
Ziyue Wang, Aomufei Yuan, Yongfu Zhu +10
Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce eac…