2 papers
cs.CV2026
SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models
Caoyuan Ma, Tian Gu, Wenpu Liu +11
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both…
cs.LG2026
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning
Ziyue Wang, Aomufei Yuan, Yongfu Zhu +10
Reinforcement Learning from Verifiable Rewards (RLVR) has become the dominant approach for improving mathematical reasoning in large language models, yet current methods reduce eac…