9 papers
Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
Lei Yang, Xinze Liu, Dayan Wu +7
The paper introduces a method to detect and correct hallucinated objects in large vision‑language models by identifying a hidden grounding pattern and using a lightweight verifier…
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR
Chuanyu Qin, Chenxu Yang, Qingyi Si +3
Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…
Online Self-Calibration Against Hallucination in Vision-Language Models
Minghui Chen, Chenxu Yang, Hengjie Zhu +3
Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment…
Near-Future Policy Optimization
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…
EasyVideoR1: Easier RL for Video Understanding
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve i…
Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models
Xiyu Liu, Qingyi Si, Zhengxiao Liu +3
While locate-then-edit knowledge editing efficiently updates knowledge encoded within Large Language Models (LLMs), a critical generalization failure mode emerges in the practical…