2 papers
cs.CV2026
Video-KTR: Reinforcing Video Reasoning via Key Token Attribution
Ziyue Wang, Sheng Jin, Zhongrong Zuo +5
Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequen…
cs.CV2025
FIRM: Flexible Interactive Reflection reMoval
Xiao Chen, Xudong Jiang, Yunkang Tao +4
Removing reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfacto…