10 papers
Improving Vision-language Models with Perception-centric Process Reward Models
Yingqian Min, Kun Zhou, Yifan Li +6
Recent advancements in reinforcement learning with verifiable rewards (RLVR) have significantly improved the complex reasoning ability of vision-language models (VLMs). However, it…
Towards Long-horizon Agentic Multimodal Search
Yifan Du, Zikang Liu, Jinbiao Peng +5
Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, managing the heterogeneous informa…
A Survey of Large Language Models
Wayne Xin Zhao, Kun Zhou, Junyi Li +19
Language is essentially a complex, intricate system of human expressions governed by grammatical rules. It poses a significant challenge to develop capable AI algorithms for compre…
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
Yifan Li, Yukai Gu, Yingqian Min +6
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation…
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
Yifan Du, Kun Zhou, Yingqian Min +3
We study how different Chain-of-Thought (CoT) designs affect the acquisition of the generalizable visual reasoning ability in vision-language models (VLMs). While CoT data, especia…
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
Jiyang Tang, Hengyi Li, Yifan Du +1
Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of v…