3 papers
cs.CV2026
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
Chenhao Qiu, Yechao Zhang, Xin Luo +2
Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long vi…
cs.CV2026
UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging Scenarios
Ruidi Fan, Yang Zhou, Siyuan Wang +3
Lip synchronization aims to generate realistic talking videos that match given audio, which is essential for high-quality video dubbing. However, current methods have fundamental d…
cs.CV2025
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…