3 papers
cs.CV2026
Multiple Hypothesis Flow Estimation for Video Frame Interpolation under Matching Ambiguity
Zibo Su, Jing Kong, Ruixing Wang +2
Many flow-based video frame interpolation (VFI) methods synthesize an intermediate frame by estimating optical flow fields, warping the two input frames, and blending the warped ob…
cs.CV2026
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
Jiahua Li, Zhanhe Zhang, Chenghao Xu +4
Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model…
cs.CV2025
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
Zibo Su, Kun Wei, Jiahua Li +3
Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-Engl…