3 papers
cs.CL2025
ChronusOmni: Improving Time Awareness of Omni Large Language Models
Yijing Chen, Yihan Wu, Kaisi Guan +4
Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target v…
cs.MM2025
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
Yuyue Wang, Xin Cheng, Yihan Wu +3
The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being co…
eess.AS2024
Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization
Yihan Wu, Yichen Lu, Yifan Peng +3
Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-wor…