Showing cs.MMShow all
3 papers · 1 filter
cs.MM2026
FATE: Frame-Level Audio-Visual Temporal Embedding
Kaisi Guan, Bingzi Zhang, Xihua Wang +4
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations…
cs.MM2025
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
Yuyue Wang, Xin Cheng, Yihan Wu +3
The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being co…
cs.MM2025
EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
Xiaochuan Liu, Xin Cheng, Yuchong Sun +4
Imitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building aliv…