2 papers
cs.GR2025
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
Yihuan Huang, Jiajun Liu, Yanzhen Ren +3
Recent talking head synthesis works typically adopt speech features extracted from large-scale pre-trained acoustic models. However, the intrinsic many-to-many relationship between…
cs.MM2025
Audio-visual Event Localization on Portrait Mode Short Videos
Wuyang Liu, Yi Chai, Yongpeng Yan +1
Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long vide…