61 citations · 99 across the 26 of their papers we have counts for
Showing cs.MMShow all
2 papers · 1 filter
cs.MM2023
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li, Ying Shan +1
Speaker diarization in real-world videos presents significant challenges due to varying acoustic conditions, diverse scenes, the presence of off-screen speakers, etc. This paper bu…
cs.MM2023★ 1 cited
Unified Pretraining Target Based Video-music Retrieval With Music Rhythm And Video Optical Flow Information
Tianjun Mao, Shansong Liu, Yunxuan Zhang +2
Background music (BGM) can enhance the video's emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrie…