Showing cs.MMShow all
3 papers · 1 filter
cs.MM2025
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
Yaru Chen, Peiliang Zhang, Fei Li +4
Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on…
cs.MM2023
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
Jinzheng Zhao, Yong Xu, Xinyuan Qian +6
Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide compl…
cs.MM2023
Audio Visual Speaker Localization from EgoCentric Views
Jinzheng Zhao, Yong Xu, Xinyuan Qian +1
The use of audio and visual modality for speaker localization has been well studied in the literature by exploiting their complementary characteristics. However, most previous work…