1 paper
Langyu Wang, Bingke Zhu, Yingying Chen +3
The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the…