1 paper · 1 filter
Jinxing Zhou, Dan Guo, Yuxin Mao +3
Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identific…