1 paper
Pengcheng Zhao, Jinxing Zhou, Yang Zhao +2
The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics…