1 paper
Xiang He, Xiangxi Liu, Yang Li +5
The audio-visual event localization task requires identifying concurrent visual and auditory events from unconstrained videos within a network model, locating them, and classifying…