1 paper
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo +3
Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in th…