1 citations · 2 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2021★ 1 cited
Rethinking the constraints of multimodal fusion: case study in Weakly-Supervised Audio-Visual Video Parsing
Jianning Wu, Zhuqing Jiang, Shiping Wen +2
For multimodal tasks, a good feature extraction network should extract information as much as possible and ensure that the extracted feature embedding and other modal feature embed…
cs.CV2021
Taylor saves for later: disentanglement for video prediction using Taylor representation
Ting Pan, Zhuqing Jiang, Jianan Han +3
Video prediction is a challenging task with wide application prospects in meteorology and robot systems. Existing works fail to trade off short-term and long-term prediction perfor…
cs.CV2020
Crowd Counting via Hierarchical Scale Recalibration Network
Zhikang Zou, Yifan Liu, Shuangjie Xu +3
The task of crowd counting is extremely challenging due to complicated difficulties, especially the huge variation in vision scale. Previous works tend to adopt a naive concatenati…