1 citations · 1 across the 2 of their papers we have counts for
4 papers
Rethinking the constraints of multimodal fusion: case study in Weakly-Supervised Audio-Visual Video Parsing
Jianning Wu, Zhuqing Jiang, Shiping Wen +2
For multimodal tasks, a good feature extraction network should extract information as much as possible and ensure that the extracted feature embedding and other modal feature embed…
Taylor saves for later: disentanglement for video prediction using Taylor representation
Ting Pan, Zhuqing Jiang, Jianan Han +3
Video prediction is a challenging task with wide application prospects in meteorology and robot systems. Existing works fail to trade off short-term and long-term prediction perfor…
MöbiusE: Knowledge Graph Embedding on Möbius Ring
Yao Chen, Jiangang Liu, Zhe Zhang +2
In this work, we propose a novel Knowledge Graph Embedding (KGE) strategy, called MöbiusE, in which the entities and relations are embedded to the surface of a Möbius ring. The pro…
Crowd Counting via Hierarchical Scale Recalibration Network
Zhikang Zou, Yifan Liu, Shuangjie Xu +3
The task of crowd counting is extremely challenging due to complicated difficulties, especially the huge variation in vision scale. Previous works tend to adopt a naive concatenati…