148 citations · 166 across the 4 of their papers we have counts for
7 papers
Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
Sijie Song, Xudong Lin, Jiaying Liu +2
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods wh…
MSL: Multi-Task Self-Supervised Learning for Skeleton Based Action Recognition
Lilang Lin, Sijie Song, Wenhan Yan +1
In this paper, we address self-supervised representation learning from human skeletons for action recognition. Previous methods, which usually learn feature presentations from a si…
Fashion Meets Computer Vision: A Survey
Wen-Huang Cheng, Sijie Song, Chieh-Yun Chen +2
Fashion is the way we present ourselves to the world and has become one of the world's largest industries. Fashion, mainly conveyed by vision, has thus attracted much attention fro…
Modality Compensation Network: Cross-Modal Adaptation for Action Recognition
Sijie Song, Jiaying Liu, Yanghao Li +1
With the prevalence of RGB-D cameras, multi-modal video data have become more available for human action recognition. One main challenge for this task lies in how to effectively le…
Unsupervised Person Image Generation with Semantic Parsing Transformation
Sijie Song, Wei Zhang, Jiaying Liu +1
In this paper, we address unsupervised pose-guided person image generation, which is known challenging due to non-rigid deformation. Unlike previous methods learning a rock-hard di…
Temporal Bilinear Networks for Video Action Recognition
Yanghao Li, Sijie Song, Yuqi Li +1
Temporal modeling in videos is a fundamental yet challenging problem in computer vision. In this paper, we propose a novel Temporal Bilinear (TB) model to capture the temporal pair…