440 citations · 942 across the 16 of their papers we have counts for
4 papers · 1 filter
Dense Contrastive Visual-Linguistic Pretraining
Lei Shi, Kai Shuang, Shijie Geng +5
Inspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior p…
Contrastive Visual-Linguistic Pretraining
Lei Shi, Kai Shuang, Shijie Geng +6
Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently. Such approaches can achieve superior performance due to the high-l…
Character Matters: Video Story Understanding with Character-Aware Relations
Shijie Geng, Ji Zhang, Zuohui Fu +3
Different from short videos and GIFs, video stories contain clear plots and lists of principal characters. Without identifying the connection between appearing people and character…
OOGAN: Disentangling GAN with One-Hot Sampling and Orthogonal Regularization
Bingchen Liu, Yizhe Zhu, Zuohui Fu +2
Exploring the potential of GANs for unsupervised disentanglement learning, this paper proposes a novel GAN-based disentanglement framework with One-Hot Sampling and Orthogonal Regu…