40 citations · 278 across the 23 of their papers we have counts for
18 papers · 1 filter
Fine-grained Iterative Attention Network for TemporalLanguage Localization in Videos
Xiaoye Qu, Pengwei Tang, Zhikang Zhou +3
Temporal language localization in videos aims to ground one video segment in an untrimmed video based on a given sentence query. To tackle this task, designing an effective model t…
Large-Scale Adversarial Training for Vision-and-Language Representation Learning
Zhe Gan, Yen-Chun Chen, Linjie Li +3
We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-…
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng +3
Recent Transformer-based large-scale pre-trained models have revolutionized vision-and-language (V+L) research. Models such as ViLBERT, LXMERT and UNITER have significantly lifted…
3D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training
Yu Cheng, Bo Yang, Bo Wang +1
Estimating 3D poses from a monocular video is still a challenging task, despite the significant progress that has been made in recent years. Generally, the performance of existing…
Adversarial Robustness: From Self-Supervised Pre-Training to Fine-Tuning
Tianlong Chen, Sijia Liu, Shiyu Chang +3
Pretrained models from self-supervision are prevalently used in fine-tuning downstream tasks faster or for better accuracy. However, gaining robustness from pretraining is left une…
BachGAN: High-Resolution Image Synthesis from Salient Object Layout
Yandong Li, Yu Cheng, Zhe Gan +3
We propose a new task towards more practical application for image generation - high-quality image synthesis from salient object layout. This new setting allows users to provide th…