37 citations · 46 across the 2 of their papers we have counts for
3 papers
cs.CV2021★ 37 cited
Unifying Multimodal Transformer for Bi-directional Image and Text Generation
Yupan Huang, Hongwei Xue, Bei Liu +1
We study the joint learning of image-to-text and text-to-image generations, which are naturally bi-directional tasks. Typical existing works design two separate task-specific model…
cs.CV2021
Learning Fine-Grained Motion Embedding for Landscape Animation
Hongwei Xue, Bei Liu, Huan Yang +3
In this paper we focus on landscape animation, which aims to generate time-lapse videos from a single landscape image. Motion is crucial for landscape animation as it determines ho…
cs.CV2021★ 9 cited
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
Hongwei Xue, Yupan Huang, Bei Liu +4
Vision-Language Pre-training (VLP) aims to learn multi-modal representations from image-text pairs and serves for downstream vision-language tasks in a fine-tuning fashion. The dom…