activity
20192023
most citedUnifying Multimodal Transformer for Bi-directional Image and Text Generation

37 citations · 79 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV20232 cited

TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

Jingye Chen, Yupan Huang, Tengchao Lv +3

The diffusion model has been proven a powerful generative model in recent years, yet remains a challenge in generating visual text. Several methods alleviated this issue by incorpo…

cs.CV2021

A Picture is Worth a Thousand Words: A Unified System for Diverse Captions and Rich Images Generation

Yupan Huang, Bei Liu, Jianlong Fu +1

A creative image-and-text generative AI system mimics humans' extraordinary abilities to provide users with diverse and comprehensive caption suggestions, as well as rich image cre…

cs.CV202137 cited

Unifying Multimodal Transformer for Bi-directional Image and Text Generation

Yupan Huang, Hongwei Xue, Bei Liu +1

We study the joint learning of image-to-text and text-to-image generations, which are naturally bi-directional tasks. Typical existing works design two separate task-specific model…

cs.CV20219 cited

Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training

Hongwei Xue, Yupan Huang, Bei Liu +4

Vision-Language Pre-training (VLP) aims to learn multi-modal representations from image-text pairs and serves for downstream vision-language tasks in a fine-tuning fashion. The dom…

cs.CV202125 cited

Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning

Zhicheng Huang, Zhaoyang Zeng, Yupan Huang +3

We study joint learning of Convolutional Neural Network (CNN) and Transformer for vision-language pre-training (VLPT) which aims to learn cross-modal alignments from millions of im…

cs.MM20202 cited

Reinforcing Short-Length Hashing

Xingbo Liu, Xiushan Nie, Qi Dai +2

Due to the compelling efficiency in retrieval and storage, similarity-preserving hashing has been widely applied to approximate nearest neighbor search in large-scale image retriev…