activity
20162024
most citedVariational Autoencoder for Deep Learning of Images, Labels and Captions

371 citations · 514 across the 13 of their papers we have counts for

collaborators

6 papers

cs.CV202225 cited

NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis

Chenfei Wu, Jian Liang, Xiaowei Hu +6

In this paper, we present NUWA-Infinity, a generative model for infinite visual synthesis, which is defined as the task of generating arbitrarily-sized high-resolution images or lo…

cs.CV20215 cited

MLP Architectures for Vision-and-Language Modeling: An Empirical Study

Yixin Nie, Linjie Li, Zhe Gan +6

We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks,…

cs.CV20218 cited

Injecting Semantic Concepts into End-to-End Image Captioning

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +5

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Re…

stat.ML2016

Unsupervised Learning with Truncated Gaussian Graphical Models

Qinliang Su, Xuejun Liao, Chunyuan Li +2

Gaussian graphical models (GGMs) are widely used for statistical modeling, because of ease of inference and the ubiquitous use of the normal distribution in practical approximation…

cs.CV201618 cited

Semantic Compositional Networks for Visual Captioning

Zhe Gan, Chuang Gan, Xiaodong He +5

A Semantic Compositional Network (SCN) is developed for image captioning, in which semantic concepts (i.e., tags) are detected from the image, and the probability of each tag is us…

stat.ML2016371 cited

Variational Autoencoder for Deep Learning of Images, Labels and Captions

Yunchen Pu, Zhe Gan, Ricardo Henao +4

A novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of t…