11 citations · 22 across the 5 of their papers we have counts for
7 papers · 1 filter
Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete Tokens
Yuxiao Chen, Jianbo Yuan, Yu Tian +5
Contrastive learning-based vision-language pre-training approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal a…
HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention
Shijie Geng, Jianbo Yuan, Yu Tian +2
The success of large-scale contrastive vision-language pretraining (CLIP) has benefited both visual recognition and multimodal content understanding. The concise design brings CLIP…
AE-StyleGAN: Improved Training of Style-Based Auto-Encoders
Ligong Han, Sri Harsha Musunuri, Martin Renqiang Min +3
StyleGANs have shown impressive results on data generation and manipulation in recent years, thanks to its disentangled style latent space. A lot of efforts have been made in inver…
A Good Image Generator Is What You Need for High-Resolution Video Synthesis
Yu Tian, Jian Ren, Menglei Chai +4
Image and video synthesis are closely related areas aiming at generating content from noise. While rapid progress has been demonstrated in improving image-based models to handle la…
Semantic Graph Convolutional Networks for 3D Human Pose Regression
Long Zhao, Xi Peng, Yu Tian +2
In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convol…
Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao, Xi Peng, Yu Tian +2
We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…