activity
20212023
most citedDeepViT: Towards Deeper Vision Transformer

349 citations · 675 across the 30 of their papers we have counts for

collaborators

30 papers

cs.CV20232 cited

COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

Sihan Chen, Xingjian He, Handong Li +3

Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…

cs.CV20224 cited

Diffusion Probabilistic Model Made Slim

Xingyi Yang, Daquan Zhou, Jiashi Feng +1

Despite the recent visually-pleasing results achieved, the massive computational cost has been a long-standing flaw for diffusion probabilistic models (DPMs), which, in turn, great…

cs.CV2022

AvatarGen: A 3D Generative Model for Animatable Human Avatars

Jianfeng Zhang, Zihang Jiang, Dingdong Yang +6

Unsupervised generation of 3D-aware clothed humans with various appearances and controllable geometries is important for creating virtual human avatars and other AR/VR applications…

cs.CV202273 cited

Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

Qibin Hou, Cheng-Ze Lu, Ming-Ming Cheng +1

This paper does not attempt to design a state-of-the-art method for visual recognition but investigates a more efficient way to make use of convolutions to encode spatial features.…

cs.CV202214 cited

MagicMix: Semantic Mixing with Diffusion Models

Jun Hao Liew, Hanshu Yan, Daquan Zhou +1

Have you ever imagined what a corgi-alike coffee machine or a tiger-alike rabbit would look like? In this work, we attempt to answer these questions by exploring a new task called…

cs.LG2022

Reachability-Aware Laplacian Representation in Reinforcement Learning

Kaixin Wang, Kuangqi Zhou, Jiashi Feng +2

In Reinforcement Learning (RL), Laplacian Representation (LapRep) is a task-agnostic state representation that encodes the geometry of the environment. A desirable property of LapR…