349 citations · 675 across the 30 of their papers we have counts for
30 papers
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
Sihan Chen, Xingjian He, Handong Li +3
Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…
Diffusion Probabilistic Model Made Slim
Xingyi Yang, Daquan Zhou, Jiashi Feng +1
Despite the recent visually-pleasing results achieved, the massive computational cost has been a long-standing flaw for diffusion probabilistic models (DPMs), which, in turn, great…
AvatarGen: A 3D Generative Model for Animatable Human Avatars
Jianfeng Zhang, Zihang Jiang, Dingdong Yang +6
Unsupervised generation of 3D-aware clothed humans with various appearances and controllable geometries is important for creating virtual human avatars and other AR/VR applications…
Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition
Qibin Hou, Cheng-Ze Lu, Ming-Ming Cheng +1
This paper does not attempt to design a state-of-the-art method for visual recognition but investigates a more efficient way to make use of convolutions to encode spatial features.…
MagicMix: Semantic Mixing with Diffusion Models
Jun Hao Liew, Hanshu Yan, Daquan Zhou +1
Have you ever imagined what a corgi-alike coffee machine or a tiger-alike rabbit would look like? In this work, we attempt to answer these questions by exploring a new task called…
Reachability-Aware Laplacian Representation in Reinforcement Learning
Kaixin Wang, Kuangqi Zhou, Jiashi Feng +2
In Reinforcement Learning (RL), Laplacian Representation (LapRep) is a task-agnostic state representation that encodes the geometry of the environment. A desirable property of LapR…