55 citations · 151 across the 15 of their papers we have counts for
12 papers
Many-to-many Image Generation with Auto-regressive Diffusion Models
Ying Shen, Yizhe Zhang, Shuangfei Zhai +3
Recent advancements in image generation have made significant progress, yet existing models present limitations in perceiving and generating an arbitrary number of interrelated ima…
Scalable Pre-training of Large Autoregressive Image Models
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai +5
This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language…
BOOT: Data-free Distillation of Denoising Diffusion Models with Bootstrapping
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2
Diffusion models have demonstrated excellent potential for generating diverse images. However, their performance often suffers from slow generation due to iterative denoising. Know…
TRACT: Denoising Diffusion Models with Transitive Closure Time-Distillation
David Berthelot, Arnaud Autef, Jierui Lin +6
Denoising Diffusion models have demonstrated their proficiency for generative sampling. However, generating good samples often requires many iterations. Consequently, techniques su…
GAUDI: A Neural Architect for Immersive 3D Scene Generation
Miguel Angel Bautista, Pengsheng Guo, Samira Abnar +9
We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle thi…
Position Prediction as an Effective Pretraining Strategy
Shuangfei Zhai, Navdeep Jaitly, Jason Ramapuram +7
Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of thei…