119 citations · 229 across the 20 of their papers we have counts for
6 papers · 1 filter
MAGVIT: Masked Generative Video Transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn +8
We introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spat…
MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis
Tianhong Li, Huiwen Chang, Shlok Kumar Mishra +3
Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for eac…
A simple, efficient and scalable contrastive masked autoencoder for learning visual representations
Shlok Mishra, Joshua Robinson, Huiwen Chang +4
We introduce CAN, a simple, efficient and scalable method for self-supervised learning of visual representations. Our framework is a minimal and conceptually clean synthesis of (C)…
Visual Prompt Tuning for Generative Transfer Learning
Kihyuk Sohn, Yuan Hao, José Lezama +5
Transferring knowledge from an image synthesis model trained on a large dataset is a promising direction for learning generative image models from various domains efficiently. Whil…
Improved Masked Image Generation with Token-Critic
José Lezama, Huiwen Chang, Lu Jiang +1
Non-autoregressive generative transformers recently demonstrated impressive image generation performance, and orders of magnitude faster sampling than their autoregressive counterp…
MaskGIT: Masked Generative Image Transformer
Huiwen Chang, Han Zhang, Lu Jiang +2
Generative transformers have experienced rapid popularity growth in the computer vision community in synthesizing high-fidelity and high-resolution images. The best generative tran…