8 papers
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
Philippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan +9
Vision Transformer (ViT) autoencoders have emerged as compelling tokenizers for images, offering improved reconstruction over convolutional tokenizers. However, existing ViT tokeni…
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
Tanush Yadav, Mohammadreza Salehi, Jae Sung Park +6
Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessential task for video understa…
Posterior Augmented Flow Matching
George Stoica, Sayak Paul, Matthew Wallingford +6
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each train…
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
Xiang Fan, Sharath Girish, Vivek Ramanujan +6
Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, i…
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
Vivek Ramanujan, Kushal Tirumala, Armen Aghajanyan +2
Current image generation methods are based on a two-stage training approach. In stage 1, an auto-encoder is trained to compress an image into a latent space; in stage 2, a generati…
Contrastive Flow Matching
George Stoica, Vivek Ramanujan, Xiang Fan +3
Unconditional flow-matching trains diffusion models to transport samples from a source distribution to a target distribution by enforcing that the flows between sample pairs are un…