3 papers
cs.CV2025
Spectral Image Tokenizer
Carlos Esteves, Mohammed Suhail, Ameesh Makadia
Image tokenizers map images to sequences of discrete tokens, and are a crucial component of autoregressive transformer-based image generation. The tokens are typically associated w…
cs.CV2025
Factorized Video Autoencoders for Efficient Generative Modelling
Mohammed Suhail, Carlos Esteves, Leonid Sigal +1
Latent variable generative models have emerged as powerful tools for generative tasks including image and video synthesis. These models are enabled by pretrained autoencoders that…
cs.CV2025
Direct Motion Models for Assessing Generated Videos
Kelsey Allen, Carl Doersch, Guangyao Zhou +9
A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular…