6 papers
Representation Fréchet Loss for Visual Generation
Jiawei Yang, Zhengyang Geng, Xuan Ju +2
We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: deco…
Latent Denoising Makes Good Tokenizers
Jiawei Yang, Tianhong Li, Lijie Fan +2
Despite their fundamental role, it remains unclear what properties could make tokenizers more effective for generative modeling. We observe that modern generative models share a co…
Autoregressive Image Generation without Vector Quantization
Tianhong Li, Yonglong Tian, He Li +2
Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can fa…
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Lijie Fan, Tianhong Li, Siyang Qin +6
Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-ima…
Denoising Vision Transformers
Jiawei Yang, Katie Z Luo, Jiefeng Li +6
We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts, which hurt the performance of ViTs i…
Self-Correcting Self-Consuming Loops for Generative Model Training
Nate Gillman, Michael Freeman, Daksh Aggarwal +4
As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite t…