4 citations · 5 across the 5 of their papers we have counts for
4 papers · 1 filter
Latent Denoising Makes Good Tokenizers
Jiawei Yang, Tianhong Li, Lijie Fan +2
Despite their fundamental role, it remains unclear what properties could make tokenizers more effective for generative modeling. We observe that modern generative models share a co…
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Lijie Fan, Tianhong Li, Siyang Qin +6
Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-ima…
Autoregressive Image Generation without Vector Quantization
Tianhong Li, Yonglong Tian, He Li +2
Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can fa…
Denoising Vision Transformers
Jiawei Yang, Katie Z Luo, Jiefeng Li +6
We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts, which hurt the performance of ViTs i…