3 papers
cs.CV2026
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
Qingfeng Li, Haoxian Zhang, Xu He +4
Visual tokenizers play a central role in latent image generation by bridging high-dimensional images and tractable generative modeling. However, most existing tokenizers are still…
cs.CV2025
Exploring Representation Invariance in Finetuning
Wenqiang Zu, Shenghao Xie, Hao Chen +9
Foundation models pretrained on large-scale natural images are widely adapted to various cross-domain low-resource downstream tasks, benefiting from generalizable and transferable…
cs.CV2024
Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
Shenghao Xie, Wenqiang Zu, Mingyang Zhao +6
Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing…