3 papers
cs.CV2025
D-AR: Diffusion via Autoregressive Models
Ziteng Gao, Mike Zheng Shou
This paper presents Diffusion via Autoregressive models (D-AR), a new paradigm recasting the image diffusion process as a vanilla autoregressive procedure in the standard next-toke…
cs.CV2024
Factorized Visual Tokenization and Generation
Zechen Bai, Jianxiong Gao, Ziteng Gao +4
Visual tokenizers are fundamental to image generation. They convert visual data into discrete tokens, enabling transformer-based models to excel at image generation. Despite their…
cs.CV2024
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
Zechen Bai, Tong He, Haiyang Mei +6
We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoni…