6 papers
Hita: Holistic Tokenizer for Autoregressive Image Generation
Anlin Zheng, Haochen Wang, Yucheng Zhao +4
Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, becaus…
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
Anlin Zheng, Xin Wen, Xuanyang Zhang +5
In this work, we present a novel direction to build an image tokenizer directly on top of a frozen vision foundation model, which is a largely underexplored area. Specifically, we…
Reconstructive Visual Instruction Tuning
Haochen Wang, Anlin Zheng, Yucheng Zhao +4
This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to co…
Progressive End-to-End Object Detection in Crowded Scenes
Anlin Zheng, Yuang Zhang, Xiangyu Zhang +2
In this paper, we propose a new query-based detection framework for crowd detection. Previous query-based detectors suffer from two drawbacks: first, multiple predictions will be i…
Detection in Crowded Scenes: One Proposal, Multiple Predictions
Xuangeng Chu, Anlin Zheng, Xiangyu Zhang +1
We propose a simple yet effective proposal-based object detector, aiming at detecting highly-overlapped instances in crowded scenes. The key of our approach is to let each proposal…
Complementary Segmentation of Primary Video Objects with Reversible Flows
Jia Li, Junjie Wu, Anlin Zheng +3
Segmenting primary objects in a video is an important yet challenging problem in computer vision, as it exhibits various levels of foreground/background ambiguities. To reduce such…